Computing technologies for language models

The RAG framework enhances language models by integrating specialized data sources to generate accurate and contextually relevant responses, addressing hallucinations and expanding their capabilities to output forms, opinions, and recommendations.

WO2025221579A1PCT designated stage Publication Date: 2025-10-23RETEQ INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/024161
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-15
Filing Date
2025-04-10
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Conventional language models, such as chatbots powered by LLMs, often hallucinate and fail to output responses containing forms, opinions based on personal communications or private comments, or recommendations, limiting their functionality.

Method used

A system that enhances language models with a Retrieval Augmented Generation (RAG) framework, utilizing specialized data sources like forms, personal communications, and private comments to generate accurate and contextually relevant responses, including populated forms, opinions, and recommendations.

Benefits of technology

Enables language models to output forms, opinions based on personal communications, and recommendations, improving their functionality and accuracy by grounding responses in verifiable facts, reducing hallucinations, and enhancing computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025024161_23102025_PF_FP_ABST
    Figure US2025024161_23102025_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure enables various computing technologies that enable a chatbot programmed to output a response containing (i) a form (e.g., a data file, a PDF file, an HTML file) that has been at least partially populated, (ii) an opinion as to whether a legislative text (structured or unstructured) is appropriate or applicable to a set of parameters, (iii) an opinion based on an index of personal communications, (iv) an opinion based on an index of private comments for a set of parameters, or (v) a recommendation for training. Such configuration improves computer functionality and is technologically beneficial, because of its enablement of the language model to output forms, opinions, or recommendations, which the language model is not known to do.
Need to check novelty before this filing date? Find Prior Art

Description

TITLECOMPUTING TECHNOLOGIES FOR LANGUAGE MODELSCROSS-REFERENCE TO RELATED PATENT APPLICATION

[0001] This patent application claims a benefit of priority to US Provisional Patent Application 63 / 634,213 filed 15 April 2024, which is incorporated by reference herein for all purposes.TECHNICAL FIELD

[0002] This disclosure relates to language models.BACKGROUND

[0003] Conventionally, a chatbot powered by a language model (e.g., a Large Language Model (LLM), a Small Language Model (SLM)) may interact with a computing terminal (e.g., a desktop computer, a laptop computer, a smartphone) operated by a user. One example of such chatbot is branded as OpenAI ChatGPT.

[0004] Although this technology is sometimes useful, this technology often hallucinates by generating seemingly realistic, but factually incorrect or nonsensical information. As such, the chatbot may augmented by employing a Retrieval Augmented Generation (RAG) framework to retrieve data from an index for inputting into the language model to provide more accurate and factual responses, which minimizes at least some hallucinations.

[0005] Although the RAG framework is sometimes useful, this approach is still technologically problematic, because this approach is not known to be able to output a response containing at least one of (i) a form (e.g., a data file, a Portable Document Format (PDF) file, a Hypertext Markup Language (HTML) file) that has been at least partially populated, (ii) an opinion as to whether a text, whether structured or unstructured, is appropriate or applicable to a set of parameters, (iii) an opinion based on an index of personal communications (e.g., chat messages, email messages), (iv) an opinion based on an index of private comments for a set of parameters, or (v) a recommendation for training.SUMMARY

[0006] This disclosure enables a language model to output a response containing at least one of (i) a form (e.g., a data file, a PDF file, a HTML file) that has been at least partially populated, (ii) an opinion as to whether a text, whether structured or unstructured, is appropriate or applicable to a set of parameters, (iii) an opinion based on an index of personal communications (e.g., chat messages, email messages), (iv) an opinion based on an index of private comments for a set of parameters, or (v) a recommendation for training. Such configuration improves computer functionality and is technologically beneficial, because of its enablement of the language model to output forms, opinions, or recommendations, which the language model is not known to do, especially when specialized data (e.g., forms, personal communications, private comments) is used for augmenting user queries. For example, there may be a system, comprising: a computing instance (e.g., a physical server) programmed to: receive a user query from a computing terminal (e.g., a desktop computer, a laptop computer, a smartphone), generate a prompt based on the user query, submit the prompt to a language model (e.g., an LLM, an SLM), receive a response to the prompt from the language model, and cause the response to be presented (e.g., visually) on the computing terminal responsive to the user query.DESCRIPTION OF DRAWINGS

[0007] FIG. 1 shows a diagram of an embodiment of a topology according to this disclosure.

[0008] FIG. 2 shows a diagram of an embodiment of a computing instance according to this disclosure.

[0009] FIG. 3 shows a flowchart of an embodiment of an algorithm to form an index based on a set of forms according to this disclosure.

[0010] FIG. 4 shows a flowchart of an embodiment of an algorithm to use an index to output a form that has been at least partially populated according to this disclosure.

[0011] FIG. 5 shows a flowchart of an embodiment of an algorithm to form an index based on a legislative text according to this disclosure.

[0012] FIG. 6 shows a flowchart of an embodiment of an algorithm to use an index to output an opinion as to whether a legislative (or another type) text is appropriate or applicable to a set of parameters according to this disclosure.

[0013] FIG. 7 shows a flowchart of an embodiment of an algorithm to form an index based on a set of personal communications according to this disclosure.

[0014] FIG. 8 shows a flowchart of an embodiment of an algorithm to use an index to output an opinion based on an index of personal communications according to this disclosure.

[0015] FIG. 9 shows a flowchart of an embodiment of an algorithm to form an index based on a set of private comments according to this disclosure.

[0016] FIG. 10 shows a flowchart of an embodiment of an algorithm to use an index to output an opinion based on an index of private comments according to this disclosure.

[0017] FIG. 11 shows a flowchart of an embodiment of an algorithm to generate a recommendation for training according to this disclosure.

[0018] FIG. 12 shows a screenshot of an embodiment of a Graphical User Interface (GUI) of an application program hosted on a computing terminal of FIG. 1 according to this disclosure.

[0019] FIG. 13 shows a screenshot of an embodiment of a GUI of an application program hosted on a computing terminal of FIG. 1 according to this disclosure.

[0020] FIG. 14 shows a screenshot of an embodiment of a GUI of an application program hosted on a computing terminal of FIG. 1 according to this disclosure.

[0021] FIG. 15 shows a screenshot of an embodiment of a GUI of an application program hosted on a computing terminal of FIG. 1 according to this disclosure.DETAILED DESCRIPTION

[0022] As explained above, this disclosure enables a language model to output a response containing at least one of (i) a form (e.g., a data file, a PDF file, a HTML file) that has been at least partially populated, (ii) an opinion as to whether a text, whether structured or unstructured, is appropriate or applicable to a set of parameters, (iii) an opinion based on an index of personal communications (e.g., chat messages, email messages), (iv) an opinion based on an index of private comments for a set of parameters,or (v) a recommendation for training. Such configuration improves computer functionality and is technologically beneficial, because of its enablement of the language model to output forms, opinions, or recommendations, which the language model is not known to do, especially when specialized data (e.g., forms, personal communications, private comments) is used for augmenting user queries. This disclosure is now described more fully with reference to various figures that are referenced above, in which some embodiments of this disclosure are shown. For example, there may be a system, comprising: a computing instance (e.g., a physical server) programmed to: receive a user query from a computing terminal (e.g., a desktop computer, a laptop computer, a smartphone), generate a prompt based on the user query, submit the prompt to a language model (e.g., an LLM, an SLM), receive a response to the prompt from the language model, and cause the response to be presented (e.g., visually) on the computing terminal responsive to the user query.

[0023] This disclosure may, however, be embodied in many different forms and should not be construed as necessarily being limited to only embodiments disclosed herein. Rather, these embodiments are provided so that this disclosure is thorough and complete, and fully conveys various concepts of this disclosure to skilled artisans.

[0024] Various terminology used herein can imply direct or indirect, full or partial, temporary or permanent, action or inaction. For example, when an element is referred to as being "on," "connected" or "coupled" to another element, then the element can be directly on, connected or coupled to the other element or intervening elements can be present, including indirect or direct variants. In contrast, when an element is referred to as being "directly connected" or "directly coupled" to another element, there are no intervening elements present.

[0025] Likewise, as used herein, a term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless specified otherwise, or clear from context, "X employs A or B" is intended to mean any of the natural inclusive permutations. That is, if X employs A; X employs B; or X employs both A and B, then "X employs A or B" is satisfied under any of the foregoing instances.

[0026] Similarly, as used herein, various singularforms "a," "an" and "the" are intended to include various plural forms (e.g., two, three, four) as well, unless context clearlyindicates otherwise. For example, a term "a" or "an" shall mean "one or more," even though a phrase "one or more" is also used herein.

[0027] Moreover, terms "comprises," "includes" or "comprising," "including" when used in this specification, specify a presence of stated features, integers, steps, operations, elements, or components, but do not preclude a presence and / or addition of one or more other features, integers, steps, operations, elements, components, or groups thereof. Furthermore, when this disclosure states that something is "based on" something else, then such statement refers to a basis which may be based on one or more other things as well. In other words, unless expressly indicated otherwise, as used herein "based on" inclusively means "based at least in part on" or "based at least partially on."

[0028] Additionally, although terms first, second, and others can be used herein to describe various elements, components, regions, layers, subsets, diagrams, or sections, these elements, components, regions, layers, subsets, diagrams, or sections should not necessarily be limited by such terms. Rather, these terms are used to distinguish one element, component, region, layer, subset, diagram, or section from another element, component, region, layer, subset, diagram, or section. As such, a first element, component, region, layer, subset, diagram, or section discussed below could be termed a second element, component, region, layer, subset, diagram, or section without departing from this disclosure.

[0029] Also, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in an art to which this disclosure belongs. As such, terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in a context of a relevant art and should not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0030] FIG. 1 shows a diagram of an embodiment of a topology according to this disclosure. In particular, there is a topology 100, which includes a network 102, a computing instance 104, a computing terminal 106, a language model 108, a data source 110 (first), a data source 112 (second), a data source 114 (third), and a data source 116 (fourth). Note that the data source 110, the data source 112, the data source 114, or thedata source 116 may be omitted, or at least two of the data source 110, the data source 112, the data source 114, or the data source 116 may be a single data source.

[0031] The network 102 may be a local area network (LAN), a wide area network (WAN), a cellular network (e.g., Verizon, AT&T), a satellite network (e.g., Starlink, Viasat), or another suitable network. For example, the network 102 may be or include Internet. The network 102 may be a single network (e.g., a WAN) or a set of networks (e.g., a LAN and a WAN communicably coupled to each other).

[0032] The computing instance 104 may be a physical server or a set of physical servers running an application program or a set of application programs thereon, as further described below. The computing instance 104 may be hosted in a data center or a set of data centers, which may be technologically advantageous to enable computational distribution for computational resiliency or computational redundancy. For example, the computing instance 104 may be or include a cloud computing instance (e.g., Amazon Web Services (AWS), Google Cloud) operative as disclosed herein.

[0033] The computing terminal 106 may be a physical computer hosting an application thereon. The physical computer may be embodied as a desktop computer, a laptop computer, a tablet computer, a smartphone, a wearable computer, or another suitable computing form factor. The computing terminal 106 may host an operating system (OS), such as Windows, MacOS, iOS, Android, Linux, or another suitable OS, and an application, such as a browser, a mobile app, or another suitable software form factor, running on or hosted via the OS.

[0034] The language model 108 (e.g., an LLM, an SLM) may be initially trained on a quantity of unlabeled content (e.g., text, unstructured text, imagery (graphics, photos, videos), sounds) using a self-supervised learning algorithm or a semi-supervised learning algorithm to understand a set of corresponding data relationships. Then, the language model 108 may be further trained by fine-tuning or refining the set of corresponding data relationships via a supervised learning algorithm or a reinforcement learning algorithm. Once the language model 108 is sufficiently trained to enable operations as disclosed herein, the language model 108 is structured to have a data structure and organized to have a data organization. As such, the data structure and the data organization collectively enable the language model 108 to be prompted to perform a task, as disclosedherein. The language model 108 may be a general purpose model, which may excel at a range of tasks (e.g., generating a content for a user consumption) and may be prompted, i.e., programmed to receive a prompt (e.g. a request, a command, a query), to do something or accomplish a certain task. The language model 108 may be embodied as or accessible via a chatbot, whether internal to the language model 108 or external to the language model 108, whether internal to the computing instance 104 or external to the computing instance 104, where a user operating the computing terminal 106 chats with the chatbot in a human-readable form, where the human-readable form may be an unstructured text, over the network 102, for the chatbot to output a content (e.g., a text, an unstructured text, an image (graphics, photos, videos), a sound), i.e., to do something or accomplish a certain task. For example, the language model 108 may be embodied as or accessible via a ChatGPT Al chatbot developed by OpenAI and released in November 2022, based on OpenAI's GPT-3.5 and GPT-4 models and fine-tuned using supervised and reinforcement learning algorithms, which may include improvements thereto. The language model 108 may be embodied as or accessible via a Google Bard / Gemini Al chatbot developed by Google and released in March 2023, based on a PaLM model. The chatbot may be omitted and the language model 108 may be accessed directly. For example, the language model 108 may host or be accessed via an Application Programming Interface (API), whether via the chatbot or directly. For example, the API may be a Representational State Transfer (REST) API.

[0035] The data source 110 may be an API, a data service, a data file, or a database (e.g., relational, graph, vector) outputting a copy of a set of records for a set of physical land lots in a geographic area (e.g., a block, a street, a neighborhood, a town, a country, a state, a country, a user-selected area, a user-drawn area) collected or updated over a period of time (e.g., a day, a week, a month), where the set of records one-to-one corresponds to the set of physical land lots (e.g., domain-specific data). Note that the data source 110 may be a single data source 110 or a set of data sources 110, which may distributively geolocated for redundancy or resiliency. Each of such records may contains a set of fields, any of which may be populated with data (e.g., text, images (graphics, photos, videos), sounds) identifying a respective physical land lot (e.g., by a identifier, a longitude and a latitude, a nickname) and various attributes or characteristics of therespective physical land lot (e.g., a length, a width, an area, a current owner or renter name, a plumbing system type identifier, an electrical wiring type identifier) or any physical property thereon (e.g., a house, a shed), and, if appropriate, then in a corresponding unit of measurement (e.g., meters, feet, acres, dates, times, volumes, coordinates). For example, such attributes or characteristics may include a total square footage of a home on the respective physical land lot (if any), a set of details on any recent renovations or updates to the home, an age and a condition of major home appliances (e.g., a hot water heater), a description of any special features or amenities of the home, a number of bedrooms and bathrooms, a set of photos or videos of the home, a number of days the home has been listed, an amount of homeowners association fees (if applicable), a listing price and any relevant pricing information, a location details and property features, a unique identifier (e.g., alphanumeric) that uniquely identifies the respective physical land lot over the period of time, or any other suitable attributes or characteristics of the respective physical land lot. For example, the data source 110 may be an API, a data service, a data file, or a database hosted on a server or a cloud computing instance for a Multiple Listing Service (MLS), where the copy of the set of records may be a copy of a set of MLS listings output from the API, the data service, the data file, or the database, which may include a copy of a set of comments (e.g., text, images (graphics, photos, videos), sounds) for the set of records that are not publicly visible unless securely accessing the data source 110 (e.g., passwords, biometrics). The set of comments may be sourced from a set of computing terminals (e.g., desktop computers, laptop computers, smartphones, tablet computers, wearables) other than the computing instance 104 and the computing terminal 106. The set of comments may be input into the data source 110 and may be correspond to the set of records in various forms of correspondence (e.g., one-to-one, one-to-many, many-to-one, many-to-many).

[0036] The data source 112 may be an API, a data service, a data file, or a database (e.g., relational, graph, vector) outputting a copy of a set of forms (e.g., static forms, dynamic forms, populatable forms, populated forms, semi-populated forms, in data file format, in PDF format, HTML format) relevant or related to the copy of the set of records one-to-one corresponding to the set of physical land lots in the geographic area collected or updated over the period of time (e.g., domain-specific data). Note that the data source112 may be a single data source 112 or a set of data sources 112, which may distributively geolocated for redundancy or resiliency. Each of such forms may have a set of fields (whether static or dynamic), any of which may be populatable or populated with data (e.g., text, image (graphics, photos, videos), sounds) relevant or related to the copy of the set of records one-to-one corresponding to the set of physical land lots in the geographic area collected or updated over the period of time. For example, the data source 112 may be an API, a data service, a data file, or a database hosted on a server or a cloud computing instance for a government web portal (e.g., municipal level, state level, federal level, agency level) or a library web portal (e.g., Westlaw, Lexis), where the copy of the set of forms may be a copy of a set government forms output from the API, the data service, the data file, or the database, where the copy of government forms may be relevant or related to the copy of the set of records one-to-one corresponding to the set of physical land lots in the geographic area collected or updated over the period of time. Note that the set of government forms is not required and there may be other types of forms, such as non-government forms (e.g., banking forms, accounting forms).

[0037] The data source 114 may be an API, a data service, a data file, or a database (e.g., relational, graph, vector) outputting a copy of a set of legislative texts (e.g., laws, statutes, regulations, codes), which may unstructured (e.g., descriptive or freeform) or structured (e.g., within a JavaScript Object Notation (JSON) files), whether formatted (e.g., listings, headings, bolding, underlining, italicizing) or plain text (e.g., domain-specific data). Note that the data source 114 may be a single data source 114 or a set of data sources 114, which may distributively geolocated for redundancy or resiliency. Each of such texts may be on municipal level, state level, federal level, agency level, or another suitable governmental level. Each of such texts may be hierarchically organized (e.g., by titles, sub-titles, chapters, sub-chapters, clauses, sub-clauses). The copy of the set of legislative texts may be relevant or related to the copy of the set of records one-to-one corresponding to the set of physical land lots in the geographic area collected or updated over the period of time. For example, the data source 112 may be an API, a data service, a data file, or a database hosted on a server or a cloud computing instance for a government web portal (e.g., on municipal level, state level, federal level, agency level) or a library web portal (e.g., Westlaw, Lexis), where the copy of the set of legislative texts may be a copy of theset of legislative texts output from the API, the data service, the data file, or the database, where the set of legislative texts may be relevant or related to the copy of the set of records one-to-one corresponding to the set of physical land lots in the geographic area collected or updated over the period of time or the copy of the set of forms.

[0038] The data source 116 may be an API, a data service, a data file, or a database (e.g., relational, graph, vector) outputting a copy of a set of personal communications (e.g., email messages, text messages, chat messages, Over-The-Top (OTT) messages), a copy of an archive of the set of personal communications, or a copy of an index of the set of personal communications (e.g., domain-specific data). For example, the set of personal communications may total 1000, but can be less or more. Note that the data source 116 may be a single data source 116 or a set of data sources 116, which may distributively geolocated for redundancy or resiliency. For example, the single data source 116 may output the copy of the set of personal communications, the copy of the archive of the set of personal communications, and the copy of the index of the set of personal communications. Likewise, for example, the set of data sources 116, which may distributively geolocated for redundancy or resiliency, may output the copy of the set of personal communications, the copy of the archive of the set of personal communications, and the copy of the index of the set of personal communications, where at least one data source 116 of the set of data sources 116 may output at least two of the copy of the set of personal communications, the copy of the archive of the set of personal communications, and the copy of the index of the set of personal communications.

[0039] The set of personal communications is intraorganizational or internal to a private electronic communication domain (e.g., a corporate email system, an enterprise chat system) and among / between user profiles within the private electronic communication domain (e.g., employee profiles). For example, the set of personal communications can include corporate email messages (e.g., Microsoft Outlook messages), corporate chat messages (e.g., Microsoft Teams messages), corporate text messages (e.g., MessageDesk messages, Heymarket messages, Avochato messages, Twilio messages), or other suitable corporate personal communications. For example, the set of personal communications can exclude external marketing communications (e.g., marketing emails, marketing chat messages) or customer communications (e.g.,customer emails, customer chat messages), although these configurations are not required and may be included in some embodiments.

[0040] At least some personal communications of the copy of the set of personal communications may contain text (e.g., unstructured, descriptive), images (graphics, photos, videos), or sounds. The set of personal communications may be relevant or related to the copy of the set of records one-to-one corresponding to the set of physical land lots in the geographic area collected or updated over the period of time or the copy of the set of forms or the copy of the set of legislative texts. Additionally or alternatively, the set of personal communications may be relevant or related to or a set of records recording a set of transactions or events involving a set of physical land lots in a geographic area collected or updated over a period of time that is different in value or format from the copy of the set of records one-to-one corresponding to the set of physical land lots in the geographic area collected or updated over the period of time sourced from the data source 110 (e.g., different physical properties or different land lots).

[0041] FIG. 2 shows a diagram of an embodiment of a computing instance according to this disclosure. In particular, the computing instance 104 hosts an application program or a set of application programs 104.1 , as referenced in context of FIG. 1. Also, the computing instance 104 hosts a vector search instance 104.2, a database 104.3 (e.g., a knowledge graph), and a natural language processing (NLP) engine (e.g., a task- dedicated logic that can be started, paused, or stopped) 104.4. As such, the computing instance 104 hosts or contains a RAG framework collectively enabled at least via the vector search instance 104.2 and the database 104.3.

[0042] The application program or the set of application programs 104.1 interfaces (e.g., communicates, commands) with the vector search instance 104.2, the database 104.3 (e.g. , relational, graph, vector), and the NLP engine 104.4. The application program or the set of application programs 104.1 may interface with the language model 108 and host or invoke the RAG framework, which combines at least some generative capabilities of the language model 108 with a retrieval function that searches relevant data sources to provide the language model 108 with up-to-date and context-specific information, which may allow the language model 108 to generate more accurate, relevant, and trustworthy responses by grounding its outputs in verifiable facts (also known as a knowledge cutoffissue), rather than just relying on its own training data which can be outdated or incomplete, while also minimizing frequent model retraining, which may save time and computational resources. As such, the RAG framework may have a retriever function, and the application program or the set of application programs 104.1 may call, instantiate, or invoke the retriever function to use the vector search instance 104.2 to scan through the database 104.3 (e.g., a knowledge base) storing a set of vector embeddings contained in or obtained from a copy of a set of data (e.g., texts, documents, messages, forms) stored in the database 104.3, as sourced from at least one of the data source 110, the data source 112, the data source 114, or the data source 116, and retrieve at least some of most relevant vector embeddings to help answer a given query by augmenting an original query with the at least some of most relevant vector embeddings, before inputting the original query with the at least some of most relevant vectors into the language model 108, as disclosed herein. Such processing may include indexing the copy of the set of data (e.g., documents or text chunks are converted into vector embeddings and stored in the database 104.3 enabling similarity-based retrieval), query vectorization (e.g., a new user query is converted into a vector embedding to be compatible with indexed vectors), retrieval of relevant vectors (e.g., vector similarity algorithms (e.g., cosine similarity, Euclidean distance) identify whatever vectors in the knowledge base that are most similar or relevant to a respective query vector to be subsequently retrieved to provide context for the language model 108 generative task). Therefore, the vector search instance 104.2 provides semantic retrieval (e.g., beyond keyword matching), efficient knowledge access (e.g., rapid and relevant retrieval), minimizes hallucination (e.g., grounded information), or data type flexibility (e.g., text, imagery (graphics, photos, videos), sounds).

[0043] The NLP engine 104.4 is programmed to semantically or linguistically interpret and understanding human language input (e.g., text, unstructured text, descriptive text), which may occur via intent classification (e.g., analyzes an input to determine an underlying intent or meaning to understand what is attempted to achieve or communicate), entity extraction (e.g., identifies and extracts key entities (e.g., people, places, organizations) in an input to gather context), and contextual understanding (e.g., maintain context across multiple user interactions to understand fuller or more precise meaningand intent). Although the NLP engine 104.4 is shown as being hosted or contained within the computing instance 104 external to the RAG framework, this configuration is not required and the RAG framework can host or contain the NLP engine 104.4.

[0044] The application program or the set of application programs 104.1 accesses, calls, or invokes the NLP engine 104.4 and the RAG framework, which contains the vector search instance 104.2 and the database 104.3, to prompt the language model 108 based on the copy of the set of data (e.g., texts, documents, messages, forms) stored in the database 104.3, as sourced from at least one of the data source 110, the data source 112, the data source 114, or the data source 116. For example, the RAG framework enhances the language model 108 by integrating real-time information retrieval with the language model 108 response generation process, where the RAG framework involves several logical components (e.g., software modules, application programs, engines) working together to perform indexing (indexer), retrieval (retriever), augmentation (augmenter), and generation (generator). The first step, indexing, prepares external data for efficient retrieval. This step involves collecting unstructured (e.g., text documents), semi-structured (e.g., JSON files), or structured data (e.g., databases) and converting this information into the set of embeddings - numerical representations that capture the semantic meaning of the text. These embeddings are generated using embedding models, which may be based on transformer architectures, and are stored in the database 104.3. The database 104.3 may be optimized for similarity searches, allowing for quick location of relevant information. The next step, retrieval, is triggered when a user operating the computing terminal 106 submits a query. The query is converted into its own embedding using the same embedding model used during indexing. A similarity search may then performed by the vector search instance 104.2 in the database 104.3 to identify the most relevant pieces of information based on semantic closeness in the vector space. Logical components responsible for retrieval include similarity scoring algorithms (e.g., cosine similarity or dot product) and retrieval methods, such as dense retrieval (using neural embeddings) or hybrid retrieval (combining dense and keyword-based search). These components ensure that only the most contextually relevant data is fetched. Once the relevant information is retrieved, the augmentation step combines this data with the user’s original query to form an augmented prompt. This step relies on prompt engineeringtechniques to structure the input in a way that maximizes the language model 108 ability to use both its pre-trained knowledge and the retrieved information. Logical components here may include query expansion algorithms, which may enrich the query with synonyms or related terms, and reranking mechanisms that refine the list of retrieved documents based on relevance scores. Next, in the generation step, the augmented prompt is passed to the language model 108, which generates the response by synthesizing its internal knowledge with the retrieved external data. The logical component responsible for this synthesis may be a decoder-based transformer model that processes the input prompt and produces coherent and contextually accurate text as output. These operations may further employ advanced RAG techniques to enhance the efficiency, accuracy, and relevance of both the retrieval and generation processes mentioned above. These approaches may address challenges such as retrieving contextually relevant data, refining search results, and improving the overall quality of generated responses. Some of these advanced RAG techniques may include improved retrieval strategies, re-ranking mechanisms, query expansion, and dynamic embeddings.

[0045] Improved retrieval methods focus on enhancing how relevant information is identified from large datasets. Dense retrieval leverages deep learning models to map both queries and documents into dense vector representations, enabling semantic similarity searches that go beyond simple keyword matching. This approach ensures that semantically related content is retrieved even if such content does not contain exact query terms. Hybrid search combines dense retrieval with traditional keyword-based (sparse) methods to balance semantic understanding and precision. By fusing results from both approaches — often using weighted scoring — hybrid search improves recall and precision, making it particularly effective for complex queries.

[0046] Re-ranking refines the initial set of retrieved documents by applying additional processing to prioritize the most relevant ones. This approach can involve simple scoring based on query-document similarity or more sophisticated machine learning models like cross-encoders or Bidirectional Encoder Representations from Transformers (BERT)- based re-rankers. These models jointly encode the query and document to provide a more precise relevance score, ensuring that the most contextually appropriate documents are ranked higher.

[0047] Query expansion enriches the user’s original query by adding related terms or concepts to improve retrieval performance. Techniques, such as synonym expansion or conceptual expansion can help capture broader meanings or alternative phrasings of a query, increasing the likelihood of retrieving relevant information. This is particularly useful in cases where the initial query might be too narrow or ambiguous.

[0048] Dynamic embeddings use context-aware models to generate embeddings that better capture nuanced meanings in language. Unlike static embeddings, which remain fixed regardless of context, dynamic embeddings adapt based on the specific query or document being processed. This allows for more accurate semantic representation and improves both retrieval and generation quality.

[0049] The application program or the set of application programs 104.1 may use canonical reference as metadata for any text chunk for the RAG framework. This approach is technologically advantageous, because this approach improves the language model 108 distinguishing content where similar language is used under different contexts. For example, this reference may be global like US / Arizona / Title 32 / Section 2198.01 / or US / Arizona / Form A / Page X / Line Y / . Likewise, the application program or the set of application programs 104.1 may use extra match context in the RAG framework that is not a fixed length document chunk. To keep context particularly in regulation text, the application program or the set of application programs 104.1 may enable an atomic unit of text chunks like a whole regulation article or a whole feedback text even when a search is based on a smaller chunk of text (e.g., a few sentences). This approach is technologically advantageous, because this approach enables both atomic context and global reference of text to improve the language model 108 ability to distinguish relevant information to provide better answers to user queries as measured by expert judgment.

[0050] FIG. 3 shows a flowchart of an embodiment of an algorithm to form an index based on a set of forms according to this disclosure. In particular, an algorithm 300 includes a set of steps 302-306 performed by the topology 100 to perform an algorithm of FIG. 4.

[0051] In step 302, the application program or the set of application programs 104 access a first data source (e.g., a form repository) hosting a first set of data (e.g., a set of populatable computerized forms) and a second data source (e.g., a listing repository)hosting a second set of data (e.g., a set of computerized listings of a set of physical land lots or physical properties) over the network 102. For example, the first data source may host a set of populatable computerized forms, each embodied in a Portable Document Format (PDF) format, LaTeX format, or another suitable format. For example, the set of populatable computerized forms may be embodied each as a form-fillable PDF that adheres to specific standards for compatibility with a particular electronic system, where the form-fillable PDF may include an interactive field (e.g., a checkbox, a radio button, a text fillable field), or be programmed for data embedding (e.g., completed forms can be saved with embedded data, enabling modifications and re-use, or other suitable configurations (e.g., embedded fonts). For example, at least some populatable computerized forms of the set of populatable computerized forms may be not populated at all (e.g., empty fields) or at least some populatable computerized forms of the set of populatable computerized forms may be at least partially prepopulated (e.g., some fields) with default values (e.g., binary, numeric, alphabetic, alphanumeric, text, calendar dates), each of which may be edited by the language model 108, as further described below. For example, a computerized listing may be a virtual listing manifested as a profile describing a physical land lot, including a physical property thereon or its attributes. The first data source may be the data source 112 accessed over the network 102 shown in FIG. 1. The second data source may be the data source 110 accessed over the network 102 shown in FIG. 1. The application program or the set of application programs 104 may access the first data source or the second data source as commanded by the computing terminal 106 or as programmed otherwise.

[0052] In step 304, the application program or the set of application programs 104 download a first copy of the first set of data from the first data source and a second copy of the second set of data from the second data source over the network 102. The first copy may be a single file or an archive. The second file may be a single file or an archive.

[0053] In step 306, the application program or the set of application programs 104 form a first index (e.g., via the vector search instance 104.2 and the database 104.3) of the first copy and a second index (e.g. , via the vector search instance 104.2 and the database 104.3) of the second copy. For example, as explained above, the first index or the second index may be formed by collecting unstructured (e.g., text documents), semi-structured(e.g., JSON files), the set of populatable computerized forms, or structured data (e.g., databases) and converting this information into the set of embeddings - numerical representations that capture the semantic meaning of the text. These embeddings are generated using embedding models, which may be based on transformer architectures, and are stored in the database 104.3. The first index and the second index are searchable and the database 104.3 may be optimized for similarity searches, allowing for quick location of relevant information.

[0054] FIG. 4 shows a flowchart of an embodiment of an algorithm to use an index to output a form that has been at least partially populated according to this disclosure. In particular, an algorithm 400 includes a set of steps 402-416 performed by the topology 100 based on the algorithm 300 shown in FIG. 3.

[0055] In step 402, the application program or the set of application programs 104 receive a user query (e.g., an unstructured text) about a first subset (e.g., a form that is populatable) of a first set of data (e.g., a set of forms that are populatable) in context of a second subset (e.g., a computerized listing) of a second set of data (e.g., a set of computerized listings of a set of physical land lots) from the computing terminal 106 over the network 102. For example, the user query may be “Which field in Form 123 issued by State of Arizona do I populate if I would like to buy 123 Main Street in Phoenix Arizona?” or “How do I populate field C in form D52 issued by City of New York if I would like to rent an apartment at 10 Ocean Parkway in Brooklyn New York?” or another suitable query.

[0056] In step 404, the application program or the set of application programs 104 uses the NLP engine 104.4 to parse (e.g., linguistically) the user query in context of the first index and the second index. The user query is converted into its own embedding using the same embedding model used during indexing, as described above.

[0057] In step 406, the application program or the set of application programs 104 uses the NLP engine 104.4 to form a semantic interpretation of the user query in context of the first index and the second index, which may be via the RAG framework. The semantic interpretation may reference the first subset from the first index or the second subset from the second index or contain a copy of the first subset sourced from the first index or a copy of the second subset from the second index. A similarity search may then performed by the vector search instance 104.2 in the database 104.3 to identify the mostrelevant pieces of information based on semantic closeness in the vector space. Logical components responsible for retrieval include similarity scoring algorithms (e.g., cosine similarity or dot product) and retrieval methods, such as dense retrieval (using neural embeddings) or hybrid retrieval (combining dense and keyword-based search). These components ensure that only the most contextually relevant data is fetched. Once the relevant information is retrieved, the augmentation step combines this data with the user’s original query to form an augmented prompt based on the first index and the second index (e.g., information relevant to the user query sourced from the first index and the second index). This step relies on prompt engineering techniques to structure the input in a way that maximizes the language model 108 ability to use both its pre-trained knowledge and the retrieved information. Logical components here may include query expansion algorithms, which may enrich the query with synonyms or related terms, and reranking mechanisms that refine the list of retrieved documents based on relevance scores.

[0058] In step 408, the application program or the set of application programs 104 feed the semantic interpretation into the language model 108 (e.g., by prompting) over the network 102. The augmented prompt is passed to the language model 108, which enables the language model 108 to generate the response by synthesizing its internal knowledge with the retrieved external data. The logical component responsible for this synthesis may be a decoder-based transformer model that processes the input prompt and produces coherent and contextually accurate text as output.

[0059] In step 410, the application program or the set of application programs 104 receive a response to the user query from the language model 108 over the network 102. The response to the user query may be embodied in several ways, as exemplified below.

[0060] In step 412, the application program or the set of application programs 104 receives the response with a phrased response (e.g., an unstructured text) over the network 102. For example, if the user query may be “Which field in Form 123 issued by State of Arizona do I populate if I would like to buy 123 Main Street in Phoenix Arizona?” or “How do I populate field C in form D52 issued by City of New York if I would like to rent an apartment at 10 Ocean Parkway in Brooklyn New York?” then the phrased response may respectively state that “If you would like buy 123 Main Street in Phoenix Arizona, then you need to populate fields 1-3, 7 and 10 in Form 123 issued by State of Arizona byrespectively inserting 123 Main Street in Phoenix Arizona into these fields” or “If you would like to rent an apartment at 10 Ocean Parkway in Brooklyn New York, then field C in form D52 issued by City of New York should be populated by inserting your building number.”

[0061] In step 414, the application program or the set of application programs 104 receives the response with a copy of the first subset (e.g., a form) having at least one field populated (e.g., text, images (graphics, photos, videos), data) by the language model 108 with a piece of information (e.g., text, image (graphics, photos, videos), sound) sourced from a copy of the second subset data (e.g., a computerized listing) over the network 102. Such population can include adding new data or editing preexisting data (e.g., default data). For example, such population may be customized as situationally appropriate, with the piece of information, based on the user query. For example, the response may include a data file (e.g., a PDF file, a word processor file, a spreadsheet file) containing a set of labels and a set of fields associated with the set of labels, where at least some fields of the set of fields are populated, customized as situationally appropriate, as populated by the language model 108 with the piece of information (e.g., a physical address) sourced from the copy of the second subset data.

[0062] In step 416, the application program or the set of application programs 104 present the response to the user query 416 to the computing terminal 106 over the network 102. When the response includes the copy of the first subset having at least one field populated by the language model 108 with the piece of information sourced from the copy of the second subset data, the response may include an user input element (e.g., a button) or a reference (e.g., a Uniform Resource Locator (URL), a hyperlink) to access (e.g., download) the copy of the first subset having at least one field populated by the language model 108 with the piece of information sourced from the copy of the second subset data. For example, the response may be presented on the computing terminal 106 to enable the user input element or the reference to be selected or activated such that the copy of the first subset having at least one field populated by the language model 108 with the piece of information sourced from the copy of the second subset data can be downloaded. For example, when the copy of the first subset having at least one field populated by the language model 108 with the piece of information sourced from the copy of the second subset data can be downloaded is a data file (e.g., a PDF file, a Latex file),then the data file can be downloaded onto the computing terminal 106. Such configuration improves computer functionality and is technologically beneficial, because of its enablement of the language model 108 to output forms (e.g., as populated), which the language model is not known to do, especially when specialized data (e.g., forms, personal communications, private comments) is used for augmenting user queries.

[0063] FIG. 5 shows a flowchart of an embodiment of an algorithm to form an index based on a legislative text according to this disclosure. In particular, an algorithm 500 includes a set of steps 502-506 performed by the topology 100, similar to the algorithm 300, to perform an algorithm of FIG. 6.

[0064] In step 502, the application program or the set of application programs 104 access a first data source (e.g., a legislative repository) hosting a first set of data (e.g., a set of computerized legislations embodied in a text format) and a second data source (e.g., a listing repository) hosting a second set of data (e.g., a set of computerized listings of a set of physical land lots) over the network 102. For example, the first data source may host a set of legislative texts (e.g., hierarchically organized by sections, clauses, or paragraphs), each embodied in raw text, unstructured text, formatted text, or descriptive text, whether static or dynamic (e.g., hierarchically expandable and collapsable by user activation). The first data source may be the data source 114 accessed over the network 102 shown in FIG. 1 . The second data source may be the data source 110 accessed over the network 102 shown in FIG. 1. The application program or the set of application programs 104 may access the first data source or the second data source as commanded by the computing terminal 106 or as programmed otherwise.

[0065] In step 504, the application program or the set of application programs 104 download a first copy of the first set of data from the first data source and a second copy of the second set of data from the second data source over the network 102. The first copy may be a single file or an archive. The second file may be a single file or an archive.

[0066] In step 506, the application program or the set of application programs 104 form a first index (e.g., via the vector search instance 104.2 and the database 104.3) of the first copy and a second index (e.g. , via the vector search instance 104.2 and the database 104.3) of the second copy. For example, as explained above, the first index or the second index may be formed by collecting unstructured (e.g., text documents), semi-structured(e.g., JSON files), the set of populatable computerized forms, or structured data (e.g., databases) and converting this information into the set of embeddings - numerical representations that capture the semantic meaning of the text. These embeddings are generated using embedding models, which may be based on transformer architectures, and are stored in the database 104.3. The first index and the second index are searchable and the database 104.3 may be optimized for similarity searches, allowing for quick location of relevant information.

[0067] FIG. 6 shows a flowchart of an embodiment of an algorithm to use an index to output an opinion as to whether a legislative (or another type) text is appropriate or applicable to a set of parameters according to this disclosure. In particular, an algorithm 600 includes a set of steps 602-614 performed by the topology 100 based on the algorithm 500 shown in FIG. 5, similar to the algorithm 400.

[0068] In step 602, the application program or the set of application programs 104 receive a user query (e.g., an unstructured text) about a first subset (e.g., a legislative text) of a first set of data (e.g., a set of legislative texts) in context of a second subset (e.g., a computerized listing) of a second set of data (e.g., a set of computerized listings of a set of physical land lots) from the computing terminal 106 over the network 102. For example, the user query may be “Does section 5 of NYC building code apply to a buyer of 123 Main Street in Queens NY?” or “Which section of statute 123 in applies to a seller of 3 First Avenue, Denver Colorado?” or another suitable query.

[0069] In step 604, the application program or the set of application programs 104 uses the NLP engine 104.4 to parse (e.g., linguistically) the user query in context of the first index and the second index. The user query is converted into its own embedding using the same embedding model used during indexing, as described above.

[0070] In step 606, the application program or the set of application programs 104 uses the NLP engine 104.4 to form a semantic interpretation of the user query in context of the first index and the second index, which may be via the RAG framework. The semantic interpretation may reference the first subset from the first index or the second subset from the second index or contain a copy of the first subset sourced from the first index or a copy of the second subset from the second index. A similarity search may then performed by the vector search instance 104.2 in the database 104.3 to identify the mostrelevant pieces of information based on semantic closeness in the vector space. Logical components responsible for retrieval include similarity scoring algorithms (e.g., cosine similarity or dot product) and retrieval methods, such as dense retrieval (using neural embeddings) or hybrid retrieval (combining dense and keyword-based search). These components ensure that only the most contextually relevant data is fetched. Once the relevant information is retrieved, the augmentation step combines this data with the user’s original query to form an augmented prompt based on the first index and the second index (e.g., information relevant to the user query sourced from the first index and the second index). This step relies on prompt engineering techniques to structure the input in a way that maximizes the language model 108 ability to use both its pre-trained knowledge and the retrieved information. Logical components here may include query expansion algorithms, which may enrich the query with synonyms or related terms, and reranking mechanisms that refine the list of retrieved documents based on relevance scores.

[0071] In step 608, the application program or the set of application programs 104 feed the semantic interpretation into the language model 108 (e.g., by prompting) over the network 102. The augmented prompt is passed to the language model 108, which enables the language model 108 to generate the response by synthesizing its internal knowledge with the retrieved external data. The logical component responsible for this synthesis may be a decoder-based transformer model that processes the input prompt and produces coherent and contextually accurate text as output.

[0072] In step 610, the application program or the set of application programs 104 receive a response to the user query from the language model 108 over the network 102.

[0073] In step 612, the application program or the set of application programs 104 receives the response as a phrased response stating yes, no, can, or cant. For example, if the user query inquired whether a legislative text is applicable to a computerized listing, then the phrased response may be yes or no (or some other suitable semantic equivalent thereof). Likewise, if the user query inquired whether a legislative text allows something physical to be done (e.g., build a swimming pool, remove trees, add floors) to a physical land lot identified in a computerized listings, then the phrased response may be can or cannot (or some other suitable semantic equivalent thereof).

[0074] In step 612, the application program or the set of application programs 104 present the response to the user query on the computing terminal 106 over the network 102. Such configuration improves computer functionality and is technologically beneficial, because of its enablement of the language model to output opinions, which the language model is not known to do, especially when specialized data (e.g., forms, personal communications, private comments) is used for augmenting user queries.

[0075] FIG. 7 shows a flowchart of an embodiment of an algorithm to form an index based on a set of personal communications according to this disclosure. In particular, an algorithm 700 includes a set of steps 702-706 performed by the topology 100, similar to the algorithm 300 or the algorithm 500, to perform an algorithm of FIG. 8.

[0076] In step 702, the application program or the set of application programs 104 access an archive of internal or intraorganizational communications hosted by a data source. The archive of internal or intraorganizational communications may be a data file, a compressed file, a ZIP file, a virtual directory, a virtual folder, or another suitable logical form. The data source may be the data source 116 accessed over the network 102 shown in FIG. 1.

[0077] The archive of internal or intraorganizational communications may contain a set of personal communications (e.g., email messages, text messages, chat messages, OTT messages), a copy of an archive (e.g., a data file, a compressed file, a ZIP file, a virtual directory, a virtual folder) of the set of personal communications, or a copy of an index of the set of personal communications. For example, the set of personal communications may total 1000, but can be less or more.

[0078] The set of personal communications is intraorganizational or internal to a private electronic communication domain (e.g., a corporate email system, an enterprise chat system) and among / between user profiles within the private electronic communication domain (e.g., employee profiles). For example, the set of personal communications can include corporate email messages (e.g., Microsoft Outlook messages), corporate chat messages (e.g., Microsoft Teams messages), corporate text messages (e.g., MessageDesk messages, Heymarket messages, Avochato messages, Twilio messages), or other suitable corporate personal communications. For example, the set of personal communications may span a date range (e.g., 1 / 1 / 2000 to 1 / 1 / 2011 ),which may be not effectively or practically used for internal or intraorganizational communications (e.g., stored offsite, magnetic tape). For example, the set of personal communications can exclude external marketing communications (e.g., marketing emails, marketing chat messages) or customer communications (e.g., customer emails, customer chat messages), although these configurations are not required and may be included in some embodiments.

[0079] At least some personal communications of the set of personal communications, the copy of the set of personal communications, or the copy of the index of the set of personal communications may contain text, images, or sounds. The set of personal communications, the copy of the set of personal communications, or the copy of the index of the set of personal communications may be relevant or related to the copy of the set of records one-to-one corresponding to the set of physical land lots in the geographic area collected or updated over the period of time, or the copy of the set of forms or the copy of the set of legislative texts, as described above. For example, the set of personal communications, the copy of the set of personal communications, or the copy of the index of the set of personal communications, or a subset of any thereof, may be relevant or related to an electronic transaction (e.g., an electronic receipt evidencing a sale or a lease of a physical land lot) or a set of electronical transactions referencing (e.g., by a physical location or address or an attribute or a characteristic of a physical land lot) at least one record of the set of records one-to-one corresponding to the set of physical land lots in the geographic area collected or updated over the period of time where the at least one record corresponds to at least one physical land lot of the set of physical land lots in the geographic area collected or updated over the period of time (e.g., a subset of personal communications related to a transaction referencing a sale or a lease of a physical land lot by its address or geographic coordinates). Likewise, for example, the set of personal communications, the copy of the set of personal communications, or the copy of the index of the set of personal communications, or a subset of any thereof, may be relevant or related to an electronic transaction (e.g. , an electronic receipt evidencing a sale or a lease of a physical land lot) or a set of electronical transactions referencing (e.g., by a physical location or address or an attribute or a characteristic of a physical land lot) the copy of the set of forms or the copy of the set of legislative texts (e.g., a subset of personalcommunications related to a transaction referencing a form or a legislative text related to a sale or a lease of a physical land lot by its address or geographic coordinates), as described above. Additionally or alternatively, the set of personal communications, the copy of the set of personal communications, or the copy of the index of the set of personal communications, or a subset of any thereof, may be relevant or related to or a set of records recording a set of transactions or events involving a set of physical land lots in a geographic area collected or updated over a period of time that is different in value from the copy of the set of records one-to-one corresponding to the set of physical land lots in the geographic area collected or updated over the period of time sourced from the data source 110 (e.g., physical land lots different from those mentioned above).

[0080] In step 704, the application program or the set of application programs 104 download a copy of the archive from the data source over the network 102. The copy may be embodied as a data file, a compressed file, a ZIP file, or another suitable logic.

[0081] In step 706, the application program or the set of application programs 104 form an index (e.g., the vector search instance 104.2) of the copy. For example, as explained above, the index may be formed by collecting unstructured (e.g., text documents), semistructured (e.g., JSON files), the set of populatable computerized forms, or structured data (e.g., databases) and converting this information into the set of embeddings - numerical representations that capture the semantic meaning of the text. For example, this process may involve transcription of audio or video or Optical Character Recognition (OCR) of PDF files or images if present in at least some personal communications of the set of personal communications. For example, the set of personal communications may be normalized to text files with metadata that provide context. For example, a PDF file may be converted into an annotated text file reading like ‘Form A, page 2, section 3, Line 4 says that ...’. For example, there may be extraction of key entities, terms, and relationships for knowledge graph construction. Note that these operations are not limited to the algorithm 700 and may be similarly performed in the algorithm 300 or the algorithm 500 or the algorithm 900, when an appropriate index is being formed. Regardless, these embeddings are generated using embedding models, which may be based on transformer architectures, and are stored in the database 104.3. The first index and thesecond index are searchable and the database 104.3 may be optimized for similarity searches, allowing for quick location of relevant information.

[0082] FIG. 8 shows a flowchart of an embodiment of an algorithm to use an index to output an opinion based on an index of personal communications according to this disclosure. In particular, an algorithm 800 includes a set of steps 802-816 performed by the topology 100 based on the algorithm 700 shown in FIG. 7, similar to the algorithm 400 or the algorithm 600.

[0083] In step 802, the application program or the set of application programs 104 receive a user query (e.g., an unstructured text) about a subset (e.g., a computerized listing) of a set of data (e.g., a set of computerized listings) from the computing terminal 106 over the network 102. For example, the user query may be “How do I complete form X for a seller of 123 Main Street, Trenton New Jersey?” or “What section of a local building code is applicable to a seller of 123 Green Avenue in Dallas Texas if there is a heated pool?” or another suitable query.

[0084] In step 804, the application program or the set of application programs 104 uses the NLP engine 104.4 to parse (e.g., linguistically) the user query in context of the index, as disclosed in context of FIG. 7. The user query is converted into its own embedding using the same embedding model used during indexing, as described above.

[0085] In step 806, the application program or the set of application programs 104 uses the NLP engine 104.4 to form a semantic interpretation of the user query in context of the index, as disclosed in context of FIG. 7, which may be via the RAG framework. The semantic interpretation may reference the index. A similarity search may then performed by the vector search instance 104.2 in the database 104.3 to identify the most relevant pieces of information based on semantic closeness in the vector space. Logical components responsible for retrieval include similarity scoring algorithms (e.g., cosine similarity or dot product) and retrieval methods, such as dense retrieval (using neural embeddings) or hybrid retrieval (combining dense and keyword-based search). These components ensure that only the most contextually relevant data is fetched. Once the relevant information is retrieved, the augmentation step combines this data with the user’s original query to form an augmented prompt based on the index (e.g., information relevant to the user query sourced from the index). This step relies on prompt engineeringtechniques to structure the input in a way that maximizes the language model 108 ability to use both its pre-trained knowledge and the retrieved information. Logical components here may include query expansion algorithms, which may enrich the query with synonyms or related terms, and reranking mechanisms that refine the list of retrieved documents based on relevance scores.

[0086] In step 808, the application program or the set of application programs 104 identifies a communication (e.g., an email, a text message, a chat message, an OTT message) within the copy of the archive based on the index relevant to the semantic interpretation, which may be via the RAG framework. For example, the communication may reference prior transactions (e.g., land lot sales or leases) or events involving semantically similar topics, in similar geographic locales. A similarity search may then performed by the vector search instance 104.2 in the database 104.3 to identify the most relevant pieces of information based on semantic closeness in the vector space. Logical components responsible for retrieval include similarity scoring algorithms (e.g., cosine similarity or dot product) and retrieval methods, such as dense retrieval (using neural embeddings) or hybrid retrieval (combining dense and keyword-based search). These components ensure that only the most contextually relevant data is fetched. Once the relevant information is retrieved, the augmentation step combines this data with the user’s original query to form an augmented prompt based on the index (e.g., information relevant to the user query sourced from the index). This step relies on prompt engineering techniques to structure the input in a way that maximizes the language model 108 ability to use both its pre-trained knowledge and the retrieved information. Logical components here may include query expansion algorithms, which may enrich the query with synonyms or related terms, and reranking mechanisms that refine the list of retrieved documents based on relevance scores.

[0087] In step 810, the application program or the set of application programs 104 generate a prompt, as augmented, to address the user query based on the semantic interpretation and the communication.

[0088] In step 812, the application program or the set of application programs 104 inputs the prompt, as augmented, into the language model 108 (e.g., by prompting) over the network 102. The augmented prompt is passed to the language model 108, whichenables the language model 108 to generate the response by synthesizing its internal knowledge with the retrieved external data (e.g., based on the index and the communication). The logical component responsible for this synthesis may be a decoderbased transformer model that processes the input prompt and produces coherent and contextually accurate text as output.

[0089] In step 814, the application program or the set of application programs 104 receive a response (e.g., a phrased response, a recommendation, a suggestion, a listing of possibilities, a ranking of possibilities based on a criterion) to the user query from the language model 108 over the network 102. For example, the response may answer the user query based on the index and the communication. For example, the response may include a guidance saying that “You can complete form X for a seller of 123 Main Street, Trenton New Jersey by filling out his full legal name and his current mailing address different from 123 Main Street, Trenton New Jersey” or “If there is a heated pool in 123 Green Avenue in Dallas Texas, then section 456 a local building code is applicable to the seller” or another suitable recommendation.

[0090] In step 816, the application program or the set of application programs 104 presents the response to the user query on the computing terminal 106 over the network 102. Such configuration improves computer functionality and is technologically beneficial, because of its enablement of the language model to output recommendations, which the language model is not known to do, especially when specialized data (e.g., forms, personal communications, private comments) is used for augmenting user queries.

[0091] FIG. 9 shows a flowchart of an embodiment of an algorithm to form an index based on a set of private comments according to this disclosure. In particular, an algorithm 900 includes a set of steps 902-906 performed by the topology 100, similar to the algorithm 300 or the algorithm 500 or the algorithm 700, to perform an algorithm of FIG. 10.

[0092] In step 902, the application program or the set of application programs 104 access (e.g., periodically) a data source (e.g., an MLS service) hosting a set of records (e.g., a set of private comments about a physical land lot or a physical property thereon, a set of public or private parameters of a physical land lot or a physical property thereon, a public or private set of attributes of a physical land lot or a physical property thereon)for a set of property identifiers (e.g., by address, geographical coordinates, name, unique name), which may be for a set of data (e.g., a set of computerized listings). For example, the set of private comments may be referred to as "Private Remarks" or "Agent-Only Remarks" and maybe sections within a computerized listing that are visible only to some user profiles (e.g., real estate agents) and not to general public. For example, the set of private comments may be showing instructions (e.g., "Sellers work night shifts, no morning showings"), issues with property access (e.g., "Key sticks, lift handle"), seller preferences for closing timelines, or others. For example, the set of private comments may exclude certain content (e.g., prohibited content), such as material facts, sold status, misleading terms, or public-facing information. For example, the set of private comments may remain confidential through a combination of access controls, computing privileges, and strict enforcement mechanisms. These comments are viewable only by authorized users, such as licensed real estate agents and brokers, who must log in to the MLS service using unique credentials. Access is further restricted to members of the MLS service or affiliated organizations, ensuring that private remarks are not exposed to the general public. At the system level, private comments are stored in designated fields within the MLS service database that are flagged as restricted. These fields are programmed to display only to users with specific permissions, such as cooperating agents, while being excluded from public-facing reports or client-accessible views. To enforce these restrictions, the MLS service may implement rules that prohibit the sharing of private remarks with unauthorized parties. Violations can result in penalties or fines for the responsible agents. Additionally, MLS organizations often employ data security measures such as encryption to safeguard sensitive information and conduct regular monitoring or audits to ensure compliance with access policies. These combined mechanisms ensure that private comments remain secure and are used appropriately within the professional real estate community. The data source may be the data source 110 accessed over the network 102, as shown in FIG. 1. The set of records may correspond to the set of property identifiers in various types of correspondence (e.g., one- to-one, one-to-many, many-to-one, many-to-many). For example, one property identifier may have one set of private user comments, whether associated with one user profile or multiple user profiles, and another property identifier have another set of private usercomments, whether associated with one user profile or multiple user profiles. For example, the application program or the set of application programs 104 may periodically access the data source by data scrubbing or spidering across the set of records, the set of property identifiers, or the set of data.

[0093] In step 904, the application program or the set of application programs 104 download a copy of the set of records, which may also include a copy of the set of property identifiers or a copy of the set of data storing the set of property identifiers. The copy may be embodied as a data file, a compressed file, a ZIP file, a productivity suite file, a word processor file, a spreadsheet file, a text file, a raw text file, a structured text file, a CSV file, a JSON file, or another suitable logic.

[0094] In step 906, the application program or the set of application programs 104 form an index (e.g., the vector search instance 104.2) of the copy. For example, as explained above, the index may be formed by collecting unstructured (e.g., text documents), semistructured (e.g., JSON files), the set of populatable computerized forms, or structured data (e.g., databases) and converting this information into the set of embeddings - numerical representations that capture the semantic meaning of the text. These embeddings are generated using embedding models, which may be based on transformer architectures, and are stored in the database 104.3. The first index and the second index are searchable and the database 104.3 may be optimized for similarity searches, allowing for quick location of relevant information.

[0095] FIG. 10 shows a flowchart of an embodiment of an algorithm to use an index to output an opinion based on an index of private comments according to this disclosure. In particular, an algorithm 1000 includes a set of steps 1002-1012 performed by the topology 100 based on the algorithm 900 shown in FIG. 9, similar to the algorithm 400 or the algorithm 600 or the algorithm 800.

[0096] In step 1002, the application program or the set of application programs 104 receive a user query (e.g., an unstructured text) for a recommendation in context of a factor (e.g., timing, price, special circumstances) associated with a property identifier (e.g., a computerized listing) from the computing terminal 106 over the network 102. The factor can be expressed as a time, a date, a numeric string, a text string (e.g., unstructured, descriptive), an alphanumeric string, or another suitable content. The property identifiermay be associated with the set of records referenced in FIG. 10. For example, the user query may be “How do I motivate a seller of 123 Main Street in Scottsdale Arizona to sell his house quicker?” or “What would motivate a seller of apartment 6 in 34 Apple Drive in Omaha Nebraska to lower her asking price?” or another suitable query.

[0097] In step 1004, the application program or the set of application programs 104 uses the NLP engine 104.4 to parse (e.g., linguistically) the user query in context of the index, as disclosed in context of FIG. 8.

[0098] In step 1006, the application program or the set of application programs 104 uses the NLP engine 104.4 to form a semantic interpretation of the user query in context of the index, as disclosed in context of FIG. 9, which may be via the RAG framework. The semantic interpretation may reference the factor, the property identifier or contain a copy of the factor or a copy of the property identifier. For example, there may be a set of private comments that reference a property, as noted above, that are relevant in motivating a seller to sell his house quicker, as noted above, or lower her asking price, as noted above. The semantic interpretation may reference the index. A similarity search may then performed by the vector search instance 104.2 in the database 104.3 to identify the most relevant pieces of information based on semantic closeness in the vector space. Logical components responsible for retrieval include similarity scoring algorithms (e.g., cosine similarity or dot product) and retrieval methods, such as dense retrieval (using neural embeddings) or hybrid retrieval (combining dense and keyword-based search). These components ensure that only the most contextually relevant data is fetched. Once the relevant information is retrieved, the augmentation step combines this data with the user’s original query to form an augmented prompt based on the index (e.g., information relevant to the user query sourced from the index). This step relies on prompt engineering techniques to structure the input in a way that maximizes the language model 108 ability to use both its pre-trained knowledge and the retrieved information. Logical components here may include query expansion algorithms, which may enrich the query with synonyms or related terms, and reranking mechanisms that refine the list of retrieved documents based on relevance scores.

[0099] In step 1008, the application program or the set of application programs 104 the semantic interpretation is input into the language model 108, as an augmented prompt,to address the user query based on the semantic interpretation. The augmented prompt is passed to the language model 108, which enables the language model 108 to generate the response by synthesizing its internal knowledge with the retrieved external data. The logical component responsible for this synthesis may be a decoder-based transformer model that processes the input prompt and produces coherent and contextually accurate text as output.

[0100] In step 1010, the application program or the set of application programs 104 receive a response from the language model 108 over the network 102, where the response is the recommendation inquired about in the user query. For example, the response may include a recommendation stating that “The seller of 123 Main Street in Scottsdale Arizona may be motivated to sell his house quicker if you promise to close within two weeks” or “The seller of apartment 6 in 34 Apple Drive in Omaha Nebraska may be motivated to lower her asking price if you promise to promise to pay in cash at closing” or another suitable recommendation.

[0101] In step 1012, the application program or the set of application programs 104 presents the response to the user query on the computing terminal 106 over the network 102. Such configuration improves computer functionality and is technologically beneficial, because of its enablement of the language model to output recommendations, which the language model is not known to do, especially when specialized data (e.g., forms, personal communications, private comments) is used for augmenting user queries.

[0102] FIG. 11 shows a flowchart of an embodiment of an algorithm to generate a recommendation for training according to this disclosure. In particular, an algorithm 1100 includes a set of steps 1102-1116 performed by the topology 100, sim ilar to the algorithm 400 or the algorithm 600 or the algorithm 800 or the algorithm 1000.

[0103] In step 1102, the application program or the set of application programs 104 track (e.g., continuously or over periods of time) a collection of prompts (e.g., a set of unstructured texts) submitted (e.g., serially, parallel) by a set of users (e.g., via a set of user profiles) to the language model 108 over the network 102 from the computing terminals 106, as related to a first set of data (e.g., a set of forms or legislative texts) in context of a second set of data (e.g., a set of computerized listings) by semantic topics (e.g., a form or a field thereof X, a law or a section thereof Y, a physical land lot or aphysical property thereon Z or an attribute or a description of any of foregoing). The collection of prompts may be a log, a list, an array, a tree, a graph, an archive, or another suitable data structure storing the collection of prompts. For example, the log or the archive may be a data file, a text file, a set of records, or another suitable storage form factor.

[0104] In step 1104, the application program or the set of application programs 104 determine whether a subset of the collection satisfies a numeric threshold and a semantic similarity threshold, which may be via the RAG framework including the NLP engine 104.4. The numeric threshold may be preset and indicate a minimum or floor amount for the subset (e.g., has a minimum amount of prompts been submitted). For example, the minimum or floor amount for the subset may be at least 25 prompts or at least 50 prompts. The semantic similarity threshold may be preset and indicate a semantic relationship between a group of prompts in the subset of the collection (e.g., are prompts semantically similar). For example, the prompts may be semantically similar if those prompts are prompting about a similar topic or a theme, such as some forms or fields thereof, some legislative texts or sections thereof, or some physical land lots or attributes thereof. In terms of semantic similarity, note that a similarity search may then performed by the vector search instance 104.2 in the database 104.3 to identify the most relevant pieces of information based on semantic closeness in the vector space. Logical components responsible for retrieval include similarity scoring algorithms (e.g., cosine similarity or dot product) and retrieval methods, such as dense retrieval (using neural embeddings) or hybrid retrieval (combining dense and keyword-based search). These components ensure that only the most contextually relevant data is fetched. Once the relevant information is retrieved, the augmentation step combines this data with the user’s original query to form an augmented prompt. This step relies on prompt engineering techniques to structure the input in a way that maximizes the language model 108 ability to use both its pre-trained knowledge and the retrieved information. Logical components here may include query expansion algorithms, which may enrich the query with synonyms or related terms, and reranking mechanisms that refine the list of retrieved documents based on relevance scores.

[0105] In step 1106, the application program or the set of application programs 104 identify a common semantic thread (e.g., a topic or a theme) in the subset, which may be via the RAG framework. For example, once there is the minimum or floor amount of prompts submitted to the language model 108, and those prompts are sufficiently semantically similar, then the application program or the set of application programs 104 determine what is semantically common (the common semantic thread) to these prompts, such as by topic, theme, concept, or inquiry. In terms of semantic similarity, note that a similarity search may then performed by the vector search instance 104.2 in the database 104.3 to identify the most relevant pieces of information based on semantic closeness in the vector space. Logical components responsible for retrieval include similarity scoring algorithms (e.g., cosine similarity or dot product) and retrieval methods, such as dense retrieval (using neural embeddings) or hybrid retrieval (combining dense and keywordbased search). These components ensure that only the most contextually relevant data is fetched. Once the relevant information is retrieved, the augmentation step combines this data with the user’s original query to form an augmented prompt. This step relies on prompt engineering techniques to structure the input in a way that maximizes the language model 108 ability to use both its pre-trained knowledge and the retrieved information. Logical components here may include query expansion algorithms, which may enrich the query with synonyms or related terms, and reranking mechanisms that refine the list of retrieved documents based on relevance scores.

[0106] In step 1108, the application program or the set of application programs 104 generate a prompt, which may be augmented as disclosed herein, for submission to the language model 108 based on the common semantic thread to request or identify a recommendation or an opportunity for training a subset of the set of users (e.g., via or associated with user profiles) who submitted the subset of the collection, to minimize the subset from repeating itself at a later point in time.

[0107] In step 1110, the application program or the set of application programs 104 submit the prompt to the language model 108 over the network 102. The prompt is passed to the language model 108, which enables the language model 108 to generate the response by synthesizing its internal knowledge with the retrieved external data. The logical component responsible for this synthesis may be a decoder-based transformermodel that processes the input prompt and produces coherent and contextually accurate text as output.

[0108] In step 1112, the application program or the set of application programs 104 receive a response to the prompt from the language model 108 over the network 102. The response may include a descriptive text expressing the recommendation or the opportunity for training the subset of the set of users who submitted the subset of the collection. For example, the descriptive text may state that “Since I see that 90% of users are asking about Arizona law dealing with easements, I suggest that you consider conducting a training session on this topic.”

[0109] In step 1114, the application program or the set of application programs 104 presents the response to the user query on the computing terminal 106 over the network 102. Such configuration improves computer functionality and is technologically beneficial, because of its enablement of the language model to output recommendations, which the language model is not known to do, especially when specialized data (e.g., forms, personal communications, private comments) is used for augmenting user queries.

[0110] In step 1116, which is optional, the application program or the set of application programs 104 may generate the prompt based on the application program or the set of application programs 104 gathering a set of metadata for the set of users (e.g., via user profiles) and having the application program or the set of application programs 104 consider the set of metadata in generating the prompt such that the prompt is at least partially generated based on the set of metadata. For example, the set of metadata can include personal demographics, prior history of training, computing privilege rights, and other suitable user attributes. This way, the recommendation is more customized or personalized.

[0111] FIG. 12 shows a screenshot of an embodiment of a Graphical User Interface (GUI) of an application program hosted on a computing terminal of FIG. 1 according to this disclosure. In particular, there is a screen 1200 showing a first viewing pane (left) for selection of a content to be presented in a second viewing pane (right). The first viewing pane shows four headings, although more or less is possible. These headings are dashboard, users, library, and feedbacks. The dashboard heading is currently selected in the first viewing pane such that the second viewing pane shows a first filter selector byuser type, a second filter selector by date / time, and a table presenting a set of rows disclosing a set of users by email address (or another identifier) on an Y-axis and a set of columns disclosing a set of parameters associated with the set of users on an X-axis, along with all user data, such that a cellular structure is presented. The table changes in content when the first selector or the second filter selector are selected for respective filtering. The screen 1200 may be used for consumption of statistical information, which corresponds to information relevant for recommendation or training, as disclosed herein.

[0112] FIG. 13 shows a screenshot of an embodiment of a GUI of an application program hosted on a computing terminal of FIG. 1 according to this disclosure. In particular, there is a screen 1300 where the feedbacks heading is selected in the first viewing pane, thereby causing the second viewing pane to present a table similar to the table shown in the screen 1200, but now having a selection column populated a set of checkboxes, a timestamp column populated with a set of timestamps, an email column populated with a set of email addresses corresponding to the set of timestamps, a question (e.g., the user query for generating the prompt) column populated with a set of questions submitted by a set of user profiles associated with the set of email addresses and the set of timestamps, and an answer column (e.g., the response from the language model 108) populated with a set of answers to the set of questions presented to the set of user profiles associated with the set of email addresses and the set of timestamps. Also, the table, the screen 1300 shows a third filter selector by sentiment and a fourth filter selector by state. As such, the table changes in content when the third selector or the fourth filter selector are selected for respective filtering. The screen 1300 may be used for consumption of statistical information, which corresponds to information relevant for recommendation or training, as disclosed herein.

[0113] FIG. 14 shows a screenshot of an embodiment of a GUI of an application program hosted on a computing terminal of FIG. 1 according to this disclosure. In particular, there is a screen 1 00 where the library heading is selected in the first viewing pane, thereby causing the second viewing pane to present a table similar to the table of FIG. 13 and the table of FIG. 14, but now enabling a disclosure of a set of identifiers (e.g., a set of names) for a set of files (e.g., a set of documents) that have been uploaded for creation of indexes, as disclosed herein. Note that any of these files can be deleted ornew files added, as enabled via a respective set of virtual buttons presented above the table. Likewise, the table can be exported into a spreadsheet file, as enabled via a respective virtual button presented above the table.

[0114] FIG. 15 shows a screenshot of an embodiment of a GUI of an application program hosted on a computing terminal of FIG. 1 according to this disclosure. In particular, there is a screen 1500 enabling a user profile that is logged into the application program or the set of application programs 104.1 from the application program hosted on the computing terminal 108 to submit the user query, as disclosed herein. Note that the first viewing pane presents a history of prior user queries and the second viewing pane presents a list of templated or dynamically generated user queries for single-click activation above a text box programmed to receive the user query, as disclosed herein.

[0115] Features described with respect to certain embodiments may be combined in or with various other embodiments in any permutational or combinatory manner. Different aspects or elements of example embodiments, as disclosed herein, may be combined in a similar manner. The skilled artisan will understand that typically there is no limit on the number of items or terms in any combination, unless otherwise apparent from the context.

[0116] Hereby, all issued patents, published patent applications, and non-patent publications (including identified articles, web pages, websites, products and manuals thereof) that are mentioned in this disclosure are herein incorporated by reference in their entirety for all purposes, to same extent as if each individual issued patent, published patent application, or non-patent publication were specifically and individually indicated to be incorporated by reference. If any disclosures are incorporated herein by reference and such disclosures conflict in part and / or in whole with the present disclosure, then to the extent of conflict, and / or broader disclosure, and / or broader definition of terms, the present disclosure controls. If such disclosures conflict in part and / or in whole with one another, then to the extent of conflict, the later-dated disclosure controls.

[0117] Various embodiments of the present disclosure may be implemented in a data processing system suitable for storing and / or executing program code that includes at least one processor coupled directly or indirectly to memory elements through a system bus. The memory elements include, for instance, local memory employed during actual execution of the program code, bulk storage, and cache memory which provide temporarystorage of at least some program code in order to reduce the number of times code must be retrieved from bulk storage during execution.

[0118] I / O devices (including, but not limited to, keyboards, displays, pointing devices, DASD, tape, CDs, DVDs, thumb drives and other memory media, etc.) can be coupled to the system either directly or through intervening I / O controllers. Network adapters may also be coupled to the system to enable the data processing system to be-come coupled to other data processing systems or remote printers or storage devices through intervening private or public networks. Modems, cable modems, and Ethernet cards are just a few of the available types of network adapters.

[0119] The present disclosure may be embodied in a system, a method, and / or a computer program product. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure. The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non- exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing.

[0120] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wirelesstransmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0121] Computer readable program instructions for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. A code segment or machine-executable instructions may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, among others. The computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0122] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or step diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each step of the flowchart illustrations and / or step diagrams, and combinations of steps in the flowchart illustrations and / or step diagrams, can be implemented by computer readable program instructions. The various illustrative logical steps, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer soft-ware, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, steps, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0123] The flowchart and step diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each step in the flowchart or step diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the step may occur out of the order noted in the figures. For example, two steps shown in succession may, in fact, be executed substantially concurrently, or the steps may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each step of the step diagrams and / or flowchart illustration, and combinations of steps in the step diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.

[0124] Words such as “then,” “next,” etc. are not intended to limit the order of the steps; these words are simply used to guide the reader through the description of the methods. Although process flow diagrams may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function.

[0125] Features or functionality described with respect to certain example embodiments may be combined and sub-combined in and / or with various other example embodiments. Also, different aspects and / or elements of example embodiments, as disclosed herein, may be combined and sub-combined in a similar manner as well. Further, some example embodiments, whether individually and / or collectively, may be components of a larger system, wherein other procedures may take precedence over and / or otherwise modify their application. Additionally, a number of steps may be required be-fore, after, and / or concurrently with example embodiments, as disclosed herein. Note that any and / or all methods and / or processes, at least as disclosed herein, can be at least partially performed via at least one entity or actor in any manner.

[0126] Although preferred embodiments have been depicted and described in detail herein, skilled artisans know that various modifications, additions, substitutions and the like can be made without departing from spirit of this disclosure. As such, these are considered to be within the scope of the disclosure, as defined in the following claims.

Claims

CLAIMSWhat is claimed is:1 . A system, comprising: a computing instance programmed to: receive a user query from a computing terminal, generate a prompt based on the user query, submit the prompt to a language model, receive a response to the prompt from the language model, and cause the response to be presented on the computing terminal responsive to the user query, wherein at least one of:(a) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a form that is populatable, wherein the first set of data is a set of forms that are populatable, wherein the second subset is a computerized listing of a physical property or a physical land lot, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the user query is augmented with a set of information sourced from a first index and a second index to generate the prompt, wherein the first index indexes the set of forms, wherein the second index indexes the second set of data, wherein the response includes at least one of a copy of the form or a reference to the copy where the copy includes a field that is populated by the language model based on the user query and the set of information;(b) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a legislative text, wherein the first set of data is a set of legislative texts, wherein the second subset is a computerized listing of a physical property or a physical land lot, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the user query is augmented with a set of information sourced from a first index and a second index to generate the prompt, wherein the first index indexes the set of forms, wherein the second index indexes the second set of data, wherein the response includes (i) a yes opinion or a semantic equivalent thereofgenerated by the language model based on the user query and the set of information, (ii) a no opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (iii) a can opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, or (iv) a cant opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information;(c) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a computerized listing of a physical property or a physical land lot, wherein the first set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the second subset is a form or a legislative text, wherein the second set of data is a set of forms that are populatable or a set of legislative texts, wherein the user query is augmented with a set of information sourced from an index of an archive of internal or intraorganizational communications including a set of personal communications to generate the prompt, wherein the index indexes the set of personal communications, wherein the response includes at least one of (i) a phrased content generated by the language model based on the user query and the set of information, (ii) a recommendation generated by the language model based on the user query and the set of information, (iii) a suggestion generated by the language model based on the user query and the set of information, (iv) a listing of possibilities generated by the language model based on the user query and the set of information, or (v) a ranking of possibilities generated by the language model based on the user query and the set of information based on a criterion;(d) wherein the user query requests a recommendation in context of a factor associated with an identifier of a property, wherein the user query is augmented with a set of information sourced from an index of a set of records for a set of identifiers of a set of physical properties or physical land lots to generate the prompt, wherein the set of records contains a set of private comments about the property and a set of parameters of the property, wherein the index indexes the set of private comments about the property and the set of parameters of the property, wherein the response includes the recommendation in context of the factor associated with the identifier of the property; or(e) wherein the prompt is tracked by a semantic topic as the prompt is related to a first set of data in context of a second set of data, wherein the first set of data is a set of forms that are populatable or a set of legislative texts, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the prompt is a first prompt, wherein the first prompt is determined to satisfy a first threshold and a second threshold such that a common semantic thread is identified with a second prompt submitted to the language model and a third prompt is generated based on the common semantic thread to request or identify a recommendation or an opportunity for training a set of users who submitted the first prompt and the second prompt, wherein the second prompt is related to the first set of data in context of the second set of data, wherein the third prompt is submitted to the language model, wherein the response is a first response, wherein the language model outputs a second response to the third prompt such that the second response is presentable on the computing terminal.

2. The system of claim 1 , wherein at least two of:(a) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a form that is populatable, wherein the first set of data is a set of forms that are populatable, wherein the second subset is a computerized listing of a physical property or a physical land lot, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the user query is augmented with a set of information sourced from a first index and a second index to generate the prompt, wherein the first index indexes the set of forms, wherein the second index indexes the second set of data, wherein the response includes at least one of a copy of the form or a reference to the copy where the copy includes a field that is populated by the language model based on the user query and the set of information;(b) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a legislative text, wherein the first set of data is a set of legislative texts, wherein the second subset is a computerized listing of a physical property or a physical land lot, wherein thesecond set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the user query is augmented with a set of information sourced from a first index and a second index to generate the prompt, wherein the first index indexes the set of forms, wherein the second index indexes the second set of data, wherein the response includes (i) a yes opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (ii) a no opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (iii) a can opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, or (iv) a cant opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information;(c) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a computerized listing of a physical property or a physical land lot, wherein the first set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the second subset is a form or a legislative text, wherein the second set of data is a set of forms that are populatable or a set of legislative texts, wherein the user query is augmented with a set of information sourced from an index of an archive of internal or intraorganizational communications including a set of personal communications to generate the prompt, wherein the index indexes the set of personal communications, wherein the response includes at least one of (i) a phrased content generated by the language model based on the user query and the set of information, (ii) a recommendation generated by the language model based on the user query and the set of information, (iii) a suggestion generated by the language model based on the user query and the set of information, (iv) a listing of possibilities generated by the language model based on the user query and the set of information, or (v) a ranking of possibilities generated by the language model based on the user query and the set of information based on a criterion;(d) wherein the user query requests a recommendation in context of a factor associated with an identifier of a property, wherein the user query is augmented with a set of information sourced from an index of a set of records for a set of identifiers of a set of physical properties or physical land lots to generate the prompt, wherein theset of records contains a set of private comments about the property and a set of parameters of the property, wherein the index indexes the set of private comments about the property and the set of parameters of the property, wherein the response includes the recommendation in context of the factor associated with the identifier of the property; or(e) wherein the prompt is tracked by a semantic topic as the prompt is related to a first set of data in context of a second set of data, wherein the first set of data is a set of forms that are populatable or a set of legislative texts, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the prompt is a first prompt, wherein the first prompt is determined to satisfy a first threshold and a second threshold such that a common semantic thread is identified with a second prompt submitted to the language model and a third prompt is generated based on the common semantic thread to request or identify a recommendation or an opportunity for training a set of users who submitted the first prompt and the second prompt, wherein the second prompt is related to the first set of data in context of the second set of data, wherein the third prompt is submitted to the language model, wherein the response is a first response, wherein the language model outputs a second response to the third prompt such that the second response is presentable on the computing terminal.

3. The system of claim 1 , wherein at least three of:(a) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a form that is populatable, wherein the first set of data is a set of forms that are populatable, wherein the second subset is a computerized listing of a physical property or a physical land lot, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the user query is augmented with a set of information sourced from a first index and a second index to generate the prompt, wherein the first index indexes the set of forms, wherein the second index indexes the second set of data, wherein the response includes at least one of a copy of the form or a reference to the copy where the copy includes a field that is populated by the language model based on the user query and the set of information;(b) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a legislative text, wherein the first set of data is a set of legislative texts, wherein the second subset is a computerized listing of a physical property or a physical land lot, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the user query is augmented with a set of information sourced from a first index and a second index to generate the prompt, wherein the first index indexes the set of forms, wherein the second index indexes the second set of data, wherein the response includes (i) a yes opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (ii) a no opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (iii) a can opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, or (iv) a cant opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information;(c) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a computerized listing of a physical property or a physical land lot, wherein the first set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the second subset is a form or a legislative text, wherein the second set of data is a set of forms that are populatable or a set of legislative texts, wherein the user query is augmented with a set of information sourced from an index of an archive of internal or intraorganizational communications including a set of personal communications to generate the prompt, wherein the index indexes the set of personal communications, wherein the response includes at least one of (i) a phrased content generated by the language model based on the user query and the set of information, (ii) a recommendation generated by the language model based on the user query and the set of information, (iii) a suggestion generated by the language model based on the user query and the set of information, (iv) a listing of possibilities generated by the language model based on the user query and the set of information, or (v) a ranking of possibilities generated by the language model based on the user query and the set of information based on a criterion;(d) wherein the user query requests a recommendation in context of a factor associated with an identifier of a property, wherein the user query is augmented with a set of information sourced from an index of a set of records for a set of identifiers of a set of physical properties or physical land lots to generate the prompt, wherein the set of records contains a set of private comments about the property and a set of parameters of the property, wherein the index indexes the set of private comments about the property and the set of parameters of the property, wherein the response includes the recommendation in context of the factor associated with the identifier of the property; or(e) wherein the prompt is tracked by a semantic topic as the prompt is related to a first set of data in context of a second set of data, wherein the first set of data is a set of forms that are populatable or a set of legislative texts, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the prompt is a first prompt, wherein the first prompt is determined to satisfy a first threshold and a second threshold such that a common semantic thread is identified with a second prompt submitted to the language model and a third prompt is generated based on the common semantic thread to request or identify a recommendation or an opportunity for training a set of users who submitted the first prompt and the second prompt, wherein the second prompt is related to the first set of data in context of the second set of data, wherein the third prompt is submitted to the language model, wherein the response is a first response, wherein the language model outputs a second response to the third prompt such that the second response is presentable on the computing terminal.

4. The system of claim 1 , wherein at least four of:(a) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a form that is populatable, wherein the first set of data is a set of forms that are populatable, wherein the second subset is a computerized listing of a physical property or a physical land lot, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the user query is augmented with a set of information sourced from a first index and a second index to generate the prompt, whereinthe first index indexes the set of forms, wherein the second index indexes the second set of data, wherein the response includes at least one of a copy of the form or a reference to the copy where the copy includes a field that is populated by the language model based on the user query and the set of information;(b) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a legislative text, wherein the first set of data is a set of legislative texts, wherein the second subset is a computerized listing of a physical property or a physical land lot, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the user query is augmented with a set of information sourced from a first index and a second index to generate the prompt, wherein the first index indexes the set of forms, wherein the second index indexes the second set of data, wherein the response includes (i) a yes opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (ii) a no opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (iii) a can opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, or (iv) a cant opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information;(c) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a computerized listing of a physical property or a physical land lot, wherein the first set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the second subset is a form or a legislative text, wherein the second set of data is a set of forms that are populatable or a set of legislative texts, wherein the user query is augmented with a set of information sourced from an index of an archive of internal or intraorganizational communications including a set of personal communications to generate the prompt, wherein the index indexes the set of personal communications, wherein the response includes at least one of (i) a phrased content generated by the language model based on the user query and the set of information, (ii) a recommendation generated by the language model based on the user query and the set of information, (iii)a suggestion generated by the language model based on the user query and the set of information, (iv) a listing of possibilities generated by the language model based on the user query and the set of information, or (v) a ranking of possibilities generated by the language model based on the user query and the set of information based on a criterion;(d) wherein the user query requests a recommendation in context of a factor associated with an identifier of a property, wherein the user query is augmented with a set of information sourced from an index of a set of records for a set of identifiers of a set of physical properties or physical land lots to generate the prompt, wherein the set of records contains a set of private comments about the property and a set of parameters of the property, wherein the index indexes the set of private comments about the property and the set of parameters of the property, wherein the response includes the recommendation in context of the factor associated with the identifier of the property; or(e) wherein the prompt is tracked by a semantic topic as the prompt is related to a first set of data in context of a second set of data, wherein the first set of data is a set of forms that are populatable or a set of legislative texts, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the prompt is a first prompt, wherein the first prompt is determined to satisfy a first threshold and a second threshold such that a common semantic thread is identified with a second prompt submitted to the language model and a third prompt is generated based on the common semantic thread to request or identify a recommendation or an opportunity for training a set of users who submitted the first prompt and the second prompt, wherein the second prompt is related to the first set of data in context of the second set of data, wherein the third prompt is submitted to the language model, wherein the response is a first response, wherein the language model outputs a second response to the third prompt such that the second response is presentable on the computing terminal.

5. The system of claim 1 , wherein:(a) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a form that is populatable, wherein the first set of data is a set of forms that are populatable,wherein the second subset is a computerized listing of a physical property or a physical land lot, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the user query is augmented with a set of information sourced from a first index and a second index to generate the prompt, wherein the first index indexes the set of forms, wherein the second index indexes the second set of data, wherein the response includes at least one of a copy of the form or a reference to the copy where the copy includes a field that is populated by the language model based on the user query and the set of information;(b) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a legislative text, wherein the first set of data is a set of legislative texts, wherein the second subset is a computerized listing of a physical property or a physical land lot, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the user query is augmented with a set of information sourced from a first index and a second index to generate the prompt, wherein the first index indexes the set of forms, wherein the second index indexes the second set of data, wherein the response includes (i) a yes opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (ii) a no opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (iii) a can opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, or (iv) a cant opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information;(c) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a computerized listing of a physical property or a physical land lot, wherein the first set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the second subset is a form or a legislative text, wherein the second set of data is a set of forms that are populatable or a set of legislative texts, wherein the user query is augmented with a set of information sourced from an index of an archive of internal or intraorganizational communications including a set of personal communications togenerate the prompt, wherein the index indexes the set of personal communications, wherein the response includes at least one of (i) a phrased content generated by the language model based on the user query and the set of information, (ii) a recommendation generated by the language model based on the user query and the set of information, (iii) a suggestion generated by the language model based on the user query and the set of information, (iv) a listing of possibilities generated by the language model based on the user query and the set of information, or (v) a ranking of possibilities generated by the language model based on the user query and the set of information based on a criterion;(d) wherein the user query requests a recommendation in context of a factor associated with an identifier of a property, wherein the user query is augmented with a set of information sourced from an index of a set of records for a set of identifiers of a set of physical properties or physical land lots to generate the prompt, wherein the set of records contains a set of private comments about the property and a set of parameters of the property, wherein the index indexes the set of private comments about the property and the set of parameters of the property, wherein the response includes the recommendation in context of the factor associated with the identifier of the property; and(e) wherein the prompt is tracked by a semantic topic as the prompt is related to a first set of data in context of a second set of data, wherein the first set of data is a set of forms that are populatable or a set of legislative texts, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the prompt is a first prompt, wherein the first prompt is determined to satisfy a first threshold and a second threshold such that a common semantic thread is identified with a second prompt submitted to the language model and a third prompt is generated based on the common semantic thread to request or identify a recommendation or an opportunity for training a set of users who submitted the first prompt and the second prompt, wherein the second prompt is related to the first set of data in context of the second set of data, wherein the third prompt is submitted to the language model, wherein the response is a first response, wherein the language model outputs a second response to the third prompt such that the second response is presentable on the computing terminal.

6. The system of claim 1 , wherein (a) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a form that is populatable, wherein the first set of data is a set of forms that are populatable, wherein the second subset is a computerized listing of a physical property or a physical land lot, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the user query is augmented with a set of information sourced from a first index and a second index to generate the prompt, wherein the first index indexes the set of forms, wherein the second index indexes the second set of data, wherein the response includes at least one of a copy of the form or a reference to the copy where the copy includes a field that is populated by the language model based on the user query and the set of information.

7. The system of claim 6, wherein the response includes the copy.

8. The system of claim 6, wherein the response includes the reference.

9. The system of claim 6, wherein the response includes the copy and the reference.

10. The system of claim 6, wherein the field is empty and populated by the language model by inserting a content into the field.11 . The system of claim 6, wherein the field is populated with a first content and populated by the language model by editing the first content to a second content.

12. The system of claim 6, wherein the form is a Portable Document Format (PDF) file.

13. The system of claim 1 , wherein (b) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a legislative text, wherein the first set of data is a set of legislative texts, wherein the second subset is a computerized listing of a physical property or a physical land lot, wherein the second set of data is a set of computerized listings of a set of physicalproperties or physical land lots, wherein the user query is augmented with a set of information sourced from a first index and a second index to generate the prompt, wherein the first index indexes the set of forms, wherein the second index indexes the second set of data, wherein the response includes at least one of (i) a yes opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (ii) a no opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (iii) a can opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, or (iv) a cant opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information.

13. The system of claim 12, wherein the legislative text is hierarchical.

14. The system of claim 12, wherein the response includes (i) a yes opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information.

15. The system of claim 12, wherein the response includes (ii) a no opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information.

16. The system of claim 12, wherein the response includes (iii) a can opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information.

17. The system of claim 12, wherein the response includes (iv) a cant opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information.

18. The system of claim 12, wherein the response includes at least two of (i) a yes opinion or a semantic equivalent thereof generated by the language model based on the userquery and the set of information, (ii) a no opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (iii) a can opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, or (iv) a cant opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information.

19. The system of claim 12, wherein the response includes at least three of (i) a yes opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (ii) a no opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (iii) a can opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, or (iv) a cant opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information.

20. The system of claim 12, wherein the response includes (i) a yes opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (ii) a no opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, (iii) a can opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information, and (iv) a cant opinion or a semantic equivalent thereof generated by the language model based on the user query and the set of information.21 . The system of claim 1 , wherein (c) wherein the user query is about a first subset of a first set of data in context of a second subset of a second set of data, wherein the first subset is a computerized listing of a physical property or a physical land lot, wherein the first set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the second subset is a form or a legislative text, wherein the second set of data is a set of forms that are populatable or a set of legislative texts, wherein the user query is augmented with a set of information sourced from an index of an archive ofinternal or intraorganizational communications including a set of personal communications to generate the prompt, wherein the index indexes the set of personal communications, wherein the response includes at least one of (i) a phrased content generated by the language model based on the user query and the set of information, (ii) a recommendation generated by the language model based on the user query and the set of information, (iii) a suggestion generated by the language model based on the user query and the set of information, (iv) a listing of possibilities generated by the language model based on the user query and the set of information, or (v) a ranking of possibilities generated by the language model based on the user query and the set of information based on a criterion;19. The system of claim 18, wherein the form is a Portable Document Format (PDF) file.

20. The system of claim 18, wherein the legislative text is hierarchical.21 . The system of claim 18, wherein the set of personal communications is a set of email messages.

22. The system of claim 18, wherein the set of personal communications is a set of text messages.

23. The system of claim 18, wherein the set of personal communications is a set of chat messages.

24. The system of claim 18, wherein the set of personal communications is a set of Over- The-Top (OTT) messages.

25. The system of claim 18, wherein the set of personal communications includes at least two of an email message, a text message, a chat message, or an Over-The-Top (OTT) message.

26. The system of claim 18, wherein the set of personal communications includes at least three of an email message, a text message, a chat message, or an Over-The-Top (OTT) message.

27. The system of claim 18, wherein the set of personal communications includes an email message, a text message, a chat message, and an Over-The-Top (OTT) message.

28. The system of claim 18, wherein the response includes at least one of (i) a phrased content generated by the language model based on the user query and the set of information, (ii) a recommendation generated by the language model based on the user query and the set of information, (iii) a suggestion generated by the language model based on the user query and the set of information, (iv) a listing of possibilities generated by the language model based on the user query and the set of information, or (v) a ranking of possibilities generated by the language model based on the user query and the set of information based on a criterion29. The system of claim 18, wherein the response includes at least two of (i) a phrased content generated by the language model based on the user query and the set of information, (ii) a recommendation generated by the language model based on the user query and the set of information, (iii) a suggestion generated by the language model based on the user query and the set of information, (iv) a listing of possibilities generated by the language model based on the user query and the set of information, or (v) a ranking of possibilities generated by the language model based on the user query and the set of information based on a criterion.

30. The system of claim 18, wherein the response includes at least three of (i) a phrased content generated by the language model based on the user query and the set of information, (ii) a recommendation generated by the language model based on the user query and the set of information, (iii) a suggestion generated by the language model based on the user query and the set of information, (iv) a listing of possibilities generated by the language model based on the user query and the set of information, or (v) a rankingof possibilities generated by the language model based on the user query and the set of information based on a criterion.

31. The system of claim 18, wherein the response includes at least four of (i) a phrased content generated by the language model based on the user query and the set of information, (ii) a recommendation generated by the language model based on the user query and the set of information, (iii) a suggestion generated by the language model based on the user query and the set of information, (iv) a listing of possibilities generated by the language model based on the user query and the set of information, or (v) a ranking of possibilities generated by the language model based on the user query and the set of information based on a criterion.

32. The system of claim 18, wherein the response includes (i) a phrased content generated by the language model based on the user query and the set of information, (ii) a recommendation generated by the language model based on the user query and the set of information, (iii) a suggestion generated by the language model based on the user query and the set of information, (iv) a listing of possibilities generated by the language model based on the user query and the set of information, and (v) a ranking of possibilities generated by the language model based on the user query and the set of information based on a criterion.

33. The system of claim 1 , wherein (d) wherein the user query requests a recommendation in context of a factor associated with an identifier of a property, wherein the user query is augmented with a set of information sourced from an index of a set of records for a set of identifiers of a set of physical properties or physical land lots to generate the prompt, wherein the set of records contains a set of private comments about the property and a set of parameters of the property, wherein the index indexes the set of private comments about the property and the set of parameters of the property, wherein the response includes the recommendation in context of the factor associated with the identifier of the property.

34. The system of claim 33, wherein the set of private comments is stored in a set of designated fields within a data source that are flagged as restricted and are programmed to display only as permissioned while being excluded from public-facing reports or views.

35. The system of claim 1 , wherein (e) wherein the prompt is tracked by a semantic topic as the prompt is related to a first set of data in context of a second set of data, wherein the first set of data is a set of forms that are populatable or a set of legislative texts, wherein the second set of data is a set of computerized listings of a set of physical properties or physical land lots, wherein the prompt is a first prompt, wherein the first prompt is determined to satisfy a first threshold and a second threshold such that a common semantic thread is identified with a second prompt submitted to the language model and a third prompt is generated based on the common semantic thread to request or identify a recommendation or an opportunity for training a set of users who submitted the first prompt and the second prompt, wherein the second prompt is related to the first set of data in context of the second set of data, wherein the third prompt is submitted to the language model, wherein the response is a first response, wherein the language model outputs a second response to the third prompt such that the second response is presentable on the computing terminal.

36. The system of claim 35, wherein at least one form of the set of forms is a Portable Document Format (PDF) file.

37. The system of claim 35, wherein at least one legislative text of the set of legislative texts is hierarchical.

38. The system of claim 35, wherein the semantic topic involves a form or a field thereof.

39. The system of claim 35, wherein the semantic topic involves a law text or a section thereof.

40. The system of claim 35, wherein the semantic topic involves an identifier of a physical land lot or a physical property thereon, or an attribute or a description of any of foregoing.

41. The system of claim 35, wherein the first threshold is a numeric threshold setting a minimum or floor amount of prompts submitted to the language model.

42. The system of claim 35, wherein the second threshold is a semantic threshold indicating a semantic relationship between the first prompt and the second prompt.

43. The system of claim 42, wherein the semantic relationship is semantically similar when the first prompt and the second prompt are prompting about a similar topic or a theme.

44. The system of claim 35, wherein the second prompt is submitted to the language model before the first prompt.

45. The system of claim 35, wherein the common semantic thread is a topic, a theme, a concept, or an inquiry.

46. The system of claim 35, wherein the prompt is generated based on gathering a set of metadata for the set of users and considering the set of metadata in generating the prompt such that the prompt is at least partially generated based on the set of metadata.

Citation Information

Patent Citations

  • System and method of automated real estate analysis

    US11182865B1

  • Text reduction and analysis interface to a text generation modeling system

    US11861320B1

  • Artificial intelligence communication assistance

    US20230325590A1

  • Risk assessment management system and method

    US20230342798A1