Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

8010 results about "Information retrieval" patented technology

Information retrieval (IR) is the activity of obtaining information system resources that are relevant to an information need from a collection of those resources. Searches can be based on full-text or other content-based indexing. Information retrieval is the science of searching for information in a document, searching for documents themselves, and also searching for the metadata that describes data, and for databases of texts, images or sounds.

NL2SQL optimization method and device based on large model, equipment and medium

The invention discloses an NL2SQL optimization method and device based on a large model, equipment and a medium, and relates to the technical field of artificial intelligence, the method comprises the following steps: constructing a target metadata knowledge base, and obtaining an initial natural language query request; determining each target entity corresponding to the initial natural language query request, and determining missing target SQL elements in the initial natural language query request based on each target entity; generating a first cue word based on the initial natural language query request, the target SQL element and the target metadata knowledge base, and complementing the target SQL element based on the first cue word by utilizing the target large model to obtain a target natural language query request; and generating a plurality of candidate SQL statements corresponding to the target natural language query request by using the target large model, verifying each candidate SQL statement, and determining a target SQL statement from each candidate SQL statement based on a verification result. According to the method, the accuracy of the NL2SQL can be improved by utilizing a large model.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Construction of a user questions repository for structured data queries from natural language questions using a large language model

A data query system and methods are provided that are configured to intelligently generate structured data queries from natural language questions using large language models (LLMs). The system includes a processor and a computer readable medium operably coupled thereto, the computer readable medium comprising a plurality of instructions stored in association therewith that are accessible to, and executable by, the processor, to perform operations which include receiving a natural language question for structured data, converting the natural language question to embeddings, matching the embeddings to pre-generated questions from a user questions repository (UQR), determining an accuracy of the matching meets or exceeds a threshold similarity, determining, using an LLM and metadata corresponding to the pre-generated questions from the UQR, a structured data query for querying for the structured data, and querying the structured database system using the structured data query.
Owner:NICE LTD

Personalized retrieval-augmented generation system

The present disclosure relates to systems, non-transitory computer-readable media, and methods for generating personal responses through retrieval-augmented generation. In particular, the disclosed systems can generate a query embedding from a query generated by an entity and determine data context specific to the entity by comparing the query embedding with a plurality of vectorized segments of content items associated with the entity. The disclosed systems can provide the data context to a large language model and generate a personalized response informed by the data context. Subsequently, the disclosed systems can provide the personalized response for display on a client device associated with the entity.
Owner:DROPBOX INC

AI-driven bet search

Systems, methods, and computer-readable media for presenting one or more bets to a user based on a natural language query received from the user. In some embodiments, a natural language query from a user may be received. The natural language query may be translated into computer-readable data by a language processing engine. The language processing engine may use a large language model to translate the natural language query into computer-readable data. The computer-readable data may be in the form of an embedding. A bet engine may match the computer-readable data to a bet when the bet and the computer-readable data exceed a predetermined threshold with regard to similarity. The bet engine may generate a new bet corresponding to the computer-readable data. The system may then present the bet to the user.
Owner:FANDUEL LTD

Systems and methods for deployment of contextual memory management system for generating contextual data for langauge model prompts

In one implementation, a computer-implemented method involves receiving a user message corresponding to a query or a statement to AI chatbot, performing preprocessing operations resulting in generation of initial context of the user message by extracting text of the user message, metadata of the user message, and a conversation identifier, obtaining historical context pertaining to the user message from a plurality of storage mechanisms provided in differing formats including a knowledge graph, a vector database comprised of vector embeddings, and a database comprising text summaries of prior conversations between the user and the AI chatbot, generating a prompt for a LLM that instructs the LLM to generate a response to the user message that is based on and consistent with the user message, the initial content, and the historical context, and providing a final response to the user that is corresponds to an LLM-generated response.
Owner:BAYARDELLE ELIZABETH

Generating a schema graph of sub-tables in a database for queries using a large language model

The present disclosure relates to systems, non-transitory computer-readable media, and methods for linking a database schema to a natural language query. In particular, in some embodiments, the disclosed systems determine, from tables in a database schema, a subset of tables relevant to a natural language query by comparing embeddings for the tables in the database schema and embeddings for the natural language query. Additionally, in some implementations, the disclosed systems select, from a schema graph comprising nodes that represent the tables in the database schema, an additional table along a path between a pair of nodes representing a pair of tables from the subset of tables. Moreover, in some embodiments, the disclosed systems determine a set of relevant tables by appending the additional table to the subset of tables. Furthermore, in some implementations, the disclosed systems generate, from the set of relevant tables, a response for the natural language query.
Owner:ADOBE INC

Long document topic summarization using large language models

A large language model predicts a topic summarization of a long document given a set of segments from the long document that pertain to a topic of interest. The set of segments from the long document are selected by searching for similar segments from other documents that have been labeled to indicate whether or not the segment pertains to a topic of interest. The search is based on an embedding of a segment from the long document closely matching embeddings of the labeled segments. Each segment of the long document is scored based on the labels of the closest-matching similar segments. The segments from the long document are ranked by their respective score and the highest-scored segments are included in a prompt to the large language model for the model to generate a topic summarization of the long document.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Graph and vector usage for automated QA system

A method for conversational query resolution includes receiving a query input from a user. The query input is decomposed into a plurality of tasks. A knowledge graph is queried to identify one or more relevant entities based on at least one of the plurality of tasks. A vector database is searched to identify one or more text chunks that correspond to the one or more relevant entities. Content relevant to at least one of the plurality of tasks is identified from the one or more text chunks. An answer is generated based on the identified content.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Retrieval-augmented generation processing using dynamically selected number of document chunks

An apparatus comprises at least one processing device configured to obtain a query comprising search text and a context identifying documents to be searched using the search text, to generate document chunks by parsing the documents, to determine a degree of specificity of the search text, and to determine a number of the document chunks to select for retrieval-augmented generation processing based on the determined degree of specificity. The at least one processing device is also configured to select a subset of the document chunks based on similarity between the document chunks and the search text, the subset including the determined number of document chunks. The at least one processing device is further configured to generate and apply a prompt including the selected subset of the document chunks to a machine learning system to generate an output, and to provide an answer to the query based on the output.
Owner:DELL PROD LP

Knowledge base question-answering system utilizing a large language model output in re-ranking

Aspects of the subject disclosure may include systems and methods, for example, including receiving a user input in natural language, retrieving a first answer including a list of a first number of documents relevant to the user input by searching indexed documents in a knowledge base, applying the first answer to a large language model reader, resulting in a second answer, re-ranking the first answer by using the second answer, resulting in a third answer including a re-ranked list of the first number of documents, and generating a final response including the second answer and one or more documents among the third answer.
Owner:AT&T INTELLECTUAL PROPERTY I L P

Systems and methods for generating a workflow data structure

Systems and methods for generating a workflow data structure are provided. The system includes one or more processors; and one or more transitory or non-transitory computer-readable media storing instructions that are executable to cause the one or more processors to perform operations, the operations comprising: receiving input data comprising a corpus of documents, a user query, and query context data; processing the corpus of documents to generate training data; training a large language model (LLM) using the training data; classifying, using the LLM, the user query to at least one content cluster of a plurality of content clusters based on the query context data; constructing, using the LLM, a workflow data structure as a function of the classifying; and generating, using the LLM, a query response as a function of the user query, the query context data, and the workflow data structure.
Owner:A&E ENGINEERING INC

Unstructured data extraction with large language models for query resolution

An unstructured data query-response pair generation system (generation system) populates a knowledge base of query-response pairs for queries of natural language content in unstructured data by prompting a first large language model (LLM) text extracted from the unstructured data. An unstructured data chatbot (chatbot) leverages the knowledge base by augmenting prompts to a second LLM responding to user queries for natural language content in the unstructured data with query-response pairs having queries that are semantically similar to the user queries. The knowledge base and LLMs are updated based on user feedback correcting responses, continually improving quality of the generation system and chatbot.
Owner:PALO ALTO NETWORKS INC

Real-time normalization of raw enterprise data from disparate sources

Various embodiments relate to normalizing raw data by mapping the raw data to a computer-readable tag. A computer-readable tag may be an identifier that at least partially represents a category (e.g., a department) and / or the raw data itself. In response to receiving the raw data, some embodiments perform the mapping by, for example, performing natural language processing (NLP) on each particular department's raw data to associate natural language words in the raw data to its corresponding computer-readable tag and then populating, at a data structure that includes the computer-readable tag, an entry with data (representing the raw data) in a standardized format. In this way, regardless of whether different sets of raw data come from disparate sources that have diverse formats, protocols, or structures relative to each other, the normalized data and standardized form makes the data compatible.
Owner:ACTABL

Generative ai-driven multi-source data query system

Embodiments of the disclosed technologies include, in response to receiving a query, matching the query to metadata from a plurality of heterogeneous data sources, and selecting one or more data sources from the plurality of heterogeneous data sources for answering the query, by sending the query and embeddings of the matched metadata to a generative artificial intelligence (GAI), and prompting the GAI to select matching data sources. Based on the data from the GAI, generating one or more custom queries targeted to the matching data sources selected by the GAI, the custom queries formatted to be sent to the selected data sources, executing the one or more custom queries across the selected data sources, and summarizing results from the executing and providing a response to the query.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Retrieval augmented generative question and answer boosting

Systems or techniques are provided for facilitating retrieval augmented generative question and answer boosting. In various embodiments, a system can access a plain text question regarding a scientific instrument. In various aspects, the system can generate, via a large language model that references a document-graph repository, a structured or unstructured answer for the plain text question. In various instances, the document-graph repository can comprise a plurality of document-graphs that respectively correspond to a plurality of technical documents. In various cases, for a first document-graph that corresponds to a first technical document, leaf nodes of the first document-graph can represent respective text blocks written in the first technical document, and non-leaf nodes of the first document-graph can respectively represent a document title, one or more section headings, and one or more scientific instrument identifiers written in the first technical document and beneath which the respective text blocks are nested.
Owner:PPD DEVELOPMENT LP +2

Document analysis and management systems and methods

Example document analysis and management systems and methods are described. In one implementation, a document is identified for processing. An artificial intelligence engine extracts information from the document and creates multiple chunks of data associated with the document. Embeddings are performed for the multiple chunks of data to create chunk embeddings, where the chunk embeddings are represented as numerical vectors. The chunk embeddings are stored in a vector database. A large language model (LLM) generates document content insights based on the multiple chunks of data and the chunk embeddings.
Owner:SIMPLEO AI

Question-answering processing method, and device, product and storage medium

Provided in the embodiments of the present disclosure are a question-answering processing method, and a device, a product and a storage medium. In the question-answering processing method, after a query instruction is acquired, target knowledge information that matches the query instruction can be acquired from among a plurality of pieces of knowledge information in a knowledge base, and the query instruction and content-parsed text that corresponds to the target knowledge information are input into a large language model for question-answering processing, wherein the content-parsed text that corresponds to the target knowledge information is obtained by means of performing content parsing on a target document element that corresponds to the target knowledge information, and when the target document element comprises a document element of a non-text modality, content parsing is performed on the target document element before the target document element is input into the large language model, such that the document element of the non-text modality in the target document element can be understood by the large language model, so as to provide question-answering reference knowledge with a relatively high reliability for the large language model. Therefore, the accuracy of answering of the large language model for the query instruction can be improved.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Large language model interactions via intelligent prompt enrichment module and updated profile

Various embodiments of the technology described programmatically access a user query intended for a Large Language Model (LLM), analyze the user query, and determine prompt-enriching information that is combined with the user query to generate an enriched user query that is ultimately communicated to the LLM. In this manner, additional prompt-enriching information or context is added to the user query before being communicated to the LLM so that the additional prompt-enriching information, along with the user query, can be tokenized to better guide the LLM to a more accurate answer without modifying weights, parameters, or training of the LLM. Certain embodiments have the technical effect of improved accuracy relative to existing approaches by enriching user queries with prompt-enriching information to generate an enriched user query that is passed to the LLM. Based on the enriched user query, certain embodiments reduce the likelihood of hallucinations present in the LLM response.
Owner:RIOT GAMES INC

Metadata determination and storage method

Methods, systems, and techniques for metadata determination and storage. A large language model that is implemented using at least one artificial neural network receives an initial prompt that includes a query related to the metadata. The metadata is in respect of data that is part of a dataset, and the initial prompt includes context for the query. The large language model determines the metadata in response to the query using the context. Once determined, the metadata is stored in the dataset such that the metadata is associated with the data to which it relates.
Owner:ROYAL BANK OF CANADA

Pre-computation for intermediate-representation infused search and retrieval augmented generation

Provided is a process including: obtaining, with a computer system, access to a code base; decomposing, with the computer system, the code base into parts; generating, with the computer system, documentation for the parts with a language model; associating, with the computer system, the documentation with the parts; indexing, with the computer system, the documentation; obtaining, with the computer system, a query searching for content in the code base; searching, with the computer system, using the index, the code base based on the generated documentation to identify documentation corresponding to the query and, then, content in the code base associated with the identified documentation; and responding, with the computer system, to the query, by identifying the content in the code base associated with the identified documentation.
Owner:DRIVER AI INC

Natural language query generation for feature stores using zero shot learning

The present disclosure pertains to natural language techniques for querying data stored in feature stores using zero shot learning. In a particular aspect, a computer-implemented method includes receiving a natural language query for retrieving features from a feature store, generating an input prompt by appending a script to the natural language query, and then using a large language model to determine tables or databases from the feature store that are relevant to the natural language query, retrieve metadata for the tables or databases from the feature store, determine feature groups comprising features relevant to the natural language query, and generate a programming language query based on the input prompt, the metadata, and the groups. A list of features within the feature groups that are accessible within the feature store may then be retrieved by executing the programming language query on the feature store.
Owner:ORACLE INT CORP

Recommendation generation using user input

Methods, systems, and apparatuses include receiving text input via a user interface for an online system. An embedding is generated based on the text input. Supplemental text is generated using the embedding and a vector store including a standardized content items, the supplemental text having a standardized format. The standardized content items are generated by applying a large language model to a plurality of content items. A prompt is formulated including the supplemental text. A generative language model is applied to the prompt. A recommendation is output by the generative language model based on the prompt. The recommendation is provided to the user interface based on at least the text input.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Intelligent database page prefetching for faster query processing

Techniques for intelligent prefetching of database pages for improved query processing are described. Page access patterns for queries are collected over time and used to update a graph data structure to include collections of nodes corresponding to database pages that are often and / or recently accessed together. When processing a query, the graph data structure can be used to identify scenarios when prefetching may be helpful and to identify particular pages that should be prefetched. These pages are prefetched and placed in a local cache storage that can be rapidly accessed by query processors in comparison to the original storage location for this data, which may be across a network.
Owner:AMAZON TECH INC

Video Query Contextualization

Systems and methods for video query contextualization can include a router model that determines how to process and respond to the query associated with the video. The systems and methods can include obtaining an input query and video data, processing the input query and the video data with the router model to generate a video clip and routing data, and the routing data can then be utilized to determine which processing system to utilize to process the video clip and the input query. The video clip can then be processed with the determined processing system to generate a query response that may be provided to the user.
Owner:GOOGLE LLC

Systems and methods for automatically generating source code

A computer-implemented method is disclosed. The method includes: receiving a request for retrieval of data satisfying one or more criteria, the request including at least one data request parameter; searching a database storing example queries based on the request to identify at least one matching query; providing, to a large language model (LLM), an input prompt to generate a query purporting to retrieve data satisfying the one or more criteria, the input prompt including the at least one data request parameter and the at least one matching query as an example; and receiving, from the LLM, a result including the generated query.
Owner:SHOPIFY INC

Logical text passage generation and retrieval for retrieval-augmented generation

Techniques for logical text passage generation and retrieval for retrieval-augmented generation. The techniques involve processing markup language documents to generate logical text passages and their corresponding embeddings. These embeddings are indexed for efficient retrieval. Upon receiving a user utterance, a user query is formed and transformed into an embedding to query the index. Relevant text passages are identified and used to prompt a large language model (LLM), which generates a completion. This completion is then sent as a response to the user. The process effectively bridges user queries with relevant information through advanced embedding and natural language processing techniques, enabling accurate and contextually appropriate interactions within a user-agent dialogue framework.
Owner:AMAZON TECH INC

Music segment tagging, sharing, and image generation

A method of automated generation of contextually-relevant images for a music segment includes receiving at least one of basic metadata information and lyric information for the music segment, generating a first prompt for a computer-implemented machine-learning language model based on the at least one of the basic metadata information and the lyric information, receiving context information from the computer-implemented machine-learning language model in response to the first prompt, generating a second prompt for the computer-implemented machine-learning language model based on the context information, generating a third prompt by providing the second prompt as an input to the computer-implemented machine-learning language model, and generating an image descriptive of the music segment by providing the third prompt as an input to a computer-implemented machine-learning image generation model.
Owner:HOOK MEDIA LLC

Method for Classifying and Controlling Transmission of a File

A method and system for classifying a video file within an environment in which the file is located and when a file is classified as sensitive, controlling transmission of the file outside the environment. Classifying the video comprises analysing, using at least one machine learning model, the video to recognise any individuals in the video; obtaining a transcript of any speech in the video and generating, using the analysis, obtained transcript and a database of individuals linked to the environment, a labelled transcript which identifies each individual linked to the environment that is in the video. Information about each identified individual may be obtained from a connected database. A first generative AI model generates a text-based summary of the video by using the labelled transcript and information about identified individuals as prompts. A second generative AI model then determines a sensitivity classification of the video using the generated text-based summary.
Owner:VARONIS SYSTEMS INC