Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

5482 results about "Information retrieval" patented technology

Information retrieval (IR) is the activity of obtaining information system resources that are relevant to an information need from a collection of those resources. Searches can be based on full-text or other content-based indexing. Information retrieval is the science of searching for information in a document, searching for documents themselves, and also searching for the metadata that describes data, and for databases of texts, images or sounds.

NL2SQL optimization method and device based on large model, equipment and medium

The invention discloses an NL2SQL optimization method and device based on a large model, equipment and a medium, and relates to the technical field of artificial intelligence, the method comprises the following steps: constructing a target metadata knowledge base, and obtaining an initial natural language query request; determining each target entity corresponding to the initial natural language query request, and determining missing target SQL elements in the initial natural language query request based on each target entity; generating a first cue word based on the initial natural language query request, the target SQL element and the target metadata knowledge base, and complementing the target SQL element based on the first cue word by utilizing the target large model to obtain a target natural language query request; and generating a plurality of candidate SQL statements corresponding to the target natural language query request by using the target large model, verifying each candidate SQL statement, and determining a target SQL statement from each candidate SQL statement based on a verification result. According to the method, the accuracy of the NL2SQL can be improved by utilizing a large model.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Systems and methods for deployment of contextual memory management system for generating contextual data for langauge model prompts

In one implementation, a computer-implemented method involves receiving a user message corresponding to a query or a statement to AI chatbot, performing preprocessing operations resulting in generation of initial context of the user message by extracting text of the user message, metadata of the user message, and a conversation identifier, obtaining historical context pertaining to the user message from a plurality of storage mechanisms provided in differing formats including a knowledge graph, a vector database comprised of vector embeddings, and a database comprising text summaries of prior conversations between the user and the AI chatbot, generating a prompt for a LLM that instructs the LLM to generate a response to the user message that is based on and consistent with the user message, the initial content, and the historical context, and providing a final response to the user that is corresponds to an LLM-generated response.
Owner:BAYARDELLE ELIZABETH

Unstructured data extraction with large language models for query resolution

An unstructured data query-response pair generation system (generation system) populates a knowledge base of query-response pairs for queries of natural language content in unstructured data by prompting a first large language model (LLM) text extracted from the unstructured data. An unstructured data chatbot (chatbot) leverages the knowledge base by augmenting prompts to a second LLM responding to user queries for natural language content in the unstructured data with query-response pairs having queries that are semantically similar to the user queries. The knowledge base and LLMs are updated based on user feedback correcting responses, continually improving quality of the generation system and chatbot.
Owner:PALO ALTO NETWORKS INC

Retrieval augmented generative question and answer boosting

Systems or techniques are provided for facilitating retrieval augmented generative question and answer boosting. In various embodiments, a system can access a plain text question regarding a scientific instrument. In various aspects, the system can generate, via a large language model that references a document-graph repository, a structured or unstructured answer for the plain text question. In various instances, the document-graph repository can comprise a plurality of document-graphs that respectively correspond to a plurality of technical documents. In various cases, for a first document-graph that corresponds to a first technical document, leaf nodes of the first document-graph can represent respective text blocks written in the first technical document, and non-leaf nodes of the first document-graph can respectively represent a document title, one or more section headings, and one or more scientific instrument identifiers written in the first technical document and beneath which the respective text blocks are nested.
Owner:PPD DEVELOPMENT LP +2

Large language model interactions via intelligent prompt enrichment module and updated profile

Various embodiments of the technology described programmatically access a user query intended for a Large Language Model (LLM), analyze the user query, and determine prompt-enriching information that is combined with the user query to generate an enriched user query that is ultimately communicated to the LLM. In this manner, additional prompt-enriching information or context is added to the user query before being communicated to the LLM so that the additional prompt-enriching information, along with the user query, can be tokenized to better guide the LLM to a more accurate answer without modifying weights, parameters, or training of the LLM. Certain embodiments have the technical effect of improved accuracy relative to existing approaches by enriching user queries with prompt-enriching information to generate an enriched user query that is passed to the LLM. Based on the enriched user query, certain embodiments reduce the likelihood of hallucinations present in the LLM response.
Owner:RIOT GAMES INC

Logical text passage generation and retrieval for retrieval-augmented generation

Techniques for logical text passage generation and retrieval for retrieval-augmented generation. The techniques involve processing markup language documents to generate logical text passages and their corresponding embeddings. These embeddings are indexed for efficient retrieval. Upon receiving a user utterance, a user query is formed and transformed into an embedding to query the index. Relevant text passages are identified and used to prompt a large language model (LLM), which generates a completion. This completion is then sent as a response to the user. The process effectively bridges user queries with relevant information through advanced embedding and natural language processing techniques, enabling accurate and contextually appropriate interactions within a user-agent dialogue framework.
Owner:AMAZON TECH INC

Chunk synthesis for retrieval augmented generation assistants

A query answering system may access a collection of data sources to populate an index. A query answering system derives content from a collection of data sources to create synthetic chunks that are each representative of a portion of content from one or more of the data sources. A query answering system populates the index with the synthetic chunks. A query answering system identifies a subset of the synthetic chunks as relevant to a user query, generates a large language model (LLM) prompt that includes the subset of the synthetic chunks from the index and the user query, provides the LLM prompt to an LLM., and generates a response to the user query based on output of the LLM.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Method and system for verifying authenticity of a document

ActiveUS20180154676A1Digital data information retrievalPaper-money testing devicesSymbolic SystemsDigital copy
A system and a method for verifying authenticity of a physical copy and a digital copy of a document are disclosed. The method comprises registering a document in a repository by storing details related to the document in a location of the repository. A symbology for the document is generated. The symbology is an identifier of the location of the repository comprising the document. The symbology is associated with either a physical or a digital copy of the document. The digital copy of the document is printed to generate a printed copy. The printed copy or the physical copy of the document is scanned to generate a scanned image. The document and the details related to the document present at the location of the repository are accessed. The scanned image is compared with the document stored in the repository to determine the authenticity of either the physical copy or the digital copy of the document.
Owner:VERIDOC SYSTEMS LLC +1

AI-based video summary generation for content consumption

A data processing system implements receiving content and a call requesting a generative model to generate a video summary of the content; constructing a prompt including the content and instructions to the model to identify semantic context of the content, to identify a text data item, an audio data item, and / or a video data item embedded in the content to generate a text transcript of the audio data item and / or the video data item, or a textual description of the video data item, to summarize the text data item, the text transcripts, and / or the textual description as a summary of the content based on the semantic context, and to generate the video summary based on the summary and a portion of the text data item, the audio data item, and / or the video data item; providing the first prompt to the generative model; providing the video summary to a client device for presentation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Optimizing retrieval-augmented generation systems through enhanced document selection

A method includes applying a document ranking layer of a document selection large language model (LLM) to a document list including multiple reference documents to obtain a ranked document list. The method further includes selecting a subset of reference documents from the ranked document list and processing a user prompt and the document subset by a field LLM to generate an answer. The method further includes ranking the answer with an answer score by a ranking LLM. The method further includes ranking the document subset by the ranking LLM to obtain a ranked document subset. The method further includes calculating a loss function of a preference optimization layer of the document selection LLM based on the answer score and updating at least one training parameter of a foundation model of the document selection LLM based on the loss function of the preference optimization layer.
Owner:INTUIT INC

Permission-based ai system responses

A method and apparatus are disclosed for generating permission-based large language model responses by using a query received from a user to identify a plurality of documents that are semantically similar to the query, using an access token received from the user to identify user accessible documents from the plurality of documents that the user is permitted to access, processing the user accessible documents to define a context of user accessible documents that is associated with the query, and then submitting the query and the context of user accessible documents to a large language model (AI system) to generate an AI system response to the query.
Owner:JIVE SOFTWARE LLC

Chunk traceability and user-based access control in horizontal retrieval augmented generation

In one implementation, a device may maintain access permissions that control whether a user is allowed to access a particular document. The device may match a plurality of document chunks from a retrieval augmented generation system to a prompt issued by the user for input to a language model. The device may form, based on the access permissions, a modified set of document chunks by excluding a particular document chunk from the plurality of document chunks based on the particular document chunk having a data lineage from the particular document. The device may augment the prompt using the modified set of document chunks prior to input to the language model.
Owner:CISCO TECHNOLOGY INC

Machine question and answer dialogue method and device

The invention discloses a machine question and answer dialogue method and device. The method comprises the following steps: receiving a user request of a target user, and obtaining question information in a target round call corresponding to the user request from the user request; a memory document of the target user is obtained, a preset model is adopted to analyze the question information and the memory document, answers corresponding to the question information are generated, information in the memory document is updated through movement of the sliding window, and the memory document at least comprises one of a user portrait and historical dialogue information; and outputting the answer. According to the method and the device, the technical problem that the accuracy of the obtained answers is relatively low due to the fact that the answers can be recommended only according to several adjacent rounds of dialogues in the related technology is solved.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Automatic official document generation method based on dynamic rules

The invention discloses an automatic official document generation method based on a dynamic rule, and relates to the technical field of official document generation, and the method comprises the following steps: receiving official document writing demand data input by a target user, carrying out semantic analysis on the official document writing demand data, extracting structured parameter data based on a semantic analysis result, and generating an official document according to the structured parameter data. Performing official document type identification on the structured parameter data to confirm an official document type identification code; a dynamic rule engine is used for matching the corresponding format verification rule base and the content logic knowledge graph, and an official document writing framework is generated; receiving business data corresponding to the target user, extracting a key information triple from the business data, inputting the key information triple into a preset intelligent mapping model for semantic level association analysis, and generating a preliminary official document manuscript in combination with an official document writing framework; and verifying the preliminary official document manuscript, and adjusting and optimizing the preliminary official document manuscript according to a verification result to generate a final official document manuscript. The method has the effect of improving the document generation efficiency.
Owner:HEBEI WORLDEYES INFORMATION TECH

Retrieval enhancement generation method and system based on composite knowledge base

The invention discloses a retrieval enhancement generation method and system based on a composite knowledge base, and belongs to the technical field of retrieval enhancement generation. Obtaining an original document, dividing the original document into text blocks according to paragraphs, and storing the text blocks into a vector database to form a text vector library; constructing a knowledge hypergraph based on the text blocks, and storing the knowledge hypergraph into a vector database to form a knowledge hypergraph database; obtaining cases and storing the cases into a vector database to form a case library; and obtaining a question input by a user, and generating a final answer by using a large language model based on a composite knowledge base formed by the text vector base, the knowledge hypergraph database and the case base. By constructing a composite knowledge base, the problems of context breakage, information loss, lack of professional expression and the like existing when an existing retrieval enhancement generation method is used for processing long documents are solved.
Owner:NORTHEASTERN UNIV CHINA

Text sentiment analysis method, system and equipment based on multi-granularity sentiment modeling and medium

The invention discloses a text sentiment analysis method, system and equipment based on multi-granularity sentiment modeling and a medium, and the text sentiment analysis method comprises the following steps: obtaining a to-be-analyzed initial text, and carrying out standardized preprocessing on the initial text to obtain text data; performing coarse-grained sentiment analysis on the text data by using a chapter-level encoder to generate chapter-level sentiment tags; performing fine-grained sentiment analysis on the text data by using a sentence-level encoder to generate a sentence-level sentiment tag; performing local correction on the sentence-level emotion label based on the chapter-level emotion label by using a cross-layer attention mechanism to obtain an updated sentence-level emotion label; extracting entity features and attribute tags in the text data, and associating the entity features, the attribute tags and the updated sentence-level emotion tags to generate an emotion triple; and carrying out conflict analysis on the emotion triad to obtain a structured emotion label of the initial text. According to the invention, the context consistency and accuracy of the sentiment analysis result can be improved.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY

Systems and methods for search alteration in non-linear media

Aspects of the present application help to alter searches in non-linear media. In some embodiments, a server may receive a part of a user search query related to non-linear content. The server may access progress data that is associated with the user profile and perform a search on a database using the progress data and the search query to produce a number of candidate content items (e.g., search results). The server may access a data structure that defines possible paths through the non-linear media, and may filter the candidate content items based on the progress data and the data structure that defines the possible paths through the non-linear media to generate output of a filtered candidate content item.
Owner:ADEIA GUIDES INC

Systems and method for enhanced conversational performance of large language models using adaptive retrieval-augmented generation

Systems and methods for enhanced conversational performance of large language models using adaptive retrieval-augmented generation are disclosed. A method may include: (1) receiving a query from a user; (2) retrieving a plurality of summaries of historical conversations from a database of historical conversation summaries similar to the query; (3) generating a first prompt comprising the query and the plurality of summaries; (4) submitting the first prompt to a first large language model (LLM); (5) receiving, from the first LLM, a first response; (6) presenting the first response to the user; (7) generating a second prompt for a summary of the query and the first response; (8) submitting the second prompt to a second LLM; and (9) saving a second response to the second prompt from the second LLM to the database of historical conversation summaries, wherein the second response comprises the summary.
Owner:JPMORGAN CHASE BANK NA +1

Systems and methods for dynamic evaluation of metadata consistency and data reliability

Systems and methods for dynamically evaluating metadata consistency and data reliability in a data management system are disclosed herein. The system may retrieve first metadata and second metadata. The system may retrieve a metadata ruleset. Based on the metadata ruleset, the system may generate a first metadata consistency metric indicating a first measure of consistency. The system may determine to process each record of the first metadata as a batch. The system may generate a second metadata consistency metric indicating a second measure of consistency. The system may determine to process each record of the second metadata independently.
Owner:CAPITAL ONE SERVICES LLC

Ai-based video summary generation for content consumption

A data processing system implements receiving content and a call requesting a generative model to generate a video summary of the content; constructing a prompt including the content and instructions to the model to identify semantic context of the content, to identify a text data item, an audio data item, and / or a video data item embedded in the content to generate a text transcript of the audio data item and / or the video data item, or a textual description of the video data item, to summarize the text data item, the text transcripts, and / or the textual description as a summary of the content based on the semantic context, and to generate the video summary based on the summary and a portion of the text data item, the audio data item, and / or the video data item; providing the first prompt to the generative model; providing the video summary to a client device for presentation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Natural language query method and device in power field, terminal equipment and storage medium

The invention discloses a natural language query method and device in the electric power field, terminal equipment and a storage medium, and belongs to the field of natural language process.The method comprises the steps that a natural language query statement input by a user is received, semantic disambiguation is conducted on the natural language query statement according to the historical context of the natural language query statement, and the semantic disambiguation result is obtained; obtaining a target entity; according to the natural language query statement, the target entity and a preset task target, constructing a cue word, and inputting the cue word into a pre-trained large language model to obtain a first SQL statement; executing the first SQL statement to obtain structured service data, and constructing a first cause and effect graph according to the structured service data; and optimizing the first SQL statement according to the cause and effect graph to obtain a second SQL statement, and executing the second SQL statement to obtain a query result. The problem of low power data query accuracy in the power field can be solved.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO LTD

An end-to-end approach to determining high-quality digital content recommendations

Embodiments of the disclosed technologies are capable of evaluating content recommendations. The embodiments describe creating a prompt using a search query and a content recommendation output by a machine learning model in response to the search query. The embodiments further describe causing a LLM to generate an evaluation of the content recommendation and the search query using the prompt. The evaluation includes a relevance score of the content recommendation and the search query. The embodiments further describe training the machine learning model to generate an updated content recommendation in response to the search query. The training includes using the relevance score of the content recommendation and the search query.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Systems, apparatuses, methods, and non-transitory computer-readable storage media for adaptive information retrieval for question-answering

Methods and systems for retrieving relevant information in response to an input question. The method includes obtaining text content related to the input question and partitioning the content into one or more paragraphs based on predefined rules. The method further involves extracting one or more evidence spans that are relevant to the input question by inputting the text content and the question into a trained language model. A semantic search is then performed on both the paragraphs and the extracted evidence spans, ranking the candidate passages based on their relevance to the input question. Each candidate passage may comprise either a paragraph or an evidence span that addresses the question. The disclosed methods and systems improve the quality and relevance of retrieved information by combining heuristic-based content partitioning with machine learning-based evidence extraction.
Owner:HUAWEI TECH CO LTD

Memory processing method and system of generative model, information interaction method and system, medium and equipment

The invention discloses a memory processing method and system for a generative model, an information interaction method and system, a medium and equipment. The method comprises the steps of obtaining interaction information between a user and a generative model, and storing the interaction information in a first database of a basic memory area; extracting and summarizing the interaction information in the first memory period to obtain memory fragment information, and storing the memory fragment information in a second database of a fragment memory area; summarizing the memory fragment information in the second memory period into topic information according to topic relevancy, and storing the topic information in a third database of a topic memory area; summarizing the topic information exceeding the second memory period as topic conclusion information, and storing the topic conclusion information in a vector retrieval engine of a historical memory area; and in response to received user query information, querying related information in the basic memory area, the fragment memory area, the topic memory area and the historical memory area, and generating reply information. According to the method, the information storage, processing and retrieval capabilities of the model can be improved.
Owner:RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD

Sticker search icon providing dynamic previews

Examples described herein relate to techniques for facilitating selection of stickers for inclusion in messages within the context of an interaction system. According to some examples, message content is detected and a set of candidate stickers is identified based on the message content. A search icon is dynamically replaced with a representation of respective ones of the set of candidate stickers. At a first point in time, the search icon represents a first candidate sticker of the set of candidate stickers. At a second point in time, the search icon represents a second candidate sticker of the set of candidate stickers.
Owner:SNAP INC

Micro-report generation method and device and storage medium

The invention discloses a micro-report generation method and device and a storage medium, and relates to the technical field of data processing, and the micro-report generation method comprises the steps that a natural language question input by a user is received, the natural language question is matched with a preset question example library, and a target question example is determined; obtaining a business index in the natural language question, and determining a data model corresponding to the business index; based on the target problem example, the business index and the data model, generating a calculation semantic configuration; generating a target SQL query by combining the corresponding model information, the data source information and the authority control rule based on the calculation semantic configuration; and generating the micro-report by executing the target SQL query. According to the method, the target question example is determined by matching the natural language question with the preset question library, and the semantic configuration is dynamically generated and calculated in combination with the specific business index and the data model, so that the target SQL query is generated, and the effect of immediately responding to the dynamic query according to the natural language question is achieved.
Owner:CHINA MERCHANTS BANK

Method and system for large language model (LLM)-selection for response generation to user queries

Disclosed herein, is a method and system for selecting a LLM for response generation to user queries. The method includes receiving a user query from a user device. The method includes determining, for the user query, a query type from a set of query types through a fine-tuned text classification model. The method includes retrieving a plurality of document embeddings based on the user query and the query type from a vector database through a semantic search technique. The method includes preparing a prompt using the user query and the plurality of document embeddings. The method includes inputting the prompt to an LLM selected from a set of LLMs based on the query type. The method includes generating, via the selected LLM, a response to the user query based on the prompt.
Owner:L&T TECH SERVICES LTD