Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

974results about "Document management systems" patented technology

Intelligent monitoring management method and system based on archive digitization

The invention discloses an intelligent monitoring management method and system based on archive digitization, and relates to the technical field of data management, and the method comprises the steps: collecting and preprocessing multi-source archive data, employing a multi-mode BERT model to carry out the feature fusion of different data sources, and generating a unified semantic representation; semantic labeling is performed on archive data through a multi-label classification model, a semantic graph of archive content is constructed by using a graph database, an association relationship between archives is represented, a semantic index tree is constructed based on the semantic graph, and rapid positioning and calling of the archive content are optimized; and recording the change of each file version, positioning the change position based on a semantic index tree, identifying the semantic change of the file through a semantic difference comparison algorithm, recording hash, carrying out granularity division on the file content through the semantic boundary of each level of node in the index tree, and generating a user access strategy. According to the invention, dynamic perception and risk early warning of user behaviors are realized, and the intellectualization and safety of the archive management system are effectively improved.
Owner:XIAN XINCHUANG TECH CO LTD

Semantic-tree-based ai content management platform

A data processing system implements receiving a call requesting a generative model to generate a semantic tree for a source content; constructing a first prompt including the source content and instructions to the model to analyze a semantic structure of the source content and to generate a semantic outline and content chunks of the source content, the semantic outline including one or more topics each connected with one or more of the content chunks, to compute one summary for each of the content chunks, to apply indices to reference each topic node of the semantic tree to one of the topics, and to apply indices to reference each leaf node of the semantic tree to one of the content chunks and the respective summary; providing the first prompt to the model and receiving the semantic tree of the source content; and storing the semantic tree in a database.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Enhanced document retrieval with semantic depth and syntactic structure

Certain aspects of the present disclosure describe a method of information retrieval. In certain aspects, the method includes identifying a set of relevant nodes of a document graph embedding semantic units associated with the document based on a document search query. The method further includes reconstructing a structural context for each relevant node in the set of relevant nodes. The method further includes processing the set of relevant nodes and the structural context of each relevant node with a large language model to generate a contextual response to the document search query.
Owner:INTUIT INC

Methods and Systems for Improved Document Processing and Information Retrieval

Disclosed are methods, systems, devices, apparatus, media, and other implementations for improved search-time content retrieval performed by an information retrieval platform. The implementations include a method including determining by a searching system, at a first time instance, one or more first search results for a first query comprising one or more query terms associated with a first concept, modifying the query, based on the determined one or more first search results, to include one or more modified query terms associated with one or more concepts hierarchically related to the first concept, and determining at a subsequent time instance one or more subsequent search results for the modified query comprising the one or more modified query terms.
Owner:PRYON INC

Optimizing retrieval-augmented generation systems through enhanced document selection

A method includes applying a document ranking layer of a document selection large language model (LLM) to a document list including multiple reference documents to obtain a ranked document list. The method further includes selecting a subset of reference documents from the ranked document list and processing a user prompt and the document subset by a field LLM to generate an answer. The method further includes ranking the answer with an answer score by a ranking LLM. The method further includes ranking the document subset by the ranking LLM to obtain a ranked document subset. The method further includes calculating a loss function of a preference optimization layer of the document selection LLM based on the answer score and updating at least one training parameter of a foundation model of the document selection LLM based on the loss function of the preference optimization layer.
Owner:INTUIT INC

Heterogeneous document set-oriented cross-modal semantic alignment and logic consistency verification system

The invention relates to document verification, in particular to a heterogeneous document set-oriented cross-modal semantic alignment and logic consistency verification system, which is used for heterogeneous document input, supports multi-format document input and comprises multi-modal elements including texts, pictures, tables and charts. The document analysis module is used for carrying out structured extraction on document contents; extracting multi-modal elements, identifying and classifying various elements in the document, and establishing position and type labels of a foundation; the knowledge graph construction module is used for uniformly modeling heterogeneous elements into a multi-modal knowledge graph; the graph neural network semantic alignment module is used for realizing accurate cross-modal semantic alignment by using a specially designed graph neural network based on the multi-modal knowledge graph; the hybrid consistency verification engine is used for performing logic consistency verification in combination with a symbol logic verification mechanism and a semantic consistency verification mechanism; according to the method, the defect that accurate cross-modal semantic alignment and logic consistency verification are difficult to carry out on professional documents with multi-modal elements can be effectively overcome.
Owner:ANHUI GAOSHAN TECH CO LTD

Computer implemented method for question answering

A computer-implemented method of generating an answer from an input query and input documents, comprising extracting input query entities from the input query and input document entities from the input documents, sampling a schema of in-domain queries with the input query to generate a query sampled schema, generating an entity-document graph from the input documents and input document entities, generating a hyper-relational knowledge graph by extracting, for each input query entity, a document title and relation to an input document entity of the input document entities from the input documents in the entity-document graph, sampling the hyper-relational knowledge graph with the query sampled schema to generate a query focused hyper-relational knowledge graph, predicting an answer to the input query by inputting the query focused hyper-relational knowledge graph and input query into a pretrained neural network, outputting the answer.
Owner:FUJITSU LTD

Retrieval-augmented generation for large language models

A document preparation method involves creating a hierarchical representation of an input document without summarizing or omitting any content. The method uses a generative language model to generate the hierarchical representation and stores it in a repository for later use by a client generative language model. This allows for more accurate and complete generation of text, enabling the use of retrieval units to enhance the output of the client generative language model while efficiently exploiting its limited context window.
Owner:POMA AI GMBH

Intelligent extraction and indexing system for file metadata

The invention relates to the technical field of archive information management, and discloses an archive metadata intelligent extraction and indexing system, which comprises a multi-modal preprocessing module for obtaining and preprocessing original multi-modal archive data; the context entity recognition module is used for performing entity recognition and standardization according to the context vector; the cross-modal fusion module is used for carrying out confidence weighted multi-modal fusion and logic verification; the archive association module is used for carrying out association identification and consistency detection between archives; the intelligent indexing module is used for carrying out hierarchical intelligent indexing and quality feedback on the consistency constrained file metadata set; the quality evaluation module is used for carrying out metadata quality evaluation and active repair on the standardized indexing result; the knowledge graph module is used for constructing a time sequence knowledge graph and intelligent retrieval service; according to the method, collaborative extraction of cross-modal information is realized by constructing a confidence-weighted multi-modal fusion model and a bidirectional attention mechanism.
Owner:SHANDONG ZHENGTU INFORMATION POLYTRON TECH INC

BOM table generation method based on full life cycle management, computer equipment and computer readable storage medium

The invention relates to a BOM table generation method based on full life cycle management, computer equipment and a computer readable storage medium, and the method comprises the following steps: S1, generating a client matrix according to a product design document and an engineering material, and obtaining information in the client matrix; the customer matrix is presented in the form of a two-dimensional table, column information in the two-dimensional table represents serial numbers of product parts and / or components of different assembly levels, and row information at least comprises names, numbers, specifications and models of the product parts and / or components and parent component identifiers; s2, generating an in-plant BOM matrix according to the material key information, the in-plant material number, the in-plant inventory status and the supply chain constraint condition, inputting the in-plant BOM matrix into a preset management system to generate an electronic drawing and document database, and establishing a structured D-BOM; s3, comparing the client matrix information obtained in the step S1 with the D-BOM established in the step S2; s4, carrying out adjustment and optimization on the D-BOM; and S5, importing the optimized D-BOM into a client matrix, and generating a final BOM.
Owner:RI SHAN COMPUTER ACCESSORY (JIASHAN) CO LTD

Large language model (LLM)-based knowledge resource retriever and ranker

Disclosed herein are a system, method, and computer program product embodiments for retrieving and ranking knowledge resources relevant to a query from knowledge base(s). For example, a query for resources from knowledge base(s) may be received. Based on the query, a first set of candidate resources are obtained from the knowledge base(s) having a lexical similarity to the query search terms, and a second set of candidate resources are obtained from the knowledge base(s) having a semantical similarity to the search terms. For each of the first and second sets of candidate resources, a confidence level indicating the relevance of the candidate resource to the query is determined. The sets of candidate resources are ranked based on at least the confidence levels to generate a ranked list of candidate resources. A query response comprising at least a subset of the ranked list candidate resources is provided to a GUI.
Owner:SAP SE

Semantic-tree-based ai content management platform

A data processing system implements receiving a call requesting a generative model to generate a semantic tree for a source content; constructing a first prompt including the source content and instructions to the model to analyze a semantic structure of the source content and to generate a semantic outline and content chunks of the source content, the semantic outline including one or more topics each connected with one or more of the content chunks, to compute one summary for each of the content chunks, to apply indices to reference each topic node of the semantic tree to one of the topics, and to apply indices to reference each leaf node of the semantic tree to one of the content chunks and the respective summary; providing the first prompt to the model and receiving the semantic tree of the source content; and storing the semantic tree in a database.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Generating probabilistic data structures for lookup tables in computer memory for multi-token searching

Methods, systems, and non-transitory computer readable storage media are disclosed for optimizing computer memory usage for lookup lists in computer memory via probabilistic data structures. For example, the disclosed system generates a probabilistic data structure (e.g., a Bloom filter) to represent data in a lookup list including multi-token items by hashing items of the lookup list to sets of bit values in a bit vector. The disclosed system classifies text content in a digital document by utilizing a maximum number of tokens from multi-token items in the lookup list to select and compare sets of sequential tokens in the digital document to the probabilistic data structure. The disclosed system also iteratively reduces the number of tokens in sets of sequential tokens for subsequent comparisons. Furthermore, in some aspects, the disclosed system causes a computing device to modify a digital document and / or database operations based on the classifications.
Owner:ONETRUST LLC

System and method for combining language models with natural language processing for summarization

A data processing system and method include receiving a set of documents to summarize a trend across the set of documents, inputting the set of documents into an information extraction model, executing the information extraction model to extract a first plurality of text segments, determining a second plurality of text segments based on the first plurality of text segments, determining a third plurality of text segments from the second plurality of text segments, generating a compressed representation of the set of documents from the third plurality of text segments to include in a prompt for a language model, inputting the prompt into the language model, and executing the language model to generate the summary based on the prompt for the set of documents.
Owner:SAS INSTITUTE INC

Optimized large language model inference from structured data via intermediate documents

ActiveUS20250370996A1Distillation corrosion inhibitionNatural language data processingLinguistic modelTheoretical computer science
Techniques for optimized LLM inference from structured data via intermediate documents (“LLMs”) are disclosed. In an example method, a computing device accesses a database including a set of collections. The computing device generates one or more documents based on a first collection of the set of collections. The computing device determines one or more portions of the one or more documents based on at least one topic of the one or more documents. The computing device generates an embedded representation of each portion. The computing device receives, from a client device, a first query, including at least a first topic. The computing device determines a first portion of the one or more portions based on the first topic. The computing device generates a response based on the first query and the first portion and outputs the response to the client device.
Owner:ZOOM COMMUNICATIONS INC

Optimized large language model inference from structured data via intermediate documents

ActiveUS12530348B2Distillation corrosion inhibitionNatural language data processingLinguistic modelTheoretical computer science
Techniques for optimized LLM inference from structured data via intermediate documents (“LLMs”) are disclosed. In an example method, a computing device accesses a database including a set of collections. The computing device generates one or more documents based on a first collection of the set of collections. The computing device determines one or more portions of the one or more documents based on at least one topic of the one or more documents. The computing device generates an embedded representation of each portion. The computing device receives, from a client device, a first query, including at least a first topic. The computing device determines a first portion of the one or more portions based on the first topic. The computing device generates a response based on the first query and the first portion and outputs the response to the client device.
Owner:ZOOM COMMUNICATIONS INC

Detecting evasive prompts for generative artificial intelligence systems

A genetic algorithm is implemented to generate prompts that evade content filters of generative artificial intelligence (AI) systems. The genetic algorithm applies grammar operations to mutate candidate prompts, communicates the candidate prompts to generative AI systems, and selects candidate prompts that successfully evade content filters according to corresponding responses. A disambiguation model that corrects grammar in prompts is tested on the selected prompts to determine if grammar is properly corrected. Once tested, the disambiguation model is deployed in an ensemble with a classifier that outputs verdicts for prompts with grammar corrected by the disambiguation model.
Owner:PALO ALTO NETWORKS INC

File positioning management method and system based on artificial intelligence

The invention provides a file positioning management method and system based on artificial intelligence, and relates to the technical field of artificial intelligence. According to the method, files are collected from multiple sources and subjected to standardization processing, an element set is generated in combination with multi-modal analysis of texts, images, audios, videos and tables, cross-modal alignment is achieved through a semantic representation model, hierarchical indexes of semantics, keywords and relations are constructed, and a unique traceability identifier is generated; in the query stage, intention recognition and joint retrieval are carried out, a result subjected to permission verification and traceability information labeling is output, online optimization and incremental reconstruction are executed based on user feedback, and comprehensiveness, accuracy, traceability and self-adaptive optimization of file positioning are achieved.
Owner:ZUNYI NORMAL COLLEGE

Context-aware information retrieval

Certain aspects of the disclosure provide for information retrieval that exploits context derived from document structure. Source documents can be preprocessed to identify fields and determine context attributes related to each field based on the structural layout of a source document. Resource documents can also be preprocessed to segment a resource document into passages and determine context related to the passages based on structural layout. Queries pertaining to a field can be enhanced by adding context metadata associated with the field. A query embedding can be generated and compared with previously generated passage embeddings to locate candidate matches based on similarity. A machine learning model can be provided with the top-ranked passages and tasked with re-ranking the passages based on relevancy to the original query. The highest re-ranked passage or set of passages can be output in response to the query.
Owner:INTUIT INC

Database management workflow with controlling documents

Methods and systems are disclosed for database management with controlling documents. Tasks addressed include: identification of requirements in the documents, tracking changes as documents evolve, mapping documents or requirements to database entries, identifying gaps between documents and the database, and proposing database updates. Disclosed embodiments address these tasks using a combination of sequential program logic, machine-learning tools, and client interaction. Workflows address one or more tasks. Examples pertaining to regulatory documents are presented. Variations are disclosed.
Owner:SAP SE

Dynamic multimodal prompt generation for efficient content moderation

Aspects of the disclosure include methods and systems for content moderation, and specifically dynamic multimodal prompt generation for efficient content moderation. A method includes receiving, by a prompt generation system, a request for a decision corresponding to content. The method includes generating, by an encoder of the prompt generation system, an embedding of the content, and retrieving, by an embedding based retrieval (EBR) module of the prompt generation system, K retrieved chunks from a database, the K retrieved chunks having a Kth closest distance to the embedding in an embedding space. A dynamic prompt comprising a prompt template, multiple retrieved chunks of the K retrieved chunks, and the content is generated and input to a pre-trained large language model. The LLM generates the decision, which is returned responsive to the request.
Owner:MICROSOFT CORP

LLM framework for large scale applications

A method includes a computer receiving a user query. The computer generates a summary of the user query using a first large language model. The computer determines a user issue from a first database based on the summary. The computer determines digital document from a second database based on the user issue. The computer generates a prompt based on the digital document and a prompt template. The computer generates a response based on the prompt using a second large language model.
Owner:DOORDASH INC

Systems and Methods for Node Graph Data Storage Across Disparate Data Sources

The present disclosure describes a method for updating a node graph data structure, comprising storing a node graph data structure comprising a plurality of entity nodes and a plurality of attribute nodes; receiving, from a data source during a plurality of time periods, a plurality of data files comprising data for a first entity; identifying a plurality of edges between a first entity node of the plurality of entity nodes that identifies the first entity and an attribute node of the plurality of attribute nodes that identifies a first attribute of the first entity, each of the plurality of edges corresponding to a value and a different time period; and updating a value stored in a data structure for an edge that corresponds to a time period associated with a data file.
Owner:MICHIGAN HEALTH INFORMATION NETWORK SHARED SERVICES

Tokenization systems and methods for redaction

A tokenization system receives a request for redaction of sensitive textual content in a document, identifies a portion of the document as the sensitive textual content, and edits the document, including replacing the sensitive textual content thus identified with tokens, each token having a token value and a pattern that identifies a start and an end of the token value. The editing produces a transformed version of the document with the tokens and without the sensitive textual content. The tokenization system may then communicates the transformed version of the document with the tokens and without the sensitive textual content to the client computing system, an automated recognition service, or a redaction plug-in to a frontend application.
Owner:OPEN TEXT CORPORATION

Multi-modal document comparison method and system, storage medium and program product

The invention discloses a multi-modal document comparison method and system, a storage medium and a program product, and relates to the technical field of data processing, and the method comprises the following steps: constructing a first index database based on a first document, and constructing a second index database based on a second document; traversing each first document block in the first document, retrieving a matching block matched with the first document block in each second document block of the second document based on the second index database, and determining a block difference based on the matching block; traversing each second document block, and determining a deleted block in each second document block based on the first index database; and writing the block difference and the deleted block into a difference knowledge graph, and generating a document comparison report based on the difference knowledge graph. According to the method, comprehensive and structured presentation of semantic-level difference information is realized through the architecture of the multi-modal block index-bidirectional matching retrieval-difference mapping knowledge domain, so that the semantic-level difference recognition accuracy of a multi-modal document comparison technology is improved.
Owner:SHANGHAI QINGCHENG JIZHI TECHNOLOGY CO LTD

Multi-aspect vector search index creation and retrieval method and system

A computer-implemented method for creating and utilizing a multi-aspect vector search index is disclosed. The method involves generating multi-dimensional vectors representing digital documents within a document repository, enabling advanced content retrieval through comprehensive vector-based indexing. The approach allows for sophisticated search capabilities by mapping documents across multiple aspects, facilitating precise and efficient content identification and extraction from extensive document collections.
Owner:XILLIO AI BV

Method and system of converting unstructured digital documents to a structure format using a secure API

In one aspect, a computerized method for document extraction workflow for unstructured documents includes the steps of implementing a text mining operation on a set of digital documents the incoming documents. This is done by defining a document type of each digital document. Based on the document type, the method defines a set of data dictionaries to extract any data from each digital document. The method uses the defined set of data dictionaries to extract any data from each digital document.
Owner:YERRAMSETTY VENKATA SAI RAMAN +1