Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1480results about "Document management systems" patented technology

Intelligent monitoring management method and system based on archive digitization

The invention discloses an intelligent monitoring management method and system based on archive digitization, and relates to the technical field of data management, and the method comprises the steps: collecting and preprocessing multi-source archive data, employing a multi-mode BERT model to carry out the feature fusion of different data sources, and generating a unified semantic representation; semantic labeling is performed on archive data through a multi-label classification model, a semantic graph of archive content is constructed by using a graph database, an association relationship between archives is represented, a semantic index tree is constructed based on the semantic graph, and rapid positioning and calling of the archive content are optimized; and recording the change of each file version, positioning the change position based on a semantic index tree, identifying the semantic change of the file through a semantic difference comparison algorithm, recording hash, carrying out granularity division on the file content through the semantic boundary of each level of node in the index tree, and generating a user access strategy. According to the invention, dynamic perception and risk early warning of user behaviors are realized, and the intellectualization and safety of the archive management system are effectively improved.
Owner:XIAN XINCHUANG TECH CO LTD

Enhanced searching using fine-tuned machine learning models

An advanced search system leverages a pre-trained large language model to enhance user query responses. The system, equipped with hardware processors, a search query via an interface and accesses a pre-trained large language model designed to respond to the search query. The system fine-tunes the model to generate a task-specific generative model. The system employs the task-specific generative model to generate a search result to the search query and analyzes the search result based on a performance metric associated with the task-specific generative model. The system refines the task-specific generative model based on the analyzing of the search result.
Owner:SNOWFLAKE INC

Semantic-tree-based ai content management platform

A data processing system implements receiving a call requesting a generative model to generate a semantic tree for a source content; constructing a first prompt including the source content and instructions to the model to analyze a semantic structure of the source content and to generate a semantic outline and content chunks of the source content, the semantic outline including one or more topics each connected with one or more of the content chunks, to compute one summary for each of the content chunks, to apply indices to reference each topic node of the semantic tree to one of the topics, and to apply indices to reference each leaf node of the semantic tree to one of the content chunks and the respective summary; providing the first prompt to the model and receiving the semantic tree of the source content; and storing the semantic tree in a database.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Ensemble augmentation with enhanced knowledge extraction techniques

Methods, systems, apparatuses, devices, and computer program products are described. A system may obtain a set of documents associated with a knowledge base for retrieval-augmented generation (RAG). The system may generate multiple representations of the information included in the documents using multiple knowledge extraction pipelines. For example, the system may generate a set of metadata-based vector embeddings based on the documents, a set of knowledge graphs based on the documents, and a set of hierarchical tree representations based on the documents. The system may receive a user query and may retrieve contextual information from the set of vector embeddings, the set of knowledge graphs, and the set of hierarchical tree representations to augment the user query for a large language model (LLM) prompt. The system may input the prompt to the LLM, and the LLM may output a response based on the user query and the contextual information.
Owner:SALESFORCE INC

Enhanced document retrieval with semantic depth and syntactic structure

Certain aspects of the present disclosure describe a method of information retrieval. In certain aspects, the method includes identifying a set of relevant nodes of a document graph embedding semantic units associated with the document based on a document search query. The method further includes reconstructing a structural context for each relevant node in the set of relevant nodes. The method further includes processing the set of relevant nodes and the structural context of each relevant node with a large language model to generate a contextual response to the document search query.
Owner:INTUIT INC

Intelligent archive description method and device and storage medium

The invention discloses an intelligent archive description method and device and a storage medium, which are applied to edge nodes, and the method comprises the steps: carrying out the hierarchical feature extraction and fusion of pre-collected archive information through a pre-deployed lightweight dual model, and generating a description item set with a confidence label; uploading the low-confidence bibliographic items in the bibliographic item set to a cloud server; acquiring description correction data transmitted by the cloud server, wherein the description correction data is generated by performing multi-modal comparison analysis on the low-confidence description item based on the cloud server; and integrating the bibliographic items which are not marked as the low-confidence bibliographic items with the bibliographic correction data, and generating a standard bibliographic file according to a standard bibliographic rule. And high-quality generation of the standard description file in a complex scene is realized.
Owner:BEIJING ZHONGYOU TECHNOLOGY CO LTD

Causal reasoning system

A research assistant system described herein includes a research assistant tool and associated components and a graphical user interface to guide user input to research, discover, and evidence answers for complex research questions. The research assistant system may include the graphical user interface (“GUI” or “user interface”) for presentation on a user device associated with a user. The user interface may provide prompts and guidance for collaboration and exploration of research concepts iteratively. A concept may include a search term, entities, and / or propositions / statements.
Owner:BRIDGEWATER ASSOCIATES EC IP LLC

Document graph

A method, an apparatus, and a computer-readable storage medium for generating a document graph. A plurality of electronic documents is received. Each electronic document has a predetermined document type. A machine learning model is selected from the plurality of machine learning models based on the predetermined document type. The selected machine learning model is instructed to extract a plurality of document portions from each electronic document in the plurality of electronic documents in accordance with the predetermined document type. A relationship between two or more document portions is defined based on the content of each document portion, and the document portions are associated based on the relationship. A graph structure having a plurality of nodes is generated. Each node includes at least one document portion. Each node is connected to another node in accordance with the relationship between document portions included in the nodes. The graph structure is stored.
Owner:DOCUSIGN INC

Intelligent archive opening identification method based on large model

The invention discloses an intelligent archive opening and identifying method based on a large model, particularly relates to the technical field of archive data auditing, and is used for solving the problems of insufficient cross-modal data analysis capability, lagging rule updating and low man-machine cooperation efficiency in the prior art. Fusing cross-modal features of texts, images and metadata through a hybrid expert model to generate multi-modal feature vectors, and dynamically allocating the multi-modal feature vectors to a rule network, a semantic network and a domain network for cooperative processing based on attention weights; the rule network parameters are optimized through gradient projection constraint, and regulation-driven real-time adaptation is achieved; matching sensitive data in combination with a multi-dimensional feature matrix of auditing personnel, and optimizing task allocation accuracy; removing redundant links by utilizing value flow analysis to generate a lightweight process, and recording as a tamper-proof evidence chain through a block chain evidence storage solidification operation; the auditing efficiency and accuracy are improved, and the compliance traceability is guaranteed.
Owner:CHONGQING SHIJI KEYI TECH DEV CO LTD

Methods and Systems for Improved Document Processing and Information Retrieval

Disclosed are methods, systems, devices, apparatus, media, and other implementations for improved search-time content retrieval performed by an information retrieval platform. The implementations include a method including determining by a searching system, at a first time instance, one or more first search results for a first query comprising one or more query terms associated with a first concept, modifying the query, based on the determined one or more first search results, to include one or more modified query terms associated with one or more concepts hierarchically related to the first concept, and determining at a subsequent time instance one or more subsequent search results for the modified query comprising the one or more modified query terms.
Owner:PRYON INC

Intelligent document processing method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an intelligent document processing method, device, equipment and medium, comprising: receiving a document processing request and generating a task identifier, creating and sending a document processing task, constructing an intermediate document structure and generating a plurality of subtasks, and processing the sub-tasks, writing states and page contents into a cache, regularly querying task states, updating accessible identifiers, responding to an access request, loading the page contents, and integrating pages to generate a complete document after all the sub-tasks are completed. According to the method, the document generation task is divided into a plurality of sub-tasks which can be independently processed, and page-level asynchronous loading is realized in combination with a state cache and a regular query mechanism, so that the response speed and the system processing efficiency are remarkably improved while the generation accuracy is ensured; the problems that in the prior art, document generation delay is high, loading is slow, and user waiting time is long are effectively solved.
Owner:CHINA PING AN LIFE INSURANCE CO LTD

Optimizing retrieval-augmented generation systems through enhanced document selection

A method includes applying a document ranking layer of a document selection large language model (LLM) to a document list including multiple reference documents to obtain a ranked document list. The method further includes selecting a subset of reference documents from the ranked document list and processing a user prompt and the document subset by a field LLM to generate an answer. The method further includes ranking the answer with an answer score by a ranking LLM. The method further includes ranking the document subset by the ranking LLM to obtain a ranked document subset. The method further includes calculating a loss function of a preference optimization layer of the document selection LLM based on the answer score and updating at least one training parameter of a foundation model of the document selection LLM based on the loss function of the preference optimization layer.
Owner:INTUIT INC

Heterogeneous document set-oriented cross-modal semantic alignment and logic consistency verification system

The invention relates to document verification, in particular to a heterogeneous document set-oriented cross-modal semantic alignment and logic consistency verification system, which is used for heterogeneous document input, supports multi-format document input and comprises multi-modal elements including texts, pictures, tables and charts. The document analysis module is used for carrying out structured extraction on document contents; extracting multi-modal elements, identifying and classifying various elements in the document, and establishing position and type labels of a foundation; the knowledge graph construction module is used for uniformly modeling heterogeneous elements into a multi-modal knowledge graph; the graph neural network semantic alignment module is used for realizing accurate cross-modal semantic alignment by using a specially designed graph neural network based on the multi-modal knowledge graph; the hybrid consistency verification engine is used for performing logic consistency verification in combination with a symbol logic verification mechanism and a semantic consistency verification mechanism; according to the method, the defect that accurate cross-modal semantic alignment and logic consistency verification are difficult to carry out on professional documents with multi-modal elements can be effectively overcome.
Owner:ANHUI GAOSHAN TECH CO LTD

Computer implemented method for question answering

A computer-implemented method of generating an answer from an input query and input documents, comprising extracting input query entities from the input query and input document entities from the input documents, sampling a schema of in-domain queries with the input query to generate a query sampled schema, generating an entity-document graph from the input documents and input document entities, generating a hyper-relational knowledge graph by extracting, for each input query entity, a document title and relation to an input document entity of the input document entities from the input documents in the entity-document graph, sampling the hyper-relational knowledge graph with the query sampled schema to generate a query focused hyper-relational knowledge graph, predicting an answer to the input query by inputting the query focused hyper-relational knowledge graph and input query into a pretrained neural network, outputting the answer.
Owner:FUJITSU LTD

Retrieval-augmented generation for large language models

A document preparation method involves creating a hierarchical representation of an input document without summarizing or omitting any content. The method uses a generative language model to generate the hierarchical representation and stores it in a repository for later use by a client generative language model. This allows for more accurate and complete generation of text, enabling the use of retrieval units to enhance the output of the client generative language model while efficiently exploiting its limited context window.
Owner:POMA AI GMBH

Intelligent extraction and indexing system for file metadata

The invention relates to the technical field of archive information management, and discloses an archive metadata intelligent extraction and indexing system, which comprises a multi-modal preprocessing module for obtaining and preprocessing original multi-modal archive data; the context entity recognition module is used for performing entity recognition and standardization according to the context vector; the cross-modal fusion module is used for carrying out confidence weighted multi-modal fusion and logic verification; the archive association module is used for carrying out association identification and consistency detection between archives; the intelligent indexing module is used for carrying out hierarchical intelligent indexing and quality feedback on the consistency constrained file metadata set; the quality evaluation module is used for carrying out metadata quality evaluation and active repair on the standardized indexing result; the knowledge graph module is used for constructing a time sequence knowledge graph and intelligent retrieval service; according to the method, collaborative extraction of cross-modal information is realized by constructing a confidence-weighted multi-modal fusion model and a bidirectional attention mechanism.
Owner:SHANDONG ZHENGTU INFORMATION POLYTRON TECH INC

Document authentication certification with blockchain and distributed ledger techniques

Embodiments are described herein for document authentication certification using information stored on a distributed ledger such as a blockchain. A distributed ledger may securely store document data describing the document. Use of a distributed ledger may provide an immutable, readily auditable record of the history of the document. Each user participating in the system may be assigned a unique identifier to be used for conducting transactions on the distributed ledger network. A user may also be provided with a digital security token such as a cryptographic key that is useable to authenticate the user and enable access to the document data stored on the distributed ledger(s).
Owner:UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)

System and Method for Enhancing Generative Artificial Intelligence (AI) Model-Based Document Search with Image Retrieval

A method, computer program product, and computing system for generating a plurality of chunks for a plurality of text portions of a document, wherein the document includes the plurality of text portions and a plurality of images. Each chunk is indexed using a word embedding. Each of the plurality of images is indexed based upon, at least in part, a position of a respective image relative to a corresponding chunk. An image placeholder is generated for each of the plurality of images. A plurality of image-enhanced embeddings is generated by inserting the image placeholder for each of the plurality of images into a respective word embedding for the corresponding chunk. The plurality of image-enhanced embeddings are provided for processing a query using a generative artificial intelligence (AI) model.
Owner:DELL PROD LP

Management method of multi-mode enterprise knowledge base system

The invention provides a management method of a multi-modal enterprise knowledge base system, which comprises the following steps: integrating multi-modal data in an enterprise, constructing a private knowledge base, and combining the private knowledge base with a public knowledge base; performing intention analysis on a multi-modal retrieval demand submitted by the user based on the joint result, performing condition retrieval on the private knowledge base and the public knowledge base based on an intention analysis result, and outputting multi-modal query data; and carrying out quality inspection on the condition retrieval process, carrying out dynamic optimization on a joint result based on a quality inspection result to obtain a final multi-modal enterprise knowledge base, and carrying out deployment management on the multi-modal enterprise knowledge base. The management effect of the multi-modal knowledge in the enterprise is greatly improved, and the retrieval accuracy of the knowledge is also improved.
Owner:SHENZHEN YUNZHIYIN TECH CO LTD

BOM table generation method based on full life cycle management, computer equipment and computer readable storage medium

The invention relates to a BOM table generation method based on full life cycle management, computer equipment and a computer readable storage medium, and the method comprises the following steps: S1, generating a client matrix according to a product design document and an engineering material, and obtaining information in the client matrix; the customer matrix is presented in the form of a two-dimensional table, column information in the two-dimensional table represents serial numbers of product parts and / or components of different assembly levels, and row information at least comprises names, numbers, specifications and models of the product parts and / or components and parent component identifiers; s2, generating an in-plant BOM matrix according to the material key information, the in-plant material number, the in-plant inventory status and the supply chain constraint condition, inputting the in-plant BOM matrix into a preset management system to generate an electronic drawing and document database, and establishing a structured D-BOM; s3, comparing the client matrix information obtained in the step S1 with the D-BOM established in the step S2; s4, carrying out adjustment and optimization on the D-BOM; and S5, importing the optimized D-BOM into a client matrix, and generating a final BOM.
Owner:RI SHAN COMPUTER ACCESSORY (JIASHAN) CO LTD

Attention embedded transformer network driven document data extraction

Attention embedded transformer network driven document data extraction is provided. For example, a system integrates one or more processors with a data repository to identify a document of a first type received from a client device. The system determines a portion of the document based on a boundary established by a digital overlay. The system generates, via a trained machine learning model, a query using the portion of the document determined based on the boundary, wherein the query is designed to facilitate an extraction of data relating to the first type. The system inputs the query into a trained attention embedded transformer network model to extract data from the document, the extracted data including at least the extraction of data relating to the first type. The system displays, via the client device, the extracted data.
Owner:ADP INC

Large language model (LLM)-based knowledge resource retriever and ranker

Disclosed herein are a system, method, and computer program product embodiments for retrieving and ranking knowledge resources relevant to a query from knowledge base(s). For example, a query for resources from knowledge base(s) may be received. Based on the query, a first set of candidate resources are obtained from the knowledge base(s) having a lexical similarity to the query search terms, and a second set of candidate resources are obtained from the knowledge base(s) having a semantical similarity to the search terms. For each of the first and second sets of candidate resources, a confidence level indicating the relevance of the candidate resource to the query is determined. The sets of candidate resources are ranked based on at least the confidence levels to generate a ranked list of candidate resources. A query response comprising at least a subset of the ranked list candidate resources is provided to a GUI.
Owner:SAP SE

Integrated Management & Governance of Document Portfolio

An integrated document portfolio management and governance system and method are disclosed. The system includes a computing unit having an application interface adapted to present and / or formulate at least one input query. The system further includes aa central controller having a backend server communicably connected to the application interface of the computing unit. The backend server includes a data receiving component adapted to receive document dataset, each comprising a plurality of data elements, from a plurality of data sources in one or more formats. The backend server further includes a data ingestion module adapted to detect, normalize, and aggregate the plurality of data elements of the document dataset and subsequently store them within a central data repository. Furthermore, the backend server includes an ontology generator module adapted to create and maintain a dynamic ontology for the ingested datasets in real-time, wherein the plurality of data elements is categorized and contextualized in accordance with the dynamic ontology. Additionally, the backend server includes a governance module adapted to enforce & monitor data compliance policies and a data analysis module adapted to analyze the ingested data and generate actionable insights, wherein the actionable insights include one or more predictive analysis, data accuracy status, governance status, operational inefficiency, risk indicators, compliance gaps, and risk lineage and integrity. In operation, a user formulates an input query towards the central controller which in response is configured to automatically manage, govern & monitor the received data and subsequently visualize one or more actionable insights and / or compliance gaps onto the application interface of the computing unit.
Owner:BJONTEGARD BERNT ERIK

Semantic-tree-based ai content management platform

A data processing system implements receiving a call requesting a generative model to generate a semantic tree for a source content; constructing a first prompt including the source content and instructions to the model to analyze a semantic structure of the source content and to generate a semantic outline and content chunks of the source content, the semantic outline including one or more topics each connected with one or more of the content chunks, to compute one summary for each of the content chunks, to apply indices to reference each topic node of the semantic tree to one of the topics, and to apply indices to reference each leaf node of the semantic tree to one of the content chunks and the respective summary; providing the first prompt to the model and receiving the semantic tree of the source content; and storing the semantic tree in a database.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Interactive patent visualization systems and methods

An interactive, dynamic GUI for visualization of patent documents including content-dense graphics illustrating the number, content size, type of a multiplicity of patent documents (issued or granted patent versus published pending application), distributed over time, with comparison to similar patent documents, market events, and expert insights based upon content of specification or detailed description and claims, all within a predetermined technology sector having at least one sub-sector or category within the technology sector.
Owner:SPORE INC

Generating probabilistic data structures for lookup tables in computer memory for multi-token searching

Methods, systems, and non-transitory computer readable storage media are disclosed for optimizing computer memory usage for lookup lists in computer memory via probabilistic data structures. For example, the disclosed system generates a probabilistic data structure (e.g., a Bloom filter) to represent data in a lookup list including multi-token items by hashing items of the lookup list to sets of bit values in a bit vector. The disclosed system classifies text content in a digital document by utilizing a maximum number of tokens from multi-token items in the lookup list to select and compare sets of sequential tokens in the digital document to the probabilistic data structure. The disclosed system also iteratively reduces the number of tokens in sets of sequential tokens for subsequent comparisons. Furthermore, in some aspects, the disclosed system causes a computing device to modify a digital document and / or database operations based on the classifications.
Owner:ONETRUST LLC