Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

25 results about "Document summary" patented technology

Knowledge graph multi-mode document analysis and image table semantization knowledge recall method

The invention discloses a knowledge graph multi-modal document analysis and image table semantization knowledge recall method, and belongs to the technical field of knowledge engineering and information retrieval. The invention provides an innovative scheme for fusing a visual language model, semantic abstract generation and knowledge graph modeling. The method comprises the following steps: constructing a vertical domain knowledge graph by adopting a BERT-BiLSTM-CRF model; according to the method, multi-modal document analysis is realized through models such as DocLayout-YOLO, TableMaster, UniMERNet and the like; the method comprises the following steps of: segmenting an image into 16 * 16 block sequences by adopting a vit-gpt2-image-adaptation model, and realizing image semantization through 768-dimensional vector space mapping and Transform coding; constructing a document summary tree based on DBSCAN clustering and LLM recursive summary; and designing a hybrid retrieval space fusing semantic vectors and structured vectors, and reordering by adopting a double-attention mechanism. According to the method, the knowledge base document retrieval recall rate is increased to 99%, the question and answer accuracy rate reaches 90% or above, the index construction time is shortened by 60%, and the problem that semantic understanding and recall of non-text elements in complex documents are difficult is effectively solved.
Owner:云鼎科技股份有限公司

Text governance method and device, equipment and medium

The invention relates to the technical field of data processing, and discloses a text governance method and device, equipment and a medium, and the method comprises the steps: collecting a preset enterprise external rule file in real time through a preset double mechanism of a timing scheduling task and an event triggering task; screening out files which do not accord with a preset category, labeling the files according to level labels, and storing the files into an internal database to form an internal specification file; extracting an internal specification data abstract, calculating a correlation coefficient between each service line label and the abstract, and determining a service line corresponding to the internal specification; constructing a knowledge graph by taking the file, the abstract, the label and the level as nodes and taking the correlation coefficient as an edge; when external specification changes are monitored, change content is collected, corresponding internal specification target files are identified, service labels with correlation coefficients exceeding a threshold value are screened to form a set, and corresponding service line internal specifications are collected; and by taking the change external specification level as a level, updating the internal specification corresponding to the service label with the level lower than the level in the set according to the change content. The accuracy of text governance can be improved.
Owner:CHINA MERCHANTS FINANCE HLDG CO LTD

Dynamic Query Classification and Routing in Multi-Model AI Architectures

The present disclosure pertains to dynamic query classification and routing in multi-model AI (Artificial Intelligence) architectures. A method includes receiving a user query at a computing device, generating an embedding representation of the query using a pre-trained language model, and applying a trained classifier to the embedding representation to determine the query's type. The classifier categorizes the query into one of three types: information retrieval, document retrieval, or document summary. A response is then generated based on the classified query type, enhancing the system's ability to provide relevant and accurate results to users.
Owner:OPEN TEXT CORPORATION

Predictive multi-modal content retrieval and display processes

Described herein are examples of a system comprising a processor and memory storing instructions to execute a predictive multi-modal retrieval subsystem configured to analyze digital content, extract metadata, predict information needs, and retrieve relevant content; a multi-document viewing subsystem configured to index documents, establish relationships, and present aggregated content in a unified interface; a data management subsystem configured to organize files, generate folder structures, and provide interactive navigation; a file summary LLM configured to process documents and generate summaries with metadata; and a user context blob configured to maintain user context data across sessions. The system includes a method for receiving digital content, processing through OCR to extract text, analyzing with the file summary LLM to generate metadata and summaries, predicting information needs by simulating task progression and identifying related documents, and presenting retrieved content through a unified interface.
Owner:FILELASSO INC

Bid document abstract generation method and system based on large language model

The invention provides a bidding file abstract generation method and system based on a large language model, and relates to the technical field of natural language process.The method comprises the steps that firstly, domain terms are extracted from an input to-be-processed bidding file set to construct a term library, and file units are split to form a domain semantic chain; inputting the term library into a pre-trained large language model to adjust the semantic analysis weight so as to obtain a semantic analysis model adaptive to the field; inputting semantic nodes in the semantic chain into the semantic analysis model to extract core expressions so as to form a dynamic abstract fragment set; determining an abstract fragment association sequence according to a semantic node connection relationship, and joining to obtain a preliminary fusion abstract text; and finally, semantic feedback of the reader is received, abstract fragment expressions and sequences are adjusted, a semantic chain is updated, and a final abstract is generated and output, so that the high-quality bidding document abstract can be efficiently generated.
Owner:NEW COMM INVESTMENT (CHENGDU) BIG DATA CO LTD

Text processing device, text processing method, and program

The summary of a document describing an incident related to a specific information security vulnerability should include information about that incident. [Solution] The extraction unit extracts damage information, including keywords related to the content of the damage case, from a document describing an example of damage caused by a specific vulnerability in information security. The generation unit generates an explanation of the damage case using the damage information extracted from the document. The output unit outputs information based on the generated explanation.
Owner:NEC CORP

Generating predicted document summary-consistency metrics using machine learning models and an expanding granularity analysis

The present disclosure is directed toward systems, methods, and non-transitory computer readable media that generate a preliminary predicted document-summary consistency portraying elements from a text prompt utilizing a generation diffusion model and refine the preliminary predicted document-summary consistency to generate a predicted document-summary consistency. In particular, the disclosed systems receive, via an interaction with a user device, a text prompt specifying elements to portray within a predicted document-summary consistency. Furthermore, the disclosed systems generate an image generation prompt from the text prompt. Moreover, the disclosed systems utilize the generation diffusion model to generate a preliminary predicted document-summary consistency depicting the elements from the text prompt. In addition, the disclosed systems refine the preliminary predicted document-summary consistency to generate the predicted document-summary consistency.
Owner:ADOBE INC

Multi-source document clustering method, system and equipment based on machine learning

The invention belongs to the technical field of natural language processing, particularly relates to a multi-source document clustering method, system and equipment based on machine learning, and aims to solve the problems of weak representation capability, poor semantic consistency, insufficient label interpretation and the like in document clustering. The method comprises the steps of obtaining a plurality of documents to be clustered; taking the document and the target document abstract word number as input data of a first language model, and generating a document abstract according with the target document abstract word number through the first language model; performing semantic analysis on the document abstract, and converting the document abstract into an embedded vector based on a semantic analysis result; dividing the document into a plurality of document clusters based on the similarity between the embedding vector of any document and the embedding vectors of other documents; generating a clustering label for representing a document clustering semantic topic; and establishing a document set under the same document cluster according to the cluster labels, and displaying the document set in a structured view. According to the scheme, the cross-domain document set can be effectively normalized and clustered.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Systems and methods for risk factor predictive modeling with document summarization

A system and method for document summarization generates summarized articles and risk factor categorizations for display at a graphical user interface (GUI) dashboard. A transaction monitoring system includes an adverse media dashboard pipeline for processing risk factor alerts and generating document summarizations for display at a user device. Document summarization extracts several sentences from a source text and stacks the sentences to create a summary. The method creates a vector representation of each sentence using a machine learning word embedding model and generates a sentence similarity matrix by computing cosine similarity values. A sentence graph creation algorithm creates a graph corresponding to the sentence similarity matrix and calculates importance scores used in selecting sentences for the document summary. The GUI dashboard includes first, second, third and fourth dashboard regions for displaying alert report records, media records, and graphical user interface layouts of document summaries and risk factor visualizations.
Owner:BANK OF MONTREAL

Literature summary text generation method, sample generation method and related devices

The invention provides a document summary text generation method, a sample generation method and related devices. The method comprises the steps of obtaining a plurality of target text blocks arranged according to a text sequence for a received literature; wherein each target text block comprises at least one text statement; generating a summary outline according to the target text block summary; wherein the summary outline comprises a plurality of titles; the title corresponds to a text block identifier; the text block identifier is used for identifying a text block; respectively generating summary sub-texts corresponding to a plurality of titles in the summary outline to form a document summary text of the document; wherein the summary sub-text is generated according to the text block corresponding to the corresponding title. The content accuracy of the literature summary text can be improved to a certain degree.
Owner:ALI HEALTH TECH CO LTD

A multi-document summarization generation method, system and device

This invention provides a method, system, and device for generating multi-document summaries, relating to the field of computer technology. The method includes: employing an improved Siamese neural network to evaluate the relevance scores between document paragraphs and headings in a multi-document context; combining a hierarchical clustering deduplication method based on relevance scores to extract a subset of paragraphs from the document paragraph set whose importance is greater than a set importance level and whose redundancy is less than a set redundancy level; and processing the paragraph subset using a multi-level Transformer-based summarization device to generate multi-document summary text. The multi-level Transformer-based summarization device has cross-document learning capabilities. This invention can generate concise, high-quality multi-document summary text.
Owner:MINZU UNIVERSITY OF CHINA

Large model retrieval enhancement generation method based on document density soft clustering

The invention provides a large model retrieval enhancement generation method based on document density soft clustering. The method comprises the following steps: generating document abstracts and embedded vector sets of all documents by utilizing a large language model based on an existing document library; clustering the existing documents by using an improved density soft clustering algorithm to obtain a document cluster; obtaining a text block embedding vector set of an existing document based on semi-overlapping sliding window division and a vector model, clustering text blocks of each document cluster by adopting a k-means clustering algorithm, obtaining a plurality of corresponding text block clusters, and obtaining a hierarchical retrieval library; obtaining a user question, and generating a pseudo answer vector set by using at least two large language models; performing first-level matching, second-level matching and third-level matching in sequence by utilizing the hierarchical retrieval library, and screening out a target text block; and carrying out merging and deduplication processing on the target text blocks, and generating a final answer by utilizing a large language model based on the optimal target text blocks and the user questions. The problem that the retrieval precision is influenced by unstable clustering caused by Gaussian distribution hypothesis is solved.
Owner:ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY

Document processing method and related device

PCT designated stageWO2026012150A1Semantic analysisText processingDocument summarizationLinguistic model
Disclosed are a document processing method and a related device, relating to the field of artificial intelligence. The method comprises: acquiring a document processing request; and acquiring a first fragment set of a first document according to the document processing request, and processing the first fragment set to obtain a summary result of the first document. The first fragment set is from a part of a fragment graph network constructed from a plurality of fragments of the first document, and the fragment graph network is configured for indicating an association relationship between the plurality of fragments of the first document. Therefore, performing document summarization by selecting from the fragment graph network a small number of representative fragments capable of expressing content of the first document ensures document summary quality, reduces the fragments input into a large language model, shortens the duration for processing long documents, and improves processing efficiency.
Owner:HUAWEI TECH CO LTD

A method and system for generating document content guide based on region tagging

The application discloses a document content guide generation method and system based on region annotation, and belongs to the field of artificial intelligence technology. The method comprises the following steps: reading a document submitted by a user; marking a region that needs to be marked; extracting context content above and below the marked region in the document, and inputting the marked region, the context content and a preset prompt word into a large language model to obtain an annotation of the marked region; storing the manually added annotation and / or the annotation generated according to the large language model to generate an annotation list; repeating the above steps until all the regions to be marked in the document are added with annotations after being marked, and the added annotations are stored in the annotation list; and generating a document content guide according to the annotation list. The application can make the large language model accurately capture important parts of an article through annotation information, so that the powerful language understanding capability of the large language model can be better utilized to generate an accurate document summary.
Owner:SHIP INFORMATION RES CENT (NO 714 RES INST OF CHINA STATE SHIPBUILDING CORP)

Knowledge search method and device based on intelligent agent, equipment, medium and program

The embodiment of the invention provides an agent-based knowledge search method, device, equipment, medium and program, the method is executed by a knowledge search agent, and the knowledge search agent comprises a tool determination module, a search tool, a result summarization model and a knowledge base. The method comprises the steps that when a tool determination module determines to use a search tool according to current-round input information, search parameters corresponding to the search tool are obtained according to the current-round input information, the first-round input information comprises original search terms, and non-first-round input information comprises the original search terms and document abstracts of documents inquired before the current round; the search tool obtains a search result of the round from a knowledge base and / or external data according to the search parameters; when it is determined that the search tool is not used, the result summarization model generates a target search result according to the document content of the document inquired before the round and the original search word. According to the method, the search time delay and the reasoning cost of the knowledge search agent are reduced.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Regulation intelligent question answering method and device based on two-stage retrieval and large model

The application discloses a kind of rules and regulations intelligent question and answer method and device based on two-stage search and big model, belongs to the intelligent question and answer field based on rules and regulations file, including the following steps: file abstract extraction is carried out to each file based on big model question and answer form, abstract vector library is constructed;Again, all files are cut and blocked, and all text blocks are vectorized using sparse vector model and dense vector model, and content vector library is constructed;The most relevant file to the user's question is retrieved in the first stage, and the two vector libraries are retrieved using a hybrid retrieval method.The top K most relevant files are obtained by weighting calculation of the two retrievals;The most relevant text block to the user's question is retrieved in the second stage, and the top k most relevant text blocks are obtained by hybrid retrieval;When answering, the top k most relevant text blocks are used as reference text, and the big model answers based on the reference text.The application can generate more accurate answers.
Owner:10TH RES INST OF CETC

Bid document abstract generation method and system based on large language model

The application provides a bidding document abstract generation method and system based on a large language model, and relates to the technical field of natural language processing. First, the inputted to-be-processed bidding document set is extracted to construct a term library, and the document unit is split to form a domain semantic chain. Then, the term library is inputted into a pre-trained large language model to adjust the semantic parsing weight, so as to obtain a domain-adapted semantic parsing model. Then, the semantic nodes in the semantic chain are inputted into the semantic parsing model to extract core expressions and form a dynamic abstract fragment set. Then, the abstract fragment association order is determined according to the connection relationship of the semantic nodes, and a preliminary fusion abstract text is obtained. Finally, the reader semantic feedback is received to adjust the abstract fragment expression and order, the semantic chain is updated to generate a final abstract, and the final abstract is outputted, so that a high-quality bidding document abstract can be efficiently generated.
Owner:NEW COMM INVESTMENT (CHENGDU) BIG DATA CO LTD

Intelligent document creation and review generated by a large language model

The systems, methods, and non-transitory computer-readable media relates to the intelligent generation and completion of digital documents. For example, the disclosed systems provide a document creation interface for entering inputs to generate and modify digital documents. In some instances, the disclosed systems receive a document generation prompt, generate a digital document (e.g., using a large language model), and provide an indication of the generated digital document within the document creation interface. Moreover, the disclosed systems can further generate a document summary of the digital document and provide the document summary for a recipient device via a document review interface. Additionally, the disclosed systems can further generate a suggested document modification element and modify the digital document in response to a user interaction with the suggested document modification element using the large language model.
Owner:DROPBOX INC

Legal document retrieval method and system based on multi-granularity index and hierarchical sorting

PendingCN122309706AData miningQuery statement
This invention provides a legal document retrieval method and system based on multi-granularity indexing and hierarchical ranking, relating to the field of legal artificial intelligence technology. The method includes: receiving a user-uploaded original query statement; rewriting the original query statement into multiple sub-queries; determining the target recall results of each sub-query statement relative to multiple index categories based on a pre-built multi-indexed legal document database; aggregating the target recall results of all sub-queries relative to multiple index categories into a candidate key statement set; determining the target relevance score between the original query statement and the document summary associated with each candidate key statement in the candidate key statement set using a pre-trained legal relevance analysis model combined with multi-dimensional business characteristics; and determining the query result corresponding to the original query statement based on the target relevance score. This invention significantly improves the accuracy, robustness, and legal interpretability of legal document retrieval.
Owner:BEIJING MEGA INTELLIGENT TECHNOLOGY CO LTD

Text abstract generation method and device, and storage medium

This application discloses a text summarization method, device, and storage medium, relating to the field of data processing technology. The method vectorizes each paragraph of text in a received document to obtain a paragraph text vector; based on a local entity library, it extracts entity words and relationships between them from each paragraph to construct a local knowledge graph; it determines the embedding vector of the local knowledge graph using a pre-trained graph neural network, and concatenates the embedding vector and the paragraph text vector to obtain a fusion vector; based on the similarity between the fusion vector and the local text vector, it determines the text type corresponding to each paragraph, and clusters all paragraphs in the document based on the text type to obtain clustered text clusters; based on the clustered text clusters corresponding to the document, it generates a document summary, ensuring the completeness of the summary's description of the core topic.
Owner:SHENZHEN ZHUOXUN INFORMATION TECH CO LTD

A document abstract optimization generation method and system

PendingCN122262325ASemantic analysisBiological modelsDocument summarizationEngineering
This invention proposes a document summarization optimization generation method and system, belonging to the field of data processing technology. The method includes: based on a historical document dataset, segmenting sentence sequences and constructing a document graph structure; combining a graph neural network to obtain target node features; then combining a contrastive learning model to construct positive and negative sample pairs and calculate the contrastive loss; inputting the target node features into a preset large model to generate a candidate summary sequence; constructing a joint objective optimization function and optimizing it to obtain an optimized graph neural network, an optimized contrastive learning model, and an initial optimized large model; performing knowledge distillation on the initial optimized large model to obtain a target optimized large model; and generating an optimized document summary from the document to be processed using the optimized model. This invention achieves structured processing of documents and fully utilizes structured information through joint optimization and knowledge distillation, fully considering the logical structure of document sentences and the deep semantic relationships between sentences to generate accurate document summaries.
Owner:GUANGDONG POWER GRID CO LTD

Document processing method, document summary generation method and device

The application provides a document processing method, a document abstract generation method and device. The document processing method comprises: obtaining a document set to be processed and a keyword set; inserting keywords in the keyword set into each document to be processed in the document set to be processed respectively to obtain a sequence to be tested; determining the perplexity of each sequence to be tested, and determining a first score result of each document to be processed based on the perplexity of each sequence to be tested; screening the document set to be processed based on the first score result of each document to be processed to obtain a target document. The document abstract generation method comprises: extracting the target document from the document set to be processed based on each keyword in the keyword set; and generating an abstract based on the target document. The application can effectively improve the effectiveness of the target document, thereby ensuring the generation effect of the abstract.
Owner:TSINGHUA UNIVERSITY

Textual summaries in information systems based on personalized prior knowledge

ActiveUS12670198B2PersonalizationIndividual knowledge
Examples of the present disclosure describe systems and methods for providing textual summaries based on personalized prior knowledge. In examples, a user request for a summary of document or an entity is received. The document or documents associated with the entity are separated into segments and semantic embeddings are created for each segment. The semantic embeddings, a context of the user request, and relevant information available in the requesting user's previous knowledge base are provided as input to a personal knowledge system. Based on the input, the personal knowledge system outputs an indication of the segments that should be summarized. The indicated segments are provided to a summarization system. The summarization system generates a document summary or an entity summary and provides the summary to the user.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Collaborative dialogue driven specification and software automatic co-construction method and device

ActiveCN121935357ASemantic analysisVersion controlCode generationSpecification document
The invention discloses a specification and software automatic co-construction method and device based on collaborative dialogue driving. The method comprises the following steps: step S1, conference content capturing and semantic processing; s2, automatically generating and updating a specification and a project intermediate file; s3, constructing a file abstract and an index; step S4, code generation; and S5, project version iteration is carried out. The conference discussion can be automatically converted into the specification document and the project code, a complete assembly line from the conference to software implementation is formed, the software development efficiency is remarkably improved, the consistency of the specification and the code is kept, the manual maintenance workload is reduced, and the project knowledge utilization efficiency is improved.
Owner:HANGZHOU RAPID INTELLIGENT TECHNOLOGY CO LTD

Model adjustment method and device, and data processing method

PendingCN122334393AReward valueData mining
This application provides a model adjustment method and apparatus, relating to the field of artificial intelligence technology. The model adjustment method includes: obtaining training data; the training data includes sample documents and a summary of sample documents; inputting the sample documents into an initial model to obtain a predicted document summary; updating the parameters of the initial model based on the predicted document summary, the sample document summary, and a reward function to obtain a target model; the reward function includes at least a first reward function; the reward value of the first reward function represents the keyword matching degree between the predicted document summary and the sample document summary.
Owner:LENOVO (BEIJING) LTD