Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

43 results about "Document summary" patented technology

Knowledge retrieval method based on multilayer index

ActiveCN120633843ASemantic analysisKnowledge representationKnowledge classificationLinguistic model
The invention discloses a knowledge retrieval method based on a multilayer index, which comprises the following steps of: obtaining a knowledge base of corresponding classification according to a knowledge retrieval statement in combination with a knowledge classification label; obtaining a plurality of retrieval target documents through vector similarity analysis according to the knowledge retrieval statements in combination with the document abstracts in the knowledge base of the corresponding classification; and obtaining a plurality of similar text blocks through vector similarity analysis according to the knowledge retrieval statement in combination with the text blocks of the plurality of retrieval target documents, and outputting the similar text blocks as knowledge retrieval. According to the method, a three-layer index structure containing knowledge classification tags, document abstracts and text blocks is constructed, the classification tags are dynamically matched through a large language model to reduce the retrieval range, high-correlation documents are rapidly positioned in combination with document abstract vector similarity analysis, and context logic relations are reserved based on text block hierarchical vector matching. The problems of fuzzy classification, low efficiency of document recall and context breakage in the traditional technology are solved, and accurate retrieval from macroscopic to microscopic is realized.
Owner:INSPUR GENERSOFT CO LTD

CAD file summary information automatic extraction method and device, medium and product

The invention provides a CAD file summary information automatic extraction method and device, a medium and a product, and relates to the technical field of computer picture word process.The method comprises the steps that a target CAD file is analyzed through a CAD file analysis development tool, and first word information and geometric attribute information corresponding to the first word information are obtained; retrieving the first character information according to the retrieval keyword to obtain second character information; according to the corresponding geometric attribute information in the second text information, carrying out positioning adjustment display through a CAD file analysis development tool to obtain a picture file; calling a first large language model to perform character recognition and arrangement processing on the picture to form a text list; and calling the second large language model to obtain summary information of the target CAD file according to a preset prompt template and the text list. According to the scheme, when a large batch of CAD files are processed, the processing efficiency is remarkably improved.
Owner:SUZHOU XINKUAIZHUANG TECH CO LTD

File disassembling method based on multi-dimensional file association analysis and automatic interpretation

The invention discloses a file disassembling method based on multi-dimensional file association analysis and automatic interpretation, and relates to the crossing field of natural language processing and big data analysis. Comprising the steps of (1) carrying out multi-source heterogeneous data collection, (2) carrying out intelligent cleaning on the collected data and constructing a knowledge graph, (3) carrying out deep correlation analysis, (4) carrying out automatic disassembly and file simulation, and (5) carrying out intelligent application and feedback: automatically generating a Markdown document containing a file abstract, clause comparison and an enterprise portrait through a structured report engine, and carrying out intelligent application and feedback on the Markdown document. GIS map overlay analysis is supported, clauses in an applicable file are matched based on an enterprise tag system, declaration materials are pre-checked through an OCR + rule engine, user feedback cases are collected, model parameters are incrementally trained monthly, it is ensured that adaptation is completed within 24 hours after a new file is released, and a collection-analysis-application-optimization closed-loop iteration mechanism is formed.
Owner:天元大数据信用管理有限公司

Purchase review document abstract generation method and system in combination with AI large model

The invention provides a purchase review document abstract generation method and system combined with an AI large model, and relates to the technical field of large language models.The method comprises the steps that firstly, a set containing a plurality of purchase review document units with review time stamps is obtained, and the content theme feature and the review logic feature of each purchase review document unit are extracted; calling a pre-trained AI large model to carry out combined abstract generation processing on the content theme features and the review logic features to generate a preliminary abstract text, then carrying out semantic coherence optimization processing on the preliminary abstract text to obtain a semantic coherence optimized abstract text, and finally, carrying out semantic coherence optimization processing on the semantic coherence optimized abstract text. And generating a final procurement document abstract containing the theme keyword set based on the optimized abstract text, and outputting the final procurement document abstract to the target storage terminal according to a preset format, so that the procurement document abstract is automatically, efficiently and accurately generated, and the document processing efficiency and the information utilization value are improved.
Owner:SICHUAN XIXING ELECTRIC POWER TECHNOLOGY CONSULTING CO LTD

Knowledge graph multi-mode document analysis and image table semantization knowledge recall method

The invention discloses a knowledge graph multi-modal document analysis and image table semantization knowledge recall method, and belongs to the technical field of knowledge engineering and information retrieval. The invention provides an innovative scheme for fusing a visual language model, semantic abstract generation and knowledge graph modeling. The method comprises the following steps: constructing a vertical domain knowledge graph by adopting a BERT-BiLSTM-CRF model; according to the method, multi-modal document analysis is realized through models such as DocLayout-YOLO, TableMaster, UniMERNet and the like; the method comprises the following steps of: segmenting an image into 16 * 16 block sequences by adopting a vit-gpt2-image-adaptation model, and realizing image semantization through 768-dimensional vector space mapping and Transform coding; constructing a document summary tree based on DBSCAN clustering and LLM recursive summary; and designing a hybrid retrieval space fusing semantic vectors and structured vectors, and reordering by adopting a double-attention mechanism. According to the method, the knowledge base document retrieval recall rate is increased to 99%, the question and answer accuracy rate reaches 90% or above, the index construction time is shortened by 60%, and the problem that semantic understanding and recall of non-text elements in complex documents are difficult is effectively solved.
Owner:云鼎科技股份有限公司

AI-based interactive application security test vulnerability studying and judging method and system

The invention discloses an interactive application security test vulnerability study and judgment method and system based on AI, and relates to the field of electric digital data processing, and the method comprises the steps: carrying out the instrumentation of a key function of a target system, and collecting execution environment information; and identifying user input from HTTP request data in the execution environment information, and tracking a propagation path of the user input to obtain a polluted data propagation chain with a vulnerability type mark. And packaging the execution environment information, the pollution data propagation chain, the vulnerability type mark and the related function call stack into a to-be-analyzed vulnerability information packet, and sending a system description document and the vulnerability information packet combination to a large language model to obtain a document summary. And the document summary and the pollution data initial entry point information are combined and sent to a model, and a context supplement request is acquired. And sending the supplement request and a preset vulnerability bypassing case combination to a model to obtain a research and judgment result. And finally updating the research and judgment result to the vulnerability detail information. By implementing the method, the verification cost of security test vulnerability research and judgment can be reduced.
Owner:HANGZHOU XIAODAO TECH CO LTD

Text governance method and device, equipment and medium

The invention relates to the technical field of data processing, and discloses a text governance method and device, equipment and a medium, and the method comprises the steps: collecting a preset enterprise external rule file in real time through a preset double mechanism of a timing scheduling task and an event triggering task; screening out files which do not accord with a preset category, labeling the files according to level labels, and storing the files into an internal database to form an internal specification file; extracting an internal specification data abstract, calculating a correlation coefficient between each service line label and the abstract, and determining a service line corresponding to the internal specification; constructing a knowledge graph by taking the file, the abstract, the label and the level as nodes and taking the correlation coefficient as an edge; when external specification changes are monitored, change content is collected, corresponding internal specification target files are identified, service labels with correlation coefficients exceeding a threshold value are screened to form a set, and corresponding service line internal specifications are collected; and by taking the change external specification level as a level, updating the internal specification corresponding to the service label with the level lower than the level in the set according to the change content. The accuracy of text governance can be improved.
Owner:CHINA MERCHANTS FINANCE HLDG CO LTD

Dynamic Query Classification and Routing in Multi-Model AI Architectures

The present disclosure pertains to dynamic query classification and routing in multi-model AI (Artificial Intelligence) architectures. A method includes receiving a user query at a computing device, generating an embedding representation of the query using a pre-trained language model, and applying a trained classifier to the embedding representation to determine the query's type. The classifier categorizes the query into one of three types: information retrieval, document retrieval, or document summary. A response is then generated based on the classified query type, enhancing the system's ability to provide relevant and accurate results to users.
Owner:OPEN TEXT CORPORATION

Predictive multi-modal content retrieval and display processes

Described herein are examples of a system comprising a processor and memory storing instructions to execute a predictive multi-modal retrieval subsystem configured to analyze digital content, extract metadata, predict information needs, and retrieve relevant content; a multi-document viewing subsystem configured to index documents, establish relationships, and present aggregated content in a unified interface; a data management subsystem configured to organize files, generate folder structures, and provide interactive navigation; a file summary LLM configured to process documents and generate summaries with metadata; and a user context blob configured to maintain user context data across sessions. The system includes a method for receiving digital content, processing through OCR to extract text, analyzing with the file summary LLM to generate metadata and summaries, predicting information needs by simulating task progression and identifying related documents, and presenting retrieved content through a unified interface.
Owner:FILELASSO INC

Financial bill and bill full-automatic matching method and device, equipment and medium

The invention discloses a full-automatic matching method, device, equipment and medium for financial bills and receipts under a fuzzy condition, and the method comprises the steps: obtaining financial bill information and receipt information, and carrying out the feature extraction of the financial bill information and the receipt information, obtaining a bill characteristic value corresponding to each bill and a bill characteristic value corresponding to each bill; according to each document feature value, generating a document abstract corresponding to each document; according to the bill feature values, the bill feature values and the bill abstracts, matching the bills with the bills one by one, and outputting a first matching result; and if the unmatched bills and the unmatched bills exist, performing fuzzy matching according to the corresponding bill feature values and the corresponding bill feature values, and outputting a second matching result. According to the method, a two-step matching method is introduced, so that the accuracy of full-automatic matching of financial bills and receipts under fuzzy conditions is improved.
Owner:HUNAN IRON & STEEL GRP TECH RES INST CO LTD

Document processing method, device, electronic device and storage medium

The present disclosure provides a document processing method, apparatus, electronic device, and storage medium, relating to the field of data processing technology. A specific implementation scheme comprises: obtaining a document collection, the document collection including multiple candidate documents, and determining segmented paragraphs and document summaries for the candidate documents; obtaining paragraph word vectors for the segmented paragraphs and summary word vectors for the document summaries, generating a vector index based on the paragraph word vectors and summary word vectors, and storing the index in a search engine cluster; and generating an inverted index based on the candidate documents and document summaries, and storing the index in the search engine cluster.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Bid document abstract generation method and system based on large language model

The invention provides a bidding file abstract generation method and system based on a large language model, and relates to the technical field of natural language process.The method comprises the steps that firstly, domain terms are extracted from an input to-be-processed bidding file set to construct a term library, and file units are split to form a domain semantic chain; inputting the term library into a pre-trained large language model to adjust the semantic analysis weight so as to obtain a semantic analysis model adaptive to the field; inputting semantic nodes in the semantic chain into the semantic analysis model to extract core expressions so as to form a dynamic abstract fragment set; determining an abstract fragment association sequence according to a semantic node connection relationship, and joining to obtain a preliminary fusion abstract text; and finally, semantic feedback of the reader is received, abstract fragment expressions and sequences are adjusted, a semantic chain is updated, and a final abstract is generated and output, so that the high-quality bidding document abstract can be efficiently generated.
Owner:NEW COMM INVESTMENT (CHENGDU) BIG DATA CO LTD

Judgment document summary generation method based on three-stage GRPO reinforcement learning

The present invention provides a method for generating judicial document summaries based on three-stage GRPO reinforcement learning, which belongs to the field of data processing technology. The method specifically includes: step 1, modeling a three-stage thinking chain; step 2, performing data distillation and stratification on the original judicial document dataset according to the three-stage thinking chain to obtain different types of datasets, wherein the types include high correlation, medium correlation, and low correlation; step 3, using the high-correlation dataset to perform SFT supervised fine-tuning training on the large language model; step 4, using the entire dataset to perform multi-stage GRPO reinforcement learning training on the trained large language model to obtain a target model; step 5, inputting the target judicial document into the target model to generate a target summary. The solution of the present invention improves the efficiency, accuracy, and adaptability of summary generation.
Owner:湖南工商大学

Text processing device, text processing method, and program

The summary of a document describing an incident related to a specific information security vulnerability should include information about that incident. [Solution] The extraction unit extracts damage information, including keywords related to the content of the damage case, from a document describing an example of damage caused by a specific vulnerability in information security. The generation unit generates an explanation of the damage case using the damage information extracted from the document. The output unit outputs information based on the generated explanation.
Owner:NEC CORP

Intelligent interaction method and system for document

The application discloses a kind of intelligent interaction method and system of document.The method includes receiving the document uploaded by user, generates document list;Receive the reading request of target document in the document list by user, generate conversational reading interaction interface, the conversational reading interaction interface includes tool area, document outline area, document current reading area and conversation window area;The document summary of the target document is generated and shown in the conversation window area by document assistant, and the document assistant includes ChatGPT natural language processing model;The content selection of the target document in the document current reading area by user is monitored in real time and the question of user is obtained;The answer of the question is generated by the document assistant and is shown in the conversation window area.The application provides a kind of document reading scheme that can be intelligently interacted, improves user reading experience, and can better meet the reading demand of user.
Owner:SHENZHEN MAIFENG TECH CO LTD

Generating predicted document summary-consistency metrics using machine learning models and an expanding granularity analysis

The present disclosure is directed toward systems, methods, and non-transitory computer readable media that generate a preliminary predicted document-summary consistency portraying elements from a text prompt utilizing a generation diffusion model and refine the preliminary predicted document-summary consistency to generate a predicted document-summary consistency. In particular, the disclosed systems receive, via an interaction with a user device, a text prompt specifying elements to portray within a predicted document-summary consistency. Furthermore, the disclosed systems generate an image generation prompt from the text prompt. Moreover, the disclosed systems utilize the generation diffusion model to generate a preliminary predicted document-summary consistency depicting the elements from the text prompt. In addition, the disclosed systems refine the preliminary predicted document-summary consistency to generate the predicted document-summary consistency.
Owner:ADOBE INC

Multi-source document clustering method, system and equipment based on machine learning

The invention belongs to the technical field of natural language processing, particularly relates to a multi-source document clustering method, system and equipment based on machine learning, and aims to solve the problems of weak representation capability, poor semantic consistency, insufficient label interpretation and the like in document clustering. The method comprises the steps of obtaining a plurality of documents to be clustered; taking the document and the target document abstract word number as input data of a first language model, and generating a document abstract according with the target document abstract word number through the first language model; performing semantic analysis on the document abstract, and converting the document abstract into an embedded vector based on a semantic analysis result; dividing the document into a plurality of document clusters based on the similarity between the embedding vector of any document and the embedding vectors of other documents; generating a clustering label for representing a document clustering semantic topic; and establishing a document set under the same document cluster according to the cluster labels, and displaying the document set in a structured view. According to the scheme, the cross-domain document set can be effectively normalized and clustered.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Systems and methods for risk factor predictive modeling with document summarization

A system and method for document summarization generates summarized articles and risk factor categorizations for display at a graphical user interface (GUI) dashboard. A transaction monitoring system includes an adverse media dashboard pipeline for processing risk factor alerts and generating document summarizations for display at a user device. Document summarization extracts several sentences from a source text and stacks the sentences to create a summary. The method creates a vector representation of each sentence using a machine learning word embedding model and generates a sentence similarity matrix by computing cosine similarity values. A sentence graph creation algorithm creates a graph corresponding to the sentence similarity matrix and calculates importance scores used in selecting sentences for the document summary. The GUI dashboard includes first, second, third and fourth dashboard regions for displaying alert report records, media records, and graphical user interface layouts of document summaries and risk factor visualizations.
Owner:BANK OF MONTREAL

Literature summary text generation method, sample generation method and related devices

The invention provides a document summary text generation method, a sample generation method and related devices. The method comprises the steps of obtaining a plurality of target text blocks arranged according to a text sequence for a received literature; wherein each target text block comprises at least one text statement; generating a summary outline according to the target text block summary; wherein the summary outline comprises a plurality of titles; the title corresponds to a text block identifier; the text block identifier is used for identifying a text block; respectively generating summary sub-texts corresponding to a plurality of titles in the summary outline to form a document summary text of the document; wherein the summary sub-text is generated according to the text block corresponding to the corresponding title. The content accuracy of the literature summary text can be improved to a certain degree.
Owner:ALI HEALTH TECH CO LTD

A multi-document summarization generation method, system and device

This invention provides a method, system, and device for generating multi-document summaries, relating to the field of computer technology. The method includes: employing an improved Siamese neural network to evaluate the relevance scores between document paragraphs and headings in a multi-document context; combining a hierarchical clustering deduplication method based on relevance scores to extract a subset of paragraphs from the document paragraph set whose importance is greater than a set importance level and whose redundancy is less than a set redundancy level; and processing the paragraph subset using a multi-level Transformer-based summarization device to generate multi-document summary text. The multi-level Transformer-based summarization device has cross-document learning capabilities. This invention can generate concise, high-quality multi-document summary text.
Owner:MINZU UNIVERSITY OF CHINA

A multi-document summary generation method for crisis help information

The present invention discloses a multi-document summary generation method for crisis help information; the method is divided into: an extractive summary stage and a generative summary stage. In the extractive summary stage, the sentence graph structure is used to fuse contextual information in a deep neural network, and the subgraph structure is used to optimize the information extraction to effectively extract key information from numerous documents; in the generative summary stage, the powerful sequence generation ability of the BART model and the characteristics of the pointer generation network in accurately copying key information and avoiding duplicate content generation are used to further streamline and summarize the extracted information, and generate an accurate and compact summary of the crisis help information. The method of the present invention is suitable for processing and summarizing massive help information in crisis scenarios. The two-stage method of the present invention can quickly and accurately provide clear information summaries for rescue and emergency management, greatly improving the efficiency and effectiveness of crisis response.
Owner:FUDAN UNIVERSITY

Large model retrieval enhancement generation method based on document density soft clustering

The invention provides a large model retrieval enhancement generation method based on document density soft clustering. The method comprises the following steps: generating document abstracts and embedded vector sets of all documents by utilizing a large language model based on an existing document library; clustering the existing documents by using an improved density soft clustering algorithm to obtain a document cluster; obtaining a text block embedding vector set of an existing document based on semi-overlapping sliding window division and a vector model, clustering text blocks of each document cluster by adopting a k-means clustering algorithm, obtaining a plurality of corresponding text block clusters, and obtaining a hierarchical retrieval library; obtaining a user question, and generating a pseudo answer vector set by using at least two large language models; performing first-level matching, second-level matching and third-level matching in sequence by utilizing the hierarchical retrieval library, and screening out a target text block; and carrying out merging and deduplication processing on the target text blocks, and generating a final answer by utilizing a large language model based on the optimal target text blocks and the user questions. The problem that the retrieval precision is influenced by unstable clustering caused by Gaussian distribution hypothesis is solved.
Owner:ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY

Document summary generation graphical user interface for electronic devices

1. Name of the design product: Graphical user interface for generating document summaries for electronic devices. 2. Purpose of this design product: for use in an electronic device. 3. The key design point of this design product lies in the graphical user interface. 4. The picture or photo that best illustrates the design points: Interface change state diagram 1. 5. The other surfaces of this design product are conventional designs, and the rear view, left view, right view, top view and bottom view are omitted. 6. Purpose of Graphical User Interface: An interactive interface used to generate document summaries. 7. Description of the changing status of the graphical user interface: Click the "Generate summary with one click" button in the main view to display the interface change status diagram 1; click the blue "Generate summary of current document" text in the box in the interface change status diagram 1, or click the "Generate summary" button at the bottom of the interface change status diagram 1 to display the interface change status diagram 2; after the summary in the interface change status diagram 2 is generated, the interface change status diagram 3 is automatically displayed; click the "Re-answer" button in the interface change status diagram 3 to display the interface change status diagram 4. 8. Other situations that require explanation: The blurred areas in each interface are content screens.
Owner:BEIJING QIHOOD TECHNOLOGY CO LTD

Document processing method and related device

PCT designated stageWO2026012150A1Semantic analysisText processingDocument summarizationLinguistic model
Disclosed are a document processing method and a related device, relating to the field of artificial intelligence. The method comprises: acquiring a document processing request; and acquiring a first fragment set of a first document according to the document processing request, and processing the first fragment set to obtain a summary result of the first document. The first fragment set is from a part of a fragment graph network constructed from a plurality of fragments of the first document, and the fragment graph network is configured for indicating an association relationship between the plurality of fragments of the first document. Therefore, performing document summarization by selecting from the fragment graph network a small number of representative fragments capable of expressing content of the first document ensures document summary quality, reduces the fragments input into a large language model, shortens the duration for processing long documents, and improves processing efficiency.
Owner:HUAWEI TECH CO LTD

A method and system for generating document content guide based on region tagging

The application discloses a document content guide generation method and system based on region annotation, and belongs to the field of artificial intelligence technology. The method comprises the following steps: reading a document submitted by a user; marking a region that needs to be marked; extracting context content above and below the marked region in the document, and inputting the marked region, the context content and a preset prompt word into a large language model to obtain an annotation of the marked region; storing the manually added annotation and / or the annotation generated according to the large language model to generate an annotation list; repeating the above steps until all the regions to be marked in the document are added with annotations after being marked, and the added annotations are stored in the annotation list; and generating a document content guide according to the annotation list. The application can make the large language model accurately capture important parts of an article through annotation information, so that the powerful language understanding capability of the large language model can be better utilized to generate an accurate document summary.
Owner:SHIP INFORMATION RES CENT (NO 714 RES INST OF CHINA STATE SHIPBUILDING CORP)

An AI-based interactive application security testing vulnerability research and judgment method and system

The application relates to an AI-based interactive application security test vulnerability judgment method and system, and relates to the field of electric digital data processing.The method comprises the following steps: inserting a plug into a key function of a target system, and collecting execution environment information; identifying user input from HTTP request data in the execution environment information, tracking the propagation path to obtain a pollution data propagation chain with a vulnerability type mark; encapsulating the execution environment information, the pollution data propagation chain, the vulnerability type mark and the related function call stack into a to-be-analyzed vulnerability information package; combining a system specification document with the vulnerability information package and sending the same to a large language model to obtain a document summary; combining the document summary and pollution data initial entry point information and sending the same to the model to obtain a context supplement request; combining the supplement request and a preset vulnerability bypass case and sending the same to the model to obtain a judgment result; and finally updating the judgment result to vulnerability detail information. By implementing the method, the verification cost of security test vulnerability judgment can be reduced.
Owner:HANGZHOU XIAODAO TECH CO LTD

Knowledge archiving methods, apparatus, equipment, storage media, and computer program products

This application discloses a knowledge archiving method, apparatus, device, storage medium, and computer program product, relating to the field of data processing technology. The disclosed knowledge archiving method includes: classifying unstructured target requirement documents by business type using a vector machine algorithm to obtain a first classification result; generating a document summary of the target requirement document using a basic large model, and classifying the target requirement document by business type based on the document summary to obtain a second classification result, wherein the basic large model is a large language model trained using historical requirement documents and historical document summaries; determining the target classification result of the target requirement document based on the first and second classification results; and archiving the target requirement document, document summary, and target classification result in a structured form to a preset knowledge base. This application can efficiently archive unstructured business requirement documents into structured document knowledge.
Owner:CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1

Multi-source information analysis method and system based on deep learning

The invention discloses a multi-source information analysis method based on deep learning, and belongs to the field of network information analysis. According to the method, firstly, multi-document abstract extraction is conducted, sentiment classification is conducted on input documents, the quality and information richness of document abstracts are improved, then information updating of multi-granularity nodes is achieved through a graph attention mechanism, and therefore the problem that a cross-document relation is difficult to model in the prior art is solved. Furthermore, the information popularity in a specific time period is deeply analyzed by combining the popularity analysis of the network main information questions with two aspects of emotional polarity and information popularity, so that the emotional evolution process of the network information can be known, and a basis is provided for scientific decision and active management.
Owner:JIANGSU YIQICE NETWORK TECH CO LTD