Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

23 results about "Document representation" patented technology

RAG intelligent retrieval question-answering system and method based on enhanced metadata

The invention discloses an RAG intelligent retrieval question-answering system and method based on enhanced metadata, and relates to the technical field of information processing and intelligent retrieval, multi-source heterogeneous knowledge data is preprocessed to obtain unified knowledge data, and structured metadata is extracted from the unified knowledge data based on different text forms; vectorizing a document text in the structured metadata by combining with embedding of the knowledge graph to obtain document representation, and outputting the document representation, the structured metadata and the enhanced keyword set as an enhanced metadata object; labeling a display relationship between different enhanced metadata objects, and constructing to obtain a knowledge database; according to the intelligent knowledge service system and method, restrictive conditions and question intentions are extracted from user questions, mixed retrieval is performed from a knowledge database based on the restrictive conditions and the question intentions, a candidate literature semantic set is output, then structured statistical visualization reports and structured answers are output, and accurate, explainable and multifunctional intelligent knowledge services are achieved.
Owner:SHANDONG UNIV

Full-text retrieval method and system fusing various types of documents

The invention provides a full-text retrieval method and system fusing various types of documents, and relates to the technical field of information retrieval, and the method comprises the following steps: obtaining document representation through document content extraction and structure recognition, generating a cross-modal semantic vector by using word embedding and nonlinear transformation, constructing a hierarchical index and a cross-document association graph, and obtaining a full-text retrieval result; the basic correlation score is calculated after the query request is received, and the comprehensive score of the candidate content segments is calculated based on the association graph to determine the optimal retrieval result, so that unified representation and retrieval of heterogeneous documents are realized, the cross-document retrieval precision and relevance are improved, and the processing capability of a retrieval system on complex queries is enhanced.
Owner:BEIJING CHANGFA TECH CO LTD

Multi-modal file label automatic generation method

The invention discloses a multi-modal file label automatic generation method, and relates to the technical field of label generation, and the method comprises the steps: firstly constructing a multi-level file label category set as a label matching reference library, then uploading a target multi-modal file, preprocessing a text extracted according to the file, and constructing a file representation vector; finally, the computing power of the current equipment is judged, labels are generated according to scenes, if the computing power is limited, the labels are extracted step by step, and a multi-level candidate label set is generated; if the computing power is sufficient, combining the multi-level label category sets, and directly extracting a candidate label set; thirdly, modifying the word segmentation result by calculating the cohesion degree of adjacent words in the initial word segmentation result, and automatically extracting the multi-level file label again; and finally, calculating a comprehensive cosine similarity value of the generated candidate file tags, and outputting an optimal file tag. According to the invention, effective utilization of resources can be realized, and the accuracy of the generated file tag is improved.
Owner:BEIJING ELECTRONICS SCI & TECH INST

Long text abstract generation method based on hierarchical graph comparison theme

PendingCN122021560ASemantic analysisText processingDocument representationInformation coverage
The invention discloses a long text abstract generation method based on hierarchical graph comparison themes, which comprises the following steps of: 1, preprocessing an original document, dividing sentence sequences, and obtaining global context-aware sentence and document representation through a hierarchical encoder network; 2, deducing document-level and sentence-level topic distribution by using a neural topic model; and 3, constructing a supervision graph based on a standard abstract to perform graph comparison learning so as to close the topic representation of a document and a key sentence and push redundant information. According to the method, the deep semantic structure of the long document can be effectively captured, so that the theme consistency and the information coverage degree of the abstract can be improved, and the redundancy is reduced.
Owner:ANHUI AGRICULTURAL UNIVERSITY

A modeling method, apparatus, equipment and medium for a power business decision model

This application relates to the field of mathematical modeling technology and discloses a modeling method for a power business decision-making model. The method includes: constructing a power corpus segmentation library and a power vector knowledge base to determine the power business problem text; performing a hybrid retrieval of the power business problem text using the power corpus segmentation library and the power vector knowledge base to obtain a first similarity score corresponding to candidate corpus segments and a second similarity score corresponding to candidate document representation vectors; filtering the candidate corpus segments and candidate document representation vectors based on the first and second similarity scores to determine the target reference text; and inputting the power business problem text and the target reference text into a target large language model to obtain the target decision model. The target decision model includes a description of the power business problem, an objective function, and its constraints. Its beneficial effect is that, based on hybrid retrieval and the standardized output of the large language model, the reliability and accuracy of the target decision model are significantly improved.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO LTD +1

Limit classification processing using graphs and neural networks

Systems and methods are provided for learning a classifier for annotating documents with predicted labels under an extreme classification where there are more than one million labels. The learning includes receiving a joint graph comprising documents and labels as nodes. A multi-dimensional vector representation of the documents (i.e., document representations) is generated based on graph convolution of the joint graph. Each document representation varies in dependence on adjacent nodes to accommodate context. The document representations are feature transformed using a residual layer. A per-label document representation is generated from the transformed document representations based on adjacent label attention. The classifier is trained for each of the more than one million labels based on joint learning using training data and the per-label document representations. The trained classifier performs highly efficiently compared to other classifiers trained using disjoint graphs of documents and labels.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Multi-modal long document retrieval method based on hierarchical representation and double-branch retrieval

PendingCN121958480Aachieve effectivenessAchieve positioningSemantic analysisCharacter and pattern recognitionDocument structuringDocument representation
The invention belongs to the technical field of document retrieval, and particularly relates to a multi-modal long document retrieval method based on hierarchical representation and double-branch retrieval. According to the method, three-level hierarchical document representation composed of an element level, a page level and a chapter and section level is constructed by designing hierarchical features oriented to a document structure; a double-branch retrieval mechanism is used for conducting semantic-level matching and feature-level matching on queries, and a final evidence set is generated through a fusion strategy. According to the method, the cross-page semantic association and the fine-grained visual clues can be utilized at the same time, effective retrieval and information positioning of the multi-modal long document are achieved, the retrieval requirements of documents with different layout forms and different lengths in a real scene are met, and the method has important practical application value.
Owner:FUDAN UNIVERSITY

Intellectual property information analysis and management method based on multi-dimensional data

The invention discloses an intellectual property information analysis and management method based on multi-dimensional data, and the method comprises the following steps: obtaining patent original data, constructing a patent record, and carrying out the preprocessing of an intellectual property data set; dividing the patent text data into a plurality of coding units according to a preset field range, and constructing vectorization representation for each patent; generating an adjusted sub-query set for the retrieval input based on a query decomposition algorithm; obtaining a sub-query mark sequence and a comprehensive candidate result set; and calculating a comprehensive sorting score, and generating a retrieval result list. According to the method, by introducing a patent retrieval framework combining multi-dimensional query decomposition and ConstBERT fixed multi-vector document representation, executable sub-query configuration can be stably generated and fine-grained semantic matching can be realized in a scene in which a long retrieval type, multiple limits and exclusion conditions coexist; and the accuracy, controllability, traceability and large-scale application performance of patent retrieval are improved.
Owner:BEIJING AUGUST MELON TECHNOLOGY CO LTD

A Multimodal Document Retrieval Method and Device Based on Cross-Modal Mutual Attention Mechanism

This application relates to the field of document retrieval technology, and particularly to a multimodal document retrieval method and apparatus based on a cross-modal mutual attention mechanism. The method includes: modeling a multimodal representation of a document; obtaining a target document-perceived multimodal document representation based on a multimodal mutual attention mechanism; fusing the document's self-attention vectorized representation and multimodal enhanced vectorized representation to obtain a unified multimodal enhanced representation of the document; calculating and ranking the relevance scores between the target document and at least one candidate document; and retrieving relevant documents. This application's embodiments, based on a cross-modal mutual attention mechanism, can obtain matching documents by acquiring a unified multimodal enhanced representation of the document and calculating relevance scores, thereby fully utilizing the multimodal information of the document, enhancing the relevance between different modalities of the document, and thus improving the matching degree of the document retrieval results, making the retrieval results more accurate and reliable.
Owner:TSINGHUA UNIVERSITY

Event-based document retrieval method and device, electronic equipment and storage medium

The application provides an event-based document retrieval method and device, electronic equipment and storage medium, wherein the method comprises: obtaining a user query statement of a user for a document set to be retrieved; inputting the user query statement into a pre-trained large language model to obtain a document retrieval result; wherein the large language model is obtained by training and optimizing a training sample data set composed of a document representation and a document identifier, the document representation is obtained by events and event relationships in the document set to be retrieved, and the document identifier is obtained by mapping the events in the document set to be retrieved to an event hierarchy. The method considers the relevance between document contents, effectively represents the document to be retrieved by using events and event relationships, and significantly improves the document retrieval performance of the large language model; the event hierarchy is used to construct a document identifier with a clear semantic structure, effectively strengthening the connection between the document identifier and the document content.
Owner:TSINGHUA UNIVERSITY

A RAG intelligent retrieval question and answer system and method based on enhanced metadata

The application discloses an RAG intelligent retrieval question and answer system and method based on enhanced metadata, relates to the technical field of information processing and intelligent retrieval, and obtains unified knowledge data after preprocessing of multi-source heterogeneous knowledge data, extracts structured metadata from the unified knowledge data based on the difference in text form; the literature text in the structured metadata is vectorized in combination with embedding of a knowledge graph to obtain document representation, and then the document representation, the structured metadata and an enhanced keyword set are output as an enhanced metadata object; a display relationship between different enhanced metadata objects is marked, and a knowledge database is constructed; restrictive conditions and a question intention are extracted from a user question, mixed retrieval is performed from the knowledge database based on the restrictive conditions and the question intention, a candidate literature semantic set is output, and then a structured statistical visual report and a structured answer are output, so that precise, interpretable and multifunctional intelligent knowledge service is realized.
Owner:SHANDONG UNIV

A knowledge encoding and activating method and device based on a large language model

This invention discloses a knowledge encoding and activation method and apparatus based on a large language model, belonging to the field of natural language processing technology. First, the invention generates pseudo-queries for the input document set, describing the document content from multiple semantic dimensions to construct a document semantic representation similar to the user query. Then, the pseudo-queries are used to train the large language model, encoding the document semantic knowledge into the model parameters to form a parameterized knowledge representation. Further, reinforcement learning training guides the language model to activate internal knowledge related to the user query during the generation process. Finally, in the inference stage, the model's internal knowledge is activated based on the user query to generate the output result. This method reduces reliance on external retrieval modules, improves the semantic matching ability between document representation and user query, and thus enhances the accuracy and stability of the generated results.
Owner:SHANXI UNIV

Document multi-level label generation method and device, computer device and readable storage medium

PendingCN122153058AMetadata text retrievalBiological modelsDocument representationEngineering
The application discloses a document multi-level label generation method and device, computer equipment and a readable storage medium, relates to the technical field of natural language processing, and can be specifically applied to the financial and medical fields. Through multi-level clustering marking processing, different granularity feature labels of the document can be accurately captured, and the demand of the financial and medical fields for generating document labels that can finely represent the document is better met. The method comprises the following steps: obtaining a target document to be generated with a document label, and obtaining a pre-generated multi-level label codebook; generating an initial document representation vector for the target document based on a document content summary and a document internal matching diagram of the target document; performing hierarchical iterative processing on the initial document representation vector by using the multi-level label codebook, determining the nearest clustering center for the initial document representation vector from a plurality of clustering centers included in each level, generating a multi-level document label corresponding to the target document, and outputting the multi-level document label.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Information extraction traceability method and system based on native multi-modal large model, and readable storage medium

The invention discloses an information extraction traceability method and system based on a native multi-modal large model and a readable storage medium, and belongs to the technical field of artificial intelligence and multi-modal learning. Comprising the steps of determining a file representation form for a to-be-processed original image, performing coordinate labeling by adopting a mantissa outward rounding mode, and constructing a space labeling data set; inputting an image by using the spatial annotation data set and adopting a dynamic resolution input strategy, and training a multi-modal large language model; inputting a sample image to the trained multi-modal large language model, outputting text content and coordinate information of the sample image, and performing coordinate post-processing on the coordinate information to realize edge compensation; and outputting the text content and the post-processed coordinate information. According to the method, the problems of coordinate dislocation, information loss and the like in a traditional OCR + language model two-stage architecture can be solved, and meanwhile, the defects of precision deviation and non-uniform format of an existing MLLM in space coordinate generation are overcome.
Owner:BEIJING YIDAO BOSHI TECH

Consistent entity tagging with multi-protocol data access

ActiveCN115812198BData accessDocument representation
A technique for consistent entity tagging is provided for data access utilizing multiple protocols. In one example, a file storage system is configured to process data according to (multiple) file storage protocols and (multiple) object storage protocols. Object storage protocols can utilize entity tags indicating whether an object (using a file representation in the file storage system) has been modified. In the case of modifying a file using a file storage protocol, an indication that the file lacks a valid entity tag can be stored. If an object storage operation is performed to retrieve an object, and if the object corresponds to a valid entity tag, that entity tag can be returned as part of the response. If the object does not correspond to a valid entity tag, the file storage system can generate a new entity tag and return the newly generated entity tag as part of the response.
Owner:EMC IP HLDG CO LLC

Multi-label classification method and device for collaborative attention and prototype alignment

The invention discloses a multi-label classification method and device for collaborative attention and prototype alignment, and belongs to the technical field of natural language processing. The method comprises the following steps: firstly, obtaining lexical element level context representation by utilizing a pre-training semantic encoder, calculating a maximum correlation score of lexical elements and a label space based on label embedding, and generating a filtering mask to sparise a sequence and suppress redundant noise; then constructing a label attention branch and a sentence-level hierarchical self-attention branch in parallel, realizing cross-branch fine-grained semantic alignment through bidirectional collaborative attention, and completing feature fusion by adopting adaptive weighting; on the basis of fusion representation, a label prototype gating fusion and momentum type online updating mechanism is introduced, a learnable label prototype is used as a semantic center to continuously guide document representation to gather towards related labels, and long-tail distribution and semantic drift are relieved. According to the method, the robustness of low-frequency label prediction can be enhanced while the classification precision and the sorting quality are improved, and the unnecessary attention calculation overhead is reduced.
Owner:ZHEJIANG SCI-TECH UNIV

An abnormality detection method, device, system and storage medium for data migration

Embodiments of the present application provide a data migration detection method, device, system and storage medium, the method comprising: obtaining information digests of next layer sub-files of a source file and a target file respectively; the source file represents a file package of migrated data; the target file represents a file package obtained after migration of the source file; the information digest of the sub-file comprises a size of the sub-file, a storage location of the sub-file and a number of next layer files of the sub-file; generating an information digest of the source file based on the information digests of the next layer sub-files of the source file; generating an information digest of the target file based on the information digests of the next layer sub-files of the target file; and determining a data migration detection result according to the information digest of the source file and the information digest of the target file.
Owner:CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

Deep learning model explanation method, apparatus, and device

The embodiment of the application provides a deep learning model explanation method, device and equipment, the method comprises the following steps: loading a target deep learning model according to a loading mode corresponding to the target deep learning model, obtaining a model file of the target deep learning model; performing convolution, uniformization, pooling and activation processing on the model file of the target deep learning model, obtaining an analysis file corresponding to the model file; reconstructing the analysis file, obtaining a reconstructed file corresponding to the analysis file; and representing the reconstructed file as a structure description file containing only structure description and a weight file arranged according to the order of operators in the structure description file. The method provided by the application can express models under different training frameworks as unified network structure representation and unified parameter arrangement order, improving the efficiency of deep learning model deployment.
Owner:GREE ELECTRIC APPLIANCE INC OF ZHUHAI +1

Hierarchical graph neural network for cross-architectural software reverse engineering

Recovering symbols from a stripped binary includes representing the stripped binary as a plurality of graph of graphs (GoG) representations, converting the plurality of GoGs into a plurality of expressive representations of each function in the stripped binary, training a machine learning (ML) model using the expressive representations, and determining a missing symbol of at least one of the functions based on an output of the ML model. Information relating to functions is and interactions between functions are used to train an ML to determine similarity of two functions. Based on similarities, missing symbols may be inferred and used to enable updates to the stripped binary file where source code is not available.
Owner:SIEMENS CORP +1

Mirror image generation method and device, electronic equipment and storage medium

The invention relates to a mirror image generation method and device, electronic equipment and a storage medium, and relates to the technical field of computers. The mirror image generation method comprises the steps of obtaining a first file list; wherein the first file list comprises newly added files and / or changed files in the container running process; the container is created based on a basic mirror image; the changed file represents a file which exists in the basic mirror image and has changed content; the newly added file represents a file which does not exist in the basic mirror image; deleting a target file from the first file list to obtain a second file list; and generating a target mirror image based on the files in the second file list and the basic mirror image. By means of the method and device, the target file can be automatically removed when the mirror image is generated, and the size of the space occupied by the mirror image generated in the container running process is reduced.
Owner:SHANGHAI INFINIGENCE AI INTELLIGENT TECHNOLOGY CO LTD

Decarbonizing BERT with topics for efficient document classification

Various embodiments of the teachings herein include a computer-implemented method of fine-tuning Natural Language Processing (NLP) models. Some examples include: providing a training data set including a multitude of training text documents; providing a NLP model including a Neural Network (NN) based Topic Model (TM) having scalable TM parameters and a parallel large-scale pre-trained Language Model (LM) having scalable LM parameters; and fine-tuning the NLP model by jointly training the NN-based TM and the parallel large-scale pre-trained LM using a projected vector comprising a combination and projection of a document topic proportion generated by the NN-based TM based on the scalable TM parameters from an input training text document of the multitude of training text documents, and of a contextualized document representation generated by the large-scale pre-trained LM based on the scalable LM parameters from the same input training text document.
Owner:DRIMCO GMBH