Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

141 results about "Document retrieval" patented technology

Document retrieval is defined as the matching of some stated user query against a set of free-text records. These records could be any type of mainly unstructured text, such as newspaper articles, real estate records or paragraphs in a manual. User queries can range from multi-sentence full descriptions of an information need to a few words.

Enterprise-level large model agent application system supporting multi-modal collaboration

The invention relates to the technical field of artificial intelligence, and discloses an enterprise-level large-model agent application system supporting multi-modal collaboration, and the system comprises a user interaction terminal, an agent engine server, a knowledge engine server, a plug-in integration center, and a distributed storage unit. The agent engine server is responsible for intention recognition and task arrangement of a multi-modal input signal, and dynamically loads a differential reasoning strategy based on an environment isolation mechanism. And the knowledge engine server constructs a cross-modal semantic anchor point, analyzes an unstructured document into a tetrad knowledge unit, and realizes accurate recall of images and texts by using a hybrid retrieval algorithm. And the plug-in integration center executes outbound replacement and inbound restoration of the sensitive data through the context-aware dynamic desensitization gateway. According to the method, through a multi-modal semantic association and closed-loop verification mechanism, the problems of low complex document retrieval precision and leakage of external calling data are solved, and the service processing capacity and safety of the system are improved.
Owner:LINGRUIDA (XIAMEN) TECHNOLOGY CO LTD

Full-text retrieval method and system fusing various types of documents

The invention provides a full-text retrieval method and system fusing various types of documents, and relates to the technical field of information retrieval, and the method comprises the following steps: obtaining document representation through document content extraction and structure recognition, generating a cross-modal semantic vector by using word embedding and nonlinear transformation, constructing a hierarchical index and a cross-document association graph, and obtaining a full-text retrieval result; the basic correlation score is calculated after the query request is received, and the comprehensive score of the candidate content segments is calculated based on the association graph to determine the optimal retrieval result, so that unified representation and retrieval of heterogeneous documents are realized, the cross-document retrieval precision and relevance are improved, and the processing capability of a retrieval system on complex queries is enhanced.
Owner:BEIJING CHANGFA TECH CO LTD

Multimodal Data Ingestion And Retrieval For Agent Systems

Techniques for multimodal document retrieval are disclosed herein. Multimodal documents that include both textual and graphical components are retrieved from a knowledge base by a multimodal retrieval augmented generation (RAG) agent in response to a query. The documents and / or components or chunks thereof are retrievable by the RAG agent from the knowledge base using the semantic summaries and / or vector search of embeddings in the knowledge base that are generated from text extracted from processing non-textual components of the data. The RAG agent classifies the query type to determine whether to use a semantic match for text or image summaries, full text semantic search, vector cosine similarity search, and / or other multimodal vector search. The RAG agent performs types of searches selected based on the modality used to generate the response to the query.
Owner:ORACLE INT CORP

Multimodal Data Ingestion And Retrieval For Agent Systems

Techniques for multimodal document retrieval are disclosed herein. Multimodal documents that include both textual and graphical components are retrieved from a knowledge base by a multimodal retrieval augmented generation (RAG) agent in response to a query. The documents and / or components or chunks thereof are retrievable by the RAG agent from the knowledge base using the semantic summaries and / or vector search of embeddings in the knowledge base that are generated from text extracted from processing non-textual components of the data. The RAG agent classifies the query type to determine whether to use a semantic match for text or image summaries, full text semantic search, vector cosine similarity search, and / or other multimodal vector search. The RAG agent performs types of searches selected based on the modality used to generate the response to the query.
Owner:ORACLE INT CORP

Cerebral stroke knowledge question-answering system construction method and system based on knowledge graph and large language model

The invention relates to the technical field of medical health information services, in particular to a cerebral apoplexy knowledge question-answering system construction method and system based on a knowledge graph and a large language model. Natural language input of a user is analyzed through a query agent, a query intention and constraint conditions are recognized, and a structured execution plan is generated; a user state management tool is forcibly activated, a static clinical portrait and a dynamic rehabilitation log are loaded, and a personalized context is constructed; a plurality of tools such as knowledge graph query, authoritative literature retrieval and rehabilitation plan generation are scheduled, and accurate retrieval and reasoning of heterogeneous knowledge are completed; through double verification of fact consistency and clinical risks, error or high-risk suggestions are intercepted and replaced with risk early warning. The problems of'illusion 'risk, insufficient individuation, poor interpretability and the like of a traditional single model are solved to a large extent, high-credibility, individuation and traceable rehabilitation knowledge service can be provided for the stroke patient and a caregiver of the stroke patient, and rehabilitation safety and effect are guaranteed.
Owner:DALIAN UNIV

RAG-based pdf intelligent retrieval and generation method and system

The application discloses a kind of PDF intelligent retrieval and generation method and system based on RAG, by obtaining the document data of input, using the classification model established in advance to parse document data, extract text content and image content to form first data set;Using deep learning model to the image content in first data set carries out feature extraction, while the text content in first data set applies natural language processing technology to carry out semantic analysis, obtains multimodal feature set;According to multimodal feature set, application information integration algorithm is uniformly encoded and is handled to generate second data set, if detecting the integrity of fusion feature vector in second data set is lower than preset threshold value, then supplementary context semantic analysis fills in missing information;Using preset index construction mechanism to the clustering processing of fusion feature vector in second data set, generates the retrieval index library containing classification index structure.The application improves the accuracy and comprehensiveness of document retrieval.
Owner:HUNAN ZHIXUE YOUKE INFORMATION TECHNOLOGY CO LTD +1

Machine learning-based literature search and retrieval and related machine learning model training methods

Described herein are systems and methods for performing literature retrieval and related machine learning model training methods. An example computer-implemented method of training a machine learning model configured for literature retrieval I includes receiving a plurality of full-text articles; extracting, from the plurality of full-text articles, a plurality of positive sentence-citation pairs, each positive sentence-citation pair comprising a respective citing sentence and at least one cited article that is associated with the respective citing sentence; creating a labeled dataset comprising the plurality of positive sentence-citation pairs; and training a machine learning model using the labeled dataset.
Owner:FLORIDA STATE UNIV RES FOUND INC

Multi-hop reasoning method based on dynamic reasoning guidance and multistage self-feedback retrieval

A multi-hop reasoning method based on dynamic reasoning guidance and multistage self-feedback retrieval belongs to the field of natural language processing, and comprises the following steps: deconstructing a multi-hop reasoning process into a target-oriented sequence decision problem, carrying out dynamic reasoning guidance by using a large language model, generating a sub-problem sequence matched with a reasoning progress in real time, and carrying out multi-level self-feedback retrieval on the sub-problem sequence; target document retrieval is guided, and sub-questions are dynamically generated; according to the generated sub-questions, obtaining associated documents by adopting a three-level collaborative retrieval mechanism; and performing information refining on the associated document through a large language model, fusing the refined information into an inference chain, and performing inference to generate an answer. The invention further discloses a multi-hop reasoning system, a storage medium and a computer program product. The method aims at solving the complex multi-hop problem that multiple dispersed knowledge fragments need to be integrated, high-accuracy and high-efficiency reasoning is achieved, the retrieval requirement is dynamically generated through an explicit thinking chain guiding mechanism, and evidence obtaining is optimized and redundant information is filtered in combination with a three-level self-feedback retrieval mechanism.
Owner:XI AN JIAOTONG UNIV

Retrieval optimization method adaptive to multi-dimensional storage of power documents

A retrieval optimization method adaptive to multi-dimensional storage of power documents relates to the technical field of information retrieval, and comprises the following steps: firstly, preprocessing user query, identifying power business scenes and technical types of the user query, and extracting a query keyword set and a query vector; then, a three-level progressive retrieval strategy is adopted to retrieve in an electric power document library composed of a metadatabase and a vector data at the first level, candidate documents are screened based on keyword matching and business scenes; in the second stage, further filtering is carried out through similarity calculation of a query vector and a technical abstract vector; in the third stage, final accurate screening is completed in combination with full-text vector similarity and electric power professional rules; finally, information integration and structured output are conducted on the result, the problems that in traditional power document retrieval, the result is inaccurate, and efficiency is low are solved, and retrieval precision and response speed are remarkably improved.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Domain-specific retrieval language models

Various examples, systems, and methods are disclosed relating to domain-specific document retrieval that incorporates custom vocabulary integration and embedding model updates. A computing system can extract multiple segments from a collection of documents and generate queries that correspond to at least one segment. The computing system can identify terms that satisfy a uniqueness criterion and input the terms into a tokenizer to create a vocabulary dataset. The vocabulary dataset, the document segments, and the queries can be used to update an embedding model to support retrieval and semantic alignment within private documents.
Owner:NVIDIA CORP

Intelligent question answering method and system based on large model and retrieval enhancement

The invention discloses an intelligent question-answering method and system based on a large model and retrieval enhancement, and relates to the field of intelligent question-answering. The method comprises the following steps: performing preliminary retrieval on a preset knowledge base to obtain document fragments, and constructing a candidate version set; if it is determined that the query request does not belong to the cross-version query type, calculating a version consistency score, and determining a target version with the highest score; obtaining a target document fragment corresponding to the target version, and performing vectorization processing on the query request to generate a problem vector; performing vector retrieval in a preset vector knowledge base on the basis of the problem vector to obtain a problem vector result, and performing secondary retrieval in the preset knowledge base on the basis of the target version to obtain a document retrieval result; merging the problem vector result and the document retrieval result to obtain an initial candidate document; and generating a first final cue word, and inputting the first final cue word into a preset large language model to obtain a first target answer. By implementing the technical scheme provided by the invention, the accuracy of intelligent question answering is improved.
Owner:BEIJING ZHIBAO HUIZHONG DIGITAL TECHNOLOGY CO LTD

Literature comprehensive retrieval system and retrieval enhancement generation method thereof

The invention discloses a literature comprehensive retrieval system and a retrieval enhancement generation method thereof, and relates to the technical field of literature synthesis, and the method comprises the following steps: a demand anchoring layer carries out the retrieval of an initial retrieval request or a recursive feedback follow-up problem and a learning key point according to breadth and depth parameters; generating a corresponding number of query statements meeting a preset retrieval rule according to the query quantity; the resource acquisition layer retrieves and collects literatures and metadata in multiple channels according to query statements; the organization processing layer extracts a set number of documents, performs integration and structured extraction on contents of the documents, generates follow-up questions and learning key points, feeds back the follow-up questions and the learning key points to the demand anchoring layer, and synchronously retains the learning key points and the documents; and the result output layer converts the retained content into a target format. In this way, through combination of hierarchical cooperation and a recursion mechanism, traditional single retrieval limitation is broken, deviation caused by fuzzy requirements is avoided, retrieval accuracy, processing efficiency, information integrity and achievement usability are taken into consideration, and literature retrieval efficiency is improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Technical supervision document retrieval method and related device

The invention provides a technical supervision document retrieval method and a related device, and belongs to the field of technical supervision document retrieval. The method comprises: acquiring a user query intention; retrieving from a database according to the query intention of the user to obtain a natural language answer facing the user question; the database construction method comprises the steps of obtaining an existing technical supervision document; performing analysis and knowledge extraction on the technical supervision document to obtain a knowledge multi-tuple; after the knowledge multi-tuple is verified, time dimension attributes are added to the knowledge multi-tuple; and respectively storing the knowledge multi-tuples added with the time dimension attributes according to data types to obtain a database. According to the technical supervision document retrieval method and device, the problem of low accuracy of technical supervision document retrieval is solved.
Owner:DATANG HYDROPOWER SCI & TECH RES INST CO LTD +2

Systems and Methods for Prompt-Based Query Generation for Diverse Retrieval

An example method for prompt-based query generation is provided. The method includes receiving, by a computing device, at least two prompts associated with a retrieval task to be performed on a corpus of documents associated with the task. The method includes applying, based on the at least two prompts and the corpus of documents, a large language model to generate a synthetic training dataset comprising a plurality of query-document pairs, wherein each query-document pair comprises a synthetically generated query and a document from the corpus of documents. The method includes training, on the plurality of query−document pairs from the synthetic training dataset, a document retrieval model to take an input query associated with the retrieval task and predict an output document retrieved from the corpus of documents. The method includes providing, by the computing device, the trained document retrieval model.
Owner:GOOGLE LLC

Knowledge graph multi-mode document analysis and image table semantization knowledge recall method

The invention discloses a knowledge graph multi-modal document analysis and image table semantization knowledge recall method, and belongs to the technical field of knowledge engineering and information retrieval. The invention provides an innovative scheme for fusing a visual language model, semantic abstract generation and knowledge graph modeling. The method comprises the following steps: constructing a vertical domain knowledge graph by adopting a BERT-BiLSTM-CRF model; according to the method, multi-modal document analysis is realized through models such as DocLayout-YOLO, TableMaster, UniMERNet and the like; the method comprises the following steps of: segmenting an image into 16 * 16 block sequences by adopting a vit-gpt2-image-adaptation model, and realizing image semantization through 768-dimensional vector space mapping and Transform coding; constructing a document summary tree based on DBSCAN clustering and LLM recursive summary; and designing a hybrid retrieval space fusing semantic vectors and structured vectors, and reordering by adopting a double-attention mechanism. According to the method, the knowledge base document retrieval recall rate is increased to 99%, the question and answer accuracy rate reaches 90% or above, the index construction time is shortened by 60%, and the problem that semantic understanding and recall of non-text elements in complex documents are difficult is effectively solved.
Owner:云鼎科技股份有限公司

Data query and analysis system and method based on multi-model fusion

PendingCN121681646ADatabase management systemsBiological modelsCross language retrievalCorrelation analysis
The invention relates to the technical field of intelligent information retrieval, and discloses a data query and analysis system and method based on multi-model fusion, and the system comprises a data collection module, an intention perception module, a fusion retrieval module and a correlation analysis module. The method corresponds to the system. According to the method, firstly, the technical problems of multi-modal data fusion and cross-language retrieval are solved through unified representation processing on unstructured literatures; further, through the dynamic perception and strategy adaptation of the query intention, the spanning from a single retrieval mode to intelligent strategy selection is realized; and finally, through a multi-modal interaction enhancement network and an evidence explanation mechanism, not only is the accuracy of literature contribution degree evaluation improved, but also a traceable recommendation basis is provided for a user, the retrieval effect is optimized in the breadth dimension and the depth dimension, and especially in a medical scientific research topic selection scene, the retrieval efficiency is improved. Researchers can be effectively assisted in finding high-value literatures and research directions, and the efficiency and quality of literature retrieval are improved.
Owner:YIZHIKANG (SUZHOU) SMART TECHNOLOGY CO LTD

Document question and answer method and device, storage medium and program product

The invention discloses a document question and answer method and device, a storage medium and a program product, and the method comprises the steps: carrying out the content retrieval of a first document database based on an inquiry statement, so as to determine a corresponding intermediate retrieval result; performing content retrieval on a second document database based on the inquiry statement and the intermediate retrieval result to determine a retrieval matching result, the first document database and the second document database being databases constructed based on at least one document; detecting whether the retrieval matching result meets a reply condition of the inquiry statement or not; and when the retrieval matching result meets a reply condition, constructing a reply answer according to the retrieval matching result. Therefore, the question answering system can adapt to questions with different semantic complexities and logic levels through a document retrieval mechanism mixed with the neural symbols, and high-quality answers can still be efficiently produced when complex questions are processed.
Owner:AISPEECH CO LTD

A method and system for constructing an investment risk assessment index system

This invention discloses a method and system for constructing an investment risk assessment index system. The method includes obtaining basic data for constructing the target index system through relevant literature retrieval in a literature database; identifying investment risk factors by using a risk factor identification model to perform routine project risk factor identification on the basic data; mining risk factors for offshore photovoltaic power generation projects based on preset criteria and marine characteristic data; and determining investment risk assessment indicators based on the investment risk factors and the offshore photovoltaic power generation project risk factors. This invention can consider the characteristics of photovoltaic power generation projects and the features of the marine environment, and construct a scientific, systematic, applicable, and comprehensive investment risk assessment index system for offshore photovoltaic power generation projects according to principles.
Owner:华能(临高)新能源有限公司 +1

Multi-document retrieval enhancement generation method and device, equipment, storage medium and computer program product

The invention relates to the technical field of natural language processing, and discloses a multi-document retrieval enhancement generation method and device, equipment, a storage medium and a computer program product.The method comprises the steps that in response to input query information, multi-document retrieval is conducted in a directory forest based on the query information, and a multi-document retrieval result is obtained, the directory forest is used for representing semantic association among the multiple documents, inputting the multi-document retrieval result into a preset large model, and obtaining a response result output by the preset large model; according to the method, the semantic association among the documents is represented by pre-constructing the directory forest, and multi-document retrieval is performed in the directory forest based on the query information, so that the multi-document linkage capability generated by retrieval enhancement can be improved, and multi-document collaborative retrieval can be realized.
Owner:BEIJING QIHOOD TECHNOLOGY CO LTD

Intelligent retrieval system for unstructured documents

The application relates to the technical field of document retrieval, in particular to an intelligent retrieval system for unstructured documents. The system comprises a data acquisition module for acquiring unstructured documents; a document feature analysis module for determining representative feature values in combination with keyword semantic importance, paragraph quantity and frequency, and constructing a theme consistency feature vector based on local and global dimensional theme distribution; a document classification module for clustering by comprehensively calculating and measuring distance of themes, keywords and consistency features, and selecting representative documents to construct a knowledge graph; and a retrieval module for generating a retrieval result based on the knowledge graph in combination with a large language model. The application solves the problem of serious homogenization of unstructured document retrieval results, improves the efficiency of intelligent retrieval of unstructured documents by clustering and deduplication and combining with a knowledge graph to enhance semantic association.

Document retrieval method and device, equipment, storage medium and computer program product

The invention relates to the technical field of computers, and discloses a document retrieval method, device and equipment, a storage medium and a computer program product.The method comprises the steps that in response to a document retrieval request, the document retrieval request is converted into a problem retrieval vector, and the vector similarity between the problem retrieval vector and a document image retrieval vector is calculated, the document image retrieval vector is obtained by analyzing a document into a document image set and encoding the document image set, and a document retrieval result corresponding to the document retrieval request is generated according to the vector similarity; according to the document retrieval method, the document is analyzed into the document picture set in advance, the document picture set is encoded to obtain the document image retrieval vector, and the document retrieval is performed through the problem retrieval vector and the document image retrieval vector, so that the document retrieval in the form of the document image is realized; therefore, more space structure information, fine-grained information and context information can be reserved, and the document retrieval precision is improved.
Owner:BEIJING QIHOOD TECHNOLOGY CO LTD

Hand-drawing shape-based document retrieval

The present disclosure provides methods, apparatuses and computer program products for hand-drawing shape-based document retrieval. An input hand-drawing shape may be obtained. A hand-drawing shape feature of the hand-drawing shape may be extracted through a feature extracting model. At least one target document may be retrieved by using the hand-drawing shape feature and a feature index library associated with a plurality of candidate documents, at least one document page in the target document locally matching the hand-drawing shape.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Intelligent retrieval method for legal documents based on deep semantic matching

This invention discloses an intelligent legal document retrieval method based on deep semantic matching, belonging to the field of legal document retrieval technology. It includes establishing multiple different retrieval modes; quantifying the relationship strength between keywords and legal documents based on the order and frequency of keywords corresponding to user-inputted search terms in historical search records, and updating the keyword ranking; analyzing search records and classifying search behavior into one of the retrieval modes; and ranking legal documents according to the correlation strength between user-inputted search terms and corresponding keywords for user selection. This invention establishes multiple retrieval modes for different search habits, with data not shared between modes. When initiating a search, users are distinguished into different retrieval modes, and different retrieval rules are applied to different modes to achieve rapid retrieval, further improving retrieval efficiency and enabling fast search results for users with different search habits.
Owner:SICHUAN LEZHENG TECH CO LTD

Rail transit document retrieval method and device, electronic equipment and storage medium

The invention provides a rail transit document retrieval method and device, electronic equipment and a storage medium. The rail transit document retrieval method comprises the steps of obtaining an initial document and a retrieval keyword; constructing a knowledge graph corresponding to the initial document, wherein the knowledge graph is used for representing a relationship among contents in the initial document; based on the knowledge graph and the content of the initial document, constructing an RAG vector corresponding to the initial document; and according to the retrieval keyword and the RAG vector, retrieving and outputting a retrieval result. By constructing the knowledge graph and the RAG vector corresponding to the whole document, the retrieval result corresponding to the retrieval keyword can be retrieved in the document in any format according to the retrieval keyword input by the user, so that the complicated PDF document can be efficiently understood and processed, and the processing difficulty of PDF data is reduced. Therefore, the method provided by the embodiment of the invention can be suitable for documents in various formats, and is particularly suitable for the field of rail transit with a large number of PDF documents.
Owner:CHINA STATE RAILWAY GRP CO LTD +3

Hybrid multi-modal document retrieval enhancement method, system and equipment and storage medium

The invention discloses a hybrid multi-modal document retrieval enhancement method, system and device and a storage medium. The method comprises the following steps: performing format analysis on an unstructured document; generating a text vector, a visual vector and a structured table vector based on an analysis result, and constructing three index libraries for storing the text vector, the visual vector and the structured table vector respectively; in response to the natural language queried by the user, recalling candidate text paragraphs, candidate pictures and candidate tables related to the semantics of the natural language queried by the user; merging at least two types of candidate elements in adjacent candidate text paragraphs, candidate pictures and candidate tables in the same page into a mixed fragment; assembling a plurality of mixed fragments to form a multi-modal context; and inputting the multi-modal context into the large language model to obtain an answer. According to the method, adjacent candidate elements of the same page are merged, spatial position association of elements in a document is considered, and cross-modal information is automatically integrated, so that the obtained answers pay more attention to context association.
Owner:MERIT DATA CO LTD

Document retrieval method and device, electronic equipment and storage medium

The invention discloses a document retrieval method and device, electronic equipment and a storage medium. The method comprises the steps of receiving a question of a user through a large language model; converting the question into a corresponding vector, and extracting one or more keywords from the question; retrieving the vector in a pre-constructed vector database to obtain a retrieval result corresponding to the vector; searching for each keyword in a pre-constructed document database to obtain a search result corresponding to each keyword; and based on the retrieval result corresponding to the vector and the retrieval result corresponding to each keyword, generating an answer corresponding to the question, and returning the answer corresponding to the question to the user through a large language model. According to the embodiment of the invention, the deep understanding capability of fine-grained semantics can be improved, so that the accuracy and compliance of retrieval recall can be enhanced.
Owner:DIGITAL GUANGDONG NETWORK CONSTR CO LTD

Document retrieval control methods, systems and servers

This invention provides a document retrieval control method, system, and server, relating to the field of document retrieval technology. This method combines a semantic model with document analysis to perform in-depth analysis, thus solving the problems of low efficiency in manual classification and insufficient semantic understanding in traditional RAG systems. Furthermore, this method can fully understand user needs and output the most relevant knowledge base based on the semantic model, reducing the retrieval scope and improving retrieval efficiency and accuracy. This solves the problems of insufficient recommendation accuracy, low intelligence, poor retrieval accuracy, and long response time in traditional RAG systems.
Owner:HANG ZHOU LING XIN SHU KE XIN XI JI SHU YOU XIAN GONG SI

Multi-model paper retrieval method for academic questions and answers

The invention discloses a multi-model paper retrieval method oriented to academic questions and answers, which relates to the technical field of natural language processing and information retrieval, and comprises the following steps: firstly, constructing a unified corpus and a training data set, and respectively encoding by using models in a first model set and a second target model to generate a document vector set; screening difficult negative samples based on the initial retrieval result of the second model, constructing a comparative learning sample, performing fine adjustment on the comparative learning sample, and recoding a corpus; and taking the fine-tuned second target model and the models in the first model set as a model group, respectively performing target query and performing similarity retrieval by adopting each model in the model group to obtain a corresponding original similarity matrix, screening documents based on the original similarity matrix, and generating a first target document list of the target query. According to the method, the accuracy and robustness of academic literature retrieval can be effectively improved.
Owner:SOUTHWEST PETROLEUM UNIV

Multi-modal long document retrieval method based on hierarchical representation and double-branch retrieval

PendingCN121958480Aachieve effectivenessAchieve positioningSemantic analysisCharacter and pattern recognitionDocument structuringDocument representation
The invention belongs to the technical field of document retrieval, and particularly relates to a multi-modal long document retrieval method based on hierarchical representation and double-branch retrieval. According to the method, three-level hierarchical document representation composed of an element level, a page level and a chapter and section level is constructed by designing hierarchical features oriented to a document structure; a double-branch retrieval mechanism is used for conducting semantic-level matching and feature-level matching on queries, and a final evidence set is generated through a fusion strategy. According to the method, the cross-page semantic association and the fine-grained visual clues can be utilized at the same time, effective retrieval and information positioning of the multi-modal long document are achieved, the retrieval requirements of documents with different layout forms and different lengths in a real scene are met, and the method has important practical application value.
Owner:FUDAN UNIVERSITY

Method for calculating high-entropy material structure descriptor by using large language model

The invention relates to the technical field of material informatics and artificial intelligence, in particular to a method for training and calculating a high-entropy material structure descriptor based on a large language model, and the method is used for a machine learning task of high-entropy material design and performance prediction. According to the 3D chemical structure model of the high-entropy material or the element information in the chemical structural formula, the corresponding information is extracted through the large language model, and the structure descriptor for the high-entropy material machine learning task is automatically calculated according to the element information with the highest utilization rate in the literature and the knowledge base, so that the method has the characteristics of simple operation, reliable data, high speed and the like; and the structure descriptor with the highest literature recognition degree can be obtained through calculation.
Owner:UNIV OF SHANGHAI FOR SCI & TECH