Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

348 results about "Document retrieval" patented technology

Document retrieval is defined as the matching of some stated user query against a set of free-text records. These records could be any type of mainly unstructured text, such as newspaper articles, real estate records or paragraphs in a manual. User queries can range from multi-sentence full descriptions of an information need to a few words.

Scientific and technical literature intelligent retrieval method based on generative artificial intelligence and related equipment

The invention provides a scientific and technological literature intelligent retrieval method and related equipment based on generative artificial intelligence, and the method comprises the steps: carrying out the multi-layer semantic annotation of medical scientific and technological literatures, constructing a symptom-disease dynamic association map, and building a medical scientific and technological literature knowledge base; performing medical context analysis and multi-modal feature extraction on user query, and generating a unified retrieval vector in combination with Boolean operation nested analysis and semantic alignment processing; performing evidence grading retrieval and clinical scene matching based on the unified retrieval vector to obtain a preliminary candidate literature set, and optimizing the preliminary candidate literature set into a target candidate literature set through fine-grained semantic recalculation; and calculating a retrieval prior probability according to the evaluation dimension, performing knowledge weighted fusion on the candidate literature, and generating a medical science and technology literature recommendation report. According to the method, the medical term association relationship is deeply understood through the association map, the result is ensured to be matched with the patient characteristics through evidence grading retrieval and clinical scene matching and screening, and the accuracy of document retrieval is improved.
Owner:FUDAN UNIVERSITY

Prompt-based data structure and document retrieval

A knowledge management system may generate a plurality of prompts based on divisions of documents of unstructured text, each prompt relevant to a division of unstructured text. At least one prompt is generated such that a corresponding division of unstructured text is a response to said at least one prompt. The system may generate prompt embeddings for the plurality of prompts corresponding to the plurality of documents of unstructured text. The system may generate prompt-embedding clusters to group similar prompts from one or more documents of unstructured text. The system may receive a query. The system may convert the query to one or more query embeddings. The system may identify one or more prompts that are relevant to the query based on comparing the one or more query embeddings to the prompt embeddings. The system may identify one or more documents in one or more prompt-embedding clusters.
Owner:PIENOMIAL INC

Dynamic depth document retrieval for enterprise language model systems

Systems and methods for resource-efficient retrieval of information using a generative AI model are disclosed. An input query requesting information from a set of documents is used in a prompt for a generative AI model to generate a search query to identify the documents relevant to the input query and their respective relevancy scores. The input query is used as an input another model to determine a depth score indicating a predicted number of documents needed to retrieve the information. Based on the depth score and the relevancy scores of the relevant documents, the system extracts grounding data from the identified relevant documents to generate an answer synthesis prompt for the generative AI model. The generative AI model processes the second to produce a response to the input query including the requested information.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Document retrieval method based on multistage index and feature clustering

The invention relates to the technical field of document retrieval and information processing in the data processing technology, in particular to a document retrieval method based on multistage indexing and feature clustering, which comprises the following steps: performing high-dimensional space mapping on multi-modal features such as texts and images through a quantum embedding layer to generate cross-modal joint feature representation; a first-level index of a multi-level index architecture is dynamically initialized based on a meta-clustering algorithm, and semantic blocks of a second-level index are divided in combination with a multi-head self-attention mechanism. And an optimal transmission matrix is generated by using a Sinkhorn algorithm to align cross-node feature distribution. The multi-target mixed retrieval strategy is fused with vector retrieval, keyword retrieval and graph retrieval results, and weight distribution is dynamically adjusted. Through collaborative optimization of quantum calculation, federated learning and causal reasoning, a closed-loop technical architecture from feature analysis to dynamic index construction is formed, the problems of insufficient cross-modal fusion, static clustering deviation and semantic association deficiency are solved, and the precision, efficiency and dynamic adaptability of heterogeneous document retrieval are improved.
Owner:TAIJI COMPUTER CORPORATION LIMITED

Multi-intelligent-agent physical examination enhancement generation method and device, equipment and storage medium

The invention provides a multi-agent query enhancement generation method, device and equipment and a storage medium, and the method comprises the steps: carrying out the sub-problem decomposition processing of an original query problem through a problem rewriting agent, and obtaining a group of sub-query problem sets; performing retrieval method selection processing on each sub-query problem through the retrieval agent according to the entity type contained in each sub-query problem in the sub-query problem set to obtain a corresponding retrieval method; performing document retrieval processing on the knowledge base by using a retrieval method to obtain a group of candidate documents, and performing correlation screening processing on the candidate documents through a document selection agent to obtain related target documents; and inputting the target document and the original query question into an answer generation agent for answer generation processing to obtain an answer result. According to the method, a multi-agent cooperation framework is adopted, joint optimization of all agents is achieved through end-to-end reinforcement learning training, and the retrieval and answer quality of complex questions is effectively improved.
Owner:SHENZHEN FUTURE QINGYAN INTELLIGENT TECHNOLOGY CO LTD

Enhanced document retrieval with semantic depth and syntactic structure

Certain aspects of the present disclosure describe a method of information retrieval. In certain aspects, the method includes identifying a set of relevant nodes of a document graph embedding semantic units associated with the document based on a document search query. The method further includes reconstructing a structural context for each relevant node in the set of relevant nodes. The method further includes processing the set of relevant nodes and the structural context of each relevant node with a large language model to generate a contextual response to the document search query.
Owner:INTUIT INC

Literature semantic search method and system based on elastic search

The invention discloses a literature semantic search method and system based on elastic search, and relates to data retrieval. The literature semantic search method comprises the steps that vectorization processing is conducted on a noun phrase list by means of a text2vec-based-multilingual model trained based on a CoSENT method; according to the semantic vector, performing approximate nearest neighbor search in a second retrieval module to obtain first candidate data; inputting the query text data into a first retrieval module, and performing keyword matching through a BM25 algorithm to obtain second candidate data; fusing the first candidate data and the second candidate data to obtain third candidate data; a Sequence Matcher algorithm is adopted to calculate character string similarity between expansion words in the third candidate data, a similarity threshold value is set based on the length of the longest common subsequence, duplicate removal is carried out, and fourth candidate data is obtained; and performing weight distribution based on positions and similarity scores on the fourth candidate data, and enhancing the distinction degree of the extension words by expanding a score interval to obtain extension word recommendation list data. According to the method, the accuracy of document retrieval is remarkably improved.
Owner:CHINA EDUCATIONAL PUBLICATIONS IMPORT & EXPORT CORP LTD

Document retrieval method and automatic question answering method

Embodiments of the present description provide a document retrieval method and an automatic question answering method. The document retrieval method comprises: obtaining data for retrieval; retrieving at least one candidate document from among a plurality of documents in a knowledge base on the basis of the data for retrieval; selecting at least one reference document from among the at least one candidate document on the basis of the association relationship between the data for retrieval and the at least one candidate document; and on the basis of the at least one reference document, updating the data for retrieval, to obtain updated data for retrieval, and retrieving a target document from among the plurality of documents by means of the updated data for retrieval. A reference document is obtained by means of coarse ranking retrieval and fine ranking retrieval, and thus, the accuracy of the reference document is guaranteed; the reference document is used for updating data for retrieval, so that positive and negative feedback interactions are achieved in the retrieval pipeline, making the data for retrieval more accurate, effectively solving retrieval errors caused by expression diversity and indirectness, and improving the accuracy of document retrieval.
Owner:ALIBABA (CHINA) CO LTD

Retrieval method and device based on document segmentation and document retrieval system

The invention provides a retrieval method and device based on document segmentation and a document retrieval system. The method comprises the following steps: acquiring a to-be-segmented document; based on an NLP algorithm, calculating the semantic similarity between the partial texts of the to-be-segmented document to obtain a first semantic relevancy; according to the first semantic relevancy of all the partial texts, the document to be segmented is segmented, a plurality of semantic text blocks are obtained, and each semantic text block comprises at least one partial text; under the condition that a query request is received, calculating semantic similarity between a query text corresponding to the query request and each semantic text block based on an NLP algorithm to obtain a plurality of second semantic relevancy, and determining the semantic text block with the highest second semantic relevancy of the query text corresponding to the query request as a target semantic text block, and displaying the target semantic text block in a display interface. According to the scheme, the problem that in the prior art, the accuracy rate is low during text retrieval is solved.
Owner:中国邮政储蓄银行股份有限公司

Water conservancy design file retrieval system and method based on local lightweight large model

The invention discloses a water conservancy design archive retrieval system and method based on a local lightweight large model, and the method comprises the steps: S1, constructing a Python automatic preprocessing assembly line, extracting texts for PDF and Word multi-format archives, correcting metadata, and outputting standardized data; s2, constructing a full-text retrieval and semantic retrieval dual-mode cross-document retrieval service by relying on a Weavi ate local vector database and a lightweight text embedding model; s3, analyzing a user query intention through a local large model, synchronously triggering metadata accurate retrieval and content semantic retrieval, and generating a structured result; and S4, integrating the core module into a local area network Web platform, adopting Docker containerization deployment, and combining an RBAC permission model and JWT authentication to guarantee security. The system comprises a preprocessing module, a cross-document retrieval module, an intelligent agent module and a background management module, and collaboration is achieved through a standardized API. According to the method, the problem of archive fragmentation is solved, multi-mode retrieval breaks through keyword limitation, an intelligent agent reduces manual intervention, a localized architecture prevents secret-related leakage, background management adapts to an existing I T environment, and full-process intelligent archive service is provided for water conservancy design.
Owner:ZHONGSHAN WATER CONSERVANCY PROJECT SURVEY & CONSULT CO LTD

Enterprise-level large model agent application system supporting multi-modal collaboration

The invention relates to the technical field of artificial intelligence, and discloses an enterprise-level large-model agent application system supporting multi-modal collaboration, and the system comprises a user interaction terminal, an agent engine server, a knowledge engine server, a plug-in integration center, and a distributed storage unit. The agent engine server is responsible for intention recognition and task arrangement of a multi-modal input signal, and dynamically loads a differential reasoning strategy based on an environment isolation mechanism. And the knowledge engine server constructs a cross-modal semantic anchor point, analyzes an unstructured document into a tetrad knowledge unit, and realizes accurate recall of images and texts by using a hybrid retrieval algorithm. And the plug-in integration center executes outbound replacement and inbound restoration of the sensitive data through the context-aware dynamic desensitization gateway. According to the method, through a multi-modal semantic association and closed-loop verification mechanism, the problems of low complex document retrieval precision and leakage of external calling data are solved, and the service processing capacity and safety of the system are improved.
Owner:LINGRUIDA (XIAMEN) TECHNOLOGY CO LTD

Dynamic depth document retrieval for enterprise language model systems

Systems and methods for resource-efficient retrieval of information using a generative AI model are disclosed. An input query requesting information from a set of documents is used in a prompt for a generative AI model to generate a search query to identify the documents relevant to the input query and their respective relevancy scores. The input query is used as an input another model to determine a depth score indicating a predicted number of documents needed to retrieve the information. Based on the depth score and the relevancy scores of the relevant documents, the system extracts grounding data from the identified relevant documents to generate an answer synthesis prompt for the generative AI model. The generative AI model processes the second to produce a response to the input query including the requested information.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Intelligent document duplicate checking system and method based on vector database and large language model

The invention discloses an intelligent document duplicate checking system and method based on a vector database and a large language model, and the method comprises the following steps: S1, collecting document data in various formats, and carrying out the preprocessing of the document data; s2, semantic coding is carried out through a large language model, and a document semantic vector is generated; s3, storing the document semantic vector into a vector database, constructing a vector index and recording historical query data; s4, carrying out preliminary candidate document retrieval, and carrying out approximate nearest neighbor retrieval based on outlier identification; s5, calculating the similarity between the candidate document and the document to be subjected to duplicate checking, and screening a final high-similarity document; s6, generating a duplicate checking report, and recording user operation behaviors; and S7, receiving user feedback, and dynamically updating the document semantic vector and the vector index. According to the method, efficient and accurate intelligent document duplicate checking is realized by utilizing the large language model and the vector database, the semantic matching capability is improved, the duplicate checking efficiency is optimized, and the intelligence and adaptability of a duplicate checking system are improved.
Owner:BEIJING RONGJIA HECHUANG TECHNOLOGY CO LTD

Full-text retrieval method and system fusing various types of documents

The invention provides a full-text retrieval method and system fusing various types of documents, and relates to the technical field of information retrieval, and the method comprises the following steps: obtaining document representation through document content extraction and structure recognition, generating a cross-modal semantic vector by using word embedding and nonlinear transformation, constructing a hierarchical index and a cross-document association graph, and obtaining a full-text retrieval result; the basic correlation score is calculated after the query request is received, and the comprehensive score of the candidate content segments is calculated based on the association graph to determine the optimal retrieval result, so that unified representation and retrieval of heterogeneous documents are realized, the cross-document retrieval precision and relevance are improved, and the processing capability of a retrieval system on complex queries is enhanced.
Owner:BEIJING CHANGFA TECH CO LTD

Multimodal Data Ingestion And Retrieval For Agent Systems

Techniques for multimodal document retrieval are disclosed herein. Multimodal documents that include both textual and graphical components are retrieved from a knowledge base by a multimodal retrieval augmented generation (RAG) agent in response to a query. The documents and / or components or chunks thereof are retrievable by the RAG agent from the knowledge base using the semantic summaries and / or vector search of embeddings in the knowledge base that are generated from text extracted from processing non-textual components of the data. The RAG agent classifies the query type to determine whether to use a semantic match for text or image summaries, full text semantic search, vector cosine similarity search, and / or other multimodal vector search. The RAG agent performs types of searches selected based on the modality used to generate the response to the query.
Owner:ORACLE INT CORP

Intelligent literature screening method and system based on zero sample voting device

The invention discloses an intelligent literature screening method and system based on zero sample voters. The method comprises the following steps: expanding a literature set based on seed literatures, constructing a classification system comprising a plurality of zero sample voters, screening literatures related to a target topic through a majority voting mechanism, and iteratively expanding the literatures to obtain more literatures related to the target topic. The zero sample voting device integrates subject term matching, language model pre-training, multi-large language model voting and a semantic embedding vector technology, adaptively determines a threshold value in combination with an improved binary search algorithm, and realizes high-precision classification without training data. The system supports parallel processing and intelligent API management, and the processing efficiency of a large-scale literature set is remarkably improved. The method is suitable for the fields of academic research, literature retrieval and the like, can quickly construct a large-scale high-correlation literature information base, and particularly meets the retrieval requirements of emerging technologies or interdisciplinary themes.
Owner:NANJING UNIV OF POSTS & TELECOMM

Retrieval enhancement generation system and method for self-defined multi-level intention recognition

The invention discloses a retrieval enhancement generation system and method for customizing multilevel intention recognition. The retrieval enhancement generation system comprises a multilevel intention label and document management module, a multilevel intention recognition module and a retrieval enhancement generation module. The multi-level intention label and document management module can self-define and build a corresponding relationship among a multi-level intention label, an intention label and a document level by level to form a label document corresponding relationship form; the multi-level intention recognition module recognizes intentions of the user level by level according to the multi-level intention labels and the multi-level intention labels established by the document management module, and gives corresponding intention labels; and the retrieval enhancement generation module performs vector similarity search in a document range delineated by the intention tag, and transmits a vector similarity search result to the large language model for result generation. Through step-by-step subdivision of the user intention, the document retrieval range is gradually reduced to the corresponding intention label, so that the answer generated by the large language model in the generation stage is more accurate.
Owner:SHENZHEN XICHEN SOFTWARE TECHNOLOGY CO LTD

Literature information retrieval and analysis system and method based on AI intelligence

ActiveCN120492636AEnergy efficient computingText database indexingUser needsDocument representation
The invention relates to the technical field of literature information retrieval, and particularly discloses a literature information retrieval analysis system and method based on AI intelligence. The system comprises a representation vector output module, a preliminary retrieval result output module, a retrieval result updating module, a dynamic knowledge graph construction module, an enhanced literature representation output module, a recommendation list output module and a knowledge graph feedback updating module. A representation vector output module extracts literature text features; a preliminary retrieval result output module executes semantic enhancement retrieval and query expansion; a retrieval result updating module optimizes a preliminary result; a dynamic knowledge graph construction module constructs a graph based on the retrieval result; the recommendation list output module provides personalized recommendation, and the knowledge graph feedback updating module updates the knowledge graph according to the user interaction flow data. According to the method, the accuracy, intelligence and dynamic adaptability of literature retrieval are comprehensively improved, and user requirements can be deeply met.
Owner:ZOUPING KEHUI INFORMATION CONSULTING CO LTD +1

Document retrieval method and system based on electric power semantic enhancement and electronic equipment

The invention relates to a document retrieval method and system based on electric power semantic enhancement and electronic equipment, belongs to the technical field of natural language processing, and solves the problem of low retrieval accuracy caused by low complex knowledge utilization rate and insufficient electric power professional semantic understanding in the prior art. Comprising the following steps: receiving user query content, and obtaining a query embedding vector by utilizing a modal joint embedding model; based on the electric power knowledge graph, utilizing a large language model and a text embedding model to obtain a structured query vector of user query content; according to the query embedded vector and the structured query vector, obtaining a plurality of candidate documents and document-level similarity scores and page-level similarity scores thereof, and further obtaining a comprehensive similarity score of each candidate document by using a double-path prediction model; and obtaining a total score according to the document-level similarity score, the page-level similarity score and the comprehensive similarity score of each candidate document, and selecting a plurality of candidate documents with the highest total score as a retrieval result. And the retrieval precision is improved.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Automatic network protocol testing method and system based on large language model

The invention provides an automatic network protocol testing method and system based on a large language model. According to the method, through three core technologies of structured protocol specification extraction, hybrid test case generation and retrieval feedback enhanced code generation, full-process automation from a protocol specification document to an executable test code is realized. The method comprises the following steps: firstly, carrying out preprocessing and structured analysis on an RFC document, and extracting protocol field information and state machine description; a complete test case covering normal, abnormal and boundary conditions is automatically generated based on protocol elements; and finally, an initial test code is generated in combination with a test equipment API document retrieval result, and multi-round iterative optimization is performed through running log analysis and an experience knowledge base. Compared with the prior art, according to the scheme, an end-to-end automatic test work chain is achieved, new protocol testing can be adapted without manual protocol modeling, meanwhile, a large language model and expert experience are fused through a feedback mechanism, and the test efficiency and the code quality are remarkably improved.
Owner:TSINGHUA UNIVERSITY

Method and system for intelligent retrieval and review of multi-modal electric power engineering document

The invention relates to a method and a system for intelligent retrieval and review of a multi-modal electric power engineering document, belongs to the technical field of natural language processing, and solves the problems of poor relevance and incomplete review of an existing multi-modal document retrieval result. The method comprises the following steps: performing logic region division on each electric power engineering document, extracting a multi-modal fusion vector of each logic region, and aggregating to obtain a document feature vector; calculating the multi-level similarity between the received query word and each electric power engineering document, and obtaining a plurality of electric power engineering documents with the highest similarity as retrieval documents; and dynamically expanding each rule in the rule base, and performing compliance review on each retrieval document by utilizing the expanded rule according to the document feature vector of the retrieval document to generate a review report. The efficient and accurate retrieval of the multi-modal electric power engineering document is realized, and the comprehensiveness and flexibility of review are improved.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Method and device for acquiring literature information and automatically constructing literature knowledge graph

The invention relates to the technical field of natural language processing, in particular to a literature information acquisition and literature knowledge graph automatic construction method and device.The method comprises the steps that at least one literature retrieval word provided by a user is used for retrieving at least one related literature, generating text data of different data types according to the literature content of the at least one related literature; based on the target large language model, constructing a knowledge graph according to the text data of different data types; and integrating the at least one target knowledge graph and the knowledge graph to construct a final knowledge graph. Therefore, the technical problems that a literature extraction technology in related technologies cannot fully understand technical terms and complex concepts in the scientific field, and integration of heterogeneous data and mining of cross-domain knowledge are difficult to effectively cope with are solved.
Owner:TSINGHUA UNIVERSITY

Method, device and equipment for generating reply information based on large language model

The invention provides a method, a device and equipment for generating reply information based on a large language model, and relates to the technical field of artificial intelligence, in particular to the fields of document retrieval, natural language processing and large language models. According to the implementation scheme, in response to a received question text of a user, a semantic vector of the question text and event information related to a specific field are obtained; based on at least two of the semantic vector of the question text, the at least one piece of argument information and the event category, obtaining a plurality of candidate documents in a document library of a specific field; for a candidate document in the plurality of candidate documents, determining quality evaluation information of the candidate document based on the event category; and determining at least one target document in the plurality of candidate documents based on the relevancy between the candidate documents and the question text and the quality evaluation information of the candidate documents, so as to obtain reply information for replying the question text based on the at least one target document.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Retrieval enhancement generation method and system based on sparse graph neural network

The invention discloses a retrieval enhancement generation method and system based on a sparse graph neural network, and aims to solve the problems of large-scale knowledge graph retrieval enhancement generation and a multi-hop reasoning task. The method comprises the following steps: embedding a query vector, receiving a user query, and encoding the user query into a query embedded vector through a sentence embedding model; performing sub-graph retrieval, performing hash processing on the query embedded vector by using a hierarchical hash index to generate a composite hash key, and matching entities in the knowledge graph layer by layer according to the queried hash key to construct a sub-graph; dynamic sparse message transmission: executing dynamic sparse message transmission update entity embedding on the constructed subgraph; document retrieval sorting is carried out, document correlation scores are calculated according to final entity embedding, and first K documents are selected to form a retrieval result document set; and answer generation: inputting the query and retrieval result document set into a generation model, and outputting a final answer. The method is suitable for large-scale knowledge graph application, and an efficient retrieval enhancement generation solution is provided.
Owner:GUANGZHOU ZHONGKE YIDE TECH CO LTD

Index optimization and compression storage system and method for large-scale literature set

The invention discloses an index optimization and compression storage system and method for a large-scale literature set, and the method comprises the following steps: S1, collecting and preprocessing literature data, and generating a standardized text data set; s2, carrying out keyword semantic vector coding, and constructing a keyword semantic vector matrix; s3, constructing an initial Gaussian mixture model to obtain a clustering center, a covariance matrix and a weight; s4, introducing a sea elephant optimization algorithm to optimize clustering parameters, and outputting an optimal clustering result; s5, constructing a semantic clustering structure, and generating an index tree structure; s6, performing bitmap compression and inverted coding, and constructing an index table supporting Boolean logic; and S7, dynamically accessing the newly added literature, and completing incremental updating of the index structure. The method is used for improving the index construction efficiency and the storage compression rate of a large-scale literature set, and efficient and semantic literature retrieval service capable of being incrementally updated is achieved.
Owner:CENTRAL COMPILATION & TRANSLATION PRESS CO LTD

Document retrieval method and device, equipment, medium and program product

The invention discloses a document retrieval method and device, equipment, a medium and a program product, and relates to the technical field of document processing. The method comprises the steps of obtaining a query statement of a document retrieval party, at least two candidate documents and candidate abstract vectors and candidate full-text vectors of the candidate documents; performing keyword extraction on the query statement to obtain a query keyword, and performing vectorization on the query statement to obtain a query statement vector; according to the query keyword, the query statement vector, each candidate document and a candidate abstract vector of each candidate document, performing coarse screening on each candidate document to obtain at least two coarse screening documents; and according to the query statement vector and the candidate full-text vector of each coarse screening document, performing fine arrangement on each coarse screening document to obtain a target retrieval document. According to the technical scheme provided by the embodiment of the invention, the accuracy of document retrieval is improved.
Owner:AGRICULTURAL BANK OF CHINA

Intelligent data report generation system and method based on large language model

The invention discloses an intelligent data report generation system and method based on a large language model, and relates to the field of computer software and artificial intelligence. A natural language processing module in the system is used for analyzing user requirements and judging whether a data source of the user requirements is a data storage module or an enterprise knowledge base module; the database query module is used for generating structured data of the SQL statement query data storage module; the database query module queries a database; the knowledge document retrieval module retrieves unstructured data of the enterprise knowledge base module; the knowledge document retrieval module retrieves a knowledge base by using an RAG technology; the report generation module is used for generating a visual chart according to a user demand and a query result; the data storage module is used for storing the structured data and the generated visual chart; and the enterprise knowledge base module stores the vectorized unstructured knowledge document. According to the method, the data report making and querying efficiency and accuracy can be improved.
Owner:ZHEJIANG GONGSHANG UNIVERSITY

Retrieval enhancement generation method and data set generation method for time-sensitive problems

The invention discloses a retrieval enhancement generation method for a time-sensitive problem, which comprises the following steps of: mixed time perception retrieval: enhancing document retrieval by adding time constraint on the basis of semantic relevance, and guiding by a time card to ensure that the retrieved document not only conforms to the meaning of query, but also conforms to the semantic relevance; the time context is met; the progressive multi-step reflection comprises the following steps of: firstly, acquiring and evaluating an initial document set by applying mixed time perception retrieval; if a document is retrieved, generating a final answer by using a large language model; otherwise, entering a reflection stage, and summarizing useful time information in the retrieved document into a context; and merging document sets accumulated in all iterations to generate a final answer. According to the method, a new framework integrating dynamic knowledge updating and time reasoning into the retrieval and generation process is provided, and accurate and timely response can be made to time-related problems.
Owner:NAT UNIV OF DEFENSE TECH

Multi-view knowledge intensive retrieval enhancement generation system and method

The invention relates to the field of retrieval enhancement, in particular to a multi-view knowledge-intensive retrieval enhancement generation system and method, which are characterized in that structural vectors and semantic topics are extracted from professional corpora through principal component analysis and non-negative matrix factorization technologies, and a multi-dimensional professional view set is constructed. After a user query is received, a potential intention is identified, a view angle weight vector is generated, the query is rewritten according to the view angle weight vector, multiple groups of view angle retrieval requests are constructed, targeted document retrieval is executed, and a structured prompt input language generation model is constructed based on a view angle weight reordering result and fusion of an original query and a multi-view angle retrieval result. The method is suitable for scenes of law assistance, intelligent diagnosis, academic questions and answers and the like, so that the accuracy, the interpretation and the reliability of retrieval and generation in the complex field are remarkably improved.
Owner:BEIHANG UNIV

Literature retrieval method based on large language model

The invention discloses a literature retrieval method based on a large language model, and belongs to the technical field of large language models. The method comprises the steps that the large language model obtains a natural language text from a user; according to the natural language text, a domain knowledge base of the target domain is constructed, and the domain knowledge base comprises a core definition corresponding to the target domain, a multi-dimensional lexicon, a theme retrieval formula, an ambiguity lexicon and a constraint rule base; performing literature retrieval according to a theme retrieval formula in the domain knowledge base to obtain a first literature set of the target domain; based on a core definition, an ambiguous word library and a constraint rule in the domain knowledge base, screening literatures in the first literature set to obtain a second literature set; and generating a literature analysis report based on the second literature set. According to the method, large-range and high-accuracy literature retrieval can be realized.
Owner:WUHAN UNIV