Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

264 results about "Document retrieval" patented technology

Document retrieval is defined as the matching of some stated user query against a set of free-text records. These records could be any type of mainly unstructured text, such as newspaper articles, real estate records or paragraphs in a manual. User queries can range from multi-sentence full descriptions of an information need to a few words.

Enhanced document retrieval with semantic depth and syntactic structure

Certain aspects of the present disclosure describe a method of information retrieval. In certain aspects, the method includes identifying a set of relevant nodes of a document graph embedding semantic units associated with the document based on a document search query. The method further includes reconstructing a structural context for each relevant node in the set of relevant nodes. The method further includes processing the set of relevant nodes and the structural context of each relevant node with a large language model to generate a contextual response to the document search query.
Owner:INTUIT INC

Water conservancy design file retrieval system and method based on local lightweight large model

The invention discloses a water conservancy design archive retrieval system and method based on a local lightweight large model, and the method comprises the steps: S1, constructing a Python automatic preprocessing assembly line, extracting texts for PDF and Word multi-format archives, correcting metadata, and outputting standardized data; s2, constructing a full-text retrieval and semantic retrieval dual-mode cross-document retrieval service by relying on a Weavi ate local vector database and a lightweight text embedding model; s3, analyzing a user query intention through a local large model, synchronously triggering metadata accurate retrieval and content semantic retrieval, and generating a structured result; and S4, integrating the core module into a local area network Web platform, adopting Docker containerization deployment, and combining an RBAC permission model and JWT authentication to guarantee security. The system comprises a preprocessing module, a cross-document retrieval module, an intelligent agent module and a background management module, and collaboration is achieved through a standardized API. According to the method, the problem of archive fragmentation is solved, multi-mode retrieval breaks through keyword limitation, an intelligent agent reduces manual intervention, a localized architecture prevents secret-related leakage, background management adapts to an existing I T environment, and full-process intelligent archive service is provided for water conservancy design.
Owner:ZHONGSHAN WATER CONSERVANCY PROJECT SURVEY & CONSULT CO LTD

Enterprise-level large model agent application system supporting multi-modal collaboration

The invention relates to the technical field of artificial intelligence, and discloses an enterprise-level large-model agent application system supporting multi-modal collaboration, and the system comprises a user interaction terminal, an agent engine server, a knowledge engine server, a plug-in integration center, and a distributed storage unit. The agent engine server is responsible for intention recognition and task arrangement of a multi-modal input signal, and dynamically loads a differential reasoning strategy based on an environment isolation mechanism. And the knowledge engine server constructs a cross-modal semantic anchor point, analyzes an unstructured document into a tetrad knowledge unit, and realizes accurate recall of images and texts by using a hybrid retrieval algorithm. And the plug-in integration center executes outbound replacement and inbound restoration of the sensitive data through the context-aware dynamic desensitization gateway. According to the method, through a multi-modal semantic association and closed-loop verification mechanism, the problems of low complex document retrieval precision and leakage of external calling data are solved, and the service processing capacity and safety of the system are improved.
Owner:LINGRUIDA (XIAMEN) TECHNOLOGY CO LTD

Dynamic depth document retrieval for enterprise language model systems

Systems and methods for resource-efficient retrieval of information using a generative AI model are disclosed. An input query requesting information from a set of documents is used in a prompt for a generative AI model to generate a search query to identify the documents relevant to the input query and their respective relevancy scores. The input query is used as an input another model to determine a depth score indicating a predicted number of documents needed to retrieve the information. Based on the depth score and the relevancy scores of the relevant documents, the system extracts grounding data from the identified relevant documents to generate an answer synthesis prompt for the generative AI model. The generative AI model processes the second to produce a response to the input query including the requested information.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Full-text retrieval method and system fusing various types of documents

The invention provides a full-text retrieval method and system fusing various types of documents, and relates to the technical field of information retrieval, and the method comprises the following steps: obtaining document representation through document content extraction and structure recognition, generating a cross-modal semantic vector by using word embedding and nonlinear transformation, constructing a hierarchical index and a cross-document association graph, and obtaining a full-text retrieval result; the basic correlation score is calculated after the query request is received, and the comprehensive score of the candidate content segments is calculated based on the association graph to determine the optimal retrieval result, so that unified representation and retrieval of heterogeneous documents are realized, the cross-document retrieval precision and relevance are improved, and the processing capability of a retrieval system on complex queries is enhanced.
Owner:BEIJING CHANGFA TECH CO LTD

Multimodal Data Ingestion And Retrieval For Agent Systems

Techniques for multimodal document retrieval are disclosed herein. Multimodal documents that include both textual and graphical components are retrieved from a knowledge base by a multimodal retrieval augmented generation (RAG) agent in response to a query. The documents and / or components or chunks thereof are retrievable by the RAG agent from the knowledge base using the semantic summaries and / or vector search of embeddings in the knowledge base that are generated from text extracted from processing non-textual components of the data. The RAG agent classifies the query type to determine whether to use a semantic match for text or image summaries, full text semantic search, vector cosine similarity search, and / or other multimodal vector search. The RAG agent performs types of searches selected based on the modality used to generate the response to the query.
Owner:ORACLE INT CORP

Retrieval enhancement generation system and method for self-defined multi-level intention recognition

The invention discloses a retrieval enhancement generation system and method for customizing multilevel intention recognition. The retrieval enhancement generation system comprises a multilevel intention label and document management module, a multilevel intention recognition module and a retrieval enhancement generation module. The multi-level intention label and document management module can self-define and build a corresponding relationship among a multi-level intention label, an intention label and a document level by level to form a label document corresponding relationship form; the multi-level intention recognition module recognizes intentions of the user level by level according to the multi-level intention labels and the multi-level intention labels established by the document management module, and gives corresponding intention labels; and the retrieval enhancement generation module performs vector similarity search in a document range delineated by the intention tag, and transmits a vector similarity search result to the large language model for result generation. Through step-by-step subdivision of the user intention, the document retrieval range is gradually reduced to the corresponding intention label, so that the answer generated by the large language model in the generation stage is more accurate.
Owner:SHENZHEN XICHEN SOFTWARE TECHNOLOGY CO LTD

Document retrieval method and system based on electric power semantic enhancement and electronic equipment

The invention relates to a document retrieval method and system based on electric power semantic enhancement and electronic equipment, belongs to the technical field of natural language processing, and solves the problem of low retrieval accuracy caused by low complex knowledge utilization rate and insufficient electric power professional semantic understanding in the prior art. Comprising the following steps: receiving user query content, and obtaining a query embedding vector by utilizing a modal joint embedding model; based on the electric power knowledge graph, utilizing a large language model and a text embedding model to obtain a structured query vector of user query content; according to the query embedded vector and the structured query vector, obtaining a plurality of candidate documents and document-level similarity scores and page-level similarity scores thereof, and further obtaining a comprehensive similarity score of each candidate document by using a double-path prediction model; and obtaining a total score according to the document-level similarity score, the page-level similarity score and the comprehensive similarity score of each candidate document, and selecting a plurality of candidate documents with the highest total score as a retrieval result. And the retrieval precision is improved.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Document retrieval method and device, equipment, medium and program product

The invention discloses a document retrieval method and device, equipment, a medium and a program product, and relates to the technical field of document processing. The method comprises the steps of obtaining a query statement of a document retrieval party, at least two candidate documents and candidate abstract vectors and candidate full-text vectors of the candidate documents; performing keyword extraction on the query statement to obtain a query keyword, and performing vectorization on the query statement to obtain a query statement vector; according to the query keyword, the query statement vector, each candidate document and a candidate abstract vector of each candidate document, performing coarse screening on each candidate document to obtain at least two coarse screening documents; and according to the query statement vector and the candidate full-text vector of each coarse screening document, performing fine arrangement on each coarse screening document to obtain a target retrieval document. According to the technical scheme provided by the embodiment of the invention, the accuracy of document retrieval is improved.
Owner:AGRICULTURAL BANK OF CHINA

Retrieval enhancement generation method and data set generation method for time-sensitive problems

The invention discloses a retrieval enhancement generation method for a time-sensitive problem, which comprises the following steps of: mixed time perception retrieval: enhancing document retrieval by adding time constraint on the basis of semantic relevance, and guiding by a time card to ensure that the retrieved document not only conforms to the meaning of query, but also conforms to the semantic relevance; the time context is met; the progressive multi-step reflection comprises the following steps of: firstly, acquiring and evaluating an initial document set by applying mixed time perception retrieval; if a document is retrieved, generating a final answer by using a large language model; otherwise, entering a reflection stage, and summarizing useful time information in the retrieved document into a context; and merging document sets accumulated in all iterations to generate a final answer. According to the method, a new framework integrating dynamic knowledge updating and time reasoning into the retrieval and generation process is provided, and accurate and timely response can be made to time-related problems.
Owner:NAT UNIV OF DEFENSE TECH

Multi-view knowledge intensive retrieval enhancement generation system and method

The invention relates to the field of retrieval enhancement, in particular to a multi-view knowledge-intensive retrieval enhancement generation system and method, which are characterized in that structural vectors and semantic topics are extracted from professional corpora through principal component analysis and non-negative matrix factorization technologies, and a multi-dimensional professional view set is constructed. After a user query is received, a potential intention is identified, a view angle weight vector is generated, the query is rewritten according to the view angle weight vector, multiple groups of view angle retrieval requests are constructed, targeted document retrieval is executed, and a structured prompt input language generation model is constructed based on a view angle weight reordering result and fusion of an original query and a multi-view angle retrieval result. The method is suitable for scenes of law assistance, intelligent diagnosis, academic questions and answers and the like, so that the accuracy, the interpretation and the reliability of retrieval and generation in the complex field are remarkably improved.
Owner:BEIHANG UNIV

Patent literature retrieval system, method and equipment and storage medium

The invention discloses a patent literature retrieval system, method and equipment and a storage medium, and relates to the field of data processing. The system comprises a data acquisition module for acquiring a retrieval request and historical behavior data of a user; the intention analysis module is used for analyzing the current request through a semantic analysis model and a technical field topic model so as to determine a user intention and an intention matching degree; the user behavior factor generation engine extracts dominant and recessive behavior characteristics from the historical data, and inputs the dominant and recessive behavior characteristics into a pre-trained user portrait model to generate user behavior factors; a recommendation generation engine recall candidate patents according to the intention of the user, and an authority calculation module calculates the authority of each patent according to a patent citation network; and the patent recommendation module fuses the user behavior factors, the user intention matching degree and the authority degree, and comprehensively sorts candidate patents to generate a recommended patent list. The method can meet the precision requirements of different users, and reduces the deviation between the patent recommendation result and the real expectation of the user.
Owner:QIZHI TECH CO LTD

Multimodal Data Ingestion And Retrieval For Agent Systems

Techniques for multimodal document retrieval are disclosed herein. Multimodal documents that include both textual and graphical components are retrieved from a knowledge base by a multimodal retrieval augmented generation (RAG) agent in response to a query. The documents and / or components or chunks thereof are retrievable by the RAG agent from the knowledge base using the semantic summaries and / or vector search of embeddings in the knowledge base that are generated from text extracted from processing non-textual components of the data. The RAG agent classifies the query type to determine whether to use a semantic match for text or image summaries, full text semantic search, vector cosine similarity search, and / or other multimodal vector search. The RAG agent performs types of searches selected based on the modality used to generate the response to the query.
Owner:ORACLE INT CORP

Cerebral stroke knowledge question-answering system construction method and system based on knowledge graph and large language model

The invention relates to the technical field of medical health information services, in particular to a cerebral apoplexy knowledge question-answering system construction method and system based on a knowledge graph and a large language model. Natural language input of a user is analyzed through a query agent, a query intention and constraint conditions are recognized, and a structured execution plan is generated; a user state management tool is forcibly activated, a static clinical portrait and a dynamic rehabilitation log are loaded, and a personalized context is constructed; a plurality of tools such as knowledge graph query, authoritative literature retrieval and rehabilitation plan generation are scheduled, and accurate retrieval and reasoning of heterogeneous knowledge are completed; through double verification of fact consistency and clinical risks, error or high-risk suggestions are intercepted and replaced with risk early warning. The problems of'illusion 'risk, insufficient individuation, poor interpretability and the like of a traditional single model are solved to a large extent, high-credibility, individuation and traceable rehabilitation knowledge service can be provided for the stroke patient and a caregiver of the stroke patient, and rehabilitation safety and effect are guaranteed.
Owner:DALIAN UNIV

RAG-based pdf intelligent retrieval and generation method and system

The application discloses a kind of PDF intelligent retrieval and generation method and system based on RAG, by obtaining the document data of input, using the classification model established in advance to parse document data, extract text content and image content to form first data set;Using deep learning model to the image content in first data set carries out feature extraction, while the text content in first data set applies natural language processing technology to carry out semantic analysis, obtains multimodal feature set;According to multimodal feature set, application information integration algorithm is uniformly encoded and is handled to generate second data set, if detecting the integrity of fusion feature vector in second data set is lower than preset threshold value, then supplementary context semantic analysis fills in missing information;Using preset index construction mechanism to the clustering processing of fusion feature vector in second data set, generates the retrieval index library containing classification index structure.The application improves the accuracy and comprehensiveness of document retrieval.
Owner:HUNAN ZHIXUE YOUKE INFORMATION TECHNOLOGY CO LTD +1

Machine learning-based literature search and retrieval and related machine learning model training methods

Described herein are systems and methods for performing literature retrieval and related machine learning model training methods. An example computer-implemented method of training a machine learning model configured for literature retrieval I includes receiving a plurality of full-text articles; extracting, from the plurality of full-text articles, a plurality of positive sentence-citation pairs, each positive sentence-citation pair comprising a respective citing sentence and at least one cited article that is associated with the respective citing sentence; creating a labeled dataset comprising the plurality of positive sentence-citation pairs; and training a machine learning model using the labeled dataset.
Owner:FLORIDA STATE UNIV RES FOUND INC

Two-step literature retrieval method and system based on keyword extension

The invention relates to the technical field of literature retrieval, in particular to a two-step literature retrieval method and system based on keyword extension. Comprising the following steps: calibrating an initial word, inputting a user keyword, querying a preset high-frequency subject word bank, and executing a calibration rule; preliminary retrieval: performing retrieval in an AuthorKeywords field of a target database by using the calibrated seed keyword to obtain a preliminary retrieval result set; and extracting and purifying extended keywords. According to the two-step method provided by the invention, through automatic expansion and parallel field retrieval, wider coverage can be realized in a single process, it is expected that operation rounds required by a user for repeatedly trying different keyword combinations can be obviously reduced, and the automatic expansion keyword extraction, filtering and multi-field parallel retrieval mechanism can improve the efficiency of the user. According to the method provided by the invention, a relatively wide literature range can be covered by single execution, so that a relatively comprehensive retrieval result can be expected to be obtained more efficiently through a structured automatic process, and the time cost of repeated trial and error and screening of researchers is saved.
Owner:BEIJING TECH & BUSINESS UNIV

Automated report generation using retrieval augmented system and large language model

A method includes creating a document retrieval and large language model architecture including at least one vector database including vectorized data corresponding to one or more documents from one or more document storage locations and a large language model. The method also includes receiving a query to generate a report associated with a current project using the large language model. The method also includes returning, in response to the query, a relevant context generated using the at least one vector database. The method also includes generating and outputting, using the large language model and based on the relevant context, one or more portions of the report.
Owner:HAMILTON SUNDSTRAND CORP

Multi-hop reasoning method based on dynamic reasoning guidance and multistage self-feedback retrieval

A multi-hop reasoning method based on dynamic reasoning guidance and multistage self-feedback retrieval belongs to the field of natural language processing, and comprises the following steps: deconstructing a multi-hop reasoning process into a target-oriented sequence decision problem, carrying out dynamic reasoning guidance by using a large language model, generating a sub-problem sequence matched with a reasoning progress in real time, and carrying out multi-level self-feedback retrieval on the sub-problem sequence; target document retrieval is guided, and sub-questions are dynamically generated; according to the generated sub-questions, obtaining associated documents by adopting a three-level collaborative retrieval mechanism; and performing information refining on the associated document through a large language model, fusing the refined information into an inference chain, and performing inference to generate an answer. The invention further discloses a multi-hop reasoning system, a storage medium and a computer program product. The method aims at solving the complex multi-hop problem that multiple dispersed knowledge fragments need to be integrated, high-accuracy and high-efficiency reasoning is achieved, the retrieval requirement is dynamically generated through an explicit thinking chain guiding mechanism, and evidence obtaining is optimized and redundant information is filtered in combination with a three-level self-feedback retrieval mechanism.
Owner:XI AN JIAOTONG UNIV

Systems and methods for role-based access control (RBAC) using large language model (LLM) embeddings

Systems and methods for Role-Based Access Control (RBAC) using Large Language Model (LLM) embeddings are described. In an illustrative, non-limiting embodiment, an Information Handling System (IHS) may include a processor and a memory coupled to the processor. The memory may store program instructions that, upon execution, generate a plurality of distinct unstructured natural language documents from a same portion of structured data of an enterprise, with each document created for a corresponding role in the enterprise. The IHS may concatenate a role context vector that defines an access privilege for a document with a document vector associated with the document to produce a role-integrated document vector. The IHS may also apply pre-attention and post-attention layers to the role context vector to manage access control during document retrieval based on user roles.
Owner:DELL PROD LP

Document retrieval enhancement method and device, electronic equipment and storage medium

The invention relates to a document retrieval enhancement method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining query information in response to a retrieval request of a user, carrying out semantic analysis on the query information to obtain a query key attribute, and converting the query key attribute into a query semantic structure, performing matching in the knowledge graph based on the query semantic structure to obtain a matching entity node and a matching entity node relationship, and obtaining a plurality of candidate documents based on the matching entity and the matching entity relationship, and sorting the plurality of candidate documents to obtain a sorting result based on the query key attribute, the semantic similarity between the candidate documents and the relationship strength of the entity relationship in the knowledge graph corresponding to each candidate document, and outputting the plurality of candidate documents as retrieval results according to the sorting result. By adopting the technical scheme, the document retrieval accuracy can be improved, the document information extraction efficiency can be improved, and the intelligent sorting capability of retrieval results can be enhanced, so that the user experience of document retrieval and analysis can be remarkably optimized.
Owner:BEIJING TIELAN TECHNOLOGY CO LTD

Retrieval optimization method adaptive to multi-dimensional storage of power documents

A retrieval optimization method adaptive to multi-dimensional storage of power documents relates to the technical field of information retrieval, and comprises the following steps: firstly, preprocessing user query, identifying power business scenes and technical types of the user query, and extracting a query keyword set and a query vector; then, a three-level progressive retrieval strategy is adopted to retrieve in an electric power document library composed of a metadatabase and a vector data at the first level, candidate documents are screened based on keyword matching and business scenes; in the second stage, further filtering is carried out through similarity calculation of a query vector and a technical abstract vector; in the third stage, final accurate screening is completed in combination with full-text vector similarity and electric power professional rules; finally, information integration and structured output are conducted on the result, the problems that in traditional power document retrieval, the result is inaccurate, and efficiency is low are solved, and retrieval precision and response speed are remarkably improved.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Domain-specific retrieval language models

Various examples, systems, and methods are disclosed relating to domain-specific document retrieval that incorporates custom vocabulary integration and embedding model updates. A computing system can extract multiple segments from a collection of documents and generate queries that correspond to at least one segment. The computing system can identify terms that satisfy a uniqueness criterion and input the terms into a tokenizer to create a vocabulary dataset. The vocabulary dataset, the document segments, and the queries can be used to update an embedding model to support retrieval and semantic alignment within private documents.
Owner:NVIDIA CORP

Intelligent question answering method and system based on large model and retrieval enhancement

The invention discloses an intelligent question-answering method and system based on a large model and retrieval enhancement, and relates to the field of intelligent question-answering. The method comprises the following steps: performing preliminary retrieval on a preset knowledge base to obtain document fragments, and constructing a candidate version set; if it is determined that the query request does not belong to the cross-version query type, calculating a version consistency score, and determining a target version with the highest score; obtaining a target document fragment corresponding to the target version, and performing vectorization processing on the query request to generate a problem vector; performing vector retrieval in a preset vector knowledge base on the basis of the problem vector to obtain a problem vector result, and performing secondary retrieval in the preset knowledge base on the basis of the target version to obtain a document retrieval result; merging the problem vector result and the document retrieval result to obtain an initial candidate document; and generating a first final cue word, and inputting the first final cue word into a preset large language model to obtain a first target answer. By implementing the technical scheme provided by the invention, the accuracy of intelligent question answering is improved.
Owner:BEIJING ZHIBAO HUIZHONG DIGITAL TECHNOLOGY CO LTD

Literature comprehensive retrieval system and retrieval enhancement generation method thereof

The invention discloses a literature comprehensive retrieval system and a retrieval enhancement generation method thereof, and relates to the technical field of literature synthesis, and the method comprises the following steps: a demand anchoring layer carries out the retrieval of an initial retrieval request or a recursive feedback follow-up problem and a learning key point according to breadth and depth parameters; generating a corresponding number of query statements meeting a preset retrieval rule according to the query quantity; the resource acquisition layer retrieves and collects literatures and metadata in multiple channels according to query statements; the organization processing layer extracts a set number of documents, performs integration and structured extraction on contents of the documents, generates follow-up questions and learning key points, feeds back the follow-up questions and the learning key points to the demand anchoring layer, and synchronously retains the learning key points and the documents; and the result output layer converts the retained content into a target format. In this way, through combination of hierarchical cooperation and a recursion mechanism, traditional single retrieval limitation is broken, deviation caused by fuzzy requirements is avoided, retrieval accuracy, processing efficiency, information integrity and achievement usability are taken into consideration, and literature retrieval efficiency is improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Document retrieval method based on multi-field information and outlier detection

The invention relates to a document retrieval method based on multi-field information and outlier detection. The method comprises the following steps: acquiring document query information for retrieving documents, and acquiring field similarity between the document query information and different field information of each candidate document; if at least one high outlier exists in the current field similarity between the document query information and any one piece of field information of the current candidate document, obtaining the document similarity between the document query information and the current candidate document according to each high outlier; if the high outlier does not exist in the current field similarity, obtaining the document similarity between the document query information and the current candidate document according to at least one current field similarity; and obtaining a document retrieval result corresponding to the document query information according to the document similarity between the document query information and each candidate document. By adopting the method, the correlation of document retrieval results can be improved.
Owner:SHENZHEN LANLING SOFTWARE CO LTD

Technical supervision document retrieval method and related device

The invention provides a technical supervision document retrieval method and a related device, and belongs to the field of technical supervision document retrieval. The method comprises: acquiring a user query intention; retrieving from a database according to the query intention of the user to obtain a natural language answer facing the user question; the database construction method comprises the steps of obtaining an existing technical supervision document; performing analysis and knowledge extraction on the technical supervision document to obtain a knowledge multi-tuple; after the knowledge multi-tuple is verified, time dimension attributes are added to the knowledge multi-tuple; and respectively storing the knowledge multi-tuples added with the time dimension attributes according to data types to obtain a database. According to the technical supervision document retrieval method and device, the problem of low accuracy of technical supervision document retrieval is solved.
Owner:DATANG HYDROPOWER SCI & TECH RES INST CO LTD +2

Systems and Methods for Prompt-Based Query Generation for Diverse Retrieval

An example method for prompt-based query generation is provided. The method includes receiving, by a computing device, at least two prompts associated with a retrieval task to be performed on a corpus of documents associated with the task. The method includes applying, based on the at least two prompts and the corpus of documents, a large language model to generate a synthetic training dataset comprising a plurality of query-document pairs, wherein each query-document pair comprises a synthetically generated query and a document from the corpus of documents. The method includes training, on the plurality of query−document pairs from the synthetic training dataset, a document retrieval model to take an input query associated with the retrieval task and predict an output document retrieved from the corpus of documents. The method includes providing, by the computing device, the trained document retrieval model.
Owner:GOOGLE LLC

Knowledge graph multi-mode document analysis and image table semantization knowledge recall method

The invention discloses a knowledge graph multi-modal document analysis and image table semantization knowledge recall method, and belongs to the technical field of knowledge engineering and information retrieval. The invention provides an innovative scheme for fusing a visual language model, semantic abstract generation and knowledge graph modeling. The method comprises the following steps: constructing a vertical domain knowledge graph by adopting a BERT-BiLSTM-CRF model; according to the method, multi-modal document analysis is realized through models such as DocLayout-YOLO, TableMaster, UniMERNet and the like; the method comprises the following steps of: segmenting an image into 16 * 16 block sequences by adopting a vit-gpt2-image-adaptation model, and realizing image semantization through 768-dimensional vector space mapping and Transform coding; constructing a document summary tree based on DBSCAN clustering and LLM recursive summary; and designing a hybrid retrieval space fusing semantic vectors and structured vectors, and reordering by adopting a double-attention mechanism. According to the method, the knowledge base document retrieval recall rate is increased to 99%, the question and answer accuracy rate reaches 90% or above, the index construction time is shortened by 60%, and the problem that semantic understanding and recall of non-text elements in complex documents are difficult is effectively solved.
Owner:云鼎科技股份有限公司

Document retrieval method and system, storage medium and electronic equipment

The invention discloses a document retrieval method and system, a storage medium and electronic equipment. The method comprises the steps of obtaining a target query text input by a target object; a pre-trained generative document retrieval model is utilized to retrieve a target document identification lexical element sequence with the highest semantic relevance with the target query text from a target database, and the target database comprises a plurality of documents and document identification lexical element sequences corresponding to the documents; the generative document retrieval model is obtained by training the sequence generation model by using a comparative learning technology and a progressive learning strategy; and feeding back a target document corresponding to the target document identification lexical element sequence to the target object. According to the method and the device, the technical problem that the related retrieval model cannot quickly and accurately retrieve the document related to the query text semantics in the complex corpus is solved.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD