Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

70 results about "Relevance (information retrieval)" patented technology

In information science and information retrieval, relevance denotes how well a retrieved document or set of documents meets the information need of the user. Relevance may include concerns such as timeliness, authority or novelty of the result.

Full-text retrieval method and system fusing various types of documents

The invention provides a full-text retrieval method and system fusing various types of documents, and relates to the technical field of information retrieval, and the method comprises the following steps: obtaining document representation through document content extraction and structure recognition, generating a cross-modal semantic vector by using word embedding and nonlinear transformation, constructing a hierarchical index and a cross-document association graph, and obtaining a full-text retrieval result; the basic correlation score is calculated after the query request is received, and the comprehensive score of the candidate content segments is calculated based on the association graph to determine the optimal retrieval result, so that unified representation and retrieval of heterogeneous documents are realized, the cross-document retrieval precision and relevance are improved, and the processing capability of a retrieval system on complex queries is enhanced.
Owner:BEIJING CHANGFA TECH CO LTD

Systems, apparatuses, methods, and non-transitory computer-readable storage media for adaptive information retrieval for question-answering

Methods and systems for retrieving relevant information in response to an input question. The method includes obtaining text content related to the input question and partitioning the content into one or more paragraphs based on predefined rules. The method further involves extracting one or more evidence spans that are relevant to the input question by inputting the text content and the question into a trained language model. A semantic search is then performed on both the paragraphs and the extracted evidence spans, ranking the candidate passages based on their relevance to the input question. Each candidate passage may comprise either a paragraph or an evidence span that addresses the question. The disclosed methods and systems improve the quality and relevance of retrieved information by combining heuristic-based content partitioning with machine learning-based evidence extraction.
Owner:HUAWEI TECH CO LTD

Advanced Routing And Multi-Index Fusion For Enhanced Retrieval Augmented Generation

Techniques for multi-index retrieval in knowledge databases to enhance retrieval augmented generation (RAG) systems. The techniques involve a query processor that receives a query from a RAG agent, compares it to index summaries, and selects target indexes. The processor then searches these indexes, fuses the retrieved content items, and reranks the results before sending them back to the RAG agent. This approach combines query routing, multi-index fusion, and reranking to improve information retrieval for RAG applications. The technique offers several advantages, including enhanced retrieval efficiency through specialized indices, increased relevance and precision of retrieved information, scalability for large datasets and high query volumes, and optimized querying across diverse data sources. The techniques address challenges in managing extensive, distributed datasets and are compatible with existing RAG frameworks, providing a solution for complex information retrieval tasks.
Owner:ORACLE INT CORP

User participation degree prediction method based on distillation multi-modal retrieval enhancement

The invention discloses a user participation degree prediction method based on distillation multi-modal retrieval enhancement, and belongs to the technical field of social media data mining and user behavior analysis. According to the method, the correlation of the UGC is evaluated by introducing the self-enhanced distillation module, and the top-K related UGC is selected in combination with the selected retriever, so that the interference of irrelevant information on the related UGC is effectively avoided, and the noise caused by irrelevant artifacts can be filtered out when the related UGC is retrieved, thereby keeping the original feature representation of the related UGC. The heterogeneous graph construction module enhances the interactive representation capability between UGCs through multi-relation modeling, and can more accurately optimize a prediction result in user participation prediction. According to the method, information retrieval between the related UGC and the unrelated UGC can be better balanced, excessive diffusion of unrelated information is avoided, and the performance of the model in user participation degree prediction is improved. Particularly, when complex multi-modal data containing a large number of irrelevant UGCs is processed, the mechanism can effectively enhance the distinction degree of the relevant UGCs, and the prediction accuracy is improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Computing systems and methods for generating a response to a query based on a corpus of documents

Systems and method for generating a response to a query. The method includes using a first large language model (LLM) to generate synthetic information related to a query; generating an amended query based on the synthetic information related to the query; using an information retrieval system to retrieve, from a plurality of chunks, a set of chunks that are relevant to the amended query, wherein each chunk of the plurality of chunks is all or a portion of a document in a corpus of documents; using a second LLM to rank the set of chunks based on a relevance to the query; selecting a subset of chunks from the set of chunks based on the ranking; and using a third LLM to generate a response to the query based on the subset of chunks.
Owner:THE TORONTO DOMINION BANK

Identifying relevance of documents for automated retrieval models using large language models

There are provided systems and methods for identifying relevance of documents for automated retrieval models using large language models. An online transaction processor or other service provider may provide computing services and platforms to entities, which may include chatbots, information retrieval systems, question-and-answer systems, and the like. To provide better retrieval model training and refinement, the service provider may generate training data from user interaction logs, which may include user feedback that may be used to determine if documents are relevant to queries, and therefore should be retrieved for answering those queries by automated retrieval models. An LLM may be used as a judge to determine whether chatbot responses reference certain document. If not references, the query may be analyzed to determine whether certain retrieved documents are relevant. Data pairs may be generated for the training data from these processes and used for model refinement.
Owner:PAYPAL INC

Self-adaptive multi-feature fusion hybrid retrieval sorting method and system

The invention discloses a self-adaptive multi-feature fusion hybrid retrieval sorting method and system, and belongs to the technical field of information retrieval. The method comprises the steps that after user query is received, a mixed retrieval process and an intention recognition process are executed in parallel; performing multi-dimensional feature extraction on the candidate documents obtained by the mixed retrieval, wherein the multi-dimensional feature extraction comprises semantic correlation features, keyword matching features, document authority features and timeliness features; and according to the identified query type, adaptively selecting a fusion weight, and carrying out weighted fusion on the multi-dimensional feature vector to calculate a final score and sort the final score. According to the method, the problem that weight distribution is rigid in traditional mixed retrieval is solved through an intention self-adaptive dynamic weight mechanism, meanwhile, by introducing multi-dimensional service features, the ranking result not only ensures the correlation, but also meets the quality requirement under a service scene, and the accuracy and practicability of a retrieval system are remarkably improved.
Owner:叶绍琛

RAG mixed retrieval method and device

The invention provides an RAG mixed retrieval method and device, and relates to the technical field of information retrieval, and the method comprises the following steps: obtaining a query text and a term library, and calculating to obtain a vector retrieval weight and a keyword retrieval weight according to the query text and the term library; obtaining a document library, and performing vector retrieval on the document library according to the query text and the vector retrieval weight to obtain a first candidate document set; performing keyword retrieval on the document library according to the query text and the keyword retrieval weight to obtain a second candidate document set; merging the first candidate document set and the second candidate document set to obtain a fused document set; and calculating, screening and sorting comprehensive feature scores of the documents in the fused document set to obtain a final document sequence, and inputting the final document sequence into the large language model to obtain an answer text. According to the method, the limitation of a single scoring dimension in professional field retrieval is overcome, and the correlation between the answer text and the query text is improved through multi-dimensional feature joint evaluation.
Owner:NANJING DOLPHIN INTELLIGENT TECH CO LTD

Multipath information retrieval method and system based on language model and uncertainty evaluation

The invention provides a multi-path information retrieval method and system based on a language model and uncertainty evaluation. The multi-path information retrieval method comprises the following steps: receiving a user query; performing semantic analysis on user query by using a pre-trained language model so as to output a quantitative uncertainty index in an unsupervised manner; dynamically determining a quantity ratio for controlling query generation and a balance parameter for controlling result sorting based on the index; based on the user query and the strategy parameters, generating and executing two groups of differential queries at least comprising exploratory queries and utilizable queries to obtain multiple paths of retrieval results; and combining multiple paths of results, and carrying out dynamic balance reordering between the correlation and diversity of the results according to balance parameters by adopting an ordering mechanism. According to the method, the retrieval strategy can be adaptively adjusted according to the uncertainty of query, and the coverage degree and diversity of results are remarkably improved.
Owner:SHANGHAI BAOSIGHT SOFTWARE CO LTD

Information retrieval method, device and equipment based on large model questions and answers and medium

The invention provides an information retrieval method, device and equipment based on large model questions and answers and a medium, and the method comprises the following steps: firstly, receiving a natural language question input by a user through a terminal, and forwarding the natural language question to a semantic analysis engine after authentication of an API (Application Program Interface) gateway; then oral vocabularies are removed through a normalization module, abbreviation is completed, a retrieval category is determined in combination with an intention recognition module, key information is extracted through a slot extraction module, and meanwhile permission-semantic coupling retrieval is completed in combination with a company organization structure; then adopting a multi-path recall mode of vector semantic retrieval and keyword supplementary retrieval to obtain a candidate result, and performing correlation rearrangement and permission filtering to obtain a target retrieval result; and finally, performing fragment positioning on different types of contents, and displaying a result with specific position information at a front end in a preset form. According to the method, the retrieval accuracy and efficiency can be improved, cross-modal unified retrieval is realized, authority security is guaranteed, and efficient and accurate information retrieval requirements in large enterprises are met.
Owner:CHINA LIFE INSURANCE CO LTD SHANGHAI DATA CENT

Policy tracing dynamic retrieval method based on multi-level intention recognition and contrastive learning

The present invention discloses a policy traceability dynamic retrieval method based on multi-level intent recognition and comparative learning, which involves the fields of natural language processing, information retrieval and generation technology, and introduces multi-level intent recognition, dynamic field classification, embedded multi-round retrieval, comparative learning mechanism and generation optimization strategy. The present invention can accurately capture query intent, dynamically select the most relevant knowledge base, and generate high-precision traceability results through multi-level verification and optimization, significantly improving the accuracy, relevance and added value of retrieval, and meeting the high-precision scenario requirements such as complex policy tracing and compliance inspection.
Owner:BEIJING BIG DATA CENT

Hybrid retrieval method, system and equipment based on multi-algorithm fusion and storage medium

The invention provides a mixed retrieval method, system and device based on multi-algorithm fusion and a storage medium, and the method comprises the steps: obtaining various types of document data, preprocessing the document data, and obtaining plain text data; segmenting the plain text data into a plurality of pieces of block node data, and generating a plurality of derivative problems for each piece of block node data by using a large language model; carrying out vectorization coding on the block node data and the derivation problem, and storing the block node data and the derivation problem in a vector database; receiving a user query question, and generating a plurality of related preset questions for the query question by using the large language model; converting the query question and the preset question into a vector form, executing keyword retrieval and vector retrieval, and obtaining corresponding candidate results in a vector database; and carrying out weighted fusion on the candidate results through a weighted reciprocal rearrangement algorithm to obtain most relevant retrieval data. According to the method, the efficiency of information retrieval and the correlation of results can be improved, and the retrieval requirement of a user in a complex semantic scene is met.
Owner:XIAMEN INTRETECH

Dynamic retrieval strategy adaptive adjustment method and system based on reinforcement learning

The invention relates to the technical field of artificial intelligence and information retrieval, in particular to a dynamic retrieval strategy self-adaptive adjustment method and system based on reinforcement learning. According to the method, a retrieval process is modeled into a Lukov decision process, and a strategy model is trained to realize autonomous optimization and dynamic adjustment of a retrieval strategy. The method comprises the following steps: defining a state space comprising a query type, context correlation, a historical retrieval effect and knowledge base characteristics; defining an action space including vector retrieval, keyword retrieval and mixed retrieval; constructing a dual-target reward function fusing the retrieval income and the retrieval cost; training a strategy model by utilizing a historical retrieval track; calling the strategy model in real time to dynamically adjust the retrieval mode when the retrieval enhancement generation system runs; and the model is iteratively optimized through continuous feedback. According to the method, the problems of low recall rate and low calculation efficiency of a traditional fixed retrieval strategy under complex query are solved, dynamic balance between retrieval precision and resource consumption is realized, and the method is suitable for intelligent retrieval scenes of multi-field and multi-modal knowledge bases.
Owner:SHANGHAI CAIJIANG INTELLIGENT TECH CO LTD

Bidirectional information retrieval enhancement generation method for large language model

The invention discloses a bidirectional information retrieval enhancement generation method for a large language model, and belongs to the technical field of artificial intelligence. In order to overcome the defects that noise is introduced and key evidences are omitted due to the fact that traditional RAG only executes'query-document 'one-way retrieval, a two-way semantic perception retrieval enhancement generation model and a two-stage training framework are constructed, wherein in the first stage, the positive / negative example distance is increased in an embedded space in a contrast learning self-supervision mode; in the second stage, fine-grained correlation discrimination is carried out on query-document bidirectional sentences through supervised dichotomy, and probabilistic correlation scores are output; in the reasoning stage, the bidirectional probabilities are fused according to Bayesian to obtain final relevancy, document reordering is carried out, and plug and play can be achieved without fine adjustment of LLM in the whole process. According to the method, the accuracy and consistency of single-hop and multi-hop questions and answers and fact checking tasks are remarkably improved, and the method has the advantages of light weight and low deployment cost.
Owner:中华人民共和国大连海关

Retrieval enhancement generation method and system based on long-term memory and multi-source knowledge iterative fusion

The invention discloses a retrieval enhancement generation method and system based on long-term memory and multi-source knowledge iterative fusion. The method comprises the following steps: firstly, after user inquiry is received, triggering long-term memory database retrieval, external knowledge base retrieval and large language model internal knowledge generation in parallel to construct a multi-source candidate knowledge base, and improving result correlation by adopting a self-adaptive retrieval method based on a dynamic top-k value; secondly, inputting user query and multi-source knowledge into the large language model for iterative fusion, and generating a final answer through conflict detection, fusion processing and confidence evaluation; and finally updating the long-term memory database. The method can improve the accuracy and stability of the generated answers in a multi-round interaction and complex task scene, reduces the risk of answer one-sidedness, inference incompleteness or fact deviation, enhances the multi-source knowledge utilization ability and long-term learning ability of the system, and improves the user experience. Therefore, the comprehensive performance and reliability of the system in applications such as intelligent question answering, information retrieval assistance and text generation are obviously expanded.
Owner:SCHOOL OF SOFTWARE ZHEJIANG UNIV (NINGBO) MANAGEMENT CENT (NINGBO SOFTWARE EDUCATION CENT) +3

Efficient document searching method and device, equipment and storage medium

The invention relates to the field of information retrieval, and discloses an efficient document search method and device, equipment and a storage medium. The method comprises the steps that a document set is preprocessed, and an index structure is constructed; analyzing and optimizing a query word input by a user to generate a target retrieval request; retrieving an initial document set containing the target retrieval request based on the index structure; and carrying out correlation scoring on the initial document set by adopting a preset multi-dimensional scoring strategy, sorting according to a scoring result, and screening out a document set most related to the target retrieval request. According to the efficient document searching method provided by the invention, the document set most relevant to the user search query is screened out by analyzing and optimizing the user search query and combining the relevance score and the behavior data sorting result, so that the problems of low retrieval efficiency and poor result relevance of a traditional document searching method are effectively solved; and the document searching efficiency and the result quality are improved.
Owner:SHANGHAI DONGPU INFORMATION TECH CO LTD

A bidirectional information retrieval augmented generation method for large language models

The application discloses a bidirectional information retrieval enhancement generation method for a large language model, and belongs to the technical field of artificial intelligence. In order to overcome the defects of traditional RAG that only performs one-way retrieval from 'query to document' and introduces noise and omits key evidence, a bidirectional semantic perception retrieval enhancement generation model and a two-stage training framework are constructed. In the first stage, the positive / negative example distance is pulled apart in the embedding space in a contrast learning self-supervised manner. In the second stage, a supervised two-classification is used to finely distinguish the relevance of the query-document bidirectional sentence pair, and an output probability correlation is obtained. In the reasoning stage, the bidirectional probability is fused according to Bayes to obtain the final correlation degree, and the document is reordered. The whole process can realize plug and play without fine-tuning the LLM. The application significantly improves the accuracy and consistency of single-hop, multi-hop question answering and fact checking tasks, and has the advantages of light weight and low deployment cost.
Owner:中华人民共和国大连海关

Information retrieval system using a hierarchical corpus encoder

A dense encoder is adapted as a hierarchical corpus encoder in an information retrieval system to use negative samples from sibling nodes in a hierarchical tree of vector embeddings for documents in a corpus. Both the encoder and hierarchical tree are co-trained using a loss function that takes the document hierarchy into account. The hierarchical corpus encoder may be used in both supervised training cases where query-document relevance judgments are present and in zero-shot cases where a query dataset is absent. The hierarchical corpus encoder demonstrates significant performance improvements over a variety of dense encoder and generative retrieval baselines, under both supervised and unsupervised scenarios, thereby establishing the effectiveness of jointly learning a document hierarchy.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

A multi-agent-based retrieval method, apparatus, device, and medium

The application relates to the technical field of artificial intelligence, in particular to a retrieval method and device based on multiple agents, equipment and a medium. Applied to a financial scene, the application realizes the intellectualization and dynamicization of the retrieval process through a multiple-agent cooperative working mechanism. The retrieval core elements are accurately extracted and a dynamic cognitive graph is updated by analyzing the agents, deep semantic mining and correlation expansion are carried out based on the cognitive seeds of semantic agents, the dependence on keywords in traditional retrieval is effectively broken, and the breadth and depth of semantic understanding are improved. The query statement is optimized by combining the extended semantics and the dynamic cognitive graph of the context agent, so that the query is more in line with the real intention and context of the user. The retrieval agent efficiently retrieves the candidate path in the updated dynamic cognitive graph, the recommendation agent filters the optimal result through cognitive consistency scoring, and the accuracy, relevance and logical coherence of the retrieval result are ensured, and the accuracy of the complex information retrieval task is significantly improved.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Intention-driven document retrieval optimization method and system based on internal reasoning of large language model

The invention belongs to the technical field of information retrieval, and relates to an intention-driven document retrieval optimization method and system based on internal reasoning of a large language model. The method comprises the following steps: reconstructing instruction fine tuning task training data; training an intention-driven correlation retrieval model based on a large language model by utilizing the reconstructed instruction fine tuning task training data; and performing document retrieval by utilizing the trained intention-driven correlation retrieval model based on the large language model. According to the method, the intrinsic reasoning ability of the large language model is utilized to drive the user intention recognition of the retrieval model, so that the ability of the retrieval model to accurately recognize the user intention and recall the documents meeting the requirements in practical application is improved, more efficient and accurate document search can be realized, and the retrieval result better meets the user requirements.
Owner:INST OF SOFTWARE - CHINESE ACAD OF SCI

Mixing retrieval method and system based on knowledge graph, medium and equipment

The invention relates to the technical field of information retrieval, and discloses a mixed retrieval method and system based on a knowledge graph, a medium and equipment, and the method comprises the steps: obtaining a query statement, carrying out the feature extraction, and generating a query vector; based on the query vector, obtaining a plurality of preliminary retrieval results through full-text retrieval; based on the query vector, searching a plurality of knowledge graph paths related to the query statement in the knowledge graph through graph query; based on the query vector, the preliminary retrieval result and the knowledge graph path, generating a query enhancement vector through a graph attention network; for each preliminary retrieval result, calculating the similarity with the query enhancement vector to obtain a similarity score; for each knowledge graph path, calculating the correlation with the query enhancement vector to obtain a path weight; and combining the similarity score with the path weight, and reordering the preliminary retrieval result. And the retrieval precision is effectively improved.
Owner:SHANDONG LUNENG SOFTWARE TECH

Rag mixed search method and device

ActiveCN121501944BLinguistic modelData mining
The application provides a RAG mixed retrieval method and device, and relates to the technical field of information retrieval, and comprises the following steps: obtaining a query text and a term library, and calculating vector retrieval weights and keyword retrieval weights according to the query text and the term library; obtaining a document library, performing vector retrieval on the document library according to the query text and the vector retrieval weights, and obtaining a first candidate document set; performing keyword retrieval on the document library according to the query text and the keyword retrieval weights, and obtaining a second candidate document set; performing merging processing on the first candidate document set and the second candidate document set, and obtaining a fusion document set; performing calculation, screening and sorting of comprehensive feature scores of each document in the fusion document set, obtaining a final document sequence, inputting the final document sequence into a large language model, and obtaining an answer text. The application overcomes the limitation of a single scoring dimension in professional field retrieval, and improves the relevance of the answer text and the query text through multi-dimensional feature joint evaluation.
Owner:NANJING DOLPHIN INTELLIGENT TECH CO LTD

Information retrieval system using a hierarchical corpus encoder

A dense encoder is adapted as a hierarchical corpus encoder in an information retrieval system to use negative samples from sibling nodes in a hierarchical tree of vector embeddings for documents in a corpus. Both the encoder and hierarchical tree are co-trained using a loss function that takes the document hierarchy into account. The hierarchical corpus encoder may be used in both supervised training cases where query-document relevance judgments are present and in zero-shot cases where a query dataset is absent. The hierarchical corpus encoder demonstrates significant performance improvements over a variety of dense encoder and generative retrieval baselines, under both supervised and unsupervised scenarios, thereby establishing the effectiveness of jointly learning a document hierarchy.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Dynamic intelligent agent retrieval tree-based timeliness news retrieval method and system, electronic device and storage medium

PendingCN122346582AEngineeringData mining
The application discloses a dynamic intelligent agent retrieval tree-based time-sensitive news retrieval method and system, an electronic device and a storage medium, and belongs to the technical field of information retrieval and artificial intelligence. The technical scheme of the application comprises the following steps: constructing a retrieval tree with a user query as a root node, the tree being expanded in semantics through multi-agent cooperation; obtaining a current news corpus and constructing a representative evaluation agent; based on the evaluation agent, dynamically selecting an optimal sub-tree from the retrieval tree that meets a preset condition; and using the optimal sub-tree to search the full news corpus to obtain a target news document. The application decouples high-cost semantic expansion and high-concurrency online retrieval, reduces query delay and computing cost, and dynamically selects an optimal sub-tree to adapt to changes in news content in real time, thereby improving the timeliness and relevance of the retrieval result.

Method and system for autonomously positioning RAG error source

The invention discloses a method and a system for autonomously positioning an RAG error source, and belongs to the field of information retrieval. According to the method, a double-path error positioning mechanism is set, differential defects of a mainstream RAG architecture are covered, and a hop count adaptive detection and noise relation quantization threshold value is designed according to graph structure characteristics (such as multi-hop reasoning and relation pruning) of GraphRAG; aiming at semantic retrieval characteristics (such as multi-entity coverage and context correlation) of the VectorRAG, designing a layered quality evaluation mechanism, and realizing accurate mapping of error types and technical essence; and error positioning is driven by an algorithm process, an automatic decision chain is formed from answer matching, retrieval target comparison and logic verification generation, and specific error types and associated elements are output.
Owner:NORTHEASTERN UNIV CHINA

User engagement prediction method based on distillation multimodal retrieval enhancement

The application discloses a user participation degree prediction method based on distillation multimodal retrieval enhancement, and belongs to the technical field of social media data mining and user behavior analysis.The application introduces a self-enhanced distillation module to evaluate the relevance of UGC, and selects top-K relevant UGC in combination with a selected retriever, so that irrelevant information is effectively avoided from interfering with relevant UGC, and the noise caused by irrelevant artifacts is filtered out when the relevant UGC is retrieved, so that the original feature representation of the relevant UGC is maintained.A heterogeneous graph construction module enhances the interaction representation capability between UGC through multi-relation modeling, and can more accurately optimize the prediction result in user participation degree prediction.The application can better balance the information retrieval between relevant UGC and irrelevant UGC, avoid excessive spread of irrelevant information, and improve the performance of the model in user participation degree prediction.In particular, when complex multimodal data containing a large amount of irrelevant UGC is processed, the mechanism of the application can effectively enhance the distinguishability of relevant UGC and improve the prediction accuracy.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

A method and system for constructing a vertical domain large language model

The application provides a construction method and system of a vertical field large language model, relating to the technical field of natural language processing, including extracting topic word information of vertical field data through data preprocessing and classification model, and applying reinforcement learning to train the large model to improve its professional ability. Further, a vertical field knowledge base is generated based on the preprocessed data and topic word information, and heuristic encoding method is used to optimize information retrieval. Finally, the trained large model and knowledge base are deployed, and an application program interface is constructed to realize intelligent service. The application effectively solves the relevance and precision problems of general large models in professional field applications by comprehensively utilizing preprocessing, classification, reinforcement learning and knowledge encoding technologies, and promotes the practical application and development of artificial intelligence in vertical fields.
Owner:INST OF SCI & TECH INFORMATION OF CHINA ACAD OF RAILWAY SCI GRP CO LTD +2

Knowledge base information retrieval method and system of service platform

The invention discloses a knowledge base information retrieval method and system of a service platform, and belongs to the technical field of data processing. The method comprises the steps that query information input by a user is recognized; identifying extended vocabularies related to the user query information from a predefined vocabulary library, and screening to obtain a target document based on the importance of the extended vocabularies in the source document; obtaining a first correlation score of each target document and the query information by using an abstract scoring algorithm; obtaining a second correlation score of each target document and the query information by using a co-word algorithm; based on the first correlation score and the second correlation score, obtaining a comprehensive correlation score of each target document and the query information through a weighting mode; and according to the comprehensive correlation scores, performing descending sort on the target documents and then displaying the target documents to a user. According to the method, the semantic association of the vocabularies and the matching of the vocabularies and the document themes are considered, and the accuracy of database retrieval results is improved.
Owner:CHINA TELECOM DIGITAL INTELLIGENCE TECH CO LTD

Machine learning based retrieval enhancement generation method for user profile sub-second invocation

PendingCN122364391ALearning basedData segment
This invention relates to the field of information retrieval technology, specifically to a retrieval enhancement generation method based on machine learning for second-level retrieval of user data. This method solves the technical problems in existing technologies where structured medical data suffers severe loss of key clinical semantics during vectorization and fixed-length segmentation, and cannot support cross-indicator correlation analysis and low-latency feedback. The method includes: acquiring user queries against a structured user database; determining the semantic loss of at least one candidate data segment corresponding to the user query in the temporal continuity and medical relevance dimensions; using the semantic loss to characterize the degree of clinical semantic loss caused by segmentation of the candidate data segment; determining the completeness of the indicator association between the candidate data segment and the user query based on the semantic loss; and adjusting the retrieval results for the user query based on the indicator association completeness.
Owner:YUBO INT BIOTECHNOLOGY (BEIJING) GRP CO LTD +1

Contextual optimization method and apparatus for retrieving an augmentation generation system

PendingCN122285861ADynamical optimizationConditional entropy
This application provides a context optimization method and apparatus for a retrieval-enhanced generation system, relating to the field of information retrieval technology. The method includes: receiving a user query and retrieving a knowledge document set related to the user query through a knowledge base retrieval; dynamically optimizing the knowledge document set by iteratively executing symmetric and asymmetric merging operations using an optimization objective function based on information bottlenecks to generate a compressed context; wherein, the symmetric merging operation is used for document pairs with low relevance scores; and the asymmetric merging operation is used for documents that identify and fuse semantic redundancy based on conditional entropy. The context optimization method and apparatus for a retrieval-enhanced generation system provided by this application, through dynamic intelligent fusion and compression of retrieved documents, can provide larger language models with higher information density and more accurate context at a lower computational cost, thereby obtaining more reliable and complete answers.
Owner:PEKING UNIV