Multi-source information retrieval fusion method, device and equipment and readable storage medium
By integrating intention analysis and multi-source information retrieval of user query data, the problem of insufficient accuracy in answering artificial intelligence models in different scenarios is solved, and more accurate and interpretable information generation is achieved.
Patent Information
- Application Number
- CN202511046230.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-07-29
AI Technical Summary
In the prior art, artificial intelligence models cannot adapt to the needs of different scenarios when generating answers, resulting in low accuracy of answers, especially in the acquisition of real-time information and specific domain knowledge.
By analyzing the user query data, we determine the query intent type, intent representation and constraints, select the matching target information source from multiple information sources, use intent representation and constraints to search, obtain information items, and generate summary answers through semantic alignment, confidence evaluation and conflict detection.
It improves the accuracy and interpretability of the answers, can meet the information needs of different scenarios, and generates comprehensive answers based on multiple heterogeneous data sources.
Smart Images

Figure CN120541310A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing technology, and in particular to a multi-source information retrieval and fusion method, apparatus, device, and readable storage medium. Background Art
[0002] With the development of artificial intelligence (AI) technology, the advantages of large AI models (such as ChatGPT and Tongyi Qianwen) in multi-domain question-answering, reasoning, and dialogue generation have made them a new direction in information processing. However, users have higher requirements for information diversity, accuracy, and timeliness in various scenarios, such as daily communication, work decision-making, and scientific research. For example, when faced with open questions, users expect to obtain general knowledge and data reasoning; when it comes to real-time information or dynamic content, users' needs shift to the latest news and data updates; and in specific fields or scientific research, users require more specialized content.
[0003] However, related technologies that rely on large models to generate answers cannot adapt to the needs of different scenarios, and the generated answers have low accuracy. Summary of the Invention
[0004] Based on this, it is necessary to provide a multi-source information retrieval fusion method, device, computer equipment, computer-readable storage medium and computer program product that can improve the accuracy of answers to the above technical problems.
[0005] In a first aspect, the present application provides a multi-source information retrieval and fusion method, comprising:
[0006] Analyze the acquired user query data to determine the query intent type, intent representation, and constraint conditions of the query data;
[0007] Determine multiple target information sources that match the query intent type from multiple preset information sources;
[0008] For each of the target information sources, searching the target information source according to the intention representation and the constraint condition to obtain an information item corresponding to the intention representation;
[0009] A plurality of the information items are fused to generate a summary answer.
[0010] In one embodiment, determining multiple target information sources that match the query intent type from multiple preset information sources includes:
[0011] Determining quality feature data of each information source in the multiple types of information sources in different quality feature dimensions, and demand preference data of the query intention type for the quality feature dimensions;
[0012] Determining a matching matrix between the query intention type and the multiple types of information sources based on the demand preference data and the quality feature data;
[0013] The information sources corresponding to the preset values of the elements in the matching matrix are determined as target information sources matching the query intention type, and multiple target information sources are obtained.
[0014] In one embodiment, for each of the target information sources, searching the target information source according to the intent representation and the constraint condition to obtain information items corresponding to the intent representation includes:
[0015] For each of the target information sources, searching the target information source according to the intention representation and the constraint condition to determine the original text that matches the intention representation;
[0016] If the type of the original text is an unstructured text type, information extraction is performed on the original text to determine the information item corresponding to the intention representation from the original text.
[0017] In one embodiment, the fusing of the plurality of information items to generate a summary answer includes:
[0018] Performing semantic alignment on the plurality of information items to obtain first candidate information items after alignment of the semantic information corresponding to the respective items;
[0019] For each of the first candidate information items, determining the confidence of the first candidate information item in each preset evaluation dimension according to a preset confidence evaluation model;
[0020] determining a composite confidence level of the first candidate information item according to the confidence level;
[0021] performing conflict detection on each of the first candidate information items to obtain conflict data between the first candidate information items;
[0022] generating a corresponding conflict handling strategy according to the composite confidence and the conflict data;
[0023] A multi-source information fusion prompt is determined according to the first candidate information item, the confidence level, the composite confidence level, and the conflict handling strategy, and a plurality of the first candidate information items are fused to generate a summary answer.
[0024] In one embodiment, generating a corresponding conflict handling strategy based on the confidence level, the composite confidence level, and the conflict data includes:
[0025] If, among the first candidate information items, there are multiple second candidate information items whose composite confidence is greater than or equal to the first preset confidence, the generated conflict handling strategy includes sorting the multiple second candidate information items in descending order according to their composite confidences, and determining a priority of each of the second candidate information items based on the sorting result;
[0026] and / or, if, among the first candidate information items, there is a third candidate information item whose composite confidence is greater than or equal to the second preset confidence, and the conflict data among the plurality of the third candidate information items is that there is no conflict, then the generated conflict handling strategy includes performing semantic fusion on the plurality of the first candidate information items to generate a neutral summary; and the second preset confidence is greater than the first preset confidence;
[0027] And / or, if there are multiple groups of fourth candidate information items among the first candidate information items, the difference between any two of which is within a preset range, and the conflicting data between the fourth candidate information items is a conflict, the generated conflict handling strategy includes displaying the conclusion, information source, and the composite confidence of each fourth candidate information item.
[0028] In one embodiment, the method further comprises:
[0029] Determining reference information of the semantic segment in the summary answer, where the reference information includes at least one of an information source, an original text, and attribute information associated with the original text;
[0030] generating an interactive reference to the semantic fragment;
[0031] In response to a triggering operation on the interactive reference, the reference information referenced by the semantic segment is displayed.
[0032] In one embodiment, determining the reference information of the semantic segment in the summary answer includes:
[0033] Splitting the summary answer according to a preset semantic format to obtain semantic segments;
[0034] For each of the semantic segments, encoding the semantic segment and the original information segment in the original text corresponding to the semantic segment to obtain a first vector of the semantic segment and a second vector of each of the original information segments;
[0035] determining, based on the first vector and each of the second vectors, a similarity between the semantic segment and the original information segment, and determining a preset number of target original information segments from the original information segments based on the similarity;
[0036] The target original information segment, the information source of the target original information segment, and the attribute information associated with the target original information segment are determined as reference information of the semantic segment in the summary answer.
[0037] In a second aspect, the present application also provides a multi-source information retrieval and fusion device, comprising:
[0038] A data analysis module is used to analyze the acquired user query data to determine the query intent type, intent representation and constraint conditions of the query data;
[0039] An information source matching module is used to determine multiple target information sources that match the query intent type from a plurality of preset information sources;
[0040] A retrieval module, configured to search each target information source according to the intention representation and the constraint conditions to obtain information items corresponding to the intention representation;
[0041] The information fusion module is used to fuse the multiple information items to obtain fused information, summarize the fused information, and generate a summary answer.
[0042] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0043] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any of the above methods when executed by a processor.
[0044] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, which implements the steps of any of the above methods when executed by a processor.
[0045] The above-mentioned multi-source information retrieval fusion method, apparatus, computer equipment, computer-readable storage medium and computer program product identify the user's query data to determine the query intent type, and determine multiple target information sources that match the query intent type from multiple types of information sources including different query channels based on the query intent type, and then use the intent representation and constraint conditions to retrieve from each target information source respectively to obtain information items in each target information source, and obtain a generative summary answer by fusing and generatively summarizing information items from different query channels. This method effectively combines and coordinates multiple heterogeneous data sources to meet the needs of different scenarios and improves the accuracy and interpretability of the answers. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0047] Figure 1 A diagram showing an application environment of a multi-source information retrieval and fusion method in one embodiment;
[0048] Figure 2 1 is a flow chart of a multi-source information retrieval and fusion method according to an embodiment;
[0049] Figure 3 204 is a flow chart of step 204 in one embodiment;
[0050] Figure 4 A flowchart of a method for generating a summary answer in one embodiment is shown;
[0051] Figure 5 A schematic flow chart of a multi-source information retrieval and fusion method according to another embodiment;
[0052] Figure 6 is a structural block diagram of a dialogue system in one embodiment;
[0053] Figure 7 It is a structural block diagram of a multi-source information retrieval and fusion device in one embodiment;
[0054] Figure 8 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0056] With the development of artificial intelligence (AI) technology, the advantages of large AI models (such as ChatGPT and Tongyi Qianwen) in multi-domain question-answering, reasoning, and dialogue generation have made them a new direction in information processing. Related technologies use large models to generate responses. However, while large models excel in generating general knowledge and natural language, they lack the ability to provide accurate information in specific fields. In particular, they lag behind in acquiring the latest events and dynamic information due to limited training data and long update cycles, resulting in outdated content. Furthermore, large models are expensive to train and have limited scalability. In other words, responses generated by large models cannot meet the needs of real-time information and specific domain knowledge, and suffer from low accuracy.
[0057] However, in addition to large-scale model-generated answers, data sources such as web searches, public libraries, personal domain libraries, and external data interfaces are also unable to adapt to different scenarios and ensure the accuracy of answers during answer generation. For example, the quality of results returned by web searches is unstable, with insufficient coverage of professional fields and a high level of noise, requiring further information screening and verification. Public databases update data slowly and lack the ability to respond to cutting-edge or dynamic information. Personal domain databases have a limited knowledge scope and are unable to meet general or cross-domain needs. External data interfaces (such as financial data APIs) can provide structured and real-time data, suitable for queries of professional information or numerical data. However, these data interfaces are often limited by call frequency, high cost, and the returned data lacks context, making them unsuitable for direct use in complex answer generation.
[0058] Therefore, in order to solve the technical problem of how to improve the adaptability and accuracy of generated answers, a multi-source information retrieval fusion method is proposed.
[0059] The multi-source information retrieval fusion method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal analyzes the acquired user query data to determine the query intent type, intent representation and constraint conditions of the query data; determines multiple target information sources that match the query intent type from the preset multiple categories of information sources; for each target information source, searches the target information source according to the intent representation and constraint conditions to obtain information items corresponding to the intent representation; fuses multiple information items to obtain fused information, summarizes the fused information, and generates a summary answer.
[0060] The terminal 102 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, IoT devices, etc. The server 104 may be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services.
[0061] In an exemplary embodiment, Figure 2 As shown in the figure, a multi-source information retrieval fusion method is provided, which is applied to Figure 1 The terminal in FIG is taken as an example to illustrate the method, including the following steps 202 to 208. Among them:
[0062] Step 202: Analyze the acquired user query data to determine the query intent type, intent representation, and constraint conditions of the query data.
[0063] Among them, the query intention type can include but is not limited to factual, comparative, and causal types. The query intention type can be achieved by pre-analyzing possible user needs, classifying and defining possible user question types, and constructing a predefined domain intention set to facilitate subsequent matching of user questions.
[0064] For example, query intent types can include knowledge and facts (e.g., common sense, definition / explanation questions, causality, application, hypothesis), logical reasoning questions, task-based questions or process operations, conversational Q&A (e.g., casual chatting / socializing / psychological consultation / emotional expression), real-time information queries, predictive questions, tutorial / guidance questions, personalized information queries, data query and analysis, and source-specific queries. It is understood that user query data can include at least one query intent type, and each query intent type can have its own corresponding intent representation and constraints.
[0065] Intent representations, or key entities, can be understood as entities that represent user intent. Key entities help the system determine the focus of the user's query, enabling more accurate location of relevant information sources and data during information retrieval. Furthermore, key entities can refer to core nouns or noun phrases in the user's query, typically representing the primary objects or topics of interest to the user. These entities are fundamental to understanding user intent and performing information retrieval tasks. For example, key entities can include objects (e.g., product names, place names), concepts (e.g., artificial intelligence and weather changes), and events.
[0066] Constraints are specific requirements or restrictions specified by users in their queries, which have a direct impact on the filtering and sorting of search results. Examples include time constraints (e.g., last month's reports, data from the past three months), location constraints (e.g., progress in Region A), type constraints (e.g., academic papers), and topic constraints (e.g., technical challenges). Constraints help the system further refine the search scope, ensuring that the returned information better meets the user's specific needs and context. By combining key entities and constraints, the system can more precisely perform information retrieval tasks, providing users with more relevant and accurate information.
[0067] For example, natural language processing models are used to identify the intent of user query data, identify the query intent type, and extract key entities and constraints from the user query data. These key entities and constraints provide a semantic foundation for subsequent information retrieval. Intent identification can be achieved through fine-tuning of large generative models or some traditional algorithms, which will not be discussed in detail here.
[0068] Step 204: Determine multiple target information sources that match the query intent type from the preset multiple types of information sources.
[0069] Multiple information sources include heterogeneous data sources, including but not limited to large models, web search, public libraries, personal domain libraries, and external data interfaces. Web search excels at retrieving real-time internet content and is suitable for providing the latest news, hot events, and dynamic data. Public libraries such as Wikipedia and open research databases provide authoritative, structured academic knowledge, historical data, and scientific facts, making them suitable for providing background knowledge. Personal domain libraries can meet users' individual needs and provide highly relevant and customized information. External data interfaces (such as financial data APIs) can provide structured and real-time data and are suitable for queries involving specialized information or numerical data. The correspondence between query intent types and target information sources is predetermined or can be determined based on the matching degree between the query intent type and multiple information sources.
[0070] For example, in a scientific research scenario, a dialogue system might be used to conduct open-ended questions and answers about specialized domain knowledge and some daily questions, while also handling various task types within the scientific research scenario. The corresponding target information source is determined based on each intent type within the scientific research scenario.
[0071] For example, for queries with a knowledge and factual intent type (common sense, definition / explanation questions, causal questions, applications, and hypotheses), if the user asks general common sense questions, covering common knowledge and popular topics, such as "Which is the tallest mountain in the world?" or "Why is the sky blue?", then the target information sources can be large models, public libraries, and personal domain libraries. If the user wants the system to explain a term, concept, or principle, such as "What is quantum computing?" or "Explain how artificial neural networks work," then the target information sources can be large models, public libraries, and personal domain libraries. If the user asks a professional question about a specific field, typically involving in-depth knowledge or terminology, such as "What is the difference between machine learning and deep learning?" or "How do antibiotics inhibit bacterial growth?", then the target information sources can be large models, public libraries, and personal domain libraries.
[0072] If the user wants to know the cause or consequence of an event or phenomenon, for example, "Why do plants need sunlight to grow?" or "What are the consequences of not exercising?", then the target information sources can be large models, public libraries, and personal domain libraries. If the user proposes hypothetical scenarios and explores the possible outcomes of "what if" situations, for example, "How long can a fish survive out of water?", then the target information sources can be large models, public libraries, and personal domain libraries.
[0073] For logical reasoning questions, if the query intent is defined as a user asking a complex question requiring reasoning or multi-step calculation, for example, "How long will it take a turtle to crawl 20 kilometers at a speed of 5 kilometers per hour?" or "If A is larger than B, B is smaller than C, and C is larger than D, which is the smallest?", the confirmed target information source can be a large model. For task-based questions or process operations, if the query intent is defined as a user requesting the completion of a task or directive operation involving a specific operational process, such as translation, writing, coding, drawing, hotel reservations, reminders, settings, etc., for example, "Please set a reminder for a meeting tomorrow at 3:00 PM" or "Open the lighting settings page," the confirmed target information source can be an external application programming interface.
[0074] For conversational Q&A (chatting / socializing / psychological consultation / emotional expression), if the query intent is defined as informal small talk or emotional expression by the user, typically without explicit knowledge requirements. For example, "Do you like watching movies?", "I feel a little tired today," or "How are you?", the confirmed target information source large model is particularly good at handling small talk and social conversations, generating natural dialogue. If the query intent is defined as the user expressing emotions or seeking emotional support, for example, "I'm feeling down today. Can you comfort me?" or "I'm feeling uncertain about the future. How should I adjust my mindset?", the confirmed target information source, or large model, can generate appropriate emotional support or comfort based on the user's emotions.
[0075] For real-time information queries, where the user wants to obtain current events or instant information, for example, "What's the weather like today?" or "Can you summarize last week's headlines?", the target information source is the big model: a big model that summarizes multiple pieces of information and generates a concise summary answer. Web searches can retrieve real-time information such as news, weather, and exchange rates. External application programming interfaces (APIs) such as weather APIs and financial data APIs can retrieve real-time data.
[0076] For predictive queries, where the user wants the system to predict possible future events or trends, for example, "What are the development trends of artificial intelligence over the next ten years?" the target information source is the big model: a reasonable prediction based on existing trends and historical data. Web searches can be used to retrieve forecast reports or trend analysis in related fields.
[0077] For queries with a tutorial / guidance intent, if the user wants to obtain step-by-step instructions or learning guidance, such as "How do I install Python?", "What are the steps for making pasta?", "How do I draw?", or "How do I learn to code?", the target information source is the macro model: it can provide detailed step-by-step instructions or learning guidance. Web searches can be used to obtain ready-made tutorials or step-by-step guidance, especially for learning tools and skills.
[0078] For queries with personalized information intent, if the user wishes to obtain an answer based on personal data or customized information, this typically requires access to a user-specific database. For example, for queries like "What was my to-do list last week?" or "What time is my doctor's appointment?", the target information source is the large model. This model can summarize multiple pieces of information to generate a concise, summarized answer. External application programming interfaces (APIs): For example, database APIs can be used to query a user's personal data, such as their schedule or task list.
[0079] For query intent types such as data query and analysis, if the user wants the system to query and analyze certain data, for example, "Analyze the sales data for the past five years," or "Show me the profit distribution for the first quarter of this year," then the target information source is web search and retrieval of relevant statistical data. External application programming interfaces (APIs), such as database APIs or data analysis interfaces, provide data query and analysis results.
[0080] For query intent types with a specified source, if the user explicitly indicates where they wish to obtain information, for example, "Please search the web for information on the following: What are the reviews of the iPhone 15?", then the target information source is web search, which retrieves statistical data in the relevant field. For example, "Please search my personal document library for the answer to the question: What are the causes of earthquakes?", then the target information source is the personal domain library: if the user has uploaded relevant domain knowledge, this can provide content. For example, "Please search wikis or authoritative open-source literature for the answer to the question: What are the causes of earthquakes?", then the target information source is the public library: This can serve as a basis for information aggregation and provide authoritative background information. For example, "Please use the large model to directly answer the question: What are the causes of earthquakes?", then the target information source is the large model: it directly answers the user's request.
[0081] Step 206 : for each target information source, search the target information source according to the intention representation and the constraint conditions to obtain information items corresponding to the intention representation.
[0082] Retrieval methods can include, but are not limited to, RAG (Retrieval-Augmented Generation) technology, web searches, and API calls to search multiple target information sources. Examples include full-text search based on keyword inverse indexing, semantic search based on sentence vectors, and field queries based on structured data. Search engines can include open source or commercial products such as Elasticsearch, Solr, and Milvus, and can also be deployed locally or in the cloud with custom recall logic.
[0083] Furthermore, when calling external data sources or knowledge services, standard interface protocols (such as RESTful API, GraphQL, plug-in protocols, etc.) can be used to achieve docking and data interaction with third-party platforms, including but not limited to public databases and file services. The encapsulation and parsing of interface calls can be implemented using existing service scheduling frameworks or middleware. Those skilled in the art can implement this based on conventional network programming methods, which will not be detailed here.
[0084] Optionally, based on the intent representation and constraints, the target information source is retrieved to obtain the corresponding original text. Information extraction is then performed on the original text to obtain information items corresponding to the intent representation. Retrieving the original text can be accomplished using existing methods and will not be elaborated on here. Information extraction can involve using an information extraction model (such as BERT-NER) to convert the original text into structured knowledge units, from which information items corresponding to the intent representation are extracted.
[0085] Step 208: fuse the multiple information items to obtain fused information and generate a summary answer.
[0086] Fusion processing can be achieved through generative models, generating a comprehensive answer (i.e., a summary answer) based on information items retrieved from multiple sources, rather than simply concatenating or listing different pieces of information. Furthermore, before performing multi-source information fusion, it is necessary to ensure that the information from different sources is semantically consistent, that is, expressing the same facts or opinions.
[0087] Optionally, when the user input contains multiple questions, the system typically needs to split these sub-questions and then independently search and process each sub-question. The final result is like a comprehensive integration of the content retrieved from multiple data sources, rather than a simple spliced output. This splitting method can be achieved using existing methods and will not be elaborated here.
[0088] Exemplarily, multiple information items are fused to obtain fused information, and generating a summary answer can be to semantically align the multiple information items obtained, identify the same object with different expressions, facilitate subsequent confidence comparison and conflict detection, and then perform confidence assessment and conflict detection. Based on the structure of conflict detection, the conflict resolution strategy, that is, the conflict handling strategy, is determined. The results of semantic alignment, confidence assessment, and conflict detection, as well as the conflict resolution strategy, are used as prompt rules when generating a generative model to constrain the generation of the final answer, thereby realizing information fusion and generative answers.
[0089] The above-mentioned multi-source information retrieval fusion method identifies the user's query data, determines the query intent type, and determines multiple target information sources that match the query intent type from multiple types of information sources including different query channels based on the query intent type. Then, the intention representation and constraint conditions are used to retrieve from each target information source respectively to obtain information items in each target information source. By fusing and generatively summarizing information items from different query channels, a generative summary answer is obtained. This method effectively combines and coordinates multiple heterogeneous data sources to meet the needs of different scenarios and improves the accuracy and interpretability of the answers.
[0090] The following provides an implementation method for determining multiple target information sources that match the query intent type. In an exemplary embodiment, Figure 3 As shown, step 204 includes steps 302 to 306. Among them:
[0091] Step 302 : Determine the quality feature data of each information source in different quality feature dimensions, and the demand preference data of the query intention type for the quality feature dimensions.
[0092] Among them, different quality feature dimensions are determined based on actual needs, and the feature dimensions may include attribute feature dimensions such as timeliness, authority, personalization and content depth.
[0093] Demand preference data can be questions from users about various general common sense, involving common knowledge and popular topics, the desire for a systematic explanation of a term, concept or principle, answers to professional questions in a specific field, the desire to know the cause or consequence of an event or phenomenon, the proposal of hypothetical scenarios, the exploration of possible results of "if" situations, the desire to reason or solve complex problems with multiple steps of calculation, the ability to complete a task or directive operation, involving specific operating procedures, etc.
[0094] Step 304: Determine a matching matrix between the query intent type and multiple types of information sources based on the demand preference data and the quality feature data.
[0095] Exemplarily, the demand preference data of the quality feature dimension of the query intent type is matched with multiple types of information sources to obtain a matching matrix between the query intent type and the multiple types of information sources. For example, based on the specific scientific research scenario, large models, web searches, public libraries, personal document libraries, and external data APIs. These data sources themselves have obvious differences in attributes such as timeliness, authority, personalization, and content depth. Based on the 10 sorted intents, the matching degree of the five data sources will be evaluated separately, focusing on the different requirements of different intents for information timeliness, authority, personalization, and content depth, so as to determine the matching degree of the five data sources for these different intents (0 is mismatch, 1 is match), as shown in the information source matching matrix in Table 1 below. A matching degree of 0 will not participate in this round of information retrieval, and the target information source that matches each query intent type will be obtained.
[0096] Table 1 Information source matching matrix
[0097]
[0098] Based on the above matching method, we can see that when the system processes different user questions, it can adaptively match different information data sources to different questions. For other systems or those with different requirements for information sources, the values in the table can be updated to flexibly adapt to various user needs, obtain the corresponding target information source, and improve the accuracy and efficiency of data source call.
[0099] Step 306 : The information source corresponding to the preset value of the element in the matching matrix is determined as the target information source matching the query intention type, and multiple target information sources are obtained.
[0100] Among them, the preset value can be 1.
[0101] In this embodiment, the priority of information sources is dynamically adjusted according to the matching degree between user intention and information source, which can adapt to the differentiated needs for information timeliness, authority, personalization, etc. in different scenarios.
[0102] In order to further obtain structured data that is easier for machines to process from the original text, in an exemplary embodiment, for each target information source, the target information source is searched based on the intent representation and constraints to obtain information items corresponding to the intent representation, including:
[0103] For each target information source, the target information source is retrieved according to the intent representation and constraints to determine the original text that matches the intent representation; if the type of the original text is unstructured text, information extraction is performed on the original text to determine the information item corresponding to the intent representation from the original text.
[0104] For example, if the original text is not structured, it is necessary to extract it using an information extraction model, such as BERT-NER, to convert the original text of the retrieved information item into structured knowledge units and determine the information item corresponding to the intent representation. If the original text is structured, no conversion is required. This method can extract useful information from unstructured text and convert it into a structured form that is easier for machines to process.
[0105] In an exemplary embodiment, a method for generating a summary answer is provided, such as Figure 4 As shown, the following steps are included:
[0106] Step 402 : semantically align multiple information items to obtain first candidate information items after their corresponding semantic information is aligned.
[0107] Semantic alignment involves vector clustering of "synonymous facts" expressed differently. Even if the information is expressed differently, if they express the same fact, they will be clustered together in the vector space. For example, "The Moho thickness is approximately 66 kilometers" and "The Moho depth is 66 kilometers" can be considered a synonymous sentence pair. This can be achieved using semantic vector models such as SBERT (Sentence-BERT) and SimCSE (Simple Contrastive Sentence Embeddings). The specific implementation can be achieved using existing methods and is not detailed here.
[0108] It should be noted that only after semantic alignment can the same object with different expressions be identified, which facilitates subsequent confidence comparison and conflict detection.
[0109] Step 404 : For each first candidate information item, determine the confidence of the first candidate information item in each preset evaluation dimension according to a preset confidence evaluation model.
[0110] Step 406: Determine the composite confidence of the first candidate information item according to the confidence.
[0111] Among them, the evaluation dimensions of the pre-set credibility assessment model include at least the source authority, data timeliness, methodological argumentation and intention matching.
[0112] Source authority source_score is used to measure the credibility of information sources, giving priority to academic / official channels, etc. The corresponding scoring standard range is 0-1, and the weight is Specifically, for example, national authoritative institutions (such as CENC and NASA) have a corresponding score of 1.00; SCI journal articles (xxx1 / 2) have a corresponding score of 0.90; high-quality preprints (such as arXiv and bioRxiv) have a corresponding score of 0.80; university / research institute project reports or databases have a corresponding score of 0.70; media reports, Wikipedia and other second-hand compilations have a corresponding score of 0.50; community Q&A / non-professional forums (such as Zhihu and Baidu) have a corresponding score of 0.30.
[0113] Data timeliness time_score is used to ensure that the latest data / research is used, which is particularly important in rapidly evolving fields (such as earthquake monitoring, climate forecasting, etc.). The corresponding scoring standard range is 0-1, and the weight is Specifically, if the time difference between the release date and the current time is ≤ 6 months, the corresponding recommendation score is 1.00; if the time difference between the release date and the current time is less than 1 year, the corresponding recommendation score is 0.85; if the time difference between the release date and the current time is less than 2-3 years, the corresponding recommendation score is 0.70; if the time difference between the release date and the current time is more than 5 years, the corresponding recommendation score is 0.40; if the time difference between the release date and the current time is more than 10 years, the corresponding recommendation score is 0.30.
[0114] Method proof method_score is used to measure whether the data acquisition / experiment / modeling behind the information is detailed and reasonable and reliable. The corresponding recommendation score range is 0-1, and the weight is Specifically, if the method description quality clearly states the data source, method, parameters, etc., the corresponding recommended score is 1; if the method description quality provides a complete method or experimental process, the corresponding recommended score is 0.90; if the method description quality briefly describes the method but lacks details, the corresponding recommended score is 0.70; if the method description quality only gives the results or conclusions without a method description, the corresponding recommended score is 0.40; if the method description quality comes from speculative content or verbal expression, the corresponding recommended score is 0.20.
[0115] Intent_score is used to determine whether the information truly answers the core of the user's question. The corresponding matching score is 0-1, and the weight is For example, using sentence vectors (such as SBERT) to calculate the semantic similarity between the text and the user question, specifically: if the similarity interval is ≥0.9, the corresponding matching score is 1.00; if the similarity interval is 0.8-0.9, the corresponding matching score is 0.90; if the similarity interval is 0.7-0.8, the corresponding matching score is 0.75; if the similarity interval is 0.6-0.7, the corresponding matching score is 0.60; if the similarity interval is <0.6, the corresponding matching score is 0.40.
[0116] Based on this, the composite confidence score confidence_score is calculated using linear weighting:
[0117] ,
[0118] Among them, the default weights (for example, ) can be modified according to actual needs and scenarios.
[0119] Step 408: Perform conflict detection on each first candidate information item to obtain conflict data between the first candidate information items.
[0120] Conflict detection can detect the numerical value, qualitative nature, and cause of information items. Inconsistent numerical values, contradictory qualitative assertions, and mutually exclusive causal inferences are considered conflicts.
[0121] For example, common conflict types include: numerical conflict: differences between indicators exceeding ±5% (e.g., "66 km" vs. "70 km"), qualitative conflict: mutually exclusive conclusions (e.g., "plate stability" vs. "risk of rupture"), and causal conflict: inconsistencies in the chain of reasoning (e.g., A→B vs. A→B). Specific conflict detection techniques are as follows: For numerical conflicts, regular expressions can be used to extract values and units from text, perform normalization conversions, and calculate the relative difference percentage. A conflict is identified when the difference between the numerical values of the same indicator exceeds a preset threshold (default ±5%, adjustable based on the domain), while also considering data confidence to avoid misjudgments. For example, if the relative difference between "66 km" and "70 km" is 6.06%, exceeding the threshold will result in a numerical conflict.
[0122] For qualitative conflicts, text vector similarity is calculated using a pre-trained language model (such as BERT), and antonym lexicons are used to identify mutually exclusive expressions (such as the semantic opposition between "stability" and "rupture").
[0123] For causal conflicts, dependency parsing is used to extract subject-verb-object structures, and logical rules are used to verify inference contradictions (for example, detecting the mutually exclusive relationship between "A causes B" and "A does not cause B"). The conflict detection results must output conflict data including the conflict type, conflicting text fragments, and the confidence level of each preset assessment dimension, which can be subsequently integrated into prompt input.
[0124] Step 410: Generate a corresponding conflict handling strategy based on the composite confidence and the conflict data.
[0125] Among them, the conflict resolution strategy can be to selectively retain information items with a composite confidence level greater than or equal to the first preset confidence level and discard information items with a confidence level less than the first preset confidence level based on the composite confidence level and the first preset confidence level. When multiple third candidate information items coexist with composite confidence levels greater than or equal to the second preset confidence level, a neutral summary answer will be jointly generated. If there are multiple groups of fourth candidate information items in the first candidate information items whose composite confidence level difference between any two is within a preset range, but the conclusions are contradictory (such as inconsistent numerical values, contradictory qualitative assertions, mutually exclusive causal inferences, etc.), the multiple conclusions and sources will be displayed for user judgment. The conflict resolution strategy will also serve as a prompt rule during model generation and as an input for multi-source information fusion prompts to constrain the generation of the final answer.
[0126] Furthermore, a corresponding conflict handling strategy is generated based on the composite confidence and the conflict data, including at least one or more of the following three situations:
[0127] Case 1: If there are multiple second candidate information items among the first candidate information items whose composite confidence is greater than or equal to the first preset confidence, the generated conflict resolution strategy includes sorting the multiple second candidate information items from largest to smallest according to their composite confidence, and determining the priority of each second candidate information item based on the sorting result.
[0128] The first preset confidence level can be set according to actual needs, for example, 0.3. When the information confidence level is too low (<0.3), the information is considered to be insufficiently credible and can be directly ignored and not used.
[0129] Exemplarily, if among the first candidate information items, there are multiple second candidate information items whose composite confidence is greater than or equal to the first preset confidence, the generated conflict handling strategy includes sorting the composite confidences of the multiple second candidate information items from large to small in the information fusion and generative answer stage, determining the priority of each second candidate information item based on the sorting result, and using the expression with the highest priority, that is, the highest composite confidence, as the main basis for the generative answer.
[0130] Case 2: If, among the first candidate information items, there are multiple third candidate information items whose composite confidence is greater than or equal to the second preset confidence, and the conflict data between the multiple third candidate information items is that there is no conflict, then the generated conflict handling strategy includes semantically fusing the multiple third candidate information items to generate a neutral summary; the second preset confidence is greater than the first preset confidence.
[0131] The second preset confidence level can be set according to actual needs, for example, 0.75. It is understood that the second scenario is based on the first scenario, after discarding candidate information items with less than the first preset confidence level from the first candidate information items. If the second scenario exists, the processing method of the second scenario is combined.
[0132] Case 3: If there are multiple groups of fourth candidate information items among the first candidate information items, the difference between any two of which has a composite confidence level within a preset range, and the conflicting data between the fourth candidate information items is a conflict, the generated conflict handling strategy includes displaying the conclusion, information source, and composite confidence level of each fourth candidate information item.
[0133] The preset range may be less than 0.1. The conflict may be manifested as one or more of qualitative contradiction, causal mutual exclusion, and high numerical deviation.
[0134] It can be understood that, based on the situation one, the situation three is processed in combination with the processing method of the situation three after discarding the candidate information items with a reliability less than the first preset reliability from the first candidate information items.
[0135] Step 412 : Determine a multi-source information fusion prompt based on the first candidate information item, the confidence level, the composite confidence level, and the conflict handling strategy, fuse the multiple first candidate information items, and generate a summary answer.
[0136] The first candidate information item can be presented in the form of an information item list or JSON (JavaScript Object Notation). For example, the information item list is filled with structured information items, such as JSON or list format, each of which includes: information content, source (e.g., web / large model / public library / personal local library / API interface), confidence level, composite confidence level, and whether there is a qualitative conflict or numerical difference. Optionally, the multi-source information fusion prompt can be as follows:
[0137] You are a researcher in a specialized field, integrating search results from various sources to produce an authoritative, credible, and clearly structured summary of your research. The following are multiple source information items (including their confidence scores and source information). Please follow the following rules to resolve conflicts and write a summary response:
[0138]
Conflict handling strategy description
[0139] 1. Confidence selection strategy:
[0140] When the information confidence is too low (i.e., less than the first preset confidence level of 0.3), the information item is considered to be insufficiently credible and should be ignored and not used. During the generation process, the second candidate information item with a higher composite confidence level will be given priority, and its expression will be considered the primary basis.
[0141] 2. Multi-confidence information coexistence strategy:
[0142] If there are multiple non-conflicting information items with high confidence (i.e., greater than or equal to the second preset confidence level of 0.75), i.e., third candidate information items, please integrate their key points and summarize them in neutral language;
[0143] Pay attention to retaining the advantages of each point of view to make the output answer comprehensive and representative;
[0144] 3. Similar confidence levels but conflicting conclusions (qualitative contradictions, causal mutual exclusion, high numerical deviation):
[0145] If multiple confidence levels are similar (i.e., the difference between any two composite confidence levels is less than 0.1) and the conclusions are substantially conflicting, i.e., the fourth candidate information item, please list each conclusion, its source, and confidence level for further judgment by the user;
[0146] You can use expressions such as "there are differences in research views", "some studies believe that..., while other studies point out..." to show differences in views and demonstrate the ability to trace back.
[0147] Please complete the following tasks:
[0148] Based on the information provided below, generate a summary scientific research answer;
[0149] Follow the three information processing rules above;
[0150] The answer style should be professional, neutral, and verifiable;
[0151] Mention the sources of different conclusions in the summary (in terms of literature or data sources);
[0152]
Information item list
[0153] <Fill in structured information items here, such as JSON or list format, each item includes:
[0154] -Information content
[0155] - Source (web / large model / public library / personal local library / API interface)
[0156] -Confidence score
[0157] - Whether there are qualitative conflicts or numerical discrepancies
[0158] >
[0159] In the above embodiment, scoring information items based on multiple dimensions such as source authority, data timeliness, methodological justification, and semantic matching with user intent can more effectively handle information conflicts and uncertainties, and can dynamically adjust the information fusion method based on the user's query intent and the characteristics of the information source to improve the reliability and credibility of the answer.
[0160] To ensure that the generative answers are traceable and explainable, after completing the fusion of multi-source information and generating summary answers, a mapping relationship is established between the generated content and the semantics of the original information to build a traceable question-answer response framework.
[0161] In an example embodiment, by determining the reference information of the semantic fragment in the summary answer, the reference information shown includes at least any one of the information source, the original text, and the attribute information associated with the original text; generating an interactive reference of the semantic fragment; and displaying the reference information referenced by the semantic fragment in response to a trigger operation for the interactive reference.
[0162] Among them, semantic fragments can be obtained by splitting the summary answer into sentences or semantic clause levels. The semantic clause level can refer to the semantic analysis of the text from the perspective of clauses in natural language processing. The semantic structure within the clause and the logical relationship between clauses are taken into account. Attribute information includes the release time and the paragraph position in the corresponding original text. Interactive references can be expressed as but not limited to numbers or links. Optionally, when the user clicks on the interactive reference, the original text fragment, source name, release time and other information of the statement are displayed, supporting structured output of traceability information for use by the front-end interface or API system.
[0163] Optionally, determining the reference information of the semantic fragments in the summary answer includes: splitting the summary answer according to a preset semantic format to obtain semantic fragments; for each semantic fragment, encoding the semantic fragment and the original information fragment in the original text corresponding to the semantic fragment to obtain a first vector of the semantic fragment and a second vector of each original information fragment; determining the similarity between the semantic fragment and the original information fragment based on the first vector and each second vector, and determining a preset number of target original information fragments from the original information fragments based on the similarity; determining the target original information fragment, the information source of the target original information fragment, and the attribute information associated with the target original information fragment as the reference information of the semantic fragment in the summary answer.
[0164] For example, the summary answer is broken down into sentences or semantic clauses to generate the generated sentences in the summary answer. Each generated sentence and the original retrieved information fragment are uniformly encoded into a sentence vector. This encoding method can use the embedding output of BERT, SBERT, SimCSE, or a locally deployed large language model. The top N sentences with the highest correlation between each sentence in the summary answer and the original information source are calculated using cosine similarity, and a mapping relationship is established. When the front-end mouse hovers over the answer sentence, clicking each [number] expands to display a "source pop-up" showing relevant fragments, links, and the original paragraph location.
[0165] In the above embodiment, by supporting the traceability method of semantic fragment dimension, the traceability information such as the information source and original content fragment is recorded for the generated answer, and users are supported to click to view the original evidence behind the answer, thereby ensuring the explainability and transparency of the system.
[0166] In an exemplary embodiment, Figure 5 As shown in the figure, a multi-source information retrieval fusion method is provided, which is applied to Figure 1 The following steps are used as an example to illustrate the terminal in the figure:
[0167] Step 502: Analyze the acquired user query data to determine the query intent type, intent representation, and constraint conditions of the query data.
[0168] Step 504: Determine multiple target information sources that match the query intent type from the preset multiple types of information sources.
[0169] Step 506 : For each target information source, search the target information source according to the intention representation and the constraint conditions to obtain information items corresponding to the intention representation.
[0170] Step 508: fuse multiple information items to generate a summary answer.
[0171] Step 510: Determine the reference information of the semantic segment in the summary answer, where the reference information includes at least one of the information source, the original text, and attribute information associated with the original text.
[0172] Step 512: Generate an interactive reference of the semantic segment, and display reference information referenced by the semantic segment in response to a triggering operation on the interactive reference.
[0173] It should be noted that the implementation of this embodiment can be achieved through the above-defined methods, which will not be elaborated here.
[0174] The following provides a dialogue system for the multi-source information retrieval and fusion method. The dialogue system includes a user input module, a natural language understanding module, a dialogue management module, a multi-source retrieval module, an information fusion and traceability module, and a system output module. Figure 6 As shown in the figure, each module is responsible for different links from user input to response generation, ensuring that the system can ultimately realize an "AI-generated but explainable" scientific research question-answering system.
[0175] Users can engage in question-and-answer conversations with the dialogue system through the user dialogue interface provided by the system's PC. User input information, i.e., user query data, can include only the most recent piece of information entered by the user, several pieces of information entered by the user, or all information previously entered by the user, without limitation.
[0176] After receiving user input information, the user input information will be passed to the natural language understanding module (NLU processing module). This module is responsible for parsing the user input, using the natural language processing model to perform semantic analysis on the user input query, building an intent classifier to identify the query intent type (such as factual, comparative, causal, etc.), extracting the key entities, constraints and query context in the query, and ultimately determining the user's intent.
[0177] The output of the NLU module is passed to the dialogue management module, which, based on the identified intent, directs it to the multi-source information retrieval module and simultaneously updates the dialogue status. The multi-source information retrieval module evaluates the match between multiple predefined information sources and the user's intent and determines which sources to call. This multi-source information search is performed using a combination of RAG technology, network search, and API calls.
[0178] After completing the retrieval, it enters the information fusion and traceability module to perform consistency checks and conflict resolution on information items retrieved from multiple sources. Based on the retrieval results, it generates appropriate language responses through natural language generation (NLG) technology, and uses the generative language model to fuse information to generate a summary answer based on the results fed back to the user. It also records traceability information such as the information source and original content fragments in the answer.
[0179] The generated response content is sent back to the user through the user dialogue interaction interface or API provided by the PC side of the system. The user can continue multiple rounds of dialogue based on the feedback, and the system repeats the above steps.
[0180] In another exemplary embodiment, the user sends natural language input to the system through a chat interface or API interface, allowing the user to enter complex questions simultaneously in a single round of dialogue. For example, the user inputs: "Please query the isotope data of granite from the OnePetrology database, and analyze the origin of granite based on the data." The user needs to first retrieve the isotope data of granite through the database API interface, and then analyze the origin of granite based on the data. Therefore, the user intention is first decomposed into two sub-intentions, namely, the intention Figure 1 :Please query the isotope data of granite from OnePetrology database, Figure 2 : Combined with the causes of granite under data analysis, the corresponding intent types are data acquisition intent and professional analysis and reasoning intent. The target information source corresponding to data acquisition intent is the specified source-call interface, and the target information source corresponding to professional analysis and reasoning intent is the target information source corresponding to knowledge and facts (common sense, definition / explanation questions, causal class, application, hypothesis).
[0181] According to the matching of target information source, Figure 1 We will meet user needs by directly calling the OnePetrology database connected to the system. Figure 2 , we will obtain relevant information items by calling large models, web searches, public libraries and personal local paper libraries.
[0182] After the system obtains information items through the following different information channels, it then uses our confidence score to finally obtain the confidence scores of different information items, as shown in Table 2:
[0183] Table 2 Confidence scores of different information items
[0184]
[0185] The "×" in "0.25×" indicates that the information is less than the preset threshold and can be discarded. By integrating the above information items through the multi-source information fusion general prompt described above, we will generate the final answer.
[0186] In the above embodiment, relevant information is retrieved from multiple heterogeneous information sources through intent recognition; confidence scores are assigned to each information item based on the authority of the source, the timeliness of the data, the degree of methodological justification, and the degree of intent matching; semantic alignment and consistency checks are performed, conflicting data is identified, and a fusion strategy is adopted to process it; a large language model is called based on the summary prompt input to generate professional, authoritative, and comprehensive answers; and a traceability mechanism is used to associate the original information items with the source for the generated content to support traceability. In other words, by constructing "multi-source fusion + confidence assessment + explainable generation", the capabilities of the intelligent question-answering system in multi-source data fusion, consistency processing, and result interpretability can be significantly improved. It can handle various types of user needs such as open questions, real-time information queries, and predictive questions, and has stronger versatility and adaptability. At the same time, it provides users with more transparent and traceable answers, significantly improving the capabilities of the intelligent question-answering system in multi-source data fusion and result interpretability.
[0187] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0188] Based on the same inventive concept, the present application also provides a multi-source information retrieval and fusion device for implementing the multi-source information retrieval and fusion method mentioned above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more multi-source information retrieval and fusion device embodiments provided below can be found in the above-mentioned limitations of the multi-source information retrieval and fusion method, and will not be repeated here.
[0189] In an exemplary embodiment, Figure 7 As shown, a multi-source information retrieval and fusion device is provided, including: a data analysis module 702, an information source matching module 704, a retrieval module 706 and an information fusion module 708, wherein:
[0190] The data analysis module 702 is used to analyze the acquired user query data to determine the query intent type, intent representation and constraint conditions of the query data.
[0191] The information source matching module 704 is used to determine multiple target information sources that match the query intention type from multiple preset types of information sources.
[0192] The retrieval module 706 is configured to search each target information source according to the intent representation and the constraint conditions to obtain information items corresponding to the intent representation.
[0193] The information fusion module 708 is used to fuse multiple information items and generate a summary answer.
[0194] The above-mentioned multi-source information retrieval and fusion device identifies the user's query data to determine the query intention type, and determines multiple target information sources that match the query intention type from multiple types of information sources including different query channels based on the query intention type, and then uses the intention representation and constraint conditions to search from each target information source respectively to obtain information items in each target information source, and obtains a generative summary answer by fusing and generatively summarizing information items from different query channels. This method effectively combines and coordinates multiple heterogeneous data sources to meet the needs of different scenarios and improves the accuracy and interpretability of the answers.
[0195] In an exemplary embodiment, the information source matching module 704 is configured to determine quality feature data of each information source in different quality feature dimensions among multiple types of information sources, and demand preference data of the query intent type on the quality feature dimensions;
[0196] Determine the matching matrix between query intent type and multiple information sources based on demand preference data and quality feature data;
[0197] The information source corresponding to the preset value of the element in the matching matrix is determined as the target information source matching the query intention type, and multiple target information sources are obtained.
[0198] In an exemplary embodiment, the retrieval module 706 is configured to search each target information source according to the intent representation and the constraint conditions to determine the original text that matches the intent representation;
[0199] If the type of the original text is an unstructured text type, information extraction is performed on the original text to determine information items corresponding to the intent representation from the original text.
[0200] In an exemplary embodiment, the information fusion module 708 includes a semantic alignment module, a confidence determination module, a conflict detection module, and a summary generation module. The semantic alignment module is used to semantically align multiple information items to obtain first candidate information items after the corresponding semantic information is aligned.
[0201] The confidence determination module is used to determine the confidence of each first candidate information item in each preset evaluation dimension according to a preset confidence evaluation model; and determine the composite confidence of the first candidate information item according to the confidence.
[0202] The conflict detection module is used to perform conflict detection on each first candidate information item to obtain conflict data between the first candidate information items; and generate a corresponding conflict handling strategy according to the composite confidence and the conflict data.
[0203] The summary generation module is used to determine the multi-source information fusion prompt based on the first candidate information item, confidence, composite confidence and conflict handling strategy, fuse multiple first candidate information items, and generate a summary answer.
[0204] In an exemplary embodiment, the conflict detection module is configured to, if there are multiple second candidate information items whose composite confidences are greater than or equal to a first preset confidence among the first candidate information items, generate a conflict handling strategy comprising sorting the multiple second candidate information items from largest to smallest according to their composite confidences, and determining a priority of each second candidate information item based on the sorting result;
[0205] and / or, if, among the first candidate information items, there is a third candidate information item whose composite confidence is greater than or equal to the second preset confidence, and the conflict data among the plurality of third candidate information items is that there is no conflict, then the generated conflict handling strategy includes performing semantic fusion on the plurality of first candidate information items to generate a neutral summary; the second preset confidence is greater than the first preset confidence;
[0206] And / or, if there are multiple groups of fourth candidate information items in the first candidate information items, the difference between any two of which has a composite confidence level is within a preset range, and the conflicting data between the fourth candidate information items is a conflict, the generated conflict handling strategy includes displaying the conclusion, information source, and composite confidence level of each fourth candidate information item.
[0207] In an exemplary embodiment, the above-mentioned multi-source information retrieval and fusion device also includes an information tracing module, which is used to determine the reference information of the semantic fragment in the summary answer, where the reference information includes at least any one of the information source, the original text, and the attribute information associated with the original text; generate an interactive reference of the semantic fragment; and display the reference information referenced by the semantic fragment in response to a trigger operation for the interactive reference.
[0208] In an exemplary embodiment, the information tracing module is used to split the summary answer according to a preset semantic format to obtain semantic segments; for each semantic segment, the semantic segment and the original information segment in the original text corresponding to the semantic segment are encoded to obtain a first vector of the semantic segment and a second vector of each original information segment; based on the first vector and each second vector, the similarity between the semantic segment and the original information segment is determined, and a preset number of target original information segments are determined from the original information segments based on the similarity; the target original information segment, the information source of the target original information segment, and the attribute information associated with the target original information segment are determined as the reference information of the semantic segment in the summary answer.
[0209] Each module in the multi-source information retrieval and fusion device described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0210] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 8As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be implemented via Wi-Fi, a mobile cellular network, near-field communication (NFC), or other technologies. When executed by the processor, the computer program implements a multi-source information retrieval and fusion method. The display unit of the computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0211] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0212] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0213] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0214] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0215] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0216] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.
[0217] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0218] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A multi-source information retrieval and fusion method, characterized in that: The method comprises: Analyze the acquired user query data to determine the query intent type, intent representation, and constraint conditions of the query data; Determine multiple target information sources that match the query intent type from multiple preset information sources; For each of the target information sources, searching the target information source according to the intention representation and the constraint condition to obtain an information item corresponding to the intention representation; A plurality of the information items are fused to generate a summary answer.
2. The method according to claim 1, characterized in that The step of determining a plurality of target information sources that match the query intent type from a plurality of preset information sources includes: Determining quality feature data of each information source in the multiple types of information sources in different quality feature dimensions, and demand preference data of the query intention type for the quality feature dimensions; Determining a matching matrix between the query intention type and the multiple types of information sources based on the demand preference data and the quality feature data; The information sources corresponding to the preset values of the elements in the matching matrix are determined as target information sources matching the query intention type, and multiple target information sources are obtained.
3. The method according to claim 1, characterized in that The step of searching each target information source according to the intent representation and the constraint condition to obtain an information item corresponding to the intent representation includes: For each of the target information sources, searching the target information source according to the intention representation and the constraint condition to determine the original text that matches the intention representation; If the type of the original text is an unstructured text type, information extraction is performed on the original text to determine the information item corresponding to the intention representation from the original text.
4. The method according to claim 1, wherein The fusing of the plurality of information items to generate a summary answer includes: Performing semantic alignment on the plurality of information items to obtain first candidate information items after alignment of the semantic information corresponding to the respective items; For each of the first candidate information items, determining the confidence of the first candidate information item in each preset evaluation dimension according to a preset confidence evaluation model; determining a composite confidence level of the first candidate information item according to the confidence level; performing conflict detection on each of the first candidate information items to obtain conflict data between the first candidate information items; generating a corresponding conflict handling strategy according to the composite confidence and the conflict data; A multi-source information fusion prompt is determined according to the first candidate information item, the confidence level, the composite confidence level, and the conflict handling strategy, and a plurality of the first candidate information items are fused to generate a summary answer.
5. The method according to claim 4, characterized in that Generating a corresponding conflict handling strategy according to the composite confidence and the conflict data includes: If, among the first candidate information items, there are multiple second candidate information items whose composite confidence is greater than or equal to the first preset confidence, the generated conflict handling strategy includes sorting the multiple second candidate information items in descending order according to their composite confidences, and determining a priority of each of the second candidate information items based on the sorting result; and / or, if, among the first candidate information items, there is a third candidate information item whose composite confidence is greater than or equal to a second preset confidence, and the conflict data among the plurality of third candidate information items is that there is no conflict, then the generated conflict handling strategy includes performing semantic fusion on the plurality of third candidate information items to generate a neutral summary; and the second preset confidence is greater than the first preset confidence; And / or, if there are multiple groups of fourth candidate information items among the first candidate information items, the difference between any two of which is within a preset range, and the conflicting data between the fourth candidate information items is a conflict, the generated conflict handling strategy includes displaying the conclusion, information source, and the composite confidence of each fourth candidate information item.
6. The method according to claim 1, characterized in that The method further comprises: Determining reference information of the semantic segment in the summary answer, where the reference information includes at least one of an information source, an original text, and attribute information associated with the original text; generating an interactive reference to the semantic fragment; In response to a triggering operation on the interactive reference, the reference information referenced by the semantic segment is displayed.
7. The method according to claim 6, characterized in that The determining of the reference information of the semantic fragment in the summary answer includes: Splitting the summary answer according to a preset semantic format to obtain semantic segments; For each of the semantic segments, encoding the semantic segment and the original information segment in the original text corresponding to the semantic segment to obtain a first vector of the semantic segment and a second vector of each of the original information segments; determining, based on the first vector and each of the second vectors, a similarity between the semantic segment and the original information segment, and determining a preset number of target original information segments from the original information segments based on the similarity; The target original information segment, the information source of the target original information segment, and the attribute information associated with the target original information segment are determined as reference information of the semantic segment in the summary answer.
8. A multi-source information retrieval and fusion device, characterized in that: The device comprises: A data analysis module is used to analyze the acquired user query data to determine the query intent type, intent representation and constraint conditions of the query data; An information source matching module is used to determine multiple target information sources that match the query intent type from a plurality of preset information sources; A retrieval module, configured to search each target information source according to the intention representation and the constraint conditions to obtain information items corresponding to the intention representation; The information fusion module is used to fuse the multiple information items and generate a summary answer.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Fusion query method and device of heterogeneous multi-source data
CN108090154A
Information retrieval method and device, storage medium and computer program product
CN118939763A
Searching method based on multi-source information fusion
CN120104859A
Interaction method and device based on artificial intelligence, equipment and intelligent agent
CN120144837A
System for fully integrated capture, and analysis of business information resulting in predictive decision making and simulation
US10860962B2
Cited By
Information screening method and device, equipment and storage medium
CN121144583A
Information screening method, device, equipment and storage medium
CN121144583B
Book semantic retrieval method and device based on intention recognition and electronic equipment
CN121188143A
Generation method and device of target reply information, computer equipment, readable storage medium and program product
CN121579763A
Domestic operating system-oriented type adaptive agent webpage search method
CN121614659A