Financial information query method and system based on atlas analysis
Through the graph analysis method, the financial query tasks are divided and dependency constructed, and multiple knowledge retrieval models are used to filter and combine query results, which solves the problem of inaccurate query results in the existing technology, and realizes efficient and professional financial information query.
Patent Information
- Application Number
- CN202510751387.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-06
AI Technical Summary
Due to the limitation of keyword matching, lack of deep understanding and scene perception, the existing financial information query system is difficult to accurately locate complex and changeable financial problems, resulting in poor comprehensiveness and accuracy of query results.
Using a graph analysis method, a task execution map is built by dividing the query tasks and dependent relationships, and multiple knowledge retrieval models are used to analyze the query subtasks, and candidate query results with accuracy scores not lower than the threshold are selected, and combined them into the final query results.
It has achieved accurate disassembly and efficient answers to complex financial problems, and improved the professionalism and user experience of knowledge management and retrieval services of financial institutions.
Smart Images

Figure CN120256647A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of intelligent question answering, and in particular, to a financial information query method and system based on graph analysis. Background Art
[0002] In the current banking business environment, with the continuous improvement of business complexity, bank employees are faced with a daunting challenge, that is, they need to master and apply the ever-growing and frequently updated rules and regulations, business operation guidelines, and external supervision policies. These information are not only widely distributed across various channels such as internal documents, external web pages, and professional reports, but also updated extremely fast, requiring employees to maintain a high degree of sensitivity to the latest developments.
[0003] However, existing financial information retrieval systems are usually limited to keyword matching, lacking the ability of in-depth understanding and scenario awareness, and it is difficult to accurately locate the specific details required by employees, especially when dealing with complex and changing financial problems or product features. Therefore, there is an urgent need for a more intelligent and flexible financial information query method to overcome the above problems and improve the overall operation efficiency of financial institutions. Summary of the Invention
[0004] Embodiments of this application provide a financial information query method and system based on graph analysis, so as to at least solve the technical problem that the comprehensiveness and accuracy of the final query result obtained are poor because the financial information query technology based on keyword matching does not deeply understand the semantics of the query task.
[0005] According to one aspect of the embodiments of this application, a financial information query method based on graph analysis is provided, including: obtaining a query task input by a target object, where the query task is used to reflect the financial problem to be queried by the target object; dividing the query task into multiple query subtasks, and sorting the multiple query subtasks according to the dependency relationship between the multiple query subtasks to obtain a task execution graph corresponding to the query task; for each query subtask, using multiple knowledge retrieval models to analyze the query subtask and the target query results of all executed query subtasks that have a dependency relationship with the query subtask, obtaining at least one first candidate query result whose accuracy score output by each knowledge retrieval model is not lower than a preset first score threshold, and determining the target query result of the query subtask based on each first candidate query result; according to the task execution graph, combining the target query results of each query subtask to obtain the final query result of the query task, and feeding back the final query result to the target object.
[0006] Optionally, the query task is divided into multiple query subtasks, and the multiple query subtasks are sorted according to the dependency relationship between the multiple query subtasks to obtain a task execution graph corresponding to the query task, including: preprocessing the query task, where the preprocessing includes at least one of the following: removing noise characters, standardizing abbreviations, and replacing synonyms; using a named entity recognition algorithm to perform entity recognition on the preprocessed query task to obtain at least one key entity in the query task and determine the attribute information of each key entity, where the key entity includes at least one of the following: financial institution, financial product name, financial market participant, economic indicator, time information, location information; using a pre-trained intent recognition model to analyze the preprocessed query task to obtain the query intent of the query task, where the intent recognition model is trained using a first fine-tuning training dataset, and the first fine-tuning dataset contains multiple query task samples related to the financial field and the query intent label corresponding to each query task sample; dividing the query task into multiple query subtasks according to the attribute information of each key entity and the query intent, and determining the dependency relationship between the query subtasks, where the dependency relationship includes: sequential dependency or parallel dependency; taking each query subtask as a graph node, and determining the directed edges between the corresponding graph nodes according to the dependency relationship between the query subtasks to construct a task execution graph.
[0007] Optionally, multiple knowledge retrieval models are used to analyze the query subtask and the target query results of all executed query subtasks that have a dependency relationship with the query subtask, to obtain at least one first candidate query result whose accuracy score output by each knowledge retrieval model is not lower than a preset first score threshold, including: determining all executed query subtasks that have a dependency relationship with the query subtask according to the task execution graph, and respectively determining the first embedding vector of the target query result of each executed query subtask and the second embedding vector of the query subtask; for each knowledge retrieval model, taking each first embedding vector and the second embedding vector as the first input vector of the knowledge retrieval model, and using the knowledge retrieval model to analyze the first input vector, to obtain the first number of primary query results within the first retrieval time range output by the knowledge retrieval model and the original confidence of each primary query result, where each knowledge retrieval model is trained using a different second fine-tuning training dataset, and each second fine-tuning training dataset contains: multiple query task samples related to the financial field, the correct query result sample of each query task sample, and the confidence of the correct query result sample; for each primary query result, determining multiple confidence calibration values of the primary query result, and determining the uncertainty of the primary query result, where at least some of the multiple confidence calibration values include: semantic relevance, information timeliness, and source credibility; calibrating the original confidence of the primary query result using the multiple confidence calibration values to obtain a calibrated confidence; using a preset weight coefficient to perform weighted summation on the calibrated confidence and the uncertainty of the primary query result to obtain the accuracy score of the primary query result; determining at least one first candidate query result whose accuracy score is not lower than the first score threshold from the first number of primary query results output by each knowledge retrieval model.
[0008] Optionally, determining multiple confidence calibration values of the primary query result includes: determining the third embedding vector of the primary query result, and calculating the vector similarity between the third embedding vector and the second embedding vector, and taking the vector similarity as the semantic relevance of the primary query result; determining the latest update time of the information contained in the primary query result and the required retrieval time range of the query task, and calculating the time fit degree between the latest update time and the required retrieval time range, and taking the time fit degree as the information timeliness of the primary query result; obtaining the credibility of the source of each piece of information in the primary query result; determining the data volume ratio of each piece of information in the primary query result to the total information of the primary query result, and performing weighted summation on the credibility of the source of each piece of information according to the data volume ratio to obtain the source credibility of the primary query result.
[0009] Optionally, determine the uncertainty of the primary selection query results, including: obtaining the probability distribution of the second number of primary selection query results obtained by multiple knowledge retrieval models through multiple forward analyses of the first input vector, where the second number is greater than the first number; determining the uncertainty of the primary selection query results according to the probability distribution and the following formula: , where represents the uncertainty of the i-th primary selection query result, represents the occurrence probability of the i-th primary selection query result, m represents the total number of multiple knowledge retrieval models, n represents the second number of primary selection query results output by a single knowledge retrieval model, and .
[0010] Optionally, determine the target query result of the query sub-task based on each first candidate query result, including: merging each first candidate query result according to the accuracy score to obtain a merged candidate query result and the accuracy score of the merged candidate query result; determining whether the accuracy score of the merged candidate query result is lower than a preset second score threshold, where the second score threshold is not lower than the first score threshold; in the case where the accuracy score of the merged candidate query result is not lower than the second score threshold, using the merged candidate query result as the target query result of the query sub-task; in the case where the accuracy score of the merged candidate query result is lower than the second score threshold, perform the following steps until a merged candidate query result with an accuracy score not lower than the second score threshold is obtained as the target query result of the query sub-task, including: The first step: determine the third embedding vector of the merged candidate query result, and use each first embedding vector, second embedding vector, and third embedding vector as the second input vector; The second step: use multiple knowledge retrieval models to analyze the second input vector again to obtain the third number of primary selection query results within the second retrieval time range output by the knowledge retrieval models and the original confidence of each primary selection query result, where the second retrieval time range is greater than the first retrieval time range, the third number is greater than the first number and not greater than the second number; The third step: determine at least one second candidate query result with an accuracy score not lower than the first score threshold from the third number of primary selection query results output by each knowledge retrieval model; The fourth step: if the accuracy score of the new merged candidate query result is lower than the second score threshold, continue to perform the first step above.
[0011] Optionally, each first candidate query result is merged according to the accuracy score to obtain a merged candidate query result and the accuracy score of the merged candidate query result, including: determining the third embedding vector of each first candidate query result; using the accuracy scores of each first candidate query result to determine the corresponding weighting coefficients, where the accuracy score is proportional to the weighting coefficient; performing weighted summation on the third embedding vectors of each first candidate query result according to the weighting coefficients to obtain a fourth embedding vector, and invoking a preset decoder to decode the fourth embedding vector to obtain the merged candidate query result; performing weighted summation on the accuracy scores of each first candidate query result according to the weighting coefficients to obtain the accuracy score of the merged candidate query result.
[0012] According to another aspect of the embodiments of the present application, there is also provided a financial information query system based on graph analysis, including: an acquisition module, configured to acquire a query task input by a target object, where the query task is used to reflect the financial problem to be queried by the target object; a graph construction module, configured to perform task partitioning on the query task to obtain a plurality of query subtasks, and perform task sorting on the plurality of query subtasks according to the dependency relationship between the plurality of query subtasks to obtain a task execution graph corresponding to the query task; a model retrieval module, configured to, for each query subtask, analyze the query subtask and the target query results of all executed query subtasks having a dependency relationship with the query subtask by using a plurality of knowledge retrieval models, obtain at least one first candidate query result whose accuracy score output by each knowledge retrieval model is not lower than a preset first score threshold, and determine the target query result of the query subtask based on each first candidate query result; a feedback module, configured to combine the target query results of each query subtask according to the task execution graph to obtain the final query result of the query task, and feedback the final query result to the target object.
[0013] According to another aspect of the embodiments of the present application, there is also provided a computer program product, which includes: a computer program, where when the computer program is executed by a processor, it implements the above-mentioned financial information query method based on graph analysis.
[0014] According to another aspect of the embodiments of the present application, there is also provided an electronic device, which includes: a memory and a processor, where a computer program is stored in the memory, and the processor is configured to execute the above-mentioned financial information query method based on graph analysis through the computer program.
[0015] In the embodiments of the present application, through the deep semantic understanding and task-driven framework, the comprehensive queries proposed by users are decomposed into several logically coherent query subtasks, and a task execution graph is constructed to optimize the execution process; an intelligent question-answering center that integrates multiple models is used to ensure that each query subtask can obtain the target query result with the highest accuracy score; finally, the target query results of each query subtask are integrated and combined to obtain the final query result of the query task, and the result is fed back to the target object, achieving the technical effect of accurately disassembling and efficiently answering complex financial problems, and achieving the purpose of greatly improving the professionalism of the knowledge management and retrieval services of financial institutions and the user experience. Furthermore, it solves the technical problem that the comprehensiveness and accuracy of the final query result obtained are poor because the financial information query technology based on keyword matching does not deeply understand the semantics of the query task. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0017] Figure 1 is a schematic flowchart of an optional financial information query method based on graph analysis according to an embodiment of the present application;
[0018] Figure 2 is a schematic structural diagram of an optional financial information query system based on graph analysis according to an embodiment of the present application;
[0019] Figure 3 is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0021] It should be noted that the terms "first", "second", etc. in the description, claims and drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances, so that the embodiments of this application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0022] Embodiment 1
[0023] According to an embodiment of the present application, a financial information query method based on graph analysis is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0024] Figure 1 is a schematic flowchart of a financial information query method based on graph analysis provided according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:
[0025] Step S102, obtain the query task input by the target object.
[0026] In the technical solution provided in the above step S102, the above target object refers to the staff or users of a financial institution, such as front-line salespersons, risk managers, product managers, and other internal or external personnel who obtain financial knowledge. And the above query task refers to the need of the target object to obtain answers to specific financial questions or relevant information through the query system of the financial institution, which includes but is not limited to: detailed consultation of financial products, interpretation of the latest regulatory policies, query of key data in industry research reports, analysis requests for specific customer cases, etc.
[0027] Step S104, perform task division on the query task to obtain multiple query subtasks, and perform task sorting on the multiple query subtasks according to the dependency relationship between the multiple query subtasks to obtain a task execution graph corresponding to the query task.
[0028] In the technical solution provided in step S104 above, the query system of the financial institution uses natural language processing technology to break down the original complex query task into a series of more specific and easily processable query subtasks. Therefore, a query subtask can be understood as an independent information retrieval or calculation unit decomposed from the original query task, and each query subtask may involve different data sources or business systems to represent different steps or aspects of implementing the original query task. Furthermore, according to the dependency relationships between the query subtasks, algorithms such as topological sorting in graph theory are used to generate a directed acyclic graph - a task execution graph to intuitively display the execution order between the query subtasks.
[0029] Step S106, for each query subtask, use multiple knowledge retrieval models to analyze the query subtask and the target query results of all executed query subtasks that have a dependency relationship with the query subtask, obtain at least one first candidate query result whose accuracy score output by each knowledge retrieval model is not lower than a preset first score threshold, and determine the target query result of the query subtask based on each first candidate query result.
[0030] In the technical solution provided in step S106 above, for each query subtask within the query task, multiple knowledge retrieval models can be used for analysis simultaneously, and the context of the query subtask, that is, the output of the executed query subtasks that have been completed and have a dependency relationship with it, is considered to obtain the possible answers (i.e., the first candidate query results) generated by each knowledge retrieval model for the query subtask and the corresponding accuracy scores for each query subtask. Among them, the possible answers include but are not limited to: specific numerical results, excerpts of policies and regulations, descriptions of product features, and conclusions of industry analysis, etc., to directly or indirectly respond to the user's initial inquiry. Furthermore, the query system intelligently filters out the target query results whose accuracy scores are not lower than the preset first score threshold by comparing and evaluating the first candidate query results generated by each model and the corresponding accuracy scores of the query subtasks, thereby ensuring the accuracy and reliability of the query results of the query subtask.
[0031] Step S108, according to the task execution graph, combine the target query results of each query subtask to obtain the final query result of the query task, and feedback the final query result to the target object.
[0032] In the technical solution provided in step S108 above, the query system serially or parallelly combines the target query results of each query subtask as needed according to the dependency relationships and task execution order defined in the task execution graph to obtain the final query result of the query task. Finally, the final query result is feedback to the target object.
[0033] Based on the solution defined in the above steps S102 to S108, it can be learned that in the embodiment of the present application, through the deep semantic understanding and task-driven framework, the comprehensive query proposed by the user is decomposed into several logically coherent query subtasks, and a task execution graph is constructed to optimize the execution process; an intelligent question-answering center that integrates multiple models is used to ensure that the target query results with the highest accuracy score can be obtained for each query subtask; finally, the target query results of each query subtask are integrated and combined to obtain the final query result of the query task, and it is fed back to the target object, achieving the technical effect of accurately disassembling and efficiently answering complex financial problems, and achieving the purpose of greatly improving the professionalism of the knowledge management and retrieval services of financial institutions and the user experience.
[0034] The following describes each step of the financial information query method based on graph analysis in combination with a specific implementation process.
[0035] As an optional implementation manner, in the technical solution provided in the above step S104, the query system can obtain the task execution graph corresponding to the query task according to the following method, including:
[0036] Step S1041, preprocess the query task.
[0037] Among them, the preprocessing performs in-depth cleaning and standardization operations on the query task received by the system, including but not limited to: using regular expressions to identify and remove noise characters (such as special symbols, whitespace, etc.) that may interfere with the model's understanding; cleaning the query task to remove elements such as HTML tags, numbers, and special characters; unifying the format of the task data to eliminate parsing exception problems caused by inconsistent encoding; standardizing abbreviations to restore their complete expressions; and replacing synonyms to enhance the richness and accuracy of semantic expressions. These preprocessing steps help reduce ambiguity in subsequent analysis and improve the overall processing efficiency.
[0038] Step S1042, use the Named Entity Recognition (NER) algorithm to perform entity recognition on the preprocessed query task, obtain at least one key entity in the query task, and determine the attribute information of each key entity.
[0039] Optionally, in the technical solution provided in step S1042 above, during the entity recognition process, the system can first perform word segmentation on the preprocessed query task and determine the part-of-speech of each word obtained after word segmentation, such as verbs, nouns, time expressions, etc.; for each word, extract its feature vector, which includes the part-of-speech, the part-of-speech of the preceding and following words, the morphological information of the word itself (such as the first letter being capitalized, whether it is a number, etc.), and based on the dictionary and financial domain knowledge; then, use a deep learning model based on the named entity recognition algorithm to analyze the feature vectors of each word to determine which entity type each word belongs to, such as financial institutions, financial product names, financial market participants, economic indicators, time information, location information. Further, the system further extracts the attribute information of each key entity from the preprocessed query task, such as the time range, geographical location, specific numerical requirements of business indicators, etc., and converts it into programmable structured parameters.
[0040] Step S1043, use the pre-trained intent recognition model to analyze the preprocessed query task to obtain the query intent of the query task.
[0041] Specifically, the above intent recognition model is trained using the first fine-tuning training dataset, and the first fine-tuning dataset contains multiple query task samples related to the financial domain and the corresponding query intent labels for each query task sample. Among them, the query task samples are a large number of query examples related to the financial domain collected in advance, such as user questions, customer service conversation records, industry forum posts, etc., to ensure coverage of various financial business and service scenarios; then, financial experts manually annotate the collected query examples to clarify the query intent category of each query example (such as transaction consultation, product recommendation, regulation interpretation, market analysis, etc.), which is used as the query intent label of the query task sample. It should be noted that the first fine-tuning training dataset after annotation can be balanced to ensure that the number of samples for each query intent is relatively uniform, avoiding the model from making biased misjudgments due to data skew.
[0042] In the technical solution provided in step S1043 above, the system first extracts features from the preprocessed query task, paying particular attention to semantic features at the lexical and sentence levels; then, the system inputs the obtained query task feature vector into the intent recognition model, and through the forward propagation process, predicts the query intent of the query task. Among them, the output layer of the model is a multi-classifier that can output a series of intent labels and their confidence levels; finally, the system can select the intent with the highest probability or meeting the preset threshold as the final intent recognition result according to the intent labels and their probability distributions output by the model.
[0043] Step S1044, dividing the query task into multiple query subtasks according to the attribute information of each key entity and the query intent, and determining the dependency relationship between each query subtask.
[0044] In the technical solution provided in the above step S1044, the system can dynamically match the predefined query subtask template according to the attribute information and query intent of each key entity, and fill the structured attribute information of each key entity into the corresponding position in the template to generate multiple query subtask instructions, so that each query subtask obtained by division focuses on one aspect of the original query task. After completing the subtask decomposition, the system can encapsulate each query subtask into an independent work unit, and clarify its input, output and execution conditions, so as to determine the logic and data dependencies between each query subtask, and obtain the dependency relationship between each query subtask, such as sequential dependency or parallel dependency.
[0045] Step S1045, taking each query subtask as a graph node, and determining the directed edges between the corresponding graph nodes according to the dependency relationships between the query subtasks, forming a directed acyclic graph, and constructing a task execution graph.
[0046] Therefore, the task execution map obtained through the above steps S1041-S1045 not only clarifies the execution order of the query subtasks, but also intuitively displays the overall picture of the query process, making it easier for the system to monitor the progress in real time and intervene in adjustments in a timely manner to ensure the efficiency and orderliness of the query process.
[0047] In an exemplary embodiment, the query task proposed by user A is: "Changes in mortgage interest rates of Bank A in the past five years". The query system will first clean up irrelevant characters (such as punctuation and stop words) in the query task and unify its format; then use the named entity recognition algorithm to perform entity recognition on the pre-processed query task, and use "Bank A" (i.e. financial institution), "mortgage interest rate" (i.e. economic indicator) and "past five years" (i.e. time range) as key entities; then, through the intent recognition model analysis, it is confirmed that the query task is a report requirement query that requires time series data.
[0048] Based on this, the query system decomposes the original query task into the following three core subtasks:
[0049] Subtask 1: "Search for documents related to mortgage interest rate adjustments over the past five years";
[0050] Subtask 2: "Analyze the annual reports of Bank A over the years and extract the relevant sections on mortgage interest rate adjustments";
[0051] Subtask 3: "Compare the results of Subtask 1 and Subtask 2 to identify the key moments and influencing factors of changes in mortgage interest rates."
[0052] Next, determine the sequential dependencies among these three subtasks, that is, subtask 2 depends on subtask 1, and subtask 3 depends on subtask 2 and subtask 1. Therefore, the execution order of these three subtasks is: subtask 1 → subtask 2 → subtask 3.
[0053] Finally, the query system can use these three core subtasks as graph nodes, determine the edges between these three graph nodes based on the dependencies (i.e., sequential dependencies) among these three core subtasks, and construct a task execution graph.
[0054] In this way, the query system not only realizes the automation and orderliness of the query process, but also ensures the efficiency and accuracy of information retrieval and processing, and finally feeds back the integrated non-performing loan rate change curve to the user.
[0055] As an alternative implementation, in the technical solution provided in step S106 above, the system can call the knowledge retrieval model according to the following steps to obtain at least one first candidate query result whose accuracy score output by each knowledge retrieval model is not lower than a preset first score threshold, including:
[0056] Step S1061, determine all executed query subtasks that have dependencies with the query subtask according to the task execution graph, and respectively determine the first embedding vectors of the target query results of each executed query subtask and the second embedding vector of the query subtask.
[0057] In the technical solution provided in step S1061 above, the query system can identify a list of historical executed subtasks that have direct or indirect dependencies with the currently to-be-executed query subtask according to the directed edges in the task execution graph, and extract the target query results of each executed query subtask in the list of historical executed subtasks from the database. Next, use a pre-trained semantic embedding model (such as BERT, Sentence-BERT) to deeply encode the target query results of each executed query subtask to generate multiple first embedding vectors of a fixed dimension; at the same time, perform similar encoding on the task description of the currently to-be-executed query subtask to obtain the corresponding second embedding vector, where the second embedding vector captures the semantic features, business intentions, and execution environment details of the currently to-be-executed query subtask.
[0058] Step S1062, for each knowledge retrieval model, use each first embedding vector and the second embedding vector as the first input vector of the knowledge retrieval model, and use the knowledge retrieval model to analyze the first input vector to obtain the first number of primary query results within the first retrieval time range output by the knowledge retrieval model and the original confidence of each primary query result.
[0059] Among them, the types of the above-mentioned knowledge retrieval models include but are not limited to: vector database retrieval models, knowledge graph query models, structured data query models, etc. The selection of the model type can be set according to the actual application scenario. In addition, each knowledge retrieval model is trained using a second fine-tuning training dataset, where the second fine-tuning training dataset contains: a large number of query task samples related to the financial field, the correct query result samples for each query task sample, and the confidence levels of the correct query result samples.
[0060] Therefore, in the technical solution provided in step S1062 above, the system can splice the generated first embedding vectors and second embedding vectors or perform weighted merging using an attention mechanism to form a first input vector containing historical context information, which can ensure that the feature vectors input into the model contain both the key features of the current query and the execution results of the previous tasks, greatly enhancing the context awareness ability of the model when processing complex query tasks. Then, the system parallelly invokes multiple knowledge retrieval models, uses the first input vector as the input, and drives each model to perform in-depth retrieval in its respective responsible financial data sources (such as research reports, financial statements). Each model can output the top first number of records highly relevant to the current query subtask within the first retrieval time range as the primary query results based on the learning results of its corresponding second fine-tuning training dataset, and each result is also attached with the original confidence level evaluated by the model.
[0061] Step S1063, for each primary query result, determine multiple confidence calibration values of the primary query result and determine the uncertainty of the primary query result; use the multiple confidence calibration values to calibrate the original confidence level of the primary query result to obtain a calibrated confidence level; use a preset weight coefficient to perform a weighted sum of the calibrated confidence level and the uncertainty of the primary query result to obtain an accuracy score of the primary query result.
[0062] Step S1064, determine at least one first candidate query result with an accuracy score not lower than the first score threshold from the first number of primary query results output by each knowledge retrieval model.
[0063] Specifically, in the technical solution provided in the above step S1063, the system takes into account that when the pre-trained large model processes complex query tasks, it may generate "hallucination" information that does not conform to the facts due to complex semantics, professional terms, and specific background knowledge involved in financial issues. At the same time, a single model may perform outstandingly in some aspects while having defects in some aspects, resulting in poor accuracy of the model output results. Therefore, the embodiments of the present application propose to introduce multiple confidence calibration values such as semantic relevance, information timeliness, and source credibility, strengthen the emphasis on authoritative information in the field, and reduce the impact of expired information or information with unknown sources on the quality of query results.
[0064] Optionally, the query system can determine multiple confidence calibration values of the primary query results according to the following steps, including:
[0065] The first step: Determine the third embedding vector of the primary query result, calculate the vector similarity between the third embedding vector and the second embedding vector, and use the vector similarity as the semantic relevance (SemanticRelevance) of the primary query result.
[0066] The second step: Determine the latest update time of the information contained in the primary query result and the required retrieval time range of the query task, and calculate the time fit between the latest update time and the required retrieval time range. Use the time fit as the information timeliness (Temporal Relevance) of the primary query result, where the latest update time is within the first retrieval time range, and the required retrieval time range is usually the preset time range between the input query task of the target object.
[0067] The third step: Obtain the credibility of the source of each piece of information in the primary query result; determine the data volume ratio of each piece of information in the primary query result to the total information of the primary query result, and perform weighted summation on the credibility of the source of each piece of information according to the data volume ratio to obtain the source credibility (Source Credibility) of the primary query result.
[0068] In the above embodiment, the system can determine multi-dimensional confidence calibration values from the semantic level, information timeliness level, and information source level respectively, where:
[0069] At the semantic level, the system can use the semantic embedding model to deeply encode the description content of the primary query result to obtain the corresponding third embedding vector, and calculate the vector similarity between the third embedding vector and the second embedding vector of the query subtask. The higher the vector similarity, the higher the semantic relevance, and the stronger the calibrated confidence; on the contrary, the lower the vector similarity, the lower the semantic relevance, and the weaker the calibrated confidence. Therefore, this vector similarity reflects the pertinence and directness of the primary query result.
[0070] At the information timeliness level, the system compares the latest update time of the information contained in the primary query results with the required retrieval time range of the query task to calculate the time fitness. Among them, if the latest update time of the information contained in the primary query results occurs within the required retrieval time range of the query task, the higher the time fitness, the stronger the calibrated confidence; conversely, if the latest update time of the information contained in the primary query results does not occur within the required retrieval time range of the query task, the lower the time fitness, the weaker the calibrated confidence. Therefore, the information timeliness reflects the timeliness and practical significance of the primary query results.
[0071] At the information source level, the system can first obtain the credibility of the sources of each piece of information in the primary query results. Among them, the credibility of the source can be comprehensively determined by the information source, the number of citations, and the matching degree with financial domain knowledge; then, determine the data volume ratio of each piece of information in the primary query results to the total information in the primary query results, and perform weighted summation on the credibility of the sources of each piece of information according to the data volume ratio to obtain the source credibility of the primary query results. Among them, the higher the source credibility, the stronger the calibrated confidence.
[0072] It should be noted that the determination processes of these three confidence calibration values can be executed in parallel according to the actual application scenario to improve the system processing efficiency.
[0073] In addition, to help the system evaluate the confidence level of each model in the output results, the embodiments of the present application propose to calculate the uncertainty of each primary query result output by the model.
[0074] Optionally, the query system can determine the uncertainty of the primary query results according to the following steps, including:
[0075] The first step: Obtain the probability distribution of the second number of primary query results obtained by multiple forward analyses of the first input vector by multiple knowledge retrieval models. This is because, since multiple knowledge retrieval models participating in the collaborative query will perform multiple forward calculations on the first input vector, and each calculation simulates different inference paths by activating different parts of the model or using different parameter sets to obtain multiple output results and their corresponding probability distributions of each knowledge retrieval model for the same input vector. Therefore, the second number is greater than the first number.
[0076] The second step: According to the probability distribution, determine the uncertainty of the primary query results according to the following formula: ,
[0077] In the formula represents the uncertainty of the i-th primary query result, represents the occurrence probability of the i-th preliminary query result, m represents the total number of multiple knowledge retrieval models, n represents the second number of preliminary query results output by a single knowledge retrieval model, and 。
[0078] Among them, according to the expression of the uncertainty of the above-mentioned preliminary query results, it can be known that if the occurrence probability of the i-th preliminary query result is very high, while the probabilities of other results are relatively low, then the uncertainty of the system for this preliminary query result will be relatively low; on the contrary, if the distributions of the preliminary query results are relatively uniform, then the uncertainty of the i-th preliminary query result will be relatively high.
[0079] Furthermore, the preset weight coefficients are used to perform weighted summation on the calibration confidence and uncertainty of the preliminary query results to obtain the accuracy score of the preliminary query results.
[0080] Since a high occurrence probability usually means that the model has confidence in this first candidate query result, but it is not necessarily always correct, especially when the model encounters ambiguous or rare scenarios. Therefore, in the embodiments of the present application, the calibration confidence and uncertainty of the first candidate query result are weighted and summed to obtain an accuracy score that can accurately reflect the accuracy of the first candidate query result.
[0081] Among them, the setting of the weight values for the calibration confidence and uncertainty can be set according to the actual application scenario. For example, in scenarios such as financial risk assessment or legal consultation that require extremely high accuracy, the fusion weight of the calibration confidence (such as 0.7) should be higher than the fusion weight of the uncertainty (such as 0.3) to ensure the reliability of the results; in scenarios such as market trend analysis or general business consultation that do not require high accuracy, the fusion weight of the uncertainty can be appropriately increased (such as 0.5), and the fusion weight of the calibration confidence can be appropriately decreased (such as 0.5). In addition, the actual contributions of the calibration confidence and uncertainty to the quality of the query results should be continuously monitored, so as to dynamically adjust the fusion weights of the calibration confidence and uncertainty.
[0082] Furthermore, after the query system obtains at least one first candidate query result whose accuracy score output by each knowledge retrieval model is not lower than the preset first score threshold, it can also determine the target query result of the query sub-task based on each first candidate query result. The specific implementation method is as follows:
[0083] Step S1064, merge each first candidate query result according to the accuracy score to obtain a merged candidate query result and the accuracy score of the merged candidate query result.
[0084] In the technical solution provided in step S1064 above, the implementation steps of the method include:
[0085] First, determine the third embedding vector of each first candidate query result, that is, use a vectorization model to convert the description content of each first candidate query result with a high accuracy score into a vector form, and obtain the third embedding vector corresponding to each first candidate query result.
[0086] Next, determine the corresponding weighting coefficient using the accuracy score of each first candidate query result. Among them, the accuracy score is proportional to the weighting coefficient, that is, the higher the accuracy score, the greater the corresponding weighting coefficient should be.
[0087] Then, perform a weighted sum of the third embedding vectors of each first candidate query result according to the weighting coefficient to obtain a fourth embedding vector, and call a preset decoder (such as a Transformer decoder, etc.) to decode the fourth embedding vector to obtain a combined candidate query result. At the same time, perform a weighted sum of the accuracy scores of each first candidate query result according to the weighting coefficient to obtain the accuracy score of the combined candidate query result.
[0088] Step S1065, determine whether the accuracy score of the combined candidate query result is lower than the second score threshold, where the second score threshold is not lower than the first score threshold.
[0089] Step S1066, in the case where the accuracy score of the combined candidate query result is not lower than the second score threshold, use the combined candidate query result as the target query result of the query sub-task.
[0090] Step S1067, in the case where the accuracy score of the combined candidate query result is lower than the second score threshold, perform the following steps until a combined candidate query result with an accuracy score not lower than the second score threshold is obtained as the target query result of the query sub-task, including:
[0091] The first step: Determine the third embedding vector of the combined candidate query result, and use each first embedding vector, second embedding vector, and third embedding vector as the enhanced second input vector.
[0092] Step 2: Analyze the second input vector again using multiple knowledge retrieval models to obtain the third number of primary query results within the second retrieval time range output by the knowledge retrieval models and the original confidence levels of each primary query result. Here, the second retrieval time range is greater than the first retrieval time range, and the third number is greater than the first number and not greater than the second number, so as to expand the retrieval time range and increase the number of output results, thereby increasing the probability of finding high-quality query results; determine at least one second candidate query result with an accuracy score not lower than the first score threshold from the third number of primary query results output by each knowledge retrieval model.
[0093] Step 3: Merge each second candidate query result according to the accuracy score to obtain a newly merged candidate query result and the accuracy score of the newly merged candidate query result.
[0094] Step 4: If the accuracy score of the newly merged candidate query result is lower than the second score threshold, continue to execute the above Step 1.
[0095] As an alternative implementation, in the technical solution provided in the above step S108, the query system can obtain the final query result of the query task according to the following method, including: determining the modal type of the target query result of each query subtask, where the modal type can be text, image, video, audio, etc.; using multimodal fusion technology (such as cross-modal encoder) to unify the target query results of different modal types into one expression framework; sorting the target query results of each query subtask according to the dependency relationship defined in the task execution graph to obtain the final query result of the query task. Among them, the final query result can be visually displayed in a structured format (such as generating a report, filling a table, or constructing a visualization chart, etc.) to facilitate the target object to understand and use.
[0096] Embodiment 2
[0097] According to the embodiments of the present application, there is also provided a graph analysis-based financial information query system for implementing the graph analysis-based financial information query method in Embodiment 1, as Figure 2 shown. The graph analysis-based financial information query system at least includes: an acquisition module 22, a graph construction module 24, a model retrieval module 26, and a feedback module 28, where:
[0098] The acquisition module 22 is used to acquire the query task input by the target object, where the query task is used to reflect the financial problem to be queried by the target object;
[0099] The graph construction module 24 is used to divide the query task into multiple query subtasks, and sort the multiple query subtasks according to the dependency relationship between the multiple query subtasks to obtain a task execution graph corresponding to the query task;
[0100] The model retrieval module 26 is used for each query subtask to analyze the query subtask and the target query results of all executed query subtasks having a dependency relationship with the query subtask by using multiple knowledge retrieval models, to obtain at least one first candidate query result whose accuracy score output by each knowledge retrieval model is not lower than a preset first score threshold, and to determine the target query result of the query subtask based on each first candidate query result;
[0101] The feedback module 28 is used to combine the target query results of each query subtask according to the task execution graph to obtain the final query result of the query task, and feedback the final query result to the target object.
[0102] The functions of each module of the financial information query system based on graph analysis are described below in combination with a specific implementation process.
[0103] As an optional implementation manner, the graph construction module 24 may construct a task execution graph corresponding to the query task according to the following steps, including:
[0104] The first step: preprocess the query task, where the preprocessing includes at least one of the following: removing noise characters, standardizing abbreviations, and replacing synonyms;
[0105] The second step: use a named entity recognition algorithm to perform entity recognition on the preprocessed query task to obtain at least one key entity in the query task, and determine the attribute information of each key entity, where the key entity includes at least one of the following: financial institution, financial product name, financial market participant, economic indicator, time information, location information;
[0106] The third step: use a pre-trained intent recognition model to analyze the preprocessed query task to obtain the query intent of the query task, where the intent recognition model is trained using a first fine-tuning training dataset, and the first fine-tuning dataset contains multiple query task samples related to the financial field and the query intent label corresponding to each query task sample;
[0107] The fourth step: divide the query task into multiple query subtasks according to the attribute information and query intent of each key entity, and determine the dependency relationship between each query subtask, where the dependency relationship includes: sequential dependency or parallel dependency;
[0108] Fifth step: Take each query subtask as a graph node, and determine the directed edges between the corresponding graph nodes according to the dependencies between the query subtasks, and construct a task execution graph.
[0109] As an optional implementation manner, the model retrieval module 26 can obtain the target query results of each query subtask according to the following steps, including:
[0110] Step S1, determine all executed query subtasks that have dependencies with the query subtask according to the task execution graph, and respectively determine the first embedding vectors of the target query results of each executed query subtask and the second embedding vector of the query subtask;
[0111] Step S2, for each knowledge retrieval model, take each of the first embedding vectors and the second embedding vector as the first input vector of the knowledge retrieval model, and use the knowledge retrieval model to analyze the first input vector to obtain the first number of primary query results within the first retrieval time range and the original confidence levels of each primary query result. Among them, each knowledge retrieval model is trained using different second fine-tuning training data sets, and each second fine-tuning training data set includes: multiple query task samples related to the financial field, the correct query result samples of each query task sample, and the confidence levels of the correct query result samples;
[0112] Step S3, for each primary query result, determine multiple confidence calibration values of the primary query result, and determine the uncertainty of the primary query result. Among them, at least the multiple confidence calibration values include: semantic relevance, information timeliness, and source credibility; use the multiple confidence calibration values to calibrate the original confidence level of the primary query result to obtain a calibrated confidence level; use a preset weight coefficient to perform weighted summation on the calibrated confidence level and the uncertainty of the primary query result to obtain the accuracy score of the primary query result;
[0113] Step S4, determine at least one first candidate query result whose accuracy score is not lower than the first score threshold from the first number of primary query results output by each knowledge retrieval model.
[0114] Step S5, merge the first candidate query results according to the accuracy scores to obtain a merged candidate query result and the accuracy score of the merged candidate query result;
[0115] Step S6, determine whether the accuracy score of the merged candidate query result is lower than a preset second score threshold, where the second score threshold is not lower than the first score threshold. Specifically, when the accuracy score of the merged candidate query result is not lower than the second score threshold, execute Step S7; otherwise, execute Step S8.
[0116] Step S7, use the merged candidate query results as the target query results of the query sub-task;
[0117] Step S8, execute the following steps until the merged candidate query results with an accuracy score not lower than the second score threshold are used as the target query results of the query sub-task, including:
[0118] The first step: Determine the third embedding vector of the merged candidate query results, and use each first embedding vector, second embedding vector, and third embedding vector as the second input vector.
[0119] The second step: Use multiple knowledge retrieval models to analyze the second input vector again, obtain the third number of primary query results within the second retrieval time range output by the knowledge retrieval models and the original confidence of each primary query result, where the second retrieval time range is greater than the first retrieval time range, the third number is greater than the first number and not greater than the second number, to expand the retrieval time range and increase the number of output results, thereby increasing the probability of finding high-quality query results; determine at least one second candidate query result with an accuracy score not lower than the first score threshold from the third number of primary query results output by each knowledge retrieval model.
[0120] The third step: Merge each second candidate query result according to the accuracy score to obtain a new merged candidate query result and the accuracy score of the new merged candidate query result.
[0121] The fourth step: If the accuracy score of the new merged candidate query result is lower than the second score threshold, continue to execute the first step above.
[0122] Optionally, in the technical solution provided in the above step S3, the model retrieval module 26 can determine multiple confidence calibration values of the primary query results according to the following method, including: determining the third embedding vector of the primary query results, calculating the vector similarity between the third embedding vector and the second embedding vector, and using the vector similarity as the semantic relevance of the primary query results; determining the latest update time of the information contained in the primary query results and the required retrieval time range of the query task, and calculating the time fit between the latest update time and the required retrieval time range, and using the time fit as the information timeliness of the primary query results; obtaining the credibility of the source of each piece of information in the primary query results; determining the data volume ratio of each piece of information in the primary query results to the total information of the primary query results, and performing a weighted sum of the credibility of the sources of each piece of information according to the data volume ratio to obtain the source credibility of the primary query results.
[0123] Optionally, in the technical solution provided in step S3 above, the model retrieval module 26 may determine the uncertainty of the primary query results according to the following method, including: obtaining the probability distribution of the second number of primary query results obtained by multiple forward analyses of the first input vector by multiple knowledge retrieval models, where the second number is greater than the first number; and determining the uncertainty of the primary query results according to the probability distribution and the following formula: ,
[0124] In the formula represents the uncertainty of the i-th primary query result, represents the occurrence probability of the i-th primary query result, m represents the total number of multiple knowledge retrieval models, and n represents the second number of primary query results output by a single knowledge retrieval model.
[0125] Optionally, in the technical solution provided in step S5 above, the model retrieval module 26 may determine the merging of each first candidate query result according to the following method, including: determining the respective third embedding vectors of each first candidate query result; using the accuracy scores of each first candidate query result to determine the corresponding weighting coefficients, where the accuracy scores are proportional to the weighting coefficients; performing weighted summation on the third embedding vectors of each first candidate query result according to the weighting coefficients to obtain a fourth embedding vector, and calling a preset decoder to decode the fourth embedding vector to obtain a merged candidate query result; and performing weighted summation on the accuracy scores of each first candidate query result according to the weighting coefficients to obtain the accuracy score of the merged candidate query result.
[0126] It should be noted that each module in the financial information query system based on graph analysis in the embodiments of the present application corresponds one-to-one to each implementation step of the financial information query method based on graph analysis in Embodiment 1. Since the details not shown in this embodiment can be referred to in Embodiment 1, they will not be elaborated here.
[0127] Embodiment 3
[0128] According to an embodiment of the present application, there is also provided a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the financial information query method based on graph analysis in Embodiment 1.
[0129] According to an embodiment of the present application, there is also provided a non-volatile storage medium, which includes a stored computer program. When the device where the non-volatile storage medium is located runs the computer program, it executes the financial information query method based on graph analysis in Embodiment 1.
[0130] According to an embodiment of the present application, a processor is further provided. The processor is used to run a computer program. When the computer program runs, it executes the financial information query method based on graph analysis in Embodiment 1.
[0131] According to an embodiment of the present application, an electronic device is further provided. The electronic device includes: a memory and a processor. Among them, a computer program is stored in the memory, and the processor is configured to execute the financial information query method based on graph analysis in Embodiment 1 through the computer program.
[0132] Specifically, when the computer program runs, it executes the following steps: obtaining a query task input by a target object, where the query task is used to reflect the financial problem to be queried by the target object; performing task division on the query task to obtain multiple query subtasks, and sorting the multiple query subtasks according to the dependency relationship between the multiple query subtasks to obtain a task execution graph corresponding to the query task; for each query subtask, using multiple knowledge retrieval models to analyze the query subtask and the target query results of all executed query subtasks that have a dependency relationship with the query subtask, obtaining at least one first candidate query result whose accuracy score output by each knowledge retrieval model is not lower than a preset first score threshold, and determining the target query result of the query subtask based on each first candidate query result; according to the task execution graph, combining the target query results of each query subtask to obtain the final query result of the query task, and feeding back the final query result to the target object.
[0133] As an alternative embodiment, the above electronic device may exist in the form of a mobile terminal, a computer terminal, or a similar computing system. Figure 3 A hardware structure block diagram of an electronic device for implementing the financial information query method based on graph analysis is shown. As Figure 3 shown, the electronic device 30 may include one or more (shown as 302a, 302b,..., 302n in the figure) processors 302 (the processor 302 may include, but is not limited to, a processing system such as a microprocessor MCU or a programmable logic device FPGA), a memory 304 for storing data, and a transmission system 306 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 3 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the electronic device 30 may further include more or fewer components than Figure 3 shown, or have a different configuration from Figure 3 shown.
[0134] It should be noted that one or more of the above-mentioned processors 302 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of the other components in the electronic device 30. As involved in the embodiments of the present application, the data processing circuit is a processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0135] The memory 304 can be used to store software programs and modules of application software, such as the program instructions / data storage system corresponding to the financial information query method based on atlas analysis in the embodiments of the present application. The processor 302 executes various functional applications and data processing by running the software programs and modules stored in the memory 304, that is, implements the vulnerability detection method of the above-mentioned application program. The memory 304 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage systems, flash memory, or other non-volatile solid-state memories. In some instances, the memory 304 can further include a memory remotely set relative to the processor 302, and these remote memories can be connected to the electronic device 30 through a network. Examples of the above-mentioned network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0136] The transmission system 306 is used to receive or send data via a network. Specific examples of the above-mentioned network can include a wireless network provided by the communication provider of the electronic device 30. In one instance, the transmission system 306 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission system 306 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0137] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the electronic device 30.
[0138] The above-mentioned embodiment numbers are only for description and do not represent the advantages or disadvantages of the embodiments.
[0139] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0140] In several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the system embodiments described above are merely illustrative. For example, the division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0141] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0142] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0143] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. The aforementioned storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs and other various media that can store program codes.
[0144] The above is only the preferred embodiment of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A financial information query method based on atlas analysis, characterized in that Including: Obtain a query task input by a target object, where the query task is used to reflect a financial problem to be queried by the target object; Perform task partitioning on the query task to obtain multiple query subtasks, and perform task sorting on the multiple query subtasks according to the dependency relationship between the multiple query subtasks to obtain a task execution graph corresponding to the query task; For each query subtask, use multiple knowledge retrieval models to analyze the query subtask and the target query results of all executed query subtasks that have a dependency relationship with the query subtask, obtain at least one first candidate query result whose accuracy score output by each knowledge retrieval model is not lower than a preset first score threshold, and determine the target query result of the query subtask based on each first candidate query result; According to the task execution graph, combine the target query results of each query subtask to obtain the final query result of the query task, and feedback the final query result to the target object.
2. The method according to claim 1, wherein Performing task partitioning on the query task to obtain multiple query subtasks, and performing task sorting on the multiple query subtasks according to the dependency relationship between the multiple query subtasks to obtain a task execution graph corresponding to the query task, including: Preprocess the query task, where the preprocessing includes at least one of the following: removing noise characters, standardizing abbreviations, and replacing synonyms; Use a named entity recognition algorithm to perform entity recognition on the preprocessed query task to obtain at least one key entity in the query task, and determine the attribute information of each key entity, where the key entity includes at least one of the following: financial institution, financial product name, financial market participant, economic indicator, time information, location information; Use a pre-trained intent recognition model to analyze the preprocessed query task to obtain the query intent of the query task, where the intent recognition model is trained using a first fine-tuning training dataset, and the first fine-tuning dataset contains multiple query task samples related to the financial field and the query intent label corresponding to each query task sample; According to the attribute information of each key entity and the query intent, divide the query task into multiple query subtasks, and determine the dependency relationship between each query subtask, where the dependency relationship includes: sequential dependency or parallel dependency; Take each query subtask as a graph node, and determine the directed edges between the corresponding graph nodes according to the dependency relationship between each query subtask, and construct the task execution graph.
3. The method according to claim 1, characterized in that, Using multiple knowledge retrieval models to analyze the query subtask and the target query results of all executed query subtasks that have a dependency relationship with the query subtask, to obtain at least one first candidate query result whose accuracy score output by each knowledge retrieval model is not lower than a preset first score threshold, including: Determine all executed query subtasks that have a dependency relationship with the query subtask according to the task execution graph, and respectively determine the first embedding vector of the target query result of each executed query subtask and the second embedding vector of the query subtask; For each of the knowledge retrieval models, use the first embedding vectors and the second embedding vector as the first input vector of the knowledge retrieval model, and analyze the first input vector by using the knowledge retrieval model to obtain the first number of primary query results within the first retrieval time range and the original confidence levels of each of the primary query results. Among them, each of the knowledge retrieval models is trained by using different second fine-tuning training data sets, and each of the second fine-tuning training data sets includes: a plurality of query task samples related to the financial field, the correct query result sample of each query task sample, and the confidence level of the correct query result sample; For each of the primary query results, determine a plurality of confidence calibration values of the primary query result, and determine the uncertainty of the primary query result. Among them, at least some of the plurality of confidence calibration values include: semantic relevance, information timeliness, and source credibility; use the plurality of confidence calibration values to calibrate the original confidence level of the primary query result to obtain a calibrated confidence level; use a preset weight coefficient to perform a weighted sum of the calibrated confidence level and the uncertainty of the primary query result to obtain the accuracy score of the primary query result; Determine at least one first candidate query result whose accuracy score is not lower than the first score threshold from the first number of primary query results output by each of the knowledge retrieval models.
4. The method according to claim 3, wherein Determining the plurality of confidence calibration values of the primary query result includes: Determine the third embedding vector of the primary query result, calculate the vector similarity between the third embedding vector and the second embedding vector, and use the vector similarity as the semantic relevance of the primary query result; Determine the latest update time of the information contained in the primary query result and the required retrieval time range of the query task, and calculate the time fit degree between the latest update time and the required retrieval time range, and use the time fit degree as the information timeliness of the primary query result. Among them, the latest update time is within the first retrieval time range, and the required retrieval time range is usually a preset time range between the time when the target object inputs the query task; Obtain the credibility of the source of each piece of information in the primary query result; determine the data volume ratio of each piece of information in the primary query result to the total information of the primary query result, and perform a weighted sum of the credibility of the source of each piece of information according to the data volume ratio to obtain the source credibility of the primary query result.
5. The method according to claim 3, wherein Determining the uncertainty of the primary query result includes: Obtain the probability distribution of the second number of primary query results obtained by performing multiple forward analyses on the first input vector by multiple of the knowledge retrieval models, where the second number is greater than the first number; According to the probability distribution, determine the uncertainty of the primary query results according to the following formula: , Wherein represents the uncertainty of the i-th preliminary search result, represents the occurrence probability of the i-th preliminary search result, m represents the total number of multiple knowledge retrieval models, n represents the second number of preliminary search results output by a single said knowledge retrieval model, and .
6. The method according to claim 3, characterized in that, Determine the target query result of the query sub-task based on each of the first candidate query results, including: Merge each of the first candidate query results according to the accuracy score to obtain a merged candidate query result and the accuracy score of the merged candidate query result; Determine whether the accuracy score of the merged candidate query result is lower than the preset second score threshold, where the second score threshold is not lower than the first score threshold; In the case where the accuracy score of the merged candidate query result is not lower than the second score threshold, use the merged candidate query result as the target query result of the query sub-task; In the case where the accuracy score of the merged candidate query result is lower than the second score threshold, perform the following steps until a merged candidate query result with an accuracy score not lower than the second score threshold is obtained as the target query result of the query sub-task, including: The first step: Determine the third embedding vector of the merged candidate query result, and use each of the first embedding vectors, the second embedding vector, and the third embedding vector as the second input vector; The second step: Analyze the second input vector again using multiple knowledge retrieval models to obtain the third number of primary query results within the second retrieval time range output by the knowledge retrieval models and the original confidence of each of the primary query results, where the second retrieval time range is greater than the first retrieval time range, the third number is greater than the first number and not greater than the second number; determine at least one second candidate query result with an accuracy score not lower than the first score threshold from the third number of primary query results output by each knowledge retrieval model; The third step: Merge each of the second candidate query results according to the accuracy score to obtain a new merged candidate query result and the accuracy score of the new merged candidate query result; The fourth step: In the case where the accuracy score of the new merged candidate query result is lower than the second score threshold, continue to perform the first step above.
7. The method according to claim 6, characterized in that, Merge each of the first candidate query results according to the accuracy score to obtain a merged candidate query result and the accuracy score of the merged candidate query result, including: Determine the respective third embedding vectors of each of the first candidate query results; Use the accuracy scores of each of the first candidate query results to determine the corresponding weighting coefficients, where the accuracy score is proportional to the weighting coefficient; Perform a weighted sum of the third embedding vectors of each of the first candidate query results according to the weighting coefficients to obtain a fourth embedding vector, and call a preset decoder to decode the fourth embedding vector to obtain the merged candidate query result; Performing weighted summation on the accuracy scores of each of the first candidate query results according to the weighted coefficients to obtain the accuracy score of the combined candidate query result.
8. A financial information query system based on atlas analysis, characterized in that, Comprising: An acquisition module, configured to acquire a query task input by a target object, where the query task is used to reflect a financial problem to be queried by the target object; A graph construction module, configured to perform task division on the query task to obtain a plurality of query subtasks, and perform task sorting on the plurality of query subtasks according to the dependency relationship between the plurality of query subtasks to obtain a task execution graph corresponding to the query task; A model retrieval module, configured to, for each query subtask, analyze the query subtask and the target query results of all executed query subtasks having a dependency relationship with the query subtask by using a plurality of knowledge retrieval models, obtain at least one first candidate query result whose accuracy score output by each of the knowledge retrieval models is not lower than a preset first score threshold, and determine the target query result of the query subtask based on each of the first candidate query results; A feedback module, configured to combine the target query results of each of the query subtasks according to the task execution graph to obtain a final query result of the query task, and feedback the final query result to the target object.
9. A computer program product, characterized in that, Comprising: A computer program, wherein when the computer program is executed by a processor, it implements the graph analysis-based financial information query method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Comprising: A memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the graph analysis-based financial information query method according to any one of claims 1 to 7 through the computer program.
Citation Information
Patent Citations
Case information semantic retrieval method and device based on knowledge graph
CN111475623A
Data retrieval method, device and equipment
CN117271754A
Resource processing method and device based on knowledge graph and large model
CN118939790A
Time sequence knowledge graph retrieval enhancement generation method and device and medium
CN120069060A
Artificial intelligence-based intelligent associated reply method and apparatus, and computer device
WO2022095357A1
Cited By
Intelligent business data analysis system and method based on knowledge graph
CN120973911A
A knowledge graph-based business data intelligent analysis system and method
CN120973911B
Information query method and electronic equipment
CN121051156A
Information retrieval methods and electronic devices
CN121051156B
Consultation information processing method and system
CN121478830A