Business question and answer intention analysis decision-making method, device and equipment, medium and product
By constructing a multimodal intent prediction network model and a dynamic knowledge graph, the problem of accuracy in intent analysis and decision-making in financial question answering is solved, achieving efficient and accurate decision support in financial scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-10
AI Technical Summary
Existing financial question-answering decision-making methods lack logical reasoning capabilities, resulting in insufficient accuracy in intent analysis and intelligent decision-making. The low efficiency of updating static full-scale knowledge graphs also affects the accuracy of decision output.
A multimodal intent prediction network model is used to perform multi-dimensional analysis of business question-and-answer requests based on text, user behavior, and context. A dynamic target knowledge graph is constructed, and intent understanding and decision-making in financial scenarios are achieved through dynamic adjustment of node edge weights.
It improves the efficiency and accuracy of intent analysis and decision-making in financial scenarios, and achieves accurate fusion of heterogeneous data and real-time decision support.
Smart Images

Figure CN121834498A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of financial technology, and in particular to a business question-and-answer intent analysis decision-making method, apparatus, equipment, medium, and product. Background Technology
[0002] With the digital transformation of the financial industry, the reliance on intelligent decision-making in various financial business scenarios has significantly increased. For example, precise recommendations, customer management, and customer churn analysis all rely heavily on intelligent decision analysis. Currently, intelligent question answering and decision-making in the financial field mainly rely on two technological paths: question answering systems based on natural language to structured query language and static full-scale knowledge graphs.
[0003] However, existing intelligent question-answering decision-making methods often lack logical reasoning capabilities in financial scenarios, resulting in insufficient accuracy in analyzing user intent and making intelligent decisions. Furthermore, static full-scale knowledge graphs require manual updates to ensure the integrity and comprehensiveness of knowledge, and the low update efficiency of knowledge graphs seriously affects the accuracy of the final decision output. Summary of the Invention
[0004] This invention provides a method, apparatus, device, medium, and product for intent analysis and decision-making in business question-and-answer sessions, in order to improve the efficiency and accuracy of intent analysis and decision-making in financial scenarios.
[0005] According to one aspect of the present invention, an intent analysis and decision-making method for business question answering is provided, the method comprising:
[0006] In response to a request from a party regarding a business question and answer request for the current functional business, the request is parsed to determine the business question and answer request text, and the historical user behavior sequence of the requester is obtained, as well as the business scenario information of the current functional business is determined.
[0007] The business question-and-answer request text, the historical user behavior sequence, and the business scenario information are input into a pre-trained multimodal intent prediction network model to obtain intent information; the intent information includes the subject object, the analysis action, and the target metric.
[0008] Based on the subject object, analysis action, and target indicator, a target knowledge graph is constructed; the target knowledge graph includes at least one node and the node edge weight, edge attribute information, and node attribute information of each node;
[0009] Based on each node in the target knowledge graph and its corresponding node edge weights, determine the key information path and at least one secondary information path;
[0010] Based on the node edge weights, edge attribute information, and node attribute information of each node in the key information path and each of the secondary information paths, a target decision result is generated, and the target decision result is fed back to the requester.
[0011] According to another aspect of the present invention, an intent analysis and decision-making apparatus for business question answering is provided, the apparatus comprising:
[0012] The question and answer request response module is used to respond to the requester's business question and answer request for the current functional business, parse the business question and answer request, determine the business question and answer request text, obtain the requester's historical user behavior sequence, and determine the business scenario information of the current functional business;
[0013] The intent information generation module is used to input the business question-and-answer request text, the historical user behavior sequence, and the business scenario information into a pre-trained multimodal intent prediction network model to obtain intent information; the intent information includes the subject object, the analysis action, and the target indicator;
[0014] The target knowledge graph construction module is used to construct a target knowledge graph based on the subject object, analysis action, and target index; the target knowledge graph includes at least one node and the node edge weight, edge attribute information, and node attribute information of each node;
[0015] The information path generation module is used to determine the key information path and at least one secondary information path based on each node in the target knowledge graph and its corresponding node edge weights.
[0016] The analysis and decision-making module is used to generate a target decision result based on the node edge weights, edge attribute information, and node attribute information of each node in the key information path and each of the secondary information paths, and to feed back the target decision result to the requesting party.
[0017] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0018] At least one processor; and
[0019] A memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the intent analysis and decision-making method for business question answering as described in any embodiment of the present invention.
[0021] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the intent analysis and decision-making method for business question answering as described in any embodiment of the present invention.
[0022] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the intent analysis and decision-making method for business question answering as described in any embodiment of the present invention.
[0023] The technical solution of this invention employs a multimodal intent prediction model to analyze and predict intent information from multiple dimensions, including question-and-answer text, user behavior sequences, and business scenarios, in business question-and-answer requests. During the intent prediction process, it bridges data silos by fusing heterogeneous data, accurately integrates the heterogeneous data, and performs intent prediction based on the characteristics of this accurate integration. This achieves high accuracy in understanding complex financial business intents in financial question-and-answer scenarios. By constructing a target knowledge graph in real time based on intent information, including the subject, analysis actions, and target indicators, and dynamically adjusting the node edge weights of the target knowledge graph, it achieves dynamic construction of target knowledge graphs for different question-and-answer scenarios and improves the accuracy of constructing the real-time generated target knowledge graph. Based on this dynamic knowledge graph, path retrieval is performed to obtain the final target decision result, improving the efficiency and accuracy of intent analysis and decision-making in financial question-and-answer scenarios.
[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1A This is a flowchart of a business question-and-answer intent analysis and decision-making method provided in Embodiment 1 of the present invention;
[0027] Figure 1B This is a schematic diagram of the model structure of a multimodal intent prediction network model provided in Embodiment 1 of the present invention;
[0028] Figure 2 This is a flowchart of a business question-and-answer intent analysis and decision-making method provided in Embodiment 2 of the present invention;
[0029] Figure 3 This is a flowchart of a business question-and-answer intent analysis and decision-making method provided in Embodiment 3 of the present invention;
[0030] Figure 4 This is a schematic diagram of the structure of a business question-and-answer intent analysis and decision-making device according to Embodiment 4 of the present invention;
[0031] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the intent analysis and decision-making method for business question answering according to embodiments of the present invention. Detailed Implementation
[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0034] Example 1
[0035] Figure 1A This is a flowchart of a business question-and-answer intent analysis and decision-making method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where intelligent intent analysis and accurate decision-making are performed on user question-and-answer requests in financial scenarios. This method can be executed by a business question-and-answer intent analysis and decision-making device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1A As shown, the method includes:
[0036] S110. In response to the requester's business Q&A request for the current functional business, parse the business Q&A request, determine the business Q&A request text, obtain the requester's historical user behavior sequence, and determine the business scenario information of the current functional business.
[0037] S120. Input the business question and answer request text, historical user behavior sequence and business scenario information into the pre-trained multimodal intent prediction network model to obtain intent information; the intent information includes the subject object, analysis action and target indicator.
[0038] S130. Based on the subject object, analysis actions, and target indicators, construct a target knowledge graph; the target knowledge graph includes at least one node and the node edge weights, edge attribute information, and node attribute information of each node.
[0039] S140. Based on each node in the target knowledge graph and its corresponding node edge weights, determine the key information path and at least one secondary information path.
[0040] S150. Based on the node edge weights, edge attribute information, and node attribute information of each node in the key information path and each secondary information path, generate the target decision result and feed it back to the requester.
[0041] The requester can be either an in-house employee or a customer. The requester can initiate a business-related Q&A request based on the current functional business page. For example, if the requester is an in-house business analyst, they can use the user profile analysis page to request access to a user's transaction records and risk rating within the bank. A possible request could be something like, "Please help me compile a profile analysis report for user XXX," or "Please help me analyze the reasons for the churn of high-end customers in XXX region." If the requester is a customer, they can initiate an asset planning-related Q&A request through the asset planning page. For example, a possible request could be, "Please help me generate an asset planning suggestion or analysis report."
[0042] The requesting party can initiate a business Q&A request based on the front-end page corresponding to different business functions. After receiving the business Q&A request, the back-end server can parse the request to obtain the requester's request identifier and the business Q&A request text. It should be noted that the requesting party can also initiate a business Q&A request via voice input on the business function front-end page. If the parsed request information is audio information, it will be converted into text. The audio information processing may include, but is not limited to, filtering interjections and merging sentence segments to obtain the business Q&A request text corresponding to the audio information.
[0043] Based on the request identifier obtained from the parsed requester, the historical user behavior sequence of the requester is acquired. Specifically, this refers to the requester's operation behavior logs over a historical time period. For example, this could include user clicks on a specific business function page or transactions with a third-party platform. It should be noted that the information and data collected during the acquisition of historical user behavior sequences are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of this data all comply with relevant national and regional laws, regulations, and standards, and necessary confidentiality measures are taken. This process does not violate public order and good morals, and corresponding access points are provided for users to choose whether to authorize or refuse authorization.
[0044] Based on the business identifier of the current functional business, the business scenario information of the current functional business can be determined. The business identifier uniquely identifies a functional business, functional module, or business module within a specific financial scenario. For example, functional businesses may include a wealth management function zone, cash management, risk assessment, and portfolio details.
[0045] The business question-and-answer request text, historical user behavior sequences, and business scenario information are input into a pre-trained multimodal intent prediction network model to obtain intent information. The multimodal intent prediction network model can be an existing neural network model, such as a convolutional neural network model or a feedforward neural network model. Furthermore, to meet the intent prediction requirements of the financial scenario in this embodiment, the multimodal intent prediction network model can also be pre-built and trained by relevant technical personnel.
[0046] In one optional embodiment, the multimodal intent prediction network model includes a modality coding layer, an attention fusion layer, and a classification layer; the classification layer has three preset slots, corresponding to three prediction categories: subject object, analysis action, and target indicator; accordingly, the business question-and-answer request text, historical user behavior sequence, and business scenario information are input into the pre-trained multimodal intent prediction network model to obtain intent information, including:
[0047] Step a1: Input the business question and answer request text, historical user behavior sequence, and business scenario information into the modality coding layer in the multimodal intent prediction network model. The modality coding layer performs text feature extraction on the business question and answer request text to generate a text feature vector, performs sequence feature extraction on the historical user behavior sequence to generate a sequence feature vector, and performs one-hot encoding on the business scenario information to generate a scenario feature vector.
[0048] Step a2: Input the text feature vector, sequence feature vector and scene feature vector into the attention fusion layer in the multimodal intent prediction network model to perform feature fusion and obtain the global fused feature vector.
[0049] Step a3: Input the global fusion feature vector into the classification layer to obtain the predicted category and its corresponding category confidence for each slot. The classification layer then generates the intent vector based on each predicted category and its corresponding category confidence.
[0050] Step a4: Perform vector parsing on the intent vector to obtain the intent information.
[0051] The modality coding layer is used to encode features from multimodal input data, including text feature coding, sequence feature coding, and scene feature coding modules. The classification layer consists of a fully connected layer and a Softmax classifier. The Softmax classifier is a parallel multi-output classifier with three slots corresponding to three prediction categories: the subject object, the analysis action, and the target indicator. Each classifier outputs a category vector and its corresponding confidence score or weight value for each prediction category in parallel.
[0052] like Figure 1B The diagram illustrates the model structure of a multimodal intent prediction network. The business question-and-answer request text is input into the text feature encoding module of the modality encoding layer for text feature extraction, generating a text feature vector. The text feature encoding module can be a pre-trained language model with a full-word mask. To further improve the accuracy of text data processing in financial scenarios, relevant financial domain information can be used to fine-tune the pre-trained language model with the full-word mask, thereby improving the accuracy of text feature extraction from business question-and-answer request texts in financial scenarios.
[0053] Historical user behavior sequences are input into the sequence feature extraction module of the modality coding layer for sequence feature extraction, generating a sequence feature vector. The sequence feature encoding module can be a bidirectional long short-term memory (LSTM) network. The forward LSM network of the bidirectional LSM network captures the positive behavioral dependencies of the historical user behavior sequences, and the backward LSM network captures the negative behavioral dependencies of the historical user behavior sequences. These are then concatenated to obtain the sequence feature vector. One-hot encoding is performed on business scenario information to generate a scenario feature vector.
[0054] Text feature vectors, sequence feature vectors, and scene feature vectors are input into the attention fusion layer of the multimodal intent prediction network model for feature fusion, resulting in a global fused feature vector. Specifically, text feature vectors, sequence feature vectors, and scene feature vectors can be weighted and fused to generate a global fused feature vector.
[0055] Furthermore, to improve the accuracy of generating the global fusion feature vector, in an optional embodiment, the text feature vector, sequence feature vector, and scene feature vector are input into the attention fusion layer of the multimodal intent prediction network model for feature fusion to obtain the global fusion feature vector. This includes: in the attention fusion layer, projecting the text feature vector onto a preset first projection matrix to generate a first vector; concatenating the sequence feature vector and scene feature vector to generate a temporary concatenated vector; projecting the temporary concatenated vector onto a preset second projection matrix to generate a second vector; and projecting the temporary concatenated vector onto a preset third projection matrix to generate a third vector; splitting the first, second, and third vectors according to a preset number of multi-head attention heads to obtain several first, second, and third sub-vectors; generating at least one weight matrix based on each first, second, and third sub-vector; and generating the global fusion feature vector based on each weight matrix.
[0056] Specifically, the number of attention heads is configured in the attention fusion layer, for example, to 8 heads, and the projection dimension is configured, for example, to 512 dimensions. Linear projection transformation is performed in the attention fusion layer based on a preset first projection matrix. The text feature vector is then projected onto a Query vector, which is the first vector. For example, if the text feature vector is a 768-dimensional vector, the first projection matrix... If the dimension is set to 768×512, then the first vector obtained after transformation by the first projection matrix is a 1×512 dimensional vector.
[0057] The sequence feature vector and scene feature vector are concatenated to generate a temporary concatenated vector. For example, if the sequence feature vector is 512-dimensional and the scene feature vector is 256-dimensional, the temporary concatenated vector will be 1×768-dimensional. This is then processed through a second projection matrix. The temporary spliced vector is projected (768×512 dimensional) to become a Key vector (1×512), also known as the second vector. This is then transformed using a third projection matrix. (768×512 dimensions) The temporary spliced vector is vector-projected and converted into a Value vector (1×512), which is also the third vector.
[0058] Based on the preset number of multi-head attention heads, the first, second, and third vectors are split according to the number of heads. Assuming the head dimension is 8, then the dimension of each head is 512 / 8 = 64, resulting in several first, second, and third sub-vectors. The first sub-vectors are Q1~Q8 (1×64), the second sub-vectors are K1~K8 (1×64), and the third sub-vectors are V1~V8 (1×64).
[0059] The similarity of the i-th head is calculated based on the first sub-vector Qi and the second sub-vector Ki, and the weight matrix of the i-th head is obtained by Softmax normalization based on the similarity. The weighted feature vector of the i-th head is obtained based on the i-th head matrix and the third sub-vector Vi. The weighted feature vectors of the 8 heads are concatenated to obtain a temporary vector, which is then transformed using a preset output projection matrix to obtain the global fusion feature vector.
[0060] The above technical solution generates query vectors based on text features and key-value vectors based on behavioral sequence features and scene features during the multi-head attention feature fusion process. This enables heterogeneous modal data to form semantic associations, improves the accuracy of feature fusion in financial semantic scenarios, and makes the generated global fusion feature vectors more in line with the complex needs of financial decision-making, thereby further improving the accuracy of subsequent prediction of user intent information.
[0061] The globally fused feature vector is input into the classification layer to obtain the predicted category and its corresponding category confidence for each slot. The classification layer then generates an intent vector based on each predicted category and its corresponding category confidence. The number of fully connected layers in the classification layer can be preset by relevant technical personnel. The fully connected layers are used to perform feature processing on the globally fused feature vector, and the processed feature vector is then input into the Softmax classifier.
[0062] Specifically, a parallel multi-output classifier is employed, with a Softmax classifier head designed for each slot in the intent vector. Each slot corresponds to a predicted category: subject object, analysis action, and target metric. Furthermore, the prediction category result for each slot also includes the confidence level of the corresponding predicted category. For example, the slot category corresponding to the subject object could include general customers, private banking customers, and corporate customers; the slot category corresponding to the analysis action could include product recommendations, data queries, attribution analysis, and complaint handling; and the slot category corresponding to the target metric could include 3-5 year stable wealth management, short-term cash management, high-yield equity products, risk assessment, fund flow analysis, and churn risk warning.
[0063] The generated intent vector can be expressed as: {Subject Object (W1), Analysis Action (W2), Target Indicator (W3)}, where W1 is the confidence level of the predicted subject object; W2 is the confidence level of the predicted analysis action; and W3 is the confidence level of the predicted target indicator. For example, the intent vector could be: {Ordinary Customer (0.9), Query Data (0.88), Short-Term Cash Management (0.95)}. Here, ordinary customer is the subject object, query data is the analysis action, and short-term cash management is the target indicator; 0.9 is the confidence level that the subject object is predicted to be ordinary customer; 0.88 is the confidence level that the analysis action is predicted to be query data; and 0.95 is the confidence level that the target indicator is short-term cash management.
[0064] In specific model prediction scenarios, the intent vector output by the model is usually structured data, for example, intent vector V. intent The expression can be as follows:
[0065] V intent ={"subject":{"High-end customers in East China":0.9},"action":{"Attribution analysis": 0.92},"target": {"Churning rate":0.88}}.
[0066] Here, subject represents the main target, and the predicted result is high-end customers in East China, with a confidence level of 0.9; action represents the analysis action, and the predicted result is attribution analysis, with a confidence level of 0.92; target represents the target indicator, and the predicted result is churn rate, with a confidence level of 0.88.
[0067] By performing vector parsing on the intent vector, intent information can be obtained. Continuing the previous example, based on the preset keywords "subject," "action," and "target," vector parsing is performed on the intent vector to obtain the prediction results of the subject, the action being analyzed, and the target indicator, along with their corresponding confidence levels.
[0068] The aforementioned technical solution generates intent vectors by comprehensively integrating and analyzing multiple dimensions such as text information, historical behavior, and business scenarios through the construction of a multimodal intent prediction network model. This improves the accuracy of intent vector generation. The design of the intent vector slots incorporates decision-making subjects, actions, and goals, thereby deeply aligning with the financial business processing flow and making intent prediction more consistent with financial scenarios. The multi-head attention mechanism of the model integrates text vectors, behavior sequence vectors, and scenario encoding vectors, thereby capturing the implicit constraints of intent and eliminating the one-sidedness of the intent prediction process caused by single-dimensional data, further improving the accuracy of the final intent information generation.
[0069] Furthermore, this embodiment also provides a model training method for a multimodal intent prediction network model. First, training sample data is prepared. This training sample data can be financial business interaction samples, obtained by acquiring user interaction events over historical time periods, including historical user input text or speech-transcribed text, historical behavior sequences, and scene context labels, as well as corresponding structured intent label data. For example, the intent label corresponding to a financial business interaction sample might be {Subject: High-end customers in East China / 0.90 (weight), Analysis Action: Attribution Analysis / 0.88 (weight), Target Metric: Churn Rate / 0.95 (weight)}. The sample size can include at least 100,000 labeled samples, covering multiple core scenarios such as customer analysis, credit approval, and risk prediction.
[0070] The loss function design can include multi-task cross-entropy loss, which includes slot classification of the subject object, analysis action, and target metric, and can also include MSE (Mean Squared Error Loss). Furthermore, attention weight prediction loss can be fused to enhance the modality fusion effect.
[0071] Specifically, the multimodal intent prediction network model is initialized, including each network layer and the layer modules contained within each layer. Training sample data with intent labels is input into the initial multimodal intent prediction network model to obtain the model's predicted intent vector. Based on the fusion loss function, the current loss value for the current iteration time period is generated through forward propagation based on the model's predicted intent vector and the real intent vector corresponding to the intent label. Based on the current loss value, the multimodal intent prediction network model is iteratively trained, continuously optimizing the model parameters until a preset model training termination condition is met. For example, the current loss value reaches a set loss threshold, the current loss value tends to stabilize, or the current iteration count reaches a set iteration count threshold. This embodiment does not impose any restrictions on these conditions.
[0072] Based on the intent vector predicted by the multimodal intent prediction network model, intent information is determined, and a target knowledge graph is constructed based on the subject object, analytical action, and target indicator in the intent information. Specifically, information retrieval can be performed from a pre-built financial database based on the subject object, analytical action, and target indicator. This involves retrieving target events related to the subject object, analytical action, target indicator, and related target events, determining the relationships between them, and using these elements as nodes, the relationships as edge attributes of the nodes, and the target event context as node attribute information to construct the target knowledge graph.
[0073] Initial edge weights are set for the nodes of the target knowledge graph, and a graph neural network model is used to update the edge weights, resulting in the updated target knowledge graph. For example, financial product X (node 1) holds customer group A (node 2), with a weight of 0.95. The node attribute information of node 2 is that customer group A is a high-end customer group in region XXX. Here, "holding" is the edge attribute information between node 1 and node 2, with an edge weight of 0.95. Another example is the XXX product default event (node 3) affecting customer group A (node 2), with a weight of 0.90. Here, "affecting" is the edge attribute information between node 3 and node 2, with an edge weight of 0.90. Yet another example is complaint ticket C001 (node 4) potentially causing customer churn (node 5), with a weight of 0.77. Here, "potentially causing" is the edge attribute information between node 4 and node 5, with an edge weight of 0.77.
[0074] For example, based on a preset path search or path ranking algorithm, the system traverses each node in the target knowledge graph and its corresponding node edge weights, filtering the top-N weighted related paths to identify core influencing factors. For instance, the identified critical path information might be: the XXX product failure incident (node 3) affects (edge attribute information, node edge weight 0.90) customer group A (node 2), which may lead to (edge attribute information, node edge weight 0.77) customer churn (node 5), with a total weight of 0.835 (weighted sum of all node edge weights). A secondary information path could be complaint ticket C001 (node 4), which may lead to (edge attribute information, node edge weight 0.77) customer churn (node 5), with a total weight of 0.77. The path with the highest total weight is the critical information path, and K secondary information paths can be obtained by ranking them according to their total weight.
[0075] Based on the node edge weights, edge attribute information, and node attribute information of each node in the key information path and each secondary information path, and using a pre-set attribution analysis report template, the target decision result is generated and fed back to the requesting party. The attribution analysis report template can be pre-set by relevant technical personnel; for example, it can include core elements, secondary elements, and recommended measures.
[0076] For example, if the business query request is to analyze the reasons for the churn of high-end customers in region XXX, then core elements are generated based on the node information of the key information path, secondary elements are generated based on the node information of the secondary information path, and suggested measures are generated based on the node information of the entire path. The node information includes node edge weights, edge attribute information, and node attribute information. The generated target decision result of "analyzing the reasons for the churn of high-end customers in region XXX" can be expressed in the following form:
[0077] I. Core Reasons (Weight: 99%) Event: YYY product default incident; Related Evidence: 1. News and Public Opinion: "Announcement of XXX Product Default by a Certain Bank" (May 1, 2023); 2. Real-time Data: Redemption volume of YYY product by high-end customers in XXX region increased by 300% year-on-year; 3. Data Warehouse Data: The churn rate of high-end customers holding YYY product reached 15.6%. II. Secondary Reasons (Weight: 72%) Event: Concentrated service quality complaints; Related Evidence: 1. Complaint Ticket: No. CT2023 (Customer C1001 complained "Service response delay exceeded 48 hours"); 2. Satisfaction Data: Customer satisfaction with high-end customers in East China decreased by 40% year-on-year. III. Recommended Measures: 1. Emergency Response: Initiate special communication with customers holding YYY products and provide alternative repayment options (such as debt-to-equity swap or installment repayment); 2. Service Optimization: Open a green channel for complaints from high-end customers and reduce response time to within 2 hours; 3. Risk Warning: Push risk warnings to customers holding similar wealth management products and provide asset reallocation suggestions.
[0078] The technical solution of this invention employs a multimodal intent prediction model to analyze and predict intent information from multiple dimensions, including question-and-answer text, user behavior sequences, and business scenarios, in business question-and-answer requests. During the intent prediction process, it bridges data silos by fusing heterogeneous data, accurately integrates the heterogeneous data, and performs intent prediction based on the characteristics of this accurate integration. This achieves high accuracy in understanding complex financial business intents in financial question-and-answer scenarios. By constructing a target knowledge graph in real time based on intent information, including the subject, analysis actions, and target indicators, and dynamically adjusting the node edge weights of the target knowledge graph, it achieves dynamic construction of target knowledge graphs for different question-and-answer scenarios and improves the accuracy of constructing the real-time generated target knowledge graph. Based on this dynamic knowledge graph, path retrieval is performed to obtain the final target decision result, improving the efficiency and accuracy of intent analysis and decision-making in financial question-and-answer scenarios.
[0079] Example 2
[0080] Figure 2 This is a flowchart of a business question-and-answer intent analysis and decision-making method provided in Embodiment 2 of the present invention. This embodiment is an optimization and improvement based on the above technical solutions.
[0081] Furthermore, the step "constructing a target knowledge graph based on the subject object, analysis action, and target indicator" is refined to "determining at least one structured entity node and one unstructured entity node based on the subject object, analysis action, and target indicator, establishing node relationships between each structured entity node and each unstructured entity node, and generating an initial knowledge graph; updating the initial knowledge graph based on the subject object and target indicator to obtain the target knowledge graph." This improves the method for constructing the target knowledge graph.
[0082] It should be noted that for parts not described in detail in the embodiments of the present invention, please refer to the descriptions in other embodiments. For example... Figure 2 As shown, the method includes the following specific steps:
[0083] S210. In response to the requester's business Q&A request for the current functional business, parse the business Q&A request, determine the business Q&A request text, obtain the requester's historical user behavior sequence, and determine the business scenario information of the current functional business.
[0084] S220. Input the business question and answer request text, historical user behavior sequence and business scenario information into the pre-trained multimodal intent prediction network model to obtain intent information; the intent information includes the subject object, analysis action and target indicator.
[0085] S230. Based on the main object, analysis actions, and target indicators, determine at least one structured entity node and one unstructured entity node, establish the node association relationship between each structured entity node and each unstructured entity node, and generate an initial knowledge graph.
[0086] S240. Based on the subject object and the target index, update the initial knowledge graph to obtain the target knowledge graph; the target knowledge graph includes at least one node and the node edge weights, edge attribute information and node attribute information of each node.
[0087] S250. Based on each node in the target knowledge graph and its corresponding node edge weights, determine the key information path and at least one secondary information path.
[0088] S260. Based on the node edge weights, edge attribute information, and node attribute information of each node in the key information path and each secondary information path, generate the target decision result and feed it back to the requester.
[0089] For example, information can be retrieved from a pre-built financial database based on the subject object, analytical action, and target indicator to determine object events related to the subject object, analytical action, and target indicator. These object events may be structured attribute data or unstructured attribute data. Therefore, based on the attribute parameters corresponding to the subject object, analytical action, target indicator, and object event, at least one structured entity node and one unstructured entity node are generated, and the node association relationship between each structured entity node and each unstructured entity node is established to generate an initial knowledge graph.
[0090] Understandably, to ensure the accuracy of constructing unstructured entity nodes and structured entity nodes, thereby ensuring the accuracy of constructing the initial knowledge graph, in one optional embodiment, at least one structured entity node and one unstructured entity node are determined based on the subject object, analysis action, and target indicator, and node association relationships are established between each structured entity node and each unstructured entity node to generate an initial knowledge graph, including:
[0091] Step b11: Generate database query statements based on the main object and target indicators, and obtain structured entity data from the pre-built data warehouse based on the database query statements.
[0092] Step b12: Generate Boolean retrieval instructions based on the subject object, analysis action, and target indicators, and obtain unstructured entity data from the pre-built distributed search analysis engine based on the Boolean retrieval instructions.
[0093] Step b2: Based on the structured entity data, generate at least one structured entity node and its corresponding node attribute information; and based on the unstructured entity data, generate at least one unstructured entity node and its corresponding node attribute information.
[0094] Step b3: Input each structured entity node and its corresponding node attribute information into the left tower module of the pre-selected dual-tower model to obtain the structured node vector output by the left tower module; and input each unstructured entity node and its corresponding node attribute information into the right tower module of the dual-tower model to obtain the unstructured node vector output by the right tower module.
[0095] Step b4: Determine the entity similarity between each structured entity node and each unstructured entity node based on the structured node vector of each structured entity node and the unstructured node vector of each unstructured entity node.
[0096] Step b5: Based on entity similarity, establish node association relationships between each structured entity node and each unstructured entity node to generate an initial knowledge graph.
[0097] Generate a database query statement based on the subject and target metric. For example, if the subject is "high-end customers in East China" and the target metric is "customer churn rate", the generated database query statement could be "SELECT client_id, region, segment, churn_rate, churn_reason_code FROMClient WHERE region='EastChina' AND segment='HighNet' AND churn_rate>0". Here, "client_id" represents the unique identifier of the customer, "region" represents the geographical region of the customer, "segment" represents the market segment of the customer, "churn_rate" represents the customer churn rate, and "churn_reason_code" represents the reason code for customer churn, used to assist in analyzing the reasons for customer churn; "region='EastChina'" is subject 1, representing the East China region; "segment='HighNet'" is subject 2, representing high-end customers; and "churn_rate>0" is the target metric, representing the customer churn rate.
[0098] Based on the database query statements generated above, structured entity data is retrieved from a pre-built data warehouse. This data warehouse can be pre-built by relevant technical personnel and stores structured data stored over historical time periods. The structured data returned by the data warehouse based on the aforementioned database query statements is shown in Table 1.
[0099] Table 1
[0100]
[0101] Based on the subject, analysis action, and target metric, Boolean search instructions are generated. These instructions can be Boolean search expressions. For example, if the subject is "high-end customers in East China," the analysis action is "analyze the reasons," and the target metric is "customer churn rate," then the generated Boolean search expressions could be: {"query": {"bool": {"must": [ {"match": {"content": "high-end customers in East China"}}, {"match": {"content": "reasons for churn"}}, {"match": {"doc_type": "complaint tickets|research reports"}}]}}}. Here, "query" represents the query operation instruction, "bool" represents a Boolean query used to combine multiple query conditions, "must" represents a condition that must be matched (documents must meet the conditions under "must" to be returned), and "match" represents a text matching query used to match text after analyzing the field. "content" represents the main content field of the document, consisting of the primary document content (the main target audience, high-end customers in East China) and the secondary document content (analyzed actions, reasons for analysis, and target metrics, customer churn rate). "doc_type" represents the document category or type, used to distinguish and manage different types of documents. In the Boolean search expression above, the document categories are complaint tickets and research reports.
[0102] Based on the Boolean retrieval command, unstructured entity data is retrieved from a pre-built distributed search and analysis engine, which stores relevant unstructured data. Continuing the previous example, the unstructured entity data returned by the distributed search and analysis engine could be: "Complaint ticket: 'High-end customer XXX complains about XX branch due to financial product default'; Industry research report: 'Attribution of High-end Customer Churn in East China: XX's declining revenue is the main reason'; Internal report: 'Analysis of High-end Customer Churn in East China: Service response delay accounts for X%'."
[0103] Based on structured entity data, generate at least one structured entity node and its corresponding node attribute information; based on unstructured entity data, generate at least one unstructured entity node and its corresponding node attribute information. Continuing the previous example, the structured entity node could be "C001", and its corresponding node attribute information could be "["EastChina", "HighNet"]". For unstructured data, a named entity recognition model can be used for entity extraction. For example, continuing the previous example, the unstructured entity nodes extracted from the aforementioned unstructured data could include "C001" and "financial product default", etc., and the node attribute information corresponding to the unstructured entity node could be the entity content context information.
[0104] Each structured entity node and its corresponding node attribute information are input into the left tower module of a pre-selected dual-tower model to obtain a structured node vector output by the left tower module. Similarly, each unstructured entity node and its corresponding node attribute information are input into the right tower module of the dual-tower model to obtain an unstructured node vector output by the right tower module. The left tower of the dual-tower model processes structured information and generates structured vector data, while the right tower processes unstructured information and generates unstructured vector data.
[0105] Based on the structured node vector of each structured entity node and the unstructured node vectors of each unstructured entity node The entity similarity (sim) between each structured entity node and each unstructured entity node is determined. The specific determination method can be as follows:
[0106]
[0107] Based on entity similarity, node associations are established between each structured entity node and each unstructured entity node to generate an initial knowledge graph. If the entity similarity between two structured entity nodes and an unstructured entity node is greater than a preset similarity threshold, then node edge relationships are established between the two structured entity nodes and the unstructured entity nodes, thereby associating structured entity nodes with unstructured entity nodes, and thus obtaining an initial knowledge graph containing several unstructured entity nodes and structured entity nodes with node edge connections.
[0108] The above technical solution combines the subject object, analysis action, and target indicator in the intent information to obtain structured and unstructured data from different data sources. It uses a dual-tower model to vectorize unstructured and structured data and constructs node edge connections based on node similarity. This effectively establishes a precise association between unstructured and structured nodes, enabling the interoperability of static data attributes and dynamic behavioral semantics. The implicit information in unstructured data can also be called effective features for reasoning and added to the initial knowledge graph, rather than isolated text data. This achieves accurate construction of the initial knowledge graph, which facilitates subsequent link tracing, accurate positioning, and decision-making.
[0109] Based on the subject object and target metrics, the initial knowledge graph is updated to obtain the target knowledge graph. For example, external resources can be acquired in real time based on the subject object and target metrics, and resource event nodes can be constructed based on the real-time acquired external resources. The initial knowledge graph can then be optimized using the resource event nodes constructed in real time to obtain the target knowledge graph.
[0110] In one optional embodiment, the initial knowledge graph is updated based on the subject object and the target metric to obtain the target knowledge graph. This includes: generating a topic subscription request based on the subject object and the target metric, subscribing to real-time events from the distributed streaming platform in real time based on the topic subscription request, and determining the event characteristics of the real-time events; determining the node edge weights, edge attribute information, and node attribute information of each node in the initial knowledge graph; updating the initial knowledge graph based on the real-time events and their corresponding event characteristics to obtain a temporary knowledge graph; and updating the node edge weights of each node in the temporary knowledge graph to obtain the target knowledge graph.
[0111] Based on the target audience and target metrics, subscription topics are generated, and subscription requests are then generated accordingly. For example, if the target audience is "high-end customers in East China" and the target metric is "customer churn rate," the generated subscription topics could include high-end customer behavior streams (Topic 1), East China financial event streams (Topic 2), and streaming risk warning streams (Topic 3). Subscription requests are then generated based on these topics, and the distributed streaming platform is used to subscribe in real-time to events related to the subscribed topics, along with the event characteristics of these events. Event characteristics can include the event's occurrence time and other event-related features.
[0112] Real-time events and their corresponding event features are added to the initial knowledge graph. Specifically, this involves determining the node correlation or node similarity between the real-time event and each node in the initial knowledge graph, and establishing edge connections between the real-time event and each node in the initial knowledge graph based on the node correlation or node similarity, thus obtaining a temporary knowledge graph. It should be noted that the edge weights corresponding to each node in the temporary knowledge graph can be preset initial weight values, which can be pre-set by relevant technical personnel. The node attribute information of the real-time event in the temporary knowledge graph can be the event features of the real-time event, and the edge attribute information can be the node relationships between the real-time event and its connected nodes.
[0113] The edge weights of each node in the temporary knowledge graph are updated to obtain the target knowledge graph. Specifically, a graph neural network model can be used to update the edge weights of each node in the temporary knowledge graph. For example, the algorithm begins to propagate information in the target knowledge graph. A node sends information to its neighboring nodes, including features such as the severity and scope of the event corresponding to the node. Other nodes receive information from multiple neighboring nodes. The model calculates the latest state of each node and recalculates the edge weights connected to that node.
[0114] The above technical solution generates topic subscription requests based on the subject object and target indicators, subscribes to events in real time from a distributed streaming platform, and continuously updates the initial knowledge graph based on the subscribed real-time events. This achieves dynamic updates to the initial knowledge graph, improves the richness and comprehensiveness of the graph content, and updates the node weight parameters of the graph in real time to obtain the final target knowledge graph. This improves the accuracy of the constructed target knowledge graph, thereby further improving the accuracy of subsequent path retrieval and decision-making based on the target knowledge graph.
[0115] This embodiment's technical solution determines at least one structured entity node and one unstructured entity node based on the subject object, analysis action, and target indicator. It then establishes node relationships between these two types of nodes to generate an initial knowledge graph. Based on the subject object and target indicator, the initial knowledge graph is updated to obtain the target knowledge graph. This technical solution achieves cross-modal entity alignment by accurately constructing structured and unstructured entity nodes and establishing relationships between them. This solves the problem of traditional knowledge graphs having only one modality, thereby improving the accuracy of the target knowledge graph construction and the richness and comprehensiveness of its content. Ultimately, this improves the accuracy of subsequent key information and intent analysis decision generation.
[0116] Example 3
[0117] Figure 3 This is a schematic diagram of the process structure of a business question-and-answer intent analysis and decision-making method provided in Embodiment 3 of the present invention. Based on the above embodiments, this embodiment provides a preferred example.
[0118] S1. In response to the requester's business Q&A request for the current function, parse the business Q&A request, determine the business Q&A request text, obtain the requester's historical user behavior sequence, and determine the business scenario information of the current function.
[0119] S2. Input the business question and answer request text, historical user behavior sequence and business scenario information into the pre-trained multimodal intent prediction network model to obtain intent information; the intent information includes the subject object, analysis action and target indicator.
[0120] The multimodal intent prediction network model includes a modality coding layer, an attention fusion layer, and a classification layer. The classification layer has three pre-defined slots, corresponding to three prediction categories: the subject object, the analysis action, and the target metric. For example, the business question-and-answer request text, historical user behavior sequences, and business scenario information are input into the modality coding layer of the multimodal intent prediction network model. The modality coding layer extracts text features from the business question-and-answer request text to generate a text feature vector, extracts sequence features from the historical user behavior sequences to generate a sequence feature vector, and performs one-hot encoding on the business scenario information to generate a scenario feature vector.
[0121] In the attention fusion layer, the text feature vector is projected onto a preset first projection matrix to generate a first vector, and the sequence feature vector and scene feature vector are concatenated to generate a temporary concatenated vector. The temporary concatenated vector is projected onto a preset second projection matrix to generate a second vector, and the temporary concatenated vector is projected onto a preset third projection matrix to generate a third vector. According to the preset number of multi-head attention heads, the first, second, and third vectors are split into several first, second, and third sub-vectors. At least one weight matrix is generated based on each first, second, and third sub-vector, and a global fusion feature vector is generated based on each weight matrix.
[0122] The global fusion feature vector is input into the classification layer to obtain the predicted category and its corresponding category confidence for each slot. The classification layer then generates an intent vector based on each predicted category and its corresponding category confidence. The intent vector is then parsed to obtain the intent information.
[0123] S31. Generate database query statements based on the main object and target indicators, and obtain structured entity data from the pre-built data warehouse based on the database query statements.
[0124] S32. Generate Boolean retrieval instructions based on the subject object, analysis action, and target indicators, and obtain unstructured entity data from the pre-built distributed search analysis engine based on the Boolean retrieval instructions.
[0125] S4. Based on the structured entity data, generate at least one structured entity node and its corresponding node attribute information; and based on the unstructured entity data, generate at least one unstructured entity node and its corresponding node attribute information.
[0126] S5. Input each structured entity node and its corresponding node attribute information into the left tower module of the pre-selected dual-tower model to obtain the structured node vector output by the left tower module. Input each unstructured entity node and its corresponding node attribute information into the right tower module of the dual-tower model to obtain the unstructured node vector output by the right tower module.
[0127] S6. Based on the structured node vectors of each structured entity node and the unstructured node vectors of each unstructured entity node, determine the entity similarity between each structured entity node and each unstructured entity node.
[0128] S7. Based on entity similarity, establish the node association relationships between each structured entity node and each unstructured entity node to generate an initial knowledge graph.
[0129] S8. Based on the subject object and target indicators, generate topic subscription requests, subscribe to real-time events from the distributed streaming platform in real time based on topic subscription requests, and determine the event characteristics of real-time events.
[0130] S9. Determine the node edge weights, edge attribute information, and node attribute information of each node in the initial knowledge graph, and update the initial knowledge graph according to real-time events and their corresponding event characteristics to obtain a temporary knowledge graph.
[0131] S10. Update the node edge weights of each node in the temporary knowledge graph to obtain the target knowledge graph.
[0132] S11. Based on each node in the target knowledge graph and its corresponding node edge weights, determine the key information path and at least one secondary information path.
[0133] S12. Based on the node edge weights, edge attribute information, and node attribute information of each node in the key information path and each secondary information path, generate the target decision result and feed it back to the requester.
[0134] Example 4
[0135] Figure 4 This is a schematic diagram of the structure of a business question-and-answer intent analysis and decision-making device provided in Embodiment 4 of the present invention. The business question-and-answer intent analysis and decision-making device provided in this embodiment of the present invention is applicable to situations where intelligent intent analysis and accurate decision-making are performed on user question-and-answer requests in financial scenarios. This business question-and-answer intent analysis and decision-making device can be implemented in hardware and / or software, such as... Figure 4 As shown, the device includes: a question-and-answer request response module 401, an intent information generation module 402, a target map construction module 403, an information path generation module 404, and an analysis and decision-making module 405. Among them,
[0136] The question and answer request response module 401 is used to respond to the requester's business question and answer request for the current functional business, parse the business question and answer request, determine the business question and answer request text, obtain the requester's historical user behavior sequence, and determine the business scenario information of the current functional business;
[0137] The intent information generation module 402 is used to input the business question-and-answer request text, the historical user behavior sequence, and the business scenario information into a pre-trained multimodal intent prediction network model to obtain intent information; the intent information includes the subject object, the analysis action, and the target indicator;
[0138] The target knowledge graph construction module 403 is used to construct a target knowledge graph based on the subject object, analysis action and target index; the target knowledge graph includes at least one node and the node edge weight, edge attribute information and node attribute information of each node;
[0139] The information path generation module 404 is used to determine the key information path and at least one secondary information path based on each node in the target knowledge graph and its corresponding node edge weights.
[0140] The analysis and decision module 405 is used to generate a target decision result based on the node edge weights, edge attribute information and node attribute information of each node in the key information path and each of the secondary information paths, and to feed back the target decision result to the requester.
[0141] The technical solution of this invention employs a multimodal intent prediction model to analyze and predict intent information from multiple dimensions, including question-and-answer text, user behavior sequences, and business scenarios, in business question-and-answer requests. During the intent prediction process, it bridges data silos by fusing heterogeneous data, accurately integrates the heterogeneous data, and performs intent prediction based on the characteristics of this accurate integration. This achieves high accuracy in understanding complex financial business intents in financial question-and-answer scenarios. By constructing a target knowledge graph in real time based on intent information, including the subject, analysis actions, and target indicators, and dynamically adjusting the node edge weights of the target knowledge graph, it achieves dynamic construction of target knowledge graphs for different question-and-answer scenarios and improves the accuracy of constructing the real-time generated target knowledge graph. Based on this dynamic knowledge graph, path retrieval is performed to obtain the final target decision result, improving the efficiency and accuracy of intent analysis and decision-making in financial question-and-answer scenarios.
[0142] Optionally, the multimodal intent prediction network model includes a modality coding layer, an attention fusion layer, and a classification layer; the classification layer has three preset slots, corresponding to three prediction categories: subject object, analysis action, and target indicator; correspondingly, the intent information generation module 402 includes:
[0143] The feature extraction unit is used to input the business question-and-answer request text, the historical user behavior sequence, and the business scenario information into the modality coding layer in the multimodal intent prediction network model. The modality coding layer performs text feature extraction on the business question-and-answer request text to generate a text feature vector, performs sequence feature extraction on the historical user behavior sequence to generate a sequence feature vector, and performs one-hot encoding on the business scenario information to generate a scenario feature vector.
[0144] The feature fusion unit is used to input the text feature vector, the sequence feature vector, and the scene feature vector into the attention fusion layer in the multimodal intent prediction network model to perform feature fusion and obtain a global fused feature vector.
[0145] The classification unit is used to input the global fusion feature vector into the classification layer to obtain the predicted category and its corresponding category confidence for each slot, and the classification layer generates an intent vector based on each predicted category and its corresponding category confidence.
[0146] The vector parsing unit is used to perform vector parsing on the intent vector to obtain intent information.
[0147] Optional, feature fusion unit, specifically used for:
[0148] In the attention fusion layer, the text feature vector is vector-projected based on a preset first projection matrix to generate a first vector, and the sequence feature vector and the scene feature vector are vector-concatenated to generate a temporary concatenated vector.
[0149] The temporary spliced vector is vector-projected based on a preset second projection matrix to generate a second vector, and the temporary spliced vector is vector-projected based on a preset third projection matrix to generate a third vector.
[0150] Based on the preset number of multi-head attention heads, the first vector, the second vector, and the third vector are split into several sub-vectors: the first sub-vector, the second sub-vector, and the third sub-vector.
[0151] At least one weight matrix is generated based on each of the first sub-vector, the second sub-vector, and the third sub-vector, and a global fusion feature vector is generated based on each of the weight matrices.
[0152] Optionally, the target map construction module 403 includes:
[0153] The initial knowledge graph construction unit is used to determine at least one structured entity node and one unstructured entity node based on the main object, analysis action and target index, and to establish the node association relationship between each structured entity node and each unstructured entity node to generate an initial knowledge graph.
[0154] The target knowledge graph construction unit is used to update the initial knowledge graph based on the subject object and the target index to obtain the target knowledge graph.
[0155] Optional, initial map building unit, specifically used for:
[0156] Based on the main object and the target metric, a database query statement is generated, and structured entity data is retrieved from a pre-built data warehouse based on the database query statement; and,
[0157] Based on the subject object, the analysis action, and the target indicator, a Boolean retrieval instruction is generated, and unstructured entity data is obtained from a pre-built distributed search analysis engine based on the Boolean retrieval instruction;
[0158] Based on the structured entity data, at least one structured entity node and its corresponding node attribute information are generated; and based on the unstructured entity data, at least one unstructured entity node and its corresponding node attribute information are generated.
[0159] Each structured entity node and its corresponding node attribute information is input into the left tower module of the pre-selected dual-tower model to obtain the structured node vector output by the left tower module. Similarly, each unstructured entity node and its corresponding node attribute information is input into the right tower module of the dual-tower model to obtain the unstructured node vector output by the right tower module.
[0160] Based on the structured node vectors of each structured entity node and the unstructured node vectors of each unstructured entity node, the entity similarity between each structured entity node and each unstructured entity node is determined.
[0161] Based on the entity similarity, the node association relationships between each structured entity node and each unstructured entity node are established, and an initial knowledge graph is generated.
[0162] Optional, target map construction unit, specifically used for:
[0163] Based on the subject object and the target metric, a topic subscription request is generated, and real-time events are subscribed to from the distributed streaming platform in real time based on the topic subscription request, and the event characteristics of the real-time events are determined.
[0164] Determine the node edge weights, edge attribute information, and node attribute information of each node in the initial knowledge graph;
[0165] The initial knowledge graph is updated based on real-time events and their corresponding event characteristics to obtain a temporary knowledge graph.
[0166] The node edge weights of each node in the temporary knowledge graph are updated to obtain the target knowledge graph.
[0167] The business question-and-answer intent analysis and decision-making device provided in the embodiments of the present invention can execute the business question-and-answer intent analysis and decision-making method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0168] Example 5
[0169] Figure 5 A schematic diagram of an electronic device 50 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0170] like Figure 5 As shown, the electronic device 50 includes at least one processor 51 and a memory, such as a read-only memory (ROM) 52 and a random access memory (RAM) 53, communicatively connected to the at least one processor 51. The memory stores computer programs executable by the at least one processor. The processor 51 can perform various appropriate actions and processes based on the computer program stored in the ROM 52 or loaded from storage unit 58 into the RAM 53. The RAM 53 can also store various programs and data required for the operation of the electronic device 50. The processor 51, ROM 52, and RAM 53 are interconnected via a bus 54. An input / output (I / O) interface 55 is also connected to the bus 54.
[0171] Multiple components in electronic device 50 are connected to I / O interface 55, including: input unit 56, such as keyboard, mouse, etc.; output unit 57, such as various types of monitors, speakers, etc.; storage unit 58, such as disk, optical disk, etc.; and communication unit 59, such as network card, modem, wireless transceiver, etc. Communication unit 59 allows electronic device 50 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0172] Processor 51 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 51 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 51 performs the various methods and processes described above, such as intent analysis decision-making methods for business question answering.
[0173] In some embodiments, the intent analysis and decision-making method for business question answering may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 58. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 50 via ROM 52 and / or communication unit 59. When the computer program is loaded into RAM 53 and executed by processor 51, one or more steps of the intent analysis and decision-making method for business question answering described above may be performed. Alternatively, in other embodiments, processor 51 may be configured to perform the intent analysis and decision-making method for business question answering by any other suitable means (e.g., by means of firmware).
[0174] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0175] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0176] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0177] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0178] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0179] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0180] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0181] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A business question-and-answer intent analysis and decision-making method, characterized in that, include: In response to a request from a party regarding a business question and answer request for the current functional business, the request is parsed to determine the business question and answer request text, and the historical user behavior sequence of the requester is obtained, as well as the business scenario information of the current functional business is determined. The business question-and-answer request text, the historical user behavior sequence, and the business scenario information are input into a pre-trained multimodal intent prediction network model to obtain intent information; the intent information includes the subject object, the analysis action, and the target metric. Based on the subject object, analysis action, and target indicator, a target knowledge graph is constructed; the target knowledge graph includes at least one node and the node edge weight, edge attribute information, and node attribute information of each node; Based on each node in the target knowledge graph and its corresponding node edge weights, determine the key information path and at least one secondary information path; Based on the node edge weights, edge attribute information, and node attribute information of each node in the key information path and each of the secondary information paths, a target decision result is generated, and the target decision result is fed back to the requester.
2. The method according to claim 1, characterized in that, The multimodal intent prediction network model includes a modality encoding layer, an attention fusion layer, and a classification layer; the classification layer has three preset slots, corresponding to three prediction categories: subject object, analysis action, and target indicator; correspondingly, the business question-and-answer request text, the historical user behavior sequence, and the business scenario information are input into the pre-trained multimodal intent prediction network model to obtain intent information, including: The business question and answer request text, the historical user behavior sequence, and the business scenario information are input into the modality coding layer in the multimodal intent prediction network model. The modality coding layer performs text feature extraction on the business question and answer request text to generate a text feature vector, performs sequence feature extraction on the historical user behavior sequence to generate a sequence feature vector, and performs one-hot encoding on the business scenario information to generate a scenario feature vector. The text feature vector, the sequence feature vector, and the scene feature vector are input into the attention fusion layer in the multimodal intent prediction network model to perform feature fusion and obtain a global fused feature vector. The global fusion feature vector is input into the classification layer to obtain the predicted category and its corresponding category confidence for each slot. The classification layer then generates an intent vector based on each predicted category and its corresponding category confidence. The intent vector is parsed to obtain the intent information.
3. The method according to claim 2, characterized in that, The step of inputting the text feature vector, the sequence feature vector, and the scene feature vector into the attention fusion layer of the multimodal intent prediction network model for feature fusion to obtain a global fused feature vector includes: In the attention fusion layer, the text feature vector is vector-projected based on a preset first projection matrix to generate a first vector, and the sequence feature vector and the scene feature vector are vector-concatenated to generate a temporary concatenated vector. The temporary spliced vector is vector-projected based on a preset second projection matrix to generate a second vector, and the temporary spliced vector is vector-projected based on a preset third projection matrix to generate a third vector. Based on the preset number of multi-head attention heads, the first vector, the second vector, and the third vector are split into several sub-vectors: the first sub-vector, the second sub-vector, and the third sub-vector. At least one weight matrix is generated based on each of the first sub-vector, the second sub-vector, and the third sub-vector, and a global fusion feature vector is generated based on each of the weight matrices.
4. The method according to claim 1, characterized in that, The step of constructing a target knowledge graph based on the subject object, analysis action, and target indicators includes: Based on the subject object, analysis action, and target indicator, at least one structured entity node and unstructured entity node are determined, and the node association relationship between each structured entity node and each unstructured entity node is established to generate an initial knowledge graph. Based on the subject object and the target indicator, the initial knowledge graph is updated to obtain the target knowledge graph.
5. The method according to claim 4, characterized in that, The step involves determining at least one structured entity node and one unstructured entity node based on the subject object, analysis action, and target indicator, establishing node association relationships between each structured entity node and each unstructured entity node, and generating an initial knowledge graph, including: Based on the main object and the target metric, a database query statement is generated, and structured entity data is retrieved from a pre-built data warehouse based on the database query statement; and, Based on the subject object, the analysis action, and the target indicator, a Boolean retrieval instruction is generated, and unstructured entity data is obtained from a pre-built distributed search analysis engine based on the Boolean retrieval instruction; Based on the structured entity data, at least one structured entity node and its corresponding node attribute information are generated; and based on the unstructured entity data, at least one unstructured entity node and its corresponding node attribute information are generated. Each structured entity node and its corresponding node attribute information is input into the left tower module of the pre-selected dual-tower model to obtain the structured node vector output by the left tower module. Similarly, each unstructured entity node and its corresponding node attribute information is input into the right tower module of the dual-tower model to obtain the unstructured node vector output by the right tower module. Based on the structured node vectors of each structured entity node and the unstructured node vectors of each unstructured entity node, the entity similarity between each structured entity node and each unstructured entity node is determined. Based on the entity similarity, the node association relationships between each structured entity node and each unstructured entity node are established, and an initial knowledge graph is generated.
6. The method according to claim 4, characterized in that, The step of updating the initial knowledge graph based on the subject object and the target metric to obtain the target knowledge graph includes: Based on the subject object and the target metric, a topic subscription request is generated, and real-time events are subscribed to from the distributed streaming platform in real time based on the topic subscription request, and the event characteristics of the real-time events are determined. Determine the node edge weights, edge attribute information, and node attribute information of each node in the initial knowledge graph; The initial knowledge graph is updated based on real-time events and their corresponding event characteristics to obtain a temporary knowledge graph. The node edge weights of each node in the temporary knowledge graph are updated to obtain the target knowledge graph.
7. A business question-and-answer intent analysis and decision-making device, characterized in that, include: The question and answer request response module is used to respond to the requester's business question and answer request for the current functional business, parse the business question and answer request, determine the business question and answer request text, obtain the requester's historical user behavior sequence, and determine the business scenario information of the current functional business; The intent information generation module is used to input the business question-and-answer request text, the historical user behavior sequence, and the business scenario information into a pre-trained multimodal intent prediction network model to obtain intent information; the intent information includes the subject object, the analysis action, and the target indicator; The target knowledge graph construction module is used to construct a target knowledge graph based on the subject object, analysis action, and target index; the target knowledge graph includes at least one node and the node edge weight, edge attribute information, and node attribute information of each node; The information path generation module is used to determine the key information path and at least one secondary information path based on each node in the target knowledge graph and its corresponding node edge weights. The analysis and decision-making module is used to generate a target decision result based on the node edge weights, edge attribute information, and node attribute information of each node in the key information path and each of the secondary information paths, and to feed back the target decision result to the requesting party.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the intent analysis and decision-making method for business question answering as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the intent analysis and decision-making method for business question answering as described in any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the intent analysis and decision-making method for business question answering according to any one of claims 1-6.