Notary Intelligent Q&A Customer Service Method and System Based on Knowledge Graph

By constructing and dynamically updating the knowledge graph, combining multimodal feature fusion and deep semantic understanding, the problem of notarized intelligent customer service system understanding user intentions and handling complex queries is solved, and efficient and personalized notarization consulting services are achieved.

CN119938816BActive Publication Date: 2025-07-25SUZHOU LIANZHENG INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411751063.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-07-25
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

The existing notarized intelligent customer service system is difficult to accurately understand user intentions when facing complex or vague user problems, lacks context understanding and multi-round dialogue capabilities, and is difficult to update and maintain knowledge bases, and cannot flexibly respond to diversified query needs.

Method used

The knowledge graph is used to construct the initial knowledge graph through naming entity recognition and semantic role annotation, dynamic updates are performed in combination with the time attenuation mechanism, multi-modal feature fusion and deep semantic understanding, and dynamic knowledge graphs are used for multi-hop reasoning to generate personalized answers.

Benefits of technology

It improves the accuracy and depth of the notarized intelligent question-and-answer system, can handle complex queries, provide more comprehensive and relevant information, ensures the timeliness and personalized adaptability of the knowledge graph, and enhances users' understanding and trust in the reasoning process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938816B_ABST
    Figure CN119938816B_ABST
Patent Text Reader

Abstract

The present invention provides a notarization intelligent question-answering customer service method and system based on a knowledge graph, which relates to the field of information technology. It includes collecting multi-source heterogeneous data in the notarization field to form an entity set and a relationship set, combining them into an initial knowledge graph, supplementing and forming a dynamic knowledge graph based on a time-decay dynamic update mechanism, performing vectorization representation to obtain graph vectors; receiving multi-modal questions, determining fusion features, performing deep semantic understanding to obtain semantic features, determining user intentions, and forming structured query expressions; based on user intentions and structured query expressions, performing multi-hop reasoning in the dynamic knowledge graph, calculating similarity using graph vectors, and combining user intentions to determine the reasoning path, performing knowledge association reasoning along the reasoning path to obtain a preliminary reasoning result, performing interpretability analysis on the preliminary reasoning result, generating an inference chain and a confidence score to form a detailed reasoning result, and finally generating a personalized answer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular, to a notarization intelligent question - answering customer service method and system based on a knowledge graph. Background Art

[0002] With the continuous advancement of construction and the rapid development of social economy, the demand for notarization services shows a continuous growth trend. Notarization services cover multiple fields such as civil, commercial, and foreign - related, and are characterized by strong professionalism and high policy nature. Traditional notarization consultation services mainly rely on manual reception and telephone consultation. When facing the increasing demand for notarization, this mode exposes problems such as insufficient human resources, limited service time, and lagging knowledge update. To improve service efficiency and quality, some notarization institutions have begun to try to introduce intelligent customer service systems to meet the public's demand for all - weather, high - efficiency, and professional notarization consultation services.

[0003] Currently, the intelligent customer service systems applied in the notarization field are mainly based on keyword matching and preset question - answer pairs. Although the service efficiency has been improved to a certain extent, there are still many limitations. First, the system's ability to understand user questions is limited, and it is difficult to accurately grasp the true intention behind complex or ambiguous expressions. Second, the preset answering methods are relatively rigid and cannot flexibly respond to diverse query needs. Third, such systems lack context understanding and multi - turn dialogue capabilities and are difficult to handle complex notarization problems that require in - depth interaction. In addition, it is difficult to update and maintain the knowledge base, and it is difficult to reflect the latest changes in laws and regulations in a timely manner.

[0004] In summary, there is an urgent need to develop a notarization intelligent question - answering customer service method based on a knowledge graph. This method can make full use of the advantages of the knowledge graph to realize the structured representation and efficient management of the complex knowledge system in the notarization field. The system can process complex queries that require multi - step logical reasoning, greatly improving the accuracy and depth of question - answering. At the same time, using the relevance of the graph, more comprehensive and relevant information can be provided to users, and the present invention can solve the problems in the prior art. Summary of the Invention

[0005] Embodiments of the present invention provide a notarization intelligent question - answering customer service method and system based on a knowledge graph, which can solve the problems in the prior art.

[0006] In the first aspect of the embodiments of the present invention,

[0007] A notarization intelligent question - answering customer service method based on a knowledge graph is provided, including:

[0008] Collect multi-source heterogeneous data in the field of notarization, preprocess the multi-source heterogeneous data to obtain preprocessed data, use a named entity recognition algorithm to extract professional terms and domain concepts from the preprocessed data to form an entity set, analyze the relationships between entities in the entity set based on semantic role labeling technology to construct a relationship set, combine the entity set and the relationship set to form an initial knowledge graph, and based on a dynamic update mechanism with time decay, extract knowledge from newly added preprocessed data and supplement it to the initial knowledge graph to form a dynamic knowledge graph, and perform vectorization representation on the dynamic knowledge graph to obtain a graph vector;

[0009] Receive a multi-modal question input by the user, extract speech text, image text, and image semantic information based on the speech and images in the multi-modal question, combine the plain text in the multi-modal question to perform multi-modal feature fusion, determine the fusion features, perform deep semantic understanding on the fusion features to obtain semantic features, input the semantic features into an intent recognition model, and through semantic mapping, adapt to predefined notarization business intent categories to determine the user intent. Based on the user intent and combined with the dynamic knowledge graph, perform entity linking and relationship extraction on the semantic features to form a structured query expression;

[0010] Based on the user intent and the structured query expression, perform multi-hop reasoning in the dynamic knowledge graph, use the graph vector to calculate similarity, and combine the user intent to determine the reasoning path. Perform knowledge association reasoning along the reasoning path to obtain a preliminary reasoning result, perform interpretability analysis on the preliminary reasoning result to generate an inference chain and a confidence score to form a detailed reasoning result, combine the pre-obtained user portrait information and the user intent to perform personalized ranking on the detailed reasoning result to obtain an ordered reasoning result, input the ordered reasoning result into a pre-constructed Transformer model, output a preliminary natural language answer, and apply text style transfer technology to finally generate a personalized answer.

[0011] In an alternative embodiment,

[0012] Using a named entity recognition algorithm to extract professional terms and domain concepts from the preprocessed data to form an entity set, analyzing the relationships between entities in the entity set based on semantic role labeling technology to construct a relationship set, and combining the entity set and the relationship set to form an initial knowledge graph includes:

[0013] Construct a labeled dataset for the notarization field, and train a deep fusion network for notarization entity recognition based on the labeled dataset. The deep fusion network for notarization entity recognition includes an encoding layer based on BERT, a feature extraction layer based on a bidirectional long short-term memory network, and a decoding layer based on a conditional random field. Use the deep fusion network for notarization entity recognition to perform named entity recognition on texts in the notarization field, and obtain an entity set, where the entity set includes entity names, entity types, occurrence frequencies, and first occurrence positions;

[0014] Perform dependency syntactic analysis on sentences containing entities in the entity set, identify the subject-predicate-object structure, and use a semantic role labeling model to perform semantic role labeling on sentences containing the subject-predicate-object structure to obtain a predicate-argument structure and determine semantic role labeling tags;

[0015] Construct a mapping rule set from semantic role labeling tags to notarization field relationship types, and convert the semantic role labeling results into candidate relationships according to the mapping rule set;

[0016] Construct a relationship classifier and train it through a preset remotely supervised dataset. Use the relationship classifier to score and filter the candidate relationships to obtain a relationship set;

[0017] Import the entities in the entity set as entity nodes into a graph database, import the relationships in the relationship set as relationship edges into the graph database, connect the corresponding entity nodes, and use graph algorithms to optimize the imported entity nodes and relationship edges to eliminate redundant relationships and merge synonymous entities to construct an initial knowledge graph.

[0018] In an alternative embodiment,

[0019] Construct a labeled dataset for the notarization field, and train a deep fusion network for notarization entity recognition based on the labeled dataset. The deep fusion network for notarization entity recognition includes an encoding layer based on BERT, a feature extraction layer based on a bidirectional long short-term memory network, and a decoding layer based on a conditional random field. Use the deep fusion network for notarization entity recognition to perform named entity recognition on texts in the notarization field, and obtain an entity set including:

[0020] Construct a labeled dataset for the notarization field, and perform masked language model pre-training on the BERT model based on the labeled dataset for the notarization field to obtain a pre-trained BERT model adapted to the notarization field;

[0021] In the deep fusion network for notarization entity recognition, initialize the encoding layer with the weights of the pre-trained BERT model; input the labeled dataset for the notarization field into the deep fusion network for notarization entity recognition, and obtain word vector representations through the encoding layer;

[0022] Input the word vector representation into the feature extraction layer to extract sequence features, where the feature extraction layer includes a forward long short-term memory network and a backward long short-term memory network, each containing an input gate, a forget gate, and an output gate;

[0023] Input the sequence features into the decoding layer, where the decoding layer includes a linear-chain conditional random field, establish a label transition probability matrix, and use the Viterbi algorithm to decode to obtain the optimal label sequence;

[0024] Adopt a joint learning strategy to optimize the encoding layer, feature extraction layer, and decoding layer simultaneously, use the AdamW optimizer for parameter update, and implement a learning rate warm-up and linear decay strategy;

[0025] After each training epoch, use a preset validation set to evaluate the performance of the notarization entity recognition deep fusion network, and combine with an early stopping strategy to save the model weights with the highest F1 score on the validation set to obtain the trained notarization entity recognition deep fusion network;

[0026] Input the notarization domain text to be recognized into the trained notarization entity recognition deep fusion network, and successively pass through the encoding layer, feature extraction layer, and decoding layer to obtain a label sequence;

[0027] According to the label sequence, identify the entities in the notarization domain text and construct an entity set.

[0028] In an optional embodiment,

[0029] Based on a time decay-based dynamic update mechanism, extract knowledge from newly added preprocessed data and supplement it to the initial knowledge graph to form a dynamic knowledge graph, including:

[0030] Construct a time decay function, which adopts an exponential decay form to control the decay speed of knowledge units;

[0031] Receive newly added preprocessed data, perform entity recognition and relationship extraction on the newly added preprocessed data to obtain new knowledge units, use the cosine similarity method to calculate the similarity between the new knowledge units and the existing knowledge units in the initial knowledge graph to obtain a similarity calculation result, and according to the similarity calculation result, determine whether the similarity calculation result is higher than a preset similarity threshold:

[0032] When the similarity calculation result is higher than the preset similarity threshold, update the existing knowledge units with the new knowledge units, and use the time decay function to calculate the fusion weights of the new knowledge units and the existing knowledge units; based on the fusion weights, perform weighted fusion on the new knowledge units and the existing knowledge units to obtain fusion knowledge units;

[0033] When the similarity is lower than the preset similarity threshold, the new knowledge unit is regarded as an added knowledge unit;

[0034] Update the fused knowledge unit to the corresponding position, add the added knowledge unit to the initial knowledge graph, update the initial knowledge graph, and record the timestamp of each update or addition;

[0035] According to a preset full-scan period, periodically conduct a full scan of the initial knowledge graph, calculate the time decay value of each knowledge unit using the time decay function, and determine whether the time decay value of each knowledge unit is less than the preset decay threshold. When the time decay value corresponding to a knowledge unit is less than the preset decay threshold, remove the corresponding knowledge unit from the initial knowledge graph;

[0036] The initial knowledge graph is updated to form a dynamic knowledge graph.

[0037] In an optional embodiment,

[0038] Receive a multimodal question input by the user. Based on the speech and image in the multimodal question, extract speech text, image text, and image semantic information, combine with the plain text in the multimodal question, perform multimodal feature fusion, determine the fusion feature, perform deep semantic understanding on the fusion feature to obtain a semantic feature, input the semantic feature into an intent recognition model, and through semantic mapping, adapt to predefined notarization business intent categories to determine the user intent, including:

[0039] Receive a multimodal question, where the multimodal question contains speech information, image information, and text information;

[0040] Use a speech recognition model to convert the speech information in the multimodal question into a first text sequence; process the image information in the multimodal question using an optical character recognition engine to obtain a second text sequence; splice the first text sequence, the second text sequence, and the text information in the multimodal question to form a combined text sequence, and map each token in the combined text sequence to a specified dimensional space through a first independent linear layer to obtain a text feature;

[0041] Use a Vision Transformer model to extract the semantic feature of the image information in the multimodal question to obtain an image semantic vector, and map the image semantic vector to the same corresponding dimensional space as the text feature through a second independent linear layer to obtain an image feature;

[0042] Calculate the correlation between the text features and the image features using the multi-head attention mechanism to obtain an attention output. Perform residual connection and layer normalization on the attention output and the text features to obtain fused features. Input the fused features into a pre-trained bi-directional encoder to obtain a contextually related representation; extract token vectors from the contextually related representation and map them to the task corresponding space through a linear layer and an activation function to obtain semantic features;

[0043] Map the semantic features to a predefined notarization business intent category space through a fully connected layer to obtain intent logical values. Apply the sigmoid function to the intent logical values to calculate the probability of each intent category, and select the intent category with the highest probability to determine the user intent.

[0044] In an alternative embodiment,

[0045] Based on the user intent and combined with a dynamic knowledge graph, perform entity linking and relation extraction on the semantic features to form a structured query expression, including:

[0046] Construct a pre-trained stacked network model. The stacked network model adopts a training objective of permutation language modeling and uses a two-stream self-attention mechanism and a segmental recurrent mechanism;

[0047] Input the semantic features into the stacked network model to obtain a bi-directional context-aware word vector representation; construct a multi-task learning architecture at the top layer of the stacked network model, including an entity recognition task and a relation extraction task; introduce a relative position encoding mechanism into the multi-task learning architecture to capture long-range dependencies; use a dynamic weight allocation strategy to adaptively adjust the importance weights of the entity recognition task and the relation extraction task in the multi-task learning architecture;

[0048] Integrate the dynamic knowledge graph, process the dynamic knowledge graph through a graph attention network to obtain enhanced entity and enhanced relation representations in the word vector representation, and form an enhanced word vector representation;

[0049] Apply adversarial training techniques to the enhanced word vector representation, train to obtain a final multi-task learning architecture, input semantic features, and identify the corresponding entity mentions; calculate the similarity between the entity mentions and the entities in the dynamic knowledge graph, select the entity with the highest similarity as the linking result to obtain a set of linked entities, and generate candidate relation pairs based on the entity pairs in the set of linked entities;

[0050] Based on the relation extraction task in the multi-task learning architecture, use the semantic features and the candidate relation pairs as inputs to predict the relations between entity pairs to obtain a set of relations;

[0051] Select the corresponding query template based on the user's intention, and fill the link entity set and the relationship set into the query template to generate a structured query expression.

[0052] In an alternative embodiment,

[0053] Based on the user's intention and the structured query expression, perform multi-hop reasoning in the dynamic knowledge graph, calculate the similarity using the graph vectors, and combine with the user's intention to determine the reasoning path, and perform knowledge association reasoning along the reasoning path to obtain the preliminary reasoning results including:

[0054] Receive the user's intention and the structured query expression as inputs, determine the starting entity for reasoning in the dynamic knowledge graph, use the starting entity for reasoning as the current node, identify the entities and relationships directly connected to the current node to form a candidate next-hop set;

[0055] Using the pre-generated graph vectors, calculate the similarity score between the current node and each node in the candidate next-hop set; convert the user's intention into a vector representation, fuse it with the graph vectors of each node in the candidate next-hop set to obtain a fused vector; combine the similarity score with the fused vector to generate a comprehensive score for each candidate next-hop node.

[0056] Based on the comprehensive score, use a reinforcement learning agent to select the next hop, where the reinforcement learning agent takes the current state as input and outputs the probability distribution of selecting the next hop, and the current state includes the current node information, the user's intention, the comprehensive score, and the selected path information;

[0057] According to the selected next hop, update the current node and the selected path information;

[0058] Apply predefined inference rules to the selected path for explicit rule-based reasoning;

[0059] Use a graph neural network model to encode the selected path for implicit reasoning based on the graph neural network, where the graph neural network model is a graph convolutional network;

[0060] Fuse the results of explicit reasoning and implicit reasoning to obtain a fused reasoning result;

[0061] Know that the preset hop count limit is reached, record the path selection during the entire reasoning process to form a complete reasoning chain; based on the length of the selected path, the similarity score of each hop, and the reliability score of the inference rules used, calculate the confidence score of the entire reasoning result;

[0062] Combine the complete reasoning chain, the finally reached node, and the confidence score to form the preliminary reasoning result.

[0063] In the second aspect of the embodiments of the present invention,

[0064] A notarization intelligent question - answering customer service system based on a knowledge graph is provided, including:

[0065] A first unit for collecting multi - source heterogeneous data in the notarization field, pre - processing the multi - source heterogeneous data to obtain pre - processed data, extracting professional terms and domain concepts from the pre - processed data by using a named entity recognition algorithm to form an entity set, analyzing the relationships between entities in the entity set based on semantic role annotation technology to construct a relationship set, combining the entity set and the relationship set to form an initial knowledge graph, based on a dynamic update mechanism with time decay, extracting knowledge from newly added pre - processed data and supplementing it to the initial knowledge graph to form a dynamic knowledge graph, and performing vector representation on the dynamic knowledge graph to obtain a graph vector;

[0066] A second unit for receiving a multi - modal question input by a user, extracting speech text, image text and image semantic information based on the speech and image in the multi - modal question, combining the pure text in the multi - modal question to perform multi - modal feature fusion, determining a fusion feature, performing deep semantic understanding on the fusion feature to obtain a semantic feature, inputting the semantic feature into an intent recognition model, and through semantic mapping, adapting to a predefined notarization service intent category to determine the user intent, and based on the user intent and the dynamic knowledge graph, performing entity linking and relationship extraction on the semantic feature to form a structured query expression;

[0067] A third unit for performing multi - hop reasoning in the dynamic knowledge graph based on the user intent and the structured query expression, calculating similarity using the graph vector, and combining the user intent to determine an inference path, performing knowledge - associated reasoning along the inference path to obtain a preliminary inference result, performing interpretability analysis on the preliminary inference result to generate an inference chain and a confidence score to form a detailed inference result, combining the pre - obtained user portrait information and the user intent to perform personalized sorting on the detailed inference result to obtain an ordered inference result, inputting the ordered inference result into a pre - constructed Transformer model to output a preliminary natural language answer, and applying text style transfer technology to finally generate a personalized answer.

[0068] In the third aspect of the embodiments of the present invention,

[0069] An electronic device is provided, including:

[0070] A processor;

[0071] A memory for storing instructions executable by the processor;

[0072] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0073] In the fourth aspect of the embodiments of the present invention,

[0074] a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0075] In the embodiments of the present invention, through named entity recognition and semantic role labeling technologies, professional terms, domain concepts and their relationships in the notarization field are accurately extracted to construct a high-quality initial knowledge graph; a time decay mechanism is adopted to automatically extract knowledge from new data and update the initial knowledge graph to ensure the timeliness and dynamics of the knowledge graph; the dynamic knowledge graph is vectorized to improve the application efficiency of the knowledge graph in subsequent reasoning, retrieval and other tasks; through feature extraction and fusion of voice, image and text information, a comprehensive understanding of multi-modal problems is achieved, and the parsing ability of user input is improved; through semantic mapping, semantic features are adapted to predefined notarization service intent categories to ensure the accuracy and effectiveness of user intent recognition; combined with the dynamic knowledge graph, entity linking and relationship extraction are performed, and finally a structured query expression is generated to provide a standardized data basis for subsequent operations; through multi-hop reasoning in the dynamic knowledge graph, combined with graph vector similarity and user intent, an efficient reasoning path is determined to improve the accuracy of knowledge association reasoning; interpretability and confidence analysis: generate an inference chain and a confidence score to provide interpretability for the inference result and enhance the user's understanding and trust in the inference process; based on the user profile and intent, personalized ranking of the inference results is performed to improve the relevance and personalized adaptability of the results. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 is a schematic flowchart of the method for notarization intelligent question-answering customer service based on a knowledge graph according to the embodiments of the present invention;

[0077] Figure 2 is a schematic structural diagram of the notarization intelligent question-answering customer service system based on a knowledge graph according to the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0078] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0079] The technical solution of the present invention will be described in detail below with specific embodiments. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0080] Figure 1 It is a schematic flowchart of the notarization intelligent question - answering customer service method based on a knowledge graph according to an embodiment of the present invention. As Figure 1 shown, the method includes:

[0081] S101. Collect multi - source heterogeneous data in the notarization field, pre - process the multi - source heterogeneous data to obtain pre - processed data, use a named entity recognition algorithm to extract professional terms and domain concepts from the pre - processed data to form an entity set, analyze the relationships between entities in the entity set based on semantic role annotation technology to construct a relationship set, combine the entity set and the relationship set to form an initial knowledge graph, based on a time - decay dynamic update mechanism, extract knowledge from the newly added pre - processed data and supplement it into the initial knowledge graph to form a dynamic knowledge graph, and perform vectorization representation on the dynamic knowledge graph to obtain a graph vector;

[0082] S102. Receive a multi - modal question input by a user, extract speech text, image text, and image semantic information based on the speech and images in the multi - modal question, combine with the plain text in the multi - modal question to perform multi - modal feature fusion, determine the fusion feature, perform deep semantic understanding on the fusion feature to obtain a semantic feature, input the semantic feature into an intent recognition model, and through semantic mapping, adapt to predefined notarization business intent categories to determine the user intent. Based on the user intent and combined with the dynamic knowledge graph, perform entity linking and relation extraction on the semantic feature to form a structured query expression;

[0083] S103. Based on the user intent and the structured query expression, perform multi - hop reasoning in the dynamic knowledge graph, use the graph vector to calculate similarity, and combine with the user intent to determine the reasoning path. Perform knowledge - associated reasoning along the reasoning path to obtain a preliminary reasoning result, perform interpretability analysis on the preliminary reasoning result to generate an inference chain and a confidence score to form a detailed reasoning result. Combine the pre - obtained user portrait information and the user intent to perform personalized sorting on the detailed reasoning result to obtain an ordered reasoning result. Input the ordered reasoning result into a pre - constructed Transformer model to output a preliminary natural language answer, and apply text style transfer technology to finally generate a personalized answer.

[0084] In a specific implementation manner, the steps are as follows:

[0085] First, collect multi-source heterogeneous data in the notarization field, including notarization laws and regulations texts, notarization case databases, notarization business manuals, etc. Preprocess the collected data, including text cleaning, word segmentation, stop word removal, etc., to obtain normalized preprocessed data. Then, use the named entity recognition algorithm to extract professional terms and domain concepts from the preprocessed data, such as "notary office", "notary public", "notarial certificate", etc., to form an entity set. Specifically, a named entity recognition model based on BiLSTM-CRF can be used and trained with a notarization field annotation dataset.

[0086] Next, analyze the relationships between entities in the entity set based on semantic role labeling technology to construct a relationship set. A semantic role labeling method based on dependency syntactic analysis can be used to extract semantic roles such as subject, verb, and object, and identify the semantic relationships between entities, such as "notary public - issue - notarial certificate". Combine the entity set and the relationship set to form an initial knowledge graph.

[0087] Then, based on a dynamic update mechanism with time decay, extract knowledge from the newly added preprocessed data and supplement it to the initial knowledge graph. Specifically, a time decay function f(t) = e (-λt) can be set, where t is the existence time of the knowledge and λ is the decay coefficient. Assign an initial weight of 1 to each triple in the graph, and the weight gradually decays over time. When the weight is below the threshold, delete the knowledge from the graph. At the same time, extract new knowledge from the newly added data, assign a higher initial weight and add it to the graph, thus forming a dynamic knowledge graph.

[0088] Perform vector representation on the dynamic knowledge graph to obtain a graph vector. Knowledge graph embedding methods such as TransE can be used to map entities and relationships to a low-dimensional vector space for subsequent similarity calculation.

[0089] In the question-answering stage, first receive the multi-modal question input by the user, including speech, image, and text. Perform speech recognition on the speech to obtain speech text, perform text recognition and scene understanding on the image to obtain image text and semantic information. Fuse the features of the speech text, image text, and image semantic information with the original text. A multi-modal feature fusion method based on the attention mechanism can be used to obtain a fused feature vector.

[0090] Perform in-depth semantic understanding on the fused features. Pre-trained language models such as BERT can be used to extract semantic features. Input the semantic features into an intent recognition model, and through semantic mapping, map the user's question to predefined notarization business intent categories, such as "apply for notarization", "notarization fees", etc., to determine the user's intent.

[0091] Based on the user's intention and the dynamic knowledge graph, entity linking and relationship extraction are performed on semantic features. For entity linking, a BERT-based entity disambiguation model can be used to link the entities mentioned in the question to the corresponding nodes in the graph. For relationship extraction, a remotely supervised relationship extraction model can be used to extract the semantic relationships between entities. Finally, structured query expressions such as SPARQL are formed.

[0092] According to the user's intention and the structured query expression, multi-hop reasoning is performed in the dynamic knowledge graph. First, the similarity between entities and relationships is calculated using graph vectors, and the initial reasoning path is determined in combination with the user's intention. Then, knowledge association reasoning is carried out along the reasoning path, and a path search algorithm based on reinforcement learning can be adopted. The reasoning process is analyzed for interpretability, generating an inference chain and a confidence score to form a detailed reasoning result.

[0093] Combining the pre-obtained user profile information (such as user preferences, historical behaviors, etc.) and the user's intention, the detailed reasoning results are sorted personalized. Ranking algorithms such as Learning to Rank can be used to obtain the ordered reasoning results. The ordered reasoning results are input into a pre-constructed Transformer model, which is fine-tuned on the Q&A data in the notarization field to output a preliminary natural language answer.

[0094] Finally, the text style transfer technology based on VAE is applied to convert the preliminary answer into a personalized expression style that conforms to the user profile, such as formal, colloquial, etc. Through the above steps, personalized notarization intelligent Q&A results are finally generated.

[0095] Exemplarily, the user asks "I want to apply for a property notarization. What materials do I need to prepare?". First, the semantic features are extracted, and the user's intention is identified as "applying for notarization". The entity linking identifies the "property notarization" entity, and the relationship extraction obtains the "materials - preparation - property notarization" relationship. Multi-hop reasoning is performed in the knowledge graph to find the "required materials" node related to "property notarization". Combining the user profile for personalized ranking and expression, the final answer is generated: "Hello, the following materials are required for applying for property notarization: 1) The original copy of the house ownership certificate; 2) The original copy of the ID card; 3) The original copy of the household register. It is recommended that you prepare these materials in advance, which can speed up the application process. If you have any other questions, feel free to let me know at any time."

[0096] In this embodiment, through named entity recognition and semantic role labeling technologies, professional terms, domain concepts and their relationships in the notarization field are accurately extracted to construct a high-quality initial knowledge graph; a time decay mechanism is adopted to automatically extract knowledge from new data and update the initial knowledge graph to ensure the timeliness and dynamics of the knowledge graph; the dynamic knowledge graph is vectorized to improve the application efficiency of the knowledge graph in subsequent reasoning, retrieval and other tasks; effective processing and fusion of multi-source heterogeneous data are realized to construct a more comprehensive and accurate knowledge graph in the notarization field; through feature extraction and fusion of voice, image and text information, a comprehensive understanding of multi-modal problems is achieved, and the parsing ability of user input is improved; deep semantic analysis is performed using the fused features to more accurately obtain the user's intention information and enhance the accuracy of intention recognition; through semantic mapping, the semantic features are adapted to predefined notarization business intention categories to ensure the accuracy and effectiveness of user intention recognition; combined with the dynamic knowledge graph, entity linking and relationship extraction are performed, and finally a structured query expression is generated to provide a standardized data basis for subsequent operations; through multi-hop reasoning in the dynamic knowledge graph, combined with graph vector similarity and user intention, an efficient reasoning path is determined to improve the accuracy of knowledge association reasoning; interpretability and confidence analysis: generate an inference chain and a confidence score to provide interpretability for the inference result and enhance the user's understanding and trust of the inference process; based on the user profile and intention, the inference results are sorted personalizedly to improve the relevance and personalized adaptability of the results; a preliminary answer is generated through the Transformer model, and a personalized answer is generated in combination with the text style transfer technology to ensure that the output conforms to the user's language habits and style.

[0097] In an alternative embodiment, a named entity recognition algorithm is used to extract professional terms and domain concepts from the preprocessed data to form an entity set, and the relationships between entities in the entity set are analyzed based on semantic role labeling technology to construct a relationship set. Combining the entity set and the relationship set to form an initial knowledge graph includes:

[0098] Construct a labeled dataset for the notarization field, and train a deep fusion network for notarization entity recognition based on the labeled dataset. The deep fusion network for notarization entity recognition includes a BERT-based encoding layer, a feature extraction layer based on a bidirectional long short-term memory network, and a decoding layer based on a conditional random field. Named entity recognition is performed on the notarization field text through the deep fusion network for notarization entity recognition to obtain an entity set, and the entity set includes entity names, entity types, occurrence frequencies, and first occurrence positions;

[0099] Perform dependency parsing on sentences containing entities in the entity set, identify the subject-predicate-object structure, use a semantic role labeling model to perform semantic role labeling on sentences containing the subject-predicate-object structure, obtain the predicate-argument structure, and determine the semantic role labeling tags.

[0100] Construct a mapping rule set from semantic role labeling tags to notarization domain relationship types, and convert the semantic role labeling results into candidate relationships according to the mapping rule set.

[0101] Construct a relationship classifier and train it through a preset remote supervision data set. Use the relationship classifier to score and filter the candidate relationships to obtain a relationship set.

[0102] Import the entities in the entity set as entity nodes into the graph database, import the relationships in the relationship set as relationship edges into the graph database, connect the corresponding entity nodes, and use graph algorithms to optimize the imported entity nodes and relationship edges to eliminate redundant relationships and merge synonymous entities to construct an initial knowledge graph.

[0103] For the construction method of the notarization domain knowledge graph, this embodiment provides a specific implementation scheme based on deep learning and natural language processing technologies. This scheme mainly includes the following steps:

[0104] First, construct a notarization domain annotation data set. By collecting a large number of notarization-related texts and inviting notarization domain experts for manual annotation, the annotation content includes entity types, entity boundaries, relationships between entities, etc. The annotation data set consists of a training set, a validation set, and a test set, where the training set accounts for about 70%, and the validation set and the test set each account for 15%.

[0105] Train a notarization entity recognition deep fusion network based on the annotation data set. This network adopts a three-layer structure: the first layer is an encoding layer based on BERT, which uses the pre-trained BERT model to encode the input text to obtain context-related word vector representations; the second layer is a feature extraction layer based on a bidirectional long short-term memory network (Bi-LSTM) to further extract sequence features; the third layer is a decoding layer based on a conditional random field (CRF) to decode the output of the Bi-LSTM to obtain the final entity label sequence. The training of the network uses the mini-batch gradient descent algorithm, the batch size is set to 32, the initial learning rate is 0.001, and the Adam optimizer is used.

[0106] Use the trained deep fusion network for notarization entity recognition to perform named entity recognition on texts in the notarization field. For the input text, first perform preprocessing of word segmentation and part-of-speech tagging, and then send it into the network for recognition. The network outputs the entity tags of each word, and entities are extracted according to the tag sequence. Post-process the recognition results, including entity deduplication, merging adjacent entities, etc. Finally, obtain an entity set, and each entity contains information such as entity name, entity type, occurrence frequency, and first occurrence position.

[0107] For example, for the input text "Zhang San and Li Si signed a housing sales contract", the entity set obtained after named entity recognition is:

[0108] {

[0109] "Zhang San": {type: "natural person", frequency: 1, first position: 0},

[0110] "Li Si": {type: "natural person", frequency: 1, first position: 3},

[0111] "Housing sales contract": {type: "contract", frequency: 1, first position: 8}

[0112] }

[0113] Next, perform dependency parsing on the sentences containing entities to identify the subject-predicate-object structure. Adopt a dependency parsing model based on neural networks, such as a transition system model based on bidirectional LSTM. The dependency parsing result of the above example sentence is:

[0114] Subject: Zhang San;

[0115] Predicate: signed;

[0116] Object: housing sales contract;

[0117] Then use a semantic role labeling model to perform semantic role labeling on the sentences containing the subject-predicate-object structure. The semantic role labeling model adopts an end-to-end neural network structure, including a word embedding layer, a BiLSTM encoding layer, and a CRF decoding layer. The model input is the word sequence and syntactic tree information, and the output is the semantic role label of each word. The semantic role labeling result of the example sentence is:

[0118] [Zhang San] A0 and [Li Si] A0 [signed] V [housing sales contract] A1;

[0119] Among them, A0 represents the doer of the action, V represents the predicate, and A1 represents the recipient of the action.

[0120] Construct a mapping rule set from semantic role labeling tags to notarization field relationship types. For example:

[0121] A0-V-A1 → Signing relationship;

[0122] A0-V-A2 → Agency relationship;

[0123] LOC-V-A1 → Location relationship;

[0124] And so on.

[0125] According to the mapping rule set, convert the semantic role annotation results into candidate relationships. For the example sentence, the obtained candidate relationships are:

[0126] (Zhang San, signed, housing sales contract)

[0127] (Li Si, signed, housing sales contract)

[0128] Construct a relationship classifier and train it. The relationship classifier adopts a convolutional neural network structure, including a word embedding layer, a convolutional layer, a pooling layer, and a fully connected layer. Use a preset remote supervision data set for training. The data set contains a large number of automatically annotated entity pairs and their relationships. During training, use the cross-entropy loss function and optimize it using the Adam optimizer.

[0129] Use the trained relationship classifier to score and filter the candidate relationships. The relationship classifier outputs the probability scores of each relationship type, and selects the relationship type with the highest probability and exceeding the threshold as the final result. For candidate relationships with probabilities lower than the threshold, they are filtered. After being processed by the relationship classifier, the final relationship set is obtained:

[0130] {(Zhang San, signed, housing sales contract), (Li Si, signed, housing sales contract)}

[0131] Import the entities in the entity set as entity nodes into a graph database, such as Neo4j. Each node contains attribute information such as entity name and type. Import the relationships in the relationship set as relationship edges into the graph database to connect the corresponding entity nodes. For example:

[0132] CREATE(Zhang San: Natural Person {name: 'Zhang San'})

[0133] CREATE(Li Si: Natural Person {name: 'Li Si'})

[0134] CREATE(Contract: Contract {name: 'Housing Sales Contract'})

[0135] CREATE(Zhang San)-[: signed]->(Contract)

[0136] CREATE(Li Si)-[: signed]->(Contract)

[0137] Optimize the imported entity nodes and relationship edges using graph algorithms. First, use community discovery algorithms such as the Louvain algorithm to partition the graph and discover tightly associated entity clusters. Then, use the PageRank algorithm to calculate the importance scores of the nodes and prune redundant nodes with low importance. For entity nodes with high similarity, such as "purchase contract" and "house sales contract", use entity alignment algorithms to merge them.

[0138] Finally, obtain the optimized initial knowledge graph, which contains the core entities and relationships in the notarization field. This knowledge graph can support subsequent knowledge reasoning, question answering, and other applications. As new notarization texts are continuously added, the knowledge graph can be continuously updated and expanded to continuously improve its coverage and accuracy.

[0139] In this embodiment, through the notarization entity recognition deep fusion network, combining BERT, bidirectional LSTM, and CRF, it can accurately identify named entities in notarization field texts, ensuring the accuracy and integrity of the entity set; through the dependency syntactic analysis and semantic role annotation model, it identifies the subject-predicate-object structure and maps it to the relationships in the notarization field, enhancing the depth of semantic understanding; through the relationship classifier to filter candidate relationships, and using graph databases and graph algorithms to optimize entity nodes and relationship edges, eliminating redundancy and merging synonymous entities, ensuring the accuracy and quality of the knowledge graph; finally, construct a high-quality and accurate notarization field knowledge graph, providing a solid foundation for the analysis and application of notarization data.

[0140] In an alternative implementation, construct a notarization field annotation dataset, and train a notarization entity recognition deep fusion network based on the annotation dataset. The notarization entity recognition deep fusion network includes an encoding layer based on BERT, a feature extraction layer based on a bidirectional long short-term memory network, and a decoding layer based on a conditional random field. Perform named entity recognition on notarization field texts through the notarization entity recognition deep fusion network to obtain an entity set including:

[0141] Construct a notarization field annotation dataset, and perform masked language model pre-training on the BERT model based on the notarization field annotation dataset to obtain a pre-trained BERT model adapted to the notarization field;

[0142] In the notarization entity recognition deep fusion network, initialize the encoding layer with the weights of the pre-trained BERT model; input the notarization field annotation dataset into the notarization entity recognition deep fusion network, and obtain word vector representations through the encoding layer;

[0143] Input the word vector representations into the feature extraction layer to extract sequence features, where the feature extraction layer includes a forward long short-term memory network and a backward long short-term memory network, each containing an input gate, a forget gate, and an output gate;

[0144] Input the sequence features into the decoding layer, which includes a linear-chain conditional random field, establish a label transition probability matrix, and use the Viterbi algorithm to decode to obtain the optimal label sequence;

[0145] Adopt a joint learning strategy to optimize the encoding layer, feature extraction layer, and decoding layer simultaneously, use the AdamW optimizer for parameter update, and implement a learning rate warm-up and linear decay strategy;

[0146] After each training epoch, evaluate the performance of the notarization entity recognition deep fusion network using a preset validation set, and combine with an early stopping strategy to save the model weights with the highest F1 score on the validation set, obtaining a trained notarization entity recognition deep fusion network;

[0147] Input the text in the notarization field to be recognized into the trained notarization entity recognition deep fusion network, and sequentially pass through the encoding layer, feature extraction layer, and decoding layer to obtain a label sequence;

[0148] According to the label sequence, identify the entities in the notarization field text and construct an entity set.

[0149] In this embodiment, first construct a notarization field annotation dataset. Specifically, collect a large number of text materials in the notarization field, including notarial deeds, notarization application materials, notarization laws and regulations, etc. Then, notarization field experts manually annotate these texts, annotating the entity types therein, such as parties, notarization matters, notarization institutions, etc. Finally, a dataset containing the text and its corresponding entity labels is obtained.

[0150] Next, based on the constructed notarization field annotation dataset, perform masked language model pre-training on the BERT model. Specifically, input the text in the annotation dataset into the BERT model, randomly mask 15% of the tokens, and let the model predict these masked tokens. In this way, the BERT model can learn the language features and context information of the notarization field text. After pre-training, a pre-trained BERT model adapted to the notarization field is obtained.

[0151] Then construct a notarization entity recognition deep fusion network. This network includes three layers: an encoding layer based on BERT, a feature extraction layer based on a bidirectional long short-term memory network, and a decoding layer based on a conditional random field. Among them, the encoding layer is initialized with the weights of the pre-trained BERT model obtained above. Input the notarization field annotation dataset into this network, and first obtain the word vector representation through the encoding layer.

[0152] The word vectors are then input into the feature extraction layer for sequence feature extraction. The feature extraction layer consists of two long short-term memory networks, a forward one and a backward one, each containing input gates, forget gates, and output gate structures. Through bidirectional processing, long-range dependencies in the context can be captured.

[0153] The extracted sequence features are then input into the decoding layer. The decoding layer uses a linear-chain conditional random field to establish a label transition probability matrix and decodes using the Viterbi algorithm to obtain the optimal label sequence.

[0154] During the training process, a joint learning strategy is adopted to optimize the three layers simultaneously. The AdamW optimizer is used for parameter updates, and a learning rate warm-up and linear decay strategy are implemented. Specifically, a small learning rate is used at the beginning of training, gradually increasing to a set maximum value as training progresses, and then linearly decaying. This strategy allows the model to converge stably at the beginning of training and fully learn later.

[0155] After each training epoch, the performance of the network is evaluated using a preset validation set. The F1 score on the validation set is calculated, and combined with an early stopping strategy, training is stopped in a timely manner when the performance on the validation set no longer improves to avoid overfitting. The model weights with the highest F1 score on the validation set are saved to obtain the final trained deep fusion network for notarial entity recognition.

[0156] When in use, the text of the notarial field to be recognized is input into the trained network. The text passes through the encoding layer in sequence to obtain word vectors, the feature extraction layer extracts sequence features, and the decoding layer obtains the label sequence. Based on the output label sequence, various entities in the text can be recognized to construct an entity set.

[0157] For a specific example, the input text is "Zhang San applied to Beijing Chang'an Notary Office for notarization of house ownership on May 1, 2021". After being processed by the network, the output label sequence is "B-PERSON O O-TIME I-TIME I-TIME O B-ORG I-ORG I-ORG O O O B-ITEM I-ITEM I-ITEM O". Here, B- indicates the start of an entity, I- indicates inside an entity, and O indicates a non-entity. According to this label sequence, the person name entity "Zhang San", the time entity "May 1, 2021", the organization entity "Beijing Chang'an Notary Office", and the matter entity "notarization of house ownership" can be recognized. Finally, a set containing these entities is constructed.

[0158] Through this method, various entities in the text of the notarial field can be effectively recognized, providing a basis for subsequent tasks such as information extraction and knowledge graph construction. This method combines the powerful language representation ability of BERT, the sequence modeling ability of long short-term memory networks, and the label dependency modeling ability of conditional random fields, and can accurately recognize professional entities in the notarial field.

[0159] In this embodiment, by pre-training the BERT model on the notarization domain annotation dataset, it is made to be more adaptable to the proprietary language characteristics of the notarization domain, improving the accuracy of entity recognition; using a deep fusion network, including a BERT encoding layer, a bidirectional LSTM feature extraction layer, and a CRF decoding layer, can more precisely extract sequence features and decode the optimal label sequence, thereby enhancing the named entity recognition effect; adopting a joint learning strategy, an AdamW optimizer, and a learning rate warm-up and linear decay strategy can effectively accelerate model convergence and improve model performance; through validation set evaluation and an early stopping strategy, it is ensured that the final model weights reach the highest F1 score on the validation set, thus guaranteeing the recognition effect; the trained model can accurately identify entities in the notarization domain text and construct an accurate entity set for subsequent processing or analysis.

[0160] In an alternative embodiment, based on a dynamic update mechanism of time decay, knowledge is extracted from the newly added preprocessed data and supplemented into the initial knowledge graph to form a dynamic knowledge graph, including:

[0161] Construct a time decay function, which adopts an exponential decay form to control the decay speed of knowledge units;

[0162] Receive the newly added preprocessed data, perform entity recognition and relationship extraction on the newly added preprocessed data to obtain new knowledge units, use the cosine similarity method to calculate the similarity between the new knowledge units and the existing knowledge units in the initial knowledge graph, obtain the similarity calculation result, and according to the similarity calculation result, judge whether the similarity calculation result is higher than a preset similarity threshold:

[0163] When the similarity calculation result is higher than the preset similarity threshold, update the existing knowledge units with the new knowledge units, and use the time decay function to calculate the fusion weights of the new knowledge units and the existing knowledge units; based on the fusion weights, perform weighted fusion on the new knowledge units and the existing knowledge units to obtain fused knowledge units;

[0164] When the similarity is lower than the preset similarity threshold, regard the new knowledge unit as a newly added knowledge unit;

[0165] Update the fused knowledge units to the corresponding positions, add the newly added knowledge units to the initial knowledge graph, update the initial knowledge graph, and record the timestamp of each update or addition.

[0166] According to the preset comprehensive scanning period, comprehensively scan the initial knowledge graph regularly, calculate the time decay value of each knowledge unit using the time decay function, and determine whether the time decay value of each knowledge unit is less than the preset decay threshold. When the time decay value corresponding to a knowledge unit is less than the preset decay threshold, remove the corresponding knowledge unit from the initial knowledge graph;

[0167] The initial knowledge graph is updated to form a dynamic knowledge graph.

[0168] The process of forming a dynamic knowledge graph by extracting knowledge from newly added preprocessed data based on the time decay-based dynamic update mechanism and supplementing it into the initial knowledge graph is as follows:

[0169] First, construct a time decay function, and use the exponential decay form to control the decay speed of knowledge units. Specifically, a benchmark decay rate α (0 < α < 1) can be set, where t represents the existence time of the knowledge unit, and the decay function can be expressed as f(t) = α^t. The smaller the value of α, the faster the decay speed. For example, α = 0.9 can be set, and the decay values after 1 day, 10 days, and 30 days are 0.9, 0.35, and 0.04 respectively.

[0170] After receiving the newly added preprocessed data, perform entity recognition and relationship extraction on it to obtain new knowledge units.

[0171] Then, use the cosine similarity method to calculate the similarity between the new knowledge units and the existing knowledge units in the initial knowledge graph. Specifically, represent the knowledge units in vector form and calculate the cosine value of the angle between the vectors as the similarity. For example, if the vectors of two knowledge units are (0.5, 0.8, 0.3) and (0.4, 0.7, 0.5) respectively, then their similarity is 0.97.

[0172] Judge whether the similarity is higher than the preset threshold according to the calculation result. If it is higher than the threshold, update the existing knowledge unit with the new knowledge unit. At the same time, use the time decay function to calculate the fusion weight, and the sum of the weights of the old and new knowledge units is 1. For example, if the existence times of the old and new knowledge units are 1 day and 10 days respectively, the weights can be set to 0.9 and 0.1. Perform weighted averaging on the old and new knowledge units based on the weights to obtain the fused knowledge unit.

[0173] If the similarity is lower than the threshold, directly add the new knowledge unit as a newly added knowledge unit. For example, if the similarity threshold is set to 0.8 and the calculated similarity is 0.7, then directly add the new knowledge unit to the knowledge graph.

[0174] Update the fused knowledge unit to the corresponding position, add the newly added knowledge unit to the initial knowledge graph, and record the time stamp of each update or addition. This can track the "age" of each knowledge unit.

[0175] According to the preset comprehensive scanning period (such as once a week), the initial knowledge graph is scanned comprehensively on a regular basis. The time decay value of each knowledge unit is calculated using a time decay function, and it is judged whether it is less than the preset decay threshold. For example, if the decay threshold is set to 0.1, and a certain knowledge unit has existed for 30 days with a decay value of 0.04, which is less than the threshold, then it is removed from the knowledge graph.

[0176] Through the above steps, the initial knowledge graph is continuously updated to form a dynamic knowledge graph. This mechanism can absorb new knowledge in a timely manner, while eliminating outdated knowledge, maintaining the timeliness and accuracy of the knowledge graph. This dynamic update mechanism based on time decay has the advantages of strong self - adaptability and good scalability. By adjusting parameters such as the decay function, similarity threshold, and decay threshold, the speed of knowledge update and elimination can be flexibly controlled to meet the needs of different fields and application scenarios. At the same time, this mechanism also provides effective technical support for the long - term maintenance and evolution of the knowledge graph.

[0177] In this embodiment, through similarity calculation and time decay function, new knowledge units are weighted - fused or newly added with existing knowledge units, realizing the dynamic update and refined processing of knowledge; adopting a time decay function in the form of exponential decay to reasonably control the decay speed of knowledge units, ensuring the timeliness and accuracy of knowledge in the knowledge graph; through regular scanning and time decay mechanism, outdated knowledge units are automatically removed to maintain the efficient management and dynamic adjustment of the knowledge graph; through continuous update and optimization, a dynamic knowledge graph that changes over time is formed to ensure the timeliness and practicality of the knowledge graph.

[0178] In an alternative implementation, a multi - modal problem input by the user is received. Based on the speech and images in the multi - modal problem, speech text, image text, and image semantic information are extracted. Combined with the pure text in the multi - modal problem, multi - modal feature fusion is performed to determine the fusion features. The fusion features are subjected to in - depth semantic understanding to obtain semantic features. The semantic features are input into an intent recognition model, and through semantic mapping, they are adapted to predefined notarization business intent categories, and the user intents are determined to include:

[0179] Receive a multi - modal problem, where the multi - modal problem contains speech information, image information, and text information;

[0180] Use a speech recognition model to convert the speech information in the multi - modal problem into a first text sequence; use an optical character recognition engine to process the image information in the multi - modal problem to obtain a second text sequence; splice the first text sequence, the second text sequence, and the text information in the multi - modal problem to form a combined text sequence, and map each token in the combined text sequence to a specified dimensional space through a first independent linear layer to obtain text features;

[0181] The semantic features of the image information in the multimodal question are extracted by using a vision Transformer model to obtain an image semantic vector, and the image semantic vector is mapped to the corresponding dimensional space identical to the text features through a second independent linear layer to obtain image features;

[0182] The multi-head attention mechanism is used to calculate the correlation between the text features and the image features to obtain an attention output. The attention output is subjected to residual connection and layer normalization with the text features to obtain a fused feature. The fused feature is input into a pre-trained bidirectional encoder to obtain a context-related representation; a token vector is extracted from the context-related representation and mapped to the task corresponding space through a linear layer and an activation function to obtain semantic features;

[0183] The semantic features are mapped to a predefined notarization service intent category space through a fully connected layer to obtain intent logic values. The sigmoid function is applied to the intent logic values to calculate the probability of each intent category, and the intent category with the highest probability is selected to determine the user intent.

[0184] This embodiment provides a notarization service intent recognition method based on multimodal input. The method first receives a multimodal question input including speech, image, and text. For the speech information, a pre-trained speech recognition model is used to convert it into a text sequence. Specifically, a speech recognition model based on Transformer, such as Wav2Vec 2.0, can be used. This model captures the context representation of speech through self-supervised learning and then is fine-tuned for the speech recognition task. For example, for the input speech "I want to handle property notarization", the model can accurately recognize and output the corresponding text sequence.

[0185] For the image information, an optical character recognition (OCR) engine is used to extract the text content in the image. An end-to-end OCR model based on a convolutional neural network and a recurrent neural network, such as CRNN (Convolutional Recurrent Neural Network), can be adopted. This model first uses a CNN to extract image features, then uses an RNN to model the feature sequence, and finally decodes through CTC (Connectionist Temporal Classification) to obtain a text result. For example, for a photo of a house property certificate, the OCR engine can recognize and output the text information on the certificate, such as "House Property Ownership Certificate", etc.

[0186] Concatenate the first text sequence obtained from speech recognition, the second text sequence obtained from OCR, and the plain text information in the original question to form a complete combined text sequence. Then, map each token in the sequence to a vector space of a specified dimension through an independent linear layer to obtain the text feature representation. For example, each token can be mapped to a 768-dimensional vector.

[0187] For the extraction of semantic information of images, a vision Transformer model such as ViT (Vision Transformer) is adopted. This model divides the input image into image patches of a fixed size, and through multiple layers of self-attention and feed-forward networks, captures the global context information of the image. The final output [CLS] token vector serves as the semantic representation of the entire image. Through another independent linear layer, map this semantic vector to the same dimensional space as the text features to obtain the image feature representation.

[0188] Next, use the multi-head attention mechanism to calculate the correlation between the text features and the image features. Multi-head attention allows the model to simultaneously focus on information in different subspaces, enhancing the ability of feature fusion. In specific implementation, 8 attention heads can be used, with the dimension of each head being 96. The output of the attention is connected with the original text features through a residual connection and undergoes layer normalization to obtain the fused feature representation.

[0189] Input the fused features into a pre-trained bidirectional encoder such as BERT (Bidirectional Encoder Representations from Transformers) to further extract the context-related semantic representation. The BERT model is pre-trained through a masked language model and a next sentence prediction task, and can effectively capture the bidirectional context information of the text. Extract the vector corresponding to the [CLS] token from the output of BERT as the sentence-level representation of the entire input.

[0190] Finally, map the above semantic features to a predefined notarization business intention category space through a fully connected layer. Assume that 10 notarization business intention categories are predefined, such as property notarization, marriage notarization, inheritance notarization, etc. The output dimension of the fully connected layer is 10, corresponding to the logit values of each intention category. Apply the sigmoid function to these logit values to convert them into probability values between 0 and 1. Select the category with the highest probability as the finally recognized user intention.

[0191] For example, for the input multimodal question "I want to handle property notarization", which includes voice, property ownership certificate pictures, and text, after the above processing, the model may output the following probability distribution: property notarization (0.92), marriage notarization (0.03), inheritance notarization (0.02), etc. At this time, the system will identify "property notarization" as the user's intention.

[0192] Through this method of multimodal feature fusion and deep semantic understanding, it is possible to make full use of the information in voice, images, and text, improving the accuracy and robustness of notarization business intention recognition. This method can flexibly handle various complex user input scenarios, providing a reliable intention understanding basis for subsequent notarization business processing.

[0193] In this embodiment, by converting voice, image, and text information into text features and image features, and using the multi-head attention mechanism to calculate the correlation, efficient fusion of different modal information is achieved, enhancing the understanding ability of multimodal questions; by extracting context-related representations through a pre-trained bidirectional encoder, it is possible to better capture semantic information and the connections between different modalities, improving the accuracy of problem processing; using the fused features to calculate semantic features, and calculating the probabilities of intention categories through a fully connected layer and a sigmoid function, it is possible to efficiently and accurately identify the user's intention, ensuring the intelligent processing of notarization business; realizing the comprehensive processing and analysis of voice, images, and text, enhancing the system's understanding and response ability to complex multimodal questions.

[0194] In an alternative embodiment, based on the user intention and combined with a dynamic knowledge graph, entity linking and relation extraction are performed on the semantic features to form a structured query expression, including:

[0195] Construct a pre-trained stacked network model, which adopts the training objective of permutation language modeling and uses a two-stream self-attention mechanism and a segmental recurrent mechanism;

[0196] Input the semantic features into the stacked network model to obtain a bidirectional context-aware word vector representation; construct a multi-task learning architecture at the top layer of the stacked network model, including an entity recognition task and a relation extraction task; introduce a relative position encoding mechanism into the multi-task learning architecture to capture long-distance dependencies; use a dynamic weight allocation strategy to adaptively adjust the importance weights of the entity recognition task and the relation extraction task in the multi-task learning architecture;

[0197] Integrate a dynamic knowledge graph, process the dynamic knowledge graph through a graph attention network, obtain enhanced entity and enhanced relation representations in the word vector representation, and form an enhanced word vector representation;

[0198] Apply adversarial training techniques to the enhanced word vector representation, train to obtain the final multi-task learning architecture, input semantic features, and identify the corresponding entity mentions; calculate the similarity between the entity mentions and the entities in the dynamic knowledge graph, select the entity with the highest similarity as the linking result to obtain a set of linked entities, and generate candidate relation pairs based on the entities in the set of linked entities;

[0199] Based on the relation extraction task in the multi-task learning architecture, use the semantic features and the candidate relation pairs as inputs to predict the relations between entity pairs and obtain a set of relations;

[0200] Select a corresponding query template based on the user intention, and fill the set of linked entities and the set of relations into the query template to generate a structured query expression.

[0201] The specific implementation of entity linking and relation extraction for semantic features based on the user intention in combination with the dynamic knowledge graph to form a structured query expression is as follows:

[0202] First, construct a pre-trained stacked network model. This model adopts the training objective of permutation language modeling and uses a two-stream self-attention mechanism and a segmental recurrence mechanism. Specifically, use a multi-layer Transformer structure as the backbone, and each layer of Transformer contains a multi-head self-attention layer and a feed-forward neural network layer. In the self-attention mechanism, a two-stream structure is adopted to calculate the attention scores of the content stream and the query stream respectively, and then the results of the two streams are fused. At the same time, a segmental recurrence mechanism is introduced to divide the long sequence into multiple segments, and recurrence calculations are performed within and between segments to capture long-distance dependencies.

[0203] Next, input the semantic features into this stacked network model to obtain a bi-directional context-aware word vector representation. Specifically, input each word token in the semantic features into the model in turn, and after the calculation of multiple layers of Transformer, obtain the contextualized representation of each word. Then, construct a multi-task learning architecture at the top layer of the stacked network model, including an entity recognition task and a relation extraction task. Introduce a relative position encoding mechanism into this multi-task learning architecture to capture long-distance dependency relationships. Specifically, on the basis of the original absolute position encoding, additionally add relative position encoding, calculate the relative distance between words, and incorporate this information into the calculation of self-attention. At the same time, use a dynamic weight allocation strategy to adaptively adjust the importance weights of the entity recognition task and the relation extraction task in the multi-task learning architecture. Dynamically adjust the weight ratio of the two tasks in the total loss according to the training losses of the two tasks.

[0204] Next, integrate the dynamic knowledge graph, process the dynamic knowledge graph through the graph attention network to obtain enhanced entity and enhanced relationship representations in the word vector representation, and form an enhanced word vector representation. Specifically, convert the entity and relationship information in the knowledge graph into a graph structure, use the graph attention network to perform message passing on the graph, and obtain the representation of each node. Then fuse these representations with the original word vectors to obtain an enhanced word vector representation.

[0205] Then, apply the adversarial training technique to the enhanced word vector representation to train and obtain the final multi-task learning architecture. During the training process, introduce adversarial samples, generate adversarial samples by adding perturbations to the word vectors, and improve the robustness of the model. Next, input semantic features and identify the corresponding entity mentions. Use the sequence annotation method to annotate each word in the input sequence to identify the boundaries of entity mentions.

[0206] After that, calculate the similarity between the entity mentions and the entities in the dynamic knowledge graph, select the entity with the highest similarity as the linking result, and obtain the linked entity set. Specifically, for each identified entity mention, calculate its similarity with all entities in the knowledge graph, and select the one with the highest similarity as the linking result. The similarity calculation can be based on the cosine similarity of word vectors.

[0207] Finally, based on the relation extraction task in the multi-task learning architecture, use the semantic features and candidate relation pairs as inputs to predict the relations between entity pairs and obtain the relation set. For each candidate relation pair, the model outputs a probability score indicating the likelihood of the relation holding. Select the relation with the highest probability as the prediction result.

[0208] Select the corresponding query template based on the user intention, fill the linked entity set and the relation set into the query template to generate a structured query expression. According to different user intentions, select the corresponding query template. For example, for the intention of asking about the capital, a template like "X is the capital of Y" can be selected. Then fill the identified and extracted entities and relations into the template slots to obtain the final structured query expression.

[0209] Through the above steps, entity linking and relation extraction based on user intention and dynamic knowledge graph are realized, converting natural language queries into structured query expressions, providing a basis for subsequent knowledge retrieval and question answering.

[0210] In an alternative embodiment, based on the user intention and the structured query expression, perform multi-hop reasoning in the dynamic knowledge graph, calculate the similarity using the graph vector, and combine the user intention to determine the reasoning path, and perform knowledge association reasoning along the reasoning path to obtain preliminary reasoning results including:

[0211] Receives the user's intention and a structured query expression as input, determines the starting entity for reasoning in the dynamic knowledge graph, uses the starting entity for reasoning as the current node, identifies the entities and relationships directly connected to the current node, and forms a candidate next-hop set;

[0212] Utilizes the pre-generated graph vectors to calculate the similarity scores between the current node and each node in the candidate next-hop set; converts the user's intention into a vector representation, fuses it with the graph vectors of each node in the candidate next-hop set to obtain a fused vector; combines the similarity scores with the fused vector to generate a comprehensive score for each candidate next-hop node;

[0213] Based on the comprehensive scores, uses a reinforcement learning agent to select the next hop, where the reinforcement learning agent takes the current state as input and outputs a probability distribution for selecting the next hop, and the current state includes the current node information, the user's intention, the comprehensive scores, and the selected path information;

[0214] Updates the current node and the selected path information according to the selected next hop;

[0215] Applies predefined reasoning rules to the selected path for rule-based explicit reasoning;

[0216] Uses a graph neural network model to encode the selected path for graph neural network-based implicit reasoning, where the graph neural network model is a graph convolutional network;

[0217] Fuses the results of explicit reasoning and the results of implicit reasoning to obtain a fused reasoning result;

[0218] Continues until reaching the preset hop count limit, records the path selections during the entire reasoning process to form a complete reasoning chain; calculates the confidence score of the entire reasoning result based on the length of the selected path, the similarity scores of each hop, and the reliability score of the reasoning rules used;

[0219] Combines the complete reasoning chain, the finally reached node, and the confidence score to form a preliminary reasoning result.

[0220] In this embodiment, through the node selection and path update mechanism in the dynamic knowledge graph, combined with the user intention and the structured query expression, it is possible to quickly identify the appropriate reasoning path, improving the reasoning efficiency; by fusing the graph vector similarity and the user intention to generate a comprehensive score, the accuracy of candidate node selection is ensured, improving the decision-making quality during the reasoning process; through the combination of rule-based reasoning and graph convolutional network for explicit and implicit reasoning, a more comprehensive reasoning ability is achieved, capable of handling complex relationships and logics; using a reinforcement learning agent to select the next-hop node and optimizing based on the comprehensive score and the selected path, ensuring that the reasoning process has an efficient decision-making mechanism; calculating the confidence of the overall reasoning result through the path length, similarity score, and reliability score of the reasoning rule, ensuring the reliability and interpretability of the reasoning result; recording the reasoning path to form a complete reasoning chain, ensuring the traceability and transparency of the reasoning process.

[0221] The specific implementation method of multi-hop reasoning in the dynamic knowledge graph according to the user intention and the structured query expression is as follows:

[0222] First, the system receives the intention and the structured query expression input by the user. For example, the user intention may be "query movies starring a certain actor", and the structured query expression may be "QUERY(actor, starring, movie)". The system determines the starting entity for reasoning in the dynamic knowledge graph based on these inputs.

[0223] Next, the system uses the pre-generated graph vectors to calculate the similarity score between the current node and each node in the candidate next-hop set. At the same time, the user intention "query movies starring a certain actor" is transformed into a vector representation and fused with the graph vectors of the candidate next-hop nodes to obtain a fused vector. Combining the similarity score with the fused vector, a comprehensive score for each candidate next-hop node is generated.

[0224] Then, the system uses a reinforcement learning agent to select the next-hop. The reinforcement learning agent takes the current state as input, including the current node, user intention, comprehensive score, and selected path information, and outputs the probability distribution of selecting the next-hop.

[0225] Apply the predefined reasoning rules to the selected path for rule-based explicit reasoning. For example, apply the rule "if A stars in B, then B is a movie".

[0226] At the same time, use a graph neural network model (such as a graph convolutional network) to encode the selected path for graph neural network-based implicit reasoning. The graph neural network may capture some implicit patterns.

[0227] Fuse the results of explicit reasoning and implicit reasoning to obtain a fused reasoning result.

[0228] The system continues multi-hop reasoning until it reaches a preset hop limit (e.g., 3 hops). During the whole process, the path selection is recorded to form a complete reasoning chain.

[0229] Based on the length of the selected path, the similarity score of each hop, and the reliability of the reasoning rules used, the system calculates the confidence score of the entire reasoning result. For example, if the path is short, the similarity of each hop is high, and the rules used are reliable, a higher confidence score is given.

[0230] Finally, the system combines the complete reasoning chain, the finally reached node (e.g., "2004"), and the confidence score to form a preliminary reasoning result.

[0231] Through this method, the system can perform multi-hop reasoning in a dynamic knowledge graph, combine the user's intention and the graph structure, find relevant knowledge, and give the reasoning process and the confidence of the result, providing a basis for subsequent question answering or decision making.

[0232] Figure 2 This is a schematic structural diagram of the notarization intelligent question answering customer service system based on the knowledge graph according to the embodiment of the present invention, as Figure 2 shown, the system includes:

[0233] The first unit is used to collect multi-source heterogeneous data in the notarization field, preprocess the multi-source heterogeneous data to obtain preprocessed data, extract professional terms and domain concepts from the preprocessed data by using a named entity recognition algorithm to form an entity set, analyze the relationships between entities in the entity set based on semantic role annotation technology to construct a relationship set, combine the entity set and the relationship set to form an initial knowledge graph, extract knowledge from the newly added preprocessed data based on a time-decay dynamic update mechanism and supplement it to the initial knowledge graph to form a dynamic knowledge graph, and perform vector representation on the dynamic knowledge graph to obtain a graph vector;

[0234] The second unit is used to receive a multi-modal question input by a user, extract speech text, image text, and image semantic information based on the speech and image in the multi-modal question, combine the pure text in the multi-modal question to perform multi-modal feature fusion to determine fusion features, perform in-depth semantic understanding on the fusion features to obtain semantic features, input the semantic features into an intention recognition model, and through semantic mapping, adapt to predefined notarization business intention categories to determine the user's intention. Based on the user's intention and combined with the dynamic knowledge graph, entity linking and relationship extraction are performed on the semantic features to form a structured query expression;

[0235] A third unit is used to perform multi-hop reasoning in a dynamic knowledge graph based on user intent and a structured query expression, calculate similarity using graph vectors, and combine with the user intent to determine an inference path. Knowledge association reasoning is carried out along the inference path to obtain a preliminary inference result. An interpretability analysis is performed on the preliminary inference result to generate an inference chain and a confidence score, forming a detailed inference result. Combining the pre-obtained user profile information and user intent, a personalized ranking is performed on the detailed inference result to obtain an ordered inference result. The ordered inference result is input into a pre-constructed Transformer model to output a preliminary natural language answer, and a text style transfer technique is applied to finally generate a personalized answer.

[0236] In a third aspect of the embodiments of the present invention,

[0237] a kind of electronic device is provided, including:

[0238] a processor;

[0239] a memory for storing instructions executable by the processor;

[0240] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0241] In a fourth aspect of the embodiments of the present invention,

[0242] a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0243] The present invention can be a method, a device, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.

[0244] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A notarization intelligent question - answering customer service method based on a knowledge graph, characterized in that, Including: Collecting multi-source heterogeneous data in the notarization field, preprocessing the multi-source heterogeneous data to obtain preprocessed data, using a named entity recognition algorithm to extract professional terms and domain concepts from the preprocessed data to form an entity set, analyzing the relationships between entities in the entity set based on semantic role labeling technology to construct a relationship set, combining the entity set and the relationship set to form an initial knowledge graph, based on a time-decay dynamic update mechanism, extracting knowledge from newly added preprocessed data and supplementing it into the initial knowledge graph to form a dynamic knowledge graph, and performing vector representation on the dynamic knowledge graph to obtain a graph vector; Receiving a multi-modal question input by a user, extracting speech text, image text, and image semantic information based on the speech and image in the multi-modal question, combining the pure text in the multi-modal question, performing multi-modal feature fusion to determine fusion features, performing deep semantic understanding on the fusion features to obtain semantic features, inputting the semantic features into an intent recognition model, and through semantic mapping, adapting to predefined notarization business intent categories to determine the user intent. Based on the user intent and the dynamic knowledge graph, entity linking and relationship extraction are performed on the semantic features to form a structured query expression; Based on the user intent and the structured query expression, performing multi-hop reasoning in the dynamic knowledge graph, calculating similarity using the graph vector, and combining the user intent to determine an inference path, performing knowledge association reasoning along the inference path to obtain a preliminary inference result, performing interpretability analysis on the preliminary inference result to generate an inference chain and a confidence score to form a detailed inference result, combining the pre-obtained user portrait information and the user intent, performing personalized ranking on the detailed inference result to obtain an ordered inference result, inputting the ordered inference result into a pre-constructed Transformer model, outputting a preliminary natural language answer, and applying text style transfer technology to finally generate a personalized answer.

2. The method according to claim 1, wherein Using a named entity recognition algorithm to extract professional terms and domain concepts from the preprocessed data to form an entity set, analyzing the relationships between entities in the entity set based on semantic role labeling technology to construct a relationship set, and combining the entity set and the relationship set to form an initial knowledge graph includes: Constructing a notarization field annotation data set, training a notarization entity recognition deep fusion network based on the annotation data set. The notarization entity recognition deep fusion network includes a BERT-based encoding layer, a feature extraction layer based on a bidirectional long short-term memory network, and a decoding layer based on a conditional random field. Performing named entity recognition on notarization field text through the notarization entity recognition deep fusion network to obtain an entity set, where the entity set includes entity names, entity types, occurrence frequencies, and first occurrence positions; Performing dependency syntactic analysis on sentences containing entities in the entity set to identify the subject-predicate-object structure, and using a semantic role labeling model to perform semantic role labeling on sentences containing the subject-predicate-object structure to obtain a predicate-argument structure and determine semantic role labeling tags; Construct a mapping rule set from semantic role labeling tags to relationship types in the notarization field. According to the mapping rule set, convert the semantic role labeling results into candidate relationships; Construct a relationship classifier and train it through a preset remotely supervised dataset. Use the relationship classifier to score and filter the candidate relationships to obtain a relationship set; Import the entities in the entity set as entity nodes into the graph database, import the relationships in the relationship set as relationship edges into the graph database, connect the corresponding entity nodes, and use graph algorithms to optimize the imported entity nodes and relationship edges to eliminate redundant relationships and merge synonymous entities to construct an initial knowledge graph.

3. The method according to claim 2, characterized in that, Construct a notarization field annotation dataset and train a deep fusion network for notarization entity recognition based on the annotation dataset. The deep fusion network for notarization entity recognition includes a BERT-based encoding layer, a feature extraction layer based on a bidirectional long short-term memory network, and a decoding layer based on a conditional random field. Perform named entity recognition on notarization field texts through the deep fusion network for notarization entity recognition to obtain an entity set including: Construct a notarization field annotation dataset and perform masked language model pre-training on the BERT model based on the notarization field annotation dataset to obtain a pre-trained BERT model adapted to the notarization field; In the deep fusion network for notarization entity recognition, initialize the encoding layer with the weights of the pre-trained BERT model; input the notarization field annotation dataset into the deep fusion network for notarization entity recognition, and obtain word vector representations through the encoding layer; Input the word vector representations into the feature extraction layer to extract sequence features, where the feature extraction layer includes a forward long short-term memory network and a backward long short-term memory network, each containing an input gate, a forget gate, and an output gate; Input the sequence features into the decoding layer. The decoding layer includes a linear chain conditional random field, establish a label transition probability matrix, and use the Viterbi algorithm to decode to obtain the optimal label sequence; Adopt a joint learning strategy to simultaneously optimize the encoding layer, the feature extraction layer, and the decoding layer, use the AdamW optimizer for parameter update, and implement a learning rate warm-up and linear decay strategy; After each training epoch, evaluate the performance of the deep fusion network for notarization entity recognition using a preset validation set, and combine an early stopping strategy to save the model weights with the highest F1 score on the validation set to obtain a trained deep fusion network for notarization entity recognition; Input the notarization field text to be recognized into the trained deep fusion network for notarization entity recognition, and successively pass through the encoding layer, the feature extraction layer, and the decoding layer to obtain a label sequence; According to the label sequence, identify the entities in the notarization field text and construct an entity set.

4. The method according to claim 1, wherein Based on a dynamic update mechanism with time decay, extract knowledge from newly added preprocessed data and supplement it to the initial knowledge graph to form a dynamic knowledge graph including: Construct a time decay function, and the time decay function adopts an exponential decay form to control the decay speed of knowledge units; Receive the newly added preprocessed data, perform entity recognition and relationship extraction on the newly added preprocessed data to obtain new knowledge units, use the cosine similarity method to calculate the similarity between the new knowledge units and the existing knowledge units in the initial knowledge graph, obtain the similarity calculation result, and according to the similarity calculation result, judge whether the similarity calculation result is higher than a preset similarity threshold: When the similarity calculation result is higher than the preset similarity threshold, update the existing knowledge units with the new knowledge units, and use the time decay function to calculate the fusion weights of the new knowledge units and the existing knowledge units; based on the fusion weights, perform weighted fusion on the new knowledge units and the existing knowledge units to obtain fused knowledge units; When the similarity is lower than the preset similarity threshold, regard the new knowledge unit as a newly added knowledge unit; Update the fused knowledge units to the corresponding positions, add the newly added knowledge units to the initial knowledge graph, update the initial knowledge graph, and record the time stamp of each update or addition; According to a preset full scan period, periodically perform a full scan on the initial knowledge graph, use the time decay function to calculate the time decay value of each knowledge unit, and judge whether the time decay value of each knowledge unit is less than a preset decay threshold. When the time decay value corresponding to a knowledge unit is less than the preset decay threshold, remove the corresponding knowledge unit from the initial knowledge graph; The initial knowledge graph is updated to form a dynamic knowledge graph.

5. The method according to claim 1, characterized in that, Receive the multimodal question input by the user. Based on the speech and image in the multimodal question, extract the speech text, image text and image semantic information, combine the pure text in the multimodal question, perform multimodal feature fusion, determine the fusion features, perform deep semantic understanding on the fusion features to obtain semantic features, input the semantic features into the intent recognition model, and through semantic mapping, adapt to the predefined notarization service intent categories to determine the user intent including: Receive a multimodal question, where the multimodal question contains speech information, image information and text information; Use a speech recognition model to convert the speech information in the multimodal question into a first text sequence; process the image information in the multimodal question using an optical character recognition engine to obtain a second text sequence; splice the first text sequence, the second text sequence and the text information in the multimodal question to form a combined text sequence, and map each token in the combined text sequence to a specified dimensional space through a first independent linear layer to obtain text features; Use a vision Transformer model to extract the semantic features of the image information in the multimodal question to obtain an image semantic vector, and map the image semantic vector to the same corresponding dimensional space as the text features through a second independent linear layer to obtain image features; Calculate the correlation between the text features and the image features using the multi-head attention mechanism to obtain the attention output. Perform residual connection and layer normalization on the attention output and the text features to obtain the fused features. Input the fused features into a pre-trained bidirectional encoder to obtain the context-associated representation; Extract the token vectors from the context-associated representation, map them to the task-corresponding space through a linear layer and an activation function to obtain the semantic features; Map the semantic features to a predefined notarization business intention category space through a fully connected layer to obtain the intention logic values. Apply the sigmoid function to the intention logic values to calculate the probabilities of each intention category, and select the intention category with the highest probability to determine the user intention.

6. The method according to claim 5, characterized in that Based on the user intention and combined with the dynamic knowledge graph, perform entity linking and relation extraction on the semantic features to form a structured query expression including: Construct a pre-trained stacked network model. The stacked network model adopts the training objective of permutation language modeling and uses a two-stream self-attention mechanism and a segmental recurrent mechanism; Input the semantic features into the stacked network model to obtain the bidirectional context-aware word vector representation; Construct a multi-task learning architecture at the top layer of the stacked network model, including an entity recognition task and a relation extraction task; Introduce a relative position encoding mechanism into the multi-task learning architecture to capture long-distance dependencies; Use a dynamic weight allocation strategy to adaptively adjust the importance weights of the entity recognition task and the relation extraction task in the multi-task learning architecture; Integrate the dynamic knowledge graph, process the dynamic knowledge graph through a graph attention network to obtain the enhanced entity and enhanced relation representations in the word vector representation, and form the enhanced word vector representation; Apply adversarial training techniques to the enhanced word vector representation, train to obtain the final multi-task learning architecture, input the semantic features, and identify the corresponding entity mentions; Calculate the similarity between the entity mentions and the entities in the dynamic knowledge graph, select the entity with the highest similarity as the linking result to obtain the linked entity set, and generate candidate relation pairs based on the entities in the linked entity set; Based on the relation extraction task in the multi-task learning architecture, use the semantic features and the candidate relation pairs as inputs to predict the relations between entity pairs to obtain the relation set; Select the corresponding query template based on the user intention, and fill the linked entity set and the relation set into the query template to generate a structured query expression.

7. The method according to claim 1, wherein Based on the user intention and the structured query expression, perform multi-hop reasoning in the dynamic knowledge graph, calculate the similarity using the graph vectors, and combine with the user intention to determine the reasoning path, and perform knowledge association reasoning along the reasoning path to obtain the preliminary reasoning result including: Receive the user intention and the structured query expression as inputs, determine the reasoning start entity in the dynamic knowledge graph, use the reasoning start entity as the current node, and identify the entities and relations directly connected to the current node to form a candidate next-hop set; Calculate the similarity scores between the current node and each node in the candidate next-hop set using pre-generated graph vectors; convert the user intention into a vector representation, fuse it with the graph vectors of each node in the candidate next-hop set to obtain fused vectors; combine the similarity scores with the fused vectors to generate a comprehensive score for each candidate next-hop node. Based on the comprehensive scores, use a reinforcement learning agent to select the next hop. The reinforcement learning agent takes the current state as input and outputs a probability distribution for selecting the next hop. The current state includes the current node information, the user intention, the comprehensive scores, and the selected path information. Update the current node and the selected path information according to the selected next hop. Apply predefined inference rules to the selected path for rule-based explicit reasoning. Use a graph neural network model to encode the selected path for graph neural network-based implicit reasoning, where the graph neural network model is a graph convolutional network. Fuse the results of explicit reasoning and implicit reasoning to obtain a fused reasoning result. Until the preset hop count limit is reached, record the path selections during the entire reasoning process to form a complete reasoning chain; calculate the confidence score of the entire reasoning result based on the length of the selected path, the similarity scores of each hop, and the reliability score of the used inference rules. Combine the complete reasoning chain, the finally reached node, and the confidence score to form a preliminary reasoning result.

8. A notarization intelligent question-answering customer service system based on a knowledge graph, for implementing the method described in any one of the foregoing claims 1-7, characterized in that, Including: The first unit is used to collect multi-source heterogeneous data in the notarization field, preprocess the multi-source heterogeneous data to obtain preprocessed data, use a named entity recognition algorithm to extract professional terms and domain concepts from the preprocessed data to form an entity set, analyze the relationships between entities in the entity set based on semantic role annotation technology to construct a relationship set, combine the entity set and the relationship set to form an initial knowledge graph, based on a time-decaying dynamic update mechanism, extract knowledge from newly added preprocessed data and supplement it to the initial knowledge graph to form a dynamic knowledge graph, and perform vector representation on the dynamic knowledge graph to obtain graph vectors. The second unit is used to receive a multi-modal question input by the user, extract speech text, image text, and image semantic information based on the speech and image in the multi-modal question, combine the pure text in the multi-modal question for multi-modal feature fusion to determine the fusion features, perform deep semantic understanding on the fusion features to obtain semantic features, input the semantic features into an intention recognition model, and through semantic mapping, adapt to predefined notarization business intention categories to determine the user intention. Based on the user intention and combined with the dynamic knowledge graph, perform entity linking and relationship extraction on the semantic features to form a structured query expression. The third unit is used to perform multi-hop reasoning in the dynamic knowledge graph based on the user intention and the structured query expression, calculate the similarity using the graph vector, and combine the user intention to determine the reasoning path. Conduct knowledge association reasoning along the reasoning path to obtain the preliminary reasoning result, perform interpretability analysis on the preliminary reasoning result, generate the reasoning chain and confidence score to form the detailed reasoning result. Combine the pre-obtained user portrait information and user intention to perform personalized ranking on the detailed reasoning result to obtain the ordered reasoning result. Input the ordered reasoning result into the pre-constructed Transformer model to output the preliminary natural language answer, and apply the text style transfer technology to finally generate the personalized answer.

9. An electronic device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Data intelligent question and answer method and system fusing domain knowledge

    CN118779438A

  • DCS intelligent decision-making method and system fusing large language model and knowledge graph

    CN118820778A