Knotarization intelligent question and answer customer service method and system based on knowledge graph

By using a knowledge graph-based method in the notarized intelligent customer service system to build and dynamically update the knowledge graph, the shortcomings of the existing system in dealing with complex problems and knowledge updates are solved, and more efficient and accurate notarization consulting services are achieved.

CN119938816AActive Publication Date: 2025-05-06SUZHOU LIANZHENG INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202411751063.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-05-06
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

The existing notarized intelligent customer service system has limited understanding ability when dealing with complex or vague user problems, rigid answering methods, lack contextual understanding and multi-round dialogue capabilities, and it is difficult to deal with complex notarization problems. It is difficult to update and maintain knowledge bases, and it is difficult to timely reflect the latest changes in regulations and policies.

Method used

The notarized intelligent Q&A customer service method based on knowledge graph is adopted to construct the initial knowledge graph through naming entity recognition and semantic role labeling technology, and the knowledge graph is dynamically updated using the time attenuation mechanism. The system can handle multimodal problems, conduct in-depth semantic understanding, entity linking and relationship extraction, generate structured query expressions, and perform multi-hop inference in the dynamic knowledge graph to generate personalized answers.

Benefits of technology

It improves the accuracy and depth of Q&A, can handle complex notarization issues, provide more comprehensive and relevant information, timely updates the knowledge base, adapts to the latest changes in regulations and policies, and improves user experience and service efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938816A_ABST
    Figure CN119938816A_ABST
Patent Text Reader

Abstract

The invention provides a notarization intelligent question and answer customer service method and system based on a knowledge graph, and relates to the technical field of information, and the method comprises the steps: collecting multi-source heterogeneous data in the notarization field, forming an entity set and a relationship set, combining to form an initial knowledge graph, carrying out a dynamic update mechanism based on time decay, and supplementing to form a dynamic knowledge graph. Carrying out vectorization expression to obtain an atlas vector; receiving a multi-modal problem, determining fusion features, performing deep semantic understanding, obtaining semantic features, determining a user intention, and forming a structured query expression; based on the user intention and the structured query expression, carrying out multi-hop reasoning in the dynamic knowledge graph, carrying out similarity calculation by utilizing graph vectors, determining a reasoning path in combination with the user intention, carrying out knowledge association reasoning along the reasoning path, obtaining a preliminary reasoning result, carrying out interpretability analysis on the preliminary reasoning result, and obtaining a final reasoning result; and generating a reasoning chain and a confidence score, forming a detailed reasoning result, and finally generating a personalized answer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a notarized intelligent question-answering customer service method and system based on a knowledge graph. Background Art

[0002] With the continuous advancement of construction and the rapid development of social economy, the demand for notarial services has shown a trend of sustained growth. Notarial business involves multiple fields such as civil, commercial, and foreign-related, and is characterized by strong professionalism and high policy. Traditional notarial consulting services mainly rely on manual reception and telephone consultation. This model, in the face of the growing demand for notarial services, exposes problems such as insufficient human resources, limited service time, and lagging knowledge updates. In order to improve service efficiency and quality, some notarial agencies have begun to try to introduce intelligent customer service systems to meet the public's demand for all-weather, efficient, and professional notarial consulting services.

[0003] The intelligent customer service systems currently used in the notarization field are mainly based on keyword matching and preset question-answer pairs. Although they have improved service efficiency to a certain extent, they still have many limitations. First, the system has limited ability to understand user questions and it is difficult to accurately grasp the true intentions behind complex or ambiguous statements. Secondly, the preset answer methods are relatively rigid and cannot flexibly respond to diverse query needs. Furthermore, such systems lack contextual understanding and multi-round dialogue capabilities, making it difficult to handle complex notarization issues that require in-depth interaction. In addition, the knowledge base is difficult to update and maintain, and it is difficult to reflect the latest changes in laws and regulations in a timely manner.

[0004] In summary, it is urgent to develop a notarization intelligent question-answering customer service method based on knowledge graph. This method can make full use of the advantages of knowledge graph to achieve structured representation and efficient management of complex knowledge systems in the notarization field. The system can handle complex queries that require multi-step logical reasoning, greatly improving the accuracy and depth of questions and answers. At the same time, by utilizing the relevance of the graph, more comprehensive and relevant information can be provided to users. The present invention can solve the problems in the prior art. Summary of the invention

[0005] The embodiments of the present invention provide a notarized intelligent question-answering customer service method and system based on knowledge graph, which can solve the problems in the prior art.

[0006] According to a first aspect of the embodiments of the present invention,

[0007] A notarized intelligent question-answering customer service method based on knowledge graph is provided, including:

[0008] Collect multi-source heterogeneous data in the field of notarization, pre-process the multi-source heterogeneous data to obtain pre-processed data, extract professional terms and domain concepts from the pre-processed data using a named entity recognition algorithm to form an entity set, analyze the relationship between entities in the entity set based on semantic role labeling technology, construct a relationship set, combine the entity set and the relationship set to form an initial knowledge graph, extract knowledge from the newly added pre-processed data and add it to the initial knowledge graph based on a dynamic update mechanism of time decay to form a dynamic knowledge graph, and vectorize the dynamic knowledge graph to obtain a graph vector;

[0009] Receive a multimodal question input by a user, extract the voice text, image text and image semantic information based on the voice and image in the multimodal question, combine the plain text in the multimodal question, perform multimodal feature fusion, determine the fusion feature, perform deep semantic understanding on the fusion feature to obtain the semantic feature, input the semantic feature into the intent recognition model, adapt the predefined notarization business intent category through semantic mapping, determine the user intent, perform entity linking and relationship extraction on the semantic feature based on the user intent combined with the dynamic knowledge graph, and form a structured query expression;

[0010] Based on user intent and structured query expressions, multi-hop reasoning is performed in the dynamic knowledge graph, and similarity is calculated using graph vectors. Combined with user intent, the reasoning path is determined, and knowledge association reasoning is performed along the reasoning path to obtain preliminary reasoning results. The preliminary reasoning results are analyzed for interpretability, and reasoning chains and confidence scores are generated to form detailed reasoning results. Combined with the pre-acquired user portrait information and user intent, the detailed reasoning results are personalized and sorted to obtain ordered reasoning results. The ordered reasoning results are input into the pre-built Transformer model, preliminary natural language answers are output, and text style transfer technology is applied to finally generate personalized answers.

[0011] In an optional embodiment,

[0012] Using a named entity recognition algorithm to extract professional terms and domain concepts from the preprocessed data to form an entity set, analyzing the relationship between entities in the entity set based on a semantic role labeling technique, constructing a relationship set, and combining the entity set and the relationship set to form an initial knowledge graph includes:

[0013] Construct a notarization field annotated dataset, and train a notarization entity recognition deep fusion network based on the annotated dataset. The notarization entity recognition deep fusion network includes a BERT-based encoding layer, a bidirectional long short-term memory network-based feature extraction layer, and a conditional random field-based decoding layer. The notarization field text is subjected to named entity recognition through the notarization entity recognition deep fusion network to obtain an entity set, which includes entity name, entity type, occurrence frequency, and first occurrence position.

[0014] Performing dependency syntactic analysis on sentences containing entities in the entity set to identify subject-predicate-object structures, using a semantic role labeling model to perform semantic role labeling on sentences containing the subject-predicate-object structures to obtain a predicate-argument structure, and determining a semantic role labeling label;

[0015] Constructing a mapping rule set from semantic role annotation labels to notarization domain relationship types, and converting semantic role annotation results into candidate relationships according to the mapping rule set;

[0016] Constructing a relationship classifier and training it with a preset remote supervision data set, using the relationship classifier to score and filter candidate relationships to obtain a relationship set;

[0017] The entities in the entity set are imported as entity nodes into the graph database, and the relationships in the relationship set are imported as relationship edges into the graph database, the corresponding entity nodes are connected, and the imported entity nodes and relationship edges are optimized using a graph algorithm to eliminate redundant relationships and merge synonymous entities, so as to construct an initial knowledge graph.

[0018] In an optional embodiment,

[0019] A notarization field annotated dataset is constructed, and a notarization entity recognition deep fusion network is trained based on the annotated dataset. The notarization entity recognition deep fusion network includes a BERT-based encoding layer, a bidirectional long short-term memory network-based feature extraction layer, and a conditional random field-based decoding layer. Named entity recognition is performed on notarization field texts through the notarization entity recognition deep fusion network, and the entity set obtained includes:

[0020] Constructing a notarization field annotated dataset, and pre-training the BERT model with a masked language model based on the notarization field annotated dataset to obtain a pre-trained BERT model adapted to the notarization field;

[0021] In the notarization entity recognition deep fusion network, the encoding layer is initialized using the weights of the pre-trained BERT model; the notarization field annotated dataset is input into the notarization entity recognition deep fusion network, and the word vector representation is obtained through the encoding layer;

[0022] Input the word vector representation into a feature extraction layer to extract sequence features, wherein the feature extraction layer includes a forward long short-term memory network and a backward long short-term memory network, each including an input gate, a forget gate, and an output gate;

[0023] Inputting the sequence features into the decoding layer, the decoding layer includes a linear chain conditional random field, establishing a label transfer probability matrix, and using the Viterbi algorithm to decode to obtain an optimal label sequence;

[0024] A joint learning strategy is adopted to optimize the encoding layer, feature extraction layer, and decoding layer at the same time. The AdamW optimizer is used to update the parameters, and the learning rate warm-up and linear decay strategies are implemented.

[0025] After each training round, the performance of the publicized entity recognition deep fusion network is evaluated using a preset validation set, and the model weight with the highest F1 score on the validation set is saved in combination with an early stopping strategy to obtain a trained publicized entity recognition deep fusion network;

[0026] Input the notarization domain text to be recognized into the trained notarization entity recognition deep fusion network, and sequentially pass through the encoding layer, feature extraction layer and decoding layer to obtain a label sequence;

[0027] According to the tag sequence, entities in the notarization domain text are identified and an entity set is constructed.

[0028] In an optional embodiment,

[0029] Based on the dynamic update mechanism of time decay, knowledge is extracted from the newly added pre-processed data and added to the initial knowledge graph to form a dynamic knowledge graph including:

[0030] Constructing a time decay function, wherein the time decay function adopts an exponential decay form to control the decay speed of the knowledge unit;

[0031] Receive newly added preprocessed data, and perform entity recognition and relationship extraction on the newly added preprocessed data to obtain a new knowledge unit, use the cosine similarity method to calculate the similarity between the new knowledge unit and the existing knowledge unit in the initial knowledge graph, and obtain a similarity calculation result, and determine whether the similarity calculation result is higher than a preset similarity threshold based on the similarity calculation result:

[0032] When the similarity calculation result is higher than the preset similarity threshold, the existing knowledge unit is updated with the new knowledge unit, and the fusion weight of the new knowledge unit and the existing knowledge unit is calculated using the time decay function; based on the fusion weight, the new knowledge unit and the existing knowledge unit are weightedly fused to obtain a fused knowledge unit;

[0033] When the similarity is lower than the preset similarity threshold, the new knowledge unit is used as a newly added knowledge unit;

[0034] Update the fused knowledge unit to the corresponding position, add the newly added knowledge unit to the initial knowledge graph, update the initial knowledge graph, and record the timestamp of each update or addition;

[0035] According to a preset full scan cycle, the initial knowledge graph is regularly scanned, the time decay value of each knowledge unit is calculated using the time decay function, and it is determined whether the time decay value of each knowledge unit is less than a preset decay threshold value. When the time decay value corresponding to a knowledge unit is less than the preset decay threshold value, the corresponding knowledge unit is removed from the initial knowledge graph;

[0036] The initial knowledge graph is updated to form a dynamic knowledge graph.

[0037] In an optional embodiment,

[0038] Receive multimodal questions input by users, extract voice text, image text and image semantic information based on the voice and image in the multimodal questions, combine with the plain text in the multimodal questions, perform multimodal feature fusion, determine the fusion features, perform deep semantic understanding on the fusion features, obtain semantic features, input the semantic features into the intent recognition model, adapt the predefined notarization business intent categories through semantic mapping, and determine that the user intent includes:

[0039] Receiving a multimodal question, wherein the multimodal question includes voice information, image information, and text information;

[0040] Using a speech recognition model to convert speech information in the multimodal problem into a first text sequence; using an optical character recognition engine to process image information in the multimodal problem to obtain a second text sequence; splicing the first text sequence, the second text sequence and the text information in the multimodal problem to form a combined text sequence, and mapping each word in the combined text sequence to a specified dimensional space through a first independent linear layer to obtain text features;

[0041] Extracting semantic features of the image information in the multimodal problem using a visual Transformer model to obtain an image semantic vector, and mapping the image semantic vector to a corresponding dimensional space having the same dimension as the text feature through a second independent linear layer to obtain an image feature;

[0042] Using a multi-head attention mechanism to calculate the correlation between the text feature and the image feature to obtain an attention output, performing a residual connection and layer normalization on the attention output and the text feature to obtain a fused feature, and inputting the fused feature into a pre-trained bidirectional encoder to obtain a context-related representation; extracting a tag vector from the context-related representation, and mapping it to a task-corresponding space through a linear layer and an activation function to obtain a semantic feature;

[0043] The semantic features are mapped to the predefined notarization business intention category space through a fully connected layer to obtain the intention logic value, the sigmoid function is applied to the intention logic value, the probability of each intention category is calculated, the intention category with the highest probability is selected, and the user intention is determined.

[0044] In an optional embodiment,

[0045] Based on the user intention and the dynamic knowledge graph, the semantic features are entity linked and relations are extracted to form a structured query expression including:

[0046] Constructing a pre-trained stacked network model, wherein the stacked network model adopts a permutation language modeling training objective, uses a two-stream self-attention mechanism and a segmented recurrence mechanism;

[0047] Inputting semantic features into the stacked network model to obtain bidirectional context-aware word vector representations; constructing a multi-task learning architecture on the top layer of the stacked network model, including entity recognition tasks and relationship extraction tasks; introducing a relative position encoding mechanism into the multi-task learning architecture to capture long-distance dependencies; using a dynamic weight allocation strategy to adaptively adjust the importance weights of entity recognition tasks and relationship extraction tasks in the multi-task learning architecture;

[0048] Integrate a dynamic knowledge graph, process the dynamic knowledge graph through a graph attention network, obtain enhanced entity and enhanced relationship representations in the word vector representation, and form an enhanced word vector representation;

[0049] Apply adversarial training technology to the enhanced word vector representation, train to obtain the final multi-task learning architecture, input semantic features, and identify corresponding entity mentions; calculate the similarity between the entity mentions and the entities in the dynamic knowledge graph, select the entity with the highest similarity as the link result, obtain a linked entity set, and generate candidate relationship pairs based on the entity pairs in the linked entity set;

[0050] Based on the relationship extraction task in the multi-task learning architecture, taking the semantic features and the candidate relationship pairs as input, predicting the relationship between entity pairs to obtain a relationship set;

[0051] A corresponding query template is selected based on the user's intention, and the linked entity set and the relationship set are filled into the query template to generate a structured query expression.

[0052] In an optional embodiment,

[0053] Based on user intent and structured query expressions, multi-hop reasoning is performed in the dynamic knowledge graph, similarity is calculated using graph vectors, and the reasoning path is determined in combination with user intent. Knowledge association reasoning is performed along the reasoning path, and preliminary reasoning results are obtained, including:

[0054] Receive user intent and structured query expressions as input, determine the inference starting entity in the dynamic knowledge graph, take the inference starting entity as the current node, identify entities and relationships directly connected to the current node, and form a candidate next hop set;

[0055] Using the pre-generated graph vector, calculate the similarity score between the current node and each node in the candidate next hop set; convert the user intention into a vector representation, and fuse it with the graph vector of each node in the candidate next hop set to obtain a fusion vector; combine the similarity score with the fusion vector to generate a comprehensive score for each candidate next hop node;

[0056] Based on the comprehensive score, a next hop is selected using a reinforcement learning agent, wherein the reinforcement learning agent takes a current state as input and outputs a probability distribution of selecting the next hop, wherein the current state includes the current node information, the user intention, the comprehensive score, and the selected path information;

[0057] According to the selected next hop, update the current node and the selected path information;

[0058] Applying predefined reasoning rules to the selected path to perform rule-based explicit reasoning;

[0059] Encoding the selected path using a graph neural network model and performing implicit reasoning based on the graph neural network, wherein the graph neural network model is a graph convolutional network;

[0060] The results of explicit reasoning and implicit reasoning are integrated to obtain a fused reasoning result;

[0061] Until the preset hop limit is reached, the path selection in the entire reasoning process is recorded to form a complete reasoning chain; based on the selected path length, the similarity score of each hop and the reliability score of the reasoning rule used, the confidence score of the entire reasoning result is calculated;

[0062] The complete reasoning chain, the final node reached, and the confidence score are combined to form a preliminary reasoning result.

[0063] According to a second aspect of the embodiments of the present invention,

[0064] Provide a notarization intelligent question-answering customer service system based on knowledge graph, including:

[0065] The first unit is used to collect multi-source heterogeneous data in the notarization field, pre-process the multi-source heterogeneous data to obtain pre-processed data, extract professional terms and domain concepts from the pre-processed data using a named entity recognition algorithm to form an entity set, analyze the relationship between entities in the entity set based on a semantic role labeling technology, construct a relationship set, combine the entity set and the relationship set to form an initial knowledge graph, extract knowledge from the newly added pre-processed data and add it to the initial knowledge graph based on a dynamic update mechanism of time decay to form a dynamic knowledge graph, and vectorize the dynamic knowledge graph to obtain a graph vector;

[0066] The second unit is used to receive a multimodal question input by a user, extract the voice text, image text and image semantic information based on the voice and image in the multimodal question, combine the plain text in the multimodal question, perform multimodal feature fusion, determine the fusion feature, perform deep semantic understanding on the fusion feature, obtain the semantic feature, input the semantic feature into the intention recognition model, adapt the predefined notarization business intention category through semantic mapping, determine the user intention, and perform entity linking and relationship extraction on the semantic feature based on the user intention combined with the dynamic knowledge graph to form a structured query expression;

[0067] The third unit is used to perform multi-hop reasoning in a dynamic knowledge graph based on user intent and structured query expressions, use graph vectors to calculate similarity, and determine the reasoning path in combination with user intent. It performs knowledge association reasoning along the reasoning path to obtain preliminary reasoning results, performs interpretable analysis on the preliminary reasoning results, generates reasoning chains and confidence scores, and forms detailed reasoning results. It combines pre-acquired user portrait information and user intent to perform personalized sorting on the detailed reasoning results to obtain ordered reasoning results, inputs the ordered reasoning results into a pre-built Transformer model, outputs preliminary natural language answers, and applies text style transfer technology to finally generate personalized answers.

[0068] According to a third aspect of the embodiments of the present invention,

[0069] An electronic device is provided, comprising:

[0070] processor;

[0071] a memory for storing processor-executable instructions;

[0072] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0073] A fourth aspect of the embodiments of the present invention is:

[0074] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0075] In the embodiment of the present invention, through named entity recognition and semantic role labeling technology, professional terms, domain concepts and their relationships in the notarization field are accurately extracted to construct a high-quality initial knowledge graph; a time decay mechanism is adopted to automatically extract knowledge from new data and update the initial knowledge graph to ensure the timeliness and dynamics of the knowledge graph; the dynamic knowledge graph is vectorized to improve the application efficiency of the knowledge graph in subsequent reasoning, retrieval and other tasks; through feature extraction and fusion of voice, image and text information, a comprehensive understanding of multimodal problems is achieved, and the parsing ability of user input is improved; through semantic mapping, semantic features are adapted to predefined notarization Business intent classification ensures the accuracy and effectiveness of user intent recognition; combined with dynamic knowledge graphs, entity linking and relationship extraction are performed to ultimately generate structured query expressions, providing a standardized data basis for subsequent operations; through multi-hop reasoning in dynamic knowledge graphs, combined with graph vector similarity and user intent, efficient reasoning paths are determined to improve the accuracy of knowledge association reasoning; explainability and confidence analysis: Generate reasoning chains and confidence scores to provide explainability for reasoning results and enhance users' understanding and trust in the reasoning process; based on user portraits and intents, personalized sorting of reasoning results is performed to improve the relevance and personalized adaptability of the results. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 It is a flowchart of a notarized intelligent question-answering customer service method based on a knowledge graph according to an embodiment of the present invention;

[0077] Figure 2 It is a structural diagram of a notarized intelligent question-and-answer customer service system based on a knowledge graph according to an embodiment of the present invention. DETAILED DESCRIPTION

[0078] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0079] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0080] Figure 1 Schematic diagram of the process of the notarized intelligent question-answering customer service method based on the knowledge graph in an embodiment of the present invention. Figure 1 As shown, the method includes:

[0081] S101. Collect multi-source heterogeneous data in the notarization field, pre-process the multi-source heterogeneous data to obtain pre-processed data, extract professional terms and domain concepts from the pre-processed data using a named entity recognition algorithm to form an entity set, analyze the relationship between entities in the entity set based on semantic role labeling technology, construct a relationship set, combine the entity set and the relationship set to form an initial knowledge graph, extract knowledge from the newly added pre-processed data and add it to the initial knowledge graph based on a dynamic update mechanism of time decay to form a dynamic knowledge graph, and vectorize the dynamic knowledge graph to obtain a graph vector;

[0082] S102. Receive a multimodal question input by a user, extract voice text, image text and image semantic information based on the voice and image in the multimodal question, combine with the plain text in the multimodal question, perform multimodal feature fusion, determine the fusion feature, perform deep semantic understanding on the fusion feature, obtain semantic features, input the semantic features into the intent recognition model, adapt the predefined notarization business intent category through semantic mapping, determine the user intent, perform entity linking and relationship extraction on the semantic features based on the user intent combined with the dynamic knowledge graph, and form a structured query expression;

[0083] S103. Based on user intent and structured query expressions, multi-hop reasoning is performed in the dynamic knowledge graph, and similarity is calculated using graph vectors. In combination with user intent, the reasoning path is determined, and knowledge association reasoning is performed along the reasoning path to obtain preliminary reasoning results. The preliminary reasoning results are analyzed for interpretability, and reasoning chains and confidence scores are generated to form detailed reasoning results. Combined with the pre-acquired user portrait information and user intent, the detailed reasoning results are personalized and sorted to obtain ordered reasoning results. The ordered reasoning results are input into the pre-built Transformer model, preliminary natural language answers are output, and text style transfer technology is applied to finally generate personalized answers.

[0084] In a specific embodiment, the steps are as follows:

[0085] First, we collect multi-source heterogeneous data in the notarization field, including notarization laws and regulations, notarization case databases, notarization business manuals, etc. We preprocess the collected data, including text cleaning, word segmentation, and stop word removal, to obtain standardized preprocessed data. Then, we use the named entity recognition algorithm to extract professional terms and domain concepts from the preprocessed data, such as "notarization office", "notary", "notarial certificate", etc., to form an entity set. Specifically, we can use a BiLSTM-CRF-based named entity recognition model and use a notarization field labeled dataset for training.

[0086] Then, the relationship between entities in the entity set is analyzed based on semantic role labeling technology to construct a relationship set. The semantic role labeling method based on dependency syntactic analysis can be used to extract semantic roles such as subject, predicate, and object, and identify semantic relationships between entities, such as "notary-issue-notarial certificate". The entity set and the relationship set are combined to form an initial knowledge graph.

[0087] Then, based on the dynamic update mechanism of time decay, knowledge is extracted from the newly added preprocessed data and added to the initial knowledge graph. Specifically, the time decay function f(t) can be set to e (-λt) , where t is the knowledge existence time and λ is the decay coefficient. Each triple in the graph is given an initial weight of 1, and the weight gradually decays over time. When the weight is lower than the threshold, the knowledge is deleted from the graph. At the same time, new knowledge is extracted from the newly added data, given a higher initial weight and added to the graph, thus forming a dynamic knowledge graph.

[0088] The dynamic knowledge graph is vectorized to obtain the graph vector. Knowledge graph embedding methods such as TransE can be used to map entities and relationships into a low-dimensional vector space to facilitate subsequent similarity calculations.

[0089] In the question-answering stage, we first receive multimodal questions input by users, including voice, image and text. We perform speech recognition on the voice to obtain speech text, and perform text recognition and scene understanding on the image to obtain image text and semantic information. We perform feature fusion of speech text, image text, image semantic information and original text, and adopt the multimodal feature fusion method of attention mechanism to obtain fused feature vector.

[0090] To conduct deep semantic understanding of fused features, pre-trained language models such as BERT can be used to extract semantic features. The semantic features are input into the intent recognition model, and through semantic mapping, user questions are mapped to pre-defined notarization business intent categories, such as "handling notarization" and "notarization fees", to determine user intent.

[0091] Based on user intent and dynamic knowledge graph, entity linking and relationship extraction are performed on semantic features. Entity linking can use the BERT-based entity disambiguation model to link the entities mentioned in the question to the corresponding nodes in the graph. Relation extraction can use the remotely supervised relationship extraction model to extract the semantic relationship between entities. Finally, a structured query expression such as SPARQL is formed.

[0092] According to user intent and structured query expressions, multi-hop reasoning is performed in the dynamic knowledge graph. First, the similarity of entities and relationships is calculated using graph vectors, and the initial reasoning path is determined in combination with user intent. Then, knowledge association reasoning is performed along the reasoning path, and a path search algorithm based on reinforcement learning can be used. The reasoning process is analyzed for interpretability, and reasoning chains and confidence scores are generated to form detailed reasoning results.

[0093] Combined with the pre-obtained user profile information (such as user preferences, historical behaviors, etc.) and user intentions, the detailed reasoning results are sorted in a personalized manner. Sorting algorithms such as Learning to Rank can be used to obtain ordered reasoning results. The ordered reasoning results are input into the pre-built Transformer model, which is fine-tuned on the notarization field question and answer data and outputs preliminary natural language answers.

[0094] Finally, the VAE-based text style transfer technology is applied to convert the preliminary answers into personalized expression styles that match the user portrait, such as formal, colloquial, etc. Through the above steps, a personalized notarized intelligent question-and-answer result is finally generated.

[0095] For example, a user asks "I want to apply for real estate notarization, what materials do I need to prepare?" First, semantic features are extracted to identify the user's intention as "apply for notarization". Entity linking identifies the "real estate notarization" entity, and relationship extraction obtains the "materials-preparation-real estate notarization" relationship. Multi-hop reasoning is performed in the knowledge graph to find the "required materials" node related to "real estate notarization". Combined with user portraits, personalized sorting and expression are performed, and finally the answer is generated: "Hello, the following materials are required for real estate notarization: 1) Original house ownership certificate; 2) Original ID card; 3) Original household registration booklet. It is recommended that you prepare these materials in advance to speed up the processing progress. If you have any other questions, please let me know."

[0096] In this embodiment, through named entity recognition and semantic role labeling technology, professional terms, domain concepts and their relationships in the notarization field are accurately extracted to construct a high-quality initial knowledge graph; a time decay mechanism is adopted to automatically extract knowledge from new data and update the initial knowledge graph to ensure the timeliness and dynamics of the knowledge graph; the dynamic knowledge graph is vectorized to improve the application efficiency of the knowledge graph in subsequent reasoning, retrieval and other tasks; effective processing and fusion of multi-source heterogeneous data are achieved to build a more comprehensive and accurate knowledge graph in the notarization field; through feature extraction and fusion of voice, image and text information, a comprehensive understanding of multimodal problems is achieved, and the parsing ability of user input is improved; deep semantic analysis using fused features can more accurately obtain user intent information and enhance the accuracy of intent recognition; semantic features are mapped through semantic mapping. Adapt to predefined notarization business intent categories to ensure the accuracy and effectiveness of user intent recognition; combine with dynamic knowledge graphs to perform entity linking and relationship extraction, and finally generate structured query expressions to provide a standardized data basis for subsequent operations; through multi-hop reasoning in dynamic knowledge graphs, combined with graph vector similarity and user intent, determine efficient reasoning paths and improve the accuracy of knowledge association reasoning; explainability and confidence analysis: generate reasoning chains and confidence scores to provide explainability for reasoning results and enhance users' understanding and trust in the reasoning process; based on user portraits and intentions, personalize the sorting of reasoning results to improve the relevance and personal adaptability of the results; generate preliminary answers through the Transformer model, and generate personalized answers in combination with text style transfer technology to ensure that the output conforms to the user's language habits and style.

[0097] In an optional implementation, using a named entity recognition algorithm to extract professional terms and domain concepts from the preprocessed data to form an entity set, analyzing the relationship between entities in the entity set based on a semantic role labeling technology, constructing a relationship set, and combining the entity set and the relationship set to form an initial knowledge graph includes:

[0098] Construct a notarization field annotated dataset, and train a notarization entity recognition deep fusion network based on the annotated dataset. The notarization entity recognition deep fusion network includes a BERT-based encoding layer, a bidirectional long short-term memory network-based feature extraction layer, and a conditional random field-based decoding layer. The notarization field text is subjected to named entity recognition through the notarization entity recognition deep fusion network to obtain an entity set, which includes entity name, entity type, occurrence frequency, and first occurrence position.

[0099] Performing dependency syntactic analysis on sentences containing entities in the entity set to identify subject-predicate-object structures, using a semantic role labeling model to perform semantic role labeling on sentences containing the subject-predicate-object structures to obtain a predicate-argument structure, and determining a semantic role labeling label;

[0100] Constructing a mapping rule set from semantic role annotation labels to notarization domain relationship types, and converting semantic role annotation results into candidate relationships according to the mapping rule set;

[0101] Constructing a relationship classifier and training it with a preset remote supervision data set, using the relationship classifier to score and filter candidate relationships to obtain a relationship set;

[0102] The entities in the entity set are imported as entity nodes into the graph database, and the relationships in the relationship set are imported as relationship edges into the graph database, the corresponding entity nodes are connected, and the imported entity nodes and relationship edges are optimized using a graph algorithm to eliminate redundant relationships and merge synonymous entities, so as to construct an initial knowledge graph.

[0103] Regarding the method for constructing a knowledge graph in the notarization field, this embodiment provides a specific implementation scheme based on deep learning and natural language processing technology. The scheme mainly includes the following steps:

[0104] First, we construct a notarization annotated dataset. We collect a large amount of notarization-related texts and invite experts in the notarization field to manually annotate them. The annotation content includes entity types, entity boundaries, and the relationship between entities. The annotated dataset consists of a training set, a validation set, and a test set, of which the training set accounts for about 70%, and the validation set and the test set each account for 15%.

[0105] The deep fusion network for entity recognition is trained based on the annotated dataset. The network adopts a three-layer structure: the first layer is the encoding layer based on BERT, which uses the pre-trained BERT model to encode the input text and obtain the context-related word vector representation; the second layer is the feature extraction layer based on the bidirectional long short-term memory network (Bi-LSTM), which further extracts sequence features; the third layer is the decoding layer based on the conditional random field (CRF), which decodes the output of Bi-LSTM to obtain the final entity label sequence. The network is trained using the mini-batch gradient descent algorithm, with the batch size set to 32, the initial value of the learning rate to 0.001, and the Adam optimizer.

[0106] Use the trained notarization entity recognition deep fusion network to perform named entity recognition on notarization field text. For the input text, first perform word segmentation and part-of-speech tagging preprocessing, and then send it to the network for recognition. The network outputs the entity label of each word, and extracts the entity according to the label sequence. Post-process the recognition results, including entity deduplication, merging adjacent entities, etc. Finally, an entity set is obtained, each entity contains information such as entity name, entity type, frequency of occurrence, and first occurrence position.

[0107] For example, for the input text "Zhang San and Li Si signed a house sales contract", the entity set obtained after named entity recognition is:

[0108] {

[0109] "Zhang San": {type: "natural person", frequency: 1, first position: 0},

[0110] "Li Si": {type: "natural person", frequency: 1, first position: 3},

[0111] "House Sales Contract": {Type: "Contract", Frequency: 1, First Position: 8}

[0112] }

[0113] Next, we perform dependency syntactic analysis on the sentences containing entities to identify the subject-verb-object structure. We use a dependency syntactic analysis model based on a neural network, such as a transfer system model based on a bidirectional LSTM. The dependency syntactic analysis result for the above example sentence is:

[0114] Subject: Zhang San;

[0115] Predicate: to sign;

[0116] Object: House sale contract;

[0117] Then, the semantic role labeling model is used to label the sentences containing subject-verb-object structure. The semantic role labeling model uses an end-to-end neural network structure, including a word embedding layer, a BiLSTM encoding layer, and a CRF decoding layer. The model input is a word sequence and syntactic tree information, and the output is a semantic role label for each word. The semantic role labeling result of the example sentence is:

[0118] [Zhang San] A0 and [Li Si] A0 [sign] V[house sale contract] A1;

[0119] A0 represents the action executor, V represents the predicate, and A1 represents the action recipient.

[0120] Construct a mapping rule set from semantic role annotation labels to notarization domain relationship types. For example:

[0121] A0-V-A1→signing relationship;

[0122] A0-V-A2 → agency relationship;

[0123] LOC-V-A1 → location relationship;

[0124] etc.

[0125] According to the mapping rule set, the semantic role labeling results are converted into candidate relations. For the example sentence, the candidate relations obtained are:

[0126] (Zhang San, signed, house sale contract)

[0127] (Li Si, signed, house sale contract)

[0128] Build a relation classifier and train it. The relation classifier uses a convolutional neural network structure, including a word embedding layer, a convolution layer, a pooling layer, and a fully connected layer. Use a preset remote supervision dataset for training, which contains a large number of automatically labeled entity pairs and their relationships. The cross entropy loss function is used during training, and the Adam optimizer is used for optimization.

[0129] Use the trained relationship classifier to score and filter candidate relationships. The relationship classifier outputs the probability score of each relationship type, and selects the relationship type with the highest probability and exceeding the threshold as the final result. Candidate relationships with probabilities below the threshold are filtered. After being processed by the relationship classifier, the final relationship set is obtained:

[0130] {(Zhang San, signed, house sale contract), (Li Si, signed, house sale contract)}

[0131] Import the entities in the entity collection as entity nodes into a graph database, such as Neo4j. Each node contains attribute information such as entity name and type. Import the relationships in the relationship collection as relationship edges into the graph database to connect the corresponding entity nodes. For example:

[0132] CREATE (Zhang San: natural person {name: 'Zhang San'})

[0133] CREATE (Li Si: natural person {name: 'Li Si'})

[0134] CREATE (contract: contract {name: 'house sale contract'})

[0135] CREATE (Zhang San)-[: Sign] → (Contract)

[0136] CREATE(Li Si)-[:Sign]→(Contract)

[0137] Use graph algorithms to optimize the imported entity nodes and relationship edges. First, use community discovery algorithms, such as the Louvain algorithm, to divide the graph into communities and discover closely related entity clusters. Then use the PageRank algorithm to calculate the importance scores of the nodes and prune redundant nodes with low importance. For entity nodes with high similarity, such as "House Purchase Contract" and "House Sales Contract", use the entity alignment algorithm to merge them.

[0138] Finally, the optimized initial knowledge graph is obtained, which contains the core entities and relationships in the notarization field. This knowledge graph can support subsequent knowledge reasoning, question-answering and other applications. As new notarization texts are continuously added, the knowledge graph can be continuously updated and expanded to continuously improve its coverage and accuracy.

[0139] In this embodiment, the notarization entity recognition deep fusion network, combined with BERT, bidirectional LSTM and CRF, can accurately identify named entities in notarization field texts and ensure the accuracy and completeness of the entity set; through dependency syntactic analysis and semantic role labeling models, the subject-predicate-object structure is identified and mapped to relationships in the notarization field, which improves the depth of semantic understanding; the candidate relationships are filtered through the relationship classifier, and the graph database and graph algorithm are used to optimize the entity nodes and relationship edges, eliminate redundancy and merge synonymous entities, and ensure the accuracy and quality of the knowledge graph; finally, a high-quality and accurate notarization field knowledge graph is constructed, which provides a solid foundation for the analysis and application of notarization data.

[0140] In an optional implementation, a notarization field annotated dataset is constructed, and a notarization entity recognition deep fusion network is trained based on the annotated dataset. The notarization entity recognition deep fusion network includes a BERT-based encoding layer, a bidirectional long short-term memory network-based feature extraction layer, and a conditional random field-based decoding layer. Named entity recognition is performed on notarization field texts through the notarization entity recognition deep fusion network, and the obtained entity set includes:

[0141] Constructing a notarization field annotated dataset, and pre-training the BERT model with a masked language model based on the notarization field annotated dataset to obtain a pre-trained BERT model adapted to the notarization field;

[0142] In the notarization entity recognition deep fusion network, the encoding layer is initialized using the weights of the pre-trained BERT model; the notarization field annotated dataset is input into the notarization entity recognition deep fusion network, and the word vector representation is obtained through the encoding layer;

[0143] Input the word vector representation into a feature extraction layer to extract sequence features, wherein the feature extraction layer includes a forward long short-term memory network and a backward long short-term memory network, each including an input gate, a forget gate, and an output gate;

[0144] Inputting the sequence features into the decoding layer, the decoding layer includes a linear chain conditional random field, establishing a label transfer probability matrix, and using the Viterbi algorithm to decode to obtain an optimal label sequence;

[0145] A joint learning strategy is adopted to optimize the encoding layer, feature extraction layer, and decoding layer at the same time. The AdamW optimizer is used to update the parameters, and the learning rate warm-up and linear decay strategies are implemented.

[0146] After each training round, the performance of the publicized entity recognition deep fusion network is evaluated using a preset validation set, and the model weight with the highest F1 score on the validation set is saved in combination with an early stopping strategy to obtain a trained publicized entity recognition deep fusion network;

[0147] Input the notarization domain text to be recognized into the trained notarization entity recognition deep fusion network, and sequentially pass through the encoding layer, feature extraction layer and decoding layer to obtain a label sequence;

[0148] According to the tag sequence, entities in the notarization domain text are identified and an entity set is constructed.

[0149] In this implementation, we first construct a notarization field annotated dataset. Specifically, we collect a large amount of text materials in the notarization field, including notarial certificates, notarization application materials, notarization laws and regulations, etc. Then, experts in the notarization field manually annotate these texts and mark out the entity types therein, such as parties, notarization matters, notarization agencies, etc. Finally, we obtain a dataset containing texts and their corresponding entity labels.

[0150] Next, based on the constructed notarization field annotated dataset, the BERT model is pre-trained with a masked language model. Specifically, the text in the annotated dataset is input into the BERT model, 15% of the tokens are randomly masked, and the model is asked to predict these masked tokens. In this way, the BERT model can learn the language features and contextual information of the notarization field text. After the pre-training is completed, a pre-trained BERT model adapted to the notarization field is obtained.

[0151] Then, a deep fusion network for notarization entity recognition is constructed. The network consists of three layers: an encoding layer based on BERT, a feature extraction layer based on a bidirectional long short-term memory network, and a decoding layer based on a conditional random field. The encoding layer is initialized using the weights of the pre-trained BERT model obtained earlier. The notarization domain annotated dataset is input into the network, and the word vector representation is first obtained through the encoding layer.

[0152] The word vector representation is then input into the feature extraction layer for sequence feature extraction. The feature extraction layer includes two long short-term memory networks, the forward and backward ones, which contain input gate, forget gate and output gate structures respectively. The long-distance dependency of the context can be captured through bidirectional processing.

[0153] The extracted sequence features are then input into the decoding layer. The decoding layer uses a linear chain conditional random field to establish a label transfer probability matrix and uses the Viterbi algorithm to decode and obtain the optimal label sequence.

[0154] During the training process, a joint learning strategy is used to optimize the three layers simultaneously. The AdamW optimizer is used to update the parameters, and the learning rate warm-up and linear decay strategies are implemented. Specifically, a smaller learning rate is used at the beginning of training, and it is gradually increased to the set maximum value as the training progresses, and then linearly decayed. This strategy can make the model converge stably in the early stage of training and fully learn in the later stage.

[0155] After each training round, the network performance is evaluated using the preset validation set. The F1 score on the validation set is calculated, and combined with the early stopping strategy, training is stopped in time when the validation set performance no longer improves to avoid overfitting. The model weight with the highest F1 score on the validation set is saved to obtain the final trained public entity recognition deep fusion network.

[0156] When in use, the notarization field text to be identified is input into the trained network. The text passes through the encoding layer to obtain word vectors, the feature extraction layer extracts sequence features, and the decoding layer obtains the label sequence. Based on the output label sequence, various entities in the text can be identified and an entity set can be constructed.

[0157] For example, the input text "Zhang San applied to Beijing Chang'an Notary Office for house ownership notarization on May 1, 2021". After network processing, the output label sequence is "B-PERSON O O-TIME I-TIME I-TIME O B-ORG I-ORG I-ORG OOO B-ITEM I-ITEM I-ITEM O". B- indicates the beginning of the entity, I- indicates the inside of the entity, and O indicates a non-entity. Based on this label sequence, the person name entity "Zhang San", the time entity "May 1, 2021", the organization entity "Beijing Chang'an Notary Office", and the event entity "house ownership notarization" can be identified. Finally, a set containing these entities is constructed.

[0158] This method can effectively identify various entities in notarization texts, providing a basis for subsequent information extraction and knowledge graph construction. This method combines the powerful language representation capabilities of BERT, the sequence modeling capabilities of long short-term memory networks, and the label dependency modeling capabilities of conditional random fields, and can accurately identify professional entities in the notarization field.

[0159] In this embodiment, by pre-training the BERT model on a notarization field labeled data set, the model is made more adaptable to the proprietary language characteristics of the notarization field and the accuracy of entity recognition is improved; a deep fusion network, including a BERT encoding layer, a bidirectional LSTM feature extraction layer and a CRF decoding layer, is used to more accurately extract sequence features and decode the optimal label sequence, thereby improving the named entity recognition effect; a joint learning strategy, an AdamW optimizer, and a learning rate warm-up and linear decay strategy are used to effectively accelerate model convergence and improve model performance; through validation set evaluation and early stopping strategies, the final model weight is ensured to reach the highest F1 score on the validation set, thereby ensuring the recognition effect; the trained model can accurately identify entities in notarization field texts and construct an accurate entity set for subsequent processing or analysis.

[0160] In an optional implementation, based on a dynamic update mechanism of time decay, knowledge is extracted from the newly added preprocessed data and added to the initial knowledge graph to form a dynamic knowledge graph, which includes:

[0161] Constructing a time decay function, wherein the time decay function adopts an exponential decay form to control the decay speed of the knowledge unit;

[0162] Receive newly added preprocessed data, and perform entity recognition and relationship extraction on the newly added preprocessed data to obtain a new knowledge unit, use the cosine similarity method to calculate the similarity between the new knowledge unit and the existing knowledge unit in the initial knowledge graph, and obtain a similarity calculation result, and determine whether the similarity calculation result is higher than a preset similarity threshold based on the similarity calculation result:

[0163] When the similarity calculation result is higher than the preset similarity threshold, the existing knowledge unit is updated with the new knowledge unit, and the fusion weight of the new knowledge unit and the existing knowledge unit is calculated using the time decay function; based on the fusion weight, the new knowledge unit and the existing knowledge unit are weightedly fused to obtain a fused knowledge unit;

[0164] When the similarity is lower than the preset similarity threshold, the new knowledge unit is used as a newly added knowledge unit;

[0165] Update the fused knowledge unit to the corresponding position, add the newly added knowledge unit to the initial knowledge graph, update the initial knowledge graph, and record the timestamp of each update or addition;

[0166] According to a preset full scan cycle, the initial knowledge graph is regularly scanned, the time decay value of each knowledge unit is calculated using the time decay function, and it is determined whether the time decay value of each knowledge unit is less than a preset decay threshold value. When the time decay value corresponding to a knowledge unit is less than the preset decay threshold value, the corresponding knowledge unit is removed from the initial knowledge graph;

[0167] The initial knowledge graph is updated to form a dynamic knowledge graph.

[0168] The dynamic update mechanism based on time decay extracts knowledge from the newly added preprocessed data and adds it to the initial knowledge graph. The process of forming a dynamic knowledge graph is as follows:

[0169] First, construct a time decay function, and use exponential decay to control the decay speed of knowledge units. Specifically, a baseline decay rate α (0<α<1) can be set, t represents the existence time of the knowledge unit, and the decay function can be expressed as f(t)=α^t. The smaller the α value, the faster the decay speed. For example, α=0.9 can be set, and the decay values ​​after 1 day, 10 days, and 30 days are 0.9, 0.35, and 0.04 respectively.

[0170] After receiving the newly added pre-processed data, entity recognition and relationship extraction are performed on it to obtain new knowledge units.

[0171] Then, the cosine similarity method is used to calculate the similarity between the new knowledge unit and the existing knowledge units in the initial knowledge graph. Specifically, the knowledge unit is represented as a vector, and the cosine value of the angle between the vectors is calculated as the similarity. For example, if the vectors of two knowledge units are (0.5, 0.8, 0.3) and (0.4, 0.7, 0.5), their similarity is 0.97.

[0172] Based on the calculation results, determine whether the similarity is higher than the preset threshold. If it is higher than the threshold, the new knowledge unit is used to update the existing knowledge unit. At the same time, the fusion weight is calculated using the time decay function, and the sum of the weights of the new and old knowledge units is 1. For example, if the existence time of the new and old knowledge units is 1 day and 10 days respectively, the weights can be set to 0.9 and 0.1. Based on the weights, the new and old knowledge units are weighted averaged to obtain the fused knowledge unit.

[0173] If the similarity is lower than the threshold, the new knowledge unit is directly added as a new knowledge unit. For example, if the similarity threshold is set to 0.8 and the calculated similarity is 0.7, the new knowledge unit is directly added to the knowledge graph.

[0174] Update the fused knowledge unit to the corresponding position, add the newly added knowledge unit to the initial knowledge graph, and record the timestamp of each update or addition. This way, the "age" of each knowledge unit can be tracked.

[0175] The initial knowledge graph is scanned regularly according to the preset full scan cycle (e.g. once a week). The time decay function is used to calculate the time decay value of each knowledge unit and determine whether it is less than the preset decay threshold. For example, if the decay threshold is set to 0.1, a knowledge unit exists for 30 days, and its decay value is 0.04, which is less than the threshold, then it will be removed from the knowledge graph.

[0176] Through the above steps, the initial knowledge graph is continuously updated to form a dynamic knowledge graph. This mechanism can absorb new knowledge in a timely manner, while eliminating outdated knowledge, and maintain the timeliness and accuracy of the knowledge graph. This dynamic update mechanism based on time decay has the advantages of strong adaptability and good scalability. By adjusting the decay function parameters, similarity threshold, decay threshold, etc., the speed of knowledge updating and elimination can be flexibly controlled to meet the needs of different fields and application scenarios. At the same time, this mechanism also provides effective technical support for the long-term maintenance and evolution of the knowledge graph.

[0177] In this embodiment, through similarity calculation and time decay function, new knowledge units are weightedly merged or added with existing knowledge units, thereby realizing dynamic updating and refined processing of knowledge; a time decay function in the form of exponential decay is used to reasonably control the decay rate of knowledge units to ensure the timeliness and accuracy of knowledge in the knowledge graph; through regular scanning and time decay mechanism, outdated knowledge units are automatically removed to maintain efficient management and dynamic adjustment of the knowledge graph; through continuous updating and optimization, a dynamic knowledge graph that changes over time is formed to ensure the timeliness and practicality of the knowledge graph.

[0178] In an optional implementation, a multimodal question input by a user is received, and based on the voice and image in the multimodal question, the voice text, image text and image semantic information are extracted, and the multimodal feature fusion is performed in combination with the plain text in the multimodal question to determine the fusion feature, and the fusion feature is deeply semantically understood to obtain the semantic feature, and the semantic feature is input into the intent recognition model, and the predefined notarization business intent category is adapted through semantic mapping to determine that the user intent includes:

[0179] Receiving a multimodal question, wherein the multimodal question includes voice information, image information, and text information;

[0180] Using a speech recognition model to convert speech information in the multimodal problem into a first text sequence; using an optical character recognition engine to process image information in the multimodal problem to obtain a second text sequence; splicing the first text sequence, the second text sequence and the text information in the multimodal problem to form a combined text sequence, and mapping each word in the combined text sequence to a specified dimensional space through a first independent linear layer to obtain text features;

[0181] Extracting semantic features of the image information in the multimodal problem using a visual Transformer model to obtain an image semantic vector, and mapping the image semantic vector to a corresponding dimensional space having the same dimension as the text feature through a second independent linear layer to obtain an image feature;

[0182] Using a multi-head attention mechanism to calculate the correlation between the text feature and the image feature to obtain an attention output, performing a residual connection and layer normalization on the attention output and the text feature to obtain a fused feature, and inputting the fused feature into a pre-trained bidirectional encoder to obtain a context-related representation; extracting a tag vector from the context-related representation, and mapping it to a task-corresponding space through a linear layer and an activation function to obtain a semantic feature;

[0183] The semantic features are mapped to the predefined notarization business intention category space through a fully connected layer to obtain the intention logic value, the sigmoid function is applied to the intention logic value, the probability of each intention category is calculated, the intention category with the highest probability is selected, and the user intention is determined.

[0184] This embodiment provides a method for notarization business intention recognition based on multimodal input. The method first receives a multimodal question input containing voice, image and text. For voice information, a pre-trained speech recognition model is used to convert it into a text sequence. Specifically, a Transformer-based speech recognition model such as Wav2Vec 2.0 can be used. The model captures the contextual representation of speech through self-supervised learning and then fine-tunes it for speech recognition tasks. For example, for the input voice "I want to apply for real estate notarization", the model can accurately recognize and output the corresponding text sequence.

[0185] For image information, an optical character recognition (OCR) engine is used to extract text content from the image. An end-to-end OCR model based on convolutional neural networks and recurrent neural networks can be used, such as CRNN (Convolutional Recurrent Neural Network). This model first uses CNN to extract image features, then uses RNN to model the feature sequence, and finally decodes the text result through CTC (Connectionist Temporal Classification). For example, for a photo of a real estate certificate, the OCR engine can recognize and output the text information on the certificate, such as "House Ownership Certificate".

[0186] The first text sequence obtained by speech recognition and the second text sequence obtained by OCR are concatenated with the plain text information in the original question to form a complete combined text sequence. Then, each token in the sequence is mapped to a vector space of a specified dimension through an independent linear layer to obtain a text feature representation. For example, each token can be mapped to a 768-dimensional vector.

[0187] For semantic information extraction of images, a visual Transformer model, such as ViT (Vision Transformer), is used. This model divides the input image into fixed-size image blocks, and captures the global context information of the image through multiple layers of self-attention and feed-forward networks. The final output [CLS] tag vector is used as the semantic representation of the entire image. Through another independent linear layer, the semantic vector is mapped to the same dimensional space as the text feature to obtain the image feature representation.

[0188] Next, the multi-head attention mechanism is used to calculate the correlation between text features and image features. Multi-head attention allows the model to focus on information in different subspaces at the same time, enhancing the ability of feature fusion. In specific implementation, 8 attention heads can be used, each with a dimension of 96. The output of the attention is residually connected with the original text features and normalized by the layer to obtain the fused feature representation.

[0189] The fused features are input into a pre-trained bidirectional encoder, such as BERT (Bidirectional Encoder Representations from Transformers), to further extract context-related semantic representations. The BERT model is pre-trained with a masked language model and next sentence prediction tasks, and can effectively capture the bidirectional context information of the text. The vector corresponding to the [CLS] tag is extracted from the output of BERT as the sentence-level representation of the entire input.

[0190] Finally, a fully connected layer is used to map the above semantic features to the predefined notarization business intent category space. Assume that 10 notarization business intent categories are predefined, such as real estate notarization, marriage notarization, inheritance notarization, etc. The output dimension of the fully connected layer is 10, corresponding to the logit value of each intent category. The sigmoid function is applied to these logit values ​​to convert them into probability values ​​between 0 and 1. The category with the highest probability is selected as the final recognized user intent.

[0191] For example, for the multimodal input question "I want to apply for real estate notarization", which includes voice, real estate certificate image and text, after the above processing, the model may output the following probability distribution: real estate notarization (0.92), marriage notarization (0.03), inheritance notarization (0.02), etc. At this time, the system will recognize "real estate notarization" as the user's intention.

[0192] This multimodal feature fusion and deep semantic understanding method can fully utilize the information in speech, images and text to improve the accuracy and robustness of notarization business intention recognition. This method can flexibly respond to various complex user input scenarios and provide a reliable intention understanding basis for subsequent notarization business processing.

[0193] In this embodiment, by converting speech, image and text information into text features and image features, and using a multi-head attention mechanism to calculate correlation, efficient fusion of different modal information is achieved, and the ability to understand multimodal problems is improved; by extracting context-related representations through a pre-trained bidirectional encoder, semantic information and the connection between different modalities can be better captured, thereby improving the accuracy of problem handling; by calculating semantic features using fused features, and calculating the probability of intent categories through a fully connected layer and a sigmoid function, user intent can be efficiently and accurately identified, ensuring the intelligent processing of notarization services; comprehensive processing and analysis of speech, images and text is achieved, and the system's ability to understand and respond to complex multimodal problems is improved.

[0194] In an optional implementation, based on the user intent and in combination with a dynamic knowledge graph, entity linking and relationship extraction are performed on the semantic features to form a structured query expression, including:

[0195] Constructing a pre-trained stacked network model, wherein the stacked network model adopts a permutation language modeling training objective, uses a two-stream self-attention mechanism and a segmented recurrence mechanism;

[0196] Inputting semantic features into the stacked network model to obtain bidirectional context-aware word vector representations; constructing a multi-task learning architecture on the top layer of the stacked network model, including entity recognition tasks and relationship extraction tasks; introducing a relative position encoding mechanism into the multi-task learning architecture to capture long-distance dependencies; using a dynamic weight allocation strategy to adaptively adjust the importance weights of entity recognition tasks and relationship extraction tasks in the multi-task learning architecture;

[0197] Integrate a dynamic knowledge graph, process the dynamic knowledge graph through a graph attention network, obtain enhanced entity and enhanced relationship representations in the word vector representation, and form an enhanced word vector representation;

[0198] Apply adversarial training technology to the enhanced word vector representation, train to obtain the final multi-task learning architecture, input semantic features, and identify corresponding entity mentions; calculate the similarity between the entity mentions and the entities in the dynamic knowledge graph, select the entity with the highest similarity as the link result, obtain a linked entity set, and generate candidate relationship pairs based on the entity pairs in the linked entity set;

[0199] Based on the relationship extraction task in the multi-task learning architecture, taking the semantic features and the candidate relationship pairs as input, predicting the relationship between entity pairs to obtain a relationship set;

[0200] A corresponding query template is selected based on the user's intention, and the linked entity set and the relationship set are filled into the query template to generate a structured query expression.

[0201] The specific implementation method of performing entity linking and relationship extraction on semantic features based on user intent and dynamic knowledge graph to form a structured query expression is as follows:

[0202] First, a pre-trained stacked network model is constructed. The model adopts the training objective of permutation language modeling, and uses a two-stream self-attention mechanism and a segmented loop mechanism. Specifically, a multi-layer Transformer structure is used as the backbone, and each layer of Transformer contains a multi-head self-attention layer and a feedforward neural network layer. In the self-attention mechanism, a two-stream structure is adopted to calculate the attention scores of the content stream and the query stream respectively, and then the results of the two streams are fused. At the same time, a segmented loop mechanism is introduced to divide long sequences into multiple segments, and loop calculations are performed within and between segments to capture long-distance dependencies.

[0203] Next, the semantic features are input into the superimposed network model to obtain a bidirectional context-aware word vector representation. Specifically, each word token in the semantic features is input into the model in turn, and after calculation by multiple layers of Transformer, the contextualized representation of each word is obtained. Then, a multi-task learning architecture is constructed on the top layer of the superimposed network model, including entity recognition tasks and relationship extraction tasks. A relative position encoding mechanism is introduced into the multi-task learning architecture to capture long-distance dependencies. Specifically, on the basis of the original absolute position encoding, an additional relative position encoding is added to calculate the relative distance between words, and this information is incorporated into the calculation of self-attention. At the same time, the dynamic weight allocation strategy is used to adaptively adjust the importance weights of the entity recognition task and the relationship extraction task in the multi-task learning architecture. According to the training loss of the two tasks, their weight ratio in the total loss is dynamically adjusted.

[0204] Next, the dynamic knowledge graph is integrated and processed through the graph attention network to obtain the enhanced entity and enhanced relationship representations in the word vector representation to form the enhanced word vector representation. Specifically, the entity and relationship information in the knowledge graph is converted into a graph structure, and the graph attention network is used to pass messages on the graph to obtain the representation of each node. These representations are then fused with the original word vector to obtain the enhanced word vector representation.

[0205] Then, adversarial training techniques are applied to the enhanced word vector representation to train the final multi-task learning architecture. During the training process, adversarial samples are introduced to generate adversarial samples by adding perturbations to the word vectors to improve the robustness of the model. Next, the semantic features are input to identify the corresponding entity mentions. Using the sequence labeling method, each word in the input sequence is labeled to identify the boundaries of entity mentions.

[0206] After that, the similarity between the entity mention and the entity in the dynamic knowledge graph is calculated, and the entity with the highest similarity is selected as the link result to obtain the linked entity set. Specifically, for each identified entity mention, its similarity with all entities in the knowledge graph is calculated, and the entity with the highest similarity is selected as the link result. The similarity calculation can be based on the cosine similarity of the word vector.

[0207] Finally, based on the relation extraction task in the multi-task learning architecture, the semantic features and candidate relation pairs are used as input to predict the relationship between entity pairs and obtain a relation set. For each candidate relation pair, the model outputs a probability score indicating the likelihood that the relation is established. The relation with the highest probability is selected as the prediction result.

[0208] Select the corresponding query template based on the user's intent, fill the linked entity set and relationship set into the query template, and generate a structured query expression. Select the corresponding query template according to different user intents. For example, for the intent of asking about the capital, you can choose a template like "X is the capital of Y". Then fill the identified and extracted entities and relationships into the template slots to obtain the final structured query expression.

[0209] Through the above steps, entity linking and relationship extraction based on user intent and dynamic knowledge graph are realized, and natural language queries are converted into structured query expressions, providing a basis for subsequent knowledge retrieval and question answering.

[0210] In an optional implementation, based on user intent and structured query expressions, multi-hop reasoning is performed in a dynamic knowledge graph, similarity is calculated using graph vectors, and the reasoning path is determined in combination with user intent, and knowledge association reasoning is performed along the reasoning path, and preliminary reasoning results are obtained, including:

[0211] Receive user intent and structured query expressions as input, determine the inference starting entity in the dynamic knowledge graph, take the inference starting entity as the current node, identify entities and relationships directly connected to the current node, and form a candidate next hop set;

[0212] Using the pre-generated graph vector, calculate the similarity score between the current node and each node in the candidate next hop set; convert the user intention into a vector representation, and fuse it with the graph vector of each node in the candidate next hop set to obtain a fusion vector; combine the similarity score with the fusion vector to generate a comprehensive score for each candidate next hop node;

[0213] Based on the comprehensive score, a next hop is selected using a reinforcement learning agent, wherein the reinforcement learning agent takes a current state as input and outputs a probability distribution of selecting the next hop, wherein the current state includes the current node information, the user intention, the comprehensive score, and the selected path information;

[0214] According to the selected next hop, update the current node and the selected path information;

[0215] Applying predefined reasoning rules to the selected path to perform rule-based explicit reasoning;

[0216] Encoding the selected path using a graph neural network model and performing implicit reasoning based on the graph neural network, wherein the graph neural network model is a graph convolutional network;

[0217] The results of explicit reasoning and implicit reasoning are integrated to obtain a fused reasoning result;

[0218] Until the preset hop limit is reached, the path selection in the entire reasoning process is recorded to form a complete reasoning chain; based on the selected path length, the similarity score of each hop and the reliability score of the reasoning rule used, the confidence score of the entire reasoning result is calculated;

[0219] The complete reasoning chain, the final node reached, and the confidence score are combined to form a preliminary reasoning result.

[0220] In this embodiment, through the node selection and path update mechanism in the dynamic knowledge graph, combined with user intent and structured query expressions, the appropriate reasoning path can be quickly identified, thereby improving the reasoning efficiency; by generating a comprehensive score through the fusion of graph vector similarity and user intent, the accuracy of candidate node selection is ensured, and the decision quality in the reasoning process is improved; by combining explicit and implicit reasoning through rule reasoning and graph convolutional networks, a more comprehensive reasoning capability is achieved, which can handle complex relationships and logic; a reinforcement learning agent is used to select the next hop node, and optimization is performed based on the comprehensive score and the selected path to ensure that the reasoning process has an efficient decision-making mechanism; the confidence of the overall reasoning result is calculated through the path length, similarity score and reliability score of the reasoning rule, thereby ensuring the reliability and interpretability of the reasoning result; the reasoning path is recorded to form a complete reasoning chain, thereby ensuring the traceability and transparency of the reasoning process.

[0221] According to user intent and structured query expressions, the specific implementation of multi-hop reasoning in a dynamic knowledge graph is as follows:

[0222] First, the system receives the user's input intent and structured query expression. For example, the user intent may be "query movies starring a certain actor", and the structured query expression may be "QUERY (actor, starring, movie)". Based on these inputs, the system determines the inference starting entity in the dynamic knowledge graph.

[0223] Next, the system uses the pre-generated graph vector to calculate the similarity score between the current node and each node in the candidate next hop set. At the same time, the user's intention "query movies starring a certain actor" is converted into a vector representation and fused with the graph vector of the candidate next hop node to obtain a fused vector. The similarity score is combined with the fused vector to generate a comprehensive score for each candidate next hop node.

[0224] The system then uses a reinforcement learning agent to select the next hop. The reinforcement learning agent takes the current state as input, including the current node, user intent, comprehensive score, and selected path information, and outputs a probability distribution for selecting the next hop.

[0225] Apply predefined inference rules to the selected path to perform explicit rule-based reasoning. For example, apply the rule "if A stars in B, then B is a movie".

[0226] At the same time, a graph neural network model (such as a graph convolutional network) is used to encode the selected path and perform implicit reasoning based on the graph neural network. The graph neural network may capture some implicit patterns.

[0227] The results of explicit reasoning and implicit reasoning are integrated to obtain a fused reasoning result.

[0228] The system continues to perform multi-hop reasoning until the preset hop limit (such as 3 hops) is reached. The path selection is recorded throughout the process to form a complete reasoning chain.

[0229] Based on the length of the selected path, the similarity score of each hop, and the reliability of the inference rules used, the system calculates the confidence score of the entire inference result. For example, if the path is short, the similarity of each hop is high, and the rules used are reliable, a higher confidence score is given.

[0230] Finally, the system combines the complete reasoning chain, the final node reached (such as "2004"), and the confidence score to form a preliminary reasoning result.

[0231] Through this method, the system can perform multi-hop reasoning in a dynamic knowledge graph, combine user intent and graph structure, find relevant knowledge, and give the reasoning process and result confidence, providing a basis for subsequent question-answering or decision-making.

[0232] Figure 2 FIG. 1 is a schematic diagram of the structure of a notarized intelligent question-answering customer service system based on a knowledge graph according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0233] The first unit is used to collect multi-source heterogeneous data in the notarization field, pre-process the multi-source heterogeneous data to obtain pre-processed data, extract professional terms and domain concepts from the pre-processed data using a named entity recognition algorithm to form an entity set, analyze the relationship between entities in the entity set based on a semantic role labeling technology, construct a relationship set, combine the entity set and the relationship set to form an initial knowledge graph, extract knowledge from the newly added pre-processed data and add it to the initial knowledge graph based on a dynamic update mechanism of time decay to form a dynamic knowledge graph, and vectorize the dynamic knowledge graph to obtain a graph vector;

[0234] The second unit is used to receive a multimodal question input by a user, extract the voice text, image text and image semantic information based on the voice and image in the multimodal question, combine the plain text in the multimodal question, perform multimodal feature fusion, determine the fusion feature, perform deep semantic understanding on the fusion feature, obtain the semantic feature, input the semantic feature into the intention recognition model, adapt the predefined notarization business intention category through semantic mapping, determine the user intention, and perform entity linking and relationship extraction on the semantic feature based on the user intention combined with the dynamic knowledge graph to form a structured query expression;

[0235] The third unit is used to perform multi-hop reasoning in a dynamic knowledge graph based on user intent and structured query expressions, use graph vectors to calculate similarity, and determine the reasoning path in combination with user intent. It performs knowledge association reasoning along the reasoning path to obtain preliminary reasoning results, performs interpretable analysis on the preliminary reasoning results, generates reasoning chains and confidence scores, and forms detailed reasoning results. It combines pre-acquired user portrait information and user intent to perform personalized sorting on the detailed reasoning results to obtain ordered reasoning results, inputs the ordered reasoning results into a pre-built Transformer model, outputs preliminary natural language answers, and applies text style transfer technology to finally generate personalized answers.

[0236] According to a third aspect of the embodiments of the present invention,

[0237] An electronic device is provided, comprising:

[0238] processor;

[0239] a memory for storing processor-executable instructions;

[0240] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0241] A fourth aspect of the embodiments of the present invention is:

[0242] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0243] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0244] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A notarized intelligent question-answering customer service method based on knowledge graph, characterized in that: include: Collect multi-source heterogeneous data in the field of notarization, pre-process the multi-source heterogeneous data to obtain pre-processed data, extract professional terms and domain concepts from the pre-processed data using a named entity recognition algorithm to form an entity set, analyze the relationship between entities in the entity set based on semantic role labeling technology, construct a relationship set, combine the entity set and the relationship set to form an initial knowledge graph, extract knowledge from the newly added pre-processed data and add it to the initial knowledge graph based on a dynamic update mechanism of time decay to form a dynamic knowledge graph, and vectorize the dynamic knowledge graph to obtain a graph vector; Receive a multimodal question input by a user, extract the voice text, image text and image semantic information based on the voice and image in the multimodal question, combine the plain text in the multimodal question, perform multimodal feature fusion, determine the fusion feature, perform deep semantic understanding on the fusion feature to obtain the semantic feature, input the semantic feature into the intent recognition model, adapt the predefined notarization business intent category through semantic mapping, determine the user intent, perform entity linking and relationship extraction on the semantic feature based on the user intent combined with the dynamic knowledge graph, and form a structured query expression; Based on user intent and structured query expressions, multi-hop reasoning is performed in the dynamic knowledge graph, and similarity is calculated using graph vectors. Combined with user intent, the reasoning path is determined, and knowledge association reasoning is performed along the reasoning path to obtain preliminary reasoning results. The preliminary reasoning results are analyzed for interpretability, and reasoning chains and confidence scores are generated to form detailed reasoning results. Combined with the pre-acquired user portrait information and user intent, the detailed reasoning results are personalized and sorted to obtain ordered reasoning results. The ordered reasoning results are input into the pre-built Transformer model, preliminary natural language answers are output, and text style transfer technology is applied to finally generate personalized answers.

2. The method according to claim 1, characterized in that Using a named entity recognition algorithm to extract professional terms and domain concepts from the preprocessed data to form an entity set, analyzing the relationship between entities in the entity set based on a semantic role labeling technique, constructing a relationship set, and combining the entity set and the relationship set to form an initial knowledge graph includes: Construct a notarization field annotated dataset, and train a notarization entity recognition deep fusion network based on the annotated dataset. The notarization entity recognition deep fusion network includes a BERT-based encoding layer, a bidirectional long short-term memory network-based feature extraction layer, and a conditional random field-based decoding layer. The notarization field text is subjected to named entity recognition through the notarization entity recognition deep fusion network to obtain an entity set, which includes entity name, entity type, occurrence frequency, and first occurrence position. Performing dependency syntactic analysis on sentences containing entities in the entity set to identify subject-predicate-object structures, using a semantic role labeling model to perform semantic role labeling on sentences containing the subject-predicate-object structures to obtain a predicate-argument structure, and determining a semantic role labeling label; Constructing a mapping rule set from semantic role annotation labels to notarization domain relationship types, and converting semantic role annotation results into candidate relationships according to the mapping rule set; Constructing a relationship classifier and training it with a preset remote supervision data set, using the relationship classifier to score and filter candidate relationships to obtain a relationship set; The entities in the entity set are imported as entity nodes into the graph database, and the relationships in the relationship set are imported as relationship edges into the graph database, the corresponding entity nodes are connected, and the imported entity nodes and relationship edges are optimized using a graph algorithm to eliminate redundant relationships and merge synonymous entities, so as to construct an initial knowledge graph.

3. The method according to claim 2, characterized in that A notarization field annotated dataset is constructed, and a notarization entity recognition deep fusion network is trained based on the annotated dataset. The notarization entity recognition deep fusion network includes a BERT-based encoding layer, a bidirectional long short-term memory network-based feature extraction layer, and a conditional random field-based decoding layer. Named entity recognition is performed on notarization field texts through the notarization entity recognition deep fusion network, and the entity set obtained includes: Constructing a notarization field annotated dataset, and pre-training the BERT model with a masked language model based on the notarization field annotated dataset to obtain a pre-trained BERT model adapted to the notarization field; In the notarization entity recognition deep fusion network, the encoding layer is initialized using the weights of the pre-trained BERT model; the notarization field annotated dataset is input into the notarization entity recognition deep fusion network, and the word vector representation is obtained through the encoding layer; Input the word vector representation into a feature extraction layer to extract sequence features, wherein the feature extraction layer includes a forward long short-term memory network and a backward long short-term memory network, each including an input gate, a forget gate, and an output gate; Inputting the sequence features into the decoding layer, the decoding layer includes a linear chain conditional random field, establishing a label transfer probability matrix, and using the Viterbi algorithm to decode to obtain an optimal label sequence; A joint learning strategy is adopted to optimize the encoding layer, feature extraction layer, and decoding layer at the same time. The AdamW optimizer is used to update the parameters, and the learning rate warm-up and linear decay strategies are implemented. After each training round, the performance of the publicized entity recognition deep fusion network is evaluated using a preset validation set, and the model weight with the highest F1 score on the validation set is saved in combination with an early stopping strategy to obtain a trained publicized entity recognition deep fusion network; Input the notarization domain text to be recognized into the trained notarization entity recognition deep fusion network, and sequentially pass through the encoding layer, feature extraction layer and decoding layer to obtain a label sequence; According to the tag sequence, entities in the notarization domain text are identified and an entity set is constructed.

4. The method according to claim 1, characterized in that: Based on the dynamic update mechanism of time decay, knowledge is extracted from the newly added pre-processed data and added to the initial knowledge graph to form a dynamic knowledge graph including: Constructing a time decay function, wherein the time decay function adopts an exponential decay form to control the decay speed of the knowledge unit; Receive newly added preprocessed data, and perform entity recognition and relationship extraction on the newly added preprocessed data to obtain a new knowledge unit, use the cosine similarity method to calculate the similarity between the new knowledge unit and the existing knowledge unit in the initial knowledge graph, and obtain a similarity calculation result, and determine whether the similarity calculation result is higher than a preset similarity threshold based on the similarity calculation result: When the similarity calculation result is higher than the preset similarity threshold, the existing knowledge unit is updated with the new knowledge unit, and the fusion weight of the new knowledge unit and the existing knowledge unit is calculated using the time decay function; based on the fusion weight, the new knowledge unit and the existing knowledge unit are weightedly fused to obtain a fused knowledge unit; When the similarity is lower than the preset similarity threshold, the new knowledge unit is used as a newly added knowledge unit; Update the fused knowledge unit to the corresponding position, add the newly added knowledge unit to the initial knowledge graph, update the initial knowledge graph, and record the timestamp of each update or addition; According to a preset full scan cycle, the initial knowledge graph is regularly scanned, the time decay value of each knowledge unit is calculated using the time decay function, and it is determined whether the time decay value of each knowledge unit is less than a preset decay threshold value. When the time decay value corresponding to a knowledge unit is less than the preset decay threshold value, the corresponding knowledge unit is removed from the initial knowledge graph; The initial knowledge graph is updated to form a dynamic knowledge graph.

5. The method according to claim 1, characterized in that Receive multimodal questions input by users, extract voice text, image text and image semantic information based on the voice and image in the multimodal questions, combine with the plain text in the multimodal questions, perform multimodal feature fusion, determine the fusion features, perform deep semantic understanding on the fusion features, obtain semantic features, input the semantic features into the intent recognition model, adapt the predefined notarization business intent categories through semantic mapping, and determine that the user intent includes: Receiving a multimodal question, wherein the multimodal question includes voice information, image information, and text information; Using a speech recognition model to convert speech information in the multimodal problem into a first text sequence; using an optical character recognition engine to process image information in the multimodal problem to obtain a second text sequence; splicing the first text sequence, the second text sequence and the text information in the multimodal problem to form a combined text sequence, and mapping each word in the combined text sequence to a specified dimensional space through a first independent linear layer to obtain text features; Extracting semantic features of the image information in the multimodal problem using a visual Transformer model to obtain an image semantic vector, and mapping the image semantic vector to a corresponding dimensional space having the same dimension as the text feature through a second independent linear layer to obtain an image feature; Using a multi-head attention mechanism to calculate the correlation between the text feature and the image feature to obtain an attention output, performing a residual connection and layer normalization on the attention output and the text feature to obtain a fused feature, and inputting the fused feature into a pre-trained bidirectional encoder to obtain a context-related representation; extracting a tag vector from the context-related representation, and mapping it to a task-corresponding space through a linear layer and an activation function to obtain a semantic feature; The semantic features are mapped to the predefined notarization business intention category space through a fully connected layer to obtain the intention logic value, the sigmoid function is applied to the intention logic value, the probability of each intention category is calculated, the intention category with the highest probability is selected, and the user intention is determined.

6. The method according to claim 5, characterized in that Based on the user intention and the dynamic knowledge graph, the semantic features are entity linked and relations are extracted to form a structured query expression including: Constructing a pre-trained stacked network model, wherein the stacked network model adopts a permutation language modeling training objective, uses a two-stream self-attention mechanism and a segmented recurrence mechanism; Inputting semantic features into the stacked network model to obtain bidirectional context-aware word vector representations; constructing a multi-task learning architecture on the top layer of the stacked network model, including entity recognition tasks and relationship extraction tasks; introducing a relative position encoding mechanism into the multi-task learning architecture to capture long-distance dependencies; using a dynamic weight allocation strategy to adaptively adjust the importance weights of entity recognition tasks and relationship extraction tasks in the multi-task learning architecture; Integrate a dynamic knowledge graph, process the dynamic knowledge graph through a graph attention network, obtain enhanced entity and enhanced relationship representations in the word vector representation, and form an enhanced word vector representation; Apply adversarial training technology to the enhanced word vector representation, train to obtain the final multi-task learning architecture, input semantic features, and identify corresponding entity mentions; calculate the similarity between the entity mentions and the entities in the dynamic knowledge graph, select the entity with the highest similarity as the link result, obtain a linked entity set, and generate candidate relationship pairs based on the entity pairs in the linked entity set; Based on the relationship extraction task in the multi-task learning architecture, taking the semantic features and the candidate relationship pairs as input, predicting the relationship between entity pairs to obtain a relationship set; A corresponding query template is selected based on the user's intention, and the linked entity set and the relationship set are filled into the query template to generate a structured query expression.

7. The method according to claim 1, characterized in that Based on user intent and structured query expressions, multi-hop reasoning is performed in the dynamic knowledge graph, similarity is calculated using graph vectors, and the reasoning path is determined in combination with user intent. Knowledge association reasoning is performed along the reasoning path, and preliminary reasoning results are obtained, including: Receive user intent and structured query expressions as input, determine the inference starting entity in the dynamic knowledge graph, take the inference starting entity as the current node, identify entities and relationships directly connected to the current node, and form a candidate next hop set; Using the pre-generated graph vector, calculate the similarity score between the current node and each node in the candidate next hop set; convert the user intention into a vector representation, and fuse it with the graph vector of each node in the candidate next hop set to obtain a fusion vector; combine the similarity score with the fusion vector to generate a comprehensive score for each candidate next hop node; Based on the comprehensive score, a next hop is selected using a reinforcement learning agent, wherein the reinforcement learning agent takes a current state as input and outputs a probability distribution of selecting the next hop, wherein the current state includes the current node information, the user intention, the comprehensive score, and the selected path information; According to the selected next hop, update the current node and the selected path information; Applying predefined reasoning rules to the selected path to perform rule-based explicit reasoning; Encoding the selected path using a graph neural network model and performing implicit reasoning based on the graph neural network, wherein the graph neural network model is a graph convolutional network; The results of explicit reasoning and implicit reasoning are integrated to obtain a fused reasoning result; Until the preset hop limit is reached, the path selection in the entire reasoning process is recorded to form a complete reasoning chain; based on the selected path length, the similarity score of each hop and the reliability score of the reasoning rule used, the confidence score of the entire reasoning result is calculated; The complete reasoning chain, the final node reached, and the confidence score are combined to form a preliminary reasoning result.

8. A notarized intelligent question-answering customer service system based on knowledge graph, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to collect multi-source heterogeneous data in the notarization field, pre-process the multi-source heterogeneous data to obtain pre-processed data, extract professional terms and domain concepts from the pre-processed data using a named entity recognition algorithm to form an entity set, analyze the relationship between entities in the entity set based on a semantic role labeling technology, construct a relationship set, combine the entity set and the relationship set to form an initial knowledge graph, extract knowledge from the newly added pre-processed data and add it to the initial knowledge graph based on a dynamic update mechanism of time decay to form a dynamic knowledge graph, and vectorize the dynamic knowledge graph to obtain a graph vector; The second unit is used to receive a multimodal question input by a user, extract the voice text, image text and image semantic information based on the voice and image in the multimodal question, combine the plain text in the multimodal question, perform multimodal feature fusion, determine the fusion feature, perform deep semantic understanding on the fusion feature, obtain the semantic feature, input the semantic feature into the intention recognition model, adapt the predefined notarization business intention category through semantic mapping, determine the user intention, and perform entity linking and relationship extraction on the semantic feature based on the user intention combined with the dynamic knowledge graph to form a structured query expression; The third unit is used to perform multi-hop reasoning in a dynamic knowledge graph based on user intent and structured query expressions, use graph vectors to calculate similarity, and determine the reasoning path in combination with user intent. It performs knowledge association reasoning along the reasoning path to obtain preliminary reasoning results, performs interpretable analysis on the preliminary reasoning results, generates reasoning chains and confidence scores, and forms detailed reasoning results. It combines pre-acquired user portrait information and user intent to perform personalized sorting on the detailed reasoning results to obtain ordered reasoning results, inputs the ordered reasoning results into a pre-built Transformer model, outputs preliminary natural language answers, and applies text style transfer technology to finally generate personalized answers.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Data intelligent question and answer method and system fusing domain knowledge

    CN118779438A

  • DCS intelligent decision-making method and system fusing large language model and knowledge graph

    CN118820778A

  • Information interaction method, apparatus and device and storage medium

    WO2022222286A1

  • Reading type examination question generation system and method based on commonsense reasoning

    WO2023225858A1

Cited By

  • Game strategy retrieval method and device based on event-driven knowledge graph embedding

    CN120104814A

  • Game strategy retrieval method and device based on event-driven knowledge graph embedding

    CN120104814B

  • Rule base dynamic construction method and device based on large language model and medium

    CN120179811A

  • Customer service data quality inspection method and device based on dynamic reasoning, equipment and medium

    CN120216707A

  • Customer service data quality inspection method, device, equipment and medium based on dynamic reasoning

    CN120216707B