Expressway toll auxiliary question and answer method and system based on knowledge graph
By building a knowledge graph-based highway toll collection auxiliary question-answering system, the problems of information missing and semantic alignment in intelligent question-answering systems when dealing with complex problems are solved, and the accuracy and speed of answer output are improved.
Patent Information
- Application Number
- CN202511211926.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-28
AI Technical Summary
When dealing with complex and highly professional questions, existing intelligent question-answering systems suffer from information missing and semantic misalignment, resulting in reduced answer accuracy.
Build a knowledge graph-based highway toll auxiliary question-answering system. By obtaining information related to highway tolls, build a domain knowledge graph and perform knowledge enhancement, combine it with a large language model to output answers, and select different output methods according to the difficulty of the question.
It improves the accuracy and speed of answer output and meets the needs of high-precision and in-depth information.
Smart Images

Figure CN120744141A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of assisted question-answering, and specifically relates to a knowledge graph-based highway toll assisted question-answering method and system. Background Art
[0002] With the rapid development of the information age, the scale of data and knowledge has grown exponentially, making the storage, management, and retrieval of knowledge increasingly complex. Against this backdrop, the rapid advancement of artificial intelligence (AI) technology has created unprecedented opportunities for the research and application of intelligent question-answering systems. These systems aim to understand users' natural language questions and provide accurate and efficient answers. Simultaneously, advances in AI technology have provided new opportunities for the development of intelligent question-answering systems. However, as data continues to grow, traditional question-answering systems continue to face numerous challenges in handling complex and highly specialized questions. These challenges include knowledge fragmentation, inconsistent data structures, and inefficient information retrieval, making it difficult to meet the demand for high-precision, in-depth information.
[0003] In addressing these challenges, knowledge graphs and large language models have emerged as the two core pillars of intelligent question-answering technology. Knowledge graphs organize and connect knowledge in a structured manner, empowering question-answering systems with stronger logical reasoning capabilities, thereby providing accurate and verifiable answers. Large language models, with their exceptional natural language understanding and generation capabilities, can parse complex contexts and generate coherent and fluent answers.
[0004] However, in the actual process of intelligent question answering, information is easily missing and semantics cannot be aligned during the interaction of data, which leads to large deviations in the model's answer output. The answer actually output by the model does not correspond to the target question or a counterfactual answer appears, which greatly reduces the accuracy of the answer. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention provides a highway toll collection auxiliary question-answering method and system based on knowledge graph, which is used to solve the technical problems in the existing technology.
[0006] In a first aspect, the present invention provides the following technical solution: a knowledge graph-based highway toll collection auxiliary question-answering method, comprising: Obtaining public information related to highway tolls, and constructing a highway toll information dataset based on the public information; Constructing a domain knowledge graph based on the highway toll information dataset; Performing knowledge enhancement on the domain knowledge graph to obtain an enhanced knowledge graph; Obtaining a question text related to highway tolls sent by a user, performing complex analysis on the question text, and outputting complex analysis results; If the result of the complex analysis is that the question text is a question of simple difficulty, a first answer is output for the question text based on the enhanced knowledge graph to output the target answer. If the result of the complex analysis is that the question text is a question of medium difficulty, a second answer is output for the question text based on the enhanced knowledge graph to output the target answer. If the result of the complex analysis is that the question text is a question of complex difficulty, a third answer is output for the question text based on the enhanced knowledge graph to output the target answer.
[0007] Compared with the existing technology, the beneficial effects of the present invention are: the present invention constructs a domain knowledge graph based on the highway toll information data set, which can effectively extract deep information in the data to highlight the professional background information and spatiotemporal constraint information of the knowledge graph, and then enhances the knowledge graph to improve the integrity of the knowledge graph and enhance the entity representation, thereby improving the accuracy of the answer output. Then, by determining the complexity of the question text and selecting different methods to output the answer, the speed of answer output can be increased while improving the effectiveness and accuracy of the answer output.
[0008] Preferably, the step of constructing a highway toll information dataset based on the public information includes: Converting the public information into text data, and removing punctuation, repeated text, stop words, and irrelevant tags from the text data to obtain clean text data; The clean text data is segmented to obtain a segmented data set, and the segmented data set is annotated with time and space according to preset annotation rules to obtain a highway toll information data set.
[0009] Preferably, the step of constructing a domain knowledge graph based on the highway toll information dataset includes: Performing word embedding representation on the text words and the corresponding annotation information in the highway toll information dataset to obtain text word embedding vectors and annotation embedding vectors respectively; Combine the attention mechanism and the bidirectional long short-term memory network to embed the text word vector , the annotation embedding vector Perform semantic feature extraction to obtain text features and annotation features : ; ; Where, is a bidirectional long short-term memory network, is the attention mechanism; The text features and the annotation features are sequentially fused and decoded to obtain a sentence sequence set; Inputting the sentence sequence set into the BERT model and determining a sentence vector sequence by forward propagation, inputting the sentence vector sequence into the bidirectional long short-term memory network, concatenating the forward and reverse hidden states at each moment to obtain a forward final state and a reverse final state, and concatenating the forward final state and the reverse final state to obtain a sentence sequence embedding representation; Determine the sentence relationship probability by using the sentence sequence embedding representation as a classification feature of the text relationship to output the sentence relationship probability, and determine the sentence relationship based on the magnitude of the sentence relationship probability; Determine an entity feature vector based on the sentence sequence set : ; Where, Indicates context, Characteristic words The feature vector in the context, Indicates the number of feature words in the context, is a set of statement sequences, is a sequence of sentence vectors, They are sentence sequence weight, sentence vector weight, and context feature weight respectively; For the entity feature vector The entities in are similarity calculated for each other, and the entities with similarity greater than the similarity threshold are merged to obtain the standard entity vector; The entity attributes in the standard entity vector are obtained, and the entity attributes, standard entity vector and sentence relationship are visualized in the form of a knowledge graph to obtain a domain knowledge graph.
[0010] Preferably, the step of performing knowledge enhancement on the domain knowledge graph to obtain an enhanced knowledge graph includes: Obtain a template question, and embed the template question and the domain knowledge graph to obtain question embedding and entity embedding; Mapping the question embedding into the same complex space as the entity embedding to obtain a mapped question embedding; The remaining entities in the domain knowledge graph except the entities involved in the template question are taken as candidate entities, and the first score value of each candidate entity is calculated. : ; Where, 、 Respectively represent entity embeddings in complex space, for The complex conjugate of For the The mapping problem in the complex space is embedded. is the number of complex spaces, is the real part of the complex number; Arrange the plurality of entities to be selected in descending order according to the first scoring value, and select the first plurality of entities to be selected as candidate entities; Map the question embedding to the sentence relationship in the domain knowledge graph to obtain the question relationship embedding , encode and embed the sentence relationship to obtain the relationship embedding , based on the problem relationship embedding, determine the final problem mapping relationship : ; Where, For the The embedding of the sentence relationship corresponding to the candidate entity, is the screening threshold, for The first relationship among Based on the final problem mapping relationship Determine the enhanced knowledge graph.
[0011] Preferably, the mapping relationship based on the final question The steps to determine the enhanced knowledge graph include: Based on the final problem mapping relationship Calculate the second score : ; Where, is a tunable hyperparameter, for Middle The embedding of the sentence relationship corresponding to the candidate entity, is the embedding of the shortest path from the entity in the template problem to each candidate entity; Arrange the candidate entities in descending order according to the second scoring value, and select the first several candidate entities as target entities; Extracting evidence text related to the template question from the highway toll information dataset to obtain an evidence pool, and performing semantic parsing on the evidence questions in the evidence pool to obtain an abstract concept evidence graph; The template question and the target entity are converted into an abstract concept entity graph, the abstract concept evidence graph and the same concept nodes in the abstract concept entity graph are fused to obtain an abstract concept graph, and the abstract concept graph is fused into the domain knowledge graph to obtain an enhanced knowledge graph.
[0012] Preferably, the step of performing complex analysis on the question text to output complex analysis results includes: Perform dependency parsing on the question text to obtain the part-of-speech tag, word vector representation, and word dependency relationship of each word; Convert part-of-speech tags, word vector representations, and word dependencies into a graph structure, interpreting word vector representations as nodes, word dependencies as edges, and part-of-speech tags as nodes. The nodes in the graph structure are input into the multi-layer CCN for update aggregation to obtain the updated node feature representation : ; Where, is the activation function, The first The set of neighbor nodes of a node, 、 Respectively The learnable weights and biases of the layers, For the The updated node feature representation of the layer, is the normalization coefficient; Calculate and update node feature representation The maximum depth and span variance of the node are expressed as follows: Perform concatenation, full connection layer processing, and activation function processing in sequence to output a first complex score; Performing word segmentation and encoding processing on the question text to obtain a context representation, calculating a sentence perplexity based on the context representation to obtain a sentence perplexity, and normalizing the sentence perplexity to obtain a second complexity score; Performing a weighted fusion of the first complexity score and the second complexity score to obtain a final complexity score; If the final complexity score is less than the first scoring threshold, the question text is a question of simple difficulty; if the final complexity score is not less than the first scoring threshold and less than the second scoring threshold, the question text is a question of medium difficulty; if the final complexity score is not less than the second scoring threshold, the question text is a question of complex difficulty.
[0013] Preferably, the step of outputting a first answer to the question text based on the enhanced knowledge graph to output a target answer includes: Extracting entities and relations from the question text, and matching the extracted entities with relations to obtain original triples; Quickly matching the original triples in the enhanced knowledge graph to output target triples; The target triple is output as a grammatically correct and complete natural language sentence according to a preset sentence template to obtain the target answer.
[0014] Preferably, the step of outputting a second answer to the question text based on the enhanced knowledge graph to output a target answer includes: Obtain a reference text, divide the reference text into a plurality of text blocks, and vectorize the text blocks and the triples in the enhanced knowledge graph to obtain text vectors and graph vectors; Extracting original triples of the question text and vectorizing the original triples to obtain original vectors; The text vector, the graph vector, and the original vector are fused into prompt information, and the prompt information is input into a large language model for answer output to obtain a target answer.
[0015] Preferably, the step of outputting a third answer to the question text based on the enhanced knowledge graph to output a target answer includes: Obtaining an answer output model and a training data set, extracting a training question from the training data set, extracting a first triple of the training question and extracting a corresponding second triple from the enhanced knowledge graph, and combining the first triple with the second triple to obtain combined data; Extracting semantic features of the combined data through an attention mechanism to obtain combined features; Add the combined data and the combined features to the dataset to be processed, input the dataset to be processed into the answer output model and perform bias score calculation to obtain a bias score : ; Where, Indicates that only the combined data in the dataset to be processed is input into the answer output model and the answer distribution outputted after prediction by the answer output model; Calculate the observation results of the dataset to be processed under different attention and counterfactual outcomes : ; ; Where, is a linear layer, They are the interference corresponding to the first attention and the second attention respectively; Based on the observation results , the counterfactual result With the bias score Calculating model bias : ; Where, is the number of data in the dataset to be processed; Based on the model deviation value Determine the loss function : ; Where, represents the cross entropy between the model output and the true answer, is the model deviation value Cross entropy between the attribute and the true answer; The answer output model is optimized and trained by minimizing the loss function to obtain an optimized model, and the question text is input into the optimized model to output a target answer.
[0016] In a second aspect, the present invention provides the following technical solution: a knowledge graph-based highway toll collection auxiliary question-answering system, the system comprising: A construction module, configured to obtain public information related to highway toll collection and construct a highway toll collection information dataset based on the public information; A graph module, configured to construct a domain knowledge graph based on the highway toll information dataset; An enhancement module, configured to perform knowledge enhancement on the domain knowledge graph to obtain an enhanced knowledge graph; An analysis module is used to obtain a question text related to highway tolls sent by a user, perform complex analysis on the question text, and output complex analysis results; An output module is used to output a first answer to the question text based on the enhanced knowledge graph to output a target answer if the complex analysis result shows that the question text is a question of simple difficulty; to output a second answer to the question text based on the enhanced knowledge graph to output a target answer if the complex analysis result shows that the question text is a question of medium difficulty; and to output a third answer to the question text based on the enhanced knowledge graph to output a target answer if the complex analysis result shows that the question text is a question of complex difficulty.
[0017] In a third aspect, the present invention provides the following technical solution: a computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor; when the processor executes the computer program, the knowledge graph-based highway toll collection auxiliary question-and-answer method as described above is implemented.
[0018] In a fourth aspect, the present invention provides the following technical solution: a storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned knowledge graph-based highway toll collection auxiliary question-answering method. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 This is a flow chart of the knowledge graph-based highway toll collection auxiliary question-answering method provided in Example 1 of the present invention; Figure 2 This is a structural block diagram of the knowledge graph-based highway toll collection auxiliary question-answering system provided in the second embodiment of the present invention; Figure 3 A schematic diagram of the hardware structure of a computer provided in another embodiment of the present invention.
[0021] The embodiments of the present invention will be further described below with reference to the accompanying drawings. DETAILED DESCRIPTION
[0022] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the embodiments of the present invention, and should not be construed as limiting the present invention.
[0023] Example 1 In the first embodiment of the present invention, Figure 1 As shown in FIG, a knowledge graph-based highway toll collection auxiliary question answering method includes: S1. Obtaining public information related to highway toll collection, and constructing a highway toll collection information dataset based on the public information; Specifically, the public information here refers to the public information related to highway toll collection, such as the highway section, distance, toll standard, toll collection time, etc.
[0024] Wherein, the step S1 includes: S11, converting the public information into text data, and removing punctuation marks, repeated text, stop words, and irrelevant tags in the text data to obtain clean text data; Specifically, after converting public information into text data, there may be missing fields, repeated information, or unrecognizable characters. Therefore, you can use functions in Python to remove punctuation, repeated text, stop words, and irrelevant tags from the text data to obtain clean text data.
[0025] S12. Performing word segmentation processing on the clean text data to obtain a word segmentation dataset, and performing time and space annotation on the word segmentation dataset according to preset annotation rules to obtain a highway toll information dataset; Specifically, the time annotation here is mainly used to annotate the data in the word segmentation dataset by time, such as the specific time it takes to travel a section of highway, the opening and closing times of the highway, the toll collection time, etc. The spatial annotation is specifically used to annotate the location of the highway.
[0026] S2. Constructing a domain knowledge graph based on the highway toll information dataset; Wherein, the step S2 includes: S21, performing word embedding representation on the text words and the corresponding annotation information in the highway toll information dataset to obtain text word embedding vectors and annotation embedding vectors respectively; Specifically, the word embedding representation here can be implemented through the BERT model.
[0027] S22, combine the attention mechanism and the bidirectional long short-term memory network to embed the text word vector , the annotation embedding vector Perform semantic feature extraction to obtain text features and annotation features : ; ; Where, is a bidirectional long short-term memory network, is the attention mechanism; Specifically, in this step, adding an attention mechanism on the basis of the bidirectional long short-term memory network can effectively capture the semantic information of keywords in the sentence, that is, it can better focus on the time entities and place entities in the article.
[0028] S23, fusing and decoding the text features and the annotation features in sequence to obtain a sentence sequence set; Specifically, the fusion process here can be implemented through the Transformer layer, and the encoding process here is specifically the CRF encoding process. During the decoding process, the label sequence with the highest score will be selected as the optimal path for parsing, and then the corresponding type of entity will be extracted according to the label set.
[0029] S24. Input the sentence sequence set into the BERT model and determine a sentence vector sequence by forward propagation, input the sentence vector sequence into the bidirectional long short-term memory network, concatenate the forward and reverse hidden states at each moment to obtain a forward final state and a reverse final state, and concatenate the forward final state and the reverse final state to obtain a sentence sequence embedding representation; Specifically, the BERT model here is a pre-trained model.
[0030] S25, using the sentence sequence embedding representation as a classification feature of the text relationship to determine the sentence relationship probability, to output the sentence relationship probability, and determining the sentence relationship based on the magnitude of the sentence relationship probability; Specifically, the sentence sequence is embedded as a classification feature of the text relationship, and then input into the fully connected layer and the softmax classification layer to obtain the probability distribution of the sentence relationship, that is, the sentence relationship probability, and the sentence relationship corresponding to the maximum sentence relationship probability is used as the output sentence relationship.
[0031] S26: Determine entity feature vector based on the sentence sequence set : ; Where, Indicates context, Characteristic words The feature vector in the context, Indicates the number of feature words in the context, is a set of statement sequences, is a sequence of sentence vectors, They are sentence sequence weight, sentence vector weight, and context feature weight respectively; Specifically, the sum of the sentence sequence weight, sentence vector weight, and context feature weight is 1.
[0032] S27, the entity feature vector The entities in are similarity calculated for each other, and the entities with similarity greater than the similarity threshold are merged to obtain the standard entity vector; Specifically, in actual situations, entities may be repeated, so by merging entities with high similarity into one, a standard entity vector can be obtained.
[0033] S28. Obtain entity attributes in the standard entity vector, and visualize the entity attributes, standard entity vector, and sentence relationships in the form of a knowledge graph to obtain a domain knowledge graph; Specifically, the extracted entity attributes, standard entity vectors, and sentence relationships are visualized in the form of a graph, where concepts such as entities and attributes are presented in the form of nodes, while relationships are presented in the form of edges. This method can better reflect the structure and logical relationships of knowledge, and can perform effective information retrieval and knowledge reasoning.
[0034] S3. performing knowledge enhancement on the domain knowledge graph to obtain an enhanced knowledge graph; Wherein, the step S3 includes: S31. Obtain a template question, and embed the template question and the domain knowledge graph to obtain question embedding and entity embedding; Specifically, the question embedding here reveals the relational semantics. For the template question, the entities and questions in it can form triples with the candidate entities in the domain knowledge graph.
[0035] S32, mapping the question embedding into the same complex space as the entity embedding to obtain a mapped question embedding; Specifically, by mapping to the complex space, the complex vectors used can preserve the properties of the dot product and have linear space and time complexity.
[0036] S33: Calculate the first score of each of the entities in the domain knowledge graph except the entities involved in the template question as candidate entities. : ; Where, 、 Respectively represent entity embeddings in complex space, for The complex conjugate of For the The mapping problem in the complex space is embedded. is the number of complex spaces, is the real part of the complex number.
[0037] S34, arranging the plurality of entities to be selected in descending order according to the first scoring values, and selecting the first plurality of entities to be selected as candidate entities; Specifically, for determining candidate entities, the number thereof is much smaller than the number of entities to be selected.
[0038] S35. Map the question embedding to the sentence relationship in the domain knowledge graph to obtain the question relationship embedding , encode and embed the sentence relationship to obtain the relationship embedding , based on the problem relationship embedding, determine the final problem mapping relationship : ; Where, For the The embedding of the sentence relationship corresponding to the candidate entity, is the screening threshold, for The first relationship among Specifically, for the above formula, by selecting relationships greater than or equal to the screening threshold from all relationships, if it is less than the screening threshold, only the relationship ranked first is selected.
[0039] S36, based on the final question mapping relationship Determine the enhanced knowledge graph; Wherein, the step S36 includes: S361, based on the final question mapping relationship Calculate the second score : ; Where, is a tunable hyperparameter, for Middle The embedding of the sentence relationship corresponding to the candidate entity, is the embedding of the shortest path from the entity in the template problem to each candidate entity.
[0040] S362: Arrange the candidate entities in descending order according to the second scoring value, and select the first several candidate entities as target entities.
[0041] S363: Extracting evidence text related to the template question from the highway toll information dataset to obtain an evidence pool, and performing semantic analysis on the evidence questions in the evidence pool to obtain an abstract concept evidence graph; Among them, in the process of determining the evidence pool, the question words in the question are directly replaced with the target entity to obtain a declarative statement, and the target entity is directly output. For the target entity corresponding to each question, several statements are generated, and the text sentences related to the declarative statement are extracted one by one as evidence text. Then, with the assistance of numerous evidence texts, the first several evidence texts are screened out and stored in the evidence pool. At the same time, the semantic structure of the evidence is represented by abstract concept nodes and relationship edges, that is, the abstract concept evidence graph presents the semantics of the evidence in a graphical way, in which the nodes are high-level abstractions of the corresponding text concepts.
[0042] S364: converting the template question and the target entity into an abstract concept entity graph, fusing the same concept nodes in the abstract concept evidence graph and the abstract concept entity graph to obtain an abstract concept graph, and fusing the abstract concept graph into the domain knowledge graph to obtain an enhanced knowledge graph; Specifically, for abstract concept entity graphs, by merging the same nodes, the coherence of evidence association is strengthened and the reasoning process between questions and candidate entities is clarified.
[0043] S4. Obtaining a question text related to highway tolls sent by a user, performing complex analysis on the question text, and outputting a complex analysis result; Wherein, the step S4 includes: S41, performing dependency syntax parsing on the question text to obtain a part-of-speech tag, a word vector representation, and word dependencies for each word; Specifically, dependency syntax parsing can be performed through natural language processing tools.
[0044] S42, converting the part-of-speech tags, word vector representations, and word dependencies into a graph structure, interpreting the word vector representations as nodes, the word dependencies as edges, and the part-of-speech tags as nodes; S43: Input the nodes in the graph structure into the multi-layer CCN for update aggregation to obtain updated node feature representation : ; Where, is the activation function, The first The set of neighbor nodes of a node, 、 Respectively The learnable weights and biases of the layers, For the The updated node feature representation of the layer, is the normalization coefficient; Specifically, in this step, as the number of network layers is gradually updated, a new representation is obtained by aggregating information with the neighboring nodes, that is, an updated node feature representation.
[0045] S44. Calculate and update node feature representation The maximum depth and span variance of the node are expressed as follows: Perform concatenation, full connection layer processing, and activation function processing in sequence to output a first complex score; Specifically, the edge features are mapped to the hidden space through linear transformation and mapped to the same dimension as the node features so that the information of nodes and edges can be fused in the graph convolution layer. The input node features are transformed and aggregated respectively through the GCN layer to capture the local syntactic structure information. Then the attention mechanism is introduced to capture the global syntactic relationship and importance difference. The maximum depth and span variance of the node features in the graph are calculated through the linear layer as the global features of the syntactic structure. The maximum depth, the span variance and the updated node feature representation are combined. After splicing, it passes through the fully connected layer and activation function, and finally outputs the first complex score.
[0046] S45. Segmenting and encoding the question text to obtain a context representation, calculating a sentence perplexity based on the context representation to obtain a sentence perplexity, and normalizing the sentence perplexity to obtain a second complexity score; Specifically, the word segmentation process is first performed through the word segmenter, and then input into the Transformer layer for encoding to obtain the context representation. Then, the perplexity calculation process of the masked language model (MLM) is performed. The vocabulary distribution of the masked position is calculated through the BERT model to generate the overall sentence perplexity, which is normalized to the second complex score.
[0047] S46: Perform weighted fusion on the first complexity score and the second complexity score to obtain a final complexity score.
[0048] S47. If the final complexity score is less than the first scoring threshold, the question text is a simple difficulty question; if the final complexity score is not less than the first scoring threshold and less than the second scoring threshold, the question text is a medium difficulty question; if the final complexity score is not less than the second scoring threshold, the question text is a complex difficulty question.
[0049] S5. If the result of the complex analysis is that the question text is a question of simple difficulty, a first answer is output for the question text based on the enhanced knowledge graph to output a target answer. If the result of the complex analysis is that the question text is a question of medium difficulty, a second answer is output for the question text based on the enhanced knowledge graph to output a target answer. If the result of the complex analysis is that the question text is a question of complex difficulty, a third answer is output for the question text based on the enhanced knowledge graph to output a target answer.
[0050] The step of outputting a first answer to the question text based on the enhanced knowledge graph to output a target answer includes: S511 , extracting entities and relations from the question text, and matching the extracted entities with relations to obtain original triples.
[0051] S512: Quickly match the original triples in the enhanced knowledge graph to output target triples; Specifically, if the original triple cannot be matched in the enhanced knowledge graph, the process proceeds to the second answer output step.
[0052] S513: Output the target triple into a grammatically correct and complete natural language sentence according to a preset sentence template to obtain a target answer; Specifically, the retrieved target triples are input and a grammatically correct and complete natural language sentence is output. In this process, sentence templates are predefined, entities and relations are dynamically inserted, templates are selected according to the relationship type, and the entity list is converted into natural language to obtain the target answer.
[0053] The step of outputting a second answer to the question text based on the enhanced knowledge graph to output a target answer includes: S521. Obtain a reference text, divide the reference text into several text blocks, and vectorize the text blocks and the triples in the enhanced knowledge graph respectively to obtain text vectors and graph vectors.
[0054] S522: Extract original triples of the question text and vectorize the original triples to obtain original vectors.
[0055] S523: Fusing the text vector, the graph vector, and the original vector into prompt information, and inputting the prompt information into a large language model to output an answer, thereby obtaining a target answer; Specifically, the Large Language Model (LLM) is trained based on a large amount of text data and can generate natural and fluent language text. In order to adapt to the knowledge of a specific field, the model will use text data from related fields for fine-tuning, so that it can generate more professional and relevant answers. The model needs to understand the contextual relationship between the input user query and the retrieved information, and ensure that the generated answer is relevant to the question by analyzing the key entities and relationships in the query and the relationship with the retrieved information. The model effectively integrates the retrieved prompt information to ensure that the answer contains relevant entities, relationships and other background information. The model generates coherent natural language text based on the input information. The generated answer not only needs to answer the user's question, but also should conform to the grammatical and semantic logic of the language to make the answer more natural.
[0056] The step of outputting a third answer to the question text based on the enhanced knowledge graph to output a target answer includes: S531. Obtain an answer output model and a training data set, extract a training question from the training data set, extract a first triple of the training question and extract a corresponding second triple from the enhanced knowledge graph, and combine the first triple with the second triple to obtain combined data. Specifically, the answer output model here is the LLM model.
[0057] S532. Extracting semantic features of the combined data through an attention mechanism to obtain combined features.
[0058] S533: Add the combined data and the combined features to the dataset to be processed, input the dataset to be processed into the answer output model and perform bias score calculation to obtain a bias score. : ; Where, Indicates that only the combined data in the dataset to be processed is input into the answer output model and the answer distribution outputted after prediction by the answer output model; Specifically, by allocating attention to the corresponding samples to focus on the combined data, ignoring the combined features to evaluate the bias impact of individual question branches on the model, and further obtaining the bias score on the answer distribution.
[0059] S534, calculating the observation results of the to-be-processed data set under different attention levels and counterfactual outcomes : ; ; Where, is a linear layer, They are the interference corresponding to the first attention and the second attention respectively; S535, based on the observation results , the counterfactual result With the bias score Calculating model bias : ; Where, is the number of data in the dataset to be processed; Specifically, by calculating the model deviation value, the participation of important combination features is increased to eliminate the influence of the deviation.
[0060] S536: Based on the model deviation value Determine the loss function : ; Where, represents the cross entropy between the model output and the true answer, is the model deviation value The cross entropy between the attribute and the true answer.
[0061] S537. Optimize and train the answer output model by minimizing the loss function to obtain an optimized model, and input the question text into the optimized model to output a target answer.
[0062] The knowledge graph-based highway toll auxiliary question-answering method provided in the first embodiment of the present invention constructs a domain knowledge graph based on the highway toll information data set, which can effectively extract deep information in the data to highlight the professional background information and spatiotemporal constraint information of the knowledge graph, and then enhances the knowledge graph to improve the integrity of the knowledge graph and enhance the entity representation, thereby improving the accuracy of the answer output. Then, by determining the complexity of the question text and selecting different methods to output the answer, the speed of answer output can be improved while improving the effectiveness and accuracy of the answer output.
[0063] Example 2 like Figure 2 As shown, in the second embodiment of the present invention, a highway toll collection auxiliary question-answering system based on a knowledge graph is provided, and the system includes: Construction module 1, for obtaining public information related to highway tolls, and constructing a highway toll information dataset based on the public information; Graph module 2, used to construct a domain knowledge graph based on the highway toll information dataset; Enhancement module 3, used to perform knowledge enhancement on the domain knowledge graph to obtain an enhanced knowledge graph; Analysis module 4, used to obtain the question text related to highway tolls sent by the user, perform complex analysis on the question text, and output complex analysis results; Output module 5 is configured to output a first answer to the question text based on the enhanced knowledge graph to output a target answer if the complex analysis result indicates that the question text is a question of simple difficulty; output a second answer to the question text based on the enhanced knowledge graph to output a target answer if the complex analysis result indicates that the question text is a question of medium difficulty; and output a third answer to the question text based on the enhanced knowledge graph to output a target answer if the complex analysis result indicates that the question text is a question of complex difficulty; The building block 1 comprises: a removal submodule, configured to convert the public information into text data and remove punctuation marks, repeated text, stop words, and irrelevant tags from the text data to obtain clean text data; The data set submodule is used to perform word segmentation processing on the clean text data to obtain a word segmentation data set, and to perform time and space annotation on the word segmentation data set according to preset annotation rules to obtain a highway toll information data set.
[0064] The atlas module 2 includes: An embedding representation submodule, configured to perform word embedding representation on the text words and the corresponding annotation information in the highway toll information dataset to obtain text word embedding vectors and annotation embedding vectors respectively; Extraction submodule, used to combine the attention mechanism and the bidirectional long short-term memory network to embed the text word vector , the annotation embedding vector Perform semantic feature extraction to obtain text features and annotation features : ; ; Where, is a bidirectional long short-term memory network, is the attention mechanism; A sequence submodule, configured to sequentially fuse and decode the text features and the annotation features to obtain a sentence sequence set; a propagation submodule, configured to input the sentence sequence set into the BERT model and determine a sentence vector sequence by forward propagation, input the sentence vector sequence into the bidirectional long short-term memory network, concatenate the forward and reverse hidden states at each moment to obtain a forward final state and a reverse final state, and concatenate the forward final state and the reverse final state to obtain an embedded representation of the sentence sequence; A probability submodule, configured to use the sentence sequence embedding representation as a classification feature of the textual relationship to determine the sentence relationship probability, thereby outputting the sentence relationship probability, and determining the sentence relationship based on the magnitude of the sentence relationship probability; Entity feature submodule, used to determine the entity feature vector based on the statement sequence set : ; Where, Indicates context, Characteristic words The feature vector in the context, Indicates the number of feature words in the context, is a set of statement sequences, is a sequence of sentence vectors, They are sentence sequence weight, sentence vector weight, and context feature weight respectively; Merging submodule, used to merge the entity feature vector The entities in are similarity calculated for each other, and the entities with similarity greater than the similarity threshold are merged to obtain the standard entity vector; The visualization submodule is used to obtain entity attributes in the standard entity vector, and visualize the entity attributes, standard entity vectors and sentence relationships in the form of a knowledge graph to obtain a domain knowledge graph.
[0065] The enhancement module 3 includes: The embedding encoding submodule is used to obtain a template question and embed the template question with the domain knowledge graph to obtain question embedding and entity embedding; A spatial submodule for mapping the question embedding into the same complex space as the entity embedding to obtain a mapped question embedding; The candidate submodule is used to take the remaining entities in the domain knowledge graph except the entities involved in the template question as candidate entities and calculate the first score value of each candidate entity : ; Where, 、 Respectively represent entity embeddings in complex space, for The complex conjugate of For the The mapping problem in the complex space is embedded. is the number of complex spaces, is the real part of the complex number; an arranging submodule, configured to arrange the plurality of entities to be selected in descending order according to the first scoring value, and select the first plurality of entities to be selected as candidate entities; The mapping submodule is used to map the question embedding to the sentence relationship in the domain knowledge graph to obtain the question relationship embedding , encode and embed the sentence relationship to obtain the relationship embedding , based on the problem relationship embedding, determine the final problem mapping relationship : ; Where, For the The embedding of the sentence relationship corresponding to the candidate entity, is the screening threshold, for The first relationship among Enhancement submodule for mapping relations based on the final question Determine the enhanced knowledge graph.
[0066] The enhancer module comprises: Scoring unit, for mapping relationships based on the final question Calculate the second score : ; Where, is a tunable hyperparameter, for Middle The embedding of the sentence relationship corresponding to the candidate entity, is the embedding of the shortest path from the entity in the template problem to each candidate entity; an arranging unit, configured to arrange the candidate entities in descending order according to the second scoring value, and select the first several candidate entities as target entities; an evidence unit, configured to extract evidence text related to the template question from the highway toll information dataset to obtain an evidence pool, and perform semantic analysis on the evidence question in the evidence pool to obtain an abstract concept evidence graph; An abstraction unit is used to convert the template problem and the target entity into an abstract concept entity graph, fuse the abstract concept evidence graph with the same concept nodes in the abstract concept entity graph to obtain an abstract concept graph, and fuse the abstract concept graph into the domain knowledge graph to obtain an enhanced knowledge graph.
[0067] The analysis module 4 includes: A parsing submodule, configured to perform dependency syntactic parsing on the question text to obtain the part-of-speech tag, word vector representation, and word dependency relationship of each word; The node submodule is used to convert part-of-speech tags, word vector representations, and word dependencies into a graph structure, interpreting word vector representations as nodes, word dependencies as edges, and part-of-speech tags as nodes. Aggregation submodule, used to input the nodes in the graph structure into the multi-layer CCN for update aggregation to obtain updated node feature representation : ; Where, is the activation function, The first The set of neighbor nodes of a node, 、 Respectively The learnable weights and biases of the layers, For the The updated node feature representation of the layer, is the normalization coefficient; Splicing submodule, used to calculate and update node feature representation The maximum depth and span variance of the node are expressed as follows: Perform concatenation, full connection layer processing, and activation function processing in sequence to output a first complex score; a representation submodule, configured to perform word segmentation and encoding processing on the question text to obtain a context representation, calculate a sentence perplexity based on the context representation to obtain a sentence perplexity, and normalize the sentence perplexity to obtain a second complexity score; a fusion submodule, configured to perform weighted fusion of the first complex score and the second complex score to obtain a final complex score; The complexity output submodule is used to determine that if the final complexity score is less than a first scoring threshold, the question text is a simple difficulty question; if the final complexity score is not less than the first scoring threshold and less than a second scoring threshold, the question text is a medium difficulty question; and if the final complexity score is not less than the second scoring threshold, the question text is a complex difficulty question.
[0068] The output module 5 includes: A first extraction submodule is used to extract entities and relations from the question text, and match the extracted entities with relations to obtain original triples; A matching submodule, configured to quickly match the original triples in the enhanced knowledge graph to output target triples; The first output submodule is used to output the target triple as a natural language sentence that conforms to grammar and has complete information according to a preset sentence template to obtain a target answer.
[0069] The output module 5 includes: A reference submodule, configured to obtain a reference text, segment the reference text into a plurality of text blocks, and vectorize the text blocks and the triples in the enhanced knowledge graph to obtain text vectors and graph vectors; A second extraction submodule is used to extract original triples of the question text and vectorize the original triples to obtain original vectors; The prompt submodule is used to fuse the text vector, the graph vector and the original vector into prompt information, input the prompt information into the large language model for answer output, and obtain the target answer.
[0070] The output module 5 includes: A triple submodule is configured to obtain an answer output model and a training data set, extract a training question from the training data set, extract a first triple of the training question and a corresponding second triple from the enhanced knowledge graph, and combine the first triple with the second triple to obtain combined data; A feature combination submodule, configured to extract semantic features of the combined data through an attention mechanism to obtain combined features; A bias submodule is used to add the combined data and the combined features to the data set to be processed, input the data set to be processed into the answer output model and perform bias score calculation to obtain a bias score. : ; Where, Indicates that only the combined data in the dataset to be processed is input into the answer output model and the answer distribution outputted after prediction by the answer output model; The attention interference submodule is used to calculate the observation results of the processed data set under different attention and counterfactual outcomes : ; ; Where, is a linear layer, They are the interference corresponding to the first attention and the second attention respectively; Deviation submodule, for , the counterfactual result With the bias score Calculating model bias : ; Where, is the number of data in the dataset to be processed; Loss submodule, used to calculate the bias value based on the model Determine the loss function : ; Where, represents the cross entropy between the model output and the true answer, is the model deviation value Cross entropy between the attribute and the true answer; The second output submodule is used to optimize the answer output model by minimizing the loss function to obtain an optimized model, and input the question text into the optimized model to output a target answer.
[0071] In other embodiments of the present invention, embodiments of the present invention provide the following technical solutions: a computer comprising a memory 102, a processor 101, and a computer program stored on the memory 102 and executable on the processor 101; when the processor 101 executes the computer program, the knowledge graph-based highway toll collection auxiliary question-and-answer method as described above is implemented.
[0072] Specifically, the processor 101 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present invention.
[0073] Memory 102 may include a large-capacity memory for data or instructions. By way of example, and not limitation, memory 102 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 102 may include removable or non-removable (or fixed) media. Where appropriate, memory 102 may be internal or external to the data processing device. In certain embodiments, memory 102 is non-volatile memory. In certain embodiments, memory 102 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM) or a flash memory (FLASH), or a combination of two or more of these. Under appropriate circumstances, the RAM can be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM can be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0074] The memory 102 may be used to store or cache various data files that need to be processed and / or used for communication, as well as possible computer program instructions executed by the processor 101 .
[0075] The processor 101 implements the above-mentioned knowledge graph-based highway toll collection auxiliary question-answering method by reading and executing computer program instructions stored in the memory 102.
[0076] In some embodiments, the computer may further include a communication interface 103 and a bus 100. Figure 3 As shown, the processor 101 , the memory 102 , and the communication interface 103 are connected via a bus 100 and communicate with each other.
[0077] The communication interface 103 is used to implement communication between the various modules, devices, units, and / or equipment in the embodiments of the present invention. The communication interface 103 can also implement data communication with other components such as external devices, image / data acquisition equipment, databases, external storage, and image / data processing workstations.
[0078] Bus 100 includes hardware, software, or both, and couples components of a computer device to each other. Bus 100 includes, but is not limited to, at least one of the following: a data bus, an address bus, a control bus, an expansion bus, and a local bus. By way of example and not limitation, bus 100 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Bus 100 may include one or more buses, where appropriate. Although embodiments of the present invention describe and illustrate a particular bus, the present invention contemplates any suitable bus or interconnect.
[0079] The computer can execute the knowledge graph-based highway toll auxiliary question and answer method of the present invention based on the knowledge graph based on the highway toll auxiliary question and answer system obtained, thereby realizing the knowledge graph-based highway toll auxiliary question and answer.
[0080] In some further embodiments of the present invention, in combination with the above-mentioned highway toll collection auxiliary question and answer method based on knowledge graph, the embodiments of the present invention provide the following technical solutions: a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned highway toll collection auxiliary question and answer method based on knowledge graph is implemented.
[0081] Those skilled in the art will appreciate that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device), or in conjunction with such instruction execution system, apparatus, or device. For purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by an instruction execution system, apparatus, or device, or in conjunction with such instruction execution system, apparatus, or device.
[0082] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0083] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the aforementioned embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following technologies known in the art may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0084] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0085] The above-described embodiments merely represent several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person of ordinary skill in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, and these variations and improvements fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.
Claims
1. A knowledge graph-based highway toll collection auxiliary question-answering method, characterized in that: include: Obtaining public information related to highway tolls, and constructing a highway toll information dataset based on the public information; Constructing a domain knowledge graph based on the highway toll information dataset; Performing knowledge enhancement on the domain knowledge graph to obtain an enhanced knowledge graph; Obtaining a question text related to highway tolls sent by a user, performing complex analysis on the question text, and outputting complex analysis results; If the result of the complex analysis is that the question text is a question of simple difficulty, a first answer is output for the question text based on the enhanced knowledge graph to output the target answer. If the result of the complex analysis is that the question text is a question of medium difficulty, a second answer is output for the question text based on the enhanced knowledge graph to output the target answer. If the result of the complex analysis is that the question text is a question of complex difficulty, a third answer is output for the question text based on the enhanced knowledge graph to output the target answer.
2. The highway toll collection auxiliary question-answering method based on knowledge graph according to claim 1 is characterized in that: The step of constructing a highway toll information dataset based on the public information includes: Converting the public information into text data, and removing punctuation, repeated text, stop words, and irrelevant tags from the text data to obtain clean text data; The clean text data is segmented to obtain a segmented data set, and the segmented data set is annotated with time and space according to preset annotation rules to obtain a highway toll information data set.
3. The highway toll collection auxiliary question-answering method based on knowledge graph according to claim 1 is characterized in that: The step of constructing a domain knowledge graph based on the highway toll information dataset includes: Performing word embedding representation on the text words and the corresponding annotation information in the highway toll information dataset to obtain text word embedding vectors and annotation embedding vectors respectively; Combine the attention mechanism and the bidirectional long short-term memory network to embed the text word vector , the annotation embedding vector Perform semantic feature extraction to obtain text features and annotation features : ; ; Where, is a bidirectional long short-term memory network, is the attention mechanism; The text features and the annotation features are sequentially fused and decoded to obtain a sentence sequence set; Inputting the sentence sequence set into the BERT model and determining a sentence vector sequence by forward propagation, inputting the sentence vector sequence into the bidirectional long short-term memory network, concatenating the forward and reverse hidden states at each moment to obtain a forward final state and a reverse final state, and concatenating the forward final state and the reverse final state to obtain a sentence sequence embedding representation; Determine the sentence relationship probability by using the sentence sequence embedding representation as a classification feature of the text relationship to output the sentence relationship probability, and determine the sentence relationship based on the magnitude of the sentence relationship probability; Determine an entity feature vector based on the sentence sequence set : ; Where, Indicates context, Characteristic words The feature vector in the context, Indicates the number of feature words in the context, is a set of statement sequences, is a sequence of sentence vectors, They are sentence sequence weight, sentence vector weight, and context feature weight respectively; For the entity feature vector The entities in are similarity calculated for each other, and the entities with similarity greater than the similarity threshold are merged to obtain the standard entity vector; The entity attributes in the standard entity vector are obtained, and the entity attributes, standard entity vector and sentence relationship are visualized in the form of a knowledge graph to obtain a domain knowledge graph.
4. The highway toll collection auxiliary question-answering method based on knowledge graph according to claim 1 is characterized in that: The step of performing knowledge enhancement on the domain knowledge graph to obtain an enhanced knowledge graph includes: Obtain a template question, and embed the template question and the domain knowledge graph to obtain question embedding and entity embedding; Mapping the question embedding into the same complex space as the entity embedding to obtain a mapped question embedding; The remaining entities in the domain knowledge graph except the entities involved in the template question are taken as candidate entities, and the first score value of each candidate entity is calculated. : ; Where, 、 Respectively represent entity embeddings in complex space, for The complex conjugate of For the The mapping problem in the complex space is embedded. is the number of complex spaces, is the real part of the complex number; Arrange the plurality of entities to be selected in descending order according to the first scoring value, and select the first plurality of entities to be selected as candidate entities; Map the question embedding to the sentence relationship in the domain knowledge graph to obtain the question relationship embedding , encode and embed the sentence relationship to obtain the relationship embedding , based on the problem relationship embedding, determine the final problem mapping relationship : ; Where, For the The embedding of the sentence relationship corresponding to the candidate entity, is the screening threshold, for The first relationship among Based on the final problem mapping relationship Determine the enhanced knowledge graph.
5. The highway toll collection auxiliary question-answering method based on knowledge graph according to claim 4 is characterized in that: Based on the final problem mapping relationship The steps to determine the enhanced knowledge graph include: Based on the final problem mapping relationship Calculate the second score : ; Where, is a tunable hyperparameter, for Middle The embedding of the sentence relationship corresponding to the candidate entity, is the embedding of the shortest path from the entity in the template problem to each candidate entity; Arrange the candidate entities in descending order according to the second scoring value, and select the first several candidate entities as target entities; Extracting evidence text related to the template question from the highway toll information dataset to obtain an evidence pool, and performing semantic parsing on the evidence questions in the evidence pool to obtain an abstract concept evidence graph; The template question and the target entity are converted into an abstract concept entity graph, the abstract concept evidence graph and the same concept nodes in the abstract concept entity graph are fused to obtain an abstract concept graph, and the abstract concept graph is fused into the domain knowledge graph to obtain an enhanced knowledge graph.
6. The highway toll collection auxiliary question-answering method based on knowledge graph according to claim 1 is characterized in that: The step of performing complex analysis on the question text to output complex analysis results includes: Perform dependency parsing on the question text to obtain the part-of-speech tag, word vector representation, and word dependency relationship of each word; Convert part-of-speech tags, word vector representations, and word dependencies into a graph structure, interpreting word vector representations as nodes, word dependencies as edges, and part-of-speech tags as nodes. The nodes in the graph structure are input into the multi-layer CCN for update aggregation to obtain the updated node feature representation : ; Where, is the activation function, The first The set of neighbor nodes of a node, 、 Respectively The learnable weights and biases of the layers, For the The updated node feature representation of the layer, is the normalization coefficient; Calculate and update node feature representation The maximum depth and span variance of the node are expressed as follows: Perform concatenation, full connection layer processing, and activation function processing in sequence to output a first complex score; Performing word segmentation and encoding processing on the question text to obtain a context representation, calculating a sentence perplexity based on the context representation to obtain a sentence perplexity, and normalizing the sentence perplexity to obtain a second complexity score; Performing a weighted fusion of the first complexity score and the second complexity score to obtain a final complexity score; If the final complexity score is less than the first scoring threshold, the question text is a question of simple difficulty; if the final complexity score is not less than the first scoring threshold and less than the second scoring threshold, the question text is a question of medium difficulty; if the final complexity score is not less than the second scoring threshold, the question text is a question of complex difficulty.
7. The highway toll collection auxiliary question-answering method based on knowledge graph according to claim 1 is characterized in that: The step of outputting a first answer to the question text based on the enhanced knowledge graph to output a target answer includes: Extracting entities and relations from the question text, and matching the extracted entities with relations to obtain original triples; Quickly matching the original triples in the enhanced knowledge graph to output target triples; The target triple is output as a grammatically correct and complete natural language sentence according to a preset sentence template to obtain the target answer.
8. The highway toll collection auxiliary question-answering method based on knowledge graph according to claim 1 is characterized in that: The step of outputting a second answer to the question text based on the enhanced knowledge graph to output a target answer includes: Obtain a reference text, divide the reference text into a plurality of text blocks, and vectorize the text blocks and the triples in the enhanced knowledge graph to obtain text vectors and graph vectors; Extracting original triples of the question text and vectorizing the original triples to obtain original vectors; The text vector, the graph vector, and the original vector are fused into prompt information, and the prompt information is input into a large language model for answer output to obtain a target answer.
9. The highway toll collection auxiliary question-answering method based on knowledge graph according to claim 1 is characterized in that: The step of outputting a third answer to the question text based on the enhanced knowledge graph to output a target answer includes: Obtaining an answer output model and a training data set, extracting a training question from the training data set, extracting a first triple of the training question and extracting a corresponding second triple from the enhanced knowledge graph, and combining the first triple with the second triple to obtain combined data; Extracting semantic features of the combined data through an attention mechanism to obtain combined features; Add the combined data and the combined features to the dataset to be processed, input the dataset to be processed into the answer output model and perform bias score calculation to obtain a bias score : ; Where, Indicates that only the combined data in the dataset to be processed is input into the answer output model and the answer distribution outputted after prediction by the answer output model; Calculate the observation results of the dataset to be processed under different attention and counterfactual outcomes : ; ; Where, is a linear layer, They are the interference corresponding to the first attention and the second attention respectively; Based on the observation results , the counterfactual result With the bias score Calculating model bias : ; Where, is the number of data in the dataset to be processed; Based on the model deviation value Determine the loss function : ; Where, represents the cross entropy between the model output and the true answer, is the model deviation value Cross entropy between the attribute and the true answer; The answer output model is optimized and trained by minimizing the loss function to obtain an optimized model, and the question text is input into the optimized model to output a target answer.
10. A knowledge graph-based highway toll collection auxiliary question-answering system, characterized in that: The system comprises: A construction module, configured to obtain public information related to highway toll collection and construct a highway toll collection information dataset based on the public information; A graph module, configured to construct a domain knowledge graph based on the highway toll information dataset; An enhancement module, configured to perform knowledge enhancement on the domain knowledge graph to obtain an enhanced knowledge graph; An analysis module is used to obtain a question text related to highway tolls sent by a user, perform complex analysis on the question text, and output complex analysis results; An output module is used to output a first answer to the question text based on the enhanced knowledge graph to output a target answer if the complex analysis result shows that the question text is a question of simple difficulty; to output a second answer to the question text based on the enhanced knowledge graph to output a target answer if the complex analysis result shows that the question text is a question of medium difficulty; and to output a third answer to the question text based on the enhanced knowledge graph to output a target answer if the complex analysis result shows that the question text is a question of complex difficulty.
Citation Information
Patent Citations
Device fault intelligent question-answering method based on large model enhancement
CN120086341A
System, method, computer readable storage medium, and computer program product for crop breeding questions and answers
CN120104811A
Wind power fault diagnosis operation and maintenance method based on multi-source data and knowledge retrieval enhancement
CN120297410A
Sparse Factor Analysis for Analysis of User Content Preferences
US20140279727A1