Intelligent tour guide man-machine conversation method and system based on knowledge graph retrieval

By constructing a knowledge graph in the tour guide field and fine-tuning the pre-trained model using LoRA technology, and combining the knowledge graph for content verification and correction, the problem of model forgetting and false content in tour guide commentary is solved, and efficient, accurate and professional tour guide commentary services are achieved.

CN120163250APending Publication Date: 2025-06-17XUSHI YIDONG CULTURE TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510268954.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the tour guide commentary, the existing technology has problems such as forgetting models, insufficient utilization of new data, slow retrieval of false content and knowledge graphs, resulting in insufficient performance of AI dialogue tour guide models and unable to meet the real-time, accurate and professional tour guide commentary needs.

Method used

The intelligent tour guide human-computer dialogue method based on knowledge graph retrieval is adopted. By building a knowledge graph in the tour guide field, the pre-trained model is fine-tuned using LoRA technology, and content checksum correction is performed in combination with the knowledge graph, and the knowledge graph retrieval algorithm is optimized to improve the response speed.

Benefits of technology

Effectively integrate new information and pre-trained knowledge, reduce model forgetting and hallucination phenomena, improve the accuracy and professionalism of answers, meet real-time response needs, and enhance user trust and satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163250A_ABST
    Figure CN120163250A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent tour guide man-machine conversation method and system based on knowledge graph retrieval, and relates to the technical field of artificial intelligence and natural language processing. The method comprises the following steps: collecting relevant data of scenic spots in a preset range, carrying out structured processing on all the relevant data of the scenic spots, and constructing a knowledge graph in the tour guide field; establishing indexes for entities and relationships in the knowledge graph, and introducing an adaptive query optimization mechanism; selecting a pre-trained language model and performing LoRA fine tuning on the pre-trained language model to obtain a fine tuning language model; a user initiates a dialogue request; the dialogue request accesses the fine tuning language model to obtain a preliminary answer result, and accesses the knowledge graph to obtain knowledge graph information; and matching and fusing the preliminary answer result and the knowledge graph information, and outputting a corrected answer result. According to the invention, knowledge can be effectively fused between the tour guide explanation field and the general field, and the accuracy and comprehensiveness of answers are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and natural language processing technology, and in particular to an intelligent tour guide human-computer dialogue method and system based on knowledge graph retrieval. Background Art

[0002] With the rapid development of artificial intelligence technology, conversation models based on deep learning have been widely used in various fields. However, in specific fields such as tour guide interpretation, the existing technology has the following problems:

[0003] When fine-tuning a pre-trained model, the model tends to forget the newly added fine-tuning data and still tends to output what was learned during pre-training, resulting in failure to fully utilize the new information.

[0004] The model may generate false content that is inconsistent with the facts, affecting user experience and trust.

[0005] Traditional knowledge graph retrieval methods have the problem of slow prediction speed in practical applications and cannot meet the needs of real-time response.

[0006] These problems limit the performance of the AI ​​conversational tour guide model and cannot meet tourists' needs for real-time, accurate and professional tour guide explanations. Summary of the invention

[0007] Purpose of the invention: To propose an intelligent tour guide human-computer dialogue method based on knowledge graph retrieval, and further propose an intelligent tour guide system for implementing this method, so as to solve the above-mentioned problems existing in the prior art.

[0008] In order to solve the above technical problems, the present invention proposes an intelligent tour guide human-computer dialogue method based on knowledge graph retrieval, comprising the following steps:

[0009] Collect relevant data of scenic spots within a predetermined range, perform structured processing on all relevant data of scenic spots, and construct a knowledge graph in the field of tour guides; index the entities and relationships in the knowledge graph, and introduce an adaptive query optimization mechanism;

[0010] Select a pre-trained language model and perform LoRA fine-tuning on it to obtain a fine-tuned language model;

[0011] The user initiates a dialogue request; the dialogue request accesses the fine-tuned language model through the first channel to obtain a preliminary answer result; the dialogue request accesses the knowledge graph through the second channel to obtain knowledge graph information;

[0012] The preliminary answer result and the knowledge graph information are matched and integrated to output a revised answer result.

[0013] Further, the scenic spot related data includes basic information, historical background, cultural connotations, and natural scenery;

[0014] The basic information includes at least: scenic spot name, geographical location, opening hours, ticket price;

[0015] The historical background includes at least: construction era, relevant historical events, important figures;

[0016] The cultural connotations include at least: legend stories, cultural allusions, folk customs;

[0017] The natural scenery includes at least: geographical features, animal and plant resources, climate environment.

[0018] Further, the process of structuring all the scenic spot related data includes:

[0019] Identifying entities from all the scenic spot related data;

[0020] Combining the identified entities in pairs, and extracting the relationship between the two entities by combining the context of the two entities;

[0021] Constructing triples in the form of "entity 1 - relationship - entity 2" to form the basic unit of the knowledge graph;

[0022] Storing all the triples using a graph database to construct a complete knowledge graph.

[0023] Further, the process of identifying entities from all the scenic spot related data includes:

[0024] Organizing all the scenic spot related data into an input sequence X, and taking the maximization of the conditional probability as the goal, identifying entities from the input sequence X, and organizing the identified entities into an output sequence Y:

[0025]

[0026] where is the normalization factor; is the feature function; is the weight of the feature function; and are the i-th and (i - 1)-th identified entities respectively; m is the current identification round; M is the total number of identification rounds; i is the number of currently identified entities; n is the total number of identified entities.

[0027] Further, the process of extracting the relationship between two entities includes:

[0028] Including the entity and the entity Context text Expressed as a feature vector :

[0029]

[0030] In the formula, Is the embedding function;

[0031] Train a convolutional neural network as a relation extraction model; input the feature vector Into the relation extraction model to extract the relationship between entity And entity ; Among them, the structure of the relation extraction model h is expressed as: ; Among them, the structure of the relation extraction model h is expressed as:

[0032]

[0033] Relation probability distribution Is expressed as:

[0034]

[0035] In the formula, 、 Are weights; 、 Are bias vectors; Is the activation function; Is the normalization function.

[0036] Furthermore, the relation extraction model uses the cross-entropy loss function To train the convolutional neural network, and the expression of the cross-entropy loss function is as follows:

[0037]

[0038] In the formula, Is the indicator function, which takes the value of 1 when the true relation is k, otherwise 0.

[0039] Furthermore, by constructing the "entity 1-relation-entity 2" triple To form the basic unit of the knowledge graph.

[0040] Furthermore, Llama 3.1 is selected as the pre-trained language model and fine-tuned with LoRA. This process includes:

[0041] Add trainable low-rank matrices A and B to the self-attention layer and the feed-forward network layer in the pre-trained language model. The size of matrix A is [d, r], and the size of matrix B is [r, d], where d is the dimension of the hidden layer and r is the rank of the low-rank matrix;

[0042]

[0043] Wherein, C is an adaptive adjustment term guided by the knowledge graph, used to integrate professional tour guide knowledge; , used to dynamically adjust the fine-tuning strength; W represents the parameter matrix of the pre-trained model.

[0044] In the above way, the scale of parameter update is reduced from [d × d] to [d × r + r × d].

[0045] Furthermore, the process of matching and integrating the preliminary answer result and the knowledge graph information includes:

[0046] Match the preliminary answer result generated by the fine-tuned language model with the knowledge graph information generated by the knowledge graph, and measure the credibility through the knowledge consistency score S:

[0047]

[0048] Wherein, It is measured by calculating the cosine similarity between the two; the concept consistency judges whether it belongs to the same concept set based on the knowledge points in the same field. If it belongs to the same concept set, a higher weight is given, and if it does not belong, the score is reduced; the semantic matching score is based on BERT to calculate the matching degree between the user's question and the model's answer.

[0049] Furthermore, according to the calculated knowledge consistency score S, the following correction strategies are executed:

[0050] If S < the first threshold, trigger the knowledge enhancement mode, retrieve the knowledge graph again and generate an answer;

[0051] If the first threshold ≤ S < the second threshold, trigger the content correction mode; the content correction mode can be:

[0052] Syntactic structure optimization: Use the syntactic reconstruction method based on dependency parsing to adjust the word order of the answer to make it more in line with the expression habits of professional tour guide knowledge.

[0053] Multimodal fusion: If the knowledge graph contains relevant images, route maps or historical background information, embed the corresponding auxiliary materials in the text answer to improve the intuitiveness and comprehensibility of the content.

[0054] Standardization of professional terms: Based on the domain-specific term dictionary, standardize the terms in the answer to ensure that they conform to industry practices and improve readability and professionalism.

[0055] If S ≥ the second threshold, directly output the answer and pop up a dialog box to allow the user to feedback whether this answer is accurate. If the user feedbacks an error, the error sample will be automatically stored and a penalty term will be added in the next fine-tuning.

[0056] In addition, the present invention also provides an intelligent tour guide human-machine dialogue system, which includes at least one processor and a memory communicatively connected to the processor; wherein, the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned intelligent tour guide human-machine dialogue method based on knowledge graph retrieval.

[0057] Compared with the prior art, the present invention has at least the following beneficial effects:

[0058] (1) By using the LoRA technology to fine-tune the pre-trained model based on the knowledge graph, the model can introduce new knowledge while retaining the general knowledge learned during pre-training. In this way, the model realizes the effective integration of knowledge between a specific field (such as tour guide explanation) and the general field, prevents the model from forgetting the newly added data during the fine-tuning process, and improves the accuracy and comprehensiveness of the answer.

[0059] (2) The present invention introduces a verification mechanism to verify and correct the answer content generated by the model using the knowledge graph, significantly reducing the probability of the model having hallucinations. By verifying the content after generating the answer, the reliability and consistency of the answer are ensured, and the user's trust and satisfaction are improved.

[0060] (3) By using the professional knowledge in the knowledge graph and combining LoRA fine-tuning, the model has higher professionalism and reliability in the field of tour guide explanation. The model can provide detailed and accurate scenic spot introductions and answers, improving the quality of the answer and meeting the user's need for professional knowledge.

[0061] (4) The present invention can respond to the user's questions in real time, provide accurate and professional tour guide explanations, reduce the hallucination phenomenon, and improve the reliability of the answer. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is the overall architecture diagram of the present invention, showing the relationships between modules and the data flow.

[0063] Figure 2 is the knowledge graph construction flow chart, describing the whole process from data collection to knowledge graph establishment.

[0064] Figure 3 is the LoRA fine-tuning flow chart, describing the steps and parameter settings during the fine-tuning process.

[0065] Figure 4It is a flow chart for model inference and verification, showing the process of generating answers and verifying and correcting them. Detailed implementation manners

[0066] In the following description, a large number of specific details are given to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, some technical features known to the public are not described to avoid confusion with the present invention.

[0067] Example 1:

[0068] This example discloses an intelligent tour guide human-machine dialogue method based on knowledge graph retrieval. Its overall architecture is shown in Figure 1 . The purpose of this method is to solve the problems of the model forgetting newly added data, generating hallucinations, and slow knowledge graph retrieval speed during the fine-tuning process in the prior art. By performing LoRA (Low-Rank Adaptation) fine-tuning on the pre-trained model Llama3.1 8b instruct based on the knowledge graph, it is ensured that the model effectively integrates new information, improves the training speed and prediction performance, and at the same time reduces the occurrence of hallucination phenomena.

[0069] The specific implementation steps are as follows:

[0070] (1) Construct a knowledge graph in the tour guide field

[0071] The process of constructing the knowledge graph is shown in Figure 2 . First, collect rich information related to each scenic spot, including historical background, cultural allusions, geographical location, characteristic scenery, etc. Structure this information to establish a knowledge graph containing entities (such as scenic spots, historical figures, events) and relationships (such as "located in", "built in", "related to"). This provides a comprehensive and accurate professional knowledge basis for the model to ensure that reliable information can be provided to users during the dialogue process.

[0072] The present invention proposes a feasible entity recognition scheme:

[0073] Organize all scenic spot-related data into an input sequence X, and take the maximization of the conditional probability as the goal to identify entities from the input sequence X, and organize the identified entities into an output sequence Y:

[0074]

[0075] In the formula, is the normalization factor; is the feature function; is the weight of the feature function; , are the i-th and i-1-th recognized entities respectively; m is the current recognition round; M is the total recognition rounds; i is the number of entities currently recognized; and n is the total number of recognized entities.

[0076] The present invention proposes a feasible solution for extracting the relationship between two entities:

[0077] Will contain entities and entities Context text Represented as a feature vector :

[0078]

[0079] In the formula, is the embedding function;

[0080] Train a convolutional neural network as a relation extraction model; the relation extraction model uses a cross entropy loss function The cross entropy loss function expression obtained by training the convolutional neural network is as follows:

[0081]

[0082] In the formula, is an indicator function that takes the value 1 when the true relationship is k and 0 otherwise.

[0083] The feature vector Input into the relation extraction model to extract entities and entities The relationship between ; Among them, the structure of the relation extraction model h is expressed as:

[0084]

[0085] Relationship Probability Distribution It is expressed as:

[0086]

[0087] In the formula, , is the weight; , is the bias vector; is the activation function; is a normalization function.

[0088] Based on the above process, the "entity 1-relationship-entity 2" triple is constructed. .

[0089] (2)Select a pre-trained model and perform LoRA fine-tuning

[0090] The process of LoRA fine-tuning is shown in Figure 3 . Select Llama3.1 8b instruct as the basic pre-trained language model. This model has powerful natural language understanding and generation capabilities, but may lack expertise in specific domains. Therefore, LoRA technology is used to fine-tune it.

[0091] During the LoRA fine-tuning process, most of the parameters of the pre-trained model are frozen, and trainable low-rank matrix parameters are added only to specific layers (such as the self-attention layer and the feed-forward network layer). This method significantly reduces the number of parameters to be updated and the computational cost of training.

[0092] The parameter settings and training details of LoRA fine-tuning are as follows:

[0093] Select the layers for fine-tuning: Mainly fine-tune the self-attention layers and feed-forward network layers in the Transformer model.

[0094] Addition of low-rank matrices: In the above layers, add trainable low-rank matrices A and B. The size of matrix A is [d, r], and the size of matrix B is [r, d], where d is the dimension of the hidden layer and r is the rank of the low-rank matrix. In this way, the scale of parameter update is reduced from the original [d × d] to [d × r + r × d].

[0095] Selection of the rank (r): Set r to 4, which is a compromise choice that can capture enough new information without introducing too many parameters.

[0096] Alpha parameter: Set Alpha to 16 to scale the output of the LoRA layer to match the parameter scale of the pre-trained model and balance the influence of new and old information.

[0097] Learning rate and optimizer: Use the AdamW optimizer, and set the initial learning rate to 1e-4. During training, adopt a learning rate scheduling strategy to gradually reduce the learning rate. Combine the weight decay strategy, and set the weight decay coefficient to 0.01 to prevent overfitting.

[0098] Training batch size and number of epochs: According to the scale of the training dataset, set the batch size to 32 and the number of training epochs to 5. Since the number of parameters to be fine-tuned by LoRA is small, the training speed is fast and the training can be completed in a short time.

[0099] Loss function: The cross-entropy loss function is adopted, and a penalty term for hallucinated content is added to the loss function. When the content generated by the model is inconsistent with the knowledge graph, the loss value will increase, prompting the model to adjust the generation strategy and reduce the occurrence of hallucinations.

[0100] Early stopping strategy: During the training process, monitor the change of the loss value on the validation set. If the loss value does not decrease for two consecutive epochs, stop the training early to prevent overfitting.

[0101] In fine-tuning training, convert the information in the knowledge graph into training samples, combine them with the tour guide dialogue data to form a training set containing a large number of question-and-answer pairs. The model learns professional knowledge and dialogue skills in the tour guide field through these data and can better answer questions raised by users.

[0102] (3)Optimize knowledge graph retrieval

[0103] The knowledge graph retrieval optimization method based on multi-layer index fusion in this embodiment, compared with traditional inverted index and hash mapping, combines a vector database (such as FAISS, HNSW) + keyword index + semantic index to improve the retrieval speed and recall rate. The optimization strategies include efficient entity indexing, using entity embedding vectors calculated based on Transformer to store the entities in the knowledge graph in the vector database (FAISS) to improve query accuracy. The hierarchical index structure includes an inverted index (keyword matching, fast preliminary screening), a semantic index (using BERT to calculate semantic similarity, optimizing sorting), and vector retrieval (FAISS / HNSW extracts the most relevant knowledge). The adaptive query optimization mechanism caches and optimizes high-frequency knowledge points based on the user input history and query frequency to reduce the number of database accesses. Compared with the traditional inverted index method, the retrieval speed is increased by 30%, the recall rate is increased by 25%, and the knowledge matching accuracy is increased by 18%.

[0104] Specific optimization measures include:

[0105] Establish an efficient index structure: Index the key entities and relationships in the knowledge graph to facilitate quick location of relevant information.

[0106] Use a caching mechanism: For high-frequency queries, use a caching mechanism to reduce the time of repeated retrieval.

[0107] Parallel retrieval: Utilize multi-threading or multi-processing technology to achieve parallel retrieval and accelerate the retrieval speed.

[0108] (4)Model inference and hallucination reduction

[0109] See Figure 4, the present invention proposes an improved hallucination suppression mechanism to ensure the reliability of the content generated by the AI dialogue tour guide model. The credibility is measured by the knowledge consistency score S, where S = λ1 × Sim(generated text, knowledge graph) + λ2 × concept consistency + λ3 × semantic matching score. Here, Sim(generated text, knowledge graph) is measured by the cosine similarity calculated by Transformer, the concept consistency is determined based on whether the knowledge points in the same field belong to the same concept set, and the semantic matching score is based on the matching degree of the user's question and the model's answer calculated by BERT. The hallucination correction strategy includes: if S < 0.7, the model triggers the knowledge enhancement mode, retrieves the knowledge graph again and generates an answer; if 0.7 ≤ S < 0.9, the model triggers the content correction mode, adjusts the answer structure to make it more in line with the known knowledge points; if S ≥ 0.9, the answer is directly output. The user feedback enhancement mechanism allows users to select "Is this answer accurate?" If the user feedbacks an error, the system automatically stores the error sample and adds a penalty term in the next fine-tuning. Through autoregressive credibility regulation, the hallucination content is reduced by 40%. By adopting a multi-layer knowledge verification mechanism, the accuracy of the model's answer is increased by 25%. Combining user feedback optimization, the knowledge system of the AI tour guide is continuously evolving.

[0110] Specific measures include:

[0111] Content matching: Match the answer generated by the model with the information in the knowledge graph, calculate the similarity, and judge the credibility of the answer.

[0112] Correction strategy: If it is found that the answer is inconsistent with the knowledge graph, the model will refer to the knowledge graph information and adjust the answer content.

[0113] Feedback mechanism: Introduce a user feedback mechanism. If the user is not satisfied with the answer, the model will record it and optimize it in subsequent training.

[0114] (5) Solve the model forgetting problem

[0115] To prevent the model from forgetting the general knowledge learned during pre-training during the fine-tuning process, the following strategies are adopted:

[0116] Contrastive learning: Add positive and negative sample pairs during training to let the model learn to distinguish correct and incorrect answers, improve the model's discrimination ability, and enhance the retention of general knowledge.

[0117] Knowledge distillation: Use the output of the pre-trained model as the teacher model to guide the learning process of the fine-tuning model, so that it does not forget the original knowledge while learning new knowledge.

[0118] Mixed training data: In the training dataset, a portion of general domain data is retained and mixed with the tour guide domain data for training to ensure the model's mastery of general knowledge.

[0119] The present invention sets the knowledge importance weight Ltotal = Lfine-tuning + γLdistillation based on the knowledge distillation strategy of domain importance, where Lfine-tuning represents the fine-tuning loss of the current task, Ldistillation represents the similarity loss with the output of the pre-trained model, and γ is the knowledge retention intensity, which is dynamically adjusted according to the task importance. Soft target distillation is used to reduce forgetting. When the new model generates answers, it compares with the answers of the original pre-trained model and calculates the KL divergence to prevent the model from deviating completely from the original knowledge system. The mixed training strategy uses 30% of the original pre-trained data + 70% of the tour guide domain data, enabling the model to specialize in domain knowledge while still retaining general conversation capabilities. Experimental results show that this method reduces the forgetting rate by 35% and improves the answer accuracy by 22% in the tour guide commentary task.

[0120] (6) Model deployment and application

[0121] In terms of model deployment and application, techniques such as model pruning, quantization, and parallel computing are adopted to improve the inference speed of the model. Specific measures include:

[0122] Model pruning: Remove parameters and neurons that have less impact on the model performance, reduce the model size, and improve the inference speed.

[0123] Quantization technology: Quantize the model weights from 32-bit floating point numbers to 16-bit or even 8-bit, reducing the computational complexity and memory occupancy, and adapting to low-resource environments.

[0124] Parallel computing: Utilize multi-GPU parallel computing, deploy the model distributively, and accelerate the inference process to meet the requirements of real-time interaction.

[0125] High-performance server deployment: Deploy the model on a server with high-performance computing capabilities and utilize GPU acceleration to ensure real-time response when users interact with the model.

[0126] Through the above methods, the AI dialogue tour guide model of the present invention has significant improvements in performance and efficiency. The model can provide professional, accurate, and real-time tour guide commentary services for users, enhancing the user experience. At the same time, due to the adoption of LoRA technology and knowledge graph optimization, the training and inference costs of the model are greatly reduced, and it has broad application prospects.

[0127] Example 2:

[0128] This embodiment details a construction method of an AI dialogue tour guide model based on knowledge graph retrieval and LoRA (Low-Rank Adaptation) accelerated fine-tuning, aiming to improve the professionalism, accuracy, and real-time response ability of the model in the tour guide field.

[0129] First, construct a knowledge graph for the tour guide field. By collecting detailed information about scenic spots across the country, including basic information (such as scenic spot name, geographical location, opening hours, ticket price), historical background (such as construction era, relevant historical events, important figures), cultural connotations (such as legends, cultural allusions, folk customs), and natural scenery (such as geographical features, animal and plant resources, climate environment), etc., clean this data to remove redundant and incorrect information to ensure data quality.

[0130] During the data structuring process, perform entity recognition to extract entities such as scenic spots, people, and events; through relationship extraction, identify the relationships between entities, such as "located in", "built in", "related to", etc. Then, organize these entities and relationships into triples (entity 1 - relationship - entity 2) to form the basic units of the knowledge graph. Use a graph database (such as Neo4j) to store the triples and construct a complete knowledge graph.

[0131] Next, select a pre-trained model and perform LoRA fine-tuning. Llama3.1 8b instruct is selected as the base pre-trained language model, which has powerful natural language understanding and generation capabilities and is suitable for dialogue generation tasks. During LoRA fine-tuning, an Adaptive Rank Adjustment (ARA) mechanism is proposed to optimize the adaptability of the pre-trained model in the tour guide explanation task. The traditional LoRA method uses fixed low-rank matrix parameters (Rank r), which may lead to a trade-off between computational efficiency and knowledge fusion ability in different task scenarios. Therefore, the present invention proposes: W = W0 + αAB + βC, where C is an adaptive adjustment term guided by the knowledge graph, used to fuse professional tour guide knowledge and improve the accuracy of content generation. β = g(knowledge density, historical similarity, user query complexity), which is used to dynamically adjust the fine-tuning intensity, so that scenic spots with different knowledge backgrounds have different fine-tuning intensities. Compared with the fixed RankLoRA, this method can automatically enhance the knowledge fusion ability in tasks that require high knowledge fidelity, and reduce unnecessary parameter adjustments in general dialogue tasks, thereby reducing computational resource consumption. Introducing the knowledge graph-guided dynamic adaptation parameter C ensures that the generated content conforms to professional domain knowledge and avoids the spread of incorrect information. Adaptively adjust the fine-tuning degree of the model between different scenic spot types (historical culture vs. natural scenery), so that it can generate accurate professional content while maintaining the understanding of general knowledge.

[0132] In terms of parameter settings, the rank r is set to 4 to balance the model capacity and the number of parameters; the scaling factor α is set to 16 to adjust the influence degree of the low-rank matrix on the model; the AdamW optimizer is used, and the initial learning rate is set to 1e-4; the batch size is 32, and the number of training epochs is 5. The training data includes question-answer pairs generated from the knowledge graph and collected dialogue data in the tour guide field. During training, in order to reduce the hallucination content generated by the model, a penalty term is added to the loss function. When the content output by the model is inconsistent with the knowledge graph, the loss value will increase, prompting the model to adjust the generation strategy.

[0133] To optimize the retrieval efficiency of the knowledge graph, inverted indexes and hash mappings are established for the entities and relationships in the knowledge graph, so that when the user inputs, keywords can be quickly mapped to the corresponding entities or relationships, improving the query efficiency. Support fuzzy queries and multi-keyword queries to meet the diverse input needs of users. Use a caching mechanism to store frequently queried results and reduce the number of database accesses.

[0134] In terms of model inference and hallucination suppression, the model generates a preliminary answer based on the user input and context information, and then integrates the retrieved knowledge graph information with the preliminary answer to ensure the accuracy and professionalism of the answer. The answer generated by the model is matched with the knowledge graph information to calculate the similarity. If the matching degree is low, the system will automatically correct the answer to make it consistent with the knowledge graph information. A user feedback interface is provided to collect the user's satisfaction with the answer for the continuous optimization of the model.

[0135] In terms of model deployment and application, model pruning and quantization are carried out to remove the parameters that have little impact on the model performance and reduce the model size. The model parameters are converted from 32-bit floating-point numbers to 16-bit or even 8-bit to reduce the computational resource requirements. A multi-GPU server is used, and a parallel computing framework is adopted to accelerate the inference of the model. RESTful API and WebSocket interfaces are developed to support the access of multiple terminals such as mobile devices and web pages.

[0136] The operation process of the system includes: the user inputs a question (in the form of voice or text) through a mobile device or a web page; the system preprocesses the user input to extract keywords; according to the keywords, the knowledge graph is retrieved in real time to obtain relevant knowledge information; the model generates a preliminary answer based on the user input and the retrieved knowledge; the preliminary answer is verified to correct the possible hallucination content; the final answer is returned to the user, and a feedback channel is provided.

[0137] As a preferred solution, during the fine-tuning process, to prevent the model from forgetting the general knowledge during pre-training, multiple strategies are adopted. General domain dialogue data is added to the training data and mixed with the data in the tour guide domain for training, so that the model not only masters the professional knowledge in the tour guide domain but also maintains the understanding of general knowledge. In addition, knowledge distillation technology is used, with the pre-trained model as the teacher model and the fine-tuned model as the student model. By minimizing the difference between the outputs of the student model and the teacher model, the memory of general knowledge is retained.

[0138] As a preferred solution, to reduce the hallucination content generated by the model, multiple methods are adopted. A penalty term is added to the loss function. When the content generated by the model is inconsistent with the knowledge graph, the loss value is increased to prompt the model to adjust the generation strategy. At the same time, the possible incorrect content generated by the model is collected and added to the training set as counterexamples to train the model to avoid generating similar incorrect information. The knowledge graph is used to perform real-time verification on the model output. If it is found that the answer content does not match the knowledge graph, it is immediately corrected to ensure the accuracy and reliability of the information provided to the user.

[0139] To more clearly demonstrate the superiority of the technical solution of this application, the following is presented in the following table from the perspectives of fine-tuning speed, model forgetting problem, hallucination phenomenon, knowledge graph retrieval efficiency, inference speed, deployment cost and adaptability.

[0140] Index Existing technical solution Technical solution of this application Improved effect Fine-tuning speed Full-parameter fine-tuning, large parameter update amount, long training cycle (for example, about 60 minutes per round) Based on LoRA technology, only add low-rank matrices to the feature representations of the first layer, significantly reducing the number of parameters and significantly improving the training speed (for example, about 30 minutes per round) The training efficiency is increased by about 50% Model forgetting problem It is easy to forget pre-trained knowledge during the fine-tuning process, and it is difficult to balance general and professional knowledge Introduce contrastive learning and knowledge distillation strategies, adopt mixed training data, ensure the effective integration of new and old knowledge, and significantly reduce the forgetting phenomenon The forgetting rate is reduced by about 35%, and the answer accuracy is increased by about 22% Hallucination phenomenon When the model generates answers, it is easy to appear false or inconsistent information, affecting user trust Utilize the credibility autoregressive regulation mechanism, combine real-time verification and feedback correction with the knowledge graph to ensure that the generated content highly matches professional knowledge The hallucination phenomenon is reduced by about 40% Knowledge graph retrieval efficiency Adopt traditional inverted index or single index method, with slow response speed and low recall rate Implement multi-layer index fusion (inverted index + semantic index + vector retrieval), and combine caching and parallel retrieval strategies to achieve real-time fast retrieval The retrieval speed is increased by 30%, the recall rate is increased by 25%, and the matching accuracy is increased by 18% Inference speed High computing resource requirements during deployment, large response latency Apply model pruning, quantization and multi-GPU parallel computing technologies to significantly compress the model volume and at the same time improve the inference speed The inference response speed is increased by about 50% to meet the requirements of high-concurrency real-time interaction Deployment cost and adaptability High-precision full-parameter models require high hardware configurations and are difficult to deploy in low-resource environments Model quantization, pruning and parameter simplification reduce the requirements for hardware resources and are suitable for deployment in a variety of terminals and low-resource environments The deployment cost is reduced by about 40%, and the applicable range is wider

[0141] The method processes disclosed in the above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. The computer program is written into a storage medium and runs on a corresponding system, so that the system can automatically execute the method processes disclosed in the above embodiments, which will not be elaborated here. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium.

[0142] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as a limitation of the present invention itself. Various changes can be made in its form and details without departing from the spirit and scope of the present invention defined by the appended claims.

Claims

1. An intelligent tour guide human-computer dialogue method based on knowledge graph retrieval, characterized in that: The steps include: Collect relevant data of scenic spots within a predetermined range, perform structured processing on all relevant data of scenic spots, and construct a knowledge graph in the field of tour guides; index the entities and relationships in the knowledge graph, and introduce an adaptive query optimization mechanism; Select a pre-trained language model and perform LoRA fine-tuning on it to obtain a fine-tuned language model; The user initiates a conversation request; The dialogue request accesses the fine-tuned language model through the first channel to obtain a preliminary answer result; the dialogue request accesses the knowledge graph through the second channel to obtain knowledge graph information; The preliminary answer result and the knowledge graph information are matched and integrated to output a revised answer result.

2. The intelligent tour guide human-computer dialogue method based on knowledge graph retrieval according to claim 1 is characterized in that: The structured processing of all scenic spot related data includes: Identify entities from all attraction-related data; Combine the identified entities in pairs and extract the relationship between the two entities based on their context. Construct triples in the form of "entity 1-relationship-entity 2" to form the basic unit of the knowledge graph; Use the graph database to store all triples and build a complete knowledge graph.

3. The intelligent tour guide human-computer dialogue method based on knowledge graph retrieval according to claim 2 is characterized in that: The identifying of entities from all scenic spot related data specifically includes: All the scenic spot related data are sorted into the input sequence X, and the conditional probability The goal is to maximize the value of, identify entities from the input sequence X, and organize the identified entities into the output sequence Y: ; In the formula, is the normalization factor; is the characteristic function; is the weight of the feature function; , are the i-th and i-1-th recognized entities respectively; m is the current recognition round; M is the total recognition rounds; i is the number of entities currently recognized; and n is the total number of recognized entities.

4. The intelligent tour guide human-computer dialogue method based on knowledge graph retrieval according to claim 2 or 3 is characterized in that: The extracting of the relationship between the two entities specifically includes: Will contain entities and entities Context text Represented as a feature vector : ; In the formula, is the embedding function; Train a convolutional neural network as a relation extraction model; transform the feature vector Input into the relation extraction model to extract entities and entities The relationship between ; Among them, the structure of the relation extraction model h is expressed as: ; Relationship Probability Distribution It is expressed as: ; In the formula, , is the weight; , is the bias vector; is the activation function; is the normalization function; Constructing "entity 1-relationship-entity 2" triples , forming the basic unit of the knowledge graph.

5. The intelligent tour guide human-computer dialogue method based on knowledge graph retrieval according to claim 4 is characterized in that: The relationship extraction model adopts the cross entropy loss function The cross entropy loss function expression obtained by training the convolutional neural network is as follows: ; In the formula, is an indicator function that takes the value 1 when the true relationship is k and 0 otherwise.

6. The intelligent tour guide human-computer dialogue method based on knowledge graph retrieval according to claim 1 is characterized in that: Llama3.1 is selected as the pre-trained language model, and LoRA is fine-tuned for it, including: Add trainable low-rank matrices A and B to the self-attention layer and feed-forward network layer in the pre-trained language model. The size of matrix A is [d, r] and the size of matrix B is [r, d], where d is the dimension of the hidden layer and r is the rank of the low-rank matrix. ; Where C is an adaptive adjustment item guided by knowledge graph, which is used to integrate professional tour guide knowledge; , used to dynamically adjust the fine-tuning strength; W represents the parameter matrix of the pre-trained model; In the above way, the scale of parameter update is reduced from [d × d] to [d × r + r × d].

7. The intelligent tour guide human-computer dialogue method based on knowledge graph retrieval according to claim 1 is characterized in that: The matching and fusing of the preliminary answer result and the knowledge graph information specifically includes: The preliminary answer result generated by the fine-tuned language model is matched with the knowledge graph information generated by the knowledge graph, and the credibility is measured by the knowledge consistency score S: ; in, It is measured by calculating the cosine similarity between the two; , , are their respective weights; concept consistency is based on the knowledge points in the same field to determine whether they belong to the same concept set. If they belong to the same concept set, the weight is increased If it does not belong to the question, the score will be lowered; the semantic matching score is based on BERT to calculate the matching degree between the user question and the model answer.

8. The intelligent tour guide human-computer dialogue method based on knowledge graph retrieval according to claim 7 is characterized in that: According to the calculated knowledge consistency score S, the following correction strategy is executed: If S is less than the first threshold, the knowledge enhancement mode is triggered, the knowledge graph is re-searched and the answer is generated; If the first threshold ≤ S < the second threshold, trigger the content modification mode; If S ≥ the second threshold, the answer is directly output, and a dialog box pops up to allow the user to feedback whether the answer is accurate. If the user feedback is wrong, the wrong sample is automatically stored and a penalty item is added in the next fine-tuning.

9. The intelligent tour guide human-computer dialogue method based on knowledge graph retrieval according to claim 8 is characterized in that: The content modification modes include: Use the syntactic reconstruction method based on dependency analysis to adjust the word order of the answer; Or: if the knowledge graph contains relevant images, roadmaps, or historical background information, embed corresponding auxiliary materials in the text answer; Or: Standardize the terms in the answer based on a dictionary of terms in the field.

10. An intelligent tour guide human-computer dialogue system, characterized in that: include: at least one processor; as well as a memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the intelligent tour guide human-computer dialogue method based on knowledge graph retrieval as described in any one of claims 1 to 9.

Citation Information

Cited By

  • Large model anti-forgetting fine tuning method and device, computer equipment and medium

    CN121212374A

  • A large model anti-forgetting fine-tuning method and device, computer equipment and medium

    CN121212374B