Tourism question-answering system based on knowledge graph enhancement

By constructing a knowledge graph-based tourism question-and-answer system, the problems of information fragmentation and insufficient knowledge organization of large language models in the tourism information service system have been solved, realizing personalized and real-time intelligent tourism consultation services and improving the reliability and user experience of the question-and-answer system.

CN121029940APending Publication Date: 2025-11-28TIBET UNIV +2

Patent Information

Application Number
CN202511159299.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

The existing tourism information service system suffers from fragmented, unstructured, and outdated information, making it difficult for users to obtain accurate, real-time, and personalized tourism information. Furthermore, the large language model question-answering system lacks structured/semantic knowledge organization, making it difficult to handle complex intentions and multilingual environments.

Method used

A knowledge graph-based tourism question-answering system is constructed, including a domain ontology module, a knowledge graph construction module, a hybrid index module, a query recognition module, a retrieval enhancement module, an answer generation module, and a personalized recommendation module. Through entity recognition, relation extraction, knowledge fusion, and semantic vector representation, intelligent tourism consultation is achieved.

Benefits of technology

It improves the reliability, adaptability, and service experience of travel Q&A, providing accurate, real-time, and personalized intelligent travel advice, and supports complex intent processing and personalized recommendations in multilingual environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029940A_ABST
    Figure CN121029940A_ABST
Patent Text Reader

Abstract

The invention provides a tourism question-answering system based on knowledge graph enhancement, and the system comprises a domain ontology module which obtains a structured data set of a tourism-related region, and constructs a tourism domain ontology model; the graph construction module is used for constructing a tourism domain knowledge graph in combination with the tourism domain ontology model; the hybrid index module is used for constructing a hybrid vector index and carrying out millisecond-level similarity search; the query identification module is used for generating a structured understanding result according to a query input by a user; the retrieval enhancement module is used for performing multi-path retrieval and retrieval optimization and outputting a query related knowledge fragment set; the answer generation module is used for generating structured tourism answers based on the query related knowledge fragment set; and the personalized recommendation module performs user portrait construction, personalized recommendation, multi-round dialogue management and active interaction. According to the invention, accurate, real-time and personalized intelligent tourism consultation service can be provided for tourists, and the reliability, adaptability and service experience of tourism questions and answers are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of tourism question and answer, in particular to a tourism question and answer system based on knowledge graph enhancement. BACKGROUND

[0002] China has vast land, from south to north, from east to west, in the vast territory, mountains, rivers, lakes, forests, deserts, valleys and plains, unique and beautiful tourist destinations are scattered everywhere. For example, Tibet Autonomous Region has unique plateau natural scenery and rich ethnic culture, and is an important tourist destination for domestic and foreign tourists. With the vigorous support of the state for the cultural industry, the tourism industry has ushered in an unprecedented development opportunity and has gradually grown into an important force to promote regional economy and cultural dissemination.

[0003] Although the tourism demand in Tibet is strong, the tourism information service system has certain lag. In addition to the relatively complete information of a few famous scenic spots, the information of other scenic spots presents the problems of fragmentation, unstructured and lagging update. Tourists, especially when they visit a destination for the first time, often have difficulty obtaining complete, accurate and timely information through traditional web search or tourism platforms.

[0004] Currently, users mainly obtain tourism information through three ways: 1. Web search: information is disorganized and unstructured, and users need to screen and integrate it themselves; 2. Online tourism platform: the platform provides more data, but the content is inconsistent across platforms, and the credibility and comparability are insufficient; 3. Social media platform: the guide is generated by users, and the content quality is uneven, lacking systematicness and authority.

[0005] At the same time, the rise of large language models makes it possible to build a new type of question and answer system. Using its powerful language understanding and generation capabilities can realize intelligent question and answer services with natural human-computer interaction. However, directly relying on general large language models has the following problems: 1) data islands or inconsistent knowledge expression, lack of structured / semantic knowledge organization. 2) Question and answer systems mostly stay in the keyword / template matching stage, and it is difficult to handle complex intentions, ambiguous questions and multi-language environments. 3) Retrieval and answer generation are separated, and the recommendation personalization and conversation continuity are weak; query fault tolerance, self-adaptation and synonym recognition ability are weak. SUMMARY

[0006] In order to overcome the shortcomings of the prior art, the purpose of the present application is to provide a tourism question and answer system based on knowledge graph enhancement, which can provide accurate, real-time and personalized intelligent tourism consulting services for tourists, greatly improving the reliability, adaptability and service experience of tourism question and answer.

[0007] To achieve the above purpose, the present application provides the following scheme: a tourism question and answer system based on knowledge graph enhancement, comprising: a domain ontology module configured to obtain a structured dataset of a travel-related region, define a domain specification description based on the structured dataset, construct a preliminary ontology based on the domain specification description, and construct a travel domain ontology model based on the preliminary ontology; a graph construction module configured to perform knowledge extraction on the structured data using the travel domain ontology model, and perform entity recognition, hybrid relation extraction, knowledge fusion, and entity-relation structure mapping based on the extracted knowledge to obtain a travel domain knowledge graph; a hybrid index module configured to generate semantic vector representations using the travel domain knowledge graph to construct a hybrid vector index and perform millisecond-level similarity search; a query identification module configured to, for a user input query, perform query normalization, intent recognition, slot extraction, and intelligent query rewriting using the hybrid vector index to generate a structured understanding result; a retrieval enhancement module configured to, based on the structured understanding result, perform multi-path retrieval and retrieval optimization in the travel domain knowledge graph and the hybrid vector index to output a set of query-related knowledge fragments; an answer generation module configured to, based on the set of query-related knowledge fragments, generate a structured travel answer using a reasoning question answering model through a prompt template in combination with the travel domain knowledge graph and user historical queries; a personalized recommendation module configured to perform user portrait construction, personalized recommendation, multi-round dialogue management, and active interaction to implement personalized intelligent travel services; The domain ontology module, the graph construction module, the hybrid index module, the query identification module, the retrieval enhancement module, the answer generation module, and the personalized recommendation module are interconnected.

[0008] Optionally, the domain ontology module comprises: a data acquisition and processing unit configured to obtain multi-source heterogeneous data of the travel-related region from official platforms, commercial platforms, social media, and professional literature, and perform deduplication, noise filtering, format unification and standardization, and multi-dimensional label addition and structure completion on the multi-source heterogeneous data to obtain a structured dataset; a system definition unit configured to perform preliminary classification of domain terms based on the structured dataset to obtain a preliminary term library and a synonym dictionary, and then perform class system design, relationship system design, and attribute system design in sequence to obtain a domain specification description; An initial ontology framework unit is configured to create a customized ontology project based on the domain specification description, establish core categories in the customized ontology project, add sub-categories and instances in the core categories, describe relationships between two core categories or instances to define object attributes, describe specific data features of instances to define data attributes, import the synonym dictionary of the unified entity, output a standardized file, and complete construction of an initial ontology; An ontology model unit is configured to perform knowledge extraction, domain named entity recognition training, relationship extraction, and entity unification and fusion for the initial ontology, complete knowledge docking of the initial ontology, and perform category set definition, attribute set definition, and relationship set definition on the initial ontology to obtain a tourism domain ontology model.

[0009] Optionally, the category system includes scenic spots, activities, transportation, accommodation, food, culture, and others, the relationship system includes spatial relationship, proximity relationship, distance relationship, time relationship, best tour period, function-service relationship, inclusion relationship, suitable object relationship, recommendation-series relationship, and tour route inclusion, and the attribute system includes basic attributes, characteristic attributes, and dynamic attributes, the basic attributes include name, description, address, and contact information, the characteristic attributes include altitude, ticket price, tour duration, suitable crowd, hotel rating, food recommendation, and activity time, and the dynamic attributes include real-time weather, passenger flow, opening and closing status, and temporary announcement.

[0010] Optionally, the graph construction module includes: A knowledge extraction unit is configured to perform knowledge extraction on the structured data by using the tourism domain ontology model to obtain extracted knowledge. An entity recognition unit is configured to select a Chinese-BERT-wwm model for fine-tuning based on the extracted knowledge, add a CRF layer to capture the dependency relationship between labels, obtain an NER model, and perform entity recognition operations. A relationship extraction unit is configured to simultaneously extract triplets in a unified format by using a rule template, a BERT-based relationship classification model, and weakly supervised learning technology, and complete hybrid relationship extraction. A knowledge fusion unit is configured to perform entity alignment based on string similarity and semantic similarity, combine context information to perform entity disambiguation, and use a graph embedding method to perform knowledge completion to complete knowledge fusion operations. A structure mapping unit is configured to use a Neo4j graph database to store entities, entity attributes, and relationships between entities, perform node and relationship modeling, construct a full-text index in the Neo4j graph database, complete graph storage operations, and obtain a tourism domain knowledge graph.

[0011] Optionally, the hybrid index module includes: The text vectorization unit is used to select the BGE-M3 model that supports multiple languages, multi-granularity coding, and multi-domain adaptation. Based on the relevant tourism regions, the selected model is fine-tuned for domain adaptation, and then entity attribute concatenation and multi-granularity feature generation are performed to obtain the text vector. The graph embedding unit is used to capture graph structure information using Node2Vec for structural embedding, combine it with BERT encoding of entity description text for semantic embedding, and concatenate or weightedly integrate the structural embedding and the semantic embedding for graph fusion representation. The index building unit is used to select a vector database, and based on the selected vector database, it performs table structure definition, dense vector index construction, sparse vector index construction, and metadata filtering to obtain a hybrid vector index to support millisecond-level similarity search.

[0012] Optionally, the query identification module includes: The query processing unit is used to perform simplified / traditional Chinese conversion, pinyin error correction, synonym expansion, abbreviation restoration, and cleaning and standardization on user-input queries to obtain standardized queries. The intent classification unit is used to perform multi-label classification output based on the normalized query using a multi-label classification model to obtain intent categories; the intent categories include attraction information, itinerary planning, cultural understanding, and practical information. The slot extraction unit is used to identify key information based on the intent category using a sequence labeling model, so as to extract location information slots, time information slots, budget information slots and crowd information slots to obtain a structured query; The query rewriting unit is used to expand the query based on the structured query by using the vector nearest neighbor results of synonyms and similar entities, remove irrelevant modifiers and colloquial interjections to simplify the query, and decompose the query by decomposing compound questions to complete the query rewriting and optimization operations and generate structured understanding results.

[0013] Optionally, the retrieval enhancement module includes: The multi-path retrieval unit is used to perform vector retrieval, graph retrieval, and keyword full-text retrieval based on the structured understanding results, thereby completing the multi-path retrieval. The retrieval optimization unit is used to fuse and score multiple retrieval results based on general relevance, thereby reordering the results in a personalized manner to obtain optimized retrieval results. The calculation expression for the fusion score is as follows: ; in, The final ranking score, For vector retrieval semantic relevance scores, The confidence level for the spectral structure. The relevance score of the keyword search text. As a score for user interest, All are dynamic weights; The knowledge fusion unit is used to automatically filter entity-fact conflicts returned by different searches based on the search optimization results, using local constraints and timeliness scoring to complete conflict detection and constraints. Then, the search optimization results are sorted and scored to prioritize the recommendation of highly reliable content and output a set of query-related knowledge fragments.

[0014] Optionally, the vector retrieval is used to recall Top-N semantically related entities, descriptions, and FAQ fragments in the Milvus vector library based on the query vector or slot key phrase in the structured understanding results, using cosine similarity, to support semantic near-synonym recall and synonym variant recognition, and to adapt to multilingual fuzzy questions. The graph retrieval is used to perform path lookup using Cypher queries based on the explicit slots in the structured understanding results, in order to perform relational reasoning and attribute filtering; The keyword full-text search is used to retrieve relevant documents or entities based on the original text or extended phrases in the structured understanding results, using word segmentation, indexing, and inverted indexing.

[0015] Optionally, the answer generation module includes: The prompt template unit is used to design dynamic prompt templates that include a set of query-related knowledge fragments, user intent, domain safety, domain culture, and precautions; The model building unit is used to select ChatGLM3-6B as the base model, and to perform INT8 quantization, LoRA fine-tuning and vLLM framework inference optimization on the base model to obtain an inference question answering model to output the initial answer to the user query. The answer generation strategy unit is used to control answer generation through temperature sampling parameters, length control strategy, and security filtering strategy; wherein, the temperature sampling parameter is set to 0.7; The answer post-processing unit is used to compare the content of the initial answer with the tourism domain knowledge graph for fact verification, and then optimize the format and segment the initial answer, and add source introduction and annotation to obtain the tourism answer.

[0016] Optionally, the personalized recommendation module includes: The user profiling unit is used to collect basic user information, travel preferences, and behavioral characteristics to obtain multi-dimensional data in order to build user profiles. The personalized recommendation unit is used to perform collaborative filtering based on the behavior of similar users, recommend content based on user interests and content features, and recommend knowledge based on the association reasoning of the knowledge graph in the tourism field, so as to obtain the recommendation output. The multi-turn dialogue management unit is used for dialogue status tracking, dialogue history maintenance, slot inheritance, and topic switching, enabling the transmission of key information and topic identification and conversion across turns. The proactive interaction unit is used to proactively ask questions when a user's query for information is insufficient, in order to improve the user profile.

[0017] This invention discloses the following technical effects by providing a tourism question-answering system based on knowledge graph enhancement: 1. By establishing an ontology system oriented towards specific tourism sectors, the knowledge structure within the industry is unified, and rich and diverse tourism information is accurately expressed, facilitating subsequent knowledge extraction and reasoning; it is also beneficial for entity disambiguation, synonym processing, and multilingual support, thereby enhancing the depth and breadth of the system's semantic understanding.

[0018] 2. By constructing a tourism domain knowledge graph with a clear structure, rich entities, and realistic relationships, it is possible to eliminate duplication and errors in heterogeneous data; support continuous knowledge updates, real-time expansion, and multi-granularity queries and reasoning; and provide high-quality data sources and knowledge structures, which are crucial for accurate question answering and recommendations.

[0019] 3. By building a hybrid index and enhancing retrieval, users can initiate searches in multiple ways and at multiple granularities (keywords / phrases / semantic relevance), with fast response and wide coverage; at the same time, it supports intelligent question answering and semantic expansion queries, and realizes the recognition of synonym variants and multilingual fuzzy questions, so as to improve the system's retrieval capabilities, the speed of intelligent question answering and the accuracy of answers.

[0020] 4. Through multi-tag classification, slot extraction, and intelligent query rewriting, it achieves standardized processing such as simplified / traditional Chinese / synonym recovery, colloquial abstraction, and question decomposition, providing high fault tolerance and adaptive optimization for user input; significantly improving the interpretability and subsequent processing efficiency of complex and diverse user input; achieving personalized understanding and supporting continuous questioning and contextual connection.

[0021] 5. Through answer generation control, it can generate natural, fluent, knowledge-rich, personalized, and trustworthy answers, avoiding cold and mechanical responses; it also supports summarization, expansion, multi-level expression, and fact-checking, improving user experience.

[0022] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a schematic diagram of the system architecture provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the spectrum construction process provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the process of generating the answer provided in the embodiments of the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0027] like Figure 1 As shown, this invention provides a tourism question-answering system based on knowledge graph enhancement, comprising: 1. Domain Ontology Module This process involves acquiring structured datasets of tourism-related regions, defining domain specification descriptions based on these datasets, constructing an initial version of the ontology based on these domain specification descriptions, and then building a tourism domain ontology model based on the initial version. The domain ontology module includes: 1.1 Data Acquisition and Processing Unit This tool is used to obtain multi-source heterogeneous data on tourism-related regions from official platforms, commercial platforms, social media, and professional literature. The multi-source heterogeneous data is deduplicated, noise filtered, format unified and standardized, and multi-dimensional labels are added and the structure is completed to obtain a structured dataset.

[0028] Deduplication: Based on SimHash, the cleaning pipeline first performs SimHash deduplication based on the content and description fields, filtering out duplicate collections with minor changes.

[0029] Noise filtering: Scan all fields and use regular expressions to remove special characters, tags, and redundant advertising text. Standardized and consistent formatting: Multi-source field mappings are uniformly mapped to the JSON standard model. For example, for scenic spot entities, a schema validation is performed uniformly before storage in MongoDB / ES / Neo4j.

[0030] Multidimensional tags and structural completion: Automatically fill in missing fields, such as completing administrative regions / latitude and longitude by location name using map API; add multiple tags, such as "suitable for photography" / "Tibetan culture" / "suitable for travel in all seasons"; merge similar entities, such as "Potala Palace" / "Potala Palace" into one record.

[0031] 1.2 System Definition Unit Based on the structured dataset, the domain terminology is initially classified to obtain a preliminary terminology database and a thesaurus. Then, the category system design, relation system design, and attribute system design are carried out in sequence to obtain the domain specification description.

[0032] The category system includes attractions, activities, transportation, accommodation, food, culture, and others; taking Tibet as an example, as shown in Table 1 below: Table 1 Core Category System

[0033] The relationship system includes spatial relationships, proximity relationships, distance relationships, temporal relationships, optimal visiting season, function-service relationships, inclusion relationships, suitable target group relationships, recommendation-connection relationships, and tour route inclusion. Taking Tibet as an example, as shown in Table 2 below: Table 2 Relationship System Modeling

[0034] The attribute system includes: Basic attributes: name, description, address, and contact information.

[0035] Key features: altitude, ticket price, tour duration, suitable audience, hotel class, recommended restaurants, and activity dates.

[0036] Dynamic attributes: real-time weather, passenger flow, on / off status, and temporary announcements.

[0037] 1.3 Initial Body Framework Units Based on the domain specification description, a custom ontology project is created. In the custom ontology project, a core category Class is established. Subclasses and instances are added to the core category. The relationship between two core categories or instances is described to define object attributes. The specific data characteristics of the instance are described to define data attributes. Then, the thesaurus of the unified entity is imported, and a standardized file is output to complete the construction of the initial version ontology.

[0038] 1.4 Ontology Model Unit This is used to perform knowledge extraction, domain named entity recognition training, relation extraction, and entity normalization and fusion on the initial version of the ontology, thereby completing the knowledge docking of the initial version of the ontology. Then, the initial version of the ontology is used to define the category set, attribute set, and relation set to obtain the tourism domain ontology model.

[0039] 2. Atlas Construction Module like Figure 2 As shown, the system utilizes the tourism domain ontology model to extract knowledge from the structured data, and then performs entity recognition, hybrid relation extraction, knowledge fusion, and entity-relation structure mapping based on the extracted knowledge to obtain a tourism domain knowledge graph. The graph construction module includes: 2.1 Knowledge Extraction Unit This is used to extract knowledge from the structured data using the tourism domain ontology model, thereby obtaining extracted knowledge.

[0040] Knowledge extraction strategy: Structured text: Direct field mapping to ontology categories and attributes.

[0041] Semi / unstructured text: Named entity recognition + syntactic relation extraction + rule templates.

[0042] 2.2 Entity Recognition Unit Based on the extracted knowledge, the Chinese-BERT-wwm model is selected for fine-tuning, and a CRF layer is added to capture the dependencies between labels, resulting in a NER model for entity recognition.

[0043] 2.3 Relation Extraction Unit This tool is used to simultaneously extract triples in a unified format using rule templates, a BERT-based relation classification model, and weakly supervised learning techniques, thus completing hybrid relation extraction.

[0044] Rule Template: Structured data / obvious field descriptions. For dataA that already has associated fields, construct triples using a condition template, for example: The field "Namtso" with the location value "Nagqu City Damxung County" is followed by a ternary set (Namtso, locatedIn, Nagqu City Damxung County). The `hotel` field in the accommodation table contains `related to the attraction` → (XX hotel, nearBy, XX attraction). Deep Learning: Relation Classification. Using BERT as input, relations are determined from sentence pairs or entity pairs. Given the NER results, sentence templates are filled in, such as "Namtso Lake is located in Nagqu City". Relation category sets are labeled, such as `locatedIn`, `nearBy`, `openTime`, `bestVisitTime`, `provides`, etc.

[0045] Remote supervision / weak supervision: Utilize a known small-scale knowledge base, map its names and expand it with high-confidence negative examples to achieve self-supervised data expansion and improve extraction coverage.

[0046] Unified triple output: All extraction methods yield triples in a uniform format: subject entity, relation, object entity or value. For example: "Lhasa Shangri-La Hotel, locatedIn, Lhasa City", "Lhasa Shangri-La Hotel, provides, free breakfast".

[0047] 2.4 Knowledge Integration Unit This method is used for entity alignment based on string similarity and semantic similarity, entity disambiguation combined with contextual information, and knowledge completion using graph embedding methods to complete the knowledge fusion operation.

[0048] 2.5 Structural Mapping Unit This tool utilizes the Neo4j graph database to store entities, entity attributes, and relationships between entities, performs node and relationship modeling, builds a full-text index in the Neo4j graph database, completes graph storage operations, and obtains a knowledge graph for the tourism field.

[0049] Point / relationship modeling: Each entity is represented by a node. Node attributes include entity type, name, alias, description, location / association, and unique attributes. Relationships in triples are mapped to Neo4j edges, each with relation type and required attributes.

[0050] Full-text search: Build a Uniq index based on name, type, etc., to accelerate node deduplication and searching. Configure the full-text index to support fuzzy matching of subsequent user-input free text.

[0051] 3. Hybrid Index Module like Figure 3 As shown, a method is used to generate semantic vector representations using the tourism domain knowledge graph to construct a hybrid vector index for millisecond-level similarity search. The hybrid index module includes: 3.1 Text Vectorization Unit The BGE-M3 model is selected to support multilingual, multi-granularity coding, and multi-domain adaptation. Based on the relevant tourism regions, the selected model is fine-tuned for domain adaptation, and then entity attribute concatenation and multi-granularity feature generation are performed to obtain text vectors.

[0052] 3.2 Spectrum Embedding Unit This is used to capture graph structure information using Node2Vec for structural embedding, combine it with BERT encoding of entity description text for semantic embedding, and concatenate or weightedly integrate the structural embedding and the semantic embedding for graph fusion representation.

[0053] 3.3 Index Building Unit The vector database is selected, and table structure definition, dense vector index construction, sparse vector index construction, and metadata filtering are performed based on the selected vector database to obtain a hybrid vector index to support millisecond-level similarity search.

[0054] 4. Query and Recognition Module like Figure 3 As shown, the system utilizes the hybrid vector index to perform query normalization, intent recognition, slot extraction, and intelligent query rewriting based on user input queries, thereby generating structured understanding results. The query recognition module includes: 4.1 Query Processing Unit This function is used to perform simplified / traditional Chinese conversion, pinyin correction, synonym expansion, abbreviation restoration, and cleaning and standardization on user-input queries to obtain standardized queries.

[0055] Traditional / Simplified Chinese Conversion and Pinyin Error Correction: The OpenCC library is used for conversion between Traditional and Simplified Chinese. Error correction is achieved using Pinyin sequences and a dictionary of incorrect characters, improving input compatibility for younger children or those with difficulty typing Chinese.

[0056] Synonym Expansion: Construct a thesaurus specifically for Tibet tourism (maintained according to ontology category B, such as "Potala Palace" = "Potala Palace" = "Pearl of the Snow Region"). When synonyms appear in a query, they will be uniformly replaced with the standard name, and synonym expressions will be added to increase recall.

[0057] Abbreviation Restoration: Maintain an abbreviation-full name lookup table (e.g., "Tibetan cuisine" → "Tibetan food and beverage"). Accurate restoration via rules or table lookup.

[0058] Cleaning and standardization: Remove redundant punctuation and spaces. All text is now in Simplified Chinese.

[0059] 4.2 Intent Classification Unit Based on the normalized query, a multi-label classification model is used to perform multi-label classification output to obtain the intent category; the intent category includes attraction information, itinerary planning, cultural understanding and practical information.

[0060] 4.3 Slot Extraction Unit Based on the intent category, the sequence labeling model is used to identify key information, extract location information slots, time information slots, budget information slots, and crowd information slots to obtain a structured query.

[0061] Sequence labeling is a task in Natural Language Processing (NLP) that aims to assign a specific label to each token (word or character) in a text in order to identify its role in the sequence.

[0062] "Slot" is a classification definition for key information, with each slot corresponding to a specific type of information: Location slot: Stores location information in the text.

[0063] Time slot: Stores time-related information, such as "May 1, 2024", "next Friday afternoon", "3 days later".

[0064] Budget slot: Stores information related to expenses, such as "5,000 yuan", "budget under 20,000 yuan", and "300 US dollars per person".

[0065] Audience slots: Store information related to the target audience, such as "a family of three", "college students", and "elderly people over 60 years old".

[0066] 4.4 Query and Rewrite Unit Based on the structured query, the query is expanded using the vector nearest neighbor results of synonyms and similar entities, irrelevant modifiers and colloquial interjections are removed to simplify the query, and the query is decomposed by decomposing compound questions to complete the query rewriting and optimization operations and generate structured understanding results.

[0067] 5. Search Enhancement Module like Figure 3 As shown, the system is used to perform multi-path retrieval and retrieval optimization based on the structured understanding results in the tourism domain knowledge graph and the hybrid vector index, to output a set of query-related knowledge fragments. The retrieval enhancement module includes: 5.1 Multi-way search unit This is used to perform vector retrieval, graph retrieval, and keyword full-text retrieval based on the structured understanding results, thus completing multi-path retrieval; wherein: The vector retrieval is used to recall Top-N semantically relevant entities, descriptions, and FAQ fragments from the Milvus vector library based on the query vector or slot key phrase in the structured understanding results, using cosine similarity, to support semantic near-synonym recall and synonym variant recognition, and to adapt to multilingual fuzzy questions.

[0068] The graph retrieval is used to perform path lookup using Cypher queries based on the explicit slots in the structured understanding results, in order to perform relational reasoning and attribute filtering.

[0069] The keyword full-text search is used to retrieve relevant documents or entities based on the original text or extended phrases in the structured understanding results, using word segmentation, indexing, and inverted indexing.

[0070] 5.2 Search Optimization Unit This method is used to fuse and score multi-way search results based on general relevance, thereby personalizing and reordering the multi-way search results to obtain optimized search results; the calculation expression for the fusion score is: ; in, The final ranking score, For vector retrieval semantic relevance scores, The confidence level for the spectral structure. The relevance score of the keyword search text. As a score for user interest, All weights are dynamic.

[0071] 5.3 Knowledge Integration Unit Based on the search optimization results, this tool automatically filters entity-fact conflicts returned by different searches using local constraints and timeliness scoring, completes conflict detection and constraint, and then sorts and scores the search optimization results to prioritize and recommend highly reliable content, outputting a set of query-related knowledge fragments.

[0072] Local constraints refer to localized rules or restrictions that are relevant to the retrieval scenario or system itself, and these conditions affect the filtering and judgment of information. They are usually related to specific application scenarios, business needs, or system settings, and are highly targeted.

[0073] Timeliness scoring: This involves evaluating and quantifying the time-sensitive validity of retrieved information to determine whether it meets current needs. Because many facts change over time, newer information is generally more reliable.

[0074] 6. Answer generation module like Figure 3 As shown, this module is used to generate structured tourism answers based on the set of query-related knowledge fragments, utilizing a reasoning question-and-answer model, prompt templates, and combining the tourism domain knowledge graph and user query history. The answer generation module includes: 6.1 Prompt Template Unit This is used to design dynamic prompt templates that include a set of relevant knowledge snippets, user intent, domain safety, domain culture, and precautions.

[0075] 6.2 Model Building Unit ChatGLM3-6B was selected as the base model. INT8 quantization, LoRA fine-tuning, and vLLM framework inference optimization were performed on the base model to obtain an inference question-answering model, which outputs the initial answer to the user's query.

[0076] Quantization is a model compression technique aimed at making models lighter and more portable, so they can run on ordinary devices.

[0077] INT8 quantization: Converts high-precision numbers into low-precision integers, such as 8-bit integers. This is equivalent to simplifying the 8 decimal places into integer parts. Although it will lose a little precision, it can significantly reduce the memory occupied by the model.

[0078] LoRA fine-tuning is a low-cost fine-tuning technique. Ordinary fine-tuning requires modifying all parameters of the model, which is time-consuming and labor-intensive; while LoRA only modifies a small number of key parameters in the model, like giving the model "targeted tutoring". It can enable the model to learn domain knowledge without spending too many computing resources, and can also retain the model's original general capabilities.

[0079] vLLM framework: It is a tool specifically designed to accelerate inference for large models, similar to an accelerator. By optimizing computation methods, such as efficiently utilizing GPU memory and batch processing requests, it significantly improves the speed at which models generate responses, potentially reducing the time from 5 seconds to less than 1 second, making dialogues smoother.

[0080] 6.3 Answer Generation Strategy Unit It is used to control answer generation through temperature sampling parameters, length control strategies, and security filtering strategies.

[0081] Temperature sampling is a core parameter for controlling the flexibility of AI responses. The higher the value, the more free and diverse the AI's responses; the lower the value, the more fixed and conservative the responses. A value of 0.7 is a common balance point, suitable for most scenarios, allowing for variations in responses without resorting to fabrication.

[0082] Length control: Dynamically adjusts based on question type. This means the AI ​​will automatically adjust the length of the answer based on the length or complexity of the question, avoiding answers that are too short and unclear or too long and verbose.

[0083] Security filtering: Filtering sensitive and inappropriate content. This means that before generating an answer, the AI ​​will automatically check whether the content contains sensitive information or inappropriate remarks, such as illegal content, discriminatory language, or malicious guidance. If so, it will be filtered out to avoid outputting harmful information.

[0084] 6.4 Answer Post-processing Unit This is used to compare the content of the initial answer with the knowledge graph of the tourism field for fact verification, and then to optimize the format and segment the initial answer, add a brief introduction and annotation of the source, and obtain the tourism answer.

[0085] 7. Personalized Recommendation Module This module is used for user profiling, personalized recommendations, multi-turn dialogue management, and proactive interaction to achieve personalized intelligent tourism services. The personalized recommendation module includes: 7.1 User Profile Unit It is used to collect basic user information, travel preferences and behavioral characteristics to obtain multi-dimensional data in order to build user profiles.

[0086] Basic information: Obtained during registration / inquiry, such as age, gender, and region.

[0087] Travel preferences: Questionnaire on travel type selection (cultural / natural) and budget / travel mode.

[0088] Behavioral analysis: historical queries, selected attraction tags, and interest ratings.

[0089] Behavioral characteristics: Query historical search logs, such as location, keywords, topic categories, etc.

[0090] Key behaviors within the platform: clicks, favorites, likes, feedback, and ratings.

[0091] 7.2 Personalized Recommendation Unit It is used for collaborative filtering based on the behavior of similar users, content recommendation based on user interests and content features, and knowledge recommendation based on the association reasoning of the knowledge graph in the tourism field, to obtain recommendation output; 7.3 Multi-turn Dialogue Management Unit It is used for dialogue state tracking, dialogue history maintenance, slot inheritance and topic switching, enabling the transmission of key information and topic identification and conversion across rounds.

[0092] Dialogue state tracking refers to the core information and progress of the current conversation, such as what questions the user has asked, what key information has been said, and what topic is being discussed. Tracking allows AI to record and update this information in real time to ensure the conversation continues smoothly.

[0093] Dialogue history: Maintaining the most recent N rounds of dialogue allows the AI ​​to remember the content of recent conversations and avoid amnesia. Maintaining the most recent N rounds: The AI ​​will not remember all conversations indefinitely, otherwise it would slow down. Usually, only the most recent 5-10 rounds are retained, which ensures contextual coherence without wasting resources.

[0094] Slot inheritance: Passing key information across rounds. A slot can be understood as a key information point that needs to be remembered, such as time, place, person, or need; inheritance allows this information to be automatically passed across multiple rounds of conversation without the user having to repeat it, making the conversation more efficient.

[0095] Topic switching: Identifying and handling topic shifts. This means enabling AI to detect changes in the topic and adjust its response accordingly, avoiding irrelevant answers.

[0096] 7.4 Active Interaction Unit This feature is used to proactively ask follow-up questions when a user's query for information is insufficient, thereby improving the user profile. Specifically, when a slot / information is detected to be incomplete, it automatically asks follow-up questions to complete the profile, such as "What month and day do you plan to depart?" or "Do you prefer budget or luxury accommodations?" Through multiple rounds of follow-up questions, the user profile is improved, significantly enhancing recommendation accuracy.

[0097] Therefore, by providing a knowledge graph-based tourism question-and-answer system, this invention can provide tourists with accurate, real-time, and personalized intelligent tourism consultation services, greatly improving the reliability, adaptability, and service experience of tourism question-and-answer.

[0098] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0099] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A tourism question-answering system based on knowledge graph enhancement, characterized in that, include: The domain ontology module is used to obtain structured datasets of tourism-related regions, define domain specification descriptions based on the structured datasets, construct an initial version of the ontology based on the domain specification descriptions, and then construct a tourism domain ontology model based on the initial version of the ontology. The knowledge graph construction module is used to extract knowledge from the structured data using the tourism domain ontology model, and then perform entity recognition, hybrid relation extraction, knowledge fusion and entity-relation structure mapping based on the extracted knowledge to obtain a tourism domain knowledge graph. The hybrid index module is used to generate semantic vector representations using the tourism domain knowledge graph to construct a hybrid vector index and perform millisecond-level similarity search. The query recognition module is used to perform query normalization, intent recognition, slot extraction, and intelligent query rewriting based on the hybrid vector index, in order to generate structured understanding results for user-input queries. The retrieval enhancement module is used to perform multi-way retrieval and retrieval optimization based on the structured understanding results in the tourism domain knowledge graph and the hybrid vector index, so as to output a set of query-related knowledge fragments; The answer generation module is used to generate structured tourism answers based on the set of query-related knowledge fragments, using a reasoning question-and-answer model, prompt templates, and combining the tourism knowledge graph and user historical queries. The personalized recommendation module is used for user profile building, personalized recommendations, multi-turn dialogue management and proactive interaction to achieve personalized intelligent tourism services. The domain ontology module, the graph construction module, the hybrid index module, the query recognition module, the retrieval enhancement module, the answer generation module, and the personalized recommendation module are interconnected.

2. The tourism question-answering system based on knowledge graph enhancement according to claim 1, characterized in that, The domain ontology module includes: The data acquisition and processing unit is used to acquire multi-source heterogeneous data on tourism-related regions from official platforms, commercial platforms, social media and professional literature, and to perform deduplication, noise filtering, format unification and standardization, and multi-dimensional label addition and structure completion on the multi-source heterogeneous data to obtain a structured dataset. The system definition unit is used to perform preliminary classification of domain terms based on the structured dataset to obtain a preliminary terminology library and a thesaurus, and then perform category system design, relation system design and attribute system design in sequence to obtain a domain specification description. An initial ontology framework unit is used to create a custom ontology project based on the domain specification description. In the custom ontology project, core categories are established, subclasses and instances are added to the core categories, the relationship between two core categories or instances is described to define object attributes, the specific data characteristics of instances are described to define data attributes, the thesaurus of synonyms of unified entities is imported, and a standardized file is output to complete the construction of the initial ontology. The ontology model unit is used to perform knowledge extraction, domain named entity recognition training, relation extraction, and entity normalization and fusion on the initial version of the ontology to complete the knowledge docking of the initial version of the ontology. Then, the initial version of the ontology is defined by defining the category set, attribute set, and relation set to obtain the tourism domain ontology model.

3. A tourism question-answering system based on knowledge graph enhancement according to claim 2, characterized in that, The category system includes attractions, activities, transportation, accommodation, food, culture, and others. The relationship system includes spatial relationships, proximity relationships, distance relationships, time relationships, best time to visit, function-service relationships, inclusion relationships, suitable audience relationships, recommendation-connection relationships, and tour route inclusion. The attribute system includes basic attributes, special attributes, and dynamic attributes. Basic attributes include name, description, address, and contact information. Special attributes include altitude, ticket price, tour duration, suitable audience, hotel rating, food recommendations, and activity times. Dynamic attributes include real-time weather, visitor flow, on / off status, and temporary announcements.

4. A tourism question-answering system based on knowledge graph enhancement according to claim 3, characterized in that, The map construction module includes: The knowledge extraction unit is used to extract knowledge from the structured data using the tourism domain ontology model to obtain extracted knowledge. The entity recognition unit is used to select the Chinese-BERT-wwm model for fine-tuning based on the extracted knowledge, and add a CRF layer to capture the dependencies between labels to obtain the NER model for entity recognition operation; The relation extraction unit is used to simultaneously extract triples in a unified format using rule templates, a BERT-based relation classification model, and weakly supervised learning techniques to complete hybrid relation extraction. The knowledge fusion unit is used to perform entity alignment based on string similarity and semantic similarity, perform entity disambiguation by combining contextual information, and perform knowledge completion using graph embedding methods to complete the knowledge fusion operation. The structure mapping unit is used to store entities, entity attributes, and relationships between entities using the Neo4j graph database, perform node and relationship modeling, build a full-text index in the Neo4j graph database, complete the graph storage operation, and obtain a knowledge graph in the tourism field.

5. A tourism question-answering system based on knowledge graph enhancement according to claim 4, characterized in that, The hybrid index module includes: The text vectorization unit is used to select the BGE-M3 model that supports multiple languages, multi-granularity coding, and multi-domain adaptation. Based on the relevant tourism regions, the selected model is fine-tuned for domain adaptation, and then entity attribute concatenation and multi-granularity feature generation are performed to obtain the text vector. The graph embedding unit is used to capture graph structure information using Node2Vec for structural embedding, combine it with BERT encoding of entity description text for semantic embedding, and concatenate or weightedly integrate the structural embedding and the semantic embedding for graph fusion representation. The index building unit is used to select a vector database, and based on the selected vector database, it performs table structure definition, dense vector index construction, sparse vector index construction, and metadata filtering to obtain a hybrid vector index to support millisecond-level similarity search.

6. A tourism question-answering system based on knowledge graph enhancement according to claim 5, characterized in that, The query identification module includes: The query processing unit is used to perform simplified / traditional Chinese conversion, pinyin error correction, synonym expansion, abbreviation restoration, and cleaning and standardization on user-input queries to obtain standardized queries. The intent classification unit is used to perform multi-label classification output based on the normalized query using a multi-label classification model to obtain intent categories; the intent categories include attraction information, itinerary planning, cultural understanding, and practical information. The slot extraction unit is used to identify key information based on the intent category using a sequence labeling model, so as to extract location information slots, time information slots, budget information slots and crowd information slots to obtain a structured query; The query rewriting unit is used to expand the query based on the structured query by using the vector nearest neighbor results of synonyms and similar entities, remove irrelevant modifiers and colloquial interjections to simplify the query, and decompose the query by decomposing compound questions to complete the query rewriting and optimization operations and generate structured understanding results.

7. A tourism question-answering system based on knowledge graph enhancement according to claim 6, characterized in that, The retrieval enhancement module includes: The multi-path retrieval unit is used to perform vector retrieval, graph retrieval, and keyword full-text retrieval based on the structured understanding results, thereby completing the multi-path retrieval. The retrieval optimization unit is used to fuse and score multiple retrieval results based on general relevance, thereby reordering the results in a personalized manner to obtain optimized retrieval results. The calculation expression for the fusion score is as follows: ; in, The final ranking score, For vector retrieval semantic relevance scores, The confidence level for the spectral structure. The relevance score of the keyword search text. As a score for user interest, All are dynamic weights; The knowledge fusion unit is used to automatically filter entity-fact conflicts returned by different searches based on the search optimization results, using local constraints and timeliness scoring to complete conflict detection and constraints. Then, the search optimization results are sorted and scored to prioritize the recommendation of highly reliable content and output a set of query-related knowledge fragments.

8. A tourism question-answering system based on knowledge graph enhancement according to claim 7, characterized in that: The vector retrieval is used to recall Top-N semantically related entities, descriptions, and FAQ fragments in the Milvus vector library based on the query vector or slot key phrase in the structured understanding results, using cosine similarity, to support semantic near-synonym recall and synonym variant recognition, and to adapt to multilingual fuzzy questions. The graph retrieval is used to perform path lookup using Cypher queries based on the explicit slots in the structured understanding results, in order to perform relational reasoning and attribute filtering; The keyword full-text search is used to retrieve relevant documents or entities based on the original text or extended phrases in the structured understanding results, using word segmentation, indexing, and inverted indexing.

9. A tourism question-answering system based on knowledge graph enhancement according to claim 8, characterized in that, The answer generation module includes: The prompt template unit is used to design dynamic prompt templates that include a set of query-related knowledge fragments, user intent, domain safety, domain culture, and precautions; The model building unit is used to select ChatGLM3-6B as the base model, and to perform INT8 quantization, LoRA fine-tuning and vLLM framework inference optimization on the base model to obtain an inference question answering model to output the initial answer to the user query. The answer generation strategy unit is used to control answer generation through temperature sampling parameters, length control strategy, and security filtering strategy; wherein, the temperature sampling parameter is set to 0.7; The answer post-processing unit is used to compare the content of the initial answer with the tourism domain knowledge graph for fact verification, and then optimize the format and segment the initial answer, and add source introduction and annotation to obtain the tourism answer.

10. A tourism question-answering system based on knowledge graph enhancement according to claim 9, characterized in that, The personalized recommendation module includes: The user profiling unit is used to collect basic user information, travel preferences, and behavioral characteristics to obtain multi-dimensional data in order to build user profiles. The personalized recommendation unit is used to perform collaborative filtering based on the behavior of similar users, to recommend content based on user interests and content features, and to recommend knowledge based on the association reasoning of the knowledge graph in the tourism field, so as to obtain the recommendation output. The multi-turn dialogue management unit is used for dialogue status tracking, dialogue history maintenance, slot inheritance, and topic switching, enabling the transmission of key information and topic identification and conversion across turns. The proactive interaction unit is used to proactively ask questions when a user's query for information is insufficient, in order to improve the user profile.

Citation Information

Patent Citations

  • Method for automatically acquiring multi-source heterogeneous data knowledge

    CN110489395A

  • Intelligent question and answer method and system based on tourism knowledge graph

    CN113535917A

  • Retrieval enhancement generation system and method based on knowledge graph

    CN117973540A

  • Tourism information consultation method and system based on big data analysis

    CN118484592A

  • Method and system for enhancing RAG questions and answers through mixed retrieval method

    CN118627625A

Cited By

  • Scenic area intelligent question answering and recommendation method based on integration of knowledge graph and large model

    CN121880532A