Tourism knowledge graph multi-dimensional construction method and system of deep learning hybrid model

Through the deep learning hybrid model combined with LLMs and MLP_Boost, the problems of inefficiency and insufficient accuracy in the construction of tourism knowledge graphs are solved, efficient and accurate knowledge graph construction is achieved, and personalized application of intelligent tourism services is supported.

CN120541234AActive Publication Date: 2025-08-26SICHUAN UNIV JINCHENG INST

Patent Information

Application Number
CN202510493697.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-26
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing tourism knowledge graph construction methods rely on expert knowledge and manual operations, resulting in inefficiency and inability to cope with the diversity and rapid changes in tourism data. The accuracy and interpretability of LLMs in generating content are low, and a large amount of manual intervention is required to correct errors.

Method used

Using a deep learning hybrid model, combining LLMs and MLP_Boost, through data extraction, knowledge extraction, knowledge fusion, knowledge evaluation and knowledge storage, Few-Shot Prompting and large-model fine-tuning technology, an accurate knowledge pre-extracted set is generated, and the algorithm and entity self-verification are improved through MLP_Boost deep learning to improve the accuracy and efficiency of the knowledge graph.

Benefits of technology

It significantly reduces the cost of manual annotation, improves the efficiency and accuracy of the construction of tourism knowledge graphs, enhances the accessibility and comprehensibility of data, provides personalized support for intelligent tourism services, and reduces the illusion problem in the LLMs information extraction process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541234A_ABST
    Figure CN120541234A_ABST
Patent Text Reader

Abstract

The invention provides a tourism knowledge graph multi-dimensional construction method and system of a deep learning hybrid model, and belongs to the technical field of knowledge graph construction, and the method comprises the steps: S1, data extraction; s2, knowledge extraction; s3, knowledge fusion; s4, knowledge evaluation; and S5, knowledge storage. According to the method, the LLMs and the deep learning technology MLPBoost are combined, so that many problems in the current knowledge graph construction process can be effectively solved, and more accurate and personalized support is provided for intelligent tourism service.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tourism knowledge graph construction, and in particular to a method and system for constructing a multi-dimensional tourism knowledge graph based on a deep learning hybrid model. Background Art

[0002] With the development of big data and artificial intelligence technologies, knowledge graphs have been widely used, particularly in the tourism sector, which has gradually attracted attention from both academia and industry. Despite this, the application of knowledge graphs in the tourism industry still faces several challenges, particularly regarding accuracy and efficiency when processing large-scale, multi-source, heterogeneous data. Furthermore, how to effectively embed entities and relationships unique to the tourism sector and build scalable knowledge graphs remains a critical issue that needs to be addressed by the academic community.

[0003] At present, many studies have proposed different knowledge graph construction methods to improve their construction efficiency and expressiveness. For example, Guoqiang Liu et al. integrated multi-source data and expert knowledge through ontology construction, knowledge fusion and other methods, and established a knowledge-driven neural network model (KPNFE). In addition, Tangzhao Wei et al. proposed a template-based semi-automatic knowledge graph construction method, which extracts triple data from multiple data sources by filling in templates, thereby improving the degree of automation of knowledge graph construction. However, although these methods have improved the efficiency of knowledge graph construction to a certain extent, they still face problems such as high training cost, large consumption of computing resources, and how to improve computing efficiency while ensuring accuracy and recall rate.

[0004] With the development of generative AI technology, the application of large language models (LLMs) in the construction of knowledge graphs has gradually become a research hotspot. LLMs have powerful natural language processing capabilities and can extract rich semantic information from large amounts of text. Therefore, they are widely used in knowledge extraction and knowledge graph construction. For example, Yichong Zhang et al. proposed a few-shot learning method based on LLMs to construct and optimize knowledge graphs in the field of traditional Chinese medicine. Shirui Pan et al. proposed a circuit diagram model that integrates LLMs with knowledge graphs, in this way realizing the complementary advantages of large models and knowledge graphs. Despite this, the black box properties of LLMs and the problem of hallucinations lead to their low accuracy and interpretability in generating content, requiring a lot of manual intervention to correct errors and improve data quality.

[0005] To address these issues, the construction of a tourism knowledge graph requires more precise domain knowledge extraction methods and more efficient knowledge graph generation mechanisms. Existing tourism knowledge graph construction methods mostly rely on expert knowledge and manual operations. This reliance can lead to inefficiency and accumulated errors, and is unable to cope with the diversity and rapid changes in tourism data.

[0006] In summary, knowledge graph technology has great application potential in the tourism field, but it also faces a series of challenges such as how to improve construction efficiency, reduce manual intervention, and improve accuracy and interpretability. Summary of the Invention

[0007] The present invention provides a multi-dimensional construction method and system for tourism knowledge graphs based on a deep learning hybrid model. By combining LLMs with deep learning technology (MLP_Boost), it can effectively solve many problems in the current knowledge graph construction process and provide more accurate and personalized support for smart tourism services.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] This specification discloses a method for constructing a multi-dimensional tourism knowledge graph using a deep learning hybrid model, including:

[0010] S1. Data extraction: Collect tourism information, including structured data, semi-structured data, and unstructured data;

[0011] S2. Knowledge Extraction: Based on a large language model and MLP-Boost deep learning model, this approach extracts valuable information from structured, semi-structured, and unstructured data and converts it into a structured form, including entities, attributes, and relationships between entities.

[0012] S3. Knowledge fusion: Based on the results of knowledge extraction in S2, entity fusion is performed to eliminate redundant information and correct erroneous data that may have been generated during the data extraction stage.

[0013] S4. Knowledge evaluation: Based on the results of knowledge fusion in S3, entity evaluation and relationship evaluation are performed using precision, recall, and F1 score.

[0014] S5. Knowledge storage: Based on the results of S4 knowledge evaluation, the structured knowledge is saved in the form of a graph to obtain a tourism knowledge graph.

[0015] In this specification, in S1, unstructured data is systematically organized and converted into structured data, and entity types and relationship types are preliminarily explored; in the tourism text data collection stage, text data is collected according to the preliminarily explored entity types and relationship types, and the collected text needs to cover every entity type and relationship type;

[0016] Use data annotation software to perform preliminary data annotation on the text data, and define the text data annotation in Json format. Each element is an object, and each object has text and list attributes to store text data and triple relationships.

[0017] In this specification, S2 includes:

[0018] S2.1. Knowledge Pre-extraction: Use the matched Few-Shot Prompting method and fine-tuning technology for weighted fusion to generate a knowledge pre-extraction set;

[0019] S2.2. Named Entity Recognition: Based on a pre-extracted knowledge set, it identifies entities with specific meanings from unstructured text and classifies them into predefined categories.

[0020] S2.3. Relationship Extraction: Based on the knowledge pre-extraction set, the semantic relationship between entities is extracted by identifying the relationship between entity pairs in the text;

[0021] S2.4. Attribute extraction: Based on a knowledge pre-extraction set, identify and extract attribute information related to a specific entity from unstructured or semi-structured text data.

[0022] In this specification, S2.1 includes:

[0023] S2.1.1. Prompt word construction: Use matching Few-Shot Prompting to obtain accurate matching examples for different entity types and different relationship types;

[0024] S2.1.2. Large Model Fine-tuning: Further train the large language model based on accurate examples of classification and matching of different entity types and different relationship types.

[0025] S2.1.3. Weight-based fusion: Combine the matched Few-Shot Prompting with the fine-tuning of the large model with limited annotated data. Under the premise of taking into account the precision and recall rate of the model, the knowledge extraction results of the two are weightedly fused to obtain the knowledge pre-extraction set.

[0026] In this specification, S2.3 includes:

[0027] S2.3.1. Multivariate Relationship Classification Based on the MLP_Boost Deep Learning Boosting Algorithm: A multivariate relationship classification model is trained using manually annotated text corpus using the MLP_Boost deep learning boosting algorithm, thereby improving the accuracy and standardization of knowledge graph triples.

[0028] S2.3.2. Relationship Verification: Use matching Few-Shot Prompting technology to provide different input and output examples for different types of tourism entity relationships, and verify them through questions and limited responses.

[0029] In this specification, in S2.4, after extracting the attribute type and attribute value, the attribute type and attribute value information are mapped to the corresponding entity by associating with the entity keyword.

[0030] In this specification, in S3, entity fusion is to merge records that appear in different forms in different tourism scene corpora but actually refer to the same entity.

[0031] In this specification, in S4, the evaluation formula is:

[0032] Where P is the precision rate, R is the recall rate, TP means that the model predicts the category and the actual category are the same, FP means that the model predicts other categories as the category to be identified; FN means that the model predicts the category to be identified as other categories.

[0033] In this specification, in S5, knowledge storage is performed using the graph-based database Neo4j.

[0034] This specification also discloses a system for constructing a multi-dimensional tourism knowledge graph using a deep learning hybrid model, which is used to implement any of the above-mentioned methods for constructing a multi-dimensional tourism knowledge graph using a deep learning hybrid model. The system comprises:

[0035] Data extraction module, used to collect tourism information, including structured data, semi-structured data and unstructured data;

[0036] The knowledge extraction module is used to extract valuable information from structured, semi-structured, and unstructured data based on a large language model and an MLP_Boost deep learning model, and convert it into a structured form, including entities, attributes, and relationships between entities;

[0037] The knowledge fusion module is used to perform entity fusion based on the results of knowledge extraction to eliminate redundant information and correct erroneous data that may be generated in the data extraction stage;

[0038] The knowledge evaluation module is used to perform entity evaluation and relationship evaluation based on the results of knowledge fusion through precision, recall and F1 score;

[0039] The knowledge storage module is used to save structured knowledge in the form of a graph based on the results of knowledge evaluation to obtain a tourism knowledge graph.

[0040] In summary, the present invention has at least the following beneficial effects:

[0041] This invention is suitable for constructing efficient and accurate tourism knowledge graphs within the smart tourism industry, improving data accessibility and understandability. This construction method can be used in decision-making support systems such as intelligent question-and-answer (Q&A), smart customer service, and recommendation systems, providing tourists with a richer, more personalized travel experience. This will help further promote tourism development and economic recovery, opening up new possibilities for innovation and optimization in the tourism industry.

[0042] The weighted fusion method of few-shot learning and fine-tuning technology of LLMs is used to generate a pre-extracted set of knowledge graph triples, which significantly reduces the cost of manual annotation and lays a solid foundation for building an accurate tourism knowledge graph.

[0043] The MLP_Boost deep learning improvement algorithm is proposed to perform multi-classification on the relationships of the pre-extracted parts, providing guarantees for the standardization and relationship verification of the knowledge graph network.

[0044] The few-shot learning technology of LLMs classification matching is used to self-verify the named entity recognition and relationship extraction results, which effectively reduces the hallucination problem in the LLMs information extraction process and improves the accuracy and credibility of LLMs named entity recognition and relationship extraction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0046] Figure 1 This is a schematic diagram of the construction of the tourism knowledge graph based on model fusion involved in the present invention.

[0047] Figure 2 Schematic diagram of tourism knowledge graph data collection and data preprocessing process.

[0048] Figure 3 Flowchart of the method for extracting knowledge graph triples from few-shot prompts based on LLMs matching.

[0049] Figure 4 Schematic diagram of the FSP_FT_Weighted_Fusion algorithm.

[0050] Figure 5Schematic diagram of the Few-Shot Prompting tourism entity type self-verification method based on matching.

[0051] Figure 6 This is a flowchart of the learning algorithm based on MLP_Boost deep learning.

[0052] Figure 7 Schematic diagram of the Few-Shot Prompting tourism entity relationship self-verification method based on matching.

[0053] Figure 8 Schematic diagram of tourism entity attribute triple extraction.

[0054] Figure 9 Schematic diagram of knowledge fusion based on synonyms and pattern matching.

[0055] Figure 10-1 Schematic diagram of the confusion matrix of the CART algorithm for a single learner.

[0056] Figure 10-2 Schematic diagram of the confusion matrix of the MLP algorithm for a single learner.

[0057] Figure 10-3 Schematic diagram of the confusion matrix of the GaussianNB algorithm for a single learner.

[0058] Figure 10-4 This is a diagram of the confusion matrix of the MLP_Boost deep learning algorithm. DETAILED DESCRIPTION

[0059] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0060] like Figure 1 As shown, this embodiment provides a multi-dimensional construction method of a tourism knowledge graph based on a deep learning hybrid model, including five steps: data extraction, knowledge extraction, knowledge fusion, knowledge evaluation, and knowledge storage.

[0061] S1. Data extraction involves data types including structured, semi-structured, and unstructured data. By applying web crawler technology, we achieve automated data collection of tourism information. Subsequently, we systematically organize and transform this unstructured data into structured data, and initially explore entity and relationship types.

[0062] S2. Knowledge extraction is a key step in the knowledge graph construction process. It extracts valuable information from unstructured, semi-structured, or structured data sources and converts it into a structured form, such as entities, attributes, and relationships between entities, on which an ontological knowledge representation is formed. This step of the invention includes four parts, namely:

[0063] S2.1. Knowledge Pre-Extraction: To obtain a set of knowledge graph triples that balances breadth and precision, this paper introduces a knowledge pre-extraction process based on LLMs. This process uses a matched Few-Shot Prompting method and fine-tuning techniques for weighted fusion to generate a pre-extracted knowledge set. This knowledge pre-extraction process includes three main steps.

[0064] S2.1.1. Prompt word construction. Few-shot prompting is a technique used in pre-trained language models in the NLP field. The core idea is to guide the model to complete a specific task by providing a small number of examples without the need for large-scale fine-tuning or training of the model.

[0065] S2.1.2. Large Model Fine-tuning. Fine-tuning a large model involves further training a pre-trained large language model using a specific dataset to adapt the model to a specific task or domain. The core goal of fine-tuning is to achieve a precise match between refined knowledge infusion and the instruction system. Through fine-tuning, the model can better adapt to the needs and characteristics of a specific domain, learning the knowledge and language patterns of that domain, thereby achieving better performance on specific tasks.

[0066] S2.1.3. Weight-based fusion. This paper combines matching Few-Shot Prompting with large model fine-tuning with limited annotated data, and proposes the FSP_FT_Weighted_Fusion algorithm. Under the premise of taking into account the precision and recall rate of the model, the knowledge extraction results of the two are weighted and fused to achieve the knowledge pre-extraction link.

[0067] S2.2. Named Entity Recognition; Named Entity Recognition (NER) is the basis of information extraction in the field of natural language processing (NLP) and a prerequisite for attribute and relationship extraction. Its goal is to identify entities with specific meanings from unstructured text and classify them into predefined categories.

[0068] S2.3. Relationship extraction: Relationship extraction is a key step in information extraction. By identifying the relationships between entity pairs in the text, the semantic relationships between entities are extracted, thereby providing support for the construction and application of knowledge graphs. This step of the invention is divided into two parts.

[0069] S2.3.1. Multivariate Relationship Classification Based on the MLP_Boost Deep Learning Boosting Algorithm: Due to issues such as non-standard relationship types, incorrect relationship logic, and diverse relationship expressions in the relationships extracted during the large model knowledge pre-extraction phase, we primarily use the MLP_Boost deep learning boosting algorithm to train a multivariate relationship classification model using manually annotated text corpus, thereby improving the accuracy and standardization of knowledge graph triples.

[0070] S2.3.2. Relationship Verification: After the triples are corrected and predicted by the multivariate relationship classification model, the matching Few-Shot Prompting technology is mainly used to provide different input and output examples for different types of tourism entity relationships. Verification is carried out through methods such as asking questions and limiting response forms to reduce the error information caused by the large model illusion phenomenon.

[0071] S2.4. Attribute extraction; The purpose of attribute extraction is to identify and extract attribute information related to a specific entity from unstructured or semi-structured text data.

[0072] S3. Knowledge Fusion. The entity fusion stage aims to eliminate redundant information and correct erroneous data that may have been generated during the data extraction stage.

[0073] S4. Knowledge Assessment. The main tasks of the knowledge assessment phase are to ensure the accuracy, completeness, and reliability of the graph, discover new knowledge, optimize construction methods, and support application development. The knowledge assessment in this paper focuses on entity recognition and relationship recognition, and is evaluated through precision, recall, and F1 score.

[0074] S5. Knowledge Storage. The primary task of knowledge storage is to store structured knowledge in the form of a graph for easy query and analysis. Its functions include providing efficient data retrieval, supporting large-scale data management, and facilitating knowledge integration and updating. Knowledge storage enables the knowledge graph to quickly respond to query requests, enabling effective knowledge organization and sharing, and providing a foundation for subsequent knowledge reasoning and application.

[0075] In some embodiments, S1. During data extraction, data collection and data preprocessing are performed: the data types involved in the present invention mainly include structured data, semi-structured data and unstructured data, and the automatic data collection of tourism information is realized by applying web crawler technology. Subsequently, these unstructured data are systematically sorted and converted to eventually form structured data, and the entity types and relationship types are preliminarily explored. In the tourism text data collection stage, it is necessary to collect text data according to the entity types and relationship types preliminarily explored. The collected text needs to cover every entity type and relationship type, and the number of high-quality texts after screening must be at least 500. In order to enrich the collection needs, the data sources are diversified, such as various tourism websites, tourism software, science encyclopedias and other websites and software, local culture websites and software, etc.

[0076] The task of the data annotation stage is to convert the raw data into a structured format that the model algorithm can understand and learn. In this stage, data annotation software is used to perform preliminary data annotation on the text data. At the same time, to meet the model input data requirements, the text data annotation is defined in Json format. Each element is an object, and each object has 'text' and 'list' attributes to store text data and triple relationships. The process of the data collection and data preprocessing stage is as follows: Figure 2 shown.

[0077] S2. Knowledge extraction; Knowledge extraction is a key step in the process of building a knowledge graph, that is, extracting valuable information from unstructured, semi-structured or structured data sources and converting it into a structured form, such as entities, attributes and the relationships between entities, on this basis forming an ontological knowledge expression. The present invention mainly uses LLMs technology for knowledge extraction. First, based on the weighted fusion method of matched Few-Shot Prompting and fine-tuning technology, a knowledge pre-extraction set is generated. Secondly, in the named entity recognition part, a large model is used to self-verify the entities and their types. Next, in the relationship extraction link, the MLP_Boost deep learning boosting algorithm is first used to perform multi-class classification on the relationships in the pre-extracted part, and then the large model is used to self-verify the relationship type. On the basis of the network data collection and cleaning in the attribute extraction part, a structured triple structure is formed through rules.

[0078] S2.1. Knowledge pre-extraction: To obtain knowledge graph triples that take into account both breadth and precision<Entity,Relation,Entity> The present invention introduces a knowledge pre-extraction link based on LLMs, and uses the matching Few-ShotPrompting method and fine-tuning technology to perform weighted fusion to generate a knowledge pre-extraction set.

[0079] S2.1.1. Prompt word construction; Few-Shot Prompting is a technology used in pre-trained language models in the NLP field. The core idea is to guide the model to complete specific tasks by providing a small number of examples without the need for large-scale fine-tuning or training of the model. In Few-Shot Prompting, there are two key components, Prompt template and examples. The Prompt template component is usually composed of a structured string and is used to insert variables to guide the model on how to understand and process the input data. The Examples component is presented to the model in combination with the actual characteristics of the data to help the model understand the task. Each example usually includes input and expected output. The present invention adopts matching Few-Shot Prompting, that is, different entity types and different relationship types are classified and matched with precise examples, rather than general examples. The information extraction method (few sample prompt extraction knowledge graph triple method based on LLMs matching) is as follows (the corresponding process is shown in Figure 3 ):

[0080] Input: Tourism text corpus collection T{item1,item2,…,item N};

[0081] 1. Clean the text corpus set T and remove symbols such as []|\r|\n|*|#|0-9|AZ|az|; 2. Construct the entity type set entity_type={e i |i=1,2,3,…,n}, where e i is the entity type, n is the number of entities; 3. Construct the relationship type set relation_type={r i |i=1,2,3,…,m}, where r i is the relationship between entities, m is the number of relationships; 4. Construct an example text set example_set = {text i |i=1,2,3,…,p}, where text i Input example text for the i-th input, p is the number of example texts; 5. Construct example text triple set example_triple = {triple i |i=1,2,3,…,p}, where p is the number of triplet example extraction results, triple i For the i-th input sample text text i The implied triple relationship, the Json logical format is:

[0082]

[0083] Where q is the sample text texti The number of entity types implied, k is text i the number of relations implied;

[0084] 6. Build a prompt word template: template{entity_type,relation_type,example_set,example_triple,item i};

[0085] 7.for i init to Nmax do; 8.According to the text item to be extracted i Generate prompt word template; 9. Call pre-trained knowledge extraction LLMs and parse the knowledge pre-extraction results; 10. end; 11. Output: a set of triples in Json format; Json = {triple i} i∈[1,N] ={nodes i ,edges i} i∈[1,N] .

[0086] In the process of pre-extracting triples of knowledge graph using matching Few-Shot Prompting, the first step is to clean the corpus data and remove characters such as *, \r\n, numbers and letters. Define the pre-extracted entity type and relationship type, where the entity type is entity_type={e i |i=1,2,3,...,n},e i is a specific entity type, such as scenic spots, hotels, festivals, traditional culture, etc., n is the number of entities; relation type relation_type = {r i |i=1,2,3,...,m},r i is the relationship between entities such as located, adjacent, popular, etc., and m is the number of relationships.

[0087] Next, define the example component, which contains the input text example example_set and the output result example example_triple, example_set = {text i |i=1,2,3,...,p},text i is the i-th input example text, p is the number of small sample texts; example_triple = {triple i |i=1,2,3,...,p},triple i For the i-th input sample text text i The implied triple relationship, which consists of a node group and a relationship edge group, is triple i{nodes i ,deges i}, node group nodes i { <e ij ,type ij >| <e ij :entity j ,type ij :e l >}, containing the entity type e of the node l and e l ∈entity_type, and entity j is a specific entity element, q is a node group nodes i The number of entities in the collection. i { <e_left iω :e ω1 ,R iω :r ω ,e_right iω :e ω2 >}, where e ω1 ,e ω2 Belong to the text extraction entity node group nodes i , r ω The relationship between the two is corresponding and r ω ∈relation_type, k is the relationship edge group edges i The number of implied relations.

[0088] The definition of the Travel Prompt template component includes elements such as roles, goals, steps, rules, input examples, and output responses. To guide the model in understanding and processing input data, structured strings are defined for inserting variables, such as structured parameter entity types, relationship types, few-shot input examples, and few-shot output examples, as shown in Table 1. To ensure the quality and diversity of examples, seven entity types and ten entity-time relationships were selected during the design of examples for entity types and relationships, covering different scenarios, data sources, and business types. Furthermore, considering the model's generalization and long-term stability, an incremental corpus approach was used to design example prompt word templates.

[0089] Table 1. Knowledge graph triple templates extracted from few-shot hints based on LLMs matching

[0090]

[0091]

[0092] S2.1.2. Fine-tuning of large models; Fine-tuning of large models refers to further training based on a pre-trained large language model using a specific data set to adapt the model to specific tasks or fields. The core goal of fine-tuning is to achieve precise matching of refined knowledge infusion and instruction system. Through fine-tuning, the model can better adapt to the needs and characteristics of a specific field, learn the knowledge and language patterns in the field, and thus achieve better performance on specific tasks. The present invention adopts a large model fine-tuning method to pre-extract the tourism knowledge graph, adopts the GLM4-9B-Chat model, and performs supervised fine-tuning (SFT) through the LLaMA Factory framework to enhance the model's knowledge extraction ability in tourism graphs. The supervised training corpus is a semi-structured triple data set that has been manually annotated. In order to extract pre-extracted triple data that meets the tourism business, the present invention proposes an extraction method based on the Chain of Thought (COT). The fine-tuned large model, guided by COT cues, completes extraction in a step-by-step process: first identifying entities in the text, then analyzing the semantic relationships between entities, and finally combining the identified entities and relationships to form complete knowledge graph triples. This hierarchical extraction strategy improves the accuracy and completeness of knowledge graph construction. Table 2 provides examples of fine-tuning instructions during training.

[0093] Table 2. Examples of LLMs fine-tuning training instructions

[0094]

[0095] S2.1.3. Weight-based Fusion: In the knowledge extraction process, the advantage of Few-Shot Prompting is that it uses a small amount of labeled data to help the model understand the task, improving model performance and reducing data preparation. Furthermore, the model can quickly adapt to new tasks without having to start training from scratch, reducing training costs while maintaining a certain extraction breadth and accuracy. However, Few-Shot Prompting also has some limitations. For example, it relies on the quality and diversity of the provided examples. If the input examples are of low quality, the extraction accuracy will not be met. Furthermore, the LLM's inability to generalize examples can lead to overfitting. Using large model fine-tuning (Fine-Tuning) can reduce the generation of inaccurate or fabricated information by the model during knowledge extraction, improve the consistency and reliability of the output, and reduce the problem of hallucinations. However, high-quality fine-tuning results rely on high-quality and large amounts of labeled data, which can be difficult to obtain or costly. Furthermore, higher accuracy also results in lower model recall. Therefore, this paper combines the matching Few-Shot Prompting with the fine-tuning of the large model with limited labeled data, and proposes the FSP_FT_Weighted_Fusion algorithm. Under the premise of taking into account the precision and recall rate of the model, the knowledge extraction results of the two are weighted and fused to realize the knowledge pre-extraction link. The fusion principle is as follows: Figure 4 shown.

[0096] The set of entities and relations extracted by Few-Shot Prompting is defined as The model voting weight corresponding to the pre-extraction error ε1 is ρ1=1 / ε1 2 Fine-Tuning of large models extracts entity and relationship sets defined as The pre-extraction error ε2, the corresponding model voting weight is ρ2 = 1 / ε2 2 The result of the fusion of the two is That is, the fused extraction result Ω is the union of Ω1 and Ω2. The extraction results of each part are shown in the following formula. The function f(Ω) identifies the fused entity set nodes and relationship set edges.

[0097]

[0098] exist Figure 4 In the left Ω1-Ω2 space, the weight of the Few-Shot Prompting model is ρ1, and the weight of the Fine-Tuning model is ρ2=0, so the pre-extraction knowledge fusion is mainly based on FSP; Figure 4In the space of Ω2-Ω1 on the right, the weight of the fine-tuning model is ρ2, and the weight of the Few-Shot Prompting model is ρ1=0, so the pre-extraction knowledge fusion is mainly based on FT; Figure 4 In the middle Ω1∩Ω2 space, the weights of the Few-Shot Prompting model and the Fine-Tuning model are ρ1, ρ2. At this time, for the cases where the same entity has different entity types, or the same relationship entity nodes have different relationship types, the pre-extracted knowledge fusion adopts a weighted voting method with a weight of max{ρ1, ρ2}=max{1 / ε1 2 ,1 / ε2 2}.

[0099] S2.2. Named Entity Recognition; Named Entity Recognition (NER) is the basis of information extraction in the field of Natural Language Processing (NLP) and is also a prerequisite for attribute and relationship extraction. Its goal is to identify entities with specific meanings from unstructured texts and classify them into predefined categories. The present invention adopts a fusion method based on matching Few-Shot Prompting and Fine-Tuning to extract defined entity types from tourism text data. In order to further optimize the illusion generated by the large model LLMS in the process of knowledge pre-extraction, that is, the phenomenon that the content generated by the model is inconsistent with the input provided by facts or experiments, an entity self-verification link based on the large model is introduced, and the matching Few-Shot Prompting method is used to provide different input and output examples for different types of tourism entities, and verification is carried out through methods such as asking questions and limiting response forms. The introduction of entity self-verification can, to a certain extent, alleviate the large model hallucination phenomenon that may cause the model to output inaccurate or misleading information, thereby improving the credibility and practicality of the model. The Few-Shot Prompting tourism entity type self-verification method based on matching is as follows (see the verification flow chart for details). Figure 5 ):

[0100] Input: Tourism text corpus T{item i} i∈[1,N]

[0101] edges i { <e_left iω ,R iω ,e_right iω >| <e_left iω :e ω1 ,R iω :r ω ,e_right iω:e ω2 >

[0102] EntityNodes{nodes i}={nodes i |<Θ ij :θ j ,type ij :e j >,e j ∈entity_type,j∈[1,len(nodes i )]} i∈[1,N] ;

[0103] 1. Construct verification type example template set Verify_set = {verify i |verify i =Ψ(e i ),i=1,2,3,…,n}, where e i is the entity type, It is e i The text description function of , n is the number of entities;

[0104] 2. Construct output response is the entity type, ∈entity_type,Res i It is e i 3. Build the prompt word template template{Verify_set,Response_set};4. for i init to Nmax do;5. for j init to len(nodes i )do;6.Judge nodes ij The actual entity type l; 7. Construct a matching template template based on the entity type l {verify l ,Res l ,item i};8.Call pre-trained verification LLMs and parse the verification results;9.End;10.End;11.Output:Model verification set in Json format

[0105] Before entity self-verification, prepare the input data tourism text corpus T{item i} i∈[1,N] And the extracted entity type set Nodes{nodes i}, each corpus text item i Corresponding to a set of entity type node sets nodesi , each group of nodes i By several entity nodes <Θ ij :θ j ,type ij :e j >composition, where θ j is the entity name, e j The entity type of tourism is shown in Table 3. First, a verification type example template set Verify_set is constructed. This set is used to match learning examples of different entity types, namely verify i =Ψ(e i ), Ψ is the entity type e i Text description function; build LLMS output response Response_set collection, which is used to respond to verification questions of different types of entities ),Right now It is e i The text question function, the number of questions is the number of entity types n; Next, use the verification sample set Verify_set and response set Response_set in the first two steps to build the prompt word template. When calling the large model verification phase, for each text corpus item i Each entity node pairs nodes ij Determine the specific value of the entity type I, and construct a matching template template{verify l ,Res l ,item i}, pass the template to the training and verification LLMs, parse the verification results. Finally, store the verification results of the large model as a model verification set in Json format {Version i} i∈[1,N] , the set consists of entity nodes ij and verification results ij {Yes,No}, the actual number of verification results should be

[0106] Table 3. Tourism entity types

[0107]

[0108] S2.3. Relationship extraction; Relationship extraction is a key link in information extraction. By identifying the relationship between entity pairs in the text, the semantic relationship between entities is extracted, thereby providing support for the construction and application of knowledge graphs. The present invention uses the weighted fusion of large-model Few-Shot Prompting and fine-tuning technology to pre-extract entity relationships in tourism corpus, and performs self-verification of LLMs on the entities in the pre-extracted triple relationships. Therefore, the pre-extracted triple relationships are verified entity relationships. In order to optimize the illusions produced by large-model LLMS in the relationship pre-extraction process, the present invention first trains the MLP_Boost deep learning boosting algorithm based on artificially labeled corpus to correct the multivariate classification of the pre-extracted partial relationships. Then, the Few-Shot Prompting technology is used to provide different input and output examples for different types of tourism entity relationships, and relationship verification is performed through methods such as asking questions and limiting response forms.

[0109] S2.3.1. Multi-relation classification based on MLP_Boost deep learning algorithm: Due to the problems of non-standard relationship types, incorrect relationship logic, and diversified relationship expressions in the relationships extracted in the large model knowledge pre-extraction stage, the MLP_Boost deep learning algorithm is mainly used to train the multi-relation classification model using manually annotated text corpus, thereby improving the knowledge graph triples.<Entity,Relation,Entity> The accuracy and standardization of the algorithm. The algorithm implementation process is as follows Figure 6 shown.

[0110] The first step is data balancing, that is, balancing the machine learning multi-classification dataset. In the tourism entity relationship classification, this paper mainly studies 10 types of entity relationships, such as location, adjacency, popularity, and provision, as shown in Table 4. In the learning sample balancing process, a combination of undersampling and oversampling is used to prevent the lack of representativeness of learning samples due to class imbalance. The multi-classification dataset includes entities, entity types, relationships, object entities, and object entity types. <e i1 ,t i1 ,r i ,e i2 ,t i2 > and other texts. In the data preprocessing phase, the entity type features and relationship type targets are mainly subjected to one-hot encoding and label encoding. In the model training phase, the supervised classification training is first performed using individual learners such as CART, GaussianNB, and MLP. Next, based on the training results, the MLP_Boost deep learning method is constructed by the boost combination method to learn the dataset S. That is, the weak classifiers CART and GaussianNB are used for training, and the recognition results of the weak classifiers are used as the model weights of the samples {ρi ,i=1,2,…,n}, are fused into the dataset S, and finally a multi-layer perceptron (MLP) is used for correction prediction. The improved combination model is evaluated using accuracy, precision, recall, and F1 score.

[0111] Table 4. Tourism entity relationship types

[0112] Subject entity Object entity relation Subject entity Object entity relation Attractions City-District-County-Town-Village lie in City-District-County-Town-Village traditional culture inheritance Attractions hotel Nearby City-District-County-Town-Village festival Will celebrate Attractions transportation places Nearby City-District-County-Town-Village gourmet specialties Include Attractions gourmet specialties You can taste City-District-County-Town-Village transportation places contain Attractions traditional culture Popularity hotel gourmet specialties supply Attractions festival Will celebrate hotel City-District-County-Town-Village lie in Attractions Attractions Adjacency hotel transportation places Nearby

[0113] S2.3.2. Relationship Verification; Triples<Entity,Relation,Entity> After the multivariate relationship classification model is corrected and predicted, the matching Few-Shot Prompting technology is mainly used to provide different input and output examples for different types of tourism entity relationships. Verification is carried out through methods such as asking questions and limiting response forms to reduce the error information caused by the large model hallucination phenomenon and improve the accuracy of the triple relationship. The self-verification (based on the matching Few-Shot Prompting tourism entity relationship self-verification method) process is as follows (the flowchart is detailed in Figure 7 ):

[0114] Input: Tourism text corpus T{item i} i∈[1,N] , the node-edge dataset after relationship classification correction {triple i} i∈[1,N] ={nodes i ,edges i} i∈[1,N] ;

[0115] 1. Construct entity relationship verification example template set Ver_set = {V i |V i =φ(e i1 ,r i ,e i2 ),i=1,2,3,…,K}, where e i1 For the entity subject, e i2 is the entity object, e i2 ∈Nodes{nodes i}, φ i yes <e i1 ,r i ,e i2 >text description function, K is the number of combinations of different entity subject-object relationships; 2. Construct the output response Res_set = {Res i |Res i =τ(e i1 ,r i ,ei2 ),i=1,2,3,…,K},τ i yes <e i1 ,r i ,e i2 >Text description letter; 3. Construct prompt word template template{Ver_set,Res_set}; 4. for i init toNmax do; 5. for j init to len(edges i )do; 6. Determine the edges of each table ij 7. Construct a matching template template{V according to the type of L L ,Res L ,item L};8.Call the pre-trained verification LLMs and parse the verification results;9.End;10.End;11.Output:Model verification set in JSON format:

[0116]

[0117] This stage mainly uses the matching Few-Shot Prompting to i} i∈[1,N] Concentrate and verify the node edge dataset {nodes i ,edges i} i∈[1,N] First, construct a set of example templates for verifying entity relationships {V i =φ(e i1 ,r i ,e i2 )}, which is used to match learning examples of different entity relationship types, e i1 For the entity subject, e i2 For the entity object, r i is the subject-guest relationship, φ is the text description function of the triple; next, we construct the LLMS output response Res_set set, which is used to respond to the verification questions τ(e i1 ,r i ,e i2 ), that is, τ is the text question function of the triple, and the number of questions is K, which is the number of combinations of different entity subject-object relationships; after completing the construction of the prompt word template, similar to entity verification, each text corpus item can be i The corresponding edges ij Determine the relationship type L under the subject and object entity constraints, and construct a matching template template{VL ,Res L ,item L}, pass the template with the example to the training and verification LLMs for verification. Finally, the verification results of each edge are stored in the form of:

[0118] And respond with YES and No.

[0119] S2.4. Attribute Extraction: The goal of attribute extraction is to identify and extract attribute information related to specific entities from unstructured or semi-structured text data. The primary data source for attribute extraction in this invention is semi-structured and structured data from platforms such as popular science encyclopedias and tourism-related websites and software. These platforms provide a wealth of information about tourism entities and corresponding attribute descriptions, but this content is typically not presented in a structured form. Therefore, attribute extraction technology is required to structure this information.

[0120] like Figure 8 As shown in the figure, the blue dashed box on the left shows the entity attribute types and their corresponding attribute values ​​extracted from unstructured or semi-structured text. The attribute extraction results are stored in the format ["attribute type", "attribute value"]. After extracting the attribute types and values, they are mapped to specific entities by associating them with entity keywords, such as the Potala Palace, a landmark entity, in the green dashed box. The final attribute extraction results are stored as triples ["entity", "attribute type", "attribute value"]. This triple format can be directly used to construct relationship networks, facilitating further data mining and information retrieval.

[0121] S3. Knowledge fusion; The entity fusion process aims to eliminate redundant information and correct erroneous data that may be generated during the data extraction stage. The entity fusion of the present invention mainly merges records that appear in different forms in different tourism scene corpora but actually refer to the same entity, and explores the knowledge fusion form of two tourism entities. First, the synonym matching method, such as entities such as attractions, culture, festivals, and food, extracts synonyms of the above four entities from standardized and authoritative data sources such as popular science encyclopedias, tourism websites, and software to form a synonym dictionary. In the synonym dictionary, a unified and standardized entity name is specified for each group of synonyms. Synonym matching and replacement are performed on the named entities generated during the knowledge extraction process. Secondly, pattern matching rules are defined according to language habits. By mining the collected text data, the contraction rules of words in entities are explored, such as "XXX city", people's usual expression habit is "XXX". Therefore, for entities of types such as city-district-county-township-village, transportation places, hotels, etc., the method of defining pattern matching rules is adopted to perform entity fusion, such as Figure 9 shown.

[0122] S4. Knowledge evaluation: The evaluation of knowledge extraction mainly includes entity evaluation and triple relationship evaluation. The number of entity types and the number of relationship types are both >> 2, which belongs to multivariate evaluation. Therefore, this invention uses accuracy, precision, recall and other methods to evaluate. The evaluation formula is: TP (True Positive) indicates that the model predicts the same category as the actual category; P is the precision rate, R is the recall rate, FP means that the model predicts other categories as the category to be identified; FN means that the model predicts the category to be identified as other categories.

[0123] S5. Knowledge storage; The knowledge graph is a structured semantic knowledge base used to describe concepts in the physical world and their relationships in symbolic form. The basic component unit is a triple such as [entity-relationship-entity] or [entity-attribute-attribute value]. Entities are linked to each other through relationships, forming a network-like knowledge structure. Therefore, the knowledge storage of the present invention is stored using the mainstream graph-based database Neo4j. Neo4j is a high-performance, NOSQL graph database that stores structured data on the network rather than in tables. At the same time, Neo4j can be regarded as a high-performance graph engine with complete transactional features and efficient retrieval performance. In the process of storing the tourism knowledge graph, nodes are first created to represent seven types of tourism-related entities such as attractions, gourmet specialties, hotels, etc. Next, edges are added to represent the relationship between the subject entity and the object entity, such as 10 types of relationships such as located, provided, popular, etc. Finally, attributes are added to enrich the graph semantics, that is, attribute values ​​are added according to the attribute types defined for different entities.

[0124] The technical concept of the present invention is as follows:

[0125] This paper constructs a tourism knowledge graph by integrating large-scale model technology with deep learning models such as MLP_Boost. The knowledge extraction process incorporates a pre-extraction step and proposes a weighted fusion method based on large-scale model Few-Shot Prompting and model fine-tuning techniques. To improve the recognition accuracy of the large-scale model, an MLP_Boost multivariate relationship classification model based on a combination of machine learning and deep learning boosting is constructed to correct and standardize the relationship extraction results. To reduce the information extraction illusion caused by the large model, a matching Few-Shot Prompting technique is used to self-verify entity and relationship types.

[0126] The knowledge graph is constructed using multi-source heterogeneous datasets and information extraction based on a fusion of LLMs and a small MLP_Boost deep learning model. By introducing LLMs and fine-tuning the weighted fusion for knowledge pre-extraction, both precision and recall are achieved. LLMs classification and matching with few-shot learning techniques are used to self-verify named entity recognition and relationship extraction results, reducing the illusion of information extraction. Compared to traditional knowledge graph networks that rely heavily on domain expert participation, the proposed construction method offers greater flexibility and accuracy while significantly reducing the workload and labor costs of manual annotation.

[0127] In one specific embodiment, a weighted fusion method using few-shot prompting and model fine-tuning techniques was used to generate a pre-extracted knowledge set. Named entity recognition employed a large model for self-validation of entities and their types. Relationship extraction primarily constructed a multivariate relationship classification model based on machine learning and deep learning, and performed large-scale model validation of relationship types. Attribute extraction integrated structured and semi-structured attribute features to add triplet attributes. Knowledge fusion employed a synonym rule dictionary for entity fusion. Experiments employed the QWen 14B model for few-shot prompting and the GLM4 model for fine-tuning. Experiments demonstrated that the proposed weighted fusion information extraction method, FSP_FT_Weighted_Fusion, achieved high precision, recall, and F1 scores after modification of the MLP_Boost multivariate relationship classification model and large-scale model self-validation. Finally, the knowledge network was visualized and stored using the Neo4j graph database. The algorithm implementation mainly relies on large model technology, adopting the Qwen2.5 14B model and using the matching Few-ShotPrompting method for knowledge extraction and verification; using the GLM4-9B-Chat model and performing supervised fine-tuning (SFT) through the LLaMA Factory framework.

[0128] Experimental data was primarily collected from tourism-related websites using web crawler technology. After data cleaning and preprocessing, it was manually annotated. The annotated text data volume exceeded 500, covering seven tourism entity types, including attractions, food, and traditional culture, and 10 relationship types. The data was stored as a JSON file containing the source text and a triplet of <entity subject: type, relationship, entity object: type>.

[0129] The named entity recognition experiment is mainly carried out from two aspects: extraction and verification of entities and their types. The methods used are Few-Shot Prompting prompt word extraction, large model fine-tuning extraction, and weighted fusion extraction proposed in this invention. Few-Shot Prompting prompt word extraction focuses on the breadth of entity recognition and is applicable to unlabeled samples; while large model fine-tuning focuses on the accuracy of entity recognition and is applicable to labeled samples. Therefore, this study weightedly fused the extraction results of the two and proposed the FSP_FT_Weighted_Fusion algorithm, and verified the extraction results. By comprehensively comparing different large model algorithms and different few-sample prompt word matching methods, the accuracy, precision, recall rate, etc. of the experimental results are evaluated, as shown in Table 5. Table 5. Evaluation table of tourism entity extraction experimental results comparing Qwen model, ChatGPT Model, and GLM Fineturning fusion model under matching few-sample prompt words and general sample prompt words

[0130] Large model algorithm Precision Recall F1 score Qwen+matched Few-Shot Prompting 0.889 0.870 0.879 Qwen+common Few-Shot Prompting 0.887 0.810 0.847 ChatGPT+matched Few-Shot Prompting 0.938 0.931 0.934 ChatGPT+common Few-Shot Prompting 0.924 0.907 0.915 GLM4+Fine_turning 0.968 0.829 0.893 Qwen_FSP_GLM_FT_Weighted_Fusion 0.943 0.927 0.935

[0131] Table 6. Comparison of NER experimental results with self-verification and no-verification

[0132] Large model algorithm Precision Recall F1 score Qwen+matched Few-Shot Prompting 0.889 0.870 0.879 Qwen+matched Few-Shot Prompting+Verification 0.974 0.866 0.917 FSP_FT_Weighted_Fusion 0.943 0.927 0.935 Qwen_FSP_GLM_FT_Weighted_Fusion+Verification 0.980 0.923 0.951

[0133] In the entity extraction experiments, two large models using the Few-Shot Prompting method were used: the free QWen14B and the commercial Chat GPT. GLM4 was used to fine-tune the large model. Comparative experiments were conducted using both matched and general learning examples, as shown in Table 5. The two large models significantly outperformed the general Few-Shot Prompting method in recognition performance. The QWen large model based on matched Few-Shot Prompting achieved a recall rate of 87%, effectively maintaining recognition breadth. Its precision was 88.9%, and its F1 score was 87.9%. GLM4+Fine_Turning, using supervised learning with labeled data, achieved a significant advantage in recognition accuracy, reaching 96.8%, with a slightly lower recall rate of 82.9%. The weighted fusion model Qwen_FSP_GLM_FT_Weighted_Fusion proposed in this study takes into account both precision and breadth in the knowledge extraction process, with an accuracy of 94.3%, a recall rate of 92.7%, and an F1 score of 93.5%, which are slightly higher than the overall recognition effect of Chat GPT.

[0134] To further improve the hallucination problem of the large model during the information extraction process, the Few-Shot Prompting method of entity type matching was used to perform self-verification of the large model. As shown in Table 6, the experimental results show that under the premise of verification, the recognition effect has been improved overall. Among them, the F1 score of the weighted fusion method FSP_FT_Weighted_Fusion reached 0.951.

[0135] MLP_Boost deep learning algorithm multi-relation classification experiment; mainly attempts to use machine learning and deep learning algorithms to train multi-relation classification models using manually annotated text corpus, thereby improving knowledge graph triples.<Entity,Relation,Entity> The accuracy and standardization of tourism entity relationship classification include 10 types of entity relationships such as location, adjacency, popularity, and provision. In order to prevent the category imbalance caused by multi-classification, the relationship types of text data are first balanced based on undersampling and oversampling to make the samples representative. <e i1 ,t i1 ,r i ,e i2 ,t i2 >Preprocessing, entity type features and relationship type targets are encoded with one-hot encoding and label encoding. The machine learning training process uses CART, GaussianNB, MLP and other algorithms for supervised classification training. The training effect is as follows Figure 10-1 、 Figure 10-2 、 Figure 10-3 、 Figure 10-4 The accuracy of the supervised learning model experimental results is shown in Table 7.

[0136] Table 7. Accuracy of supervised learning model experimental results

[0137]

[0138]

[0139] There are significant differences between the three types of individual learners in the recognition process of 0, 4, 5, and 7 categories. The experiment uses a boosting method to combine the three types of individual learners and train them on the dataset S, that is, first use CART and GaussianNB weak classifiers for training, and the recognition results of the weak classifiers {ρ_cart i} 1≤i≤n , {ρ_gnb i} 1≤i≤nThe model weights used as samples are added to the set S, and a multilayer perceptron (MLP) is used for correct prediction. The boosting model MLP_Boost (i.e., CART_GNB_MLP Boosting) achieves an accuracy of 0.978. The classification results for each relationship are shown in Table 8, showing that the classification of multivariate relationships has been further improved.

[0140] Table 8. MLP_Boost deep learning algorithm 10 classification experimental results

[0141] Precision Recall F1 score Precision Recall F1 score 0 1.00 0.88 0.94 7 0.83 0.67 0.74 1 0.83 1.00 0.91 8 1.00 1.00 1.00 2 1.00 1.00 1.00 9 1.00 1.00 1.00 3 1.00 1.00 1.00 Accuracy 0.98 4 1.00 1.00 1.00 Macro Avg 0.97 0.95 0.96 5 1.00 1.00 1.00 Weighted Avg 0.98 0.98 0.98 6 1.00 1.00 1.00

[0142] Relationship Extraction Verification: After the entity-relationship combination was modified by the MLP_Boost multi-classification model, self-verification was performed on the large model in this stage to optimize the illusion of LLMs in the relationship extraction process. To explore the relationship recognition performance of the weighted fusion model Qwen_FSP_GLM_FT_Weighted_Fusion proposed in this study, this experiment compared QWen14B and Chat GPT, and fine-tuned the large model using GLM4, as shown in Table 9. It can be seen that after weighted fusion, FSP_FT_Weighted_Fusion achieved better recognition results, with an F1 score of 0.929. After self-verification of the large model, the precision reached 0.941 and the F1 score reached 0.932, as shown in Table 10.

[0143] Table 9. Comparison of the Qwen model, ChatGPT model, and GLM Fineturning fusion model under the matching of few-sample prompt words and general sample prompt words, the evaluation table of tourism relationship extraction experimental results

[0144]

[0145]

[0146] Table 10. Comparison of RE experimental results with self-verification and no-verification

[0147] Large model algorithm Precision Recall F1 score Qwen_Few_Shot_Prompting+GLM4_Fine_turning 0.926 0.931 0.929 Qwen_Few_Shot_Prompting+GLM4_Fine_turning Verification 0.941 0.924 0.932

[0148] Knowledge Storage: Based on a large model with a small number of examples and a fine-tuning fusion approach, we extracted information from text corpora through the aforementioned experimental process and constructed a tourism knowledge graph. The extracted triples (entity, relationship, entity) and the collected and processed (entity, attribute, attribute value) were stored in a Neo4j graph database. For example, the attraction entity "XX Palace" is represented by a dot, and the connecting line between the attraction entity subject "XX Palace" and the food specialty entity object "XX Palace" represents the subject-object relationship.

[0149] This embodiment discloses a system for constructing a multi-dimensional tourism knowledge graph using a deep learning hybrid model, which is used to implement any of the above-described methods for constructing a multi-dimensional tourism knowledge graph using a deep learning hybrid model. The system for constructing a multi-dimensional tourism knowledge graph using a deep learning hybrid model includes:

[0150] Data extraction module, used to collect tourism information, including structured data, semi-structured data and unstructured data;

[0151] The knowledge extraction module is used to extract valuable information from structured, semi-structured, and unstructured data based on a large language model and an MLP_Boost deep learning model, and convert it into a structured form, including entities, attributes, and relationships between entities;

[0152] The knowledge fusion module is used to perform entity fusion based on the results of knowledge extraction to eliminate redundant information and correct erroneous data that may be generated in the data extraction stage;

[0153] The knowledge evaluation module is used to perform entity evaluation and relationship evaluation based on the results of knowledge fusion through precision, recall and F1 score;

[0154] The knowledge storage module is used to save structured knowledge in the form of a graph based on the results of knowledge evaluation to obtain a tourism knowledge graph.

[0155] The above embodiments are intended to illustrate the present invention, not to limit the present invention. Therefore, changes in illustrative values ​​or substitutions of equivalent components should still fall within the scope of the present invention.

Claims

1. A multi-dimensional construction method of tourism knowledge graph based on deep learning hybrid model, characterized by: include: S1. Data extraction: Collect tourism information, including structured data, semi-structured data, and unstructured data; S2. Knowledge Extraction: Based on a large language model and MLP-Boost deep learning model, this approach extracts valuable information from structured, semi-structured, and unstructured data and converts it into a structured form, including entities, attributes, and relationships between entities. S3. Knowledge fusion: Based on the results of knowledge extraction in S2, entity fusion is performed to eliminate redundant information and correct erroneous data that may have been generated during the data extraction stage. S4. Knowledge evaluation: Based on the results of knowledge fusion in S3, entity evaluation and relationship evaluation are performed using precision, recall, and F1 score. S5. Knowledge storage: Based on the results of S4 knowledge evaluation, the structured knowledge is saved in the form of a graph to obtain a tourism knowledge graph.

2. The method for constructing a multi-dimensional tourism knowledge graph based on a deep learning hybrid model according to claim 1 is characterized in that: In S1, unstructured data is systematically organized and converted into structured data, and entity types and relationship types are preliminarily explored. In the tourism text data collection stage, text data is collected according to the preliminarily explored entity types and relationship types. The collected text needs to cover every entity type and relationship type. Use data annotation software to perform preliminary data annotation on the text data, and define the text data annotation in Json format. Each element is an object, and each object has text and list attributes to store text data and triple relationships.

3. The method for constructing a multi-dimensional tourism knowledge graph based on a deep learning hybrid model according to claim 2 is characterized in that S2 include: S2.

1. Knowledge Pre-extraction: Use the matched Few-Shot Prompting method and fine-tuning technology for weighted fusion to generate a knowledge pre-extraction set; S2.

2. Named Entity Recognition: Based on a pre-extracted knowledge set, it identifies entities with specific meanings from unstructured text and classifies them into predefined categories. S2.

3. Relationship Extraction: Based on the knowledge pre-extraction set, the semantic relationship between entities is extracted by identifying the relationship between entity pairs in the text; S2.

4. Attribute extraction: Based on a knowledge pre-extraction set, identify and extract attribute information related to a specific entity from unstructured or semi-structured text data.

4. The method for constructing a multi-dimensional tourism knowledge graph based on a deep learning hybrid model according to claim 3 is characterized in that: S2.1 includes: S2.1.

1. Prompt word construction: Use matching Few-Shot Prompting to obtain accurate matching examples for different entity types and different relationship types; S2.1.

2. Large Model Fine-tuning: Further train the large language model based on accurate examples of classification and matching of different entity types and different relationship types. S2.1.

3. Weight-based fusion: Combine the matched Few-Shot Prompting with the fine-tuning of the large model with limited annotated data. Under the premise of taking into account the precision and recall rate of the model, the knowledge extraction results of the two are weightedly fused to obtain the knowledge pre-extraction set.

5. The method for constructing a multi-dimensional tourism knowledge graph based on a deep learning hybrid model according to claim 4 is characterized in that: S2.3 includes: S2.3.

1. Multivariate Relationship Classification Based on the MLP_Boost Deep Learning Boosting Algorithm: A multivariate relationship classification model is trained using manually annotated text corpus using the MLP_Boost deep learning boosting algorithm, thereby improving the accuracy and standardization of knowledge graph triples. S2.3.

2. Relationship Verification: Use matching Few-Shot Prompting technology to provide different input and output examples for different types of tourism entity relationships, and verify them through questions and limited responses.

6. The method for constructing a multi-dimensional tourism knowledge graph based on a deep learning hybrid model according to claim 3 is characterized in that: In S2.4, after extracting the attribute type and attribute value, the attribute type and attribute value information are mapped to the corresponding entity by associating with the entity keyword.

7. The method for constructing a multi-dimensional tourism knowledge graph based on a deep learning hybrid model according to claim 1 is characterized in that: In S3, entity fusion is the merging of records that appear in different forms in different tourism scene corpora but actually refer to the same entity.

8. The method for constructing a multi-dimensional tourism knowledge graph based on a deep learning hybrid model according to claim 1 is characterized in that: In S4, the evaluation formula is: Where P is the precision rate, R is the recall rate, TP means that the model predicts the category and the actual category are the same, FP means that the model predicts other categories as the category to be identified; FN means that the model predicts the category to be identified as other categories.

9. The method for constructing a multi-dimensional tourism knowledge graph based on a deep learning hybrid model according to claim 1 is characterized in that: In S5, knowledge storage is performed using the graph-based database Neo4j.

10. A multi-dimensional tourism knowledge graph construction system based on a deep learning hybrid model, characterized by: A method for constructing a multi-dimensional tourism knowledge graph using a deep learning hybrid model according to any one of claims 1 to 9, wherein the multi-dimensional tourism knowledge graph construction system using the deep learning hybrid model comprises: Data extraction module, used to collect tourism information, including structured data, semi-structured data and unstructured data; The knowledge extraction module is used to extract valuable information from structured, semi-structured, and unstructured data based on a large language model and an MLP_Boost deep learning model, and convert it into a structured form, including entities, attributes, and relationships between entities; The knowledge fusion module is used to perform entity fusion based on the results of knowledge extraction to eliminate redundant information and correct erroneous data that may be generated in the data extraction stage; The knowledge evaluation module is used to perform entity evaluation and relationship evaluation based on the results of knowledge fusion through precision, recall and F1 score; The knowledge storage module is used to save structured knowledge in the form of a graph based on the results of knowledge evaluation to obtain a tourism knowledge graph.

Citation Information

Patent Citations

  • Power transformer knowledge graph construction method for intelligent operation and maintenance

    CN116108190A

  • Data processing method and device, electronic equipment, medium and program product

    CN116401339A

  • Live broadcast content copywriting generation method and device, equipment and medium

    CN117744621A

  • Movie personalized recommendation method and system fusing large language model and knowledge graph

    CN118551123A

  • Weak semantic association extraction type question answering method based on few-sample prompt learning

    CN119066173A

Cited By

  • Big language model teaching quality evaluation method based on Bloom target classification system

    CN121542864A