A method for generating cultural tourism evaluation data based on multi-stage processing

Through multi-stage processing methods and deep learning optimization, the data deviation and adaptability problems in the generation of cultural and tourism evaluation data are solved, and efficient and accurate evaluation data generation is achieved to adapt to market changes and meet the high-quality development needs of the cultural and tourism industry.

CN119990315BActive Publication Date: 2025-09-02LESHAN NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510066643.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-09-02
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The existing cultural and tourism evaluation data generation methods have problems such as data deviation, insufficient emotional analysis accuracy, high computational complexity, poor adaptability of generated data to cultural and tourism scenarios, and delayed response to emergencies, resulting in inaccurate evaluation results.

Method used

A multi-stage processing method is adopted, including data collection, key question-point extraction and evaluation data generation, combined with deep learning and transfer learning technology, through large-scale cultural and tourism-related text data training models, iterative optimization and human-computer interaction mechanisms are introduced to ensure the quality and diversity of data generation.

Benefits of technology

Evaluation data with higher quality and closer to tourists' needs has been generated, which improves computing efficiency and reduces costs. It is suitable for all kinds of scenic spots, can dynamically adapt to market changes and provide efficient and accurate evaluation support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990315B_ABST
    Figure CN119990315B_ABST
Patent Text Reader

Abstract

The present invention provides a method for generating cultural and tourism evaluation data based on multi-stage processing, which belongs to the field of data processing technology. By introducing the collaborative work of data collection, key question point extraction, and evaluation data generation, it ensures that the generated data has high quality, diversity and pertinence. Key question point extraction accurately extracts the key information that tourists are concerned about and generates targeted questions from multiple dimensions, providing a clear framework for data generation; evaluation data generation is based on large-scale cultural and tourism-related text data, and optimizes the data generation process through deep learning and transfer learning technology to ensure that the output data is closer to the actual needs of tourists. At the same time, by introducing iterative optimization and human-computer interaction mechanisms, the system can dynamically adapt to changing market demands and user feedback, ensuring that the quality and diversity of data generation are effectively balanced, thereby providing more accurate and efficient evaluation data support for the cultural and tourism industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides a method for generating cultural tourism evaluation data based on multi-stage processing, which belongs to the technical field of data processing. Background Art

[0002] With the development of big data and artificial intelligence technologies, data analysis and intelligent management in the cultural and tourism industry have made significant progress. Existing data generation mostly relies on natural language processing (NLP) and large language model (LLM) technologies, which can extract effective information from massive amounts of data and help models optimize decision-making. However, research on evaluation generation lags behind, and its drawbacks are evident in the trend of cultural and tourism integration. Existing work related to this patent and its shortcomings are as follows:

[0003] In cultural and tourism data analysis, user reviews are widely used to understand tourist needs and emotional orientation. Existing methods typically crawl user review data from multiple platforms and use natural language processing techniques (such as sentiment analysis and keyword extraction) to analyze tourists' emotional attitudes, thereby providing a basis for optimizing the services of cultural and tourism projects. This method mainly relies on review data, which may lead to data bias. For example, negative or overly positive reviews may affect the reliability of analysis results. In addition, the accuracy of sentiment analysis depends on semantic processing capabilities, but current sentiment analysis technology does not fully understand the emotional expressions unique to some complex contexts or regional cultures, resulting in poor sentiment recognition, which in turn affects the accuracy and effectiveness of evaluation results.

[0004] Graph neural networks (GNNs) are used to classify and cluster cultural and tourism data. They primarily perform data classification and feature extraction by constructing undirected graphs and establishing edges based on similarities between nodes (such as geographic location and attraction type). However, their computational complexity is high, and performance bottlenecks can occur, especially when processing large-scale cultural and tourism data. The extraction and classification accuracy of node features depend on the training of the GNN, and the quality and quantity of the training data directly determine the model's effectiveness. If the training data is insufficient or of low quality, the model classification may be inaccurate, which in turn affects the evaluation results of cultural and tourism projects.

[0005] With the rapid development of modern evaluation technology, large language models (LLMs) have been widely used to generate evaluation data, particularly to increase the difficulty and challenge of evaluation tasks. Although the powerful capabilities of LLMs make them a potential tool for generating cultural and tourism-related evaluation data, this method is not directly used for cultural and tourism evaluation. Instead, it serves as a general evaluation framework to improve the diversity and accuracy of evaluation tasks. The generated data may not be suitable for cultural and tourism evaluation scenarios. The complexity of cultural and tourism data and the need to incorporate local characteristics mean that the challenging questions generated by LLMs must be combined with specific domain knowledge to ultimately generate effective cultural and tourism evaluation data.

[0006] In the tourism industry, with the development of big data technology and the widespread application of artificial intelligence, an increasing number of tasks are beginning to rely on language models (LMs) for data processing and analysis. The subsequent introduction of large language models has brought new breakthroughs in the processing of tourism data. By pre-training large amounts of data to learn language patterns, they can capture the rich semantic information contained in text data. In the tourism industry, tourism-related data often appears in text form, such as tourist reviews, travel guides, news reports, and attraction descriptions. This text data contains a wealth of potentially valuable information. The introduction of language models enables the tourism industry to better understand this data and conduct efficient analysis. However, due to the limited number of model parameters and its primary reliance on historical data for training, these models lack sensitivity to the impact of sudden events (such as epidemics and natural disasters) on tourist flow, which can easily lead to biased prediction results. Furthermore, the collection and updating of scenic area attribute data is cumbersome, and the timeliness and accuracy of the data may not be guaranteed, affecting the model's prediction accuracy.

[0007] To sum up, the current methods for generating cultural and tourism evaluation data mainly use technical means such as large language models and graph neural networks. However, due to many defects such as a single data source that is prone to bias, limited accuracy of sentiment analysis, high computational complexity, poor adaptation of generated data to cultural and tourism scenarios, and insufficient response of models to emergencies and data timeliness, it is difficult to generate accurate and effective cultural and tourism evaluation data sets, which hinders the high-quality advancement of cultural and tourism evaluation work.

[0008] Existing methods have exposed many drawbacks that need to be addressed urgently: over-reliance on user reviews leads to data bias, and negative or overly positive reviews interfere with the reliability of the results; sentiment analysis is limited by semantic processing bottlenecks, and the accuracy of complex context and regional cultural sentiment recognition is poor; GNN's computational complexity soars when faced with large-scale data, and it falls into a performance dilemma, and the insufficient quality and quantity of training data restrict the accuracy of node classification; LLM-generated data is poorly adapted to cultural and tourism scenarios, and lacks consideration of industry complexity and local characteristics; traditional language models are limited in parameter quantity and rely on historical data, resulting in a slow response to emergencies, and the cumbersome collection and update of scenic spot attribute data leads to a lack of timeliness and accuracy. Summary of the Invention

[0009] In order to solve the above problems, the present invention proposes a cultural and tourism evaluation data generation method based on multi-stage processing. By introducing the collaborative work of data collection, key question point extraction, and evaluation data generation, it ensures that the generated data has high quality, diversity and pertinence.

[0010] Key question point extraction accurately extracts the key information that tourists are concerned about and generates targeted questions from multiple dimensions, providing a clear framework for data generation;

[0011] Evaluation data generation is based on large-scale cultural and tourism-related text data. The data generation process is optimized through deep learning and transfer learning technologies to ensure that the output data is closer to the actual needs of tourists.

[0012] At the same time, by introducing iterative optimization and human-computer interaction mechanisms, the present invention enables the system to dynamically adapt to changing market demands and user feedback, ensuring an effective balance between data generation quality and diversity, thereby providing more accurate and efficient evaluation data support for the cultural and tourism industry.

[0013] The present invention proposes a method for generating cultural tourism evaluation data based on multi-stage processing, which covers three closely connected stages:

[0014] First, we use cutting-edge big model technology to build an information source library targeting the six key dimensions of the cultural tourism industry: "food, accommodation, transportation, travel, shopping, and entertainment." We extensively collect and integrate massive text data from multiple channels, including travel guides, tourist diaries, hotel and restaurant reviews, and official introductions to scenic spots. We then use advanced natural language processing algorithms to deeply explore key information points with potential evaluation value, such as specialty food ingredients, hotel transportation hub accessibility, and seasonal changes in scenic spots. Through structured integration, we form a key information collection.

[0015] Then, driven by powerful computing capabilities and a large model with deep learning architecture, the aforementioned key information set is organically integrated with contextual information such as given cultural and tourism scenarios and project details as input. Through complex neural network operations, a complete evaluation data set is generated, including (questions, answers, question types, question accuracy, answer matching, and generation reliability).

[0016] Finally, an iterative optimization mechanism is constructed based on the evaluation results of the generated indicators. If any of the three indicators fails to reach the preset high-quality threshold, the large model regeneration process is immediately initiated. After multiple iterations, all indicators are met. On this basis, manual review is introduced to conduct a comprehensive and rigorous secondary verification of the data quality finally generated by the model from multiple perspectives such as the professional depth of the questions, the accuracy of the answers, and the adaptability of the data to actual cultural and tourism scenarios. This ensures the continuous and stable generation of high-quality cultural and tourism question and answer evaluation data, injecting strong impetus into the digital upgrade and high-quality development of the cultural and tourism industry. Furthermore, this method can be used to continuously generate high-quality cultural and tourism question and answer evaluation data.

[0017] The specific steps are:

[0018] S1. Data Collection. First, collect cultural and tourism-related data from multiple sources, including official tourism documents, tourist reviews, travel blogs, online travel Q&A platforms, and social media data, to form the original data. Specific process:

[0019] S1.1. Data cleaning and preprocessing: Use efficient data cleaning technology to ensure high data quality and consistency.

[0020] S1.2. Dimension Division and Classification: Ensure that the data of each dimension can accurately reflect its characteristics. The cleaned data is expressed according to the six core dimensions of the cultural tourism industry: "Eating (E1)", "Accommodation (E2)", "Transportation (E3)", "Tourism (E4)", "Shopping (E5)", and "Entertainment (E6)" as follows:

[0021]

[0022] In each dimension, It is the set of all cultural and tourism information under this dimension; each dimension is represented as a set containing multiple data

[0023] S2. Extract key questions: We accurately extract key questions closely related to tourism from a vast amount of tourism information. By applying natural language processing technology, we deeply explore various aspects of tourists' potential concerns during their tourism journey. These key questions will serve as important guidance for subsequent data generation, ensuring that the generated data is highly targeted and practical, and can truly reflect tourists' real needs and concerns.

[0024] Specific process:

[0025] S2.1. Extraction Prompt Design: Design a prompt template P1 specifically for different dimensions and provide a one-shot example in P1. Embed the cleaned and normalized data into this prompt to form a complete model input sequence P′1 that meets the model input requirements. The language model for extracting key question points is expressed as:

[0026] G k :Ψ×Φ→Ψ

[0027] is the data space, and Φ is the parameter space of the generative model.

[0028] S2.2, key question point extraction: the input text sequence of each dimension is sent to G k In the process, the model deeply analyzes and infers semantics, vocabulary associations, and contextual logic, and finally obtains key question points of different dimensions:

[0029]

[0030] The parameters of the current model are θ∈Φ, where i is the dimension, k is the number of key information in different dimensions, and θ is the parameter for extracting key question points.

[0031] Finally, after manual screening, key question point information of different dimensions is extracted.

[0032] S3. Evaluation data generation. Its underlying architecture is based on an existing large language model. Through a multi-head self-attention mechanism, it processes different representation subspaces of key input information in parallel, thereby achieving comprehensive capture and deep understanding of semantics. During the training phase, large-scale cultural and tourism-related text data is used, covering travel notes, guides, official introductions, and tourist reviews. This massive data is used to train the model to learn rich language patterns and cultural and tourism domain knowledge. Based on the pre-trained weights of the large language model, transfer learning technology is used to fine-tune it on a specific cultural and tourism dataset, enabling the model to better adapt to the language habits and cultural and tourism scene characteristics of the local domain.

[0033] Specific process:

[0034] S3.1, Prompt construction for evaluation data generation. For a given context C of the generated question, the prompt template for evaluation data generation is P2. Where λ is the number of key question points, E i For each dimension, Combining it with the given context C according to specific splicing rules, P′2 obtains Prompt: Will Serves as input information to drive the model to generate evaluation data.

[0035] S3.2. Design a template for generating evaluation data. Its structure is (Q, A, T, L1, L2, L3). Q is a question related to culture and tourism generated based on the prompt word P′2; A is the answer to Q generated based on the prompt word; T represents the question type. Due to the particularity of culture and tourism data, the type judgment function f is used according to pre-set rules. T (Q) Divide the problem into factual (F) and planning (P), that is, T = f T (Q), T∈{F,P}; question correctness L1 is an indicator of whether the question generated by the model is a correct question based on the context. If the model judges whether the generated question is a correct question based on the context; answer matching L2 is used to judge whether the answer accurately matches the question; generation reliability L3 is determined by the large model to determine whether the question and answer can be directly extracted from the context.

[0036] S3.3, is the data space, Φ is the parameter space of the generative model, the parameters of the current model are θ∈Φ, t represents the time step, i represents the i-th iteration at the current time step, G d Data generates a language model, and the initial generated data is represented as:

[0037]

[0038] in The evaluation data generated for the i-th iteration of time step t, L i ∈{0,1}. The indicators and evaluation data are generated simultaneously by the language model.

[0039] S3.4. Perform iterative calculations and adjustments on the evaluation indicators L1, L2, and L3 of the generated data. The iterative process depends on the generation capability of the language model. The diversity of generation ensures that the model can gradually generate data that meets the evaluation criteria through multiple attempts. The convergence assumption is that within a finite number of generations, if the language model can generate data such that L1 = 1, L2 = 1, and L3 = 1, then the generated data X is considered to be t When the expected quality standard is reached, the iterative process stops and the final generated dataset X is output. * :

[0040] X * =X t , where L1=L2=L3=1

[0041] If there is any indicator L in a round t i =0,(i∈{1,2,3}), it indicates that the generated data does not meet the quality requirements. In this case, it is necessary to iteratively generate data based on the evaluation results, that is, perform i+1 iterations:

[0042]

[0043] S3.5. Manual evaluation: After the evaluation data is generated, the accuracy of the generated data is verified from the linguistic level. For each piece of data generated:

[0044] X=(Q,A,T,L i )

[0045] L i ∈{L1,L2,L3}

[0046] R i ∈{R1,R2,R3}

[0047] The verification process consists of two stages: automatic model determination L and manual review R. The probability of the two indicators being equal in all generated data is calculated. From a linguistic perspective, the data generated by the model is analyzed to see whether it meets the structural integrity and semantics of human language, thereby fully verifying the reliability and credibility of the evaluation system. The consistency index is expressed as:

[0048]

[0049] Judgment result L iWith R i Indicates the i-th quality indicator of the final evaluation data currently generated. Manual review requires the indicator L that the model automatically determines i Perform manual evaluation. Based on the automatic judgment results, add manual review results R i Verify the generated questions and answers R i ∈{0,1}. R1=1 indicates that the question has been manually verified to be qualified, R2=1 indicates that the answer has been manually verified to be qualified, and R3=1 indicates that the answer can be extracted from the context. The generated data is considered reliable by calculating the consistency index of L and R. If the probability that the two indicators are equal in all generated data is greater than 90%, the generated data is considered reliable.

[0050] Compared with the existing technology, the present invention has significant advantages. Through the collaborative work of key question point extraction and evaluation data generation, the core issues that tourists care about can be accurately refined, and higher quality, more targeted and diverse evaluation data can be generated. This method not only improves the accuracy of data generation, but also better reflects the actual needs of tourists and shows greater flexibility in responding to the impact of emergencies. The use of deep learning to optimize the data generation process improves computing efficiency and reduces costs, and is applicable to all types of scenic spots. Overall, the present invention provides an efficient, accurate and sustainable evaluation data generation solution for the cultural and tourism industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is a flow chart of the method for generating question-answer evaluation data in the field of culture and tourism of the present invention;

[0052] Figure 2 This is a flow chart for constructing the initial data of Mount Emei in the embodiment. DETAILED DESCRIPTION

[0053] The specific technical solutions of the present invention are described with reference to the accompanying drawings.

[0054] This paper proposes a method for generating cultural tourism evaluation data based on a multi-stage process. This method aims to ensure that the generated evaluation data comprehensively covers all dimensions of cultural tourism, has a high degree of accuracy and reliability, and truly and objectively reflects the actual situation, thereby precisely meeting the cultural tourism industry's urgent need for high-quality evaluation data.

[0055] correspond Figure 1 To achieve this goal, the present invention reconstructs the traditional large-scale model evaluation data generation method and divides it into three core components, each of which has a key function:

[0056] S1. Data collection. First, collect cultural and tourism related data from multiple sources, including official tourism documents, tourist reviews, travel blogs, online tourism Q&A platforms, social media data, etc. to form the original data. Figure 2 ,These data sources are diverse and heterogeneous, and can provide cultural and tourism information from multiple angles.

[0057] Specific process:

[0058] S1.1 Data Cleaning and Preprocessing: Collected cultural and tourism data often contains inconsistent formats, redundant information, irrelevant content, and noisy information. The system uses efficient data cleaning techniques, such as deduplication, standardization, and text error correction, to ensure high data quality and consistency.

[0059] S1.2. Dimension Division and Classification: Ensure that the data of each dimension can accurately reflect its characteristics. The cleaned data is expressed according to the six core dimensions of the cultural tourism industry: "Eating (E1)", "Accommodation (E2)", "Transportation (E3)", "Tourism (E4)", "Shopping (E5)", and "Entertainment (E6)" as follows:

[0060]

[0061] In each dimension, For all cultural and tourism information sets under this dimension, such as food Dining information in the dimension, accommodation Accommodation information in the dimension, etc. Each dimension can be represented as a set containing multiple data items.

[0062] S2. Key Question Point Extraction Method. We precisely extract key questions closely related to tourism from a vast amount of tourism information. By leveraging natural language processing technology, this extraction method can deeply explore various aspects of a tourist's tourism experience, including but not limited to the characteristics of the attraction, dining and accommodation, transportation convenience, and service quality. These key questions will serve as an important guide for subsequent data generation, ensuring that the generated data is highly targeted and practical, truly reflecting tourists' real needs and concerns.

[0063] Specific process:

[0064] S2.1. Extraction Prompt Design: Design a prompt template P1 specifically for different dimensions and provide a one-shot example in P1. Embed the cleaned and normalized data into this prompt to form a complete model input sequence P′1 that meets the model input requirements. The language model for extracting key question points is expressed as:

[0065] G k :Ψ×Φ→Ψ

[0066] is the data space, and Φ is the parameter space of the generative model.

[0067] S2.2, key question point extraction: the input text sequence of each dimension is sent to G k In the process, the model deeply analyzes and infers semantics, vocabulary associations, and contextual logic, and finally obtains key question points of different dimensions:

[0068]

[0069] The parameters of the current model are θ∈Φ, where i is the dimension, k is the number of key information in different dimensions, and θ is the parameter for extracting key question points.

[0070] Finally, after manual screening, the key question points of different dimensions are extracted as follows: There are 102 items in total, covering special cuisine, restaurant services, etc.; "accommodation" dimension A total of 89, covering accommodation types, guest experience, etc.; "travel" dimension There are 49 in total, including information on public transportation, self-driving tours, etc.; "Travel" dimension A total of 456, covering scenic spot culture, travel guides, etc.; "Purchase" dimension A total of 48, focusing on special products, shopping places; "Entertainment" dimension There are 20 in total, related to entertainment activities and leisure venues, which lay the foundation for the subsequent evaluation data generation.

[0071] S3. Evaluation data generation. Its underlying architecture is based on the existing large language model. Through the multi-head self-attention mechanism (Multi-Head Attention), it can process different representation subspaces of input key information in parallel, thereby achieving comprehensive capture and in-depth understanding of semantics. During the training phase, large-scale cultural and tourism-related text data was used, covering multi-source heterogeneous texts such as travel notes, guides, official introductions, and tourist reviews. The model was trained with massive data to learn rich language patterns and cultural and tourism field knowledge. Based on the pre-trained weights of the large language model, fine-tuning is performed on a specific cultural and tourism data set through transfer learning technology, so that the model can better adapt to the language habits and cultural and tourism scene characteristics of the local domain.

[0072] Specific process:

[0073] S3.1, Prompt construction for evaluation data generation. For a given context C of the generated question, the prompt template for evaluation data generation is P2. Where λ is the number of key question points, E i For each dimension, Combining it with the given context C according to specific splicing rules, P′2 obtains Prompt: Will Serves as input information to drive the model to generate evaluation data.

[0074] S3.2. Design a template for generating evaluation data. Its structure is (Q, A, T, L1, L2, L3). Q is a question related to culture and tourism generated based on the prompt word P′2; A is the answer to Q generated based on the prompt word; T represents the question type. Due to the particularity of culture and tourism data, the type judgment function f is used according to pre-set rules. T (Q) Divide the problem into factual (F) and planning (P), that is, T = f T (Q), T∈{F,P}; question correctness L1 is an indicator of whether the question generated by the model is a correct question based on the context. If the model judges whether the generated question is a correct question based on the context; answer matching L2 is used to judge whether the answer accurately matches the question; generation reliability L3 is determined by the large model to determine whether the question and answer can be directly extracted from the context.

[0075] S3.3, is the data space, Φ is the parameter space of the generative model, the parameters of the current model are θ∈Φ, t represents the time step, i represents the i-th iteration at the current time step, G d Data generates a language model, and the initial generated data is represented as:

[0076]

[0077] in The evaluation data generated for the i-th iteration of time step t, L i ∈{0,1}. The indicators and evaluation data are generated simultaneously by the language model.

[0078] S3.4. To ensure the quality and accuracy of the generated data, the evaluation indicators L1, L2, and L3 of the generated data are iteratively calculated and adjusted in rounds. The iterative process depends on the generation ability of the language model. The diversity of generation ensures that the model can gradually generate data that meets the evaluation criteria through multiple attempts. The convergence assumption is: within a finite number of generations, if the language model can generate data such that L1 = 1, L2 = 1, and L3 = 1, then the generated data X is considered to be t When the expected quality standard is reached, the iterative process stops and the final generated dataset X is output. * :

[0079] X * =X t , where L1=L2=L3=1

[0080] If there is any indicator L in a round t i =0,(i∈{1,2,3}), it indicates that the generated data does not meet the quality requirements. In this case, it is necessary to iteratively generate data based on the evaluation results, that is, perform i+1 iterations:

[0081]

[0082] S3.5. Manual evaluation: After the evaluation data is generated, in order to further verify the accuracy of the generated data from the linguistic level, for each piece of generated data:

[0083] X=(Q,A,T,L i )

[0084] L i ∈{L1,L2,L3}

[0085] R i ∈{R1,R2,R3}

[0086] The verification process consists of two stages: automatic model determination L and manual review R. The probability of the two indicators being equal in all generated data is calculated. From a linguistic perspective, the data generated by the model is analyzed to see whether it meets the structural integrity and semantics of human language, thereby fully verifying the reliability and credibility of the evaluation system. The consistency index is expressed as:

[0087]

[0088] Judgment result L i With R i Indicates the i-th quality indicator of the final evaluation data currently generated. Manual review requires the indicator L that the model automatically determines i Perform manual evaluation. Based on the automatic judgment results, add manual review results R i Verify the generated questions and answers R i ∈{0, 1}. R1=1 indicates that the question has been manually verified to be qualified, R2=1 indicates that the answer has been manually verified to be qualified, and R3=1 indicates that the answer can be extracted from the context. The generated data is considered reliable by calculating the consistency index of L and R. If the probability that the two indicators are equal in all generated data is greater than 90%, then the generated data is considered reliable.

[0089] This paper proposes a phased data generation method that combines a large language model, a key question point extractor, and an indicator optimization mechanism to form a complete and efficient evaluation data generation system from data extraction to generation to verification.

[0090] We explore key question points from the six core dimensions of "eating, accommodation, transportation, sightseeing, shopping, and entertainment", accurately locate tourists' concerns, and use this as the core guiding information for data generation, ensuring the pertinence and coverage of evaluation data from the source.

[0091] Leveraging the powerful text generation capabilities of the large language model, we use these key questions as a guide to develop content creation point by point. This ensures that the generated data closely revolves around the core concerns of tourists, avoiding the generation of broad, unfocused, and redundant information in content details.

[0092] The present invention introduces three quality evaluation indicators, L1, L2, and L3, to comprehensively evaluate the quality of generated data from data structure, semantic rationality to usability, and optimize the results through an iterative generation mechanism.

[0093] During the quality assessment process, we innovatively combine automatic model judgment and manual review to verify the language structure integrity and semantic rationality of the generated data from a human language perspective, greatly improving the reliability of the evaluation system.

Claims

1. A method for generating cultural tourism evaluation data based on multi-stage processing, characterized in that: Including data collection, key question point extraction, and evaluation data generation; Key question point extraction extracts key information that tourists are concerned about, generates targeted questions from multiple dimensions, and provides a clear framework for data generation; Evaluation data is generated based on large-scale cultural and tourism-related text data. Deep learning and transfer learning technologies are used to optimize the data generation process to ensure that the output data is closer to the actual needs of tourists. At the same time, by introducing iterative optimization and human-computer interaction mechanisms, we can dynamically adapt to changing market demands and user feedback to ensure an effective balance between data generation quality and diversity; The specific steps include: S1. Data collection: First, collect cultural and tourism-related data from multiple sources, including official tourism documents, tourist reviews, travel blogs, online tourism Q&A platforms, and social media data to form the original data; S2. Extract key questions: We accurately extract key questions closely related to tourism from a vast amount of tourism information. By applying natural language processing technology, we deeply explore various aspects that tourists may be concerned about during their tourism journey. These key questions will serve as important guidance for subsequent data generation, ensuring that the generated data is highly targeted and practical, and can truly reflect tourists' real needs and concerns. S3. Evaluation data generation. Its underlying architecture is based on an existing large language model. Through a multi-head self-attention mechanism, it processes different representation subspaces of key input information in parallel, thereby achieving comprehensive capture and in-depth understanding of semantics. During the training phase, it uses large-scale cultural and tourism-related text data, including travel notes, guides, official introductions, and tourist reviews. This massive data is used to train the model to learn rich language patterns and cultural and tourism domain knowledge. Based on the pre-trained weights of the large language model, transfer learning technology is used to fine-tune the model on a specific cultural and tourism dataset, enabling it to better adapt to the language habits and cultural and tourism scene characteristics of the local domain. S1 specifically includes the following sub-steps: S1.

1. Data cleaning and preprocessing: Use efficient data cleaning technology to ensure high data quality and consistency; S1.

2. Dimension Division and Classification: Ensure that the data in each dimension accurately reflects its characteristics. The cleaned data is expressed according to the six core dimensions of the cultural tourism industry: "Eating (E1)", "Accommodation (E2)", "Transportation (E3)", "Tourism (E4)", "Shopping (E5)", and "Entertainment (E6)" as follows: In each dimension, It is the set of all cultural and tourism information under this dimension; each dimension is represented as a set containing multiple data S2 specifically includes the following sub-steps: S2.

1. Extraction Prompt Design: Design a prompt template P1 specifically for different dimensions and provide a oneshot example in P1. Embed the cleaned and normalized data into this prompt to form a complete model input sequence P′1 that meets the model input requirements. The language model for extracting key question points is expressed as: G k :Ψ×Φ→Ψ‘ is the data space, Φ is the parameter space of the generative model; S2.2, key question point extraction: the input text sequence of each dimension is sent to G k In the process, the model deeply analyzes and infers semantics, vocabulary associations, and contextual logic, and finally obtains key question points of different dimensions: The parameters of the current model are θ∈Φ, where i is the dimension, k is the number of key information in different dimensions, and θ is the parameter for extracting key question points; Finally, after manual screening, the key question points of different dimensions were extracted; S3 specifically includes the following sub-steps: S3.1, evaluation data generation prompt construction; for a given generation question context C, the prompt template generated by the evaluation data is P2, for Where λ is the number of key question points, E i For each dimension, Combining it with the given context C according to specific splicing rules, P'2 obtains Prompt: Will Serves as input information to drive the model to generate evaluation data; S3.

2. Design a template for generating evaluation data. Its structure is (Q, A, T, L1, L2, L3). Q is a question related to culture and tourism generated based on the prompt word P'2. A is the answer to Q generated based on the prompt word. T represents the type of question. Due to the particularity of culture and tourism data, the type judgment function f is used according to pre-set rules. T (Q) Divide the problem into factual (F) and planning (P), that is, T = f T (Q), T∈{F,P}; the question correctness L1 is an indicator of whether the question generated by the model is a correct question based on the context. If the model determines whether the generated question is a correct question based on the context; the answer matching L2 is used to determine whether the answer accurately matches the question; the generation reliability L3 is determined by the large model to determine whether the question and answer can be accurately matched. Directly extracted below; S3.3, is the data space, Φ is the parameter space of the generative model, the parameters of the current model are θ∈Φ, t represents the time step, i represents the i-th iteration at the current time step, G d Data generates a language model, and the initial generated data is represented as: G d :χ×Φ→X i in The evaluation data generated for the i-th iteration of time step t, L i ∈{0,1}; indicators and evaluation data are generated simultaneously by the language model; S3.

4. Iterate and adjust the evaluation indicators L1, L2, and L3 of the generated data in rounds. The iterative process depends on the generation ability of the language model. The diversity of generation ensures that the model can gradually generate data that meets the evaluation criteria through multiple attempts. The convergence assumption is that within a finite number of generations, the language model can generate data such that L1 = 1, L2 = 1, and L3 = 1, then the generated data X is considered to be t When the expected quality standard is reached, the iterative process stops and the final generated dataset X is output. * : X * =X t , where L1=L2=L3=1 If there is any indicator L in a round t i =0,(i∈{1,2,3}), it indicates that the generated data does not meet the quality requirements. In this case, it is necessary to iteratively generate data based on the evaluation results, that is, perform i+1 iterations: S3.

5. Manual evaluation: After the evaluation data is generated, the accuracy of the generated data is verified from the linguistic level. For each piece of data generated: X=(Q,A,T,L i ) <h2 style=";text-align:left;direction:ltr">L<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> ∈{L1,L2,L3} R i ∈{R1,R2,R3} The verification process consists of two stages: automatic model determination L and manual review R; calculating the probability of the two indicators being equal in all generated data, and analyzing from a linguistic perspective whether the data generated by the model meets the structural integrity and semantics of human language, thereby comprehensively verifying the reliability and credibility of the evaluation system. The consistency index is expressed as: Judgment result L i With R i Indicates the i-th quality indicator of the final evaluation data currently generated; manual review requires the indicator L automatically determined by the model i Perform manual evaluation; based on the automatic judgment results, add manual review results R i Verify the generated questions and answers R i ∈{0,1}; where R1=1 indicates that the manual verification question is qualified, R2=1 indicates that the manual verification answer is qualified, and R3=1 indicates that the answer can be extracted from the context; by calculating the consistency index of L and R, that is, the probability that the two indicators are equal in all generated data is greater than 90%, then the generated data is considered reliable.

Citation Information

Patent Citations

  • Hotel intelligent question and answer recommendation and decision support analysis method and system

    CN110807091A

  • Question and answer method, electronic equipment and computer storage medium

    CN117851556A