Tourism evaluation data generation method based on multi-stage processing

By introducing multi-stage processing methods in the generation of cultural and tourism evaluation data, including data collection, key question-point extraction and large language model optimization, the problems of data deviation, high computing complexity and poor adaptability in the existing technology are solved, and high-quality, diverse and targeted cultural and tourism evaluation data generation are achieved.

CN119990315AActive Publication Date: 2025-05-13LESHAN NORMAL UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510066643.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The existing cultural and tourism evaluation data generation methods have problems such as single data source, limited emotional analysis accuracy, high computational complexity, poor adaptation of generated data to cultural and tourism scenarios, and insufficient response to emergencies and data timeliness, making it difficult to generate accurate and effective cultural and tourism evaluation data sets.

Method used

A method of cultural and tourism evaluation data generation based on multi-stage processing is proposed. Through the coordinated work of data collection, extraction of key question points and evaluation data generation, the generated data is ensured to be of high quality, diversity and targeted. The method includes optimizing the data generation process using large language models and deep learning techniques, and dynamically adapting to market demand and user feedback through iterative optimization and human-computer interaction mechanisms.

Benefits of technology

By accurately refining the core issues that tourists care about, and generating higher quality, diversity and targeted evaluation data, the accuracy and adaptability of data generation are improved, which can better reflect the actual needs of tourists, and show higher flexibility in responding to emergencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990315A_ABST
    Figure CN119990315A_ABST
Patent Text Reader

Abstract

The invention provides a text travel evaluation data generation method based on multi-stage processing, belongs to the technical field of data processing, and ensures that generated data has relatively high quality, diversity and pertinence by introducing cooperative work of data collection, key question point extraction and evaluation data generation. Key questioning point extraction accurately extracts key information concerned by tourists, targeted questions are generated from multiple dimensions, and a clear framework of data generation is provided; evaluation data generation is based on large-scale travel related text data, a data generation process is optimized through deep learning and transfer learning technologies, and it is ensured that output data is closer to actual demands of tourists. Meanwhile, by introducing iterative optimization and a man-machine interaction mechanism, the system can dynamically adapt to continuously changing market demands and user feedback, and it is ensured that data generation quality and diversity are effectively balanced, so that more accurate and efficient evaluation data support is provided for the text and travel industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides a method for generating cultural tourism evaluation data based on multi-stage processing, which belongs to the technical field of data processing. Background Art

[0002] With the development of big data and artificial intelligence technology, data analysis and intelligent management in the cultural and tourism industry have made significant progress. Existing data generation mostly relies on natural language processing (NLP) and large language model (LLM) technology, which can extract effective information from massive data and help the model optimize decision-making. However, the research on evaluation generation lags behind, and the disadvantages are fully revealed under the trend of cultural and tourism integration. Existing work related to this patent and its shortcomings:

[0003] In the analysis of cultural and tourism data, user reviews are widely used to understand tourists' needs and emotional orientation. Existing methods usually crawl user review data from multiple platforms and use natural language processing techniques (such as sentiment analysis and keyword extraction) to analyze tourists' emotional attitudes, thereby providing a basis for the service optimization of cultural and tourism projects. This method mainly relies on review data and may lead to data bias. For example, bad reviews or overly positive reviews may affect the reliability of the analysis results. In addition, the accuracy of sentiment analysis depends on semantic processing capabilities, but current sentiment analysis technology does not fully understand the emotional expressions unique to some complex contexts or regional cultures, resulting in poor sentiment recognition effects, which in turn affects the accuracy and effectiveness of the evaluation results.

[0004] Graph neural networks (GNNs) are used to classify and cluster cultural and tourism data. They mainly construct undirected graphs and establish edges based on the similarities between nodes (such as geographic location, type of scenic spot, etc.) to perform data classification and feature extraction. However, their computational complexity is high, especially when processing large-scale cultural and tourism data, and they face performance bottlenecks. The extraction and classification accuracy of node features depend on the training effect of the graph neural network, and the quality and quantity of training data directly determine the effect of the model. If the training data is insufficient or the data quality is low, the model classification may be inaccurate, which in turn affects the evaluation results of the cultural and tourism project.

[0005] With the rapid development of modern evaluation technology, large language models (LLMs) have been widely used to generate evaluation data, especially to increase the difficulty and challenge of evaluation tasks. Although the powerful capabilities of LLMs make them a potential tool for generating cultural and tourism related evaluation data, this method is not directly used for cultural and tourism evaluation, but as a general evaluation framework to improve the diversity and accuracy of evaluation tasks. The generated data may not be suitable for cultural and tourism evaluation scenarios. The complexity of cultural and tourism data and the need to combine local characteristics mean that it is necessary to combine the challenging problems generated by LLM with specific domain knowledge to finally form effective cultural and tourism evaluation data.

[0006] In the cultural and tourism industry, with the development of big data technology and the widespread application of artificial intelligence, more and more tasks have begun to rely on language models (LM) for data processing and analysis. The subsequent introduction of large language models has brought new breakthroughs in the processing of cultural and tourism data. Through large-scale data pre-training to learn language patterns, it can capture rich semantic information in text data. In the cultural and tourism industry, tourism-related data often appear in text form, such as tourist comments, travel guides, news reports, and scenic spot introductions. These text data contain a lot of potential valuable information. The introduction of language models enables the cultural and tourism industry to better understand these data and conduct efficient analysis. However, due to the limited number of model parameters and the main reliance on historical data for training, it lacks sensitivity to the impact of emergencies (such as epidemics, natural disasters, etc.) on tourist flow, which can easily lead to deviations in prediction results. In addition, the collection and update process of scenic spot attribute data is relatively cumbersome, and the timeliness and accuracy of the data may not be guaranteed, affecting the prediction accuracy of the model.

[0007] To sum up, the current methods for generating cultural and tourism evaluation data mainly use technical means such as large language models and graph neural networks. However, due to the existence of many defects such as a single data source that is prone to bias, limited accuracy of sentiment analysis, high computational complexity, poor adaptation of generated data to cultural and tourism scenarios, and insufficient response of the model to emergencies and data timeliness, it is difficult to generate accurate and effective cultural and tourism evaluation data sets, which hinders the high-quality advancement of cultural and tourism evaluation work.

[0008] Existing methods have exposed many drawbacks that need to be addressed: over-reliance on user reviews leads to data bias, and bad reviews or overly good reviews interfere with the reliability of results; sentiment analysis is limited by the bottleneck of semantic processing, and the accuracy of complex context and regional cultural sentiment recognition is poor; GNN's computational complexity soars when faced with large-scale data, and it falls into a performance dilemma, and the insufficient quality and quantity of training data restrict the accuracy of node classification; LLM-generated data is poorly adapted to cultural and tourism scenarios, and lacks consideration of industry complexity and local characteristics; traditional language models are limited in parameter quantity, rely on historical data, and respond slowly to emergencies, and the cumbersome collection and update of scenic spot attribute data leads to a lack of timeliness and accuracy. Summary of the invention

[0009] In order to solve the above problems, the present invention proposes a method for generating cultural and tourism evaluation data based on multi-stage processing. By introducing the collaborative work of data collection, key question point extraction, and evaluation data generation, it ensures that the generated data has high quality, diversity and pertinence.

[0010] Key question point extraction accurately extracts the key information that tourists are concerned about and generates targeted questions from multiple dimensions, providing a clear framework for data generation;

[0011] The evaluation data is generated based on large-scale cultural and tourism-related text data. The data generation process is optimized through deep learning and transfer learning techniques to ensure that the output data is closer to the actual needs of tourists.

[0012] At the same time, by introducing iterative optimization and human-computer interaction mechanisms, the present invention enables the system to dynamically adapt to changing market demands and user feedback, ensuring an effective balance between data generation quality and diversity, thereby providing more accurate and efficient evaluation data support for the cultural and tourism industry.

[0013] The present invention proposes a method for generating cultural tourism evaluation data based on multi-stage processing, which covers three closely connected stages:

[0014] First, we use cutting-edge big model technology to build an information source library for the six key dimensions of the cultural and tourism industry: "food, accommodation, transportation, travel, shopping and entertainment". We collect and integrate massive text data from multiple channels such as travel guides, tourists' travel notes, hotel and restaurant reviews, and official introductions of scenic spots. Then, we use excellent natural language processing algorithms to deeply mine key information points for potential evaluation, such as special food ingredients, hotel transportation hub accessibility, seasonal landscape changes in scenic spots, etc., and form a key information collection through structured integration.

[0015] Then, with powerful computing power and a large model of deep learning architecture as the core driving force, the above key information set is organically integrated with contextual information such as given cultural and tourism scenes and project details as input, and a complete evaluation data set containing (questions, answers, question types, question correctness, answer matching, and generation reliability) is generated through complex neural network operations;

[0016] Finally, an iterative optimization mechanism is constructed based on the evaluation results of the generated indicators. When any of the three indicators does not reach the preset high-quality threshold, the large model regeneration process is immediately started. After multiple iterations until all indicators meet the standards, manual review is introduced on this basis. The quality of the data finally generated by the model is comprehensively and rigorously verified from multiple perspectives such as the professional depth of the questions, the accuracy of the answers, and the adaptability of the data to the actual cultural and tourism scenes. This ensures that high-quality cultural and tourism question and answer evaluation data can be continuously and stably generated, injecting strong impetus into the digital upgrade and high-quality development of the cultural and tourism industry. Furthermore, this method can be used to continuously generate high-quality cultural and tourism question and answer evaluation data.

[0017] The specific steps are:

[0018] S1. Data collection. First, collect cultural and tourism-related data from multiple sources, including official tourism documents, tourist reviews, travel blogs, online tourism Q&A platforms, and social media data to form the original data. Specific process:

[0019] S1.1. Data cleaning and preprocessing: Adopt efficient data cleaning techniques to ensure the high quality and consistency of data.

[0020] S1.2. Dimension division and classification: Ensure that the data in each dimension can accurately reflect its characteristics. Classify the cleaned data according to the six core dimensions of the cultural and tourism industry: "Eat (E 1 )", "Accommodation (E 2 )", "Transportation (E 3 )", "Tour (E 4 )", "Shopping (E 5 )", "Entertainment (E 6 )", expressed as:

[0021]

[0022] Among them, each dimension is the set of all cultural and tourism information under this dimension; each dimension is represented as a set containing multiple data

[0023] S2. Extraction of key question points; Precisely extract key question points closely related to cultural and tourism from a vast amount of cultural and tourism information. By applying natural language processing techniques, deeply explore various aspects that tourists may be concerned about during the cultural and tourism process. These key question points will serve as important guidelines for subsequent data generation, ensuring that the generated data is highly targeted and practical, and can truly reflect the real needs and concerns of tourists.

[0024] Specific process:

[0025] S2.1. Extraction Prompt design: Design a Prompt template P 1 specific to different dimensions, and give a one-shot example in P 1 . Embed the cleaned and normalized data into this Prompt to form a complete model input sequence P' 1 that meets the model input requirements. The key question point extraction language model is expressed as:

[0026] G k : Ψ×Φ→Ψ

[0027] where Ψ is the data space and Φ is the parameter space of the generation model.

[0028] S2.2. Key question point extraction: Feed the input text sequences of each dimension into G k . After the model deeply analyzes and reasons the semantics, lexical associations, and context logic, finally obtain the key question points of different dimensions:

[0029]

[0030] The parameters of the current model are θ∈Φ, i is the dimension, k is the number of key information in different dimensions, and θ is the parameter for extracting key question points.

[0031] Finally, after manual screening, key question point information of different dimensions is extracted.

[0032] S3. Evaluation data generation. Its underlying architecture is based on the existing large language model. Through the multi-head self-attention mechanism, it processes different representation subspaces of input key information in parallel, thereby achieving comprehensive capture and deep understanding of semantics. In the training phase, large-scale cultural and tourism-related text data is used, covering travel notes, guides, official introductions, and tourist evaluations. Multi-source heterogeneous texts are used to train the model using massive data to learn rich language patterns and cultural and tourism field knowledge. Based on the pre-trained weights of the large language model, the model is fine-tuned on a cultural and tourism-specific data set through transfer learning technology, so that the model can better adapt to the language habits and cultural and tourism scene characteristics of the local domain.

[0033] Specific process:

[0034] S3.1, Prompt construction for evaluation data generation. For a given context C of the generated question, the prompt template for evaluation data generation is P 2 ,for Where λ is the number of key question points, E i For each dimension, Combine it with the given context C according to specific splicing rules, P′ 2 Get Prompt: Will Serves as input information to drive the model to generate evaluation data.

[0035] S3.2, design a template to generate evaluation data, the structure of which is (Q, A, T, L 1 ,L 2 ,L 3 ). Where Q is based on the prompt word P′ 2 The questions generated are related to culture and tourism; A is the corresponding answer to Q generated based on the prompt word; T represents the question type. Due to the particularity of culture and tourism data, the type judgment function f is used according to the pre-set rules. T (Q) Divide the problem into factual (F) and planning (P), that is, T = f T (Q), T∈{F,P}; question correctness L 1 It is an indicator of whether the question generated by the model is the correct question based on the context. If the model determines whether the generated question is the correct question based on the context; the answer matching degree L 2 Used to determine whether the answer matches the question exactly; generate reliability L 3The big model determines whether the questions and answers can be directly extracted from the context.

[0036] S3.3, is the data space, Φ is the parameter space of the generative model, the parameter θ∈Φ of the current model, t represents the time step, i represents the i-th iteration at the current time step, G d Data generates a language model, then the initial generated data is represented as:

[0037]

[0038] in is the evaluation data generated at the i-th iteration of time step t, L i ∈{0,1}. The indicators and evaluation data are generated simultaneously by the language model.

[0039] S3.4. Evaluation index L for generated data 1 ,L 2 ,L 3 Perform round-by-round iterative calculations and adjustments. The iterative process depends on the generation capability of the language model. The diversity of generation ensures that the model can gradually generate data that meets the evaluation criteria through multiple attempts. The convergence assumption is that within a finite number of generations, the language model can generate data such that L 1 =1,L 2 =1,L 3 =1, then the generated data X t When the expected quality standard is reached, the iteration process stops and the final generated dataset X is output. * :

[0040] X * =X t , where L 1 =L 2 =L 3 =1

[0041] If in a round t there is any indicator L i =0,(i∈{1,2,3}), it indicates that the generated data does not meet the quality requirements. In this case, it is necessary to iteratively generate data based on the evaluation results, that is, perform i+1 iterations:

[0042]

[0043] S3.5. Manual evaluation: After the evaluation data is generated, the accuracy of the generated data is verified from the language level. For each piece of data generated:

[0044] X=(Q,A,T,L i )

[0045] Li ∈{L 1 ,L 2 ,L 3}

[0046] R i ∈{R 1 ,R 2 ,R 3}

[0047] The verification process includes two stages: automatic model determination L and manual review R. The probability of the two indicators being equal in all generated data is calculated, and from the perspective of language, the data generated by the model is analyzed to see whether it meets the structural integrity and semantics of human language, thereby comprehensively verifying the reliability and credibility of the evaluation system. The consistency index is expressed as:

[0048]

[0049] Judgment result L i With R i Indicates the i-th quality indicator of the final evaluation data currently generated. Manual review requires indicators L that are automatically determined by the model i Perform manual evaluation. Based on the automatic judgment results, add manual review results R i Verify the generated questions and answers R i ∈{0,1}. 1 =1 means the manual verification question is qualified, R 2 =1 means the answer is qualified by manual verification, R 3 =1 indicates that the answer can be extracted from the context. By calculating the consistency index of L and R, that is, the probability of the two indexes being equal in all generated data is greater than 90%, the generated data is considered to be reliable.

[0050] Compared with the prior art, the present invention has significant advantages. Through the collaborative work of key question point extraction and evaluation data generation, the core issues that tourists care about can be accurately refined, and higher quality, more targeted and diverse evaluation data can be generated. This method not only improves the accuracy of data generation, but also better reflects the actual needs of tourists and shows higher flexibility in dealing with the impact of emergencies. The use of deep learning to optimize the data generation process improves computing efficiency and reduces costs, and is suitable for all types of scenic spots. Overall, the present invention provides an efficient, accurate and sustainable evaluation data generation solution for the cultural and tourism industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is a flow chart of the method for generating question-answer evaluation data in the field of culture and tourism of the present invention;

[0052] Figure 2Flowchart for constructing the initial data of Emei Mountain in the embodiment. Detailed implementation manners

[0053] The specific technical solutions of the present invention will be described in conjunction with the accompanying drawings.

[0054] The present invention proposes a method for generating cultural and tourism evaluation data based on multi-stage processing. The purpose of this method is to ensure that the generated evaluation data can comprehensively cover all dimensions of cultural and tourism, with high accuracy and reliability, truly and objectively reflect the actual situation, so as to accurately meet the urgent need of the cultural and tourism industry for high-quality evaluation data.

[0055] Corresponding to Figure 1 , to achieve this goal, the present invention reconstructs the traditional method for generating evaluation data of large models and divides it into three core components, each of which undertakes a key function:

[0056] S1. Data collection. First, collect cultural and tourism-related data from multiple sources, including official tourism documents, tourist evaluations, travel blogs, online travel Q&A platforms, social media data, etc., to form the original data. Corresponding to Figure 2 , these data sources are diverse and heterogeneous, and can provide multi-angle cultural and tourism information.

[0057] Specific process:

[0058] S1.1. Data cleaning and preprocessing: The collected cultural and tourism data often has inconsistent formats, redundant information, some irrelevant content, and noise information. The system will adopt efficient data cleaning techniques, such as deduplication, standardization, text correction, etc., to ensure the high quality and consistency of the data.

[0059] S1.2. Dimension division and classification: Ensure that the data in each dimension can accurately reflect its characteristics, and divide the cleaned data according to the six core dimensions of the cultural and tourism industry: "Eat (E 1 )", "Live (E 2 )", "Travel (E 3 )", "Tour (E 4 )", "Shop (E 5 )", "Entertainment (E 6 )", which is expressed as:

[0060]

[0061] Among them, each dimension, is all the cultural and tourism information sets under this dimension, such as the catering information in the eat dimension, the accommodation information in the live dimension, etc. Each dimension can be represented as a set containing multiple items of data

[0062] S2. Key question point extraction method. Accurately extract key question points closely related to cultural tourism from massive cultural tourism information. By using natural language processing technology, this extraction method can deeply explore various aspects that tourists may be concerned about during the cultural tourism process, including but not limited to the characteristics of cultural tourism attractions, catering and accommodation, transportation convenience, and service quality. These key question points will serve as an important guide for subsequent data generation to ensure that the generated data is highly targeted and practical, and can truly reflect the real needs and concerns of tourists.

[0063] Specific process:

[0064] S2.1. Extraction Prompt Design: Design a Prompt template P specifically for different dimensions 1 , and in P 1 A one shot example is given in the figure. The cleaned and normalized data is embedded in this prompt to form a complete model input sequence P′ that meets the model input requirements. 1 The key question point extraction language model is expressed as:

[0065] G k :Ψ×Φ→Ψ

[0066] is the data space, and Φ is the parameter space of the generative model.

[0067] S2.2. Extraction of key question points: The input text sequences of each dimension are sent to G k In the process, the model conducts in-depth analysis and reasoning on semantics, vocabulary associations, and contextual logic, and finally obtains key question points of different dimensions:

[0068]

[0069] The parameters of the current model are θ∈Φ, i is the dimension, k is the number of key information in different dimensions, and θ is the parameter for extracting key question points.

[0070] Finally, after manual screening, the key question points in different dimensions were extracted as follows: There are 102 items in total, covering special cuisine, restaurant services, etc.; "Accommodation" dimension A total of 89, covering accommodation types, guest experience, etc.; "travel" dimension There are 49 in total, including information on public transportation, self-driving tours, etc.; "Travel" dimension A total of 456, covering scenic spot culture, travel guides, etc.; "Purchase" dimension A total of 48, focusing on special products, shopping places; "Entertainment" dimension There are 20 in total, related to entertainment activities and leisure venues, which lay the foundation for the subsequent generation of evaluation data.

[0071] S3. Evaluation data generation. Its underlying architecture is based on the existing large language model. Through the multi-head self-attention mechanism (Multi-Head Attention), it can process different representation subspaces of input key information in parallel, thereby achieving comprehensive capture and deep understanding of semantics. In the training phase, large-scale cultural and tourism-related text data was used, covering multi-source heterogeneous texts such as travel notes, guides, official introductions, and tourist reviews. The model was trained using massive data to learn rich language patterns and cultural and tourism field knowledge. Based on the pre-trained weights of the large language model, the model is fine-tuned on a cultural and tourism-specific data set through transfer learning technology, so that the model can better adapt to the language habits and cultural and tourism scene characteristics of the local domain.

[0072] Specific process:

[0073] S3.1, Prompt construction for evaluation data generation. For a given context C of the generated question, the prompt template for evaluation data generation is P 2 ,for Where λ is the number of key question points, E i For each dimension, Combine it with the given context C according to specific splicing rules, P′ 2 Get Prompt: Will Serves as input information to drive the model to generate evaluation data.

[0074] S3.2, design a template to generate evaluation data, the structure of which is (Q, A, T, L 1 ,L 2 ,L 3 ). Where Q is based on the prompt word P′ 2 The questions generated are related to culture and tourism; A is the corresponding answer to Q generated based on the prompt word; T represents the question type. Due to the particularity of culture and tourism data, the type judgment function f is used according to the pre-set rules. T (Q) Divide the problem into factual (F) and planning (P), that is, T = f T (Q), T∈{F,P}; question correctness L 1 It is an indicator of whether the question generated by the model is the correct question based on the context. If the model determines whether the generated question is the correct question based on the context; the answer matching degree L 2 Used to determine whether the answer matches the question exactly; generate reliability L 3 The big model determines whether the questions and answers can be directly extracted from the context.

[0075] S3.3, is the data space, Φ is the parameter space of the generative model, the parameter θ∈Φ of the current model, t represents the time step, i represents the i-th iteration at the current time step, G d Data generates a language model, then the initial generated data is represented as:

[0076]

[0077] in is the evaluation data generated at the i-th iteration of time step t, L i ∈{0,1}. The indicators and evaluation data are generated simultaneously by the language model.

[0078] S3.4. In order to ensure the quality and accuracy of the generated data, the evaluation index L 1 ,L 2 ,L 3 Perform round-by-round iterative calculations and adjustments. The iterative process depends on the generation capability of the language model. The diversity of generation ensures that the model can gradually generate data that meets the evaluation criteria through multiple attempts. The convergence assumption is that within a finite number of generations, the language model can generate data such that L 1 =1,L 2 =1,L 3 =1, then the generated data X t When the expected quality standard is reached, the iteration process stops and the final generated dataset X is output. * :

[0079] X * =X t , where L 1 =L 2 =L 3 =1

[0080] If in a round t there is any indicator L i =0,(i∈{1,2,3}), it indicates that the generated data does not meet the quality requirements. In this case, it is necessary to iteratively generate data based on the evaluation results, that is, perform i+1 iterations:

[0081]

[0082] S3.5. Manual evaluation: After the evaluation data is generated, in order to further verify the accuracy of the generated data from the language level, for each piece of data generated:

[0083] X=(Q,A,T,L i )

[0084] L i ∈{L 1 , L 2 , L 3}

[0085] R i ∈{R 1 , R 2 , R 3}

[0086] The verification process includes two stages: automatic model determination L and manual review R. The probability of the two indicators being equal in all generated data is calculated, and from the perspective of language, the data generated by the model is analyzed to see whether it meets the structural integrity and semantics of human language, thereby comprehensively verifying the reliability and credibility of the evaluation system. The consistency index is expressed as:

[0087]

[0088] Judgment result L i With R i Indicates the i-th quality indicator of the final evaluation data currently generated. Manual review requires indicators L that are automatically determined by the model i Perform manual evaluation. Based on the automatic judgment results, add manual review results R i Verify the generated questions and answers R i ∈{0,1}. 1 =1 means the manual verification question is qualified, R 2 =1 means the answer is qualified by manual verification, R 3 =1 indicates that the answer can be extracted from the context. By calculating the consistency index of L and R, that is, the probability of the two indexes being equal in all generated data is greater than 90%, the generated data is considered to be reliable.

[0089] The present invention proposes a phased data generation method, which combines a large language model, a key question point extractor and an indicator optimization mechanism to form a complete and efficient evaluation data generation system from data extraction to generation to verification.

[0090] We explore key question points from the six core dimensions of "eating, accommodation, transportation, sightseeing, shopping and entertainment" to accurately locate tourists' concerns, and use this as the core guiding information for data generation to ensure the pertinence and coverage of evaluation data from the source.

[0091] By utilizing the powerful text generation capabilities of the large language model and taking these key questions as a guide, we can create content point by point. This ensures that the generated data is closely centered around the core concerns of tourists and avoids generating broad, unfocused, redundant information in terms of content details.

[0092] The present invention introduces L 1 , L 2 , L 3Three quality evaluation indicators are used to comprehensively evaluate the quality of generated data from data structure, semantic rationality to usability, and the results are optimized through an iterative generation mechanism.

[0093] During the quality assessment process, we innovatively combine automatic model judgment and manual review to verify the language structure integrity and semantic rationality of the generated data from a human language perspective, greatly improving the reliability of the evaluation system.

Claims

1. A method for generating cultural tourism evaluation data based on multi-stage processing, characterized in that: Including data collection, extraction of key question points, and generation of evaluation data; Key question point extraction generates targeted questions from multiple dimensions by extracting key information that tourists are concerned about, providing a clear framework for data generation; Evaluation data generation is based on large-scale cultural and tourism-related text data. The data generation process is optimized through deep learning and transfer learning technology to ensure that the output data is closer to the actual needs of tourists. At the same time, by introducing iterative optimization and human-computer interaction mechanisms, we can dynamically adapt to changing market demands and user feedback to ensure that the quality and diversity of data generation are effectively balanced.

2. According to the method for generating cultural tourism evaluation data based on multi-stage processing according to claim 1, it is characterized in that: The specific steps include: S1. Data collection: First, collect cultural and tourism-related data from multiple sources, including official tourism documents, tourist reviews, travel blogs, online tourism Q&A platforms, and social media data to form the original data; S2. Extraction of key question points: accurately extract key question points closely related to cultural tourism from massive cultural tourism information; by using natural language processing technology, deeply explore various aspects that tourists may be concerned about during the cultural tourism process. These key question points will serve as important guidance for subsequent data generation, ensuring that the generated data is highly targeted and practical, and can truly reflect the real needs and concerns of tourists; S3, evaluation data generation; its underlying architecture is based on the existing large language model, and through the multi-head self-attention mechanism, it processes the different representation subspaces of the input key information in parallel, so as to achieve comprehensive capture and deep understanding of semantics; in the training stage, it uses large-scale cultural and tourism-related text data, covering travel notes, guides, official introductions, and tourists' evaluations. Multi-source heterogeneous texts are used to train the model using massive data to learn rich language patterns and cultural and tourism field knowledge; Based on the pre-trained weights of the large language model, fine-tuning is performed on a specific cultural and tourism dataset through transfer learning technology, so that the model can better adapt to the language habits and cultural and tourism scene characteristics of the local domain.

3. According to the method for generating cultural tourism evaluation data based on multi-stage processing according to claim 2, it is characterized in that: S1 specifically includes the following sub-steps: S1.

1. Data cleaning and preprocessing: Use efficient data cleaning technology to ensure high quality and consistency of data; S1.2, Dimension division and classification: Ensure that the data of each dimension can accurately reflect its characteristics, The cleaned data is expressed as follows according to the six core dimensions of the cultural tourism industry: "eating (E1)", "accommodation (E2)", "transportation (E3)", "travel (E4)", "shopping (E5)", and "entertainment (E6)": In each dimension, It is a set of all cultural and tourism information under this dimension; each dimension is represented as a set containing multiple data.

4. According to the method for generating cultural tourism evaluation data based on multi-stage processing according to claim 3, it is characterized in that: S2 specifically includes the following sub-steps: S2.

1. Extraction prompt design: Design a prompt template P1 specifically for different dimensions, and give a oneshot example in P1. The cleaned and normalized data is embedded in this prompt to form a complete model input sequence P′1 that meets the model input requirements; the key question point extraction language model is expressed as: G k :Ψ×Φ→Ψ is the data space, Φ is the parameter space of the generative model; S2.

2. Extraction of key question points: The input text sequences of each dimension are sent to G k In the process, the model conducts in-depth analysis and reasoning on semantics, vocabulary associations, and contextual logic, and finally obtains key question points of different dimensions: The parameters of the current model are θ∈Φ, i is the dimension, k is the number of key information in different dimensions, and θ is the parameter for extracting key question points; Finally, after manual screening, key question point information of different dimensions is extracted.

5. A method for generating cultural tourism evaluation data based on multi-stage processing according to claim 4, characterized in that: S3 specifically includes the following sub-steps: S3.1, evaluation data generation prompt construction; for a given generation question context C, the prompt template for evaluation data generation is P2. Where λ is the number of key question points, E i For each dimension, Combining it with the given context C according to specific splicing rules, P'2 obtains Prompt: Will Serves as input information to drive the model to generate evaluation data; S3.

2. Design a template for generating evaluation data. Its structure is (Q, A, T, L1, L2, L3). Q is a question related to culture and tourism generated based on the prompt word P'2. A is the answer to Q generated based on the prompt word. T represents the type of question. Due to the particularity of culture and tourism data, the type judgment function f is used according to the pre-set rules. T (Q) Divide the problem into factual (F) and planning (P), that is, T = f T (Q), T∈{F,P}; question correctness L1 is an indicator of whether the generated question is a correct question based on the context. If the model determines whether the generated question is a correct question based on the context; answer matching L2 is used to determine whether the answer accurately matches the question; generation reliability L3 is determined by the large model whether the question and answer can be directly extracted from the context; S3.3, is the data space, Φ is the parameter space of the generative model, the parameter θ∈Φ of the current model, t represents the time step, i represents the i-th iteration at the current time step, G d Data generates a language model, then the initial generated data is represented as: G d : in is the evaluation data generated at the i-th iteration of time step t, L i ∈{0,1}; the indicators and evaluation data are generated simultaneously by the language model; S3.

4. Perform iterative calculation and adjustment on the evaluation indicators L1, L2, and L3 of the generated data. The iterative process depends on the generation capability of the language model. The diversity of generation ensures that the model can gradually generate data that meets the evaluation criteria through multiple attempts. The convergence assumption is that within a finite number of generations, the language model can generate data such that L1 = 1, L2 = 1, and L3 = 1, then the generated data X t When the expected quality standard is reached, the iteration process stops and the final generated dataset X is output. * : X * =X t , where L1=L2=L3=1 If in a round t there is any indicator L i =0,(i∈{1,2,3}), it indicates that the generated data does not meet the quality requirements. In this case, it is necessary to iteratively generate data based on the evaluation results, that is, perform i+1 iterations: S3.

5. Manual evaluation: After the evaluation data is generated, the accuracy of the generated data is verified from the language level. For each piece of data generated: X=(Q,A,T,L i ) <h2 style=";text-align:left;direction:ltr">L<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> ∈{L1,L2,L3} R i ∈{R1,R2,R3} The verification process includes two stages: automatic model determination L and manual review R; calculating the probability of the two indicators being equal in all generated data, and analyzing from a linguistic perspective whether the data generated by the model meets the structural integrity and semantics of human language, thereby comprehensively verifying the reliability and credibility of the evaluation system. The consistency index is expressed as: Judgment result L i With R i Indicates the i-th quality indicator of the final evaluation data currently generated; manual review requires the indicator L automatically determined by the model i Conduct manual evaluation; On the basis of automatic judgment results, add manual review results R i Verify the generated questions and answers R i ∈{0,1}; R1=1 indicates that the manual verification question is qualified, R2=1 indicates that the manual verification answer is qualified, and R3=1 indicates that the answer can be extracted from the context; by calculating the consistency index of L and R, that is, the probability that the two indicators are equal in all generated data is greater than 90%, the generated data is considered to be reliable.

Citation Information

Patent Citations

  • Hotel intelligent question and answer recommendation and decision support analysis method and system

    CN110807091A

  • Question and answer system based on text travel industry combination algorithm

    CN114297362A

  • Question and answer method, electronic equipment and computer storage medium

    CN117851556A

  • Deep reading method and system based on knowledge graph

    CN118193748A

  • Power standard question and answer model training method and system based on retrieval enhancement

    CN118733706A