Tourism recommendation method and system based on big data
Through a big data-based method, user behavior and social media data are collected to cluster points of interest, calculate destination fit and timeliness needs, and adjust the priority of travel recommendations, solving the problem of inability to adapt to short-term changes in users in the existing technology, and improving the accuracy and user experience of travel recommendations.
Patent Information
- Application Number
- CN202510549429.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology cannot effectively identify and adapt to short-term changes in user needs, especially during peak tourism or seasonal changes, resulting in inaccurate recommendation information, unable to reflect the latest tourism trends and user feedback, affecting user experience and tourism service quality.
Through a big data-based method, user behavior, social media interaction and search query data are collected, points of interest are clustered and classified, destination fit is calculated, timeliness needs and demand changes are analyzed, tourism recommendation priority is adjusted, and recommendation results are optimized in combination with real-time data.
It realizes the capture of users' immediate interests and reflects instant changes in travel trends, optimizes recommendation results, and improves the accuracy and user satisfaction of personalized recommendations.
Smart Images

Figure CN120296263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of travel recommendation, and particularly to a travel recommendation method and system based on big data. Background Art
[0002] Travel recommendation technology involves applying various computing technologies and algorithms to analyze user data, aiming to provide personalized travel suggestions and experiences. This field combines technologies such as big data analysis, machine learning, artificial intelligence, and geographic information systems to mine data such as user preferences, historical behaviors, and external conditions. The recommendation system can predict new destinations, travel times, and travel activities that a user may be interested in based on information such as the user's travel history, search habits, and social media activities.
[0003] Among them, the big data-based travel recommendation method designs and implements a travel recommendation system by collecting and analyzing large-scale user behavior data, destination information, and user feedback, uses algorithm models to identify potential travel preferences, and recommends personalized travel options and activities to users. Its purpose is to help users discover travel experiences that best suit their interests and budgets, while providing precise marketing tools for travel enterprises to more effectively attract and serve the target customer group.
[0004] The existing technology relies on static data analysis and lacks an immediate response to the dynamic changes in user behaviors and preferences, resulting in the inability to effectively identify and adapt to short-term changes in user needs in practical applications. Especially during peak travel seasons or seasonal changes, information cannot be updated in a timely manner, leading to inaccurate recommendation information. In addition, the existing technology fails to make full use of real-time data, making it impossible to reflect the latest travel trends and user feedback information during the selection and recommendation process of travel destinations, affecting other users' experiences and limiting the effectiveness of travel recommendations in dealing with emergencies and seasonal demand fluctuations, thereby reducing the quality of travel services and the adaptability of the travel market. Summary of the Invention
[0005] The purpose of the present invention is to solve the deficiencies existing in the prior art and propose a travel recommendation method and system based on big data.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A big data-based travel recommendation method, including the following steps,
[0007] S1: Based on user behavior data, social media interaction data, and search query data, collect destination information accessed by users from multiple data sources, analyze the user's browsing history, search frequency, and geographical location preferences, cluster the user's interest points, and divide the destination preferences according to the clustering results to obtain user interest classification values;
[0008] S2: Based on the user interest classification value, search for travel destinations that match the user's interest points, calculate the degree of fit between each destination and the user's preferences, screen travel destinations that meet the user's needs, determine the travel recommendation priority, and obtain the destination matching score;
[0009] S3: Based on the destination matching score, analyze historical travel data and seasonal trends, determine the timeliness requirements of the user for the destination, identify the time characteristics of the user's behavior, perform time window partitioning, and predict the demand changes of future popular travel destinations to obtain destination demand change information;
[0010] S4: Based on the destination demand change information, according to the real-time user needs and current travel trends, collect time nodes associated with travel needs, perform demand fluctuation analysis, and adjust the travel recommendation priority to optimize the ranking of destinations to obtain travel activity recommendation results.
[0011] The improvement of the present invention is that the step of clustering the user's interest points is specifically as follows:
[0012] S111: Based on user behavior data, social media interaction data, and search query data, collect destination information visited by the user from multiple data sources, analyze the user's browsing history, search frequency, and geographical location preferences to obtain a user interest point dataset;
[0013] S112: Based on the user interest point dataset, use the formula:
[0014]
[0015] Calculate the weighted distance between interest points to cluster the user's interest points to obtain a clustering result, where D ij represents the weighted distance between interest point i and interest point j, x i and x j represent the longitudes in the geographical location coordinates of interest point i and interest point j, y i and y j represent the latitudes in the geographical location coordinates of interest point i and interest point j, f i and f j represent the access frequency information between interest point i and interest point j, w1 is the weight of the longitude difference in the geographical coordinates, w2 is the weight of the latitude difference in the geographical coordinates, and w3 is the weight of the access frequency information difference.
[0016] The improvement of the present invention is that the step of obtaining the user interest classification value is specifically as follows:
[0017] S121: Based on the clustering result, analyze the characteristics of each clustering cluster, calculate the similarity between the user's interest points and the clustering cluster characteristics, and use the formula:
[0018]
[0019] Obtain the similarity between the user's interest points and the clustering clusters, where S iz represents the similarity between user interest point i and clustering cluster z, and D ij represents the weighted distance between interest point i and interest point j, and α is an adjustment coefficient used to control the influence degree of the distance on the similarity;
[0020] S122: Based on the similarity between the user's interest points and the clustering clusters, assign each user to the clustering cluster that matches their interest points, perform destination preference division, and obtain the user interest classification value.
[0021] The improvement of the present invention is that the calculation steps of the matching degree between the destination and the user preference are specifically as follows:
[0022] S211: Based on the user interest classification value, screen the tourism destinations, match the destination list that matches the user's interest points, and obtain the destination matching list;
[0023] S212: Based on the destination matching list, extract the feature data of the destination, including geographical location, cultural characteristics, and activity type, compare it with the user interest classification value, and use the formula:
[0024] MU = ∑(ws i ×vs ui ×vs di );
[0025] Obtain the matching degree MU between each destination and the user preference, where ws i represents the weight of the i-th feature, vs ui represents the i-th feature value in the user interest vector, and vs di represents the i-th feature value in the destination feature vector.
[0026] The improvement of the present invention is that the obtaining steps of the destination matching degree score are specifically as follows:
[0027] S221: Based on the matching degree between each destination and the user preference, screen the tourism destinations that meet the user's needs, and obtain the destination screening information;
[0028] S222: Based on the destination screening information, calculate the priority score of each destination according to the popularity and seasonal factors of the destination, and sort the destinations to determine the tourism recommendation priority, and obtain the destination matching degree score.
[0029] The improvement of the present invention is that the determination steps of the user's timeliness requirement for the destination are specifically as follows:
[0030] S311: Based on the destination matching degree score, collect the user's historical travel data, including the destination visit date and the stay duration. For each user and destination, count the number of visits and the stay duration within different seasons to obtain the seasonal user travel data;
[0031] S312: Based on the seasonal user travel data, according to the local festivals and climate conditions of each destination, use the formula:
[0032]
[0033] Calculate the timeliness demand index, and identify the peak and trough periods of user demand. Among them, PG ui represents the timeliness demand index of user u for destination i, VG uk is the number of visits of user u to the destination in season k, SG dk is the characteristic score of the destination in season k, N G is the number of seasons.
[0034] The improvement of the present invention is that the step of obtaining the destination demand change information is specifically:
[0035] S321: Based on the timeliness demand, analyze the user's travel date, destination, and travel duration, and identify the user behavior characteristics to obtain the user behavior data set;
[0036] S322: Based on the user behavior data set, divide the time window according to the destination demand and seasonal fluctuations, and predict the demand change of future popular tourist destinations to obtain the destination demand change information.
[0037] The improvement of the present invention is that the step of obtaining the travel activity recommendation result is specifically:
[0038] S411: Based on the destination demand change information, collect the real-time user demand and the current travel trend data, integrate the user's search and reservation information, and identify the time nodes associated with the travel demand to obtain the demand time node analysis result;
[0039] S412: Based on the demand time node analysis result, conduct demand fluctuation analysis, extract the key demand change trends, and use the formula:
[0040]
[0041] Calculate the demand fluctuation value VB of each time node t , where PB ti represents the demand intensity of the i-th time node, TB ti represents the time impact coefficient of the i-th time node, n Yis the total number of time nodes;
[0042] S413: Based on the demand fluctuation value, adjust the tourism recommendation priority, and apply a sorting algorithm to optimize the priority sorting of destinations to obtain the tourism activity recommendation result.
[0043] A tourism recommendation system based on big data, the system includes:
[0044] The user interest clustering module analyzes the user's browsing history, search frequency, and geographical location preferences based on user behavior data, social media interaction data, and search query data, clusters the user's interest points, and divides the destination preferences according to the clustering results to obtain the user interest classification value;
[0045] The tourism matching scoring module searches for tourism destinations that match the user's interest points based on the user interest classification value, calculates the degree of fit between each destination and the user's preferences, and screens the tourism destinations that meet the user's needs to obtain the destination matching degree score;
[0046] The tourism trend analysis module analyzes the historical tourism data and seasonal trends based on the destination matching degree score, determines the user's timeliness demand for the destination, divides the time window, and predicts the demand change of future popular tourist destinations to obtain the destination demand change information;
[0047] The tourism recommendation optimization module conducts demand fluctuation analysis based on the destination demand change information, according to the real-time user demand and the current tourism trend, and adjusts the tourism recommendation priority to obtain the tourism activity recommendation result.
[0048] Compared with the prior art, the advantages and positive effects of the present invention are:
[0049] In the present invention, by deeply integrating the user's social media behavior, search data, and geographical location information, the dynamic clustering and classification of tourism destination preferences are realized, the ability to capture the user's immediate interests is improved, and it is ensured that the recommended tourism options are highly consistent with the user's current preferences. By analyzing the historical tourism data and seasonal trends, the recommended destinations can be adjusted predictably to reflect the immediate changes in tourism trends, thereby optimizing the user experience and increasing the satisfaction rate. By further combining the real-time tourism data, the recommendation results not only reflect the historical behavior but also adapt to the immediate tourism demand and personal preferences, thus improving the accuracy of personalized tourism recommendations. Brief Description of the Drawings
[0050] Figure 1 is the flowchart of the tourism recommendation method based on big data proposed by the present invention;
[0051] Figure 2 is the flowchart of clustering the user's interest points in the present invention;
[0052] Figure 3 This is the flowchart for obtaining the user interest classification value in the present invention;
[0053] Figure 4 This is the flowchart for calculating the degree of fit between the destination and the user preferences in the present invention;
[0054] Figure 5 This is the flowchart for obtaining the destination matching degree score in the present invention;
[0055] Figure 6 This is the flowchart for determining the timeliness requirement of the user for the destination in the present invention;
[0056] Figure 7 This is the flowchart for obtaining the destination demand change information in the present invention;
[0057] Figure 8 This is the flowchart for obtaining the tourism activity recommendation result in the present invention. Detailed implementation manners
[0058] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0059] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality of" is two or more unless otherwise specifically defined.
[0060] Embodiment
[0061] Please refer to Figure 1 , the present invention provides a technical solution: a tourism recommendation method based on big data, including the following steps:
[0062] S1: Based on user behavior data, social media interaction data, and search query data, collect destination information accessed by the user from multiple data sources, analyze the user's browsing history, search frequency, and geographical location preferences, cluster the user's interest points, and divide the destination preferences according to the clustering results to obtain the user interest classification value;
[0063] S2: Based on the user interest classification values, search for travel destinations that match the user's interest points, calculate the degree of fit of each destination with the user's preferences, screen travel destinations that meet the user's needs, determine the travel recommendation priority, and obtain the destination matching score.
[0064] S3: Based on the destination matching score, analyze historical travel data and seasonal trends, determine the timeliness requirements of the user for the destination, identify the time characteristics of the user's behavior, perform time window partitioning, predict the demand changes of future popular travel destinations, and obtain the destination demand change information.
[0065] S4: Based on the destination demand change information, according to the real-time user needs and current travel trends, collect time nodes associated with travel needs, conduct demand fluctuation analysis, adjust the travel recommendation priority, optimize the ranking of destinations, and obtain the travel activity recommendation results.
[0066] The user interest classification values specifically include destination type, preference degree, and interest point category. The destination matching score includes the fit index and satisfaction level. The destination demand change information includes the demand growth rate, demand decay rate, and timeliness index. The travel activity recommendation results include the recommendation list and priority index.
[0067] Please refer to Figure 2 , the steps for clustering the user's interest points are specifically as follows:
[0068] S111: Based on the user behavior data, social media interaction data, and search query data, collect the destination information accessed by the user from multiple data sources, analyze the user's browsing history, search frequency, and geographical location preferences, and obtain the user interest point data set.
[0069] Collect the destination information accessed by the user from multiple data sources. First, extract the target locations visited by each user from the user's behavior data, including web browsing, application usage, etc. records. Social media data can be analyzed through the user's sharing, liking, commenting, etc. behaviors to obtain the user's interest in specific locations. Search query data can obtain the user's query preferences in the search engine through the user's search records. Combining multi-dimensional data, further weight and screen the most representative destination information. Then, by analyzing the user's browsing history, track the user's long-term and short-term interest points, calculate the attention frequency of each user to a certain destination in different time periods, and the geographical location preference is obtained through the user's location data, GPS records, or IP address, etc. Conduct spatial clustering on the user's interest points to obtain an interest point data set containing all the destination data accessed by the user.
[0070] S112: Based on the user interest point data set, use the formula:
[0071]
[0072] Calculate the weighted distance between points of interest, cluster the user's points of interest, and obtain the clustering result. Among them, D ij represents the weighted distance between point of interest i and point of interest j, and x i and x j represent the longitudes in the geographical location coordinates of point of interest i and point of interest j, and y i and y j represent the latitudes in the geographical location coordinates of point of interest i and point of interest j, and f i and f j represent the access frequency information of point of interest i and point of interest j. The frequency represents the frequency or intensity of the user's access to this point of interest and is used to measure the importance of the point of interest. w1 is the weight of the longitude difference in the geographical coordinates, w2 is the weight of the latitude difference in the geographical coordinates, and w3 is the weight of the difference in access frequency information;
[0073] Through actual data collection, the geographical location of point of interest i is (34.0522, -118.2437), and the geographical location of point of interest j is (33.9499, -118.3887). That is, the longitude difference is |34.0522 - 33.9499| = 0.1023, and the latitude difference is |-118.2437 - (-118.3887)| = 0.1450. The access frequencies of point of interest i and point of interest j are f i = 10 and f j = 8, that is, the frequency difference is |10 - 8| = 2; Substitute the data into the formula for calculation:
[0074] Assume that the weight coefficients are w1 = 0.5, w2 = 0.3, and w3 = 0.2, then:
[0075]
[0076] The result shows that the weighted distance between point of interest i and point of interest j is 0.8995, indicating the differences in geographical location and access frequency. The result will be used for further clustering analysis to determine whether they belong to the same cluster.
[0077] Please refer to Figure 3 , and the specific steps for obtaining the user interest classification value are as follows:
[0078] S121: Based on the clustering result, analyze the characteristics of each clustering cluster, calculate the similarity between the user's points of interest and the characteristics of the clustering cluster, and use the formula:
[0079]
[0080] Obtain the similarity between the user's points of interest and the clustering cluster. Among them, S izDenotes the similarity between user interest point i and cluster z, D ij Denotes the weighted distance between interest point i and interest point j. α is an adjustment coefficient used to control the influence degree of distance on similarity;
[0081] It is necessary to extract the central features of each cluster from the results of cluster analysis. The central features can be the average geographical location of interest points within the cluster, the weighted average of interest point access frequencies, etc., which are used to represent the "typical" interest points of the cluster. Then, for each user's interest points, calculate their similarity with each cluster. The method is to calculate the weighted distance or similarity between the user's interest points and the cluster center point based on the geographical location and access frequency of the user's interest points. To ensure the accuracy of similarity calculation, it is also necessary to weight the distance and frequency values so that the contributions of distance and frequency to similarity meet the actual application requirements. For example, set different weight parameters to adjust the relative influence of these two items. Finally, by comparing the similarity between each user's interest point and the center points of each cluster, the similarity between the user's interest point and the cluster is obtained. The result can be used for subsequent interest classification and destination preference division. If the weighted distance between interest point i and cluster center point j is D ij = 2.5, and the adjustment coefficient α = 0.8 is set, then the similarity S iz is calculated as follows:
[0082]
[0083] This result indicates that the similarity between interest point i and cluster z is 0.3333, meaning there is a certain degree of difference but not a perfect match. This value will be used for subsequent user interest classification.
[0084] S122: Based on the similarity between the user's interest points and the clusters, assign each user to the cluster that matches their interest points, conduct destination preference division, and obtain the user interest classification value;
[0085] For each user, calculate based on the similarity between their interest points and all clusters, and select the cluster that best matches the user's interest points. Usually, the maximum similarity principle is adopted, that is, select the cluster with the largest similarity value for assignment. If there are multiple clusters with equal similarity to the user's interest points, other criteria can be further introduced for refined selection. For example, by comparing the geographical location difference and access frequency difference between the user's interest points and the cluster. Then, after assigning each user to the corresponding cluster, conduct destination preference division based on the characteristics of the cluster. The user's interest classification value is the label or characteristic value of the cluster. In this way, users can be effectively classified by interest. The interest classification value of each user reflects their preference for a specific destination or type of interest point, and finally forms a set of user interest data.
[0086] Please refer toFigure 4 , the calculation steps for the degree of fit between the destination and the user's preferences are specifically as follows:
[0087] S211: Based on the user interest classification value, screen the tourist destinations, match the destination list that conforms to the user's interest points, and obtain the destination matching list;
[0088] First, extract all the characteristic data of the user interest classification value, including the user's preference tags, historical behavior records, and corresponding interest points. Subsequently, according to the characteristic tags in the user interest classification value, match the tags with the relevant characteristic tags in the tourist destination database. During the matching process, the method of comparing classification tags is adopted to screen the destination entries that meet the tag conditions. At the same time, the correlation score of the classification tags needs to be calculated during the comparison process. When the correlation score is lower than the set matching threshold, the corresponding destination entry is excluded. For the destination entries that meet the correlation score threshold, record them in the preliminary screening result list, and further screen based on the repeatability of the tags between the user interest classification value and the destinations in the preliminary screening results, excluding duplicate or incomplete destination entries. Finally, organize the destination entries that conform to the user interest classification value into a matching list.
[0089] S212: Based on the destination matching list, extract the characteristic data of the destination, including geographical location, cultural characteristics, and activity types, and compare them with the user interest classification value. Use the formula:
[0090] MU = ∑(ws i ×vs ui ×vs di );
[0091] Obtain the degree of fit MU between each destination and the user's preferences. Among them, ws i represents the weight of the i-th feature, expressing the key importance of the feature in the calculation of the degree of fit. vs ui represents the i-th feature value in the user interest vector, which is extracted from the user data and describes the user's preference or interest degree in this feature. vs di represents the i-th feature value in the destination feature vector, which is extracted from the description or data of the destination and describes the performance or attribute of the destination in this feature;
[0092] If there is the following data, ws i is the weight of the i-th feature, and the setting basis is the influence degree of the feature on the matching result. For example, the geographical location weight is 0.4, the cultural characteristics weight is 0.35, and the activity type weight is 0.25;
[0093] vs uiis the i-th eigenvalue in the user's interest vector. The geographic location eigenvalue is extracted from the regional information in the user's historical access data, which is 0.8. The cultural characteristics are extracted from the classification data of the user's preference selection, which is 0.7. The activity type is analyzed from the activity records of the user's participation, which is 0.6.
[0094] vs di is the i-th eigenvalue in the destination feature vector, the geographical location eigenvalue is calculated by the matching degree between the actual location information of the destination and the user's area of interest, which is 0.9, the cultural characteristics are calculated by the similarity between the destination cultural label and the user's preference, which is 0.75, and the activity type is calculated by the comparison between the coverage of the destination activity type and the user's main interest activities, which is 0.65;
[0095] Substitute the above parameters into the formula to calculate:
[0096] MU=(0.4×0.8×0.9)+(0.35×0.7×0.75)+(0.25×0.6×0.65);
[0097] MU = 0.288 + 0.18375 + 0.0975;
[0098] MU = 0.56925;
[0099] The result shows that for the current user interest classification value, the calculated degree of fit between the destination and the user preference is 0.56925, indicating that the destination has a high degree of match with the user's interest and can be used as a recommendation candidate.
[0100] See also Figure 5 , the specific steps for obtaining the destination matching score are:
[0101] S221: based on the degree of fit between each destination and the user's preference, filter the travel destinations that meet the user's needs to obtain destination screening information;
[0102] A mapping relationship between the user interest vector and the destination feature vector is established, the fit of each destination is extracted from the calculation results of the previous stage and filtered according to the threshold, a fit threshold is preset, and the destination entries with a fit greater than or equal to the fit are retained in the filter list. The calculation method of the fit threshold can be determined based on the average and standard deviation of the fit of all destinations. After the fit values are sorted, the filtering range is set by quantile, and then the destinations that meet the filtering conditions are verified one by one, including checking whether there are problems such as data duplication, label errors or missing, and duplicate entries are merged. For entries with missing feature values, the incomplete records are supplemented in the data table or eliminated. After the screening is completed, the destination filtering information that meets the user's needs is obtained.
[0103] S222: Based on the destination screening information, calculate the priority score for each destination according to the popularity and seasonal factors of the destination, sort the destinations, determine the tourism recommendation priority, and obtain the destination matching score;
[0104] Extract the popularity data and season-related data attributes for each destination. The popularity data can be quantified through methods such as the destination's historical visit records, the number of user reviews, and the review scores. The specific method is to take the logarithm normalization of the historical visit volume and calculate the popularity score by weighting according to the review scores. The seasonal factor data is based on the climate and season characteristics of the region where the destination is located, and generates the seasonal score by looking up the peak and off-peak season data within the corresponding time period. When calculating the priority score for each destination, weights are assigned to the popularity and seasonal scores respectively. The specific weight allocation can be determined through regression analysis of multi-year statistical data. The two scores are calculated through the weighted average formula to obtain the comprehensive score, and then sorted from high to low according to the comprehensive score to finally determine the tourism recommendation priority.
[0105] Please refer to Figure 6 , and the steps for determining the user's timeliness demand for the destination are specifically as follows:
[0106] S311: Based on the destination matching score, collect the user's historical travel data, including the destination visit date and stay duration. For each user and destination, count the number of visits and stay duration within different seasons to obtain the seasonal user travel data;
[0107] Analyze the user's historical travel data, including the destination visit date and stay duration. Extract the start time, end time, and total stay duration of each trip from the user's travel schedule. Classify each destination visited by the user into four seasons: spring, summer, autumn, and winter. Count the number of visits within different seasons and calculate the average stay duration of the user within each season. Use the number of visits as the index of the user's visit frequency to the destination in that season, and normalize the stay duration as the representation of the user's seasonal attractiveness to the destination to obtain the user travel data index based on seasonal division.
[0108] S312: Based on the seasonal user travel data, according to the local festivals and climate conditions of each destination, use the formula:
[0109]
[0110] Calculate the timeliness demand index and identify the peak and trough periods of user demand. Among them, PG ui represents the timeliness demand index of user u for destination i, which is used to measure the intensity of the user's interest in the destination in different seasons. VG ukis the number of visits of user u to the destination in season k, which is directly counted from the user's historical travel data, SG dk is the feature score of the destination in season k. The score is comprehensively evaluated based on the seasonal characteristics of the destination in this season, such as climate conditions and festival activities, N G is the number of seasons, including four seasons: spring, summer, autumn, and winter;
[0111] Combined with the local festivals and climate conditions of the destination, count the activities and climate attractiveness of the destination in each season, extract the date, type, and popularity data of the activities, and use climate conditions such as temperature, humidity, and precipitation for quantification. Normalize the popularity of festival activities into numerical data such as the number of participants and activity duration, and calculate the weighted average of each index in the climate conditions to form the comprehensive feature score of the destination in each season. Finally, calculate the peak and trough periods of user demand, and collect the following data. The number of visits of user u to destination i in each season is 10 times in spring, 8 times in summer, 12 times in autumn, and 6 times in winter. The corresponding seasonal feature scores of the destination are 0.8 in spring, 0.9 in summer, 0.7 in autumn, and 0.6 in winter. Then substitute into the formula for calculation:
[0112]
[0113]
[0114] This result shows that the timeliness demand index of user u for destination i is 6.8, which represents the average interest intensity of the user in this destination in different seasons. Combining this index can further identify the peak demand period (for example, the index contribution in spring is higher) and the trough period (for example, the index contribution in winter is lower).
[0115] Please refer to Figure 7 for the specific steps to obtain the destination demand change information:
[0116] S321: Based on the timeliness demand, analyze the user's travel date, destination, and travel duration to identify user behavior characteristics and obtain the user behavior data set;
[0117] Extract the start time and end time of each trip from the user's travel records, count the frequency of the user's visits to destinations and the duration of the trips, further combine the user's historical travel behavior patterns to generate a behavior record table for each of the user's trips, analyze the correspondence between the travel dates and time nodes such as holidays and weekends, classify the user's selection behavior according to the different types of destinations, extract the user's preference characteristics according to the number of visits, stay time, and time period, count the access behavior characteristics of the user to popular destinations in different time periods. In addition, extract the user's search frequency data from travel-related platforms, count the number of times the user queries destinations within different time ranges, combine information such as the actual booking time, destination type, and stay duration in the booking data, calculate the demand fluctuation of the user within each time period, and finally integrate the user behavior time period characteristics and the results of the demand fluctuation analysis to obtain the user behavior data set.
[0118] S322: Based on the user behavior data set, divide time windows according to the destination demand and seasonal fluctuations, and predict the demand changes of future popular tourist destinations to obtain destination demand change information;
[0119] According to the demand intensity and seasonal fluctuation characteristics of the destination, segment the demand change curve, count the access popularity of the destination in different seasons through historical travel data, divide time windows according to the demand peaks and valleys within the time period, analyze the differences in the user's travel behavior within the time windows, combine the user's historical behavior data and the demand fluctuation curve, use statistical analysis methods to calculate the relative demand intensity of the target time period within the time window, further combine the seasonal fluctuation characteristics of the destination and the estimated future external impacts, such as holidays, large-scale events and other factors, comprehensively construct a future demand prediction model, and finally analyze the output results of the user demand data set and the prediction model, integrate the demand change information within the target time window to obtain the demand change results of future popular tourist destinations.
[0120] Please refer to Figure 8 , and the specific steps for obtaining the travel activity recommendation results are as follows:
[0121] S411: Based on the destination demand change information, collect real-time user demand and current travel trend data, integrate the user's search and booking information, identify time nodes associated with travel demand, such as holidays, festivals or special events, to obtain the analysis results of demand time nodes;
[0122] Integrate the user's search and reservation information. Based on the access records, historical browsing behavior, and reservation frequency, clean the data to remove missing values and outliers, extract the user's search popularity, click-through rate, and reservation conversion rate within a specific time period. Map and match the data with destination-related time nodes (such as holidays, major events, etc.), conduct a preliminary classification of the time nodes, divide them into high-demand nodes, ordinary nodes, and low-demand nodes according to the attributes of the nodes, and perform a weighted scoring of the nodes in combination with the demand intensity. Finally, use the weighted average method to calculate the comprehensive demand value of each time node to completely obtain the analysis results of the demand time nodes.
[0123] S412: Based on the analysis results of the demand time nodes, conduct a demand fluctuation analysis, extract the key demand change trends, and use the formula:
[0124]
[0125] Calculate the demand fluctuation value VB of each time node t , where PB ti represents the demand intensity of the i-th time node, that is, the statistical value of the demand volume at this time node, and TB ti represents the time impact coefficient of the i-th time node, which is used to measure the influence proportion of different time nodes in the total demand fluctuation, and n Y is the total number of time nodes;
[0126] According to the analysis results of the demand time nodes, extract the demand intensity PB within a specific time period ti . The demand intensities of users at five time nodes before and after the Spring Festival collected from the monitored tourism platform data are 3000, 4000, 5000, 4500, and 3500 (unit: number of visits) respectively. The time impact coefficient TB ti is determined by the social activity importance of the time node and the demand increase and decrease rate of the adjacent dates of the node. If the time impact coefficients of the five time nodes are 0.9, 1.0, 1.1, 1.0, and 0.8 respectively, and the total number of time nodes n Y = 5, substitute the above data into the formula:
[0127]
[0128]
[0129] The result shows that the total demand fluctuation value VB t = 3900, which can reflect the average level of the comprehensive demand value of each time node within the selected time interval and serve as the basic data for subsequent demand volatility analysis.
[0130] S413: Based on the demand fluctuation value, adjust the tourism recommendation priority, apply a sorting algorithm to optimize the priority ranking of destinations, and obtain the tourism activity recommendation result;
[0131] Using the demand fluctuation value as an input variable, optimize the tourism recommendation priority through a priority adjustment algorithm, re-rank the priority of recommended destinations using breadth-first search. The breadth-first search starts from the demand fluctuation value, combines multi-dimensional features such as the historical click-through rate and user praise rate of tourism destinations to construct a weighted adjacency matrix, processes the priority scores of each destination in turn, sorts the destinations with the highest priority scores in sequence, and updates the tourism activity recommendation list based on this to obtain the tourism activity recommendation result.
[0132] A tourism recommendation system based on big data, the system includes:
[0133] The user interest clustering module analyzes the user's browsing history, search frequency, and geographical location preferences based on user behavior data, social media interaction data, and search query data, clusters the user's interest points, and divides the destination preferences according to the clustering results to obtain the user interest classification value;
[0134] The tourism matching scoring module searches for tourism destinations that match the user's interest points based on the user interest classification value, calculates the degree of fit between each destination and the user's preferences, and filters out the tourism destinations that meet the user's needs to obtain the destination matching score;
[0135] The tourism trend analysis module analyzes historical tourism data and seasonal trends based on the destination matching score, determines the user's timeliness demand for destinations, conducts time window division, and predicts the demand changes of future popular tourist destinations to obtain the destination demand change information;
[0136] The tourism recommendation optimization module conducts demand fluctuation analysis based on the destination demand change information, adjusts the tourism recommendation priority according to the real-time user demand and the current tourism trend, and obtains the tourism activity recommendation result.
[0137] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A tourism recommendation method based on big data, characterized in that, Including the following steps: S1: Based on user behavior data, social media interaction data, and search query data, collect destination information visited by users from multiple data sources, analyze user browsing history, search frequency, and geographical location preferences, cluster user interest points, and divide destination preferences according to the clustering results to obtain user interest classification values; S2: Based on the user interest classification values, search for travel destinations that match the user interest points, calculate the degree of fit between each destination and the user preferences, screen travel destinations that meet the user needs, determine the travel recommendation priority, and obtain the destination matching degree score; S3: Based on the destination matching degree score, analyze historical travel data and seasonal trends, determine the timeliness requirements of the user for the destination, identify the time characteristics of user behavior, conduct time window division, and predict the demand changes of future popular travel destinations to obtain destination demand change information; S4: Based on the destination demand change information, according to the real-time user needs and current travel trends, collect time nodes associated with travel demands, conduct demand fluctuation analysis, and adjust the travel recommendation priority, optimize the sorting of destinations, and obtain travel activity recommendation results.
2. The tourism recommendation method based on big data according to claim 1, wherein The step of clustering user interest points is specifically as follows: S111: Based on user behavior data, social media interaction data, and search query data, collect destination information visited by users from multiple data sources, analyze user browsing history, search frequency, and geographical location preferences to obtain a user interest point data set; S112: Based on the user interest point data set, use the formula: Calculate the weighted distance between interest points, cluster the user interest points, and obtain the clustering result, where D ij represents the weighted distance between interest point i and interest point j, x i and x j represent the longitudes in the geographical location coordinates of interest point i and interest point j, y i and y j represent the latitudes in the geographical location coordinates of interest point i and interest point j, f i and f j represent the access frequency information of interest point i and interest point j, w1 is the weight of the longitude difference in the geographical coordinates, w2 is the weight of the latitude difference in the geographical coordinates, and w3 is the weight of the access frequency information difference.
3. The tourism recommendation method based on big data according to claim 1, wherein, The step of obtaining the user interest classification value is specifically as follows: S121: Based on the clustering results, analyze the characteristics of each clustering cluster, calculate the similarity between the user interest points and the clustering cluster characteristics, and use the formula: Obtain the similarity between the user's point of interest and the clustering cluster, where S iz represents the similarity between the user's point of interest i and the clustering cluster z, D ij represents the weighted distance between point of interest i and point of interest j, and α is an adjustment coefficient used to control the influence degree of the distance on the similarity; S122: Based on the similarity between the user interest points and the clustering cluster, assign each user to the clustering cluster that matches their interest points, conduct destination preference division, and obtain the user interest classification value.
4. The tourism recommendation method based on big data according to claim 1, characterized in that The step of calculating the degree of fit between the destination and the user preferences is specifically as follows: S211: Based on the user interest classification values, screen travel destinations, match the destination list that matches the user interest points, and obtain the destination matching list; S212: Based on the destination matching list, extract the characteristic data of the destination, including geographical location, cultural characteristics, and activity types, and compare them with the user interest classification values, using the formula: MU = ∑(ws i × vs ui × vs di ); Obtain the matching degree MU between each destination and the user's preference, where ws i represents the weight of the i-th feature, vs ui represents the i-th eigenvalue in the user interest vector, vs di represents the i-th eigenvalue in the destination feature vector.
5. The tourism recommendation method based on big data according to claim 1, wherein The step of obtaining the destination matching degree score is specifically as follows: S221: Based on the degree of fit between each destination and the user preferences, screen travel destinations that meet the user needs to obtain destination screening information; S222: Based on the destination screening information, calculate the priority score of each destination according to the popularity of the destination and seasonal factors, sort the destinations, determine the travel recommendation priority, and obtain the destination matching degree score.
6. The tourism recommendation method based on big data according to claim 1, characterized in that, The step of determining the timeliness requirements of the user for the destination is specifically as follows: S311: Based on the destination matching score, collect historical user travel data, including the destination visit date and stay duration. For each user and destination, count the number of visits and stay durations within different seasons to obtain seasonal user travel data; S312: Based on the seasonal user travel data, according to the local festivals and climate conditions of each destination, use the formula: Calculate the timeliness demand index, and identify the peak and trough periods of user demand, where PG ui represents the timeliness demand index of user u for destination i, VG uk is the number of visits of user u to the destination in season k, SG dk is the feature score of the destination in season k, N G is the number of seasons.
7. The tourism recommendation method based on big data according to claim 1, wherein The specific steps for obtaining the destination demand change information are as follows: S321: Based on the timeliness demand, analyze the user's travel date, destination, and travel duration to identify user behavior characteristics and obtain a user behavior data set; S322: Based on the user behavior data set, divide time windows according to destination demand and seasonal fluctuations, and predict the demand changes of future popular tourist destinations to obtain destination demand change information.
8. The tourism recommendation method based on big data according to claim 1, wherein, The specific steps for obtaining the travel activity recommendation result are as follows: S411: Based on the destination demand change information, collect real-time user demands and current travel trend data, integrate the user's search and booking information, and identify the time nodes associated with travel demands to obtain the demand time node analysis result; S412: Based on the demand time node analysis result, conduct demand fluctuation analysis, extract the key demand change trends, and use the formula: Calculate the demand fluctuation value VB at each time node t , where PB ti represents the demand intensity at the i-th time node, and TB ti represents the time impact coefficient at the i-th time node, and n Y is the total number of time nodes; S413: Based on the demand fluctuation value, adjust the travel recommendation priority, and apply a sorting algorithm to optimize the priority ranking of destinations to obtain the travel activity recommendation result.
9. A tourism recommendation system based on big data, characterized in that, Execute according to the big data-based travel recommendation method described in any one of claims 1-8. The system includes: The user interest clustering module analyzes the user's browsing history, search frequency, and geographical location preferences based on user behavior data, social media interaction data, and search query data, clusters the user's interest points, and divides the destination preferences according to the clustering results to obtain the user interest classification value; The travel matching score module searches for travel destinations that match the user's interest points based on the user interest classification value, calculates the degree of fit between each destination and the user's preferences, and filters the travel destinations that meet the user's needs to obtain the destination matching score; The travel trend analysis module analyzes historical travel data and seasonal trends based on the destination matching score, determines the timeliness demand of the user for the destination, conducts time window division, and predicts the demand changes of future popular tourist destinations to obtain the destination demand change information; The travel recommendation optimization module conducts demand fluctuation analysis based on the destination demand change information, according to real-time user demands and current travel trends, and adjusts the travel recommendation priority to obtain the travel activity recommendation result.