Tourist attraction recommendation method based on social media content sentiment analysis
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本发明提供一种基于社交媒体内容情绪分析的旅游景点推荐方法,以解决现有技术基于游客的评论文本进行即时情感分析导致旅游景点推荐结果不够准确的问题,所采用的技术方案具体如下:
针对游客评价发布时间滞后导致的定位漂移问题,本发明通过获取评价数据与游客轨迹序列,利用发布时间进行轨迹回溯,并结合评价数据的发布内容进行双重校验。能够有效锁定目标景点及其对应的实际观赏时间,避免了仅依赖单一时刻定位数据造成的景点匹配错误,确保了后续分析基于准确的时空数据基础。
Smart Images

Figure CN122220623B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of next-generation information technology, specifically to a method for recommending tourist attractions based on sentiment analysis of social media content. Background Technology
[0002] With the widespread use of social media, travel recommendations based on visitor-generated content have become a mainstream technology. Existing attraction rating algorithms mainly rely on real-time sentiment analysis of content posted by visitors on social media, extracting emotional features from review texts to construct attraction rating systems.
[0003] However, in practical tourism recommendation applications, when using real-time sentiment analysis of tourist reviews to recommend tourist attractions, this single sentiment analysis technique cannot effectively eliminate the interference of the convenience of attraction facilities and services on the analysis of the attraction's actual attractiveness. Furthermore, tourists' tendency to provide retrospective reviews leads to a mismatch between the reviews and the actual conditions of the attractions visited. During peak tourist seasons, it is difficult to accurately identify and match high-quality attractions that meet tourists' psychological expectations based solely on attraction reviews, resulting in inaccurate tourism recommendation results. Summary of the Invention
[0004] This invention provides a tourist attraction recommendation method based on sentiment analysis of social media content, to solve the problem that existing technologies, which rely on real-time sentiment analysis of tourists' comments, result in inaccurate tourist attraction recommendations. The specific technical solution adopted is as follows:
[0005] One embodiment of the present invention discloses a method for recommending tourist attractions based on sentiment analysis of social media content, the method comprising the following steps: Obtain evaluation data and tourist trajectory sequences published by different historical tourists, perform trajectory backtracking based on the publication time of the evaluation data and the trajectory sequences, and determine the target attractions and the viewing time of the attractions by combining the publication content of the evaluation data; The tour resistance coefficient is calculated based on the environmental data of the target scenic spot during the viewing time at the scenic spot; content matching is performed on the evaluation text of the evaluation data to generate a burden reduction service tag; sentiment analysis is performed on the evaluation data to obtain the basic satisfaction, and the basic satisfaction is corrected based on the tour resistance coefficient and the burden reduction service tag to obtain the corrected satisfaction. Multi-level resistance intervals are constructed using the resistance coefficients of different historical tourists; the resistance coefficient of the tourists to be recommended at the recommended attractions within the planned visit time is calculated as the expected resistance; the corresponding resistance intervals are matched according to the expected resistance; and the attraction recommendation list is generated based on the corrected satisfaction corresponding to the matched resistance intervals.
[0006] Furthermore, determining the target attraction and the viewing time includes: A backtracking window is constructed based on the release time of the evaluation data, and the length of the backtracking window is adapted to the typical duration of a single tour activity; The trajectory segments that fall within the backtracking window are extracted from the tourist trajectory sequence, and the trajectory segments are matched with the areas of candidate attractions to determine the candidate attractions and the viewing time of the candidate attractions; Construct an entity word list for candidate attractions, calculate the semantic matching degree between the evaluation text of the evaluation data and the entity word list, and select target attractions from the candidate attractions based on the semantic matching degree.
[0007] Further, the semantic matching degree between the evaluation text of the evaluation data and the entity vocabulary is calculated, and based on the semantic matching degree, the target attraction is selected from the candidate attractions, including: The evaluation text of the evaluation data is segmented into words to generate an evaluation word segmentation set; The intersection-union ratio of the evaluation word segmentation set and the landscape entity set in the entity word list is determined as the semantic matching degree, and the candidate scenic spot with the highest semantic matching degree and exceeding the preset semantic matching threshold is selected as the target scenic spot.
[0008] Furthermore, the environmental data includes real-time humidity and heat index and real-time visitor flow data; the visitor resistance coefficient is calculated based on the environmental data of the target attraction during the viewing time at the attraction, including: Obtain historical humidity index data for the target attraction prior to the designated viewing time, as well as the peak visitor flow benchmark for the target attraction; Based on the real-time humidity and heat index and the historical humidity and heat index data, it is determined that the weather conditions at the target scenic spot are inappropriate during the viewing time at the scenic spot. The real-time passenger flow data is compared with the peak passenger flow benchmark of the target attraction to determine the passenger flow resistance value of the target attraction during the viewing time at the attraction; Based on the aforementioned weather inappropriateness and passenger flow resistance values, the tour resistance coefficient is determined.
[0009] Furthermore, determining the weather conditions at the target scenic spot during the viewing time includes: The highest value of the historical humidity index data of the target scenic spot within the observation period prior to the viewing time of the scenic spot is recorded as the highest humidity index of the target scenic spot. Determine the first difference between the real-time humidity index and the preset suitable humidity index benchmark value, and the second difference between the highest humidity index and the suitable humidity index benchmark value; The maximum value between the second difference and the preset correction factor is determined as the target denominator; Calculate the first ratio of the first difference to the target denominator, and truncate the first ratio. The truncated value is then determined as the weather inappropriateness of the target scenic spot during the viewing time at the scenic spot.
[0010] Furthermore, determining the visitor flow resistance value of the target scenic spot during the viewing time at the scenic spot includes: Calculate the second ratio between the real-time passenger flow data and the passenger flow peak benchmark, and perform numerical truncation on the second ratio. Determine the truncated value as the passenger flow resistance value of the target scenic spot during the viewing time at the scenic spot.
[0011] Furthermore, content matching is performed on the evaluation text of the evaluation data to generate a burden reduction service tag, including: Construct a feature word library for reducing the burden on target attractions, wherein the feature word library for reducing the burden on attractions includes at least a subset of transportation tools and a subset of privileged access channels; Detect whether the evaluation text of the evaluation data contains feature elements belonging to the feature word library of the burden reduction service; If the aforementioned feature element is detected, the burden reduction service marker is set to a first value, which indicates that the tourist has used the burden reduction service; if the aforementioned feature element is not detected, the burden reduction service marker is set to a second value, which indicates that the tourist is in a natural tour state, and the first value is not equal to the second value.
[0012] Furthermore, the process of revising the basic satisfaction level based on the tour resistance coefficient and the burden reduction service marker to obtain revised satisfaction level includes: When the burden reduction service is marked as the first value, the basic satisfaction level is used as the adjusted satisfaction level; When the burden reduction service is marked as the second value, a positive gain coefficient is constructed using the tour resistance coefficient, and the basic satisfaction is increased and corrected using the positive gain coefficient.
[0013] Furthermore, the construction of multi-level resistance intervals using historical visit resistance coefficients corresponding to different historical tourists includes: Multiple consecutive numerical ranges are preset, serving as low resistance, medium resistance, and high resistance zones, respectively. The historical tour resistance coefficients corresponding to different historical tourists are classified into the low resistance range, medium resistance range, and high resistance range, thus obtaining a multi-level resistance range.
[0014] Furthermore, based on the adjusted satisfaction level corresponding to the matched resistance range, a list of recommended attractions is generated, including: Iterate through the evaluation data of the target attractions, and based on the tour resistance coefficient corresponding to each evaluation data, allocate the corrected satisfaction corresponding to each evaluation data to the matched resistance range. Calculate the statistical average of all corrected satisfaction levels within each resistance range to obtain the graded evaluation index for the target scenic spot; Determine the resistance range to which the expected tour resistance belongs, and select the corresponding graded evaluation index as the ranking basis; The candidate attractions are arranged in descending order of the numerical values used for sorting, and a recommended attraction list is generated.
[0015] The beneficial effects of this invention are: To address the location drift issue caused by the lag in tourist review posting times, this invention acquires review data and tourist trajectory sequences, uses the posting time for trajectory backtracking, and performs double verification by combining the posting content of the review data. This effectively pinpoints the target attraction and its corresponding actual viewing time, avoiding attraction matching errors caused by relying solely on location data at a single moment, and ensuring that subsequent analysis is based on accurate spatiotemporal data.
[0016] Furthermore, existing technologies rely solely on subjective user ratings, failing to distinguish whether the ratings stem from the landscape itself or are influenced by external environmental factors (such as high temperatures or congestion) or auxiliary facilities (such as cable cars). This invention corrects basic satisfaction levels by calculating the difficulty coefficient of the tour and generating burden-reducing service markers. It decouples users' subjective evaluations from the physiological and psychological burdens during the tour, thereby uncovering the true attractiveness of the attraction after removing environmental noise.
[0017] Finally, this invention overcomes the limitations of static rating-based recommendations by constructing multi-level resistance intervals using historical data and calculating the expected resistance of tourists during the recommendation process to match the corresponding resistance intervals. This allows the recommendation system to rank attractions based on the specific context of the user's planned visit, invoking adjusted satisfaction levels for that specific context. This avoids recommending attractions only suitable for comfortable environments under adverse conditions, significantly improving recommendation accuracy and user experience. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1This is a flowchart illustrating a tourist attraction recommendation method based on social media content sentiment analysis, provided as an embodiment of the present invention. Figure 2 This is a flowchart illustrating the steps of a method for obtaining the tour resistance coefficient according to an embodiment of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Please see Figure 1 The diagram illustrates a flowchart of a tourist attraction recommendation method based on social media content sentiment analysis, according to an embodiment of the present invention. The method includes the following steps: Step S100: Obtain evaluation data and tourist trajectory sequences published by different historical tourists, perform trajectory backtracking based on the publication time of the evaluation data and the tourist trajectory sequences, and determine the target scenic spot and the viewing time of the scenic spot by combining the publication content of the evaluation data.
[0022] In outdoor tourism scenarios, tourists commonly exhibit the behavior of visiting before posting reviews. For example, tourists may post reviews retrospectively on their return trip or while resting. The posting time of tourist reviews often lags behind the actual visit time. Directly associating the posting time with information about the tourist attraction's environment can lead to time-matching errors. Therefore, this step establishes a unique mapping between review content and the visit time and attraction through trajectory backtracking and semantic verification.
[0023] As an optional implementation, to obtain review data posted by different historical tourists, an open API interface is called to retrieve the historical review list for each attraction. Regular expressions are used to remove HTML tags and URL links from the original text. The SimHash value of each review text is calculated. If the Hamming distance between two texts is less than the text duplication threshold of 3, the two texts are considered duplicates. The earliest published review is retained, and invalid short texts with a length of less than 5 characters are filtered out to obtain the review text. A time parser is used to convert all time strings corresponding to the review text into a standard format (e.g., year-month-day hour:minute:second), thus obtaining the review data posted by different historical tourists.
[0024] As an optional implementation, to obtain visitor trajectory sequences for different historical visitors, the open API interfaces or authorized web crawlers of social media platforms are used to acquire all dynamic content with geolocation tags for the target visitor within a historical time period, and the latitude and longitude coordinates of the dynamic content are extracted. Due to the sparsity of data on social media platforms, a linear interpolation algorithm is used to complete the virtual tour trajectory points between two adjacent latitude and longitude coordinates, thereby constructing a continuous visitor trajectory sequence.
[0025] As an optional implementation method, the method for determining the target attraction and the viewing time includes: Step S101: Construct a backtracking window based on the release time of the evaluation data. The length of the backtracking window is adapted to the typical duration of a single visit. In this embodiment, the duration of the backtracking window is 12 hours. The backtracking window is constructed by shifting its length backward from the release time.
[0026] Step S102: Extract trajectory segments that fall within the backtracking window from the tourist trajectory sequence, and match the trajectory segments with the areas of candidate attractions to determine the candidate attractions and their viewing times. The specific implementation process includes: Obtain candidate scenic spot geofences, which are closed polygonal regions constructed based on electronic map data, which consists of an ordered set of latitude and longitude coordinates. In this set It indicates the longitude and latitude of the boundary point of the scenic spot.
[0027] It should be understood that when the latitude and longitude coordinates of a tourist are located within the geofence of the electronic map, the ray emanating from the tourist's location will inevitably intersect with the boundary of the geofence; if the tourist's location is not within the geofence of the electronic map, the number of intersections will necessarily be even.
[0028] Specifically, spatial matching is performed using the ray projection method, which identifies the trajectory points to be matched. As an endpoint, a ray is emitted in any direction; in this embodiment, it is taken as horizontal to the right. The number of intersections between the ray and the boundary line segment of the closed polygon region is obtained. If the number of intersections is odd, the trajectory point Q is determined to be inside the geofence, i.e., the match is successful; if the number of intersections is even, the trajectory point Q is determined to be outside the geofence, i.e., the match fails.
[0029] All successfully matched trajectory points and their corresponding attractions are selected as candidate attractions, and the timestamp of the first time a tourist's trajectory point enters the geofence of the candidate attraction is marked as the viewing time of the candidate attraction.
[0030] Step S103: Construct an entity word list for candidate attractions, calculate the semantic matching degree between the evaluation text of the evaluation data and the entity word list, and select target attractions from the candidate attractions based on the semantic matching degree.
[0031] In the backtracking window, tourists may have visited multiple geographically adjacent attractions. While the trajectory data alone can yield a set of candidate attractions, it's impossible to determine the correspondence between a user's review and a specific attraction. By calculating the semantic overlap between the review text and the entity lists of each candidate attraction, a unique target object can be selected from the candidate set, enabling precise identification of the specific attraction.
[0032] As an optional implementation, selecting the target attraction from the candidate attractions includes: First, the evaluation text of the evaluation data is segmented into words to generate an evaluation word set.
[0033] The evaluation text was segmented using a Chinese word segmentation tool, and the segmented results were recorded as the evaluation word segmentation set. .
[0034] Secondly, the intersection-union ratio of the evaluation word segmentation set and the landscape entity set in the entity word list is determined as the semantic matching degree, and the candidate scenic spot with the highest semantic matching degree and exceeding the preset semantic matching threshold is selected as the target scenic spot.
[0035] Taking the k-th candidate scenic spot as an example, the formula for calculating the semantic matching degree between the evaluation text of the evaluation data and the entity vocabulary of the k-th candidate scenic spot is as follows: ; in, Let represent the semantic matching degree of the k-th candidate scenic spot, which is the semantic matching degree between the evaluation text of the evaluation data and the entity vocabulary of the k-th candidate scenic spot. To evaluate the word segmentation set, Let be the entity vocabulary for the k-th candidate attraction, where || represents the cardinality of the set, i.e., the number of elements in the set. By measuring the degree of overlap between the evaluation word segmentation set and the attraction features, we can effectively filter out the unique tourist attraction related to the evaluation content from multiple geographically close candidate attractions, ensuring the accuracy of subsequent attraction analysis. The entity vocabulary is constructed by collecting data from official guides and encyclopedia entries and extracting attraction names, aliases, and unique landscape names.
[0036] It should be noted that while the backtracking window can cover the user's entire itinerary, tourists may have track points for multiple candidate attractions within the time period corresponding to the backtracking window. For example, a tourist might visit the Yuanmingyuan and the Summer Palace, which are geographically adjacent, on the same day. Judging solely based on dwell time or entry time is highly prone to misjudgment. In such cases, it is necessary to combine the semantic information of the attractions in the evaluation text for calculation and filtering. For instance, if the evaluation word segmentation set contains words like "Dashuifa" and "Xiyanglou," and these words only exist in the entity word list corresponding to Yuanmingyuan but not in the entity word list of the Summer Palace, then the semantic match degree between this evaluation text and the entity of the Yuanmingyuan attraction is significantly greater than zero, while the semantic match degree with the entity of the Summer Palace will approach zero.
[0037] Calculate the semantic matching score for all candidate attractions, set the semantic matching threshold to 0.05, and select the candidate attraction with the highest semantic matching score that exceeds the semantic matching threshold as the target attraction. If the highest semantic matching score is lower than the semantic matching threshold, it is determined that the current evaluation content does not contain specific attraction information, and the evaluation text is marked as invalid.
[0038] It should be further noted that social media reviews are typically unstructured short texts, averaging 20-40 words in length. These texts contain a large number of generic sentiment words and function words that cannot be used to distinguish between attractions. The average length of the entity list for candidate attractions, including their official names and alternative names, is approximately 10-15 words. Due to the abundance of generic sentiment words and function words in the review text, a single review usually only mentions one or two specific scenic entity names. Therefore, the semantic match score of candidate attractions is relatively low, and consequently, the semantic match threshold is also set relatively low.
[0039] The following comparative examples further illustrate the effect of this threshold. Example 1: The evaluation text is "Today there were too many people, the queue was long, and the experience was extremely poor," which does not contain any scenic spot entities. In this case, the calculated semantic matching score is 0. Example 2: The evaluation text is "There were many people at the Great Fountain, and we couldn't even squeeze into the Western-style Buildings ruins." After processing, the total number of word segments is 20. At this time, it matches the two scenic spot entities "Great Fountain" and "Western-style Buildings." The intersection element is 2, and the union element is 28. The calculated semantic matching score is approximately 0.07, which is significantly higher than the semantic matching threshold. Therefore, the target scenic spot is identified as the Old Summer Palace.
[0040] Step S200: Calculate the tour resistance coefficient based on the environmental data of the target attraction during the viewing time at the attraction; perform content matching on the evaluation text of the evaluation data to generate a burden reduction service tag; perform sentiment analysis on the evaluation data to obtain a basic satisfaction level, and correct the basic satisfaction level based on the tour resistance coefficient and the burden reduction service tag to obtain a corrected satisfaction level.
[0041] In evaluating target attractions, visitor satisfaction is often influenced by both external environmental factors, such as high temperatures and congestion, and internal facilities, such as cable cars and VIP access. Without considering the convenience offered by these facilities, attractions with mediocre scenery but excellent amenities might be recommended to users; conversely, without considering environmental obstacles, it's impossible to identify high-quality attractions that can still achieve visitor satisfaction even under adverse conditions. This step aims to isolate environmental and facility interference to obtain a revised satisfaction level after removing these factors.
[0042] As an optional implementation method, such as Figure 2 As shown, the method for obtaining the tour resistance coefficient can be implemented by steps S201 to S204.
[0043] Step S201: Obtain historical humidity index data of the target attraction before the viewing time of the attraction and the peak visitor flow benchmark of the target attraction.
[0044] It should be understood that in cross-seasonal and cross-regional tourism recommendation scenarios, environmental data fluctuates greatly, and the highest current or seasonal values are usually used for analysis. To ensure the objectivity of comparisons between different times and different attractions, it is necessary to first conduct statistical analysis on the historical environmental data and historical visitor flow data of the attractions.
[0045] Specifically, the observation period was set to 5 years, and historical temperature, relative humidity, and visitor flow data for the target scenic spot area were obtained at multiple time points. A historical humidity index was constructed using the historical temperature and relative humidity data within the observation period. The humidity index (Humidex) was calculated according to the standard established by JM Masterton and FA Richardson in 1979. This index is a dimensionless indicator that combines temperature and relative humidity to describe human thermal perception; its value approximately corresponds to the dry heat temperature felt by the human body at the same air temperature.
[0046] It should be noted that the above calculation method is only one typical way to calculate the humidity index. Those skilled in the art will understand that any meteorological index constructed based on temperature and humidity data to characterize the thermal comfort of the human body in a humid and hot environment falls within the technical scope of the humidity index described in this invention.
[0047] Furthermore, following the same method used to obtain historical humidity and heat index data, the real-time humidity and heat index of the target scenic spot during the viewing time can be obtained. Simultaneously, the peak visitor flow benchmark for the target scenic spot is obtained. The process includes: retrieving historical visitor flow data for the target scenic spot within the observation period, extracting all historical data with the same month and day attribute as the viewing time of the target scenic spot; statistically analyzing the visitor flow distribution for each hour within the historical dates, and determining the maximum hourly visitor flow value as the peak visitor flow benchmark for the target scenic spot. .
[0048] Step S202: Compare the real-time humidity index with the historical humidity index data to determine the weather conditions at the target scenic spot during the viewing time.
[0049] Specifically, the highest historical humidity index value of the target scenic spot within the observation period prior to the viewing time of the scenic spot is recorded as the highest humidity index of the target scenic spot; a first difference between the real-time humidity index and the suitable humidity index benchmark value is determined, as well as a second difference between the highest humidity index and the suitable humidity index benchmark value; a first ratio of the first difference and the second difference is calculated, and the first ratio is truncated, and the truncated value is determined as the weather incompatibility of the target scenic spot during the viewing time of the scenic spot.
[0050] Taking any target scenic spot as an example, the formula for calculating the inappropriate weather conditions during the viewing time at that scenic spot is as follows: ; in, The weather conditions at the target attraction w during the designated viewing time are inappropriate. This represents a truncation function. If the calculated result is less than zero, it is set to zero; if it is greater than 1, it is set to 1. The truncation function effectively prevents numerical overflow caused by extreme abnormal weather data. This indicates the real-time humidity index of the target attraction. This indicates the baseline value for the suitable humidity and heat index of the target scenic spot. This represents the highest humidity index of the target scenic spot, and Max() is the maximum value function. The preset correction factor is set to 0.001 in this embodiment to avoid the denominator being zero during the calculation process.
[0051] It should be noted that, taking into account the climate characteristics and human tolerance of typical tourist attractions, the baseline value of the suitable humidity and heat index for the target attractions is [not specified]. The value is preset to a fixed value of 26. This value represents the thermal comfort threshold for the human body when outdoors; that is, when the humidity index is 26, the vast majority of tourists feel comfortable and without thermal stress.
[0052] Step S203: Compare the real-time passenger flow data with the peak passenger flow benchmark of the scenic spot to determine the passenger flow resistance value of the target scenic spot during the viewing time of the scenic spot.
[0053] Specifically, a second ratio of real-time passenger flow data to the peak passenger flow benchmark is calculated, and the second ratio is truncated. The truncated value is then determined as the passenger flow resistance value of the target attraction during the viewing time at the attraction.
[0054] Taking any target attraction as an example, the formula for calculating the visitor flow resistance value during the viewing time of that target attraction is as follows: ; in, The target attraction's visitor flow resistance value. This represents a truncation function. If the calculated result is less than zero, it is set to zero; if it is greater than 1, it is set to 1. This truncation function effectively prevents numerical overflow caused by extreme weather data. This indicates real-time visitor flow data for the target attraction. This represents the peak passenger flow benchmark and is usually not 0.
[0055] It should be noted that the specific method for obtaining the real-time visitor flow data of the target attraction is to set the evaluation release time window as a time window one hour before and after the visit time of the attraction, and record the number of evaluations released for the target attraction within the evaluation release time window as the real-time visitor flow data of the target attraction.
[0056] Step S204: Based on the aforementioned weather incompatibility and passenger flow resistance values, determine the tour resistance coefficient. The specific calculation formula is as follows: ; in, As a weight for weather and environmental resistance, Due to the congestion of tourist attractions, The weather at the target scenic spot was inappropriate. The resistance value for passenger flow at the target attraction. This represents the resistance coefficient for visiting the target attraction.
[0057] It should be noted that the resistance coefficient of a target attraction reflects the resistance level of the visit. When the calculated resistance coefficient approaches 1, it indicates that the objective environment at the time of the visit is extremely harsh, and the visit is more difficult; conversely, if the resistance coefficient approaches 0, it indicates that the environment at the time of the visit is comfortable. In outdoor sightseeing scenarios, tourist fatigue and discomfort mainly stem from the weather conditions and overcrowding at the attraction. In this embodiment, [the following is set...] To prevent overemphasis on weather conditions or overcrowding at attractions from rendering the overall resistance assessment ineffective, and to ensure that the final recommended attractions meet tourists' psychological expectations.
[0058] It should be understood that scenic spots usually set up services such as cable cars, sightseeing vehicles or fast passageways to reduce the burden on tourists and improve their experience. Using these services to reduce environmental resistance can lead to inaccurate calculations of the resistance coefficient of the target attraction. Therefore, it is necessary to correct the resistance coefficient of the target attraction based on the actual state of tourists at the attraction.
[0059] As another preferred implementation, the calculation process of the tour resistance coefficient adopts a dynamic weight allocation mechanism based on user preferences: in the aforementioned embodiments, weather environment resistance weights are used. Weight of crowded tourist attractions The values are set to fixed values (e.g., all 0.5). However, different users exhibit significant differences in their sensitivity to environmental stress. For example, some users are extremely sensitive to high temperatures but have a high tolerance for crowding, while others are the opposite. To achieve more accurate personalized recommendations, this invention introduces a user sensitivity factor to implement dynamic weighting.
[0060] The specific steps for setting the dynamic weights include: When a user requests a recommendation, the user selects a desired attraction tag, which includes at least "heat-averse", "crowd-averse", or "easygoing".
[0061] If the expected attraction label is "heat-sensitive", then increase the weighting of weather and environmental resistance. The proportion, for example, setting =0.7, =0.3, making the negative experience in hot weather have a greater impact on the user's resistance coefficient; if the user is labeled "averse to crowds", then the weight of crowded places at attractions will be increased. The proportion, for example, setting =0.3, =0.7; if there is no explicit label, the default balanced weight is maintained, i.e. .
[0062] As an optional implementation, content matching is performed on the evaluation text of the evaluation data to generate a burden reduction service tag, including: First, a feature terminology database for burden-reduction services at target attractions is constructed. This database includes at least a subset of transportation options and a subset of privileged access routes. The transportation options subset includes at least "cable car," "ropeway," or "sightseeing bus," and the privileged access route subset includes at least "VIP access," "skip-the-line," or "fast-track access." Second, the evaluation text of the evaluation data is checked to see if it contains feature elements belonging to the burden-reduction service feature terminology database. If the feature elements are detected, the burden-reduction service flag is set to a first value, which indicates that the tourist used the burden-reduction service. If the feature elements are not detected, the burden-reduction service flag is set to a second value, which indicates that the tourist was in a natural sightseeing state, and the first value is not equal to the second value.
[0063] Specifically, the evaluation word segmentation set is compared one by one with the feature word library of the burden reduction service. If the evaluation text contains any words belonging to the subset of transportation tools or the subset of privileged passages, it is determined that the tourist has used the burden reduction service and the burden reduction service flag is assigned a value of 1; if it does not contain any words belonging to the subset of transportation tools or the subset of privileged passages, the burden reduction service flag is assigned a value of 0.
[0064] As an optional implementation method, the method for obtaining satisfaction includes the following steps: When the burden reduction service is marked as a first value, the basic satisfaction level is used as the adjusted satisfaction level; when the burden reduction service is marked as a second value, a positive gain coefficient is constructed using the tour resistance coefficient, and the basic satisfaction level is increased and corrected using the positive gain coefficient.
[0065] Specifically, a sentiment analysis model using a Long Short-Term Memory (LSTM) network architecture is employed. This model has been pre-tuned based on a tourism corpus and possesses the ability to accurately identify semantic features of tourism-related verticals. The set of evaluation word segments is input into the sentiment analysis model, and the output is the confidence score of the evaluation word segments belonging to the positive sentiment category, with a value range of [0,1]. The output of the sentiment analysis model is recorded as the basic satisfaction level. .
[0066] It should be noted that a basic satisfaction score close to 1 indicates extreme satisfaction, and the corresponding evaluation word set may contain texts such as "shocking" and "exquisite"; a basic satisfaction score close to 0 indicates extreme dissatisfaction, and the corresponding evaluation word set may contain texts such as "disappointing" and "avoid pitfalls". A score of approximately 0.5 indicates a neutral or emotionally neutral evaluation.
[0067] Specifically, the baseline satisfaction level is adjusted using the visitor resistance coefficient of the target attraction and the service reduction indicator to obtain the adjusted satisfaction level. The calculation formula is as follows: ; in, To adjust satisfaction levels, Based on basic satisfaction, Adjust the weights for environmental resistance. The resistance coefficient for visiting the target attraction. Marked as a service to reduce workload.
[0068] It should be noted that when the burden reduction service flag is 1, it means that the tourist has used the relevant burden reduction service content of the target scenic spot. In this case, regardless of whether there is significant resistance to the tour at the target scenic spot, it will be forcibly reset to zero, which is consistent with the actual situation that tourists use the relevant burden reduction service content of the scenic spot to avoid environmental resistance and obtain a comfortable experience. When the burden reduction service flag is 0, it means that the tourist has not used the relevant burden reduction service content of the target scenic spot. In this case, the basic satisfaction is corrected by the environmental resistance correction weight. In this embodiment, the value is 0.6, which reflects the difficulty of tourists still giving positive reviews under high resistance environment. If the value is higher, it means that the actual attraction is more attractive, so the corrected satisfaction value should also be higher.
[0069] Furthermore, it is important to emphasize that this invention analyzes tourists' subjective experiences based on natural language processing technology. When generating service burden reduction tags, it strictly relies on explicit semantic features in the evaluation text as the sole criterion, without incorporating tourists' ticketing consumption data for verification.
[0070] The technical basis for this logic is that the evaluation text is a subjective perception mapping of the visitor's experience, rather than an objective behavioral log. If the evaluation text does not contain features related to reducing visitor burden, even if the visitor has an objective ticket purchase record (e.g., not riding due to queuing after purchasing a ticket, or the facility experience being mediocre and not remembered), it indicates that the service facility has not effectively offset the negative impact of environmental resistance on the visitor experience. Therefore, this embodiment only determines the natural visit state based on the evaluation text, ensuring that the corrected satisfaction truly reflects the visitor's degree of recognition of the core value of the landscape under the current perceived resistance.
[0071] Step S300: Construct multi-level resistance intervals using historical visit resistance coefficients; calculate the visit resistance coefficient of the tourists to be recommended at the recommended attractions within the planned visit time as the expected visit resistance; match the corresponding resistance intervals based on the expected visit resistance; and generate a list of recommended attractions based on the corrected satisfaction level corresponding to the matched resistance intervals.
[0072] The experiential value of tourist attractions is not a constant but varies with environmental pressures. Therefore, a multi-level resistance range is constructed using historical visit resistance coefficients, and the visit resistance coefficients of tourists to be recommended at the attractions within the planned visit time are matched to the corresponding resistance ranges. This determines the actual experience feedback under the same resistance level and ultimately generates a list of recommended attractions, minimizing the deviation between user expectations and actual experiences.
[0073] As an optional implementation, the steps for constructing a multi-level resistance range include: Multiple consecutive numerical ranges are preset, serving as low resistance, medium resistance, and high resistance zones, respectively. The historical tour resistance coefficients corresponding to different historical tourists are classified into the low resistance range, medium resistance range, and high resistance range, thus obtaining a multi-level resistance range.
[0074] Specifically, in order to cover common sightseeing environments, the resistance zone is divided into: Lower resistance zone: This provides an ideal tourist environment with a comfortable climate and sparse tourist flow. Intermediate resistance zone: This corresponds to a typical tourist environment; Advanced resistance zone: This is in response to adverse tourist conditions such as high temperatures, torrential rain, or extreme congestion.
[0075] As an optional implementation, the steps for generating a list of recommended attractions include: Iterate through the evaluation data of the target attractions, and based on the tour resistance coefficient corresponding to each evaluation data point, allocate the revised satisfaction level corresponding to each evaluation data point to the matched resistance range.
[0076] The statistical average of all corrected satisfaction levels within each resistance range is calculated to obtain the graded evaluation index for the target scenic spot.
[0077] Specifically, for each target attraction, the corrected satisfaction level corresponding to all evaluation texts will be calculated and assigned to the corresponding resistance coefficient within the resistance interval. The arithmetic mean of all corrected satisfaction levels assigned within each resistance interval will be calculated, and the result will be recorded as a graded evaluation index. The graded evaluation index includes a high-level resistance interval evaluation index, a medium-level resistance interval evaluation index, and a low-level resistance interval evaluation index. If there is no corrected satisfaction value within the corresponding interval, the value will be 0.
[0078] Determine the resistance range to which the expected tour resistance belongs, and select the corresponding graded evaluation index as the ranking basis.
[0079] Specifically, if the expected resistance to tourism falls into the high multi-level resistance range, the high multi-level resistance range classification evaluation index will be used as the ranking basis; if it falls into the medium multi-level resistance range, the medium multi-level resistance range classification evaluation index will be used as the ranking basis; if it falls into the low multi-level resistance range, the low multi-level resistance range classification evaluation index will be used as the ranking basis.
[0080] The step of obtaining the expected tour resistance includes: First, obtain the predicted heat and humidity index by calling the API interface of a third-party meteorological service platform to obtain the weather forecast data of the target scenic spot during the user's planned visit time, extract the predicted temperature and predicted relative humidity, and use the predicted temperature and relative humidity to calculate the predicted heat and humidity index.
[0081] Secondly, the predicted visitor flow is obtained by reading real-time reservation and ticketing data from the scenic area's ticketing management system as the predicted value. If there is no reservation data, a time period is preset, and visitor flow data for the scenic spot within the preset time period before the current time is extracted. Linear interpolation or time series algorithms are then used to obtain the predicted visitor flow at the planned time. In this embodiment, the preset time period is set to 3 days.
[0082] Finally, the predicted humidity index and predicted visitor flow are substituted into the tourism resistance coefficient calculation formula described in step S200 to calculate the expected tourism resistance for that period.
[0083] The candidate attractions are sorted in descending order according to the numerical value of the sorting criterion to generate the attraction recommendation list. Specifically, this includes: sorting all candidate attractions in descending order according to the numerical value of the sorting primary key, and selecting the top N attractions as the attraction recommendation list. In this embodiment, N is 10.
[0084] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A tourist attraction recommendation method based on sentiment analysis of social media content, characterized in that, Includes the following steps: Obtain evaluation data and tourist trajectory sequences published by different historical tourists, perform trajectory backtracking based on the publication time of the evaluation data and the trajectory sequences, and determine the target attractions and the viewing time of the attractions by combining the publication content of the evaluation data; Calculate the visitor resistance coefficient based on environmental data of the target scenic spot during the visitor's visit time. Content matching is performed on the evaluation text of the evaluation data to generate a burden reduction service tag; Sentiment analysis is performed on the evaluation data to obtain basic satisfaction, and the basic satisfaction is corrected based on the tour resistance coefficient and the burden reduction service marker to obtain corrected satisfaction. Construct multi-level resistance zones by utilizing the visitor resistance coefficients corresponding to different historical tourists; Calculate the expected tour resistance coefficient of the tourists to be recommended at the recommended attractions within the planned tour time as the expected tour resistance, match the corresponding resistance interval based on the expected tour resistance, and generate an attraction recommendation list based on the corrected satisfaction corresponding to the matched resistance interval; The environmental data includes real-time humidity index and real-time passenger flow data; The visitor resistance coefficient is calculated based on environmental data of the target attraction during the viewing time at the attraction, including: Obtain historical humidity index data for the target attraction prior to the designated viewing time, as well as the peak visitor flow benchmark for the target attraction; Based on the real-time humidity and heat index and the historical humidity and heat index data, it is determined that the weather conditions at the target scenic spot are inappropriate during the viewing time at the scenic spot. The real-time passenger flow data is compared with the peak passenger flow benchmark of the target attraction to determine the passenger flow resistance value of the target attraction during the viewing time at the attraction; Based on the aforementioned weather inappropriateness and passenger flow resistance values, the tour resistance coefficient is determined.
2. The tourist attraction recommendation method based on social media content sentiment analysis according to claim 1, characterized in that, The determination of the target scenic spot and the viewing time includes: A backtracking window is constructed based on the release time of the evaluation data, and the length of the backtracking window is adapted to the typical duration of a single tour activity; The trajectory segments that fall within the backtracking window are extracted from the tourist trajectory sequence, and the trajectory segments are matched with the areas of candidate attractions to determine the candidate attractions and the viewing time of the candidate attractions; Construct an entity word list for candidate attractions, calculate the semantic matching degree between the evaluation text of the evaluation data and the entity word list, and select target attractions from the candidate attractions based on the semantic matching degree.
3. The tourist attraction recommendation method based on social media content sentiment analysis according to claim 2, characterized in that, Calculate the semantic matching degree between the evaluation text of the evaluation data and the entity vocabulary, and based on the semantic matching degree, filter the target scenic spot from the candidate scenic spots, including: The evaluation text of the evaluation data is segmented into words to generate an evaluation word segmentation set; The intersection-union ratio of the evaluation word segmentation set and the landscape entity set in the entity word list is determined as the semantic matching degree, and the candidate scenic spot with the highest semantic matching degree and exceeding the preset semantic matching threshold is selected as the target scenic spot.
4. The tourist attraction recommendation method based on social media content sentiment analysis according to claim 1, characterized in that, The determination of unsuitable weather conditions at the target scenic spot during the viewing time includes: The highest value of the historical humidity index data of the target scenic spot within the observation period prior to the viewing time of the scenic spot is recorded as the highest humidity index of the target scenic spot. Determine the first difference between the real-time humidity index and the preset suitable humidity index benchmark value, and the second difference between the highest humidity index and the suitable humidity index benchmark value; The maximum value between the second difference and the preset correction factor is determined as the target denominator; Calculate the first ratio of the first difference to the target denominator, and truncate the first ratio. The truncated value is then determined as the weather inappropriateness of the target scenic spot during the viewing time at the scenic spot.
5. The tourist attraction recommendation method based on social media content sentiment analysis according to claim 1, characterized in that, The determination of the visitor flow resistance value of the target scenic spot during the viewing time includes: Calculate the second ratio between the real-time passenger flow data and the passenger flow peak benchmark, and perform numerical truncation on the second ratio. Determine the truncated value as the passenger flow resistance value of the target scenic spot during the viewing time at the scenic spot.
6. The tourist attraction recommendation method based on social media content sentiment analysis according to claim 1, characterized in that, The evaluation text of the evaluation data is matched for content to generate a burden reduction service tag, including: Construct a feature word library for reducing the burden on target attractions, wherein the feature word library for reducing the burden on attractions includes at least a subset of transportation tools and a subset of privileged access channels; Detect whether the evaluation text of the evaluation data contains feature elements belonging to the feature word library of the burden reduction service; If the characteristic element is detected, the burden reduction service flag is set to a first value, which indicates that the tourist has used the burden reduction service; if the characteristic element is not detected, the burden reduction service flag is set to a second value, which indicates that the tourist is in a natural tour state, and the first value is not equal to the second value.
7. The tourist attraction recommendation method based on social media content sentiment analysis according to claim 6, characterized in that, The process of revising the basic satisfaction level based on the tour resistance coefficient and the burden reduction service marker to obtain revised satisfaction level includes: When the burden reduction service is marked as the first value, the basic satisfaction level is used as the adjusted satisfaction level; When the burden reduction service is marked as the second value, a positive gain coefficient is constructed using the tour resistance coefficient, and the basic satisfaction is increased and corrected using the positive gain coefficient.
8. The tourist attraction recommendation method based on social media content sentiment analysis according to claim 1, characterized in that, The method of constructing multi-level resistance ranges using historical visit resistance coefficients corresponding to different historical tourists includes: Multiple consecutive numerical ranges are preset, serving as low resistance, medium resistance, and high resistance zones, respectively. The historical tour resistance coefficients corresponding to different historical tourists are classified into the low resistance range, medium resistance range, and high resistance range, thus obtaining a multi-level resistance range.
9. The tourist attraction recommendation method based on social media content sentiment analysis according to claim 1, characterized in that, Based on the adjusted satisfaction level corresponding to the matched resistance range, a list of recommended attractions is generated, including: Iterate through the evaluation data of the target attractions, and based on the tour resistance coefficient corresponding to each evaluation data, allocate the corrected satisfaction corresponding to each evaluation data to the matched resistance range. Calculate the statistical average of all corrected satisfaction levels within each resistance range to obtain the graded evaluation index for the target scenic spot; Determine the resistance range to which the expected tour resistance belongs, and select the corresponding graded evaluation index as the ranking basis; The candidate attractions are arranged in descending order of the numerical values used for sorting, and a recommended attraction list is generated.
Citation Information
Patent Citations
Scenic spot recommendation system based on scenic spot clustering and group emotion recognition
CN112257517A
Scenic spot recommendation method and device based on tourist portrait, and storage medium
CN115578220A