Scenic spot recommendation method based on social media multi-modal data
Through the integration of multimodal data on social media, machine learning and sentiment analysis technology, the physical information and user emotions of the scenic spots are extracted, and a personalized list of scenic spots is generated, which solves the problem of low information utilization in the existing technology and improves the accuracy and user experience of the recommendation system.
Patent Information
- Application Number
- CN202510552038.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-26
AI Technical Summary
The existing attractions recommendation system is difficult to make full use of social media multimodal data, especially pictures and travel texts, resulting in low information utilization and difficult to accurately reflect user interests and attractions characteristics.
By obtaining multimodal data from social media platforms, using machine learning landmark recognition and sentiment analysis technology, physical information of attractions is extracted from pictures, combined with emotional information from travel notes, comprehensive recommendation scores are calculated, and a personalized list of attractions is generated.
It has achieved an in-depth understanding of user interests and scenic spot characteristics, improved the accuracy and personalization of the recommendation system, and enhanced users' trust and dependence on the recommendation system.
Smart Images

Figure CN120541293A_ABST
Abstract
Description
Technical Field
[0001] The present invention discloses a method for recommending scenic spots based on multimodal data of social media, and belongs to the technical field of scenic spot recommendation. Background Art
[0002] At present, most scenic spot recommendation systems are based on analysis of user historical behavior, geographic location or displayed evaluations. The data source is single and it is difficult to fully reflect user interests and scenic spot characteristics. With the rise of social media, users have generated a large amount of multimodal data (including pictures, text, etc.). Although the picture content contains the scenic spot image information and geographic location information taken by tourists, it lacks rich language and emotional descriptions. For scenic spot recommendations, the information expressed is relatively single, and the travel text data makes up for this deficiency. The combination of these two aspects of data enriches users' preferences and emotional expressions for scenic spots. However, traditional methods have a low utilization rate of these multimodal data and it is difficult to extract valuable information from them. To this end, the present invention proposes a scenic spot recommendation method based on social media multimodal data, which captures user interests by fusing the landmark information of pictures, the emotional information of texts and other scenic spot characteristics, and improves the accuracy and personalization of scenic spot recommendations. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for recommending attractions based on multimodal data from social media, so as to solve the problem in the prior art that traditional methods for recommending attractions are difficult to extract valuable information from data.
[0004] A method for recommending attractions based on multimodal data from social media, comprising:
[0005] S1. Obtaining social media platform data and performing data preprocessing, wherein the multimodal data includes pictures and travel diary text;
[0006] S2. Extracting scenic spot entity information from images using machine learning landmark recognition methods;
[0007] S3. Count the number of pictures of all identified scenic spots, obtain points of interest and calculate scenic spot data;
[0008] S4, integrating the sentiment information of travel notes text and analyzing user preferences through sentiment scores and extracted keywords;
[0009] S5. Combine the number of pictures and sentiment analysis results to calculate the comprehensive recommendation score and generate a recommended list of attractions.
[0010] Data preprocessing includes using the perceptual hashing algorithm to calculate the hash value of each image, treating images with the same hash value as duplicate images and deleting one of the duplicate images, deleting duplicate travel text data, deleting images that do not contain scenic spot-related images, and deleting travel text data that does not contain scenic spot-related information and emotional keywords, thereby obtaining an image dataset for scene recognition and a travel text dataset containing emotional information.
[0011] S2 includes calling the image recognition API to identify and extract potential landmarks from the image, setting a landmark relevance threshold for the scenic spot, and returning the landmark name as the scenic spot entity information if the relevance of the landmark extracted from the image is higher than the set threshold.
[0012] S3 includes obtaining point of interest data of relevant attractions through the map, obtaining detailed information of the attractions and the user visit frequency of the corresponding attractions on the map, counting the maximum visit frequency of the points of interest of the attractions in the map, and for the identified attractions, counting the number of corresponding attractions in the image dataset. The detailed information includes location and name.
[0013] S4 includes using the SnowNLP module to calculate the sentiment score in the travel text data. The sentiment analysis result of the SnowNLP Chinese natural language sentiment analysis library is a decimal number, which assigns a value between 0 and 1 to the travel text data, representing the tourist's sentiment level from negative to positive, with 0 representing the most negative and 1 representing the most positive. Let x be the comment text shared by the tourist:
[0014] x={a1,a2...,a n};
[0015] Where a1, a2, ..., a n It is text attribute information, including sentiment words, emoticons, sentence types, degree adverbs, negation words, punctuation marks, place name and scenic spot keywords, and time information.
[0016] Let C be the set of emotion categories:
[0017] C={y1,y2...,y m};
[0018] Where y1, y2, ..., y m The sentiment categories include extremely negative, negative, neutral, positive and extremely positive. Sentiment scores between 0 and 0.2 are extremely negative, sentiment scores between 0.2 and 0.4 are negative, sentiment scores between 0.4 and 0.5 are neutral, sentiment scores between 0.6 and 0.8 are positive, and sentiment scores between 0.8 and 1 are positive.
[0019] The calculation formula for the posterior probability is:
[0020]
[0021] In the formula, P(y j |x) is the condition that the text attribute x belongs to the sentiment category y j The posterior probability of a i is the i-th value in x, y j is the jth value in C, P(a i |y i ) is in the emotion category y j down, a i The conditional probability of occurrence, P(y j ) is the sentiment category y j The prior probability of .
[0022] S4 includes integrating travel text data, using the pre-trained deep learning model RoBERTa to extract sentiment keywords, using a word segmenter to convert the travel text data into RoBERTa's input format, passing the preprocessed input data to RoBERTa, RoBERTa outputs the sentiment tendency of each travel text, and uses RoBERTa's attention mechanism to extract sentiment-related keywords. The attention weight reflects the importance of each word in sentiment judgment. Based on the attention weight, words with high weight are selected as sentiment-related keywords.
[0023] S5 includes:
[0024] Recommendation score = w photo P photo +W sentiment s text +W poi P poi +W image I relevance ;
[0025]
[0026]
[0027]
[0028]
[0029] Where W photo is the weight that reflects the importance of the relevance of the image content to the recommendation, P photo is the result of normalizing the number of pictures of scenic spots, W sentiment is the weight that reflects the influence of user emotions in recommendation, S text is the overall sentiment score of each attraction, W poiis the weight that reflects the influence of the real popularity of scenic spots in the recommendation, P poi is the result of normalization of the POI heat of the scenic spot, W image is the weight that reflects the importance of the relevance of the image content to the recommendation, I relevance is the normalized result of the relevance score of the landmark, sentiment score i is the sentiment score value of the i-th sentiment category, and POI is a point of interest.
[0030] Calculate W photo 、W sentiment 、W poi 、W image When the target layer is the comprehensive recommendation score, the quasi-measurement layer is the key indicators that affect the recommendation, including the number of photos, text sentiment score, POI visit frequency and image content relevance. The solution layer is the final recommendation list of candidate attractions. Based on the comparison of the importance of the quasi-measurement layer indicators pairwise by expert scores, a judgment matrix is established to calculate W respectively. photo 、W sentiment 、W poi and W image The numerical values of the four indicator weights.
[0031] Compared with the existing technology, the present invention has the following beneficial effects: through multimodal data fusion, the present invention can integrate a variety of different types of data information, conduct a comprehensive analysis and understanding of scenic spots from multiple dimensions, grasp the characteristics and connotations of scenic spots, use automatic extraction of scenic spot entity information, efficiently obtain key information of scenic spots, and provide faster and more convenient data support for subsequent recommendations and other links, making the entire recommendation system run more efficiently. The recommendation method based on sentiment analysis can deeply explore the user's preference and emotional tendencies for scenic spots, not only focusing on the user's behavioral data, but also on the user's emotional experience, thereby enhancing the user's trust and reliance on the recommendation system. Matching user interests and scenic spot characteristics means that the recommendation system can deeply understand the user's interests, hobbies, preferences, etc., and carefully compare and match them with the unique characteristics, cultural connotations, and landscape features of the scenic spot. This matching method can recommend scenic spots that truly meet the user's interests and hobbies to the user, thus achieving a match between user interests and scenic spot characteristics. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flow chart of the present invention;
[0033] Figure 2 This is the structural diagram of the SnowNLP sentiment scoring algorithm. DETAILED DESCRIPTION
[0034] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0035] A method for recommending attractions based on multimodal data from social media, comprising:
[0036] S1. Obtaining social media platform data and performing data preprocessing, wherein the multimodal data includes pictures and travel diary text;
[0037] S2. Extracting scenic spot entity information from images using machine learning landmark recognition methods;
[0038] S3. Count the number of pictures of all identified scenic spots, obtain points of interest and calculate scenic spot data;
[0039] S4, integrating the sentiment information of travel notes text and analyzing user preferences through sentiment scores and extracted keywords;
[0040] S5. Combine the number of pictures and sentiment analysis results to calculate the comprehensive recommendation score and generate a recommended list of attractions.
[0041] Data preprocessing includes using the perceptual hashing algorithm to calculate the hash value of each image, treating images with the same hash value as duplicate images and deleting one of the duplicate images, deleting duplicate travel text data, deleting images that do not contain scenic spot-related images, and deleting travel text data that does not contain scenic spot-related information and emotional keywords, thereby obtaining an image dataset for scene recognition and a travel text dataset containing emotional information.
[0042] S2 includes calling the image recognition API to identify and extract potential landmarks from the image, setting a landmark relevance threshold for the scenic spot, and returning the landmark name as the scenic spot entity information if the relevance of the landmark extracted from the image is higher than the set threshold.
[0043] S3 includes obtaining point of interest data of relevant attractions through the map, obtaining detailed information of the attractions and the user visit frequency of the corresponding attractions on the map, counting the maximum visit frequency of the points of interest of the attractions in the map, and for the identified attractions, counting the number of corresponding attractions in the image dataset. The detailed information includes location and name.
[0044] S4 includes using the SnowNLP module to calculate the sentiment score in the travel text data. The sentiment analysis result of the SnowNLP Chinese natural language sentiment analysis library is a decimal number, which assigns a value between 0 and 1 to the travel text data, representing the tourist's sentiment level from negative to positive, with 0 representing the most negative and 1 representing the most positive. Let x be the comment text shared by the tourist:
[0045] x={a1,a2...,a n};
[0046] Where a1, a2, ..., a n It is text attribute information, including sentiment words, emoticons, sentence types, degree adverbs, negation words, punctuation marks, place name and scenic spot keywords, and time information.
[0047] Let C be the set of emotion categories:
[0048] C={y1,y2...,y m};
[0049] Where y1, y2, ..., y m The sentiment categories include extremely negative, negative, neutral, positive and extremely positive. Sentiment scores between 0 and 0.2 are extremely negative, sentiment scores between 0.2 and 0.4 are negative, sentiment scores between 0.4 and 0.5 are neutral, sentiment scores between 0.6 and 0.8 are positive, and sentiment scores between 0.8 and 1 are positive.
[0050] The calculation formula for the posterior probability is:
[0051]
[0052] Where, P(y j |x) is the condition that the text attribute x belongs to the sentiment category y j The posterior probability of a i is the i-th value in x, y j is the jth value in C, P(a i |y i ) is in the emotion category y j down, a i The conditional probability of occurrence, P(y j ) is the sentiment category y j The prior probability of .
[0053] S4 includes integrating travel text data, using the pre-trained deep learning model RoBERTa to extract sentiment keywords, using a word segmenter to convert the travel text data into RoBERTa's input format, passing the preprocessed input data to RoBERTa, RoBERTa outputs the sentiment tendency of each travel text, and uses RoBERTa's attention mechanism to extract sentiment-related keywords. The attention weight reflects the importance of each word in sentiment judgment. Based on the attention weight, words with high weight are selected as sentiment-related keywords.
[0054] S5 includes:
[0055] Recommendation score = W photo P photo +W sentiment S text +W poi P poi +W image I relevance ;
[0056]
[0057]
[0058]
[0059]
[0060] Where W photo is the weight that reflects the importance of the relevance of the image content to the recommendation, P photo is the result of normalizing the number of pictures of scenic spots, W sentiment is the weight that reflects the influence of user emotions in recommendation, S text is the overall sentiment score of each attraction, W poi is the weight that reflects the influence of the real popularity of scenic spots in the recommendation, P poi is the result of normalization of the POI heat of the scenic spot, W image is the weight that reflects the importance of the relevance of the image content to the recommendation, I relevance is the normalized result of the relevance score of the landmark, sentiment score i is the sentiment score value of the i-th sentiment category, and POI is a point of interest.
[0061] Calculate W photo 、W sentiment 、W poi 、W imageWhen the target layer is the comprehensive recommendation score, the quasi-measurement layer is the key indicators that affect the recommendation, including the number of photos, text sentiment score, POI visit frequency and image content relevance. The solution layer is the final recommendation list of candidate attractions. Based on the comparison of the importance of the quasi-measurement layer indicators pairwise by expert scores, a judgment matrix is established to calculate W respectively. photo 、W sentiment 、W poi and W image The numerical values of the four indicator weights.
[0062] The technical process of the present invention is as follows Figure 1 As shown, it includes four parts: image and text data acquisition and processing, image information processing, text information processing, and weighted calculation of scenic spot recommendation scores. In the image and text data acquisition and processing, data collection and data cleaning are performed based on the Qunar online website (an example of a website in an embodiment of the present invention), and image and text data are obtained and sent to the image information processing and text information processing parts. In the image information processing, the name of the scenic spot landmark is recognized in combination with Baidu Image (Baidu Map is an example of an embodiment of the present invention), and detailed information of the scenic spot is obtained in combination with Baidu Map, and the number of pictures of each scenic spot is counted. In the text information processing, the text sentiment score is extracted by SnowNLP, and text keywords are extracted. The results of the two parts of image information processing and text information processing are input into the weighted calculation of the scenic spot recommendation score, and the scenic spot recommendation score is obtained by combining the number of photos, text sentiment score, frequency of visit to the geographical location of the scenic spot, and relevance of the photo content.
[0063] SnowNLP sentiment scoring algorithm Figure 2 As shown in the figure, a text dataset is obtained, sentence segmentation is performed, and then sentiment words and their positions are identified. Degree word processing and negation word processing are performed respectively. Then, weighted calculation is performed, and emoticon processing and sentence processing are performed respectively. Finally, the sentiment score is obtained.
[0064] The present invention is verified by taking a certain scenic spot as an example. First, a judgment matrix is established according to the expert scores as shown in Table 1.
[0065] Table 1 Judgment Matrix
[0066]
[0067] The steps to calculate the weight using the eigenvector method are as follows:
[0068] Calculate the geometric mean of each row including the number of photos
[0069] The sentiment of the text is
[0070] POI frequency is
[0071] The image relevance is
[0072] The normalized weights include,the sum of 2.34+1.14+0.80+0.35=4.90;
[0073] The number of photos is
[0074] The sentiment of the text is
[0075] POI frequency is
[0076] The image relevance is
[0077] The final weight is W photo =0.41;
[0078] W sentiment =0.32;
[0079] W poi =0.17;
[0080] W image =0.10;
[0081] The number of photos of a certain scenic spot is 21,450, and the maximum number of photos of all scenic spots is 23,315.
[0082]
[0083] The frequency of POI visits to scenic spots is 2,438,300 times / year, and the maximum POI visit frequency is 3,295,000 times / year;
[0084]
[0085] The content relevance score of scenic spot pictures is 0.60, and the maximum relevance score of all pictures is 0.93;
[0086]
[0087] Recommendation score: 0.41×0.92+0.32×0.83+0.17×0.74+0.10×0.65=0.8336.
[0088] The above results combined with the results of other scenic spots are shown in Table 2.
[0089] Table 2 Score results of recommended attractions
[0090] Attractions in a city <![CDATA[W photo ]]> <![CDATA[P phtot ]]> <![CDATA[W sentiment ]]> <![CDATA[S text ]]> <![CDATA[W poi ]]> <![CDATA[P poi ]]> <![CDATA[W image ]]> <![CDATA[I relevance ]]> Recommendation score A scenic spot 0.41 0.92 0.32 0.83 0.17 0.74 0.10 0.65 0.8336 A scenic spot 0.41 0.87 0.32 0.78 0.17 0.69 0.10 0.56 0.7800 Scenic Area B 0.41 0.81 0.32 0.72 0.17 0.67 0.10 0.54 0.7236 Scenic Area C 0.41 0.76 0.32 0.67 0.17 0.58 0.10 0.49 0.6736 D Scenic Area 0.41 0.7 0.32 0.61 0.17 0.52 0.10 0.43 0.6136 Scenic Area E 0.41 0.64 0.32 0.55 0.17 0.46 0.10 0.37 0.5536 ;
[0091] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for recommending scenic spots based on multimodal data from social media, characterized in that: include: S1. Obtaining social media platform data and performing data preprocessing, wherein the multimodal data includes pictures and travel diary text; S2. Extracting scenic spot entity information from images using machine learning landmark recognition methods; S3. Count the number of pictures of all identified scenic spots, obtain points of interest and calculate scenic spot data; S4, integrating the sentiment information of travel notes text and analyzing user preferences through sentiment scores and extracted keywords; S5. Combine the number of pictures and sentiment analysis results to calculate the comprehensive recommendation score and generate a recommended list of attractions.
2. The method for recommending scenic spots based on multimodal data from social media according to claim 1, characterized in that: Data preprocessing includes using the perceptual hashing algorithm to calculate the hash value of each image, treating images with the same hash value as duplicate images and deleting one of the duplicate images, deleting duplicate travel text data, deleting images that do not contain scenic spot-related images, and deleting travel text data that does not contain scenic spot-related information and emotional keywords, thereby obtaining an image dataset for scene recognition and a travel text dataset containing emotional information.
3. The method for recommending scenic spots based on multimodal data from social media according to claim 2, characterized in that: S2 includes calling the image recognition API to identify and extract potential landmarks from the image, setting a landmark relevance threshold for the scenic spot, and returning the landmark name as the scenic spot entity information if the relevance of the landmark extracted from the image is higher than the set threshold.
4. The method for recommending scenic spots based on multimodal data from social media according to claim 3, characterized in that: S3 includes obtaining point of interest data of relevant attractions through the map, obtaining detailed information of the attractions and the user visit frequency of the corresponding attractions on the map, counting the maximum visit frequency of the points of interest of the attractions in the map, and for the identified attractions, counting the number of corresponding attractions in the image dataset. The detailed information includes location and name.
5. The method for recommending scenic spots based on multimodal data from social media according to claim 4, characterized in that: S4 includes using the SnowNLP module to calculate the sentiment score in the travel text data. The sentiment analysis result of the SnowNLP Chinese natural language sentiment analysis library is a decimal number, which assigns a value between 0 and 1 to the travel text data, representing the tourist's sentiment level from negative to positive, with 0 representing the most negative and 1 representing the most positive. Let x be the comment text shared by the tourist: x={a1,a2…,a n }; Where a1, a2, ..., a n It is text attribute information, including sentiment words, emoticons, sentence types, degree adverbs, negation words, punctuation marks, place name and scenic spot keywords, and time information.
6. The method for recommending scenic spots based on multimodal data from social media according to claim 5, characterized in that: Let C be the set of emotion categories: C={y1,y2...,y m }; Where, y1, y2, ..., ym are sentiment categories, including extremely negative, negative, neutral, positive, and extremely positive. Sentiment scores between 0 and 0.2 are extremely negative, sentiment scores between 0.2 and 0.4 are negative, sentiment scores between 0.4 and 0.5 are neutral, sentiment scores between 0.6 and 0.8 are positive, and sentiment scores between 0.8 and 1 are positive.
7. The method for recommending scenic spots based on multimodal data from social media according to claim 6, characterized in that: The calculation formula for the posterior probability is: In the formula, P(y j |x) is the condition that the text attribute x belongs to the sentiment category y j The posterior probability of a i is the i-th value in x, y j is the jth value in C, P(a i |y i ) is in the emotion category y j down, a i The conditional probability of occurrence, P(y j ) is the sentiment category y j The prior probability of .
8. The method for recommending scenic spots based on multimodal data from social media according to claim 7, characterized in that: S4 includes integrating travel text data, using the pre-trained deep learning model RoBERTa to extract sentiment keywords, using a word segmenter to convert the travel text data into RoBERTa's input format, passing the preprocessed input data to RoBERTa, RoBERTa outputs the sentiment tendency of each travel text, and uses RoBERTa's attention mechanism to extract sentiment-related keywords. The attention weight reflects the importance of each word in sentiment judgment. Based on the attention weight, words with high weight are selected as sentiment-related keywords.
9. The method for recommending scenic spots based on multimodal data of social media according to claim 8, wherein S5 include: Recommendation score = w photo P photo +W sentiment s text +W poi P poi +W image I relevance ; Where W photo is the weight that reflects the importance of the relevance of the image content to the recommendation, P photo is the result of normalizing the number of pictures of scenic spots, W sentiment is the weight that reflects the influence of user emotions in recommendation, S text is the overall sentiment score of each attraction, W poi is the weight that reflects the influence of the real popularity of scenic spots in the recommendation, P poi is the result of normalization of the POI heat of the scenic spot, W image is the weight that reflects the importance of the relevance of the image content to the recommendation, I relevance is the normalized result of the relevance score of the landmark, sentiment score i is the sentiment score value of the i-th sentiment category, and POI is a point of interest.
10. The method for recommending scenic spots based on multimodal data from social media according to claim 9, characterized in that: Calculate W photo 、W sentiment 、W poi 、W image When the target layer is the comprehensive recommendation score, the quasi-measurement layer is the key indicators that affect the recommendation, including the number of photos, text sentiment score, POI visit frequency and image content relevance. The solution layer is the final recommendation list of candidate attractions. Based on the comparison of the importance of the quasi-measurement layer indicators pairwise by expert scores, a judgment matrix is established to calculate W respectively. photo 、W sentiment 、W poi and W image The numerical values of the four indicator weights.