Tourism preference recommendation method based on multi-modal data fusion and dynamic modeling
The multi-modal data fusion and dynamic modeling approach in tourism recommendations addresses the limitations of single-modal systems by integrating diverse data sources and user behavior analytics to provide personalized and timely recommendations.
Patent Information
- Application Number
- CN202510594218.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-15
AI Technical Summary
The existing travel recommendation system mainly relies on single modal data, cannot comprehensively and accurately capture users' travel preferences, and is difficult to adapt to the dynamic changes in user interests, resulting in the lack of comprehensiveness and attractiveness of recommended content and is susceptible to data noise.
Multimodal data fusion technology is adopted to extract the characteristics of multiple modal data through natural language processing, computer vision and audio processing technology, combine dynamic modeling technology to capture user preferences in real time, use graph neural networks and generative models to generate personalized recommended content, and introduce attention mechanisms and knowledge bases for incremental learning.
It realizes a comprehensive and meticulous portrait of users' travel preferences, can capture changes in interests in real time, generate high-quality and personalized recommended content, and improves the accuracy of recommendations and user satisfaction.
Smart Images

Figure CN120316355A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of media convergence and data analysis, and particularly relates to a tourism preference recommendation method for multimodal data fusion and dynamic modeling. Background Art
[0002] With the rapid development of the tourism industry, tourism information has grown explosively, and it is difficult for users to quickly find tourism products and services that meet their preferences in the vast amount of tourism information. Most of the existing tourism recommendation systems are based on single-modal data (such as only based on the text information of the tourism web pages browsed by users) or simple user behavior statistics for recommendation, which cannot comprehensively and accurately capture users' tourism preferences and are difficult to adapt to the dynamic changes of users' interests. Multimodal data refers to data that integrates multiple modal forms (such as text, pictures, audio, video, etc.), and each modal data has unique expression methods and information characteristics. In the tourism recommendation scenario, data in various forms, such as the text introduction of tourist attractions, beautiful pictures, promotional videos, and real audio evaluations of tourists, provide a rich and comprehensive information basis for tourism recommendation.
[0003] Currently, the mainstream tourism recommendation systems on the market mainly adopt recommendation strategies based on single-dimensional data such as user ratings and browsing history. These systems model users' interest preferences through simple algorithms (such as content-based recommendation algorithms and collaborative filtering algorithms) and make recommendations according to the model results. For example, recommend scenic spots with similar ratings based on the ratings of the tourism scenic spots browsed by users in the past. Moreover, these systems only use single-modal data and ignore the rich multimodal data such as pictures, audio, and video in tourism information, resulting in the lack of comprehensiveness and attractiveness of the recommended content, and being unable to capture the changes in users' tourism preferences in real time, so that the recommended content cannot adapt to users' current interest needs in a timely manner. At the same time, the existing systems have simple modeling methods for users' interest preferences and are easily affected by data noise, resulting in a large deviation between the recommended results and users' actual needs. Summary of the Invention
[0004] In view of the above deficiencies, the present invention discloses a tourism preference recommendation method for multimodal data fusion and dynamic modeling, which fuses multimodal data, uses dynamic modeling technology to capture users' tourism preferences in real time, combines RAG flow to generate personalized recommendation content, and at the same time introduces advanced technologies such as attention mechanism and graph neural network, and relies on the knowledge base for incremental learning, effectively improving the accuracy of tourism recommendation and user satisfaction.
[0005] The present invention is implemented by the following technical solutions:
[0006] A tourism preference recommendation method for multimodal data fusion and dynamic modeling, comprising the following steps:
[0007] (1) Collect data of text, pictures, audio, and video related to tourist attractions from travel websites, social media, and travel forums; extract semantic features such as keywords, themes, and sentiment tendencies in the text using semantic analysis algorithms in natural language processing (NLP) technology; use image recognition models in the field of computer vision (CV) to identify pictures and images to obtain image features of scenic elements, landscape features, and color styles; use audio feature extraction algorithms in audio processing technology to extract the content of audio to obtain audio features of intonation, emotion, and keywords; divide the video into several frame images and audio information, and then extract image features and audio features respectively;
[0008] (2) Vectorize the semantic features, image features, and audio features obtained in step (1) to obtain semantic feature vectors, image feature vectors, and audio feature vectors. Then, based on the fusion method of the attention mechanism, combine the semantic feature vectors, image feature vectors, and audio feature vectors to obtain a comprehensive feature vector. Specifically, let the text feature vector be Vt, the image feature vector be Vi, and the audio feature vector be Va. Then, the fused comprehensive feature vector Vm can be expressed as: v m =αv t +βv i +γv a , where α, β, and γ are weight coefficients; store the data of text, pictures, audio, and video related to tourist attractions collected and the corresponding comprehensive feature vectors to build a knowledge base;
[0009] (3) Under the user's authorization, collect the user's travel behavior data at different times through the platform's log recording system and user tracking technology. The travel behavior data includes the travel web pages browsed by the user, the travel keywords searched, and the tourist attractions collected. And use the geographic information system GIS method to obtain the user's geographical location information; extract key features such as browsing time, search keyword frequency, and collected scenic spot types from the travel behavior data, vectorize the key features and geographical location information respectively, and then combine them to build a user feature vector; build a prediction model based on the long short-term memory network (LSTM) and graph neural network (GNN), use the user feature vector as the input of the prediction model, predict the user's travel preference feature vector, and use the historically collected user feature vectors to train and optimize the prediction model;
[0010] (4) Compare according to the user's travel preference feature vector obtained in step (3) in the knowledge base described in step (1). Specifically, use the approximate nearest neighbor search (ANN) algorithm to retrieve several comprehensive feature vectors that are the same or similar to the user's travel preference feature vector. Let the user's travel preference feature vector be Up, and the comprehensive feature vector in the knowledge base be Vm. Then, the cosine similarity sim(Up, Vm) between the two can be expressed as:
[0011] (5) Use a generative model to generate several recommended copywriting based on the data corresponding to the comprehensive feature vectors retrieved in step (4). Let the generated recommended copywriting be y, then the generation process is expressed as y = G(Vm), where G is the generative model function, and the generative model includes a generative model based on GPT and a generative model based on DeepSeek; then screen and sort the recommended copywriting according to the user's geographical location information, time, and budget. During the sorting process, use a sorting algorithm to sort and push the recommended content to the user according to the matching degree between the recommended content and the user's interests, the popularity of the content, and the user's historical preferences, and rank the travel recommendations that best meet the user's needs at the top. Specifically, let the comprehensive score of the recommended copywriting be S, which can be expressed as: S = w1sim(u p , v m ) + w2H + w3P, where W1, W2, and W3 are weight coefficients, H is the popularity of the content, and P is the historical preference score of the user. The present invention adopts the RAG flow technology, that is, the Retrieval-Augmented Generation process, which combines retrieval technology and generation technology. When generating recommended content, first retrieve information related to the user's needs from a large-scale knowledge base, and then generate content based on the retrieval results, integrating the accuracy of retrieval and the flexibility of generation, and improving the accuracy and richness of the recommended content.
[0012] Further, in step (1), use the semantic analysis algorithm in natural language processing (NLP) technology, that is, the pre-trained language model BERT (Bidirectional Encoder Representations from Transformers) based on deep learning is pre-trained through historical text data to learn language representations, and then used to extract semantic features such as keywords, themes, and sentiment tendencies in the text related to tourist attractions; use the image recognition model in the field of computer vision (CV), that is, the object detection model YOLO (You Only Look Once) based on a convolutional neural network (CNN), to identify pictures and images to obtain image features such as scenic spot elements, landscape features, and color styles; use the audio feature extraction algorithm in audio processing technology, that is, the Mel Frequency Cepstral Coefficients (MFCC) extraction algorithm, to extract the content of the audio to obtain audio features such as intonation, emotion, and keywords. The BERT model learns rich language representations through pre-training on large-scale text data and can better understand the semantic information of the text; the YOLO model has fast and accurate object detection capabilities and can identify multiple objects in the image in real time.
[0013] Further, in step (1), data such as text, pictures, audio, and video related to tourist attractions are collected from tourist websites, social media, and tourist forums. After preprocessing the data, it is vectorized. The preprocessing includes deduplication, denoising, format unification, filling in missing data, tokenizing text and performing part-of-speech tagging, cropping and scaling images, and denoising and normalizing audio. In step (3), with the user's authorization, tourist behavior data of the user at different time periods is collected through the platform's log recording system and user tracking technology, and after cleaning, denoising, and normalization in sequence, key features are extracted.
[0014] Further, during the training process of the prediction model in step (3), the cross-entropy loss function L is used to measure the difference between the result predicted by the model and the true label. L can be expressed as:
[0015] where N is the number of samples, y i is the true label, is the probability predicted by the model.
[0016] Further, in step (5), the recommended content is sorted and pushed to the user, that is, it is displayed to the user in the form of a tourist recommendation page and APP push.
[0017] Further, in step (5), a user feedback mechanism is set up to collect the user's likes, comments, and improvement suggestions on the recommended content in real time, which are used to adjust and optimize the generation model.
[0018] The technical solution of the present invention has the following beneficial effects compared with the prior art:
[0019] 1. The present invention uses multi-modal data fusion technology to comprehensively and deeply mine diverse information related to tourism. For the text introduction of tourist attractions, semantic analysis algorithms in natural language processing (NLP) technology are used to extract semantic features such as keywords, themes, and sentiment tendencies; for tourist pictures, image recognition models in the field of computer vision (CV) are used to identify information such as scenic elements, landscape features, and color styles in the images; for tourist audio, audio feature extraction algorithms in audio processing technology are used to analyze the intonation, emotion, and key information of the audio content; at the same time, combining the multi-frame images and audio information of tourist videos, multi-modal fusion algorithms, such as fusion methods based on attention mechanisms, are used to fuse the features of different modalities to form a comprehensive and detailed tourist content portrait.
[0020] 2. The present invention utilizes dynamic modeling technology to capture users' interests in real time. By collecting users' travel behavior data at different time periods, such as the travel web pages browsed, the travel keywords searched, the travel attractions bookmarked, etc., and combining with the users' geographical location information (obtained by using Geographic Information System (GIS) technology), a dynamic model of users' travel preferences is constructed. A method combining time series analysis methods (such as Long Short-Term Memory Network (LSTM)) and Graph Neural Network (GNN) is adopted to predict the changing trend of users' interests. The LSTM network can handle the long-term dependencies in time series data and capture the changing patterns of users' interests over time. The GNN is used to mine the correlation between users and travel attractions, as well as the interest similarity among users.
[0021] 3. The present invention retrieves relevant travel content from the travel recommendation knowledge base according to the user's current travel preference vector. The knowledge base adopts vector retrieval technology, such as an algorithm based on Approximate Nearest Neighbor Search (ANN), to quickly and accurately retrieve the content most relevant to the user's interests, and then uses a generation model (such as the GPT-3.5 series model) to generate content based on the retrieved content to improve the quality and relevance of the generated content.
[0022] 4. In the recommendation process, the present invention fully considers the users' real-time travel behavior and geographical location information, uses GIS technology and spatio-temporal analysis methods to improve the accuracy and practicality of the recommendation. By using the RAG flow technology, combining the advantages of retrieval and generation, an attention mechanism can be introduced to generate high-quality personalized travel recommendation content, and the generation model can be optimized in real time according to user feedback. The LSTM network is used to process time series data to capture the changing patterns of users' interests over time, and the GNN is combined to mine the correlation between users and travel attractions to improve the accuracy of predicting users' interest preferences. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a flowchart of the travel preference recommendation method for multi-modal data fusion and dynamic modeling described in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The present invention will be further described below through embodiments, but it is not intended to limit the present invention. The specific experimental conditions and methods not specified in the following embodiments are usually conventional means well known to those skilled in the art.
[0025] The following embodiments are applied in conjunction with the three-level integrated knowledge base established by the applicant of the present invention. Relevant data resources can be obtained from the three-level integrated knowledge base. The three-level integrated knowledge base covers the media history of the whole Guangxi Zhuang Autonomous Region and newly added manuscript resources, and integrates various tourism-related data, such as tourism websites, social media, tourism forums, etc. It collects multi-modal data such as text, pictures, audio, and video of tourist attractions, and constructs a special tourism recommendation classification library to collect, clean, and preprocess tourism-related multi-modal data. It can also update information in real time, such as newly added tourist attractions, activities, etc., to ensure the timeliness and accuracy of the recommended content. At the same time, the knowledge base uses vector representation technology to convert tourism content into vector form for subsequent retrieval and similarity calculation.
[0026] Embodiment 1: As Figure 1 shown, a tourism preference recommendation method for multi-modal data fusion and dynamic modeling includes the following steps:
[0027] (1) Collect data of text, pictures, audio, and video related to authorized tourist attractions; use the semantic analysis algorithm in natural language processing (NLP) technology, that is, the pre-trained language model BERT (Bidirectional Encoder Representations from Transformers) based on deep learning to perform pre-training through historical text data, learn language representations, and then use it to extract semantic features such as keywords, themes, and sentiment tendencies in the text related to tourist attractions; use the image recognition model in the field of computer vision (CV), that is, the object detection model YOLO (You Only Look Once) based on convolutional neural network (CNN), to identify pictures and images to obtain image features such as scenic spot elements, landscape features, and color styles; use the audio feature extraction algorithm in audio processing technology, that is, the Mel Frequency Cepstral Coefficients (MFCC) extraction algorithm, to extract the content of audio to obtain audio features such as intonation, emotion, and keywords; collect data of text, pictures, audio, and video related to tourist attractions from tourism websites, social media, and tourism forums, and perform vectorization after preprocessing the data. The preprocessing includes duplicate removal, noise reduction, format unification, filling in missing data, word segmentation and part-of-speech tagging for text, cropping and scaling for images, and noise reduction and normalization for audio.
[0028] (2) Vectorize the semantic features, image features, and audio features obtained in step (1) to obtain semantic feature vectors, image feature vectors, and audio feature vectors. Then, based on the fusion method of the attention mechanism, combine the semantic feature vectors, image feature vectors, and audio feature vectors to obtain a comprehensive feature vector. Specifically, let the text feature vector be Vt, the image feature vector be Vi, and the audio feature vector be Va. Then, the fused comprehensive feature vector Vm can be expressed as: v m = αv t + βv i + γv a , where α, β, and γ are weight coefficients; store the data of texts, pictures, audio, and videos related to tourist attractions collected and the corresponding comprehensive feature vectors to build a knowledge base; among them, attributes or fields representing the characteristics of tourist data can be selected, such as keywords of texts, feature vectors of images, MFCC features of audio, etc., and the features are converted into vector forms using vectorization techniques. For example, the text is converted into word vectors using the word embedding technique, and the image is converted into a feature vector using a pre-trained CNN model. Finally, the vectorized data is stored in a database or storage system that supports vector operations, such as a vector retrieval library based on Faiss, for subsequent similarity calculation and retrieval;
[0029] (3) Collect the tourist behavior data of users at different time periods through the platform's log recording system and user tracking technology under user authorization. Among them, dynamic modeling technology can be used to capture users' interests in real time, focusing on the dynamic changes in users' tourist preferences; the tourist behavior data includes the tourist web pages browsed by users, the tourist keywords searched, and the tourist attractions collected, and the geographical location information of users is obtained using the Geographic Information System (GIS) method; extract key features such as browsing time, search keyword frequency, and collected scenic spot types from the tourist behavior data, vectorize the key features and geographical location information respectively, and then combine them to construct a user feature vector; build a prediction model based on the Long Short-Term Memory Network (LSTM) and the Graph Neural Network (GNN). Use the user feature vector as the input of the prediction model to predict the user's tourist preference feature vector, and use the historically collected user feature vectors to train and optimize the prediction model; collect the tourist behavior data of users at different time periods through the platform's log recording system and user tracking technology under user authorization, and extract key features after cleaning, denoising, and normalizing in turn; during the training process of the prediction model, the cross-entropy loss function L is used to measure the difference between the result predicted by the model and the true label. L can be expressed as:
[0030] where N is the number of samples, y i is the true label, is the probability predicted by the model;
[0031] (4) Compare the user's travel preference feature vector obtained in step (3) in the knowledge base described in step (1). Specifically, use the approximate nearest neighbor search (ANN) algorithm to retrieve several comprehensive feature vectors that are the same as or similar to the user's travel preference feature vector. Let the user's travel preference feature vector be Up, and the comprehensive feature vector in the knowledge base be Vm. Then the cosine similarity sim(Up, Vm) between the two can be expressed as:
[0032] (5) Use a generation model to generate several recommended texts based on the data corresponding to the comprehensive feature vectors retrieved in step (4). Let the generated recommended text be y. Then the generation process is expressed as y = G(Vm), where G is the generation model function. The generation model includes a GPT-based generation model and a Deepseek-based generation model. Then, screen and sort the recommended texts according to the user's geographical location information, time, and budget. During the sorting process, use the LambdaMART sorting algorithm to sort the recommended content according to the matching degree between the recommended content and the user's interests, the popularity of the content, and the user's historical preferences, and push it to the user. And rank the travel recommendations that best meet the user's needs at the top. Specifically, let the comprehensive score of the recommended text be S, which can be expressed as: S = w1sim(u p , v m ) + w2H + w3P, where W1, W2, and W3 are weight coefficients, H is the popularity of the content, and P is the user's historical preference score.
[0033] The method described in this embodiment can be applied by setting a suitable device. For example, set a device composed of a multi-modal data fusion unit, a user travel preference dynamic modeling unit, a RAG flow recommendation generation unit, and a personalized recommendation unit. Among them, the multi-modal data fusion unit is responsible for collecting and processing multi-modal data such as text, pictures, audio, and video of tourist attractions, extracting features using corresponding algorithms and fusing them to form a tourist content portrait; the user travel preference dynamic modeling unit is responsible for collecting user travel behavior data and geographical location information, constructing a user travel preference dynamic model, and updating the user interest preference vector in real time; the RAG flow recommendation generation unit is responsible for retrieving relevant content from the knowledge base according to the user interest preference vector, generating personalized recommended texts using the generation model, and optimizing them in combination with user feedback; the personalized recommendation unit is responsible for screening and sorting the recommended content in combination with user constraint conditions, displaying the recommendation results, and collecting user feedback.
[0034] Meanwhile, corresponding computer devices can be configured, such as configuring a processor, a memory, and a communication interface; the processor is used to execute the functional codes of the above-mentioned module units to implement relevant algorithms and logics for tourism preference recommendation; the memory is used to store information such as multi-modal tourism data, user behavior data, and model parameters; the communication interface is used to perform data interaction with a tourism platform, a user terminal, etc.; or a hard disk, an optical disc, a USB flash drive, etc. can be used to store the computer program for the processor to implement the recommended method, for long-term preservation of program codes and data.
[0035] Embodiment 2: The difference between the tourism preference recommendation method for multi-modal data fusion and dynamic modeling in this embodiment and the method described in Embodiment 1 is only that in step (5), the recommended content is sorted and pushed to the user, that is, it is displayed to the user in the form of a tourism recommendation page and APP push; and a user feedback mechanism is set up to collect the user's likes, comments, and improvement suggestions on the recommended content in real time for adjusting and optimizing the generation model. For example, in the generation process, an attention mechanism can be introduced to enable the generation model to focus on the key information in the retrieved content, improve the quality and relevance of the generated content, and combine the user's real-time feedback (such as likes, comments, etc.) to adjust and optimize the generation model in real time to improve the quality of the recommended content. In the generation process, the generation model not only based on the retrieved content, but also combines the incremental information obtained from the three-level integrated knowledge base of media convergence in Guangxi Zhuang Autonomous Region, such as the latest tourism trends, user evaluations, etc., to improve the quality and relevance of the generated content.
[0036] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A tourism preference recommendation method for multimodal data fusion and dynamic modeling, characterized in that: Including the following steps: (1) Collect data of texts, pictures, audios and videos related to tourist attractions from tourist websites, social media and tourist forums; use semantic analysis algorithms in natural language processing technology to extract semantic features such as keywords, themes and sentiment tendencies in the texts; use image recognition models in the field of computer vision to identify pictures and images to obtain image features of scenic spot elements, landscape features and color styles; use audio feature extraction algorithms in audio processing technology to extract the content of audios to obtain audio features of intonation, sentiment and keywords; divide videos into several frame images and audio information, and then extract image features and audio features respectively; (2) Vectorize the semantic features, image features, and audio features obtained in step (1) to obtain a semantic feature vector, an image feature vector, and an audio feature vector. Then, based on the fusion method of the attention mechanism, the semantic feature vector, the image feature vector, and the audio feature vector are combined to obtain a comprehensive feature vector. Specifically, let the text feature vector be Vt, the image feature vector be Vi, and the audio feature vector be Va. Then, the fused comprehensive feature vector Vm can be expressed as: v m = αv t + βv i + γv a , where α, β, and γ are weight coefficients; store the data of text, pictures, audio, and video related to tourist attractions collected and the corresponding comprehensive feature vectors to construct a knowledge base; (3) Under the authorization of the user, collect the tourist behavior data of the user at different time periods through the platform's log recording system and user tracking technology. The tourist behavior data includes the tourist web pages browsed by the user, the tourist keywords searched, and the tourist attractions collected. And use the Geographic Information System (GIS) method to obtain the geographical location information of the user; extract key features such as browsing time, search keyword frequency and collected scenic spot types from the tourist behavior data, vectorize the key features and geographical location information respectively, and then combine and construct to obtain the user feature vector; build a prediction model based on the long short-term memory network and the graph neural network, use the user feature vector as the input of the prediction model, predict to obtain the user's tourist preference feature vector, and use the historically collected user feature vectors to train and optimize the prediction model; (4) Compare the user's travel preference feature vector obtained in step (3) with the knowledge base described in step (1). Specifically, use the approximate nearest neighbor search algorithm to retrieve several comprehensive feature vectors that are the same as or similar to the user's travel preference feature vector. Let the user's travel preference feature vector be Up and the comprehensive feature vector in the knowledge base be Vm. Then the cosine similarity sim(Up, Vm) between the two can be expressed as: (5) Use a generative model to generate several recommended copywriting based on the data corresponding to the comprehensive feature vectors retrieved in step (4). Let the generated recommended copywriting be y, then the generation process is expressed as y = G(Vm), where G is the generative model function, and the generative model includes a generative model based on GPT and a generative model based on Deepseek; then screen and sort the recommended copywriting according to the user's geographical location information, time, and budget. During the sorting process, use a sorting algorithm to sort the recommended content according to the matching degree between the recommended content and the user's interests, the popularity of the content, and the user's historical preferences, and push it to the user, and rank the travel recommendations that best meet the user's needs at the top. Specifically, let the comprehensive score of the recommended copywriting be S, which can be expressed as: S = w1sim(u p , v m ) + w2H + w3P, where W1, W2, and W3 are weight coefficients, H is the popularity of the content, and P is the historical preference score of the user.
2. The tourism preference recommendation method for multi-modal data fusion and dynamic modeling according to claim 1, characterized in that: In step (1), use the semantic analysis algorithm in natural language processing technology, that is, the pre-trained language model BERT based on deep learning is pre-trained through historical text data to learn language representations, and then used to extract semantic features such as keywords, themes and sentiment tendencies in the texts related to tourist attractions; use the image recognition model in the field of computer vision, that is, the object detection model YOLO based on convolutional neural network, to identify pictures and images to obtain image features of scenic spot elements, landscape features and color styles; use the audio feature extraction algorithm in audio processing technology, that is, the Mel Frequency Cepstral Coefficient extraction algorithm, to extract the content of audios to obtain audio features of intonation, sentiment and keywords.
3. The tourism preference recommendation method for multimodal data fusion and dynamic modeling according to claim 1, wherein: In step (1), collect data of texts, pictures, audios and videos related to tourist attractions from tourist websites, social media and tourist forums, preprocess the data and then vectorize it. The preprocessing includes duplicate removal, noise removal, format unification, filling in missing data, word segmentation and part-of-speech tagging for texts, cropping and scaling for images, and noise reduction and normalization for audios; in step (3), under the authorization of the user, collect the tourist behavior data of the user at different time periods through the platform's log recording system and user tracking technology, and extract key features after cleaning, noise removal and normalization in sequence.
4. The tourism preference recommendation method for multimodal data fusion and dynamic modeling according to claim 1, characterized in that: During the training process of the prediction model described in step (3), the cross-entropy loss function L is used to measure the difference between the result predicted by the model and the true label. L can be expressed as: where N is the number of samples, and y i is the true label, and is the probability predicted by the model.
5. The tourism preference recommendation method for multi-modal data fusion and dynamic modeling according to claim 1, characterized in that: In step (5), sort the recommended content and push it to the user, that is, display it to the user in the form of a tourist recommendation page and APP push.
6. The tourism preference recommendation method for multimodal data fusion and dynamic modeling according to claim 1, wherein: In step (5), set up a user feedback mechanism to collect the likes, comments and improvement suggestions of the user on the recommended content in real time, and use them to adjust and optimize the generation model.
Citation Information
Cited By
AI tourism personalized explanation system and service system
CN121053717A
Tourism demand prediction method and device based on TFT model, and storage medium
CN121146179A