A method for generating visual tourism interest recommendation information based on tourism photos
By performing multi-dimensional analysis and cross-modal semantic fusion on travel photos, and combining spatiotemporal semantic association mapping and travel interest knowledge graph matching, a multi-level collaborative aggregated interest profile of users is constructed. This solves the problems of single and inefficient generation of interest recommendation information in existing technologies, and realizes the efficient generation of personalized travel recommendations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIAN COLLEGE
- Filing Date
- 2026-03-17
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies struggle to fully extract multi-dimensional information from travel photos, resulting in simplistic interest-based recommendations that fail to accurately reflect users' true travel interests and preferences. Furthermore, they lack efficient spatiotemporal semantic mapping and multi-level collaborative aggregation, hindering the rapid provision of personalized travel recommendation services.
By performing multi-dimensional analysis and cross-modal semantic fusion on travel photos, combined with spatiotemporal semantic association mapping and travel interest knowledge graph matching, core interest concepts are extracted and dynamic interest intensity values are evaluated to construct a multi-level collaborative aggregation of users' travel interest profiles.
It enables the accurate extraction of user preferences and travel interest profiles, improves the efficiency and personalization of visual travel interest recommendation information generation, and enhances user experience.
Smart Images

Figure CN122132632A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data management technology, and in particular to a method for generating visual tourism interest recommendation information based on tourism photos. Background Technology
[0002] In the field of generating travel interest recommendation information, existing technologies often struggle to fully extract the multi-dimensional information contained in travel photos. Insufficient integration of data such as image content, shooting location, and timestamps results in relatively simplistic semantic information that fails to comprehensively reflect users' true travel interests and preferences. This limitation in information mining leads to a lack of accurate data support for subsequent interest recommendations, making it difficult to meet users' personalized needs.
[0003] Meanwhile, existing methods lack efficient spatiotemporal semantic association mapping and multi-level collaborative aggregation mechanisms in the process of transforming tourism-related data into visual recommendation information. This not only leads to insufficient accuracy and completeness in the construction of tourism interest profile data, but also results in low efficiency in generating visual recommendation information. Consequently, they cannot quickly and accurately provide users with valuable guidance on their tourism interests, making it difficult to meet users' needs for efficient and personalized tourism recommendation services. Summary of the Invention
[0004] This invention provides a method for generating visual tourism interest recommendation information based on tourism photos, in order to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides a method for generating visualized tourism interest recommendation information based on tourism photos, comprising: S1. Obtain the user's travel photos and the corresponding shooting location information and shooting timestamp information; S2. Perform image content analysis on the travel photos, and integrate the analysis results, shooting location information, and shooting timestamp information in multiple dimensions to obtain multi-dimensional semantic information of the travel photos; S3. Based on multi-dimensional semantic information, perform spatiotemporal semantic association mapping on travel photos to obtain the user's spatiotemporal semantic layer; S4. Map the spatiotemporal semantic layer to the preset tourism interest knowledge graph to obtain the core interest concepts of the spatiotemporal semantic layer, and evaluate the dynamic interest intensity value of the core interest concepts based on the visual features in the spatiotemporal semantic layer. S5. Based on the hierarchical relationship in the tourism interest knowledge graph, multi-level collaborative aggregation of core interest concepts and dynamic interest intensity values is carried out to obtain user tourism interest profile data. S6. Based on user travel interest profile data and combined with the shooting location information, generate user-visualized travel interest recommendation information.
[0006] In a preferred embodiment, obtaining the user's travel photos and corresponding shooting location information and shooting timestamp information includes: Based on the user's digital identity, the user's travel photos are obtained by retrieving multi-source heterogeneous image storage nodes in the distributed storage space. Metadata encapsulation structure of tourist photos is parsed to obtain embedded metadata of tourist photos; Spatiotemporal information decoding of embedded metadata yields the geographic location and timestamp information of the tourist photos.
[0007] In a preferred embodiment, the step of analyzing the image content of the tourist photos and fusing the analysis results, shooting location information, and shooting timestamp information in multiple dimensions to obtain multi-dimensional semantic information of the tourist photos includes: Multi-granularity scene target analysis is performed on tourism photos to obtain the image semantic analysis results of tourism photos; Geographic semantic tags for travel photos are obtained by performing point-of-interest semantic mapping on the geographical location information of the photos taken. By performing temporal context semantic inference on the shooting timestamp information, temporal activity semantic tags for travel photos are obtained. By performing cross-modal semantic fusion of image semantic parsing results, geographic semantic tags, and temporal activity semantic tags, multi-dimensional semantic information of tourism photos can be obtained.
[0008] In a preferred embodiment, the step of performing spatiotemporal semantic association mapping on travel photos based on multi-dimensional semantic information to obtain the user's spatiotemporal semantic layer includes: Based on multi-dimensional semantic information, spatiotemporal context anchoring is performed on tourism photos to obtain preliminary spatiotemporal correlation data of users; Semantic conflict resolution is performed on the initial spatiotemporal correlation data, and topological reconstruction is performed on the spatiotemporal correlation data after conflict resolution to obtain the user's spatiotemporal semantic relationship network; Based on the spatiotemporal semantic relationship network, multi-dimensional semantic features are aggregated in a structured layered manner to obtain the user's intermediate semantic layer; Semantic integrity is verified on the intermediate semantic layer to obtain the user's spatiotemporal semantic layer.
[0009] In a preferred embodiment, mapping the spatiotemporal semantic layer to a preset tourism interest knowledge graph to obtain the core interest concepts of the spatiotemporal semantic layer includes: Based on the vectorized representation of concept nodes in the tourism interest knowledge graph, cross-space semantic alignment is performed on the spatiotemporal semantic layer to obtain the semantic nodes of the spatiotemporal semantic layer. Multi-dimensional matching of semantic nodes with interest nodes in the tourism interest knowledge graph is performed to obtain the initial matching nodes of the spatiotemporal semantic layer. Based on the semantic association strength between nodes in the tourism interest knowledge graph, the initial matching nodes are ranked by importance to obtain candidate core interest nodes in the spatiotemporal semantic layer. Based on the topological structure of the tourism interest knowledge graph, semantic concept generalization is performed on candidate core interest nodes to obtain core interest concepts in the spatiotemporal semantic layer.
[0010] In a preferred embodiment, evaluating the dynamic interest intensity value of the core interest concept based on visual features in the spatiotemporal semantic layer includes: Extract visual features associated with core interest concepts from the spatiotemporal semantic layer; Based on a pre-defined visual cognition theory, a spatiotemporal contextual saliency analysis is performed on visual features, and weights are assigned to core interest concepts based on the saliency analysis results. By analyzing the distribution patterns of the frequency of visual features over time, temporal distribution data of visual features are obtained. Based on the weight assignment results, time distribution data, and spatial distribution data of visual features in the spatiotemporal semantic layer, the dynamic interest intensity value of the core interest concept is calculated. The formula for calculating the dynamic interest intensity value is as follows: ; In the formula, This represents the dynamic interest intensity value. The total number of visual features associated with the core interest concept. For the first Weights of visual features For the first Normalized frequency of occurrence of visual features in temporal distribution data The preset time decay coefficient, For the first The time interval between the most recent occurrence of a visual feature and the current evaluation baseline time point It is a logarithmic function. This represents the spatial distribution dispersion of visual features at their corresponding geographical locations within the spatiotemporal semantic layer.
[0011] In a preferred embodiment, the step of performing spatiotemporal contextual saliency analysis on visual features based on a preset visual cognition theory, and assigning weights to core interest concepts based on the saliency analysis results, includes: The frequency of visual features in travel photos is statistically analyzed to obtain data on the occurrence frequency of visual features. Elements adjacent to visual features in the spatiotemporal semantic layer are considered as surrounding visual elements. Analyze the degree of difference between visual features and surrounding visual elements in terms of color, texture and outline to obtain visual contrast data of visual features; The degree of fit between visual features and core interest concepts is determined to obtain scene relevance data for visual features; Based on occurrence frequency data, visual contrast data, and scene relevance data, the saliency of visual features is scored to obtain the saliency score results of visual features. Based on the saliency score, the visual features are classified into weight levels, and the core interest concepts associated with the visual features are assigned corresponding weight values.
[0012] In a preferred embodiment, the step of performing multi-level collaborative aggregation of core interest concepts and dynamic interest intensity values based on the hierarchical relationships in the tourism interest knowledge graph to obtain user tourism interest profile data includes: Based on the parent-child and sibling relationships between concept nodes in the tourism interest knowledge graph, the hierarchical path tracing of core interest concepts is performed to obtain the hierarchical distribution structure of core interest concepts. Based on the hierarchical distribution structure, a multi-scale field is constructed for the dynamic interest intensity values to obtain the user's interest intensity distribution map. By performing cross-level semantic convergence on the interest intensity distribution map, we can obtain the user's hierarchical tourism interest themes. Structured profile rendering is performed on the hierarchical tourism interest themes to obtain user tourism interest profile data.
[0013] In a preferred embodiment, the step of constructing a multi-scale field based on the hierarchical distribution structure of dynamic interest intensity values to obtain a user's interest intensity distribution map includes: Based on the node topology in the hierarchical distribution structure, the hierarchical intensity field is initialized for the dynamic interest intensity value to obtain the initial hierarchical intensity field of the dynamic interest intensity value. Based on the coupling relationship between adjacent nodes in the hierarchical distribution structure, the initial hierarchical intensity field is propagated by field effect to obtain the equalized hierarchical intensity field with dynamic interest intensity values. Cross-level multi-scale fusion of the equalized hierarchical intensity field yields the user's interest intensity distribution map.
[0014] In a preferred embodiment, generating visualized travel interest recommendation information for users based on user travel interest profile data and combined with geographic location information includes: Readable recommendation content descriptions are generated from user travel interest profile data to obtain structured descriptive text for users; Based on the coordinates in the geographic location information, the structured descriptive text is annotated with points of interest to obtain the user's annotated text. Visual interactive logic is bound to the labeled text to obtain the user's visual style description text; By integrating multi-source information such as structured descriptive text, labeled content text, and visual style descriptive text, visualized travel interest recommendation information for users can be obtained.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention performs multi-dimensional analysis and cross-modal semantic fusion on travel photos, combines spatiotemporal semantic association mapping and travel interest knowledge graph matching to accurately extract core interest concepts, and uses dynamic interest intensity value quantitative evaluation to comprehensively construct a travel interest profile that fits user preferences, making the recommended information more in line with the user's personalized needs.
[0016] 2. This invention relies on technologies such as multi-level collaborative aggregation and visual interactive logic binding to achieve efficient connection from data integration and interest mining to recommendation generation, which greatly improves the generation efficiency of visual tourism interest recommendation information. At the same time, it combines shooting geographical location information to mark points of interest, making the recommendation information more intuitive and easy to understand, and enhancing the user experience. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a method for generating visual tourism interest recommendation information based on tourism photos, as provided in an embodiment of the present invention.
[0018] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0019] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0020] This application provides a method for generating visual tourism interest recommendation information based on tourist photos. The executing entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, this method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0021] Reference Figure 1 The diagram shown is a flowchart illustrating a method for generating visualized travel interest recommendation information based on travel photos, according to an embodiment of the present invention. In this embodiment, the method for generating visualized travel interest recommendation information based on travel photos includes: S1. Obtain the user's travel photos and the corresponding shooting location information and shooting timestamp information; In this embodiment of the invention, obtaining the user's travel photos and corresponding shooting location information and shooting timestamp information includes: Based on the user's digital identity, the user's travel photos are obtained by retrieving multi-source heterogeneous image storage nodes in the distributed storage space. Metadata encapsulation structure of tourist photos is parsed to obtain embedded metadata of tourist photos; Spatiotemporal information decoding of embedded metadata yields the geographic location and timestamp information of the tourist photos.
[0022] Using the user's digital identity as the sole retrieval criterion, the system traverses all multi-source heterogeneous image storage nodes within the distributed storage space. In each storage node, it retrieves the pre-stored attribution information for each image and performs a one-to-one complete matching and comparison with the user's digital identity. Images whose attribution information differs from the user's digital identity are eliminated, while images whose attribution information completely overlaps with the user's digital identity are retained. The entire set of retained images is then identified as the user's travel photos.
[0023] For the identified user travel photos, based on the preset metadata encapsulation structure hierarchy of the travel photos, the process starts from the outermost encapsulation structure and performs a layer-by-layer disassembly operation. During the disassembly process, redundant data and invalid identifiers attached to each encapsulation structure are stripped off in turn, and the process continues to disassemble towards the inner layers until the core storage area of the metadata is reached. All data content stored in the core storage area is extracted, and all extracted data content is determined as the embedded metadata of the travel photos.
[0024] For the embedded metadata of the extracted tourism photos, based on the preset spatiotemporal information storage rules, the dedicated data segment in the metadata that is specifically used to store spatiotemporal information is identified, the encoding rules and character arrangement logic of the dedicated data segment are clarified, and the character sequence in the data segment is reverse parsed and transformed according to the encoding rules. The content that can represent geographic coordinates after parsing and transformation is converted into specific shooting geographic location information, and the content that can represent time records after parsing and transformation is converted into specific shooting timestamp information.
[0025] S2. Perform image content analysis on the travel photos, and integrate the analysis results, shooting location information, and shooting timestamp information in multiple dimensions to obtain multi-dimensional semantic information of the travel photos; In this embodiment of the invention, the step of performing image content analysis on the tourist photos and fusing the analysis results, shooting location information, and shooting timestamp information in multiple dimensions to obtain multi-dimensional semantic information of the tourist photos includes: Multi-granularity scene target analysis is performed on tourism photos to obtain the image semantic analysis results of tourism photos; Geographic semantic tags for travel photos are obtained by performing point-of-interest semantic mapping on the geographical location information of the photos taken. By performing temporal context semantic inference on the shooting timestamp information, temporal activity semantic tags for travel photos are obtained. By performing cross-modal semantic fusion of image semantic parsing results, geographic semantic tags, and temporal activity semantic tags, multi-dimensional semantic information of tourism photos can be obtained.
[0026] Scene target analysis is performed on identified user travel photos at three different granularities: pixel-level, object-level, and scene-level. At the pixel-level, basic visual features such as color, texture, and brightness of the image are identified pixel by pixel. At the object-level, the outline range of various entity targets in the image is defined and the target type is distinguished based on the basic visual features. At the scene-level, the type and distribution characteristics of all entity targets are integrated to determine the overall scene category presented by the image. The analysis results of the three granularities are integrated to form complete content, and the integrated complete content is determined as the image semantic analysis result of the travel photo.
[0027] The system retrieves a pre-defined semantic mapping library of points of interest (POIs), which contains POI type names and feature descriptions corresponding to various geographical locations around the world. It then performs precise matching between the geographical location information of the acquired tourist photos and the geographical locations in the POI semantic mapping library, extracts the POI type names and feature descriptions corresponding to the successfully matched geographical locations, and determines the extracted content as the geographic semantic tags for the tourist photos.
[0028] By combining the basic time attributes such as season, time period and holidays corresponding to the obtained timestamp information of the travel photos, the preset time activity semantic association rules are retrieved. These time activity semantic association rules contain common human activity types corresponding to different time attributes. The basic time attributes of the timestamp information are matched with the time activity semantic association rules, and the common human activity types corresponding to the successfully matched time attributes are extracted. The extracted content is determined as the time activity semantic tags of the travel photos.
[0029] A correlation verification benchmark based on image semantic analysis results is established. By querying a pre-set tourism interest knowledge graph or semantic association rule base, the features of interest points corresponding to geographic semantic tags are semantically matched with scene categories and entity targets in the image analysis results. Simultaneously, the types of human activities corresponding to time activity semantic tags are logically consistent with the image analysis results. Successfully matched related information is retained and integrated. Information that is semantically inconsistent, has no related path in the knowledge graph, or has a confidence level below a pre-set threshold after vector similarity calculation is judged as irrelevant and redundant information is removed. Finally, all verified information with strong semantic associations is integrated into a unified, conflict-free multi-dimensional semantic set, which serves as a high-quality input for spatiotemporal semantic association mapping in subsequent steps.
[0030] S3. Based on multi-dimensional semantic information, perform spatiotemporal semantic association mapping on travel photos to obtain the user's spatiotemporal semantic layer; The method of performing spatiotemporal semantic association mapping on travel photos based on multi-dimensional semantic information to obtain the user's spatiotemporal semantic layer includes: Based on multi-dimensional semantic information, spatiotemporal context anchoring is performed on tourism photos to obtain preliminary spatiotemporal correlation data of users; Semantic conflict resolution is performed on the initial spatiotemporal correlation data, and topological reconstruction is performed on the spatiotemporal correlation data after conflict resolution to obtain the user's spatiotemporal semantic relationship network; Based on the spatiotemporal semantic relationship network, multi-dimensional semantic features are aggregated in a structured layered manner to obtain the user's intermediate semantic layer; Semantic integrity is verified on the intermediate semantic layer to obtain the user's spatiotemporal semantic layer.
[0031] Based on the multi-dimensional semantic information of tourism photos, the core elements corresponding to the geographic semantic tags and time activity semantic tags are extracted. The core elements of the geographic semantic tags are defined as the type of point of interest contained in the tag and its precise geographic coordinates; the core elements of the time activity semantic tags are defined as the specific activity type inferred by the tag and its determined time point or time period. The location of the point of interest associated with the geographic semantic tags is used as the spatial anchor point, and the time attribute associated with the time activity semantic tags is used as the time anchor point. The image semantic analysis results of the tourism photos are attached to the spatiotemporal anchor points to form a corresponding relationship. The spatiotemporal anchor points and corresponding attached content of all tourism photos are integrated, and the integrated content is determined as the user's initial spatiotemporal association data.
[0032] The system analyzes the spatiotemporal anchor points and image semantic analysis results of various tourist photos in the preliminary spatiotemporal correlation data. It compares the spatiotemporal anchor point correspondences and image semantic content between different tourist photos, identifying and marking logical contradictions. These "logical contradictions" are primarily based on three categories of rules: first, spatiotemporal-physical contradictions, where the time and geographical location indicated by the spatiotemporal anchor points of different photos lack reasonable physical accessibility or coexistence; second, semantic-environmental contradictions, where the image semantic analysis results significantly contradict common environmental knowledge such as season, weather, and location type implied by the spatiotemporal anchor points; and third, multi-source information conflicts, where semantic information from different dimensions of the same photo contradicts each other. The system automatically compares these contradictions using a pre-set rule base and knowledge graph. When any of the above contradictions is detected, it is marked as a logical contradiction to be resolved. Based on the original content in the multi-dimensional semantic information, the contradictory content of the marked data is corrected and information that does not conform to the original semantics is removed to complete the semantic conflict resolution. The spatiotemporal correlation data after conflict resolution is rearranged according to the chronological order and spatial proximity relationship to construct the spatiotemporal correlation links between each tourist photo. The content after rearrangement and construction of correlation links is determined as the user's spatiotemporal semantic relationship network.
[0033] Using a spatiotemporal semantic relationship network as a carrier, multi-dimensional semantic features corresponding to each tourist photo in the network are extracted. According to the type of semantic features, they are divided into scene features, object features, time features, and space features. For each type of feature, a layer-by-layer inductive integration is carried out. First, the similar features of a single tourist photo are integrated to form a single photo feature subset. Then, the similar feature subsets of all photos are integrated to form a category feature set. The feature hierarchy structure formed after multi-layer inductive integration is determined as the user's intermediate semantic layer.
[0034] Based on a pre-defined feature integrity rule base, the system compares each aggregated scene-type and object-type feature set in the intermediate semantic layer with the original image semantic analysis results of each tourist photo. If a significant scene or object identified in a photo does not appear in the corresponding feature set, it is determined to be a missing feature and is supplemented. Simultaneously, the system verifies the consistency between the time-type and spatial feature sets and the spatiotemporal anchor information of each photo, ensuring that each spatiotemporal record has a corresponding feature description. Furthermore, the system identifies and merges duplicate features with similarity exceeding a set threshold by calculating the cosine similarity between feature vectors; invalid features without corresponding nodes in the knowledge graph or with excessively low confidence are removed. This ultimately forms a comprehensive, clearly structured, and non-redundant hierarchical semantic structure—the user's spatiotemporal semantic layer.
[0035] S4. Map the spatiotemporal semantic layer to the preset tourism interest knowledge graph to obtain the core interest concepts of the spatiotemporal semantic layer, and evaluate the dynamic interest intensity value of the core interest concepts based on the visual features in the spatiotemporal semantic layer. In this embodiment of the invention, mapping the spatiotemporal semantic layer to a preset tourism interest knowledge graph to obtain the core interest concepts of the spatiotemporal semantic layer includes: Based on the vectorized representation of concept nodes in the tourism interest knowledge graph, cross-space semantic alignment is performed on the spatiotemporal semantic layer to obtain the semantic nodes of the spatiotemporal semantic layer. Multi-dimensional matching of semantic nodes with interest nodes in the tourism interest knowledge graph is performed to obtain the initial matching nodes of the spatiotemporal semantic layer. Based on the semantic association strength between nodes in the tourism interest knowledge graph, the initial matching nodes are ranked by importance to obtain candidate core interest nodes in the spatiotemporal semantic layer. Based on the topological structure of the tourism interest knowledge graph, semantic concept generalization is performed on candidate core interest nodes to obtain core interest concepts in the spatiotemporal semantic layer.
[0036] The evaluation of the dynamic interest intensity value of core interest concepts based on visual features in the spatiotemporal semantic layer includes: Extract visual features associated with core interest concepts from the spatiotemporal semantic layer; Based on a pre-defined visual cognition theory, a spatiotemporal contextual saliency analysis is performed on visual features, and weights are assigned to core interest concepts based on the saliency analysis results. By analyzing the distribution patterns of the frequency of visual features over time, temporal distribution data of visual features are obtained. Based on the weight assignment results, time distribution data, and spatial distribution data of visual features in the spatiotemporal semantic layer, the dynamic interest intensity value of the core interest concept is calculated. The formula for calculating the dynamic interest intensity value is as follows: ; In the formula, This represents the dynamic interest intensity value. The total number of visual features associated with the core interest concept. For the first Weights of visual features For the first Normalized frequency of occurrence of visual features in temporal distribution data The preset time decay coefficient, For the first The time interval between the most recent occurrence of a visual feature and the current evaluation baseline time point It is a logarithmic function. This represents the spatial distribution dispersion of visual features at their corresponding geographical locations within the spatiotemporal semantic layer.
[0037] The method, based on a pre-defined visual cognition theory, performs spatiotemporal contextual saliency analysis on visual features, and assigns weights to core interest concepts based on the saliency analysis results, including: The frequency of visual features in travel photos is statistically analyzed to obtain data on the occurrence frequency of visual features. Elements adjacent to visual features in the spatiotemporal semantic layer are considered as surrounding visual elements. Analyze the degree of difference between visual features and surrounding visual elements in terms of color, texture and outline to obtain visual contrast data of visual features; The degree of fit between visual features and core interest concepts is determined to obtain scene relevance data for visual features; Based on occurrence frequency data, visual contrast data, and scene relevance data, the saliency of visual features is scored to obtain the saliency score results of visual features. Based on the saliency score, the visual features are classified into weight levels, and the core interest concepts associated with the visual features are assigned corresponding weight values.
[0038] Among the parameters required for calculating the dynamic interest intensity value, the dynamic interest intensity value is the final calculated result and has no specific dimension. The total number of visual features associated with the core interest concept comes from the statistical count of visual features associated with the core interest concept in the spatiotemporal semantic layer, obtained by identifying and counting all visual features that meet the conditions one by one.
[0039] No. The weight of each visual feature is derived from the spatiotemporal context saliency analysis results based on the pre-set visual cognition theory. This is achieved by statistically analyzing the frequency of visual features in all travel photos, analyzing the degree of difference between visual features and surrounding visual elements, and judging the fit between visual features and the scenes corresponding to core interest concepts. The visual features are then scored and classified into levels and assigned corresponding values.
[0040] No. The normalized frequency of occurrence of a visual feature in the time distribution data is derived from the time distribution data of visual features. It is obtained by statistically analyzing the number of times a visual feature appears in different time intervals and dividing that number by the total number of times all visual features appear in the corresponding time intervals. The dimensionless value is calculated. The preset time decay coefficient is derived from the preset time decay rule and is directly set based on the general law that user interest decays over time.
[0041] No. The time interval between the most recent occurrence of a visual feature and the current evaluation baseline time point is derived from the temporal distribution data of the visual features. It is obtained by extracting the timestamp of the most recent occurrence of the visual feature and calculating the time difference between the two timestamps, with the dimension being time. The logarithmic function is a general mathematical function used only to transform subsequent calculation results and has no specific dimension. The spatial distribution dispersion of the visual features at their corresponding geographical locations in the spatiotemporal semantic layer is derived from the spatial distribution data of the visual features in the spatiotemporal semantic layer. It is obtained by statistically analyzing the distribution range and density of the geographical areas corresponding to the visual features, calculating the distance deviation between each geographical location and the distribution center, and then averaging the results.
[0042] The significance of this formula lies in its ability to quantitatively assess the dynamic intensity of a user's interest in core concepts by integrating three dimensions: the weight of visual features, the time decay effect, and the spatial distribution dispersion. The numerator, by introducing the time decay effect, corrects the product of the weight of visual features and the normalized frequency of occurrence, thus weakening the influence of early visual features on current interest. The denominator normalizes the numerator by summing the weights of all visual features. Finally, the influence of spatial distribution dispersion is introduced through a logarithmic function to reflect the moderating effect of the spatial diffusion of interest on interest intensity. Overall, this formula achieves a dynamic, multi-dimensional, and accurate characterization of user interest intensity.
[0043] The trend of this formula is that when the first... When the weight of a visual feature increases, the frequency of normalization increases, and the time interval between the most recent occurrence and the current evaluation baseline decreases, the numerator value increases, and the dynamic interest intensity value increases accordingly. When the preset time decay coefficient increases, the time decay effect intensifies, the contribution of early visual features decreases rapidly, and the dynamic interest intensity value decreases even faster over time. When the spatial distribution dispersion of visual features at their corresponding geographical locations in the spatiotemporal semantic layer increases, the value of the logarithmic function increases, and the dynamic interest intensity value increases accordingly. When the total number of visual features associated with the core interest concept increases, if the weight and normalization frequency of the newly added visual features are high, the dynamic interest intensity value will increase; if the weight and normalization frequency of the newly added visual features are low, the dynamic interest intensity value may decrease or remain stable.
[0044] The vectorized representations of all concept nodes in the tourism interest knowledge graph are extracted. These vectorized representations are standardized expressions of the semantic connotations of each concept node in the knowledge graph. Using these vectorized representations as the semantic alignment benchmark, various semantic information such as scene-type features, object-type features, time-type features, and space-type features in the spatiotemporal semantic layer are projected onto the semantic space where the vectorized representations are located according to their corresponding semantic dimensions. During the projection process, the correspondence between various semantic information and the semantic dimensions of the vectorized representations is matched one by one, so that the semantic information of the spatiotemporal semantic layer and the vectorized representations of the concept nodes in the knowledge graph form a precise correspondence mapping relationship. The independent semantic information units after the mapping are determined as semantic nodes of the spatiotemporal semantic layer.
[0045] The system retrieves all interest nodes from the tourism interest knowledge graph. These nodes encompass various tourism-related concepts, such as tourist attractions, tourism activities, and tourism cuisine. It compares the semantic nodes and interest nodes in the spatiotemporal semantic layer one by one from three dimensions: semantic connotation, semantic attributes, and application scenarios. Semantic connotation checks if the core content described by both is consistent; semantic attributes check if their feature definitions are the same; and application scenarios check if they match the applicable tourism scenarios. For example, taking the semantic node description "interior photos of buildings containing Gothic spires, stained glass windows, and long rows of benches" as an example, the system first calculates the similarity between its text description vector and the names and definition vectors of each interest node in the knowledge graph from the semantic connotation dimension. The high similarity result of the "Church" node indicates that the core content of the two nodes is consistent. Secondly, in the semantic attribute dimension, the attribute set parsed from the node, such as building type, style, and elements, is compared item by item with the predefined attribute fields of the "Church" node, such as category, common style, and typical features. The feature definition of the two nodes is determined to be the same based on the number of overlaps in key attributes. Finally, in the application scenario dimension, the context labels associated with the node are matched and verified with the typical tourism scenario labels pre-associated with the "Church" node to confirm that its applicable scenario is consistent. The overlap between the semantic node and the interest node in the three dimensions is confirmed, and interest nodes that meet the preset matching requirements are selected. The selected interest nodes are determined as the initial matching nodes of the spatiotemporal semantic layer.
[0046] This paper extracts semantic association strength data between initial matching nodes and surrounding nodes in a tourism interest knowledge graph. First, the quantification of connection frequency is based on the analysis of massive amounts of tourism-related text data or anonymized user behavior logs. The frequency of any two interest nodes appearing together within the same context window is statistically analyzed, and this frequency is normalized and smoothed to convert it into standard frequency weights. Second, the quantification of semantic relevance is based on a pre-trained tourism-related word vector model. The name and attribute description text of each interest node are mapped to high-dimensional semantic vectors. The cosine similarity between the vectors is calculated to measure their semantic similarity, resulting in semantic similarity weights. Finally, the semantic association strength data is calculated by fusing the connection frequency weights and semantic similarity weights using a preset weighting ratio. This reflects the semantic influence of the initial matching nodes in the knowledge graph. Using semantic association strength data as the sole ranking criterion, all initial matching nodes are arranged in descending order. The initial matching nodes that rank highly and whose association strength data falls within a preset range are selected as candidate core interest nodes for the spatiotemporal semantic layer.
[0047] Referring to the topological structure of the tourism interest knowledge graph, which clearly presents the hierarchical relationships and parallel associations between nodes, the upper-level semantic concepts of candidate core interest nodes are traced. These upper-level semantic concepts are general expressions of the specific content of candidate core interest nodes. The specific semantic content of candidate core interest nodes is summarized into the corresponding upper-level semantic concept category, expanding the semantic coverage of candidate core interest nodes and forming a more general and universal semantic expression. This general semantic expression is determined as the core interest concept of the spatiotemporal semantic layer.
[0048] The system traverses all visual information units within the spatiotemporal semantic layer. These visual information units contain visual elements such as color, shape, texture, and outline corresponding to travel photos. Based on the semantic features of the core interest concept, it identifies and filters visual information that is directly related to the semantic features of the core interest concept. For example, when the core interest concept is natural landscape, it filters out visual information corresponding to the outline of the landscape, vegetation, and color. It extracts the color parameters, shape, outline, texture, details, and other features corresponding to this visual information and determines the extracted features as visual features related to the core interest concept in the spatiotemporal semantic layer.
[0049] Following a pre-defined visual cognition theory, this study clarifies the criteria for determining the salience of visual features in a spatiotemporal context. These criteria include three core indicators: frequency of occurrence, visual contrast, and scene relevance. The study counts the number of times a visual feature appears in all travel photos based on its frequency of occurrence; analyzes the degree of difference between the feature and surrounding visual elements based on its visual contrast; and assesses the degree of fit between the feature and the scene corresponding to the core interest concept based on its scene relevance. Based on the analysis results of these three indicators, the extracted visual features are scored for salience. The salience levels of the visual features are then classified according to their scores, and corresponding weight values are assigned to the core interest concepts according to their salience levels, thus completing the weight assignment for the core interest concepts.
[0050] The frequency of occurrence of extracted visual features in different time intervals is statistically analyzed. These time intervals are divided into different seasons, months, and dates based on timestamp information. The frequency data of visual features in each time interval is recorded. The frequency data of all time intervals are integrated to analyze the occurrence patterns and trends of visual features in the time dimension. For example, if a certain type of visual feature appears more frequently in summer than in winter, the integrated frequency data and patterns are used to determine the temporal distribution data of the visual features.
[0051] The weighting results of core interest concepts are summarized, along with the temporal distribution data of visual features and the spatial distribution data of visual features in the spatiotemporal semantic layer. This spatial distribution data reflects the distribution range and density of the corresponding geographical areas of the visual features. The weighting results are used as the core calculation coefficients. Combined with the temporal influence factors reflected in the temporal distribution data and the spatial influence factors reflected in the spatial distribution data, the relevant data of the core interest concepts are comprehensively calculated. During the calculation process, the weighting results are used as the benchmark, and the temporal dimension weight reflected by the temporal influence factor and the spatial dimension weight reflected by the spatial influence factor are superimposed. The resulting value is determined as the dynamic interest intensity value of the core interest concept.
[0052] Visual feature extraction was performed on all collected travel photos to identify the specific types and actual manifestations of various visual features in the photos. Then, each travel photo was precisely labeled, and the occurrence of each visual feature in each photo was fully recorded. After the labeling of all travel photos was completed, the labeling results of each type of visual feature were comprehensively summarized, and the total number of times each type of visual feature appeared in all travel photos was accumulated to finally form the frequency data of visual features.
[0053] A spatiotemporal semantic layer corresponding to the travel photos is constructed. In this spatiotemporal semantic layer, the spatial location and semantic positioning of each visual feature are accurately identified. Taking the spatial location and semantic positioning of the visual features as the core, a fixed adjacent spatial range and semantic association range are defined. Within the defined range, all elements in the spatiotemporal semantic layer that are spatially and semantically adjacent to the visual feature are selected. These selected elements are directly identified as the surrounding visual elements of the visual feature.
[0054] The color representation of the visual feature itself is extracted, clarifying its hue, brightness, and saturation. Simultaneously, the same color representation extraction is performed on each surrounding visual element, comparing the specific differences in hue, brightness, and saturation between the visual feature and each surrounding visual element. Next, the texture representation of the visual feature itself is extracted, clarifying its density, direction, and shape. The same texture representation extraction is performed on each surrounding visual element, comparing the specific differences in density, direction, and shape between the visual feature and each surrounding visual element. Then, the outline representation of the visual feature itself is extracted, clarifying its shape, line thickness, and degree of closure. The same outline representation extraction is performed on each surrounding visual element, comparing the specific differences in shape, line thickness, and degree of closure between the visual feature and each surrounding visual element. Finally, the comparison results of color, texture, and outline are synthesized to quantify the overall differences between the visual feature and surrounding visual elements. The results directly form the visual contrast data of the visual feature.
[0055] The specific connotations and expressive characteristics of the core interest concept corresponding to the tourism photo are clearly defined. These expressive characteristics include the visual presentation requirements, scene expression requirements, and semantic matching requirements corresponding to the core interest concept. The visual expression, scene adaptability, and semantic connotation of the visual features are compared with each requirement of the core interest concept in detail. During the comparison process, the actual matching status of the visual features in each requirement is clarified. Then, based on the overall matching status, the fit between the visual features and the core interest concept is judged as a whole. The result of the judgment is directly used as the scene relevance data of the visual features.
[0056] The frequency of visual features, visual contrast data, and scene relevance data are analyzed using a unified evaluation dimension to ensure that the three types of data can form a synergistic evaluation basis for the salience of visual features. Based on the analyzed three types of data, the prominence of visual features is comprehensively evaluated from three perspectives: frequency of occurrence, visual difference, and scene fit. Combining the actual data performance at each level, the overall salience of the visual features is scored and ranked, and the determined scores and ranks directly form the salience score of the visual features.
[0057] The saliency scores of all visual features are ranked and divided into intervals. Different weight levels are assigned according to the scores, and each weight level corresponds to a fixed score interval. Each visual feature is assigned to its corresponding weight level based on its saliency score, thus completing the weight level division of visual features. Then, based on the weight level to which each visual feature belongs, the corresponding weight value is determined, and this weight value is directly assigned to the core interest concept that is related to the visual feature.
[0058] S5. Based on the hierarchical relationship in the tourism interest knowledge graph, multi-level collaborative aggregation of core interest concepts and dynamic interest intensity values is carried out to obtain user tourism interest profile data. In this embodiment of the invention, the step of performing multi-level collaborative aggregation of core interest concepts and dynamic interest intensity values based on the hierarchical relationship in the tourism interest knowledge graph to obtain user tourism interest profile data includes: Based on the parent-child and sibling relationships between concept nodes in the tourism interest knowledge graph, the hierarchical path tracing of core interest concepts is performed to obtain the hierarchical distribution structure of core interest concepts. Based on the hierarchical distribution structure, a multi-scale field is constructed for the dynamic interest intensity values to obtain the user's interest intensity distribution map. By performing cross-level semantic convergence on the interest intensity distribution map, we can obtain the user's hierarchical tourism interest themes. Structured profile rendering is performed on the hierarchical tourism interest themes to obtain user tourism interest profile data.
[0059] The method, based on a hierarchical distribution structure, constructs a multi-scale field from dynamic interest intensity values to obtain a user's interest intensity distribution map, including: Based on the node topology in the hierarchical distribution structure, the hierarchical intensity field is initialized for the dynamic interest intensity value to obtain the initial hierarchical intensity field of the dynamic interest intensity value. Based on the coupling relationship between adjacent nodes in the hierarchical distribution structure, the initial hierarchical intensity field is propagated by field effect to obtain the equalized hierarchical intensity field with dynamic interest intensity values. Cross-level multi-scale fusion of the equalized hierarchical intensity field yields the user's interest intensity distribution map.
[0060] Retrieve the nodes corresponding to the core interest concepts in the tourism interest knowledge graph, trace the parent nodes of each level upwards from the node until reaching the root concept node at the top level of the graph. At the same time, horizontally sort out all the sibling nodes of the node at each level, record the parent-child inheritance relationship and sibling parallel relationship between each node, clarify the position and associated objects of each node in the level, integrate all traced nodes and their corresponding relationships to form a complete hierarchical network including the top root node, intermediate level nodes, bottom core interest concept nodes and sibling nodes, and determine the hierarchical distribution structure of the core interest concepts.
[0061] Using the hierarchical distribution structure of core interest concepts as a carrier, dynamic interest intensity values are directly assigned to the core interest concept nodes in the hierarchical distribution structure. The parent node of this node is assigned an initial transmission value based on the intensity of the core interest concept node, with the transmission ratio set according to the semantic coverage of the parent node in the knowledge graph. The child nodes of this node are assigned an initial transmission value based on the intensity of the core interest concept node, with the transmission ratio set according to the semantic subdivision of the child nodes in the knowledge graph. The sibling nodes of this node are assigned an initial association value associated with the intensity of the core interest concept node, with the association ratio set according to the semantic similarity between the sibling nodes and the core interest concept node. The hierarchical structure containing the initial intensity values of each node is determined as the initial hierarchical intensity field of dynamic interest intensity values.
[0062] Based on the coupling relationship between adjacent nodes in the hierarchical distribution structure, this coupling relationship includes the inheritance coupling of parent and child nodes and the association coupling of sibling nodes. Starting from the core interest concept node of the initial hierarchical intensity field, the intensity value of the node is passed to the parent and child nodes according to the inheritance ratio of the parent and child nodes. The inheritance ratio is determined by the semantic association depth between nodes. At the same time, it is diffused to the sibling nodes according to the association ratio of the sibling nodes. The association ratio is determined by the semantic similarity between nodes. The intensity value of each node is updated in real time during the transmission and diffusion process. When the intensity value of each node no longer changes after three consecutive transmissions, the transmission stops. The hierarchical structure in which the intensity values of each node tend to be stable after transmission and diffusion is determined as the balanced hierarchical intensity field of dynamic interest intensity values.
[0063] Based on the equalized hierarchical intensity field, the intensity values of nodes at different levels are extracted. From the top-level root concept node to the bottom-level sub-nodes, the intensity values of different levels within the same semantic category are weighted and integrated. The weighting is set according to the semantic generality of the level, with higher generality resulting in greater weight. The intensity values across levels are fused at multiple scales. During fusion, the intensity characteristics of each level are preserved while eliminating intensity gaps between levels, forming a continuous intensity distribution covering all levels. This distribution clearly presents the intensity differences of user interests at different semantic granularities. This continuous intensity distribution is determined as the user's interest intensity distribution map.
[0064] Starting from the top-level root concept node of the interest intensity distribution map, the intensity values of nodes at each level are converged downwards. In each level, the node with the highest intensity value is selected as the core semantic node of that level. The core semantic node represents the travel interest direction that users are most concerned about at that level. The core semantic nodes of each level are integrated to form a hierarchical theme set, which includes a top-level general theme, intermediate sub-themes, and bottom-level specific themes. This theme set is determined as the user's hierarchical travel interest theme.
[0065] Using a hierarchical tourism interest theme framework, the semantic, visual, and spatiotemporal features corresponding to each theme are structurally integrated. Semantic features include the core connotation and related concepts of the theme, visual features include the visual elements of the tourism photos corresponding to the theme, and spatiotemporal features include the shooting time and geographical location corresponding to the theme. The core content of each theme is rendered in descending order of theme level, clarifying the relationship and weight ratio between each theme. The weight ratio is set according to the intensity value of the theme. The integrated and rendered structured content, including hierarchical themes, feature details, relationships, and weight ratios, is determined as the user's tourism interest profile data.
[0066] S6. Based on user travel interest profile data and combined with the shooting location information, generate user-visualized travel interest recommendation information.
[0067] In this embodiment of the invention, generating visualized travel interest recommendation information for users based on user travel interest profile data and combined with shooting location information includes: Readable recommendation content descriptions are generated from user travel interest profile data to obtain structured descriptive text for users; Based on the coordinates in the geographic location information, the structured descriptive text is annotated with points of interest to obtain the user's annotated text. Visual interactive logic is bound to the labeled text to obtain the user's visual style description text; By integrating multi-source information such as structured descriptive text, labeled content text, and visual style descriptive text, visualized travel interest recommendation information for users can be obtained.
[0068] Extract all core content from the user's travel interest profile data, including hierarchical travel interest themes, the weight of each theme, visual features, and spatiotemporal features. Following the order from the top-level general theme to the bottom-level specific theme, transform the core features of each theme into easy-to-understand natural language expressions. In the process of expression, clarify the relationship and weight differences between the themes, and integrate scene descriptions and activity features corresponding to travel photos to form logically coherent and well-structured text content. This text content is then identified as the user's structured descriptive text.
[0069] The system retrieves the coordinates of the location information of the tourist photos, matches them with a preset point of interest (POI) database, obtains the specific POI name and feature description corresponding to the coordinates, finds the content paragraphs related to the POI in the structured description text, and adds the POI name and feature description to the specified position of the corresponding paragraph in an embedded form. This completes the POI annotation operation on the structured description text, and the annotated text is then identified as the user's annotated content text.
[0070] To meet the presentation requirements of the annotated text, corresponding visual interaction logic was designed. This interaction logic includes hierarchical expansion logic of text content, pop-up display logic for point-of-interest (POI) annotations, and switching logic for different themes. The hierarchical structure of the text is bound to the expansion logic to achieve the effect of displaying the specific content of the lower level when the top-level theme is clicked. The POI annotations are bound to the pop-up display logic to achieve the effect of displaying detailed information about the POI when the annotation location is clicked. Different themes are bound to the switching logic to achieve seamless switching between themes. The text content after binding all the interaction logic is determined as the user's visual style description text.
[0071] This process aggregates users' structured descriptive text annotations and visual style descriptions. Using the structured descriptive text as the basic content framework, it integrates the point-of-interest (POI) annotations from the POIs and embeds the interactive logic settings from the visual style descriptions. This ensures that the integrated content retains the logical hierarchy and core information of the structured descriptive text, includes detailed content from the POIs, and possesses visual interactive features. The integrated content, including the content framework annotations and interactive logic, is then defined as the user's visual travel interest recommendation information.
[0072] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0073] This application embodiment can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for generating visual tourism interest recommendation information based on tourism photos, characterized in that, The method includes: S1. Obtain the user's travel photos and the corresponding shooting location information and shooting timestamp information; S2. Perform image content analysis on the travel photos, and integrate the analysis results, shooting location information, and shooting timestamp information in multiple dimensions to obtain multi-dimensional semantic information of the travel photos; S3. Based on multi-dimensional semantic information, perform spatiotemporal semantic association mapping on travel photos to obtain the user's spatiotemporal semantic layer; S4. Map the spatiotemporal semantic layer to the preset tourism interest knowledge graph to obtain the core interest concepts of the spatiotemporal semantic layer, and evaluate the dynamic interest intensity value of the core interest concepts based on the visual features in the spatiotemporal semantic layer. S5. Based on the hierarchical relationship in the tourism interest knowledge graph, multi-level collaborative aggregation of core interest concepts and dynamic interest intensity values is carried out to obtain user tourism interest profile data. S6. Based on user travel interest profile data and combined with the shooting location information, generate user-visualized travel interest recommendation information.
2. The method for generating visual tourism interest recommendation information based on tourism photos as described in claim 1, characterized in that, The process of obtaining the user's travel photos and corresponding shooting location information and shooting timestamp information includes: Based on the user's digital identity, the user's travel photos are obtained by retrieving multi-source heterogeneous image storage nodes in the distributed storage space. Metadata encapsulation structure of tourist photos is parsed to obtain embedded metadata of tourist photos; Spatiotemporal information decoding of embedded metadata yields the geographic location and timestamp information of the tourist photos.
3. The method for generating visual tourism interest recommendation information based on tourism photos as described in claim 1, characterized in that, The process involves analyzing the image content of tourist photos and then fusing the analysis results, shooting location information, and shooting timestamp information in multiple dimensions to obtain multi-dimensional semantic information of the tourist photos, including: Multi-granularity scene target analysis is performed on tourism photos to obtain the image semantic analysis results of tourism photos; Geographic semantic tags for travel photos are obtained by performing point-of-interest semantic mapping on the geographical location information of the photos taken. By performing temporal context semantic inference on the shooting timestamp information, temporal activity semantic tags for travel photos are obtained. By performing cross-modal semantic fusion of image semantic parsing results, geographic semantic tags, and temporal activity semantic tags, multi-dimensional semantic information of tourism photos can be obtained.
4. The method for generating visualized tourism interest recommendation information based on tourism photos as described in claim 1, characterized in that, The method of performing spatiotemporal semantic association mapping on travel photos based on multi-dimensional semantic information to obtain the user's spatiotemporal semantic layer includes: Based on multi-dimensional semantic information, spatiotemporal context anchoring is performed on tourism photos to obtain preliminary spatiotemporal correlation data of users; Semantic conflict resolution is performed on the initial spatiotemporal correlation data, and topological reconstruction is performed on the spatiotemporal correlation data after conflict resolution to obtain the user's spatiotemporal semantic relationship network; Based on the spatiotemporal semantic relationship network, multi-dimensional semantic features are aggregated in a structured layered manner to obtain the user's intermediate semantic layer; Semantic integrity is verified on the intermediate semantic layer to obtain the user's spatiotemporal semantic layer.
5. The method for generating visualized tourism interest recommendation information based on tourism photos as described in claim 1, characterized in that, The process of mapping the spatiotemporal semantic layer to a preset tourism interest knowledge graph yields the core interest concepts of the spatiotemporal semantic layer, including: Based on the vectorized representation of concept nodes in the tourism interest knowledge graph, cross-space semantic alignment is performed on the spatiotemporal semantic layer to obtain the semantic nodes of the spatiotemporal semantic layer. Multi-dimensional matching of semantic nodes with interest nodes in the tourism interest knowledge graph is performed to obtain the initial matching nodes of the spatiotemporal semantic layer. Based on the semantic association strength between nodes in the tourism interest knowledge graph, the initial matching nodes are ranked by importance to obtain candidate core interest nodes in the spatiotemporal semantic layer. Based on the topological structure of the tourism interest knowledge graph, semantic concept generalization is performed on candidate core interest nodes to obtain core interest concepts in the spatiotemporal semantic layer.
6. The method for generating visualized tourism interest recommendation information based on tourism photos as described in claim 1, characterized in that, The evaluation of the dynamic interest intensity value of core interest concepts based on visual features in the spatiotemporal semantic layer includes: Extract visual features associated with core interest concepts from the spatiotemporal semantic layer; Based on a pre-defined visual cognition theory, a spatiotemporal contextual saliency analysis is performed on visual features, and weights are assigned to core interest concepts based on the saliency analysis results. By analyzing the distribution patterns of the frequency of visual features over time, temporal distribution data of visual features are obtained. Based on the weight assignment results, time distribution data, and spatial distribution data of visual features in the spatiotemporal semantic layer, the dynamic interest intensity value of the core interest concept is calculated.
7. The method for generating visualized tourism interest recommendation information based on tourism photos as described in claim 6, characterized in that, The method, based on a pre-defined visual cognition theory, performs spatiotemporal contextual saliency analysis on visual features, and assigns weights to core interest concepts based on the saliency analysis results, including: The frequency of visual features in travel photos is statistically analyzed to obtain data on the occurrence frequency of visual features. Elements adjacent to visual features in the spatiotemporal semantic layer are considered as surrounding visual elements. Analyze the degree of difference between visual features and surrounding visual elements in terms of color, texture and outline to obtain visual contrast data of visual features; The degree of fit between visual features and core interest concepts is determined to obtain scene relevance data for visual features; Based on occurrence frequency data, visual contrast data, and scene relevance data, the saliency of visual features is scored to obtain the saliency score results of visual features. Based on the saliency score, the visual features are classified into weight levels, and the core interest concepts associated with the visual features are assigned corresponding weight values.
8. The method for generating visualized tourism interest recommendation information based on tourism photos as described in claim 1, characterized in that, Based on the hierarchical relationships in the tourism interest knowledge graph, multi-level collaborative aggregation of core interest concepts and dynamic interest intensity values is performed to obtain user tourism interest profile data, including: Based on the parent-child and sibling relationships between concept nodes in the tourism interest knowledge graph, the hierarchical path tracing of core interest concepts is performed to obtain the hierarchical distribution structure of core interest concepts. Based on the hierarchical distribution structure, a multi-scale field is constructed for the dynamic interest intensity values to obtain the user's interest intensity distribution map. By performing cross-level semantic convergence on the interest intensity distribution map, we can obtain the user's hierarchical tourism interest themes. Structured profile rendering is performed on the hierarchical tourism interest themes to obtain user tourism interest profile data.
9. The method for generating visualized tourism interest recommendation information based on tourism photos as described in claim 8, characterized in that, The method, based on a hierarchical distribution structure, constructs a multi-scale field from dynamic interest intensity values to obtain a user's interest intensity distribution map, including: Based on the node topology in the hierarchical distribution structure, the hierarchical intensity field is initialized for the dynamic interest intensity value to obtain the initial hierarchical intensity field of the dynamic interest intensity value. Based on the coupling relationship between adjacent nodes in the hierarchical distribution structure, the initial hierarchical intensity field is propagated by field effect to obtain the equalized hierarchical intensity field with dynamic interest intensity values. Cross-level multi-scale fusion of the equalized hierarchical intensity field yields the user's interest intensity distribution map.
10. The method for generating visualized tourism interest recommendation information based on tourism photos as described in claim 1, characterized in that, The process of generating visualized travel interest recommendations based on user travel interest profile data and combined with geographic location information includes: Readable recommendation content descriptions are generated from user travel interest profile data to obtain structured descriptive text for users; Based on the coordinates in the geographic location information, the structured descriptive text is annotated with points of interest to obtain the user's annotated text. Visual interactive logic is bound to the labeled text to obtain the user's visual style description text. By integrating multi-source information such as structured descriptive text, labeled content text, and visual style descriptive text, visualized travel interest recommendation information for users can be obtained.