Tibetan tourism resource intelligent pushing method and system integrating images and texts

By combining multi-scale visual analysis and cross-modal alignment algorithms with snowfield spatial clustering, the problem of insufficient matching of image and text features in the promotion of Tibetan tourism resources has been solved, realizing personalized and accurate resource promotion.

CN122152905APending Publication Date: 2026-06-05TIBET ZHIHE CULTURE TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIBET ZHIHE CULTURE TECHNOLOGY CO LTD
Filing Date
2026-03-05
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing technologies for promoting Tibetan tourism resources suffer from poor cross-modal data alignment, low matching degree between image and text features, and insufficient spatial clustering and push matching, making it difficult to achieve personalized and accurate resource push.

Method used

Image and text features are extracted using a multi-scale visual analysis model of Tibetan landscapes. Feature mapping is established by combining a cross-modal alignment algorithm for Tibetan tourism images and texts. Accurate clustering is performed using a snowfield spatial clustering recommendation algorithm. Feature database of Tibetan landscape cultural tourism intelligent analysis platform is constructed, and demand matching is achieved through an intelligent push module.

Benefits of technology

This has improved the accuracy and personalization of Tibetan tourism resource recommendations, ensuring that the data delivered aligns with user needs and achieving full-process intelligentization from data collection to delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122152905A_ABST
    Figure CN122152905A_ABST
Patent Text Reader

Abstract

The application discloses a Tibetan tourism resource intelligent pushing method and system integrating image and text, and the method comprises the following steps: collecting image and text data of Tibetan tourism resources and extracting features through a Tibetan scene tourism intelligent analysis platform; processing the data through a Tibetan scene multi-scale visual analysis model to obtain multi-dimensional features; calling a Tibetan tourism image and text cross-modal alignment algorithm to establish a correlation between image and text features and screen a matching feature pair; starting a snowfield space clustering recommendation algorithm to combine space parameters to divide resource clustering clusters; constructing a feature database on the platform and establishing an index; and pushing resources after calculating a matching degree according to user requirements. The system comprises six units, and each unit cooperatively operates. The application solves the problems of cross-modal alignment difference and the insufficient combination of clustering and pushing in the prior art through a series of special models and algorithms, realizes deep analysis and accurate pushing of Tibetan tourism resource data, and improves the pushing efficiency and the individualization level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Tibetan tourism resource promotion technology, and in particular to a method and system for intelligent promotion of Tibetan tourism resources that integrates text, images, and graphics. Background Technology

[0002] With the rapid development of the tourism industry in Tibetan areas, its rich natural landscapes and unique cultural landscapes have attracted a large number of tourists. The digitization and intelligent delivery of tourism resources has become a key direction for enhancing tourist experience and promoting the upgrading of the local cultural tourism industry. Currently, the amount of image and text data related to Tibetan tourism resources is experiencing explosive growth, including images of natural landscapes, cultural landscapes, and various types of textual tourism information. How to efficiently extract valuable features from this massive amount of data and achieve deep correlation and accurate delivery of image and text information has become a pressing issue for the industry. The development of technologies such as the Tibetan Scenic Area Cultural Tourism Intelligent Analysis Platform and multi-scale visual analysis models has provided a technical foundation for integrating image and text data and optimizing tourism resource delivery. However, existing technologies still have room for improvement in cross-modal data processing, spatial clustering analysis, and intelligent delivery matching, making it difficult to fully meet users' personalized and precise needs for accessing Tibetan tourism resources.

[0003] Existing technologies for pushing Tibetan tourism resources suffer from two main drawbacks: First, cross-modal data alignment is poor. Most technologies struggle to achieve deep correlation between image and text features. When processing text and image data of Tibetan tourism resources, the lack of a feature mapping mechanism tailored to the unique characteristics of Tibetan landscapes often results in low matching between text and image information, failing to accurately reflect the core attributes of tourism resources and thus affecting the effectiveness of subsequent push data. Second, the integration of spatial clustering and push matching is insufficient. Existing technologies often neglect the unique spatial distribution characteristics of Tibetan areas when clustering tourism resources. The clustering results cannot accurately classify different types and regions of tourism resources. Furthermore, when calculating the matching degree between user needs and resources, the feature parameters in the clustering results are not fully utilized, leading to a significant deviation between the push results and the actual needs of users, making it difficult to achieve personalized push goals. Summary of the Invention

[0004] In order to overcome the shortcomings and deficiencies of existing technologies, this invention provides a method and system for intelligently pushing Tibetan tourism resources by integrating text and images.

[0005] The technical solution adopted in this invention is an intelligent method for pushing Tibetan tourism resources by integrating images and text, including the following steps: S1, collecting image and text data related to Tibetan tourism resources through a Tibetan scenic and cultural tourism intelligent analysis platform. The image and text data includes images of Tibetan natural landscapes, images of cultural landscapes, and tourism text descriptions. During the collection process, spatial location features, visual texture features, and text semantic features of the image and text data are extracted; S2, inputting the image and text data collected in S1 into a Tibetan scenic multi-scale visual analysis model. The model performs multi-scale feature decomposition on the image data to obtain local and global features of the image at different resolutions. At the same time, semantic word segmentation is performed on the text data to extract text keyword features and semantic association features; S3, calling a Tibetan tourism image and text cross-modal alignment algorithm to perform cross-modal feature mapping between the image features and text features output in S2, establishing an association mapping relationship between image features and text features. The algorithm calculates feature similarity thresholds and filters out cross-modal feature pairs that meet the similarity requirements. S4: The snow-covered land spatial clustering recommendation algorithm is activated. The cross-modal feature pairs obtained in S3 are combined with the spatial location parameters of Tibetan tourism resources to perform spatial clustering analysis, dividing the tourism resources into different clusters. The feature center value and intra-cluster feature dispersion of each cluster are calculated. S5: Based on the clustering results of S4, a Tibetan tourism resource feature database is constructed in the Tibetan Scenic Tourism Intelligent Analysis Platform. The image and text features, spatial location parameters, and clustering feature parameters corresponding to each cluster are stored in the database, and a database index structure is established for rapid feature retrieval. S6: Based on the user's tourism demand parameters, the Tibetan tourism resource intelligent push module calls the data in the feature database, combines it with the cluster feature center values ​​from S4, calculates the matching degree between the user's demand and each cluster, and outputs the corresponding Tibetan tourism resource push information based on the matching degree ranking results.

[0006] Furthermore, the expression for the multi-scale visual analysis model of the hidden scenery is: ,in Indicated in scale Lower image coordinates Visual analysis results at the location, Indicates the first Weight parameters for each feature extraction channel, Indicates the first A nonlinear feature mapping function, Representing scale lower coordinate The original image pixel values ​​at that location, This represents the convolution operation. Representing scale Next Each convolutional kernel parameter, Represents the semantic feature fusion coefficient. Representing scale lower coordinate The semantic association feature value of the text at that location This indicates the total number of feature extraction channels.

[0007] Furthermore, the expression for the cross-modal alignment algorithm for Tibetan travel images and text is: ,in Representing an image With words Cross-modal alignment Indicates the image number 1 The first feature and text The association weights of each feature Representing image features With textual features The included angle, Indicates the number of dimensions of image features. This represents the number of dimensions of text features. Indicates the spatial alignment coefficient. This represents the spatial distance parameter between the image and the text.

[0008] Furthermore, the expression for the snowfield spatial clustering recommendation algorithm is: ,in This indicates the spatial clustering degree of tourism resource samples. This represents the total number of samples within a cluster. Indicates the spatial location weighting coefficient. Indicates the spatial distance attenuation coefficient. Indicates the first Spatial location coordinate parameters of each sample The spatial coordinates of the cluster center are represented by the following parameters. This represents the feature similarity weight coefficient. Indicates the first A set of image and text features for each sample. A set of graphic and textual features representing the center of a cluster.

[0009] Furthermore, the resource feature processing expression of the Tibetan Scenic Area Cultural Tourism Intelligent Analysis Platform is as follows: ,in This indicates the platform's priority in processing the characteristics of tourism resources. Indicates the feature processing weight coefficients. Indicates the total number of feature dimensions. Indicates the first The processing weights of each feature, The first characteristic vector of a resource is represented by the... One portion, Represents the first feature vector in the database One portion, Represents the cluster association coefficient. This parameter indicates the importance of the cluster to which the resource belongs.

[0010] Furthermore, the matching calculation expression for the intelligent push of Tibetan tourism resources is as follows: ,in Indicate user needs With clusters The matching degree of the push notifications Indicates the number of dimensions of the demand features. Indicates the first Weight parameters for each demand characteristic, Indicating the first in user requirements The values ​​of each feature, Represents clusters The Middle The mean of each feature, To represent the minimum parameter to avoid the denominator being zero, Represents clusters The Middle The standard deviation of each feature Indicates the first The matching gain coefficient of each feature.

[0011] Further, S3 includes the following sub-steps: S31, extracting the multi-scale visual feature matrix of the image and the semantic feature vector of the text from the output of the Tibetan landscape multi-scale visual analysis model, determining that the row dimension of the image feature matrix is ​​the number of image feature types and the column dimension is the number of feature dimensions, and that the dimension of the text feature vector is consistent with the column dimension of the image feature matrix; S32, inputting the image feature matrix and the text feature vector into the feature mapping module of the Tibetan tourism image-text cross-modal alignment algorithm, mapping the image features and text features to the same feature space through the feature transformation matrix within the module, generating a cross-modal feature matrix; S33, calculating the cosine similarity between each row of image features and the corresponding text features in the cross-modal feature matrix, setting an initial similarity threshold, and filtering out feature pairs with similarity greater than the initial threshold to form a preliminary aligned feature set; S34, adjusting the similarity threshold based on the spatial position correlation of features in the preliminary aligned feature set, eliminating feature pairs with mismatched spatial positions, and finally obtaining a feature pair set that meets the cross-modal alignment requirements.

[0012] Further, S4 includes the following sub-steps: S41, obtain the cross-modal aligned feature pair set output by S3, extract the spatial location coordinates of Tibetan tourism resources corresponding to each feature pair, and establish a feature-location association data table, which includes feature pair identifiers, image feature values, text feature values, and spatial location coordinate parameters; S42, input the feature-location association data table into the initialization module of the snowfield spatial clustering recommendation algorithm, set the initial value of the number of clusters and the maximum number of clustering iterations, and initialize the center position and center feature value of each cluster; S43, calculate the spatial distance and feature similarity between each feature pair and the center of each cluster through the algorithm, allocate the feature pair to the corresponding cluster according to the weighted result of distance and similarity, and update the center position and center feature value of each cluster; S44, repeat the calculation and allocation process of S43 until the number of clustering iterations reaches the maximum number or the change in the center position and center feature value of the cluster is less than the set threshold, stop the iteration and output the final clustering result.

[0013] Further, S5 includes the following sub-steps: S51, receiving the clustering results output by S4, organizing the feature pair data, spatial location parameters, and cluster center parameters corresponding to each cluster, and classifying these data according to feature type, dividing them into image feature class, text feature class, and spatial parameter class; S52, creating a Tibetan tourism resource feature database in the database module of the Tibetan Cultural Tourism Intelligent Analysis Platform, setting up an image feature table, a text feature table, a spatial parameter table, and a clustering information table for the database, and establishing associations between the tables through cluster identifiers and feature pair identifiers; S53, dividing S51 into... Each type of data is stored in its corresponding data table. The image feature table records image feature values ​​and feature dimension information, the text feature table records text feature vectors and semantic association information, the spatial parameter table records spatial location coordinates and location association information, and the cluster information table records cluster center parameters and the number of samples within each cluster. S54. A database index is constructed based on the key fields in each data table. The index includes a cluster identifier index, a feature dimension index, and a spatial location index. The index optimizes the database retrieval speed and enables fast queries on labeled clusters or labeled feature data.

[0014] An intelligent system for pushing Tibetan tourism resources integrating text and images is proposed. This system, applied to an intelligent method for pushing Tibetan tourism resources based on integrated text and images, includes: a Tibetan landscape and cultural tourism intelligent analysis data acquisition unit, which establishes a data transmission connection with the data source of Tibetan tourism resources to collect text and image data of Tibetan tourism resources, extract spatial location features, visual texture features, and textual semantic features, and transmits the collected feature data to a Tibetan landscape multi-scale visual analysis processing unit; a Tibetan landscape multi-scale visual analysis processing unit, which receives the feature data transmitted by the Tibetan landscape and cultural tourism intelligent analysis data acquisition unit, performs multi-scale decomposition and feature extraction on the feature data through a built-in Tibetan landscape multi-scale visual analysis model, generating image local features, image global features, text keyword features, and textual semantic association features, and sends the generated feature data to a Tibetan tourism text-image cross-modal alignment processing unit; and a Tibetan tourism text-image cross-modal alignment processing unit, which establishes data interaction with the Tibetan landscape multi-scale visual analysis processing unit, and calls a Tibetan tourism text-image cross-modal alignment algorithm to perform cross-modal mapping of the received image features and text features. The process involves establishing feature associations and selecting feature pairs that meet similarity requirements, then transmitting the feature pair data to the Tibetan tourism spatial clustering recommendation processing unit. This unit receives feature pair data from the Tibetan tourism image-text cross-modal alignment processing unit, performs spatial clustering analysis on the feature pairs using its built-in Tibetan tourism spatial clustering recommendation algorithm, divides tourism resource clusters, calculates clustering parameters, and transmits the clustering results data to the Tibetan tourism resource feature database storage unit. This unit, connected to the Tibetan tourism spatial clustering recommendation processing unit, receives the clustering results data, constructs a feature database, establishes data tables and index structures, stores cluster features, image-text features, and spatial location parameters, and establishes a data retrieval connection with the Tibetan tourism resource intelligent push matching unit. Finally, the Tibetan tourism resource intelligent push matching unit receives user-input tourism demand parameters, retrieves data from the Tibetan tourism resource feature database storage unit, calculates the matching degree between user demands and clusters, sorts the data based on the matching degree, generates tourism resource push information, and outputs it to the user interaction terminal.

[0015] Beneficial Effects: This invention proposes an intelligent method and system for pushing Tibetan tourism resources by integrating images and text. It uses a multi-scale visual analysis model of Tibetan landscapes to decompose and extract features from image and text data at multiple scales. Combined with a cross-modal alignment algorithm for Tibetan tourism images and text, it establishes a deep association between image and text features. A snow-covered land spatial clustering recommendation algorithm is used to achieve accurate clustering of tourism resources. A feature database is constructed based on a Tibetan landscape and cultural tourism intelligent analysis platform, and demand matching is completed through an intelligent push module. This method can efficiently mine the value of Tibetan tourism resource image and text data, improve the accuracy and personalization of resource push, and achieve intelligent operation of the entire process from data collection and feature processing to push output through the collaborative cooperation of various units of the system. This provides users with Tibetan tourism resource information that better meets their needs. This method employs a specially designed cross-modal alignment algorithm for Tibetan tourism images and text, combined with a feature mapping mechanism built based on the characteristics of Tibetan tourism resources, to improve the matching degree of image and text information, accurately reflect the core attributes of resources, and ensure the effectiveness of subsequent push data. This method utilizes a snow-covered spatial clustering recommendation algorithm to fully consider the special spatial distribution characteristics of Tibetan areas, accurately divide resource clusters, and call the clustering result feature parameters during push matching to reduce the deviation between push results and users' actual needs, effectively achieving the goal of personalized push. Attached Figure Description

[0016] Figure 1 This is a flowchart of the method steps of the present invention; Figure 2 This is a diagram showing the system unit composition of the present invention. Detailed Implementation

[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0018] like Figure 1 As shown, the intelligent push method for Tibetan tourism resources integrating text, images, and videos includes the following steps: S1. Collect image and text data related to Tibetan tourism resources through the Tibetan Landscape and Cultural Tourism Intelligent Analysis Platform. The image and text data includes images of Tibetan natural landscapes, images of cultural landscapes, and tourism text descriptions. During the collection process, extract the spatial location features, visual texture features, and textual semantic features of the image and text data. Specifically, the implementation process of step S1 is as follows: The data acquisition module of the Tibetan Scenic Area Cultural Tourism Intelligent Analysis Platform is activated. This module establishes an API interface connection with the database of the Tibetan tourism management department, the scenic area monitoring system, and the tourism information platform to collect image and text data related to Tibetan tourism resources. The image data resolution is set to 1920×1080 pixels, the single image storage format is JPEG, and the text data adopts UTF-8 encoding format. The number of characters in a single text description is controlled within the range of 200-500 characters. During the acquisition process, the spatial location features, visual texture features, and text semantic features of the image and text data are extracted using the platform's built-in feature extraction component. GPS positioning parameters are used when extracting spatial location features. The accuracy of data collection is controlled within 10 meters, and the longitude and latitude coordinates of the resource are recorded. When extracting visual texture features, gray-level co-occurrence matrix parameters are selected, and four texture indices—contrast, correlation, energy, and homogeneity—are calculated, with each index ranging from 0 to 1. When extracting textual semantic features, word segmentation tools are used to process the textual data, with the segmentation granularity set to the word level. The top 20 most frequently occurring keywords are extracted as semantic features, and the co-occurrence frequency among keywords is calculated. The co-occurrence frequency statistics window is set to 50 characters to construct textual semantic association features. After collection, all extracted feature data is stored in the platform's temporary database, with the data storage rate controlled at 100 records / second to ensure the real-time performance and integrity of data collection.

[0019] S2. Input the image and text data collected in S1 into the multi-scale visual analysis model of the Tibetan landscape. The model performs multi-scale feature decomposition on the image data to obtain local and global features of the image at different resolutions. At the same time, semantic word segmentation is performed on the text data to extract text keyword features and semantic association features. Specifically, the implementation process of step S2 is as follows: Through the data transmission component of the Zangjing Cultural Tourism Intelligent Analysis Platform, the image and text data collected in S1 and stored in the temporary database are transmitted to the input layer of the Zangjing multi-scale visual analysis model. The model input data volume is set to 500 data entries per batch, and the number of processing threads is set to 8 threads to improve processing efficiency. When the model performs multi-scale feature decomposition on the image data, three scale levels are set: low scale (image size scaled to 480×270 pixels), medium scale (image size maintained at 960×540 pixels), and high scale (image size maintained at 1920×1080 pixels). At each scale level, convolution operations are used to extract local and global features of the image. When extracting local features, the convolution kernel size is set to 3×3, and the stride is set to 1. Local detail features such as image edges, textures, and color blocks are extracted. The feature vector dimension is set to 256. For global feature extraction, average pooling is used with an 8×8 kernel size and a stride of 8, integrating local features into a global feature vector with a dimension of 128. Simultaneously, the model performs semantic word segmentation on the text data, employing a dictionary-based forward maximum matching algorithm. The dictionary includes 3000 Tibetan tourism-specific terms. After segmentation, a bag-of-words model is used to extract text keyword features, with a keyword feature vector dimension of 64. Cosine similarity is then used to calculate the semantic correlation between keywords, with a correlation value ranging from 0 to 1, thus constructing text semantic correlation features. After processing, the image local features, image global features, text keyword features, and text semantic correlation features are output to the model result cache, with a cache capacity of 1000 entries to ensure fast data retrieval in subsequent steps.

[0020] S3. Call the Tibet Travel Image-Text Cross-Modal Alignment Algorithm to perform cross-modal feature mapping between the image features and text features output by S2, establish the association mapping relationship between image features and text features, calculate the feature similarity threshold through the algorithm, and select cross-modal feature pairs that meet the similarity requirements. Specifically, the implementation process of step S3 is as follows: In the algorithm call module of the Tibetan Tourism Intelligent Analysis Platform, the execution command of the Tibetan Tourism Image-Text Cross-Modal Alignment Algorithm is triggered. The algorithm first reads image features and text features from the model result cache output by S2, with the reading rate set to 200 records / second to ensure stable data reading. When performing cross-modal feature mapping on image features and text features, the algorithm uses a feature transformation matrix to map image features and text features to the same 128-dimensional feature space. The dimension of the feature transformation matrix is ​​set to 128×(256+64), and the matrix element values ​​range from -1 to 1. Feature space transformation is achieved through matrix multiplication. After the transformation is completed, the algorithm calculates the cosine similarity between image features and text features. The similarity calculation window is set for each image-text data pair, that is, one image data corresponds to one text data for similarity calculation. A similarity threshold of 0.6 was set to filter cross-modal feature pairs with a similarity greater than 0.6. To further improve alignment accuracy, the algorithm also introduced a spatial location consistency verification mechanism. During verification, the GPS positioning coordinates of the image data were compared with the geographical location information mentioned in the text data. The geographical location information was extracted using named entity recognition technology, with an accuracy rate set above 95%. If the geographical location deviation between the two was less than 500 meters, the two were determined to be spatially consistent, and the cross-modal feature pair was retained. If the deviation was greater than 500 meters, the two were determined to be spatially inconsistent, and the cross-modal feature pair was removed. The finally selected cross-modal feature pairs were stored in the algorithm result database, which was a MySQL relational database. The table structure included fields for feature pair ID, image feature vector, text feature vector, similarity value, and spatial location deviation value, which facilitated subsequent steps for retrieval.

[0021] S4. Start the snowfield spatial clustering recommendation algorithm, combine the cross-modal feature pairs obtained in S3 with the spatial location parameters of Tibetan tourism resources, perform spatial clustering analysis, divide different tourism resource clusters, and calculate the feature center value and intra-cluster feature dispersion of each cluster. Specifically, the implementation process of step S4 is as follows: The task scheduling module of the Tibetan Tourism Intelligent Analysis Platform is used to launch the Snowland Spatial Clustering Recommendation Algorithm. The algorithm reads cross-modal feature pairs from the algorithm result database in S3, with each read set to 1000 data entries. After reading, the data is preprocessed, which only includes data format unification and converting the feature vector format into JSON format recognizable by the algorithm; normalization is not involved. The algorithm performs spatial clustering analysis based on the spatial location parameters of Tibetan tourism resources. The spatial location parameters are the longitude and latitude coordinates extracted in S1. The K-means clustering algorithm framework is used for clustering, with an initial value of 20 clusters and a maximum number of clustering iterations set to 50. The Euclidean distance is used as the clustering distance metric, and a feature similarity weighting factor is introduced, with a weighting factor value of 0.7, i.e., spatial distance... The weights are 0.3 and feature similarity weights are 0.7. The comprehensive distance between each cross-modal feature pair and the center of each cluster is calculated by weighting. Based on the comprehensive distance, the cross-modal feature pairs are assigned to the nearest cluster. After the assignment, the feature center value of each cluster is updated. The feature center value is calculated by averaging all feature vectors within the cluster. At the same time, the intra-cluster feature dispersion of each cluster is calculated by averaging the Euclidean distances between all feature vectors within the cluster and the central feature vector. The dispersion threshold is set to 0.5. If the intra-cluster feature dispersion of all clusters is less than 0.5 after a certain iteration, or if the number of iterations reaches 50, the iteration stops and the final clustering result is output. The clustering result includes the feature center value of each cluster, the intra-cluster feature dispersion, the number of cross-modal feature pairs within the cluster, and the spatial distribution range of resources within the cluster. The result is stored in the clustering result database.

[0022] S5. Based on the clustering results of S4, construct a Tibetan tourism resource feature database in the Tibetan Scenic Area Cultural Tourism Intelligent Analysis Platform. Store the image and text features, spatial location parameters, and clustering feature parameters corresponding to each cluster into the database, and establish a database index structure for fast feature retrieval. Specifically, the implementation process of step S5 is as follows: The database construction module of the Tibetan Tourism Intelligent Analysis Platform receives the clustering results output by S4. First, it classifies and organizes the clustering result data, dividing the data into two categories: clustering parameter data and feature data. The clustering parameter data includes cluster ID, feature center value, intra-cluster feature dispersion, and number of intra-cluster samples. The feature data includes cross-modal feature pair ID, local image features, global image features, text keyword features, text semantic association features, and spatial location coordinates. Subsequently, a Tibetan tourism resource feature database is built in the platform's built-in database server. The database adopts a distributed storage architecture, consisting of a master database and two slave databases. The master database is responsible for writing data, and the slave databases are responsible for reading data. The database storage engine uses InnoDB and supports transaction processing to ensure data consistency. Four core data tables are created for the database: a clustering parameter table, an image feature table, a text feature table, and a spatial location table. Each table establishes a foreign key relationship through the cluster ID and feature pair ID. The clustering parameter table fields... The database includes cluster ID, feature center value, dispersion, number of samples, and creation time. Image feature table fields include feature pair ID, cluster ID, local feature vector, and global feature vector. Text feature table fields include feature pair ID, cluster ID, keyword feature vector, and semantic relevance. Spatial location table fields include feature pair ID, cluster ID, longitude, and latitude. The categorized clustering parameter data is written to the clustering parameter table, and feature data is written to the corresponding data tables. Batch insertion is used, with 500 records inserted at a time at a rate of 50 records per second to avoid excessive database load. After data storage, a database index is built based on the cluster ID, feature pair ID, longitude, and latitude fields. The index type is a B+ tree index, with the cluster ID as the clustered index and the other fields as non-clustered indexes. After index construction, retrieval testing is performed, randomly selecting 100 cluster IDs for querying. The query response time is controlled within 0.5 seconds to ensure fast feature retrieval.

[0023] S6. Based on the user's travel demand parameters, the Tibetan tourism resource intelligent push module calls the data in the feature database, combines it with the cluster feature center value of S4, calculates the matching degree between the user's demand and each cluster, and outputs the corresponding Tibetan tourism resource push information according to the matching degree ranking result.

[0024] Specifically, the implementation process of step S6 is as follows: The Tibetan tourism resource intelligent push module receives the tourism demand parameters input by the user through the user interaction interface. The demand parameters include the user's preferred type of tourism resource (such as natural landscapes and cultural landscapes), the desired tourism area (specified by longitude and latitude range, with a range precision of 1 degree), the tourism time (seasonal information, divided into four seasons: spring, summer, autumn, and winter), and the number of tourists (three levels: 1-2 people, 3-5 people, and 6 people or more). The parameter input format is in the form of a form. After the user submits the data, the module verifies the validity of the parameters. The verification rules are: longitude range -90 to 90 degrees, latitude range -180 to 180 degrees, and the season and number of tourists must meet the set level. After the verification is passed, the module enters the matching stage. The module calls the data in the Tibetan tourism resource feature database through the database query interface. The query conditions are set to cluster data within the longitude and latitude range specified by the user. The query results include cluster ID, feature center value, and resource type within the cluster. The module first calculates the spatial location coordinates; then, it combines the cluster feature center values ​​output by S4 to calculate the matching degree between user needs and each cluster. The matching degree calculation adopts a weighted summation method, with the resource type matching weight set to 0.4, the region matching weight set to 0.3, the season matching weight set to 0.2, and the number of tourists matching weight set to 0.1. The score range of each matching dimension is 0-100 points. After weighting, the total matching degree is obtained, with the total matching degree range of 0-100 points. The clusters are sorted from high to low according to the total matching degree, and the top 10 clusters are selected as the recommendation results. Each recommendation result includes the image and text information, spatial location information, and recommendation reasons (generated based on the scores of each dimension of the matching degree) of 3-5 representative tourist resources in the cluster. Finally, the recommendation results are output through the user interaction terminal (supporting web and mobile terminals). The output format supports two forms: image and text list and map annotation. Users can click to view detailed information. The push response time is controlled within 2 seconds to ensure user experience.

[0025] Preferably, the expression of the multi-scale visual analysis model of the hidden scenery is: ,in Indicated in scale Lower image coordinates Visual analysis results at the location, Indicates the first Weight parameters for each feature extraction channel, Indicates the first A nonlinear feature mapping function, Representing scale lower coordinate The original image pixel values ​​at that location, This represents the convolution operation. Representing scale Next Each convolutional kernel parameter, Represents the semantic feature fusion coefficient. Representing scale lower coordinate The semantic association feature value of the text at that location This indicates the total number of feature extraction channels.

[0026] Specifically, the implementation process of the Tibetan scenic area multi-scale visual analysis model is as follows: The model is loaded into the model deployment module of the Tibetan scenic area cultural tourism intelligent analysis platform. During model execution, image data collected in S2 is first read. The images are scaled according to three set scale levels (low scale 480×270 pixels, medium scale 960×540 pixels, and high scale 1920×1080 pixels). Each scale level corresponds to a different feature extraction channel, with a total of eight channels. The weight parameters for each channel are preset based on the characteristics of Tibetan tourism resource images, with values ​​ranging from 0.1 to 0.3 to ensure reasonable weight allocation for different features. The convolutional kernel parameters used in the model are adjusted according to the scale level: 16 kernels at the low scale, 32 at the medium scale, and 64 at the high scale. All convolutional kernels are of uniform size. The model is set to 3×3, and image features at each scale are extracted through convolution operations. The ReLU function is used as the nonlinear feature mapping function to perform nonlinear transformation on the feature values ​​after convolution operations, avoiding excessive concentration of feature values. The semantic feature fusion coefficient is set to 0.4 to fuse text semantic association features with image features. The text semantic association feature values ​​are obtained through word segmentation and semantic analysis in S2, and the values ​​range from 0 to 1. During the model operation, the visual resolution results at the image coordinates at each scale are calculated by multi-channel feature weighted summation and semantic feature fusion. After the calculation is completed, the visual resolution results at each scale are output for subsequent cross-modal alignment processing. This model improves the resolution accuracy of details and overall features of Tibetan tourism resource images through multi-scale feature extraction and semantic fusion, laying the foundation for subsequent feature alignment.

[0027] Preferably, the expression for the cross-modal alignment algorithm for Tibetan travel images and text is: ,in Representing an image With words Cross-modal alignment Indicates the image number 1 The first feature and text The association weights of each feature Representing image features With textual features The included angle, Indicates the number of dimensions of image features. This represents the number of dimensions of text features. Indicates the spatial alignment coefficient. This represents the spatial distance parameter between the image and the text.

[0028] Specifically, the implementation process of the Tibetan tourism image-text cross-modal alignment algorithm is as follows: The algorithm is started in the algorithm running module of the Tibetan Scenic Tourism Intelligent Analysis Platform. First, it reads the image features and text features output by S2. The image features include 256-dimensional local features and 128-dimensional global features, and the text features are 64-dimensional keyword features. The algorithm first performs dimension unification processing on these features, concatenating the image local features and global features into a 384-dimensional vector, while keeping the text features unchanged at 64 dimensions. Then, a feature transformation matrix is ​​constructed. The matrix dimensions are set according to the concatenated feature dimensions to ensure that the image and text features can be mapped to the same 128-dimensional feature space. The matrix elements are obtained through training with a large amount of Tibetan tourism image-text data, and the values ​​range from -0.8 to 0.8 to avoid feature distortion caused by excessively large values. The association weight parameter is preset based on the relevance of the image and text data. For natural landscape image and text pairs, the association weight value is between 0.6 and 0.8, and for cultural landscape pairs, it is between 0.5 and 0.7. The spatial location alignment coefficient is set to 0.3 to balance the influence of feature similarity and spatial location consistency in the alignment result. When the algorithm calculates the angle between image and text features, it uses the ratio of vector dot product to modulus. The similarity calculation window is set to a single image and text pair. After filtering out feature pairs with similarity greater than 0.6, a second filtering is performed in combination with the spatial location distance parameter (deviation less than 500 meters). Finally, the algorithm outputs cross-modal feature pairs that meet the requirements. This algorithm improves the alignment accuracy of Tibetan tourism image and text data and reduces cross-modal information deviation through feature mapping and double filtering.

[0029] Preferably, the expression for the snowfield spatial clustering recommendation algorithm is: ,in This indicates the spatial clustering degree of tourism resource samples. This represents the total number of samples within a cluster. Indicates the spatial location weighting coefficient. Indicates the spatial distance attenuation coefficient. Indicates the first Spatial location coordinate parameters of each sample The spatial coordinates of the cluster center are represented by the following parameters. This represents the feature similarity weight coefficient. Indicates the first A set of image and text features for each sample. A set of graphic and textual features representing the center of a cluster.

[0030] Specifically, the implementation process of the snow-covered area spatial clustering recommendation algorithm is as follows: After the algorithm starts, it reads cross-modal feature pairs from the S3 algorithm result database, reading 1000 pairs at a time, and simultaneously extracting the spatial location coordinates (longitude and latitude) corresponding to each pair of data. In the algorithm initialization phase, the initial number of clusters is set to 20, the maximum number of clustering iterations is 50, the spatial location weight coefficient is set to 0.3, and the feature similarity weight coefficient is set to 0.7. The influence of spatial location and feature attributes in clustering is balanced through weight allocation. The spatial distance attenuation coefficient is set to 0.02 based on the distribution density of tourism resources in Tibetan areas. In areas with higher density (such as the area around Lhasa), the attenuation coefficient can be appropriately increased to 0.03 to ensure that the clustering results can reflect the distribution density of tourism resources in Tibetan areas. The algorithm considers the spatial clustering characteristics of resources. During clustering, it first calculates the spatial distance between each sample and the center of each cluster using the Euclidean distance formula. Simultaneously, it calculates the feature similarity between the sample and the cluster center, obtained by the ratio of the number of elements in the intersection and union. Then, it weights the spatial distance and feature similarity to obtain a comprehensive distance, and assigns the sample to the nearest cluster based on this comprehensive distance. After each iteration, it updates the spatial location and feature set of the cluster centers, calculates the feature dispersion within each cluster, and sets a dispersion threshold of 0.5. Iteration stops when the dispersion of all clusters is less than the threshold or the number of iterations reaches its maximum, and the clustering results are output. This algorithm achieves accurate clustering of Tibetan tourism resources through a dual consideration of spatial and feature factors, providing a classification basis for subsequent recommendations.

[0031] Preferably, the resource feature processing expression of the Tibetan Scenic Area Cultural Tourism Intelligent Analysis Platform is: ,in This indicates the platform's priority in processing the characteristics of tourism resources. Indicates the feature processing weight coefficients. Indicates the total number of feature dimensions. Indicates the first The processing weights of each feature, The first characteristic vector of a resource is represented by the... One portion, Represents the first feature vector in the database One portion, Represents the cluster association coefficient. This parameter indicates the importance of the cluster to which the resource belongs.

[0032] Specifically, the resource feature processing implementation process of the Tibetan Tourism Intelligent Analysis Platform is as follows: The platform's feature processing module receives the clustering results output by S4 and the resource feature data output by S2. First, the feature data is standardized, mapping the feature values ​​of each dimension to the range of 0-1. The feature processing weight coefficient is set to 0.6 to ensure the dominant role of feature similarity in the feature processing priority calculation. The total number of feature dimensions is set to 64 based on the feature types of Tibetan tourism resources, including multiple dimensions such as image texture, color, and text semantics. The processing weight of each feature dimension is allocated according to its influence on resource classification; the image texture feature weight is set to 0.2. -0.3, the text semantic feature weights range from 0.15 to 0.25; the clustering correlation coefficient is set to 0.4, used to associate the importance of resource features with their respective clusters; when calculating the processing priority of resource features, the platform first transforms the sum of feature similarity results using a logarithmic function to avoid excessively large values, and then combines them with clustering importance parameters for weighting to obtain priority values, ranging from 0 to 10; resource features with higher priority (greater than 7) will be prioritized for storage and retrieval optimization to ensure that core feature data can be quickly accessed during subsequent pushes. This platform improves the efficiency of resource feature processing and ensures the high efficiency of database operation through priority sorting.

[0033] Preferably, the matching calculation expression for the intelligent push of Tibetan tourism resources is: ,in Indicate user needs With clusters The matching degree of the push notifications Indicates the number of dimensions of the demand features. Indicates the first Weight parameters for each demand characteristic, Indicating the first in user requirements The values ​​of each feature, Represents clusters The Middle The mean of each feature, To represent the minimum parameter to avoid the denominator being zero, Represents clusters The Middle The standard deviation of each feature Indicates the first The matching gain coefficient of each feature.

[0034] Specifically, the matching calculation process for the intelligent recommendation of Tibetan tourism resources is as follows: After receiving the user's input demand parameters, the intelligent recommendation module first converts the parameters into a computable feature vector. The demand feature dimension is set to 8 dimensions, including resource type, region, season, number of tourists, etc. The weight parameter of each demand feature is set according to the user demand survey results: resource type weight 0.4, region weight 0.3, season weight 0.2, number of tourists weight 0.1, and other auxiliary feature weights totaling 0.1. The minimum value parameter is set to 0.001 to avoid the denominator being zero during calculation. The feature standard deviation is calculated through the features of all samples within the cluster, with a value range between 0.1 and 0.5. The gain coefficient is set according to the degree of influence of the feature on the push results. The gain coefficient for resource type and regional features is 1.2, and the gain coefficient for seasonality and number of tourists is 1.0. When calculating the matching degree, the difference between the user demand feature and the mean of the cluster feature is first normalized by the Sigmoid function to control the result between 0 and 1. Then, it is multiplied by the weight and gain coefficient of the corresponding feature, and finally summed to obtain the total matching degree. The total matching degree ranges from 0 to 100 points. The top 10 clusters are selected as the recommendation results according to the scores from high to low. This matching calculation method improves the matching accuracy of user needs and resources through multi-dimensional weight allocation and normalization processing, and realizes personalized push.

[0035] Preferably, step S3 includes the following sub-steps: S31, extracting the multi-scale visual feature matrix of the image and the semantic feature vector of the text from the output of the Tibetan landscape multi-scale visual analysis model, determining that the row dimension of the image feature matrix is ​​the number of image feature types and the column dimension is the number of feature dimensions, and that the dimension of the text feature vector is consistent with the column dimension of the image feature matrix; S32, inputting the image feature matrix and the text feature vector into the feature mapping module of the Tibetan tourism image-text cross-modal alignment algorithm, mapping the image features and text features to the same feature space through the feature transformation matrix within the module, generating a cross-modal feature matrix; S33, calculating the cosine similarity between each row of image features and the corresponding text features in the cross-modal feature matrix, setting an initial similarity threshold, and filtering out feature pairs with similarity greater than the initial threshold to form a preliminary aligned feature set; S34, adjusting the similarity threshold based on the spatial position correlation of features in the preliminary aligned feature set, eliminating feature pairs with mismatched spatial positions, and finally obtaining a feature pair set that meets the requirements of cross-modal alignment.

[0036] Specifically, the implementation process of S3 revolves around cross-modal feature alignment, and the specific steps are as follows: In S31, data is extracted from the output of the Tibetan landscape multi-scale visual analysis model. The row dimension of the image multi-scale visual feature matrix is ​​set to 8 (corresponding to 8 types of image features, including edges, textures, colors, etc.), and the column dimension is set to 256 (the number of dimensions for each type of feature). The dimension of the text semantic feature vector is kept consistent with the column dimension of the matrix, both being 256, to ensure uniformity in subsequent feature processing. In S32, the above feature data is input into the feature mapping module of the Tibetan tourism image-text cross-modal alignment algorithm. The feature transformation matrix within the module is generated through training with 5000 sets of Tibetan tourism image-text samples. The matrix element values ​​range from -0.5 to 0.5. The image is transformed through matrix multiplication. The text features are mapped to the same 256-dimensional feature space to generate a cross-modal feature matrix. In stage S33, the cosine similarity between each row of image features and the corresponding text features in the matrix is ​​calculated. The initial threshold is set to 0.6, and feature pairs with similarity exceeding the threshold are selected to form a preliminary aligned feature set. The size of the set is controlled to be 70%-80% of the total data processed each time. In stage S34, the threshold is adjusted based on the location correlation in the feature space. By comparing the GPS coordinate deviation of the resource corresponding to the feature (the deviation is allowed to be less than 300 meters), feature pairs with mismatched locations are eliminated. The accuracy of the final aligned feature pairs needs to reach more than 92% to provide high-quality feature data for subsequent spatial clustering. This process improves the cross-modal alignment accuracy through multi-stage screening and ensures reliable feature correlation.

[0037] Preferably, step S4 includes the following sub-steps: S41, obtaining the cross-modal aligned feature pair set output by S3, extracting the spatial location coordinates of Tibetan tourism resources corresponding to each feature pair, and establishing a feature-location association data table, which includes feature pair identifiers, image feature values, text feature values, and spatial location coordinate parameters; S42, inputting the feature-location association data table into the initialization module of the snowfield spatial clustering recommendation algorithm, setting the initial value of the number of clusters and the maximum number of clustering iterations, and initializing the center position and center feature value of each cluster; S43, calculating the spatial distance and feature similarity between each feature pair and the center of each cluster using the algorithm, allocating the feature pairs to the corresponding clusters based on the weighted result of distance and similarity, and updating the center position and center feature value of each cluster; S44, repeating the calculation and allocation process of S43 until the number of clustering iterations reaches the maximum number or the change in the center position and center feature value of the cluster is less than a set threshold, stopping the iteration and outputting the final clustering result.

[0038] Specifically, S4 focuses on the spatial clustering operation in the snowy region, and the implementation details are as follows: In S41, the cross-modal aligned feature pair set output by S3 is obtained, and the spatial location coordinates of Tibetan tourism resources corresponding to each feature pair are extracted (longitude accuracy to 0.001 degrees, latitude accuracy to 0.001 degrees). A feature-location association data table is constructed, which includes 12 fields (feature pair identifier, 6 types of image feature values, 4 types of text feature values, and spatial coordinates). The data storage format is CSV for easy subsequent reading. In S42, the data table is input into the snowy region spatial clustering recommendation algorithm initialization module. The initial number of clusters is set to 15, and the maximum number of clustering iterations is set to 40. During initialization, 15 feature pairs are randomly sampled as initial cluster centers. The center feature values ​​are directly related to the spatial coordinates. The corresponding sample data is continued; in stage S43, the spatial distance (using the Manhattan distance formula) and feature similarity (calculated by cosine similarity) between each feature pair and each cluster center are calculated. The spatial distance weight is set to 0.3 and the feature similarity weight is set to 0.7. After weighting, the comprehensive distance is obtained. Feature pairs are assigned to the corresponding clusters according to the principle of minimum comprehensive distance. After every 100 data points are assigned, the feature value and coordinates of the cluster center are updated once (taking the average value of all samples in the cluster); in stage S44, the operation of S43 is repeated. When the number of iterations reaches 40 or the change in the feature value and coordinates of the cluster center is less than 0.01 in 3 consecutive iterations, the clustering results are output. The number of samples in each cluster is controlled between 50 and 100 to ensure uniform cluster distribution and provide a reasonable basis for resource classification.

[0039] Preferably, step S5 includes the following sub-steps: S51, receiving the clustering results output by S4, organizing the feature pair data, spatial location parameters, and cluster center parameters corresponding to each cluster, and classifying these data according to feature type to divide them into image feature class, text feature class, and spatial parameter class; S52, creating a Tibetan tourism resource feature database in the database module of the Tibetan Cultural Tourism Intelligent Analysis Platform, setting up an image feature table, a text feature table, a spatial parameter table, and a clustering information table for the database, and establishing associations between the tables through cluster identifiers and feature pair identifiers; S53, classifying the data from S51. The various types of data are then stored in their respective data tables. The image feature table records image feature values ​​and feature dimension information, the text feature table records text feature vectors and semantic association information, the spatial parameter table records spatial location coordinates and location association information, and the cluster information table records cluster center parameters and the number of samples within each cluster. S54. A database index is constructed based on the key fields in each data table. The index includes a cluster identifier index, a feature dimension index, and a spatial location index. The index optimizes the database retrieval speed, enabling fast queries of labeled clusters or labeled feature data.

[0040] Specifically, S5 revolves around the construction of the feature database, and the implementation process is as follows: In S51, the clustering results output from S4 are received, and the data is divided into three categories according to feature type: image feature category (containing local feature vectors (256 dimensions) and global feature vectors (128 dimensions); text feature category (containing keyword feature vectors (64 dimensions) and semantic association matrix (8×8 dimensions); and spatial parameter category (containing longitude, latitude, and region code (6-digit numbers)). After classification, the data integrity is checked, and the missing rate must be less than 0.5%. In S52, a Tibetan tourism resource feature database is created in the database module of the Tibetan Cultural Tourism Intelligent Analysis Platform. Using the PostgreSQL database management system, four core data tables are created (image feature table, text feature table, spatial parameter table, and clustering information table). Each table establishes a foreign key relationship between the cluster identifier (8-digit string) and the feature pair identifier (10-digit string). The table storage engine uses BTREE, supporting efficient index queries. In stage S53, the data classified in S51 is written into the corresponding tables. Each record in the image feature table stores two sets of vector data, the text feature table records keyword vectors and semantic matrices, the spatial parameter table records coordinates and region codes, and the clustering information table records cluster center parameters (feature vectors, coordinates) and the number of samples within each cluster. Data writing adopts a batch insertion method, inserting 200 records at a time, with the writing rate controlled at 30 records / second to avoid database overload. In stage S54, indexes are built based on the key fields of the data tables. The cluster identifier is set as a clustered index, and the feature pair identifier, longitude, and latitude are set as non-clustered indexes. After the index is built, performance testing is performed. The response time for a single data query should be less than 0.3 seconds, and the response time for a batch query (100 records) should be less than 5 seconds to ensure that the required resource feature data can be quickly retrieved during subsequent pushes, improving the overall system operating efficiency.

[0041] The Tibetan Scenery Multi-Scale Visual Analysis Model is a feature extraction model designed for Tibetan tourism resource image and text data. It is used to mine the core features of images and text from different resolution dimensions. The implementation process is as follows: After deployment on the Tibetan Scenery Cultural Tourism Intelligent Analysis Platform, the collected image and text data is received. The images are scaled at three levels: low (480×270 pixels), medium (960×540 pixels), and high (1920×1080 pixels). Each scale corresponds to eight feature extraction channels (weights 0.1-0.3). Local features (256 dimensions) and global features (128 dimensions) are extracted using 3×3 convolutional kernels (16 for low scale, 32 for medium scale, and 64 for high scale). Simultaneously, word-level segmentation (with a dictionary containing 3000 Tibetan tourism-specific terms) is used for the text data. The top 20 high-frequency keywords are extracted to construct a 64-dimensional feature vector. Semantic association features are generated by combining the keyword co-occurrence frequency (50-character window). Finally, the text features and image features are fused using a semantic fusion coefficient of 0.4 to output multi-dimensional features. The purpose of this model is to provide accurate and multi-dimensional feature data for subsequent cross-modal alignment, solving the problem that traditional single-scale feature extraction cannot take into account both the details of Tibetan tourism resources (such as mural textures and natural landscape layers) and overall features, improving the completeness and relevance of feature data, and laying the foundation for subsequent accurate push.

[0042] The Tibetan Tourism Image-Text Cross-Modal Alignment Algorithm is an algorithm for associating and matching the features of Tibetan tourism resources images and text descriptions, establishing a reliable mapping relationship between image and text features. Its implementation process is as follows: After the platform's algorithm execution module starts, it first reads the image features (256-dimensional local + 128-dimensional global stitched together to form 384-dimensional) and text features (64-dimensional) output from the Tibetan landscape multi-scale visual analysis model. Using a trained feature transformation matrix (elements ranging from -0.8 to 0.8), both are mapped to the same 128-dimensional feature space. Then, association weights are set (0.6-0.8 for natural landscapes, 0.5-0.7 for cultural landscapes), and the cosine similarity of the angle between the image and text features is calculated, with an initial threshold of 0.6 to filter feature pairs. Finally, a spatial location consistency check is introduced (GPS coordinate deviation less than 300 meters, named entity recognition accuracy above 95%) to eliminate mismatched feature pairs, outputting the final alignment result. The algorithm's function is to eliminate feature bias in cross-modal image-text data, ensuring that the image and text description point to the same tourism resource. This approach addresses the issue of traditional cross-modal alignment neglecting the spatial correlation and cultural specificity of Tibetan tourism resources, improves the accuracy of image-text feature matching (final accuracy rate over 92%), provides high-quality correlation feature data for subsequent cluster analysis, and ensures the accuracy of resource classification.

[0043] The Snowland Spatial Clustering Recommendation Algorithm is a classification algorithm based on the spatial location and feature attributes of Tibetan tourism resources. It is used to divide aligned image-text feature pairs into reasonable resource clusters. The implementation process is as follows: Feature pair data (1000 records each time) and GPS coordinates (accuracy 0.001 degrees) are read from the cross-modal alignment results. Clustering parameters are initialized (number of clusters 15, maximum number of iterations 40, spatial weight 0.3, feature similarity weight 0.7, distance decay coefficient 0.02-0.03). Using the K-means framework, spatial distance is calculated using Manhattan distance, feature similarity is calculated using cosine similarity, and a weighted comprehensive distance is obtained. Feature pairs are then assigned to the nearest cluster. The cluster centers are updated every 100 data points (average value is taken). The algorithm stops when 40 iterations are reached or the change in cluster centers is less than 0.01, and the clustering results are output (50-100 samples per cluster, dispersion less than 0.5). The algorithm's function is to classify Tibetan tourism resources according to spatial distribution and feature attributes, forming structured resource clusters. This approach addresses the issue of traditional clustering neglecting the spatial aggregation characteristics of resources in Tibetan areas (such as the differences between dispersed and concentrated tourist attractions), enabling precise resource classification. This provides a clear basis for subsequent database construction and personalized recommendations, thereby improving resource retrieval efficiency.

[0044] The Tibetan Cultural Tourism Intelligent Analysis Platform algorithm is an integrated algorithm system supporting the entire process of Tibetan tourism resource data processing, including data acquisition, feature processing, and database management. Its implementation process is as follows: In the data acquisition stage, multiple data sources are connected via API interfaces. Image and text data are collected at 1920×1080 pixels (JPEG), UTF-8 encoding (200-500 characters), extracting features such as GPS (accuracy 10 meters), gray-level co-occurrence matrix texture (4 indicators), and text semantics, with a storage rate of 100 records / second. In the feature processing stage, a feature processing weight coefficient of 0.6 is used to standardize 64-dimensional features (including image texture and text semantics) to a range of 0-1, and weights are assigned according to feature influence (texture 0). The processing priority (0-10 points) is calculated using a clustering association coefficient of 0.4 (2-0.3, semantic 0.15-0.25), prioritizing data with a priority > 7. In the database management stage, a distributed PostgreSQL database (1 master, 2 slaves) is constructed, creating 4 related tables (clustering parameter table, image feature table, etc.). Batch insertion (200 rows / time, 30 rows / second) is used to store data. A B+ tree index is built based on cluster identifiers to ensure query response time < 0.3 seconds (single row) and < 5 seconds (100 rows / batch). The platform algorithm aims to achieve intelligent processing of Tibetan tourism resource data from collection to storage, integrating technologies from various stages to solve the problems of fragmented and inefficient data processing in traditional platforms. It provides a stable operating environment for models and algorithms, ensuring the real-time performance, integrity, and efficiency of data processing, and supporting the stable operation of the entire intelligent push system.

[0045] like Figure 2 As shown, an intelligent push system for Tibetan tourism resources integrating images and text is presented. This system is applied to an intelligent push method for Tibetan tourism resources integrating images and text, and includes: a Tibetan landscape and cultural tourism intelligent analysis data acquisition unit, which establishes a data transmission connection with the data source of Tibetan tourism resources to collect image and text data of Tibetan tourism resources, extract spatial location features, visual texture features, and text semantic features of the data, and transmit the collected feature data to a Tibetan landscape multi-scale visual analysis processing unit; a Tibetan landscape multi-scale visual analysis processing unit, which receives the feature data transmitted by the Tibetan landscape and cultural tourism intelligent analysis data acquisition unit, performs multi-scale decomposition and feature extraction on the feature data through a built-in Tibetan landscape multi-scale visual analysis model, generates image local features, image global features, text keyword features, and text semantic association features, and sends the generated feature data to a Tibetan tourism image and text cross-modal alignment processing unit; and a Tibetan tourism image and text cross-modal alignment processing unit, which establishes data interaction with the Tibetan landscape multi-scale visual analysis processing unit, and calls a Tibetan tourism image and text cross-modal alignment algorithm to perform cross-modal mapping of the received image features and text features. The system establishes feature associations and filters feature pairs that meet similarity requirements, then transmits the feature pair data to the Tibetan tourism resource feature database storage unit. This unit receives feature pair data from the Tibetan tourism image-text cross-modal alignment processing unit, performs spatial clustering analysis on the feature pairs using its built-in Tibetan tourism resource feature database storage algorithm, divides tourism resource clusters, calculates clustering parameters, and transmits the clustering results data to the Tibetan tourism resource feature database storage unit. This unit, connected to the Tibetan tourism resource feature database storage unit, receives the clustering results data, constructs a feature database, establishes data tables and index structures, stores cluster features, image-text features, and spatial location parameters, and establishes a data retrieval connection with the Tibetan tourism resource intelligent push matching unit. The Tibetan tourism resource intelligent push matching unit receives tourism demand parameters input by the user, retrieves data from the Tibetan tourism resource feature database storage unit, calculates the matching degree between the user's demand and the clusters, sorts the data based on the matching degree, generates tourism resource push information, and outputs it to the user's interactive terminal.

[0046] This invention presents an intelligent method and system for pushing Tibetan tourism resources, integrating text and imagery. It utilizes a multi-scale visual analysis model of Tibetan landscapes to perform multi-scale feature decomposition on text and image data of Tibetan tourism resources. This allows for the extraction of both local and global image features, as well as the mining of textual keywords and semantic associations, achieving in-depth analysis of resource data. In the resource clustering stage, a snow-covered spatial clustering recommendation algorithm combined with spatial location parameters of tourism resources is used for cluster analysis. This accurately divides different resource clusters and calculates feature parameters within each cluster, providing a reliable basis for resource classification and management. In the push matching stage, leveraging the feature database and intelligent push module built on the Tibetan landscape and cultural tourism intelligent analysis platform, resource feature data can be quickly retrieved. The system calculates the matching degree based on user needs and outputs the results, significantly improving push efficiency and accuracy. Simultaneously, the various units of the system work together to form a complete intelligent process from data collection to push output, ensuring overall operational stability.

[0047] To address the issue of poor cross-modal data alignment, this method and system specifically employ a Tibetan tourism image-text cross-modal alignment algorithm. This algorithm maps image features and text features to the same space and establishes a correlation. By calculating feature similarity, it filters feature pairs with high matching degrees and optimizes the feature mapping mechanism in conjunction with the characteristics of Tibetan tourism resources. This completely solves the problem of image-text information matching deviation and ensures the effectiveness of subsequent data processing. Regarding the insufficient integration of spatial clustering and push matching, the method utilizes a snow-covered spatial clustering recommendation algorithm that fully considers the spatial distribution characteristics of Tibetan areas. Location parameters are incorporated during clustering to improve the accuracy of cluster division. Furthermore, during the push phase, the feature center values ​​from the clustering results are directly called to calculate the user demand matching degree, strengthening the correlation between clustering and push, significantly reducing the deviation between push results and user needs, and achieving the goal of personalized push.

[0048] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0049] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for intelligently pushing Tibetan tourism resources by integrating text, images, and videos, characterized in that: Includes the following steps: S1. Collect image and text data related to Tibetan tourism resources through the Tibetan Scenic Tourism Intelligent Analysis Platform. The image and text data includes images of Tibetan natural landscapes, images of cultural landscapes, and tourism text descriptions. During the collection process, spatial location features, visual texture features, and text semantic features of the image and text data are extracted. S2. Input the image and text data collected in S1 into the Tibetan Scenic Multi-Scale Visual Analysis Model. The model performs multi-scale feature decomposition on the image data to obtain local and global features of the image at different resolutions. At the same time, semantic word segmentation is performed on the text data to extract text keyword features and semantic association features. S3. Call the Tibetan tourism image and text cross-modal alignment algorithm to perform cross-modal feature mapping between the image features and text features output in S2, establish the association mapping relationship between image features and text features, calculate the feature similarity threshold through the algorithm, and select cross-modal feature pairs that meet the similarity requirements. S4. Start the snowfield spatial clustering recommendation algorithm, combine the cross-modal feature pairs obtained in S3 with the spatial location parameters of Tibetan tourism resources, perform spatial clustering analysis, divide different tourism resource clusters, and calculate the feature center value and intra-cluster feature dispersion of each cluster. S5. Based on the clustering results of S4, construct a Tibetan tourism resource feature database in the Tibetan Scenic Tourism Intelligent Analysis Platform. Store the image and text features, spatial location parameters, and clustering feature parameters corresponding to each cluster into the database, and establish a database index structure for rapid feature retrieval. S6. According to the user's tourism demand parameters, call the data in the feature database through the Tibetan tourism resource intelligent push module, combine it with the cluster feature center value of S4, calculate the matching degree between the user's demand and each cluster, and output the corresponding Tibetan tourism resource push information based on the matching degree ranking results.

2. The method for intelligently pushing Tibetan tourism resources by integrating text and images according to claim 1, characterized in that, The expression for the multi-scale visual analysis model of the hidden landscape is: ,in Indicated in scale Lower image coordinates Visual analysis results at the location, Indicates the first Weight parameters for each feature extraction channel, Indicates the first A nonlinear feature mapping function, Representing scale lower coordinate The original image pixel values ​​at that location, This represents the convolution operation. Representing scale Next Each convolutional kernel parameter, Represents the semantic feature fusion coefficient. Representing scale lower coordinate The semantic association feature value of the text at that location This indicates the total number of feature extraction channels.

3. The method for intelligently pushing Tibetan tourism resources by integrating text and images according to claim 1, characterized in that, The expression for the cross-modal alignment algorithm for Tibetan travel images and text is: ,in Representing an image With words Cross-modal alignment Indicates the image number 1 The first feature and text The association weights of each feature Representing image features With textual features The included angle, Indicates the number of dimensions of image features. This represents the number of dimensions of text features. Indicates the spatial alignment coefficient. This represents the spatial distance parameter between the image and the text.

4. The method for intelligently pushing Tibetan tourism resources by integrating text and images according to claim 1, characterized in that, The expression for the snowfield spatial clustering recommendation algorithm is: ,in This indicates the spatial clustering degree of tourism resource samples. This represents the total number of samples within a cluster. Indicates the spatial location weighting coefficient. Indicates the spatial distance attenuation coefficient. Indicates the first Spatial location coordinate parameters of each sample The spatial coordinates of the cluster center are represented by the following parameters. This represents the feature similarity weight coefficient. Indicates the first A set of image and text features for each sample. A set of graphic and textual features representing the center of a cluster.

5. The method for intelligently pushing Tibetan tourism resources by integrating text and images according to claim 1, characterized in that, The resource feature processing expression of the Tibetan Scenic Area Cultural Tourism Intelligent Analysis Platform is as follows: ,in This indicates the platform's priority in processing the characteristics of tourism resources. Indicates the feature processing weight coefficients. Indicates the total number of feature dimensions. Indicates the first The processing weights of each feature, The first characteristic vector of a resource is represented by the... One portion, Represents the first feature vector in the database One portion, Represents the cluster association coefficient. This parameter indicates the importance of the cluster to which the resource belongs.

6. The method for intelligently pushing Tibetan tourism resources by integrating text and images according to claim 1, characterized in that, The matching calculation expression for the intelligent push of Tibetan tourism resources is as follows: ,in Indicate user needs With clusters The matching degree of the push notifications Indicates the number of dimensions of the demand features. Indicates the first Weight parameters for each demand characteristic, Indicating the first in user requirements The values ​​of each feature, Represents clusters The Middle The mean of each feature, To represent the minimum parameter to avoid the denominator being zero, Represents clusters The Middle The standard deviation of each feature Indicates the first The matching gain coefficient of each feature.

7. The method for intelligently pushing Tibetan tourism resources by integrating text and images according to claim 1, characterized in that, S3 includes the following steps: S31, extracting the multi-scale visual feature matrix of the image and the semantic feature vector of the text from the output of the Tibetan landscape multi-scale visual analysis model, determining that the row dimension of the image feature matrix is ​​the number of image feature types and the column dimension is the number of feature dimensions, and that the dimension of the text feature vector is consistent with the column dimension of the image feature matrix; S32, inputting the image feature matrix and the text feature vector into the feature mapping module of the Tibetan tourism image-text cross-modal alignment algorithm, mapping the image features and text features to the same feature space through the feature transformation matrix within the module, generating a cross-modal feature matrix; S33, calculating the cosine similarity between each row of image features and the corresponding text features in the cross-modal feature matrix, setting an initial similarity threshold, filtering out feature pairs with similarity greater than the initial threshold, and forming a preliminary aligned feature set; S34. Based on the spatial correlation of features in the initial aligned feature set, adjust the similarity threshold, remove feature pairs with mismatched spatial positions, and finally obtain a feature pair set that meets the cross-modal alignment requirements.

8. The method for intelligently pushing Tibetan tourism resources by integrating text and images according to claim 1, characterized in that, S4 includes the following sub-steps: S41, obtain the cross-modal aligned feature pair set output by S3, extract the spatial location coordinates of Tibetan tourism resources corresponding to each feature pair, and establish a feature-location association data table, which includes feature pair identifiers, image feature values, text feature values, and spatial location coordinate parameters; S42, input the feature-location association data table into the initialization module of the snowfield spatial clustering recommendation algorithm, set the initial value of the number of clusters and the maximum number of clustering iterations, and initialize the center position and center feature value of each cluster; S43, calculate the spatial distance and feature similarity between each feature pair and the center of each cluster through the algorithm, allocate the feature pair to the corresponding cluster according to the weighted result of distance and similarity, and update the center position and center feature value of each cluster; S44, repeat the calculation and allocation process of S43 until the number of clustering iterations reaches the maximum number or the change in the center position and center feature value of the cluster is less than the set threshold, stop the iteration and output the final clustering result.

9. The method for intelligently pushing Tibetan tourism resources by integrating text and images according to claim 1, characterized in that, S5 includes the following steps: S51, receiving the clustering results output from S4, organizing the feature pair data, spatial location parameters, and cluster center parameters corresponding to each cluster, and classifying these data according to feature type, dividing them into image feature class, text feature class, and spatial parameter class; S52, creating a Tibetan tourism resource feature database in the database module of the Tibetan Cultural Tourism Intelligent Analysis Platform, setting up an image feature table, text feature table, spatial parameter table, and cluster information table for the database, and establishing associations between the tables through cluster identifiers and feature pair identifiers; S53, processing the data classified in S51... Various types of data are stored in corresponding data tables. The image feature table records image feature values ​​and feature dimension information, the text feature table records text feature vectors and semantic association information, the spatial parameter table records spatial location coordinates and location association information, and the cluster information table records cluster center parameters and the number of samples within the cluster. S54. A database index is constructed based on the key fields in each data table. The index includes a cluster identifier index, a feature dimension index, and a spatial location index. The index optimizes the database retrieval speed and enables fast queries on labeled clusters or labeled feature data.

10. An intelligent push system for Tibetan tourism resources integrating text, images, and video, characterized in that: This system is applied to the intelligent push method for Tibetan tourism resources integrating images and text as described in claim 1, comprising: a Tibetan landscape and cultural tourism intelligent analysis data acquisition unit, which establishes a data transmission connection with the Tibetan tourism resource data source, for acquiring image and text data of Tibetan tourism resources, extracting spatial location features, visual texture features, and textual semantic features of the data, and transmitting the acquired feature data to a Tibetan landscape multi-scale visual analysis processing unit; a Tibetan landscape multi-scale visual analysis processing unit, which receives the feature data transmitted by the Tibetan landscape and cultural tourism intelligent analysis data acquisition unit, performs multi-scale decomposition and feature extraction on the feature data through a built-in Tibetan landscape multi-scale visual analysis model, generates image local features, image global features, text keyword features, and textual semantic association features, and sends the generated feature data to a Tibetan tourism image and text cross-modal alignment processing unit; and a Tibetan tourism image and text cross-modal alignment processing unit, which establishes data interaction with the Tibetan landscape multi-scale visual analysis processing unit, calls the Tibetan tourism image and text cross-modal alignment algorithm to perform cross-modal mapping of the received image features and text features, and establishes feature associations. The system first identifies and filters feature pairs that meet similarity requirements, then transmits the feature pair data to the Tibetan tourism resource spatial clustering recommendation processing unit. This unit receives feature pair data from the Tibetan tourism image-text cross-modal alignment processing unit, performs spatial clustering analysis on the feature pairs using its built-in Tibetan tourism resource clustering recommendation algorithm, divides tourism resource clusters, calculates clustering parameters, and transmits the clustering results data to the Tibetan tourism resource feature database storage unit. This unit, connected to the Tibetan tourism resource spatial clustering recommendation processing unit, receives the clustering results data, constructs a feature database, establishes data tables and index structures, stores cluster features, image-text features, and spatial location parameters, and establishes a data retrieval connection with the Tibetan tourism resource intelligent push matching unit. Finally, the Tibetan tourism resource intelligent push matching unit receives user-input tourism demand parameters, retrieves data from the Tibetan tourism resource feature database storage unit, calculates the matching degree between user demands and clusters, sorts the data based on the matching degree, generates tourism resource push information, and outputs it to the user's interactive terminal.