A Cultural and Tourism Content Analysis Method and System Based on Semantic Recognition

By building a grid-time and spatial module set and cross-modal semantic recognition technology, the problem that traditional cultural and tourism recommendation systems are difficult to discover micro-destinations is solved, timely and accurate identification and recommendation of emerging tourism hotspots is achieved, and the development efficiency of the cultural and tourism industry is improved.

CN120235732BActive Publication Date: 2025-08-01CHONGQING TOURISM CLOUD INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510703104.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Traditional cultural and tourism recommendation systems rely on static lists of interest points and manual predefined tags, making it difficult to discover and cover "micro destinations" in a timely and comprehensive manner, resulting in recommendation blind spots and low recall rates.

Method used

By constructing a grid space-time module set, combining text and visual data of multi-source UGC cultural and tourism materials, semantic recognition technology is used to calculate semantic drift and cross-modal semantic similarity, filter candidate hot spots and generate recommendation lists.

Benefits of technology

It has achieved timely and accurate identification of emerging tourism hotspots, improved the accuracy and scalability of the recommendation system, and helped the local cultural and tourism industry to accurately market and infrastructure optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235732B_ABST
    Figure CN120235732B_ABST
Patent Text Reader

Abstract

The present application discloses a method and system for analyzing cultural and tourism content based on semantic recognition, which relates to the technical field of digital analysis of culture and tourism. The method includes: obtaining a multi-source UGC cultural and tourism material set, and constructing a grid spatio-temporal domain module set according to a target area grid set and continuous historical time windows; allocating each multi-source UGC cultural and tourism material to the corresponding grid spatio-temporal domain module; extracting the grid text semantic center vector of the grid spatio-temporal domain module, and calculating the semantic drift amount of the target area grid, so as to screen at least one candidate hot area grid from each target area grid; calculating the corresponding grid visual semantic center vector, and calculating the cross-modal semantic similarity in combination with the second grid text semantic center vector of the candidate hot area grid, so as to screen the target hot area grid and generate a cultural and tourism destination recommendation list. Thus, it can quickly focus on geographical areas with active recent content and rising discussion heat, which helps to discover micro-destinations of emerging tourism hotspots.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cultural and tourism digital intelligence analysis, and particularly to a method and system for analyzing cultural and tourism content based on semantic recognition. Background Art

[0002] In recent years, with the growth of the "decentralized" travel needs of the young population, "in-depth travel", "off-the-beaten-path attractions", and "hidden scenic spots" have increasingly become new trends in the cultural and tourism market. In recent years, the growth rate of UGC (User-Generated Content) related to "non-popular scenic spots" has been rapid, especially the attention to "micro-destinations" - that is, those tourist spots that are remote, small in scale, and not yet promoted - is the most significant.

[0003] Traditional cultural and tourism recommendation systems mainly rely on static POI (Point-of-Interest) lists, supplemented by manually predefined "scenic spot levels" and "theme tags" (such as "historical culture" and "ecological leisure"), resulting in the inability to cover "off-the-beaten-path attractions" and "hidden beautiful scenery" emerging on social media, and it is difficult to discover emerging micro-destinations in a timely manner, leading to blind spots in cultural and tourism recommendations and being unfavorable to the development of the local cultural and tourism consumption industry. Summary of the Invention

[0004] This application provides a method, system, storage medium, computer program product, and electronic device for analyzing cultural and tourism content based on semantic recognition, so as to at least solve the problems in the current related technologies that rely on static point-of-interest lists and manually predefined tags, making it difficult to discover and cover "micro-destinations" in a timely and comprehensive manner, resulting in recommendation blind spots and low recall rates.

[0005] In a first aspect, an embodiment of the present application provides a method for analyzing cultural and tourism content based on semantic recognition, including: obtaining a multi-source UGC cultural and tourism material set, and constructing a grid spatio-temporal domain module set according to a preset target area grid set and consecutive first and second historical time windows; each multi-source UGC cultural and tourism material in the multi-source UGC cultural and tourism material set includes text UGC data and corresponding visual UGC data, which are associated through the same content ID, and record the corresponding timestamp and geographical coordinate information; each grid spatio-temporal domain module in the grid spatio-temporal domain module set is respectively used to indicate the target area grid under the corresponding historical time window; performing spatio-temporal grid matching on each of the multi-source UGC cultural and tourism materials according to the timestamp and geographical coordinate information, and respectively allocating them to the corresponding grid spatio-temporal domain modules; for each of the grid spatio-temporal domain modules, extracting the text semantic vectors corresponding to the text UGC data in the grid spatio-temporal domain module, and calculating the corresponding grid text semantic center vector; calculating the semantic drift amount of the target area grid according to the first grid text semantic center vector corresponding to the first historical time window and the second grid text semantic center vector corresponding to the second historical time window, so as to screen at least one candidate hot area grid from each target area grid; for each of the candidate hot area grids, sampling a preset number of key frames from each visual UGC data corresponding to the second historical time window of the candidate hot area grid, so as to calculate the corresponding grid visual semantic center vector, and calculating the cross-modal semantic similarity in combination with the second grid text semantic center vector of the candidate hot area grid; generating a cultural and tourism destination recommendation list according to at least one target hot area grid whose corresponding cross-modal semantic similarity exceeds a preset similarity threshold.

[0006] Second aspect, the embodiments of the present application provide a cultural and tourism content analysis system based on semantic recognition, including: a data acquisition unit, configured to acquire a multi-source UGC cultural and tourism material set, and construct a grid spatio-temporal domain module set according to a preset target area grid set and consecutive first and second historical time windows; each multi-source UGC cultural and tourism material in the multi-source UGC cultural and tourism material set includes text UGC data and corresponding visual UGC data, which are associated by the same content ID, and record corresponding timestamps and geographic coordinate information; each grid spatio-temporal domain module in the grid spatio-temporal domain module set is respectively used to indicate a target area grid under the corresponding historical time window; a material spatio-temporal matching unit, configured to perform spatio-temporal grid matching on each of the multi-source UGC cultural and tourism materials according to the timestamps and geographic coordinate information, so as to respectively allocate them to the corresponding grid spatio-temporal domain modules; a grid semantic calculation unit, configured to, for each of the grid spatio-temporal domain modules, extract the text semantic vectors corresponding to the text UGC data in the grid spatio-temporal domain module, and calculate the corresponding grid text semantic center vector; a hot grid screening unit, configured to calculate the semantic drift amount of the target area grid according to the first grid text semantic center vector corresponding to the first historical time window and the second grid text semantic center vector corresponding to the second historical time window, so as to screen at least one candidate hot area grid from each target area grid; a cross-modal alignment verification unit, configured to, for each of the candidate hot area grids, sample a preset number of key frames from each visual UGC data corresponding to the second historical time window of the candidate hot area grid, so as to calculate the corresponding grid visual semantic center vector, and calculate the cross-modal semantic similarity in combination with the second grid text semantic center vector of the candidate hot area grid; a micro-destination recommendation unit, configured to generate a cultural and tourism destination recommendation list according to at least one target hot area grid whose corresponding cross-modal semantic similarity exceeds a preset similarity threshold.

[0007] Third aspect, there is provided an electronic device, including: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the steps of the method for analyzing cultural and tourism content based on semantic recognition according to any embodiment of the present application.

[0008] Fourth aspect, the embodiments of the present application provide a storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the method for analyzing cultural and tourism content based on semantic recognition according to any embodiment of the present application are implemented.

[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the cultural and tourism content analysis method based on semantic recognition of any embodiment of the present application.

[0010] The cultural and tourism content analysis method and system based on semantic recognition provided by this application can produce at least the following technical effects:

[0011] (1) By constructing a gridded spatiotemporal domain module set and combining it with historical time window comparative analysis, we can capture the semantic evolution trend of UGC content in both time and space dimensions. Specifically, by comparing the text semantic center vectors of the same grid cell within two consecutive time windows, we can effectively identify areas with significant semantic drift, that is, areas where new gameplay or new scenes may have appeared. This allows us to quickly focus on geographical areas with recent content activity and rising discussion heat, thereby discovering emerging tourist hotspots that are not covered by traditional POI lists.

[0012] (2) After screening candidate hotspot grids through the semantic drift of text semantic vectors, keyframe visual semantic analysis is performed on such candidate hotspot grids, and cross-modal similarity is used as the criterion to avoid the interference of large-scale noise text (such as promotional advertisements or human-computer format reviews). The deep semantic alignment between the text UGC and image UGC of the grid unit is used to filter out the deviation caused by single modality analysis, ensuring that the identified hotspot areas have real tourism appeal and user interest support.

[0013] Through this technical solution, a cultural and tourism recommendation list based on actual user attention dynamics is generated, which can timely and accurately mine emerging and high-quality "micro-destinations" from the cultural and tourism social platform database, providing local governments or cultural and tourism suppliers with early identification tools for "potential hot areas", helping to formulate precise marketing strategies, optimize travel route planning and infrastructure layout, and thereby promote the brand exposure and commercial value release of niche attractions, empowering the development of the regional cultural and tourism consumption industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0015] Figure 1 A flowchart illustrating an example of a method for analyzing cultural and tourism content based on semantic recognition according to an embodiment of the present application is shown;

[0016] Figure 2 The figure shows an operation flowchart of an example for generating a cultural and tourism destination recommendation list according to an embodiment of the present application;

[0017] Figure 3 The figure shows an operation flowchart of an example for calculating the semantic drift amount according to an embodiment of the present application;

[0018] Figure 4 The figure shows an operation flowchart of an example for cross-modal semantic similarity calculation according to an embodiment of the present application;

[0019] Figure 5 The figure shows a structural block diagram of an example of a cultural and tourism content analysis system based on semantic recognition according to an embodiment of the present application;

[0020] Figure 6 The figure is a schematic structural diagram of an embodiment of an electronic device of the present application. Detailed implementation manners

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the protection scope of the present application.

[0022] It should be noted that currently, some tourists pursue the "sense of privacy" of avoiding peak tourist seasons, and some tourists seek novel experiences such as "recommended by internet influencers" and "adventure check-in". However, the information of these micro-destinations is mostly scattered in social media short reviews, niche travel forums, or short videos of influencers, and it is difficult to capture them in a timely manner through traditional crawlers or manual tagging.

[0023] Traditional cultural and tourism recommendation systems mainly rely on static POI lists, supplemented by manually predefined "scenic area grades" and "theme tags" (such as "historical culture" and "ecological leisure"). However, the number of POIs on mainstream OTA platforms mostly remains in the hundreds of thousands, while the number of "grassroots check-in points" on social platforms is conservatively estimated to exceed one million, resulting in a huge information blind spot.

[0024] In addition, in UGC, users' descriptions of the same location will increase significantly with high-value information such as festivals and events. For example, "XX ancient town" is mostly described as "beautiful lantern festival scenery" before the Spring Festival, and as "misty rain gallery" during the rainy season, and it is difficult to capture this semantic-level change through a method based solely on keyword statistics.

[0025] On the other hand, common spatiotemporal clustering methods only consider geographic coordinate density or the time window of publication, overlooking the "thematic triggers" inherent in the text—new ways of enjoying the same location during different holidays. For example, during the Sophora japonica Festival, reviews of a rural market featured frequent words like "sea of flowers," "market," and "special snacks," while these words rarely appear on weekdays. Therefore, without incorporating both semantics and time into the model, it would be difficult to identify this scene as an emerging micro-destination.

[0026] It should be understood that the purpose of the above description of the current related art is only to facilitate the public to better understand the inventive spirit and motivation of this application, and is not to be construed as limiting this application. In addition, the technical solutions described in the above-mentioned current related art are not prior art and may also be undisclosed technical solutions, such as solutions under research or in the laboratory stage.

[0027] In the technical solutions of this application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved shall comply with the provisions of relevant laws and regulations and shall not violate public order and good morals.

[0028] Figure 1 A flowchart of an example of a method for analyzing cultural and tourism content based on semantic recognition according to an embodiment of the present application is shown.

[0029] Regarding the execution subject of the method of the embodiment of the present application, it can be any controller or processor with computing or processing capabilities. Specifically, it can be implemented by the cultural tourism social UGC analysis and management platform. Through multi-source UGC fusion, spatiotemporal grid refinement management, semantic drift quantitative analysis and cross-modal verification, a multi-source, dynamic, semantic-driven recommendation system architecture is constructed, which solves the problem of insufficient recognition ability of the existing recommendation system for "niche check-in places", significantly improves the accuracy of hot spot discovery, the reliability of recommendation results and the scalability of the system, thereby providing strong technical support for the upgrading of local cultural tourism consumption and precise operation.

[0030] In some examples, it can be integrated into an electronic device or terminal through software, hardware, or a combination of software and hardware, and the type of terminal or electronic device can be diverse, such as a mobile phone, tablet computer, or desktop computer, etc.

[0031] like Figure 1 As shown, in step S110, a multi-source UGC cultural and tourism material set is obtained, and a grid spatiotemporal domain module set is constructed according to a preset target area grid set and a continuous first historical time window and a second historical time window.

[0032] Here, each piece of multi-source UGC cultural and tourism material in the multi-source UGC cultural and tourism material set contains text UGC data and corresponding visual UGC data, which are associated through the same content ID, and record the corresponding timestamp and geographical coordinate information. Each grid spatio-temporal domain module in the grid spatio-temporal domain module set is respectively used to indicate the grid of the target area in the corresponding historical time window.

[0033] In some embodiments, the cultural and tourism social UGC analysis and management platform can access multiple tourism social data sources through an authorization interface, so as to access the user-generated content publicly available from multiple sources, such as Xiaohongshu, Mafengwo, Douyin, etc., and then collect the multi-source UGC cultural and tourism materials that meet the conditions according to the information structure of "text UGC - visual UGC - metadata". Exemplarily, the text UGC data should include titles, text descriptions or comments, etc., the visual UGC data should contain pictures or video frames, and the metadata can include the published content ID, geographical coordinate information (such as GPS longitude and latitude), and the published timestamp. Furthermore, with the content ID as the primary key, the text and visual content form a complete UGC data entry through an associated manner.

[0034] Regarding the details of grid cell division, in some exemplary embodiments, the geographical area to be detected (such as a certain city or a certain tourism area) can be divided into grid cells of equal size (for example, 500m × 500m) according to longitude and latitude, generate grid IDs, and can also mask non-tourism areas such as river waters according to business requirements to reduce invalid grids.

[0035] As an additional or alternative embodiment, the target area grid set is a plurality of mutually independent target area grids generated by regular spatial division of the target area, and the division granularity of the target area grids is dynamically determined according to the characteristics of the urban functional areas of the target area. Therefore, a dynamic granularity division mechanism based on the characteristics of urban functional areas is introduced, and the system uses publicly available urban planning data (such as urban zoning maps, land use data, administrative divisions) to identify the functional areas of the target area. Exemplarily, fine-grained division can be adopted for areas with dense commercial / scenic spots to accurately capture changes in user behavior within a small range; coarse-grained division can be adopted for ecological / rural areas to reduce noise interference in low-frequency data areas. In addition, it can also be adaptively adjusted according to the spatial density distribution of real-time UGC, that is, automatically refine the grid in areas with dense and active UGC to enhance the local hotspot detection ability.

[0036] Regarding the description of the continuous first historical time window and the second historical time window, they can be preset time windows adjacent to the current time, such as "the previous 7 - 14 days" and "the most recent 7 days", etc., and the window time length can also be flexibly adjusted or defined according to business requirements.

[0037] As an attachable or replaceable implementation, the first historical time window and the second historical time window are consecutive time periods defined based on a dynamic sliding window mechanism. The end time of the first historical time window coincides with the start time of the second historical time window, and the end time of the second historical time window is the current analysis reference time point. The dynamic sliding window mechanism is configured to shift forward at a preset time interval, so that the first historical time window and the second historical time window are updated synchronously over time.

[0038] Exemplarily, the first historical time window and the second historical time window are defined by a double sliding window, where the first historical time window is , and the second historical time window is , is the time window length; is the analysis reference time point (such as the current system time) and is updated according to the sliding window step size. Thus, through the connection of the double windows, both the "recent hotspots" and the "underlying trends" can be reflected simultaneously, ensuring that the drift detection is both sensitive and robust to "new fires". In addition, the preset time interval is less than the window duration of the historical time window. For example, the preset value (such as 1 day) is the minimum moving step size, and the overall window slides forward by this interval each time, continuously rolling on a daily granularity to capture the trend changes of each day; the step size is less than the window length, taking into account both data continuity and real-time performance.

[0039] In step S120, according to the time stamps and geographical coordinate information, spatio-temporal grid matching is performed on each multi-source UGC cultural and tourism material to be respectively allocated to the corresponding grid spatio-temporal domain module.

[0040] In some implementations, for each UGC cultural and tourism material, according to its recorded geographical coordinates and time stamps, after discretization, the corresponding grid ID and time window are matched, so as to accurately map it to the corresponding grid spatio-temporal domain module. For example, a UGC released on March 15, 2025, located in grid A1, will be allocated to the T2 window module corresponding to grid A1 (for example, from March 14 to March 20), rather than the T1 window module corresponding to grid A1 (for example, from March 7 to March 13), so that each UGC is allocated to the corresponding grid spatio-temporal domain module. Thus, by adopting a geospatial index structure and a time sliding window mechanism, the matching efficiency is improved, supporting the efficient mapping and retrieval of large-scale content, and ensuring that the content is accurately allocated to its actual geographical location and time background.

[0041] In step S130, for each grid spatio-temporal domain module, the text semantic vectors corresponding to the text UGC data in the grid spatio-temporal domain module are extracted, and the corresponding grid text semantic center vector is calculated.

[0042] In some embodiments, for each grid spatio-temporal domain module, the system inputs all the text UGC data contained therein into a semantic understanding model (such as a BERT model or a RoBERTa model, etc.) to extract the semantic representation vectors of each piece of text. Subsequently, within the same grid module, the average value or weighted combination of all text semantic vectors is calculated as the text semantic center vector of this grid module. Thus, based on the vectorization of the deep semantic model, not only can keywords be identified, but also context associations can be captured, realizing the conversion of unstructured text information into a unified vector representation, enabling the system to quantify and compare the content expressions in different spatio-temporal regions.

[0043] In some examples of the embodiments of the present application, based on the BERT model, the text semantic vectors corresponding to each text UGC data in the grid spatio-temporal domain module are extracted.

[0044] Then, the arithmetic mean of all text semantic vectors corresponding to the grid spatio-temporal domain module is taken and L2 norm normalization is performed to obtain the corresponding grid text semantic center vector:

[0045] , Equation (1)

[0046] In the formula, represents the grid text semantic center vector of grid , represents the number of text UGCs within grid , and respectively represent the th text semantic vector and the th text semantic vector, represents the L2 norm operator.

[0047] It should be noted that all the text UGCs within the grid spatio-temporal domain module are generated within the same geographical grid and the same time window. Equal-weight averaging of their corresponding vectors is essentially a merger within the "same semantic entity" and can represent the mainstream tendency of user discussions in this area during this period. Through arithmetic averaging, the abnormal vectors generated by a few extreme comments or robot spamming can be diluted, enabling the center vector to more reflect the true expressions of most users rather than being skewed by a few outliers. Through the dual processing of averaging and normalization, the interference caused by a small number of mislabels, spam information, or spamming can be effectively suppressed, making the center vector more discriminative for "true hotspots". In addition, the arithmetic averaging and L2 normalization have extremely small computational amounts, support real-time updates of thousands of grids simultaneously, and can meet the deployment requirements of large-scale regions.

[0048] In step S140, the semantic drift amount of the target area grid is calculated based on the first grid text semantic center vector corresponding to the first historical time window and the second grid text semantic center vector corresponding to the second historical time window, so as to screen at least one candidate hotspot area grid from each target area grid.

[0049] In some embodiments, for each target area grid, the degree of change between its text semantic center vector in the first time window (T1) and the second time window (T2) is compared to calculate its semantic drift, which represents the degree of significant change in the area's user-focused topics. For example, if an area described as a "quiet valley" during T1 frequently appears with new semantics such as "cliff swing" and "high-altitude glass plank road" during T2, it indicates that the area may have become a "check-in" or "hotspot" due to factors such as new project construction and social media dissemination, and has the potential to become a "micro-destination."

[0050] More specifically, in one example of the present application, a threshold can be pre-set and compared with the semantic drift amount to screen out grids with significant drift as candidate hotspot areas. In another example of the present application, the drift amount can be sorted from large to small, and the top-ranked grids are selected as candidate hotspot grids.

[0051] As a further optimization implementation method, the grid range of each candidate hotspot grid can also be compared with the POI list to identify whether the candidate hotspot grid overlaps with a certain POI range. If so, it can be filtered out to avoid the interference of existing popular POIs in the discovery process of niche and popular micro-destinations.

[0052] In some examples of the embodiments of the present application, the semantic drift amounts of all target area grids are aggregated to generate a global distribution of semantic drift, thereby forming a cross-regional semantic change distribution map to reflect the breadth and intensity of changes in user interest topics in the entire target area within a specific period of time, and the dynamic threshold corresponding to the preset statistical quantile is calculated based on the global distribution of semantic drift. For example, the semantic drift dynamic threshold is set according to the distribution characteristics of the drift amount data and a specific statistical quantile (such as 75%, 85%, 90%, etc.). Furthermore, at least one target area grid whose corresponding semantic drift amount exceeds the dynamic threshold is marked as a candidate hotspot area grid. Thus, combined with the semantic drift dynamic threshold of the statistical quantile, the number of hotspot areas can be dynamically adjusted to avoid "flooding micro-hotspots" when the content is dense, and to avoid "amplifying virtual hotspots" during cold zone periods, thereby achieving a balance between the stability and sensitivity of hotspot identification.

[0053] In step S150, for each candidate hotspot area grid, a preset number of key frames are sampled from each visual UGC data of the second historical time window corresponding to the candidate hotspot area grid to calculate the corresponding grid visual semantic center vector, and the cross-modal semantic similarity is calculated in combination with the second grid text semantic center vector of the candidate hotspot area grid.

[0054] In some embodiments, for each candidate hotspot area grid, a certain amount of visual UGC (e.g., pictures or video clips) is extracted within a second historical time window closer to the current time, representative key frames are sampled from them, and converted into visual semantic vectors using an image semantic model (e.g., CLIP, BLIP, etc.). These image semantic vectors are then aggregated to form a visual semantic center vector. In addition, it should be noted that when performing cross-modal calculations, only all textual content and visual content falling within the grid within the second time window (i.e., the window closer to the current time) is taken to ensure that the current center vector only maps the latest topic and enables comparison of the same spatiotemporal dimension, thereby ensuring the accuracy and robustness of hotspot detection.

[0055] Subsequently, the platform semantically matches the visual semantic center vector with the text semantic center vector of the grid and evaluates the cross-modal semantic similarity between the two. If the similarity is high, it indicates that the graphic and text content in the area is highly consistent in expression and has a real user cognition and dissemination basis. Therefore, by introducing a cross-modal similarity graphic and text semantic alignment mechanism, false hotspots that are "hyped with text" but lack image evidence are effectively eliminated, ensuring that the remaining hotspot grids are not only "everyone is saying that it has changed", but also "everyone is taking pictures of new ways of playing or new scenes", effectively improving the authenticity and credibility of the recommended new cultural and tourism hotspots.

[0056] In step S160 , a recommended list of cultural and tourism destinations is generated based on at least one target hotspot area grid whose corresponding cross-modal semantic similarity exceeds a preset similarity threshold.

[0057] In some embodiments, the platform can generate hot spot area names or geographic tags based on UGC frequency or hot spot area grids, and can add recommendation reasons (such as "discussion popularity has increased by 200% in the past 7 days"). It can also select representative text and image content (for example, UGC text and image content with high likes in the grid area), and generate a recommendation list based on the above content.

[0058] Regarding the business application of the recommended list of cultural and tourism destinations, on the one hand, it can be pushed by the platform according to the classification of different user interest tags; on the other hand, it can also be used by cultural and tourism management departments to promote and develop potential scenic spots, and there should be no restrictions here.

[0059] Thus, a closed-loop transformation from content-driven to behavior recommendation is achieved, truly discovering potential micro-destinations with development potential from "user-generated content", forming a personalized, highly real-time, and semantically authentic recommendation mechanism, which helps improve the response efficiency of the cultural and tourism system to emerging tourism behaviors and promotes the transformation of the cultural and tourism industry from "concentrated scenic spot check-ins" to "distributed micro-destination discovery".

[0060] Figure 2 The operation flowchart of an example for generating a cultural and tourism destination recommendation list according to an embodiment of the present application is shown.

[0061] As Figure 2 shown, in step S210, for each target hot-spot area grid, the second grid text semantic center vector of the target hot-spot area grid is concatenated with the corresponding grid visual semantic center vector to obtain the corresponding grid multi-modal feature, and the attention weights between the grid multi-modal feature and each preset cultural and tourism feature label word vector are calculated, and at least one matching target cultural and tourism feature label is obtained through sorting.

[0062] In some embodiments, two types of center vectors of the adjacent time window of the target hot-spot grid that have undergone drift detection and cross-modal verification are read from the cache: the text semantic center (such as 768 dimensions) and the visual semantic center (such as 512 dimensions), and the two can be projected into the same dimensional space (such as 256 dimensions) respectively through independent lightweight projection networks (each consisting of one layer of fully connected + LayerNorm), and L2 normalization is performed to eliminate the dimensional and dimension differences.

[0063] Then, the normalized text vector and visual vector are concatenated into a multi-modal feature vector, and the text and visual signals are further mixed through one layer of residual fusion layer (such as, fully connected + ReLU + skip connection), so that the finally output multi-modal feature contains both "topic intention" and "scene perception" information.

[0064] In addition, a label vocabulary can also be maintained in the platform, in which various common cultural and tourism feature labels can be recorded, such as "night tour light show", "deep tour of ancient town", "rural idyllic experience", "food check-in", "parent-child interaction", or "outdoor adventure", etc. Furthermore, for each label phrase, the same BERT encoding and projection network as the text center are used to generate a label vector with the same dimension as the grid multi-modal feature, and new labels can be supplemented to the label vocabulary through an interface.

[0065] Furthermore, the dot product is performed on the grid multimodal vectors and all label vectors to obtain their respective matching scores. By performing softmax normalization on the scores, an attention weight distribution is obtained, which can reflect the strength of the correlation between the grid and each label. Sorting the attention weights from high to low, a preset number of labels with the highest rankings can be selected as the recommended characteristic labels for this hotspot, and the attention mechanism makes the tagging results highly interpretable.

[0066] In some examples of the embodiments of the present application, the grid multimodal features and the label phrase vectors are mapped to the same vector space to ensure the comparability of similarity calculations. Furthermore, through the attention mechanism of dot product + softmax, the correlation between each label and the grid features is automatically measured, and the higher the weight, the stronger the matching degree.

[0067] Specifically, the attention weights are calculated by the following formula:

[0068] , Equation (2)

[0069] In the formula, represents the attention weight of the th cultural and tourism characteristic label, represents the transpose of the grid multimodal features, represents the th label embedding vector of the cultural and tourism characteristic label; represents a preset temperature coefficient, which is a positive real number less than 1 and is used to adjust the smoothness of the attention distribution; represents the total number of cultural and tourism characteristic labels, represents the th label embedding vector of the cultural and tourism characteristic label, and the denominator represents the sum of the amplified matching scores of all labels to achieve softmax normalization.

[0070] In Equation (2), by introducing the temperature coefficient in the attention mechanism, the smoothness of the attention distribution is adjusted. Specifically, when is small (such as 0.05), the output difference of the exponential operation is amplified, the attention distribution is sharper, and the model only biases towards the few most relevant labels; when is large (such as 0.2), the distribution is smoother and more labels can be taken into account. Thus, can be adjusted as needed to achieve a balance in various scenarios from "highlighted recommendation" to "diverse display" and meet the personalized tagging requirements.

[0071] In step S220, according to each target hotspot area grid and the corresponding target cultural and tourism characteristic labels, a cultural and tourism destination recommendation list is generated.

[0072] Through the embodiments of the present application, each hotspot grid that has been cross-modally verified is automatically labeled with a cultural and tourism label that best represents the characteristics of the region, replacing traditional multi-step entity extraction and graph labeling, and achieving lightweight, explainable, and easily scalable labeled recommendations.

[0073] In some implementations, the specific location name can be reverse-checked based on the center coordinates of the target hotspot grid to serve as the corresponding geographic location information. In addition, a high-weight UGC text summary (such as a refined evaluation) and a representative keyframe within the grid can be selected as the text and visual material for the recommendation card. A list of recommended cultural and tourism destination cards is constructed using the above materials, and the cultural and tourism feature tags selected by the corresponding grid, such as "night tour light show", "food check-in", "parent-child interaction", etc., are prominently displayed in a prominent position on each card. In this way, an efficient conversion from massive UGC to refined cultural and tourism recommendation cards is achieved, which not only ensures the timeliness and high quality of the recommended content, but also has good interpretability and scalability.

[0074] Figure 3 An operational flowchart of an example of calculating semantic drift according to an embodiment of the present application is shown.

[0075] like Figure 3 As shown, in step S310, the semantic similarity index between the first grid text semantic center vector and the second grid text semantic center vector is calculated.

[0076] Here, the cosine similarity is used to measure the angular deviation between the overall topics of the "early" and "recent" texts in the high-dimensional semantic space, which directly reflects the migration of the focus of user discussion.

[0077] , Formula (3)

[0078] Where, represents the target area grid, Indicates that it corresponds to The first grid text semantic center vector is the average semantic center vector of all text UGC in the first historical time window; Indicates that it corresponds to The second grid text semantic center vector is the average semantic center vector of all text UGC in the second historical time window; represents the vector transpose operation, represents the L2 norm operator, Represents the semantic similarity index, which is the cosine similarity of the text semantic center vector.

[0079] Specifically, for the grid The semantic center of the early text and the recent Center for Textual Semantics Directly perform the inner product and divide by the product of the norms. Through normalization, the output value is stabilized between -1 and 1, and the risk of numerical overflow during operation is low. Thus, it can highly summarize the mainstream expression changes of all comments within the grid and filter out some extreme noises, which helps in the preliminary hot spot screening.

[0080] In step S320, based on the LDA topic model, process the first grid text semantic center vector and the second grid text semantic center vector to generate the topic distribution difference index corresponding to consecutive historical time windows.

[0081] It should be noted that semantic similarity can only reflect the alignment of the overall vectors, while the topic distribution difference can capture the structural changes of key topics (such as "night tour" and "food") in the high-dimensional space. The two complement each other. Therefore, the LDA (Latent Dirichlet Allocation) topic model is introduced to extract the topic probability distribution under each time window and use the KL divergence to quantify the distribution difference, so as to capture the structural topic evolution from the micro level.

[0082] More specifically, the LDA model can be trained offline on a large-scale cultural and tourism corpus (such as comments and travelogues). The number of topics is generally set to 20 - 30, covering various play styles (food, night tour, cultural heritage, outdoor adventure, etc.), and the topic model can also be fine-tuned regularly to ensure that emerging play styles (such as "immersive script killing" and "niche homestay experience") can be recognized in a timely manner.

[0083] , Equation (4)

[0084] In the formula, represents the total number of topics set by the LDA topic model, represents the th topic's probability distribution in the first historical time window of the grid , represents the th topic's probability distribution in the second historical time window of the grid ; represents the KL divergence, which is used to quantify the degree of change in the topic structure within the consecutive historical time windows of the grid .

[0085] Here, for all text sets of each grid in the "early stage" and "recent stage", the topic probability vector distributions are respectively obtained by using model inference, and the two are compared element by element to complete the divergence accumulation. The larger the value, the more obvious the change in the topic structure. In this way, not only can we see the change in the "expression angle" of the text semantics, but also identify the migration trends of each sub-topic, such as the increase in the proportion of "food" and the decrease in the proportion of "night view".

[0086] In step S330, the semantic similarity index and the topic distribution difference index are fused to generate a corresponding semantic drift amount.

[0087] , Equation (5)

[0088] In the formula, represents the semantic drift amount of the grid The larger the value, the more significant the new topics that have emerged in this grid recently; and represent the text similarity weight coefficient and the topic divergence weight coefficient respectively.

[0089] Regarding the description of Equation (5), first convert the similarity to a distance and then fuse it to ensure the additivity of the two indicators within the same numerical magnitude; through linear weighting, the two types of metrics of "overall semantic distance" and "topic structure change" are combined into a single drift amount.

[0090] In the embodiments of the present application, the cosine distance is used to quickly detect the overall expression shift, and the KL divergence is used to refine and capture the key topic structure changes. After fusing the two, it can cover two types of semantic drift features at the same time, which is more robust than a single indicator. In addition, the weights and can be determined based on experience or A / B testing, can adapt to different scenarios or hot topic types, and the calculation process of the semantic drift amount is explicit, which helps the operation team to review the core reasons for the grid drift to be on the list. Thus, through the dual-index fusion of cosine similarity + KL divergence, a panoramic monitoring of "semantic expression change" and "topic structure change" is achieved, greatly improving the accuracy of discovering new micro-destinations.

[0091] Figure 4 FIG. shows an operation flowchart of an example of cross-modal semantic similarity calculation according to an embodiment of the present application.

[0092] It should be noted that visual UGC (short video key frames, pictures) provides an objective "true scene mapping" for the cultural and tourism scene, making up for possible subjective omissions or noises in text descriptions. To integrate visual signals into the same semantic space, it is necessary to project the high-dimensional visual features into the same vector dimension as the text center through projection and perform normalization processing to achieve true cross-modal comparability.

[0093] As Figure 4 shown, in step S410, each key frame is input into a visual encoder to obtain a corresponding original visual feature vector.

[0094] Specifically, by locating all the key frames within the second historical time window that fall within the grid For short videos and images, from each short video or multiple images, a fixed number (e.g., 5 frames) of representative frames are extracted according to the scene switching and diversity strategy. Each frame image is sent into a visual encoder (such as ResNet, VGGNet, etc.) after unified preprocessing (scaling, normalization) to extract high-dimensional visual features.

[0095] , Equation (6)

[0096] In the formula, represents the key frame of the th frame corresponding to the second historical time window, is the key frame index, represents the total number of key frames of the preset sampling; represents the visual encoder function, represents the th frame corresponding original visual feature vector.

[0097] In Equation (6), a pre-trained Vision Transformer (ViT) model is adopted, which can capture high-level features such as scene structure, color, and object presence; the dimension can be set to 512.

[0098] In step S420, the original visual feature vector is mapped to a fusion space with the same dimension as the text center through a linear projection layer, and L2 normalization processing is performed to obtain the corresponding frame projection transformation vector.

[0099] , Equation (7)

[0100] In the formula, and respectively represent the visual projection weight matrix and the visual projection bias vector, represents the frame projection transformation vector corresponding to the th frame key frame after projection and L2 normalization.

[0101] In some embodiments, the projection network configuration can adopt a single-layer fully connected network (256×512 matrix) and bias to map the 512-dimensional original visual features to a 256-dimensional fusion space, and then immediately perform L2 normalization on the vector after the affine transformation to ensure that the output vector norm is 1, eliminate the scale difference, and provide a fair comparison benchmark for the inner product with the text center. Thus, the high-dimensional visual semantics are seamlessly mapped to the text space, realizing a true same-dimension representation of text and images, and the intensity difference of the visual signal itself is removed through the normalization operation, making the cosine metric only focus on the direction consistency.

[0102] In step S430, arithmetic mean aggregation is performed on all frame projection transformation vectors of the grid, and L2 normalization is performed again to obtain the grid visual semantic center vector.

[0103] It should be noted that single-frame visual features are often fragmented and affected by the shooting angle and time node. By averaging the multi-frame visual projection vectors sampled from the same grid and performing normalization, the "visual average scene" can be refined, denoising while highlighting the most representative visual information of the grid recently.

[0104] , Equation (8)

[0105] , Equation (9)

[0106] In the formula, represents the arithmetic mean aggregation vector of the grid , represents the grid of the grid visual semantic center vector.

[0107] Here, within the second historical time window, frames are extracted from each video according to "scene change detection". If the number of frames is too large, a fixed N frames (such as N = 5) are selected in a time-uniform or clustering center manner, and the normalized vectors after affine projection of each frame are aggregated into , and is obtained through arithmetic mean, and then L2 normalization is performed again to output . Thus, by cross-averaging the semantic vectors of different video frames, the deviation of a single perspective or accidental picture is reduced, the influence of different frame qualities, highlights / shadows is balanced, and the stability of the overall scene representation is improved.

[0108] In step S440, the cross-modal semantic similarity between the grid visual semantic center vector and the second grid text semantic center vector within the same period window is calculated in an inner product manner.

[0109] , Equation (10)

[0110] In the formula, represents the cross-modal semantic similarity of the grid .

[0111] Here, the cosine similarity is directly used to measure the directional consistency between the "text center vector" and the "visual semantic center vector" in the fusion space within the recent window. Only when both point to similar high-dimensional positions does it indicate that there have been new gameplay / topic bursts and corresponding real scenarios in this grid recently, and finally it is retained as a real new cultural and tourism hotspot. Thus, it can effectively reduce the misjudgment caused by "text spamming" or "visual dominance", and significantly reduce the false alarm rate. In addition, by quantifying the similarity value, the operation can adjust the similarity threshold and observe the score distribution to obtain different recommendation results, which is beneficial for the operation to conduct multi-level analysis of new cultural and tourism hotspots.

[0112] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of combined actions. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, each embodiment is described with emphasis. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0113] Figure 5 The structural block diagram of an example of a cultural and tourism content analysis system based on semantic recognition according to an embodiment of the present application is shown.

[0114] As Figure 5 shown, the cultural and tourism content analysis system 500 based on semantic recognition includes a data acquisition unit 510, a material spatio-temporal matching unit 520, a grid semantic calculation unit 530, a hot grid screening unit 540, a cross-modal alignment verification unit 550, and a micro-destination recommendation unit 560.

[0115] The data acquisition unit 510 is used to acquire a multi-source UGC cultural and tourism material set, and construct a grid spatio-temporal domain module set according to a preset target area grid set and continuous first and second historical time windows; each multi-source UGC cultural and tourism material in the multi-source UGC cultural and tourism material set includes text UGC data and corresponding visual UGC data, which are associated by the same content ID, and record the corresponding timestamp and geographical coordinate information; each grid spatio-temporal domain module in the grid spatio-temporal domain module set is respectively used to indicate the target area grid under the corresponding historical time window.

[0116] The material spatio-temporal matching unit 520 is used to perform spatio-temporal grid matching on each of the multi-source UGC cultural and tourism materials according to the timestamp and geographical coordinate information, so as to be respectively assigned to the corresponding grid spatio-temporal domain module.

[0117] The grid semantic computing unit 530 is used to extract the text semantic vectors corresponding to each text UGC data in each of the grid spatio-temporal domain modules, and calculate the corresponding grid text semantic center vectors.

[0118] The hot grid screening unit 540 is used to calculate the semantic drift amount of the target area grid according to the first grid text semantic center vector corresponding to the first historical time window and the second grid text semantic center vector corresponding to the second historical time window, so as to screen at least one candidate hot area grid from each target area grid.

[0119] The cross-modal alignment verification unit 550 is used to sample a preset number of key frames from each visual UGC data corresponding to the second historical time window of each of the candidate hot area grids, so as to calculate the corresponding grid visual semantic center vectors, and calculate the cross-modal semantic similarity in combination with the second grid text semantic center vector of the candidate hot area grid.

[0120] The micro-destination recommendation unit 560 is used to generate a cultural and tourism destination recommendation list according to at least one target hot area grid whose corresponding cross-modal semantic similarity exceeds a preset similarity threshold.

[0121] In some embodiments, the embodiments of the present application provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to be used to execute the steps of any one of the above-mentioned semantic recognition-based cultural and tourism content analysis methods of the present application.

[0122] In some embodiments, the embodiments of the present application further provide a computer program product, the computer program product includes a computer program stored on a non-volatile computer-readable storage medium, the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is enabled to execute the steps of any one of the above-mentioned semantic recognition-based cultural and tourism content analysis methods.

[0123] In some embodiments, the embodiments of the present application further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the steps of the semantic recognition-based cultural and tourism content analysis method.

[0124] Figure 6 It is a schematic hardware structure diagram of an electronic device for executing the semantic recognition-based cultural and tourism content analysis method provided by another embodiment of the present application, asFigure 6 As shown, the device includes:

[0125] One or more processors 610 and a memory 620, Figure 6 Taking one processor 610 as an example.

[0126] The device for executing the cultural and tourism content analysis method based on semantic recognition may further include: an input device 630 and an output device 640.

[0127] The processor 610, the memory 620, the input device 630, and the output device 640 may be connected by a bus or other means, Figure 6 Taking connection by bus as an example.

[0128] The memory 620, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the cultural and tourism content analysis method based on semantic recognition in the embodiments of the present application. The processor 610 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 620, that is, implements the cultural and tourism content analysis method based on semantic recognition in the above method embodiments.

[0129] The memory 620 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 620 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 620 may optionally include a memory remotely set relative to the processor 610, and these remote memories can be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and their combinations.

[0130] The input device 630 can receive input digital or character information, and generate signals related to the user settings and function control of the electronic device. The output device 640 may include a display device such as a display screen.

[0131] The one or more modules are stored in the memory 620, and when executed by the one or more processors 610, execute the cultural and tourism content analysis method based on semantic recognition in any of the above method embodiments.

[0132] The above-mentioned product can execute the method provided by the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference may be made to the method provided by the embodiments of the present application.

[0133] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to:

[0134] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0135] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDA, MID, and UMPC devices, etc.

[0136] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and intelligent toys and portable vehicle navigation devices.

[0137] (4) Other airborne electronic devices with data interaction functions, such as in-vehicle device installed on a vehicle.

[0138] The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0139] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A method for analyzing cultural and tourism content based on semantic recognition, characterized in that, The method includes: Obtain a multi-source UGC cultural and tourism material set, and construct a grid spatio-temporal domain module set according to a preset target area grid set and continuous first and second historical time windows; each multi-source UGC cultural and tourism material in the multi-source UGC cultural and tourism material set includes text UGC data and corresponding visual UGC data, which are associated by the same content ID, and record the corresponding timestamp and geographical coordinate information; each grid spatio-temporal domain module in the grid spatio-temporal domain module set is respectively used to indicate the target area grid under the corresponding historical time window; Perform spatio-temporal grid matching on each of the multi-source UGC cultural and tourism materials according to the timestamp and geographical coordinate information, so as to be respectively assigned to the corresponding grid spatio-temporal domain module; For each of the grid spatio-temporal domain modules, extract the text semantic vectors corresponding to the text UGC data in the grid spatio-temporal domain module, and calculate the corresponding grid text semantic center vector; Calculate the semantic drift amount of the target area grid according to the first grid text semantic center vector corresponding to the first historical time window and the second grid text semantic center vector corresponding to the second historical time window, so as to screen at least one candidate hot area grid from each target area grid; For each of the candidate hot area grids, sample a preset number of key frames from each visual UGC data corresponding to the second historical time window of the candidate hot area grid, so as to calculate the corresponding grid visual semantic center vector, and calculate the cross-modal semantic similarity in combination with the second grid text semantic center vector of the candidate hot area grid; Generate a cultural and tourism destination recommendation list according to at least one target hot area grid whose corresponding cross-modal semantic similarity exceeds a preset similarity threshold; 2. The method according to claim 1, wherein The target area grid set is a plurality of mutually independent target area grids generated by regular spatial division of the target area; the division granularity of the target area grid is dynamically determined according to the urban functional area characteristics of the target area; The first historical time window and the second historical time window are continuous time periods defined based on a dynamic sliding window mechanism, the end time of the first historical time window coincides with the start time of the second historical time window, and the end time of the second historical time window is the current analysis benchmark time point; The dynamic sliding window mechanism is configured to move forward at a preset time interval, so that the first historical time window and the second historical time window are updated synchronously with time; the preset time interval is less than the window duration of the historical time window; 3. The method according to claim 1, wherein The generating a cultural and tourism destination recommendation list according to at least one target hot area grid whose corresponding cross-modal semantic similarity exceeds a preset similarity threshold includes: For each of the target hot area grids, splice the second grid text semantic center vector of the target hot area grid and the corresponding grid visual semantic center vector to obtain the corresponding grid multi-modal feature, and calculate the attention weights between the grid multi-modal feature and each preset cultural and tourism characteristic label word vector, and obtain at least one matching target cultural and tourism characteristic label through sorting; Generate a list of recommended cultural and tourism destinations based on each target hot spot area grid and corresponding target cultural and tourism characteristic tags.

4. The method according to claim 3, characterized in that Calculating the attention weights between the grid multimodal features and the word vectors of each preset cultural and tourism characteristic tag includes: , In the formula, represents the attention weight of the th cultural and tourism feature label, represents the transpose of the grid multi-modal feature, represents the th label embedding vector of the cultural and tourism feature label; represents a preset temperature coefficient, which is a positive real number less than 1 and is used to adjust the smoothness of the attention distribution; represents the total number of cultural and tourism feature labels, represents the th label embedding vector of the cultural and tourism feature label, and the denominator represents the sum of the amplified matching scores of all labels to achieve softmax normalization.

5. The method according to claim 1 or 4, characterized in that, The semantic drift amount is calculated by the following method: Calculate the semantic similarity index between the first grid text semantic center vector and the second grid text semantic center vector: , In the formula, represents the grid of the target area, represents corresponding to the first grid text semantic center vector, which is the average semantic center vector of all text UGCs within the first historical time window; represents corresponding to the second grid text semantic center vector, which is the average semantic center vector of all text UGCs within the second historical time window; represents the vector transpose operation, represents the L2 norm operator, represents the semantic similarity index, which is the cosine similarity of the text semantic center vector; Based on the LDA topic model, process the first grid text semantic center vector and the second grid text semantic center vector to generate a topic distribution difference index corresponding to a continuous historical time window: , In the formula, represents the total number of topics set by the LDA topic model, represents the -th topic's probability distribution in the first historical time window of the grid ; represents the -th topic's probability distribution in the second historical time window of the grid ; represents the KL divergence, which is used to quantify the degree of change in the topic structure within consecutive historical time windows of the grid . Fuse the semantic similarity index and the topic distribution difference index to generate a corresponding semantic drift amount: , In the formula, represents the semantic drift amount of the grid . The larger the value is, the more significant the situation of new topics emerging in this grid recently is; and represent the text similarity weight coefficient and the topic divergence weight coefficient respectively.

6. The method according to claim 5, wherein Screening at least one candidate hot spot area grid from each target area grid includes: Summarize the semantic drift amounts of all target area grids to generate a global semantic drift distribution, and calculate a dynamic threshold corresponding to a preset statistical quantile based on the global semantic drift distribution; Mark at least one target area grid with a corresponding semantic drift amount exceeding the dynamic threshold as a candidate hot spot area grid.

7. The method according to claim 6, characterized in that, Calculating the corresponding grid visual semantic center vector and calculating the cross-modal semantic similarity in combination with the second grid text semantic center vector of the candidate hot spot area grid includes: Input each key frame into a visual encoder to obtain a corresponding original visual feature vector: , In the formula, represents the key frame of the th frame corresponding to the second historical time window, is the key frame index, represents the total number of key frames for preset sampling; represents the visual encoder function, represents the original visual feature vector corresponding to the Map the original visual feature vector to a fusion space with the same dimension as the text center through a linear projection layer and perform L2 normalization to obtain a corresponding frame projection transformation vector: , In the formula, and respectively represent the visual projection weight matrix and the visual projection bias vector, represents the frame projection transformation vector corresponding to the th key frame after projection and L2 normalization; Arithmetically average and aggregate all frame projection transformation vectors of the grid, and perform L2 normalization again to obtain the grid visual semantic center vector: , , In the formula, represents the arithmetic mean aggregation vector of the grid , and represents the grid visual semantic center vector of the grid . Calculate the cross-modal semantic similarity between the grid visual semantic center vector and the second grid text semantic center vector within the same period window by means of inner product: , In the formula, represents the cross-modal semantic similarity of the grid .

8. The method according to claim 7, wherein Extracting the text semantic vectors corresponding to each text UGC data in the grid spatio-temporal domain module includes: Based on the BERT model, extract the text semantic vectors corresponding to each text UGC data in the grid spatio-temporal domain module; Calculating the corresponding grid text semantic center vector includes: Take the arithmetic mean of all text semantic vectors corresponding to the grid spatio-temporal domain module and perform L2 norm normalization to obtain the corresponding grid text semantic center vector: , In the formula, represents the grid text semantic center vector of the grid, represents the grid the number of text UGCs within the grid, and respectively represent the th text semantic vector and the th text semantic vector.

9. A cultural and tourism content analysis system based on semantic recognition, characterized in that, The system includes: A data acquisition unit for acquiring a multi-source UGC cultural and tourism material set and constructing a grid spatio-temporal domain module set according to a preset target area grid set and continuous first and second historical time windows; each multi-source UGC cultural and tourism material in the multi-source UGC cultural and tourism material set includes text UGC data and corresponding visual UGC data, which are associated by the same content ID and record the corresponding timestamp and geographical coordinate information; each grid spatio-temporal domain module in the grid spatio-temporal domain module set is respectively used to indicate the target area grid under the corresponding historical time window; A material spatio-temporal matching unit, which is used to perform spatio-temporal grid matching on each of the multi-source UGC cultural and tourism materials according to the time stamp and geographical coordinate information, so as to be respectively assigned to the corresponding grid spatio-temporal domain module; A grid semantic calculation unit, which is used to extract the text semantic vectors corresponding to the text UGC data in each grid spatio-temporal domain module for each of the grid spatio-temporal domain modules, and calculate the corresponding grid text semantic center vector; A hot spot grid screening unit, which is used to calculate the semantic drift amount of the target area grid according to the first grid text semantic center vector corresponding to the first historical time window and the second grid text semantic center vector corresponding to the second historical time window, so as to screen at least one candidate hot spot area grid from each target area grid; A cross-modal alignment verification unit, which is used to sample a preset number of key frames from each visual UGC data corresponding to the second historical time window of each candidate hot spot area grid for each of the candidate hot spot area grids, so as to calculate the corresponding grid visual semantic center vector, and calculate the cross-modal semantic similarity in combination with the second grid text semantic center vector of the candidate hot spot area grid; A micro-destination recommendation unit, which is used to generate a cultural and tourism destination recommendation list according to at least one target hot spot area grid whose corresponding cross-modal semantic similarity exceeds a preset similarity threshold.

Citation Information

Patent Citations

  • Deep semantic understanding method based on cross-modal model

    CN116680578A

  • Speech Recognition Accuracy with Natural-Language Understanding based Meta-Speech Systems for Assistant Systems

    US20210118440A1