File travel content analysis method and system based on semantic recognition

Through the cultural and tourism content analysis method based on semantic recognition, the grid time and space domain module is constructed and the semantic drift and cross-modal semantic similarity are calculated, which solves the problem that traditional recommendation systems are difficult to discover ‘micro destinations’, and achieves efficient and accurate recommendations of cultural and tourism destinations.

CN120235732AActive Publication Date: 2025-07-01CHONGQING TOURISM CLOUD INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510703104.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Traditional cultural and tourism recommendation systems rely on static lists of interest points and manual predefined tags, making it difficult to discover and cover ‘micro destinations’ in a timely and comprehensive manner, resulting in blind spots in recommendations and low recall rates.

Method used

Using a cultural and tourism content analysis method based on semantic recognition, a multi-source UGC cultural and tourism material is obtained, a grid time and space module set is constructed, and text semantic vectors and visual semantic center vectors are extracted, semantic drift and cross-modal semantic similarity are calculated, and a recommendation list of cultural and tourism destinations is generated.

Benefits of technology

It has achieved the semantic evolution trend that captures the time and space dimensions of UGC content, quickly focused on geographical areas where content is active and discussion has increased recently, and discovered emerging tourist hotspots that cannot be covered by traditional POI lists, improving the accuracy and reliability of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235732A_ABST
    Figure CN120235732A_ABST
Patent Text Reader

Abstract

The invention discloses a text travel content analysis method and system based on semantic recognition, and relates to the technical field of text travel data intelligent analysis, and the method comprises the steps: obtaining a multi-source UGC text travel material set, and constructing a grid time-space domain module set according to a target region grid set and a continuous historical time window; each multi-source UGC text travel material is distributed to a corresponding grid time-space domain module; extracting a grid text semantic center vector of the grid time-space domain module, and calculating a semantic drift amount of the target area grids so as to screen at least one candidate hot spot area grid from each target area grid; and calculating a corresponding grid visual semantic center vector, and calculating a cross-modal semantic similarity in combination with the second grid text semantic center vector of the candidate hot spot region grid to screen a target hot spot region grid, and generating a travel destination recommendation list. Therefore, the method can quickly focus on geographic areas with active recent contents and increased discussion popularity, and is beneficial to exploring micro-destinations of newly-emerging tourism hotspots.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cultural and tourism digital intelligence analysis, and particularly to a method and system for analyzing cultural and tourism content based on semantic recognition. Background Art

[0002] In recent years, with the growth of the "decentralized" travel needs of the young population, "deep travel", "minority check-in", and "hidden scenic spots" have increasingly become new trends in the cultural and tourism market. In recent years, the growth rate of UGC (User-Generated Content) related to "non-popular scenic spots" has been rapid, especially the attention to "micro-destinations" - those tourist spots that are remote, small in scale, and not yet promoted - is the most significant.

[0003] Traditional cultural and tourism recommendation systems mainly rely on static POI (Point-of-Interest) lists, supplemented by manually predefined "scenic spot levels" and "theme tags" (such as "historical culture" and "ecological leisure"), resulting in the inability to cover "minority check-in points" and "hidden beautiful scenery" emerging on social media, and it is difficult to discover emerging micro-destinations in a timely manner, leading to blind spots in cultural and tourism recommendations and being unfavorable to the development of the local cultural and tourism consumption industry. Summary of the Invention

[0004] This application provides a method, system, storage medium, computer program product, and electronic device for analyzing cultural and tourism content based on semantic recognition, so as to at least solve the problems in the current related technologies, such as relying on static interest point lists and manually predefined tags, making it difficult to discover and cover "micro-destinations" in a timely and comprehensive manner, resulting in recommendation blind spots and low recall rates.

[0005] In a first aspect, an embodiment of the present application provides a method for analyzing cultural and tourism content based on semantic recognition, including: obtaining a multi-source UGC cultural and tourism material set, and constructing a grid spatio-temporal domain module set according to a preset target area grid set and consecutive first and second historical time windows; each multi-source UGC cultural and tourism material in the multi-source UGC cultural and tourism material set includes text UGC data and corresponding visual UGC data, which are associated by the same content ID, and the corresponding timestamp and geographical coordinate information are recorded; each grid spatio-temporal domain module in the grid spatio-temporal domain module set is respectively used to indicate the target area grid under the corresponding historical time window; performing spatio-temporal grid matching on each of the multi-source UGC cultural and tourism materials according to the timestamp and geographical coordinate information, so as to be respectively assigned to the corresponding grid spatio-temporal domain module; for each of the grid spatio-temporal domain modules, extracting the text semantic vectors corresponding to the text UGC data in the grid spatio-temporal domain module, and calculating the corresponding grid text semantic center vector; calculating the semantic drift amount of the target area grid according to the first grid text semantic center vector corresponding to the first historical time window and the second grid text semantic center vector corresponding to the second historical time window, so as to screen at least one candidate hot area grid from each target area grid; for each of the candidate hot area grids, sampling a preset number of key frames from each of the visual UGC data corresponding to the second historical time window of the candidate hot area grid, so as to calculate the corresponding grid visual semantic center vector, and calculating the cross-modal semantic similarity in combination with the second grid text semantic center vector of the candidate hot area grid; generating a cultural and tourism destination recommendation list according to at least one target hot area grid whose corresponding cross-modal semantic similarity exceeds a preset similarity threshold.

[0006] Second aspect, an embodiment of the present application provides a cultural and tourism content analysis system based on semantic recognition, including: a data acquisition unit, configured to acquire a multi-source UGC cultural and tourism material set, and construct a grid spatio-temporal domain module set according to a preset target area grid set and consecutive first and second historical time windows; each multi-source UGC cultural and tourism material in the multi-source UGC cultural and tourism material set includes text UGC data and corresponding visual UGC data, which are associated by the same content ID, and record corresponding timestamps and geographical coordinate information; each grid spatio-temporal domain module in the grid spatio-temporal domain module set is respectively used to indicate the target area grid under the corresponding historical time window; a material spatio-temporal matching unit, configured to perform spatio-temporal grid matching on each of the multi-source UGC cultural and tourism materials according to the timestamps and geographical coordinate information, so as to be respectively allocated to the corresponding grid spatio-temporal domain module; a grid semantic calculation unit, configured to, for each of the grid spatio-temporal domain modules, extract the text semantic vectors corresponding to the text UGC data in the grid spatio-temporal domain module, and calculate the corresponding grid text semantic center vector; a hot grid screening unit, configured to calculate the semantic drift amount of the target area grid according to the first grid text semantic center vector corresponding to the first historical time window and the second grid text semantic center vector corresponding to the second historical time window, so as to screen at least one candidate hot area grid from each target area grid; a cross-modal alignment verification unit, configured to, for each of the candidate hot area grids, sample a preset number of key frames from each visual UGC data corresponding to the second historical time window of the candidate hot area grid, so as to calculate the corresponding grid visual semantic center vector, and calculate the cross-modal semantic similarity in combination with the second grid text semantic center vector of the candidate hot area grid; a micro-destination recommendation unit, configured to generate a cultural and tourism destination recommendation list according to at least one target hot area grid corresponding to the cross-modal semantic similarity exceeding a preset similarity threshold.

[0007] Third aspect, there is provided an electronic device, including: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the method for analyzing cultural and tourism content based on semantic recognition according to any embodiment of the present application.

[0008] Fourth aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the method for analyzing cultural and tourism content based on semantic recognition according to any embodiment of the present application are implemented.

[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the cultural and tourism content analysis method based on semantic recognition of any embodiment of the present application.

[0010] The cultural and tourism content analysis method and system based on semantic recognition provided by this application can produce at least the following technical effects: (1) By constructing a gridded spatiotemporal domain module set and combining it with historical time window comparative analysis, the semantic evolution trend of UGC content in the temporal and spatial dimensions can be captured. Specifically, by comparing the text semantic center vectors of the same grid unit in two consecutive time windows, the region with significant semantic drift can be effectively identified, that is, the region may have new gameplay or new scenes, and the geographical area with recent active content and rising discussion heat can be quickly focused on, thereby discovering emerging tourist hotspots that cannot be covered by traditional POI lists.

[0011] (2) After screening the candidate hotspot grids through the semantic drift of the text semantic vector, key frame visual semantic analysis is performed on such candidate hotspot grids, and cross-modal similarity is used as the criterion to avoid the interference of large-scale noise text (such as promotional advertisements or human-computer format reviews). The deep semantic alignment between the text UGC and image UGC of the grid unit is used to filter out the deviation caused by single modality analysis, ensuring that the identified hotspot areas have real tourism appeal and user interest support.

[0012] Through this technical solution, a cultural and tourism recommendation list based on actual user attention dynamics is generated, which can timely and accurately mine emerging and high-quality "micro-destinations" from the cultural and tourism social platform database, providing local governments or cultural and tourism suppliers with early identification tools for "potential hot areas", helping to formulate precise marketing strategies, optimize travel route planning and infrastructure layout, thereby promoting brand exposure and commercial value release of niche attractions, and empowering the development of the regional cultural and tourism consumption industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0014] Figure 1 A flowchart showing an example of a method for analyzing cultural and tourism content based on semantic recognition according to an embodiment of the present application is shown; Figure 2Shows an operation flowchart of an example for generating a cultural and tourism destination recommendation list according to an embodiment of the present application; Figure 3 Shows an operation flowchart of an example for calculating the semantic drift amount according to an embodiment of the present application; Figure 4 Shows an operation flowchart of an example for cross-modal semantic similarity calculation according to an embodiment of the present application; Figure 5 Shows a structural block diagram of an example of a cultural and tourism content analysis system based on semantic recognition according to an embodiment of the present application; Figure 6 Is a schematic structural diagram of an embodiment of an electronic device of the present application. Detailed implementation manners

[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0016] It should be noted that currently, some tourists pursue the "sense of privacy" of avoiding peak tourist seasons, and some tourists seek novel experiences such as "recommendations from online influencers" and "adventure check-ins". However, information about these micro-destinations is mostly scattered in social media short reviews, niche travel forums, or short videos of influencers, and it is difficult to capture them in a timely manner through traditional crawlers or manual tagging.

[0017] Traditional cultural and tourism recommendation systems mainly rely on static POI lists, supplemented by manually predefined "scenic spot levels" and "theme tags" (such as "historical culture" and "ecological leisure"). However, the number of POIs on mainstream OTA platforms mostly stays in the hundreds of thousands level, while the number of "grassroots check-in spots" on social platforms is conservatively estimated to exceed one million, resulting in a huge information blind spot.

[0018] In addition, in UGC, the descriptions of the same location by users will increase significantly with high-value information such as festivals and events. For example, "XX Ancient Town" is mostly described as "beautiful lantern festival scenery" before the Spring Festival, and as "misty rain gallery" during the rainy season, and it is difficult to capture this semantic-level transformation by simply relying on keyword statistics.

[0019] On the other hand, common spatiotemporal clustering only considers geographic coordinate density or release time windows, ignoring the "theme triggers" contained in the text - new ways of playing at the same place during different holidays. For example, in the short reviews of a rural market during the "Sophora japonica Festival", "flower sea", "market", and "special snacks" became high-frequency words, while these words rarely appear on weekdays. Therefore, if "semantics" and "time" are not incorporated into the model together, it is difficult to identify the scene as an emerging micro-destination.

[0020] It should be understood that the purpose of the above description of the current related art is only to facilitate the public to better understand the inventive spirit and motivation of the present application, and is not to be regarded as a limitation of the present application. In addition, the technical solutions described in the above-mentioned current related art are not prior art, and may also be undisclosed technical solutions, such as solutions under research or in the laboratory stage.

[0021] In the technical solution of this application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved shall comply with the provisions of relevant laws and regulations and shall not violate public order and good morals.

[0022] Figure 1 A flowchart of an example of a method for analyzing cultural and tourism content based on semantic recognition according to an embodiment of the present application is shown.

[0023] Regarding the execution subject of the method of the embodiment of the present application, it can be any controller or processor with computing or processing capabilities. Specifically, it can be implemented by the cultural tourism social UGC analysis and management platform. Through multi-source UGC fusion, spatiotemporal grid refinement management, semantic drift quantitative analysis and cross-modal verification, a multi-source, dynamic, semantically driven recommendation system architecture is constructed, which solves the problem of insufficient recognition of "niche check-in places" by the existing recommendation system, significantly improves the accuracy of hot spot discovery, the reliability of recommendation results and the scalability of the system, thereby providing strong technical support for the upgrading of local cultural tourism consumption and precise operation.

[0024] In some examples, it may be integrated and configured in an electronic device or terminal by means of software, hardware, or a combination of software and hardware, and the type of the terminal or electronic device may be diverse, such as a mobile phone, a tablet computer, or a desktop computer, etc.

[0025] like Figure 1 As shown, in step S110, a multi-source UGC cultural and tourism material set is obtained, and a grid spatiotemporal domain module set is constructed according to a preset target area grid set and a continuous first historical time window and a second historical time window.

[0026] Here, each piece of multi-source UGC cultural and tourism material in the multi-source UGC cultural and tourism material set contains text UGC data and corresponding visual UGC data, which are associated through the same content ID, and record the corresponding timestamp and geographical coordinate information. Each grid spatio-temporal domain module in the grid spatio-temporal domain module set is respectively used to indicate the target area grid under the corresponding historical time window.

[0027] In some embodiments, the cultural and tourism social UGC analysis and management platform can access multiple tourism social data sources through an authorization interface, so as to access the user-generated content publicly disclosed from multiple sources, such as Xiaohongshu, Mafengwo, Douyin, etc., and then collect the multi-source UGC cultural and tourism materials that meet the conditions according to the information structure "text UGC - visual UGC - metadata". Exemplarily, the text UGC data should include the title, text description or comment, etc., the visual UGC data should contain pictures or video frames, and the metadata can include the published content ID, geographical coordinate information (such as GPS longitude and latitude), and the publishing timestamp. Furthermore, with the content ID as the primary key, the text and visual content form a complete UGC data entry through an associated manner.

[0028] Regarding the details of grid cell division, in some exemplary embodiments, the geographical area to be detected (such as a certain city or a certain tourism area) can be divided into grid cells of equal size (for example, 500m × 500m) according to longitude and latitude, generate grid IDs, and can also mask non-tourism areas such as river waters according to business requirements to reduce invalid grids.

[0029] As an additional or alternative embodiment, the set of target area grids is a plurality of mutually independent target area grids generated by regular spatial division of the target area, and the division granularity of the target area grids is dynamically determined according to the characteristics of the urban functional areas of the target area. Therefore, by introducing a dynamic granularity division mechanism based on the characteristics of urban functional areas, the system uses publicly available urban planning data (such as urban zoning maps, land use data, administrative divisions) to identify the functional areas of the target area. Exemplarily, fine-grained division can be adopted for areas with dense commercial / scenic spots to accurately capture changes in user behavior in a small area; coarse-grained division can be adopted for ecological / rural areas to reduce noise interference in low-frequency data areas. In addition, it can also be adaptively adjusted according to the spatial density distribution of real-time UGC, that is, automatically refine the grid in areas with dense and active UGC to enhance the local hot spot detection ability.

[0030] Regarding the description of the continuous first historical time window and the second historical time window, they can be preset time windows adjacent to the current time, such as "the previous 7 - 14 days" and "the most recent 7 days", etc., and the window time length can also be flexibly adjusted or defined according to business requirements.

[0031] As an attachable or replaceable implementation, the first historical time window and the second historical time window are consecutive time periods defined based on a dynamic sliding window mechanism. The end time of the first historical time window coincides with the start time of the second historical time window, and the end time of the second historical time window is the current analysis reference time point. The dynamic sliding window mechanism is configured to shift forward at a preset time interval, so that the first historical time window and the second historical time window are updated synchronously over time.

[0032] Exemplarily, the first historical time window and the second historical time window are defined by a dual sliding window, where the first historical time window is , and the second historical time window is , is the time window length; is the analysis reference time point (such as the current system time) and is updated according to the sliding window step size. Thus, through the connection of the dual windows, both the "recent hotspots" and the "underlying trends" can be reflected simultaneously, ensuring that the drift detection is both sensitive and robust to "new fires". In addition, the preset time interval is less than the window duration of the historical time window. For example, the preset value (such as 1 day) is the minimum moving step size, and the overall window slides forward by this interval each time, continuously rolling on a daily granularity to capture the trend changes of each day; the step size is less than the window length, taking into account both data continuity and real-time performance.

[0033] In step S120, according to the time stamps and geographical coordinate information, each multi-source UGC cultural and tourism material is subjected to spatio-temporal grid matching to be respectively assigned to the corresponding grid spatio-temporal domain module.

[0034] In some implementations, each UGC cultural and tourism material is discretely matched to the corresponding grid ID and time window according to its recorded geographical coordinates and time stamps, so as to be accurately mapped to the corresponding grid spatio-temporal domain module. For example, a UGC released on March 15, 2025, located in grid A1, will be assigned to the T2 window module corresponding to grid A1 (for example, from March 14 to March 20), rather than the T1 window module corresponding to grid A1 (for example, from March 7 to March 13), so that each UGC is assigned to the corresponding grid spatio-temporal domain module. Thus, by adopting a geospatial index structure and a time sliding window mechanism, the matching efficiency is improved, supporting the efficient mapping and retrieval of large-scale content, and ensuring that the content is accurately assigned to its actual geographical location and time background.

[0035] In step S130, for each grid spatio-temporal domain module, the text semantic vectors corresponding to each text UGC data in the grid spatio-temporal domain module are extracted, and the corresponding grid text semantic center vector is calculated.

[0036] In some embodiments, for each grid spatio-temporal domain module, the system inputs all the text UGC data contained therein into a semantic understanding model (such as a BERT model or a RoBERTa model, etc.) to extract the semantic representation vectors of each piece of text. Subsequently, within the same grid module, the average value or weighted combination of all text semantic vectors is calculated as the text semantic center vector of this grid module. Thus, based on the vectorization of the deep semantic model, not only can keywords be identified, but also context associations can be captured, realizing the conversion of unstructured text information into a unified vector representation, enabling the system to quantify and compare the content expressions in different spatio-temporal regions.

[0037] In some examples of the embodiments of the present application, based on the BERT model, the text semantic vectors corresponding to each text UGC data in the grid spatio-temporal domain module are extracted.

[0038] Then, the arithmetic mean of all the text semantic vectors corresponding to the grid spatio-temporal domain module is taken and L2 norm normalization is performed to obtain the corresponding grid text semantic center vector: , Equation (1) In the formula, represents the grid grid text semantic center vector of, represents the grid the number of text UGCs within, and respectively represent the th text semantic vector and the th text semantic vector, represents the L2 norm operator.

[0039] It should be noted that all the text UGCs within the grid spatio-temporal domain module are generated within the same geographical grid and the same time window. Equal-weight averaging of their corresponding vectors is essentially a merger within the "same semantic body" and can represent the mainstream tendency of user discussions in this region during this period. Through arithmetic averaging, the abnormal vectors generated by a few extreme comments or robot spamming can be diluted, enabling the center vector to more reflect the true expressions of most users rather than being skewed by a few outliers. Through the dual processing of averaging and normalization, the interference caused by a small number of mislabels, spam information, or spamming can be effectively suppressed, making the center vector more discriminative for "true hotspots". In addition, the arithmetic averaging and L2 normalization have extremely small computational amounts, support real-time updates of thousands of grids simultaneously, and can meet the deployment requirements for large-scale regions.

[0040] In step S140, calculate the semantic drift amount of the target area grid according to the first grid text semantic center vector corresponding to the first historical time window and the second grid text semantic center vector corresponding to the second historical time window, so as to screen at least one candidate hot area grid from each target area grid.

[0041] In some embodiments, for each target area grid, compare the change degree between the text semantic center vectors under the first time window (T1) and the second time window (T2) respectively, and calculate its semantic drift amount, which represents the significant change degree of this area in the user's concerned topic. For example, an area described as "quiet valley" during T1, if new semantics such as "cliff swing" and "high-altitude glass plank road" frequently appear during T2, it indicates that this area may be "checked in" or "hot" due to factors such as new project construction and social media dissemination, and has the potential to become a "micro destination".

[0042] More specifically, in an example of the embodiment of the present application, a threshold can be preset, and grids with significant drift can be screened out as candidate hot areas by comparing with the semantic drift amount. In another example of the embodiment of the present application, the grids can also be sorted from large to small according to the drift amount, and several grids ranked at the top can be selected as candidate hot grids.

[0043] As a further optimized embodiment, the grid ranges of each candidate hot grid can also be compared with the POI list to identify whether the candidate hot grid coincides with a certain POI range. If they coincide, they can be filtered out to avoid the interference of existing popular POIs on the discovery process of niche popular micro destinations.

[0044] In some examples of the embodiment of the present application, summarize the semantic drift amounts of all target area grids to generate a global semantic drift distribution, thereby constituting a cross-regional semantic change distribution map to reflect the breadth and intensity of the change of the user interest theme in the entire target area during a specific period, and calculate the dynamic threshold corresponding to the preset statistical quantile based on the global semantic drift distribution. For example, set the semantic drift dynamic threshold according to the distribution characteristics of the drift amount data and specific statistical quantiles (such as 75%, 85%, 90%, etc.). Furthermore, at least one target area grid whose corresponding semantic drift amount exceeds the dynamic threshold is marked as a candidate hot area grid. Thus, the semantic drift dynamic threshold combined with the statistical quantile can dynamically adjust the number of hot areas, avoid "submerging micro hotspots" when the content is dense, and also avoid "amplifying false hotspots" during the cold period, realizing the balance between the stability and sensitivity of hot spot recognition.

[0045] In step S150, for each candidate hotspot area grid, a preset number of key frames are sampled from each visual UGC data of the second historical time window corresponding to the candidate hotspot area grid to calculate the corresponding grid visual semantic center vector, and the cross-modal semantic similarity is calculated in combination with the second grid text semantic center vector of the candidate hotspot area grid.

[0046] In some embodiments, for each candidate hotspot area grid, a certain amount of visual UGC (such as pictures or video clips) is extracted in the second historical time window closer to the current time, representative key frames are sampled, and converted into visual semantic vectors using image semantic models (such as CLIP, BLIP, etc.), and then these image semantic vectors are aggregated to form a visual semantic center vector. In addition, it should be noted that when performing cross-modal calculations, only all text content and visual content falling on the grid in the second time window (i.e., the window closer to the current time) is taken to ensure that the current center vector only maps the latest topic, and realizes comparison of the same spatiotemporal dimension, thereby ensuring the accuracy and robustness of hotspot detection.

[0047] Subsequently, the platform semantically matches the visual semantic center vector with the text semantic center vector of the grid and evaluates the cross-modal semantic similarity between the two. If the similarity is high, it indicates that the graphic content of the area is highly consistent in expression and has a real user cognition and communication basis. Therefore, by introducing the graphic and text semantic alignment mechanism of cross-modal similarity, false hot spots that are "hyped with words" but lack image evidence are effectively eliminated, ensuring that the remaining hot spot grids not only "everyone is saying that it has changed", but also "everyone is shooting new ways of playing or new scenes", effectively improving the authenticity and credibility of the recommended new cultural and tourism hot spots.

[0048] In step S160, a recommended list of cultural and tourism destinations is generated according to at least one target hotspot area grid whose corresponding cross-modal semantic similarity exceeds a preset similarity threshold.

[0049] In some embodiments, the platform can generate hot spot area names or geographic tags based on UGC frequency or hot spot area grids, and can add recommendation reasons (such as "the discussion popularity has increased by 200% in the past 7 days"). It can also select representative content with pictures and texts (for example, UGC picture and text content with high likes in the grid area), and generate a recommendation list based on the above content.

[0050] Regarding the business application of the recommended list of cultural and tourism destinations, on the one hand, it can be pushed by the platform according to the categories of different user interest tags; on the other hand, it can also be used by cultural and tourism management departments to promote and develop potential scenic spots, and there should be no restrictions here.

[0051] Thus, a closed-loop transformation from content-driven to behavior recommendation is achieved, truly discovering potential micro-destinations with development potential from "user-generated content", forming a personalized, highly real-time, and semantically authentic recommendation mechanism, which helps improve the response efficiency of the cultural and tourism system to emerging tourism behaviors and promotes the transformation of the cultural and tourism industry from "concentrated scenic spot check-ins" to "distributed micro-destination discovery".

[0052] Figure 2 The operation flowchart showing an example of generating a cultural and tourism destination recommendation list according to an embodiment of the present application is shown.

[0053] As Figure 2 shown, in step S210, for each target hot-spot area grid, the second grid text semantic center vector of the target hot-spot area grid is concatenated with the corresponding grid visual semantic center vector to obtain the corresponding grid multi-modal feature, and the attention weights between the grid multi-modal feature and each preset cultural and tourism feature label word vector are calculated, and at least one matching target cultural and tourism feature label is obtained through sorting.

[0054] In some embodiments, two types of center vectors of the adjacent time window of the target hot-spot grid that have undergone drift detection and cross-modal verification are read from the cache: the text semantic center (such as 768 dimensions) and the visual semantic center (such as 512 dimensions), and the two can be projected into the same-dimensional space (such as 256 dimensions) respectively through independent lightweight projection networks (each consisting of one layer of fully connected + LayerNorm), and L2 normalization is performed to eliminate the dimension and dimensionality differences.

[0055] Then, the normalized text vector and visual vector are concatenated into a multi-modal feature vector, and the text and visual signals are further mixed through one layer of residual fusion layer (such as, fully connected + ReLU + skip connection), so that the finally output multi-modal feature contains both "topic intention" and "scene perception" information.

[0056] In addition, a label vocabulary table can also be maintained in the platform, and various common cultural and tourism feature labels can be recorded in the table, such as "night tour light show", "deep tour of ancient towns", "rural pastoral experience", "food check-in", "parent-child interaction", or "outdoor adventure", etc. Furthermore, for each label phrase, the same BERT encoding and projection network as the text center are used to generate a label vector with the same dimension as the grid multi-modal feature, and new labels can be added to the label vocabulary table through an interface.

[0057] Then, dot products are performed on the grid multimodal vector and all label vectors to obtain their respective matching scores. The scores are normalized by softmax to obtain the attention weight distribution, which can reflect the correlation between the grid and each label. By sorting the attention weights from high to low, a preset number of labels with the highest ranking can be taken as the recommended characteristic labels for the hotspot, and the attention mechanism makes the labeling results more interpretable.

[0058] In some examples of the embodiments of the present application, the grid multimodal features and the label phrase vectors are mapped to the same vector space to ensure the comparability of similarity calculations. Furthermore, through the dot product + softmax attention mechanism, the relevance of each label to the grid feature is automatically measured, and the higher the weight, the stronger the match.

[0059] Specifically, the attention weight is calculated by the following formula: , Formula (2) In the formula, Indicates The attention weight of cultural and tourism feature tags, represents the transpose of the multimodal features of the grid, Indicates The label embedding vector of the cultural and tourism characteristic labels; represents the preset temperature coefficient, which is a positive real number less than 1 and is used to adjust the smoothness of attention distribution; Indicates the total number of cultural and tourism feature tags. Indicates The label embedding vector of the cultural and tourism characteristic labels, the denominator It represents the sum of the up-scaled matching scores of all labels to achieve softmax normalization.

[0060] In formula (2), by introducing the temperature coefficient in the attention mechanism To adjust the smoothness of attention distribution. Specifically, when When it is small (such as 0.05), the difference in the output of the exponential operation is amplified, the attention distribution is sharper, and the model only favors the most relevant few labels; when When it is larger (such as 0.2), the distribution is smoother and more labels can be taken into account. It can be adjusted as needed to achieve a balance in various scenarios from "highlighting recommendations" to "diversified displays" to meet personalized labeling needs.

[0061] In step S220, a recommended list of cultural and tourism destinations is generated based on each target hot spot area grid and the corresponding target cultural and tourism characteristic tags.

[0062] Through the embodiments of the present application, for each hotspot grid verified across modalities, a cultural and tourism label that best represents the characteristics of the area is automatically attached, replacing traditional multi-step entity extraction and graph marking, and realizing lightweight, interpretable, and easily extensible labeling recommendations.

[0063] In some embodiments, the specific location name can be retrieved by reverse querying based on the center coordinates of the target hotspot grid as the corresponding geographical location information. In addition, a high-weight UGC text summary (such as a refined evaluation sentence) within the grid and a representative key frame can be selected as the text and visual materials for the recommendation card. Through the above materials, a list of cultural and tourism destination recommendation cards is constructed, and the cultural and tourism characteristic labels screened by the corresponding grid, such as "night tour light show", "food punching", "parent-child interaction", etc., are prominently displayed at the prominent position of each card. Thus, the efficient conversion from massive UGC to refined cultural and tourism recommendation cards is realized, ensuring both the timeliness and high quality of the recommended content, and having good interpretability and scalability.

[0064] Figure 3 An operation flowchart showing an example of calculating the semantic drift amount according to an embodiment of the present application is shown.

[0065] As Figure 3 shown, in step S310, a semantic similarity index between the first grid text semantic center vector and the second grid text semantic center vector is calculated.

[0066] Here, the cosine similarity is used to measure the angular deviation of the overall topics of the "previous period" and "recent period" texts in the high-dimensional semantic space, directly reflecting the migration of the key points of user discussion.

[0067] , Equation (3) In the formula, represents the target area grid, represents corresponding to the first grid text semantic center vector, which is the average semantic center vector of all text UGCs within the first historical time window; represents corresponding to the second grid text semantic center vector, which is the average semantic center vector of all text UGCs within the second historical time window; represents the vector transpose operation, represents the L2 norm operator, represents the semantic similarity index, which is the cosine similarity of the text semantic center vector.

[0068] Specifically, for the previous text semantic center of the grid and the recent text semantic center By directly performing the inner product and dividing by the product of the norms, through normalization, the output value is stabilized between -1 and 1, and the risk of numerical overflow is low during the operation. Thus, it is possible to highly summarize the mainstream expression changes of all comments within the grid and filter out some extreme noises, which helps in the preliminary hot spot screening.

[0069] In step S320, based on the LDA topic model, the first grid text semantic center vector and the second grid text semantic center vector are processed to generate the topic distribution difference index corresponding to the continuous historical time windows.

[0070] It should be noted that semantic similarity can only reflect the alignment of the overall vectors, while the topic distribution difference can capture the structural changes of key topics (such as "night tour" and "food") in the high-dimensional space. The two complement each other. Therefore, the LDA (Latent Dirichlet Allocation) topic model is introduced to extract the topic probability distribution under each time window and use the KL divergence to quantify the distribution difference, so as to capture the structural topic evolution from the micro level.

[0071] More specifically, the LDA model can be offline trained on a large-scale cultural and tourism corpus (such as comments and travelogues). The number of topics is generally set to 20 - 30, covering various playing methods (food, night tour, cultural heritage, outdoor adventure, etc.), and the topic model can also be fine-tuned regularly to ensure that emerging playing methods (such as "immersive script killing" and "niche homestay experience") can be identified in a timely manner.

[0072] , Equation (4) In the formula, represents the total number of topics set by the LDA topic model, represents the th topic's probability distribution in the first historical time window of the grid , represents the th topic's probability distribution in the second historical time window of the grid ; represents the KL divergence, which is used to quantify the degree of change in the topic structure within the continuous historical time windows of the grid .

[0073] Here, for all text sets of each grid in the "early stage" and "recent stage", the topic probability vector distributions are respectively obtained by using model inference, and the two are compared element by element to complete the divergence accumulation. The larger the value, the more obvious the change in the topic structure. In this way, not only can we see the change in the "expression angle" of the text semantics, but also identify the migration trends of each sub-topic, such as the increase in the proportion of "food" and the decrease in the proportion of "night view".

[0074] In step S330, the semantic similarity index and the topic distribution difference index are fused to generate a corresponding semantic drift amount.

[0075] , Equation (5) In the formula, represents the semantic drift amount of the grid . The larger the value, the more significant the new topics emerging in this grid recently; and represent the text similarity weight coefficient and the topic divergence weight coefficient respectively.

[0076] Regarding the description of Equation (5), first convert the similarity into a distance and then fuse it to ensure the additivity of the two indicators within the same numerical magnitude; through linear weighting, combine the two types of metrics of "overall semantic distance" and "topic structure change" into a single drift amount.

[0077] In the embodiments of the present application, the cosine distance is used to quickly detect the overall expression shift, and the KL divergence is used to refine and capture the key topic structure changes. After fusing the two, it can cover two types of semantic drift features at the same time, which is more robust than a single indicator. In addition, the weights and can be determined by experience or A / B testing, can adapt to different scenarios or hot spot types, and the calculation process of the semantic drift amount is explicit, which helps the operation team to review the core reasons for the grid drift to be on the list. Thus, through the dual-index fusion of cosine similarity + KL divergence, a panoramic monitoring of "semantic expression change" and "topic structure change" is realized, and the accuracy of discovering new micro-destinations is greatly improved.

[0078] Figure 4 shows an operation flowchart of an example of cross-modal semantic similarity calculation according to an embodiment of the present application.

[0079] It should be noted that visual UGC (short video key frames, pictures) provides an objective "real scene mapping" for the cultural and tourism scene, making up for possible subjective omissions or noises in text descriptions. To integrate visual signals into the same semantic space, it is necessary to project the high-dimensional visual features into the same vector dimension as the text center through projection and perform normalization processing to achieve true cross-modal comparability.

[0080] As Figure 4 shown, in step S410, each key frame is input into a visual encoder to obtain a corresponding original visual feature vector.

[0081] Specifically, by locating all the key frames within the second historical time window that fall within the grid For short videos and images, from each short video or multiple images, a fixed number (e.g., 5 frames) of representative frames are extracted according to the scene switching and diversity strategy. Each frame of the image is sent to a visual encoder (such as ResNet, VGGNet, etc.) after unified preprocessing (scaling, normalization) to extract high-dimensional visual features.

[0082] , Equation (6) In the formula, represents the key frame of the th frame corresponding to the second historical time window, is the key frame index, represents the total number of key frames of the preset sampling; represents the visual encoder function, represents the th frame corresponding to the original visual feature vector.

[0083] In Equation (6), a pre-trained Vision Transformer (ViT) model is adopted, which can capture high-level features such as scene structure, color, and object presence; the dimension can be set to 512.

[0084] In step S420, the original visual feature vector is mapped to a fusion space with the same dimension as the text center through a linear projection layer, and L2 normalization processing is performed to obtain the corresponding frame projection transformation vector.

[0085] , Equation (7) In the formula, and respectively represent the visual projection weight matrix and the visual projection bias vector, represents the frame projection transformation vector corresponding to the th frame key frame after projection and L2 normalization.

[0086] In some embodiments, the projection network configuration can adopt a single-layer fully connected network (256×512 matrix) and bias to map the 512-dimensional original visual features to a 256-dimensional fusion space, and then immediately perform L2 normalization on the vector after affine transformation to ensure that the output vector norm is 1, eliminate the scale difference, and provide a fair comparison benchmark for the inner product with the text center. Thus, the high-dimensional visual semantics are seamlessly mapped to the text space, realizing a true same-dimension representation of text and images, and the intensity difference of the visual signal itself is removed through the normalization operation, making the cosine metric only focus on the direction consistency.

[0087] In step S430, arithmetic mean aggregation is performed on all the frame projection transformation vectors of the grid, and L2 normalization processing is performed again to obtain the grid visual semantic center vector.

[0088] It should be noted that single-frame visual features are often fragmented and affected by shooting angles and time nodes. By averaging the multi-frame visual projection vectors sampled from the same grid and normalizing them, a "visual average scene" can be extracted, which can denoise and highlight the most representative visual information of the grid in the recent period.

[0089] , Equation (8) , Equation (9) In the formula, represents the arithmetic mean aggregation vector of the grid , represents the grid visual semantic center vector of the grid .

[0090] Here, within the second historical time window, frames are extracted from each video according to "scene change detection". If the number of frames is too large, a fixed N frames (such as N = 5) are selected in a time-uniform or cluster center manner, and the normalized vectors after affine projection of each frame are aggregated into , and is obtained through arithmetic mean, and then L2 normalization is performed again to output . Thus, by cross-averaging the semantic vectors of different video frames, the deviation of a single perspective or accidental picture is reduced, the influence of different frame qualities, highlights / shadows is balanced, and the stability of the overall scene representation is improved.

[0091] In step S440, the cross-modal semantic similarity between the grid visual semantic center vector and the second grid text semantic center vector within the same period window is calculated in an inner product manner.

[0092] , Equation (10) In the formula, represents the cross-modal semantic similarity of the grid .

[0093] Here, the direction consistency of the "text center vector" and the "visual semantic center vector" in the fusion space within the recent window is directly measured through cosine similarity. Only when both point to similar high-dimensional positions does it indicate that there are both new gameplay / topic bursts and real-scene correspondences in the grid recently, and finally it is retained as a real cultural and tourism new hotspot. Thus, it can effectively reduce the misjudgment caused by "text spamming" or "visual overheating", and significantly reduce the false alarm rate. In addition, through the quantization of the similarity value, the operation can adjust the similarity threshold and observe the score distribution to obtain different recommendation results, which is beneficial for the operation to conduct multi-level analysis of cultural and tourism new hotspots.

[0094] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of combined actions. However, those skilled in the art should be aware that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0095] Figure 5 FIG. shows a structural block diagram of an example of a cultural and tourism content analysis system based on semantic recognition according to an embodiment of the present application.

[0096] As Figure 5 shown, the cultural and tourism content analysis system 500 based on semantic recognition includes a data acquisition unit 510, a material spatio-temporal matching unit 520, a grid semantic calculation unit 530, a hot grid screening unit 540, a cross-modal alignment verification unit 550, and a micro-destination recommendation unit 560.

[0097] The data acquisition unit 510 is configured to acquire a multi-source UGC cultural and tourism material set, and construct a grid spatio-temporal domain module set according to a preset target area grid set and consecutive first and second historical time windows; each multi-source UGC cultural and tourism material in the multi-source UGC cultural and tourism material set includes text UGC data and corresponding visual UGC data, which are associated by the same content ID, and record the corresponding timestamp and geographic coordinate information; each grid spatio-temporal domain module in the grid spatio-temporal domain module set is respectively used to indicate the target area grid under the corresponding historical time window.

[0098] The material spatio-temporal matching unit 520 is configured to perform spatio-temporal grid matching on each of the multi-source UGC cultural and tourism materials according to the timestamp and geographic coordinate information, so as to respectively allocate them to the corresponding grid spatio-temporal domain modules.

[0099] The grid semantic calculation unit 530 is configured to, for each of the grid spatio-temporal domain modules, extract the text semantic vectors corresponding to the text UGC data in the grid spatio-temporal domain module, and calculate the corresponding grid text semantic center vector.

[0100] The hot grid screening unit 540 is configured to calculate the semantic drift amount of the target area grid according to the first grid text semantic center vector corresponding to the first historical time window and the second grid text semantic center vector corresponding to the second historical time window, so as to screen at least one candidate hot area grid from each target area grid.

[0101] The cross-modal alignment verification unit 550 is configured to sample a preset number of key frames from each of the visual UGC data corresponding to the candidate hot spot area grid in the second historical time window for each of the candidate hot spot area grids, so as to calculate the corresponding grid visual semantic center vector, and calculate the cross-modal semantic similarity in combination with the second grid text semantic center vector of the candidate hot spot area grid.

[0102] The micro-destination recommendation unit 560 is configured to generate a cultural and tourism destination recommendation list according to at least one target hot spot area grid whose corresponding cross-modal semantic similarity exceeds a preset similarity threshold.

[0103] In some embodiments, the embodiments of the present application provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to be used for executing the steps of any one of the above-mentioned cultural and tourism content analysis methods based on semantic recognition in the present application.

[0104] In some embodiments, the embodiments of the present application further provide a computer program product, the computer program product includes a computer program stored on a non-volatile computer-readable storage medium, the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is enabled to execute the steps of any one of the above-mentioned cultural and tourism content analysis methods based on semantic recognition.

[0105] In some embodiments, the embodiments of the present application further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the cultural and tourism content analysis method based on semantic recognition.

[0106] Figure 6 is a schematic hardware structure diagram of an electronic device for executing the cultural and tourism content analysis method based on semantic recognition provided by another embodiment of the present application, as Figure 6 shown, the device includes: one or more processors 610 and a memory 620, Figure 6 Taking one processor 610 as an example.

[0107] The device for executing the cultural and tourism content analysis method based on semantic recognition may further include: an input device 630 and an output device 640.

[0108] The processor 610, the memory 620, the input device 630 and the output device 640 may be connected through a bus or other means, Figure 6Take the bus connection as an example.

[0109] As a non-volatile computer-readable storage medium, the memory 620 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as program instructions / modules corresponding to the method for analyzing cultural and tourism content based on semantic recognition in the embodiments of the present application. The processor 610 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 620, that is, implements the method for analyzing cultural and tourism content based on semantic recognition in the above method embodiments.

[0110] The memory 620 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device. In addition, the memory 620 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 620 may optionally include a memory remotely provided relative to the processor 610, and these remote memories can be connected to the electronic device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0111] The input device 630 can receive input digital or character information, and generate signals related to the user settings and function control of the electronic device. The output device 640 may include a display device such as a display screen.

[0112] The one or more modules are stored in the memory 620 and, when executed by the one or more processors 610, execute the method for analyzing cultural and tourism content based on semantic recognition in any of the above method embodiments.

[0113] The above products can execute the methods provided in the embodiments of the present application, and have corresponding functional modules and beneficial effects for executing the methods. For technical details not described in detail in this embodiment, reference can be made to the methods provided in the embodiments of the present application.

[0114] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aiming to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0115] (2) Ultra-mobile personal computer devices: Such devices fall within the category of personal computers, have computing and processing capabilities, and generally also possess the feature of mobile Internet access. Such terminals include: PDA, MID, UMPC devices, etc.

[0116] (3) Portable entertainment devices: Such devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, as well as smart toys and portable in-vehicle navigation devices.

[0117] (4) Other airborne electronic devices with data interaction functions, such as in-vehicle device installed on vehicles.

[0118] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0119] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or rather the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0120] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A method for analyzing cultural and tourism content based on semantic recognition, characterized in that The method includes: Obtaining a multi-source UGC cultural and tourism material set, and constructing a grid spatio-temporal domain module set according to a preset target area grid set and consecutive first and second historical time windows; each multi-source UGC cultural and tourism material in the multi-source UGC cultural and tourism material set includes text UGC data and corresponding visual UGC data, which are associated through the same content ID, and record the corresponding timestamp and geographical coordinate information; each grid spatio-temporal domain module in the grid spatio-temporal domain module set is respectively used to indicate the target area grid under the corresponding historical time window. Performing spatio-temporal grid matching on each of the multi-source UGC cultural and tourism materials according to the timestamp and geographical coordinate information, so as to be respectively assigned to the corresponding grid spatio-temporal domain module. For each of the grid spatio-temporal domain modules, extracting the text semantic vectors corresponding to the text UGC data in the grid spatio-temporal domain module, and calculating the corresponding grid text semantic center vector. Calculating the semantic drift amount of the target area grid according to the first grid text semantic center vector corresponding to the first historical time window and the second grid text semantic center vector corresponding to the second historical time window, so as to screen at least one candidate hot area grid from each target area grid. For each of the candidate hot area grids, sampling a preset number of key frames from each visual UGC data corresponding to the second historical time window of the candidate hot area grid, so as to calculate the corresponding grid visual semantic center vector, and calculating the cross-modal semantic similarity in combination with the second grid text semantic center vector of the candidate hot area grid. Generating a cultural and tourism destination recommendation list according to at least one target hot area grid whose corresponding cross-modal semantic similarity exceeds the preset similarity threshold.

2. The method according to claim 1, wherein The target area grid set is a plurality of mutually independent target area grids generated by regular spatial division of the target area; the division granularity of the target area grid is dynamically determined according to the urban functional area characteristics of the target area. The first historical time window and the second historical time window are consecutive time periods defined based on a dynamic sliding window mechanism, the end time of the first historical time window coincides with the start time of the second historical time window, and the end time of the second historical time window is the current analysis benchmark time point. The dynamic sliding window mechanism is configured to move forward at a preset time interval, so that the first historical time window and the second historical time window are updated synchronously with time; the preset time interval is less than the window duration of the historical time window.

3. The method according to claim 1, characterized in that The generating a cultural and tourism destination recommendation list according to at least one target hot area grid whose corresponding cross-modal semantic similarity exceeds the preset similarity threshold includes: For each of the target hot area grids, splicing the second grid text semantic center vector of the target hot area grid with the corresponding grid visual semantic center vector to obtain the corresponding grid multi-modal feature, and calculating the attention weights between the grid multi-modal feature and each preset cultural and tourism characteristic label word vector, and obtaining at least one matching target cultural and tourism characteristic label through sorting. Generate a list of recommended cultural and tourism destinations based on each target hot spot area grid and the corresponding target cultural and tourism feature tags.

4. The method according to claim 3, wherein Calculating the attention weights between the grid multimodal features and the vector representations of each preset cultural and tourism feature tag includes: , Wherein, represents the attention weight of the th cultural and tourism feature label, represents the transpose of the grid multi-modal feature, represents the th label embedding vector of the cultural and tourism feature label; represents a preset temperature coefficient, which is a positive real number less than 1 and is used to adjust the smoothness of the attention distribution; represents the total number of cultural and tourism feature labels, represents the th label embedding vector of the cultural and tourism feature label, and the denominator represents the sum of the amplified matching scores of all labels to achieve softmax normalization.

5. The method according to claim 1 or 4, characterized in that, The semantic drift amount is calculated in the following manner: Calculate the semantic similarity index between the first grid text semantic center vector and the second grid text semantic center vector: , In the formula, represents the grid of the target area, represents corresponding to the first grid text semantic center vector, which is the average semantic center vector of all text UGCs within the first historical time window; represents corresponding to the second grid text semantic center vector, which is the average semantic center vector of all text UGCs within the second historical time window; represents the vector transpose operation, represents the L2 norm operator, represents the semantic similarity index, which is the cosine similarity of the text semantic center vectors; Based on the LDA topic model, process the first grid text semantic center vector and the second grid text semantic center vector to generate a topic distribution difference index corresponding to a continuous historical time window: , In the formula, represents the total number of topics set by the LDA topic model, represents the -th topic's probability distribution in the first historical time window of grid , represents the -th topic's probability distribution in the second historical time window of grid ; represents the KL divergence, which is used to quantify the degree of change in the topic structure within consecutive historical time windows of grid . Fuse the semantic similarity index and the topic distribution difference index to generate the corresponding semantic drift amount: , In the formula, represents the semantic drift amount of the grid . The larger the value, the more significant the emergence of new topics in this grid in the near future; and represent the text similarity weight coefficient and the topic divergence weight coefficient respectively.

6. The method according to claim 5, wherein Screening at least one candidate hot spot area grid from each target area grid includes: Summarize the semantic drift amounts of all target area grids to generate a global semantic drift distribution, and calculate the dynamic threshold corresponding to the preset statistical quantile based on the global semantic drift distribution; Mark at least one target area grid with a corresponding semantic drift amount exceeding the dynamic threshold as a candidate hot spot area grid.

7. The method according to claim 6, wherein Calculating the corresponding grid visual semantic center vector and calculating the cross-modal semantic similarity in combination with the second grid text semantic center vector of the candidate hot spot area grid includes: Input each key frame into the visual encoder to obtain the corresponding original visual feature vector: , In the formula, represents the key frame of the th frame corresponding to the second historical time window, is the key frame index, represents the total number of key frames for preset sampling; represents the visual encoder function, represents the original visual feature vector corresponding to the Map the original visual feature vector to a fusion space with the same dimension as the text center through a linear projection layer and perform L2 normalization to obtain the corresponding frame projection transformation vector: , Wherein, and respectively represent a visual projection weight matrix and a visual projection bias vector, represents the frame projection transformation vector corresponding to the th key frame after projection and L2 normalization; Arithmetically average and aggregate all the frame projection transformation vectors of the grid, and perform L2 normalization again to obtain the grid visual semantic center vector: , , In the formula, represents the arithmetic mean aggregation vector of the grid , and represents the grid visual semantic center vector of the grid . Calculate the cross-modal semantic similarity between the grid visual semantic center vector and the second grid text semantic center vector within the same time window in an inner product manner: , In the formula, represents the cross-modal semantic similarity of the grid .

8. The method according to claim 7, characterized in that Extracting the text semantic vectors corresponding to each text UGC data in the grid spatio-temporal domain module includes: Based on the BERT model, extract the text semantic vectors corresponding to each text UGC data in the grid spatio-temporal domain module; Calculating the corresponding grid text semantic center vector includes: Take the arithmetic mean of all text semantic vectors corresponding to the grid spatio-temporal domain module and perform L2 norm normalization to obtain the corresponding grid text semantic center vector: , In the formula, represents the grid text semantic center vector of the grid, represents the grid the number of text UGCs within the grid, and respectively represent the th text semantic vector and the th text semantic vector.

9. A cultural and tourism content analysis system based on semantic recognition, characterized in that, The system includes: A data acquisition unit for acquiring a multi-source UGC cultural and tourism material set and constructing a grid spatio-temporal domain module set according to a preset target area grid set and continuous first and second historical time windows; each multi-source UGC cultural and tourism material in the multi-source UGC cultural and tourism material set includes text UGC data and corresponding visual UGC data, which are associated by the same content ID and record the corresponding timestamp and geographical coordinate information; each grid spatio-temporal domain module in the grid spatio-temporal domain module set is used to indicate the target area grid under the corresponding historical time window; A material spatio-temporal matching unit, configured to perform spatio-temporal grid matching on each of the multi-source UGC cultural and tourism materials according to the time stamp and geographical coordinate information, so as to be respectively allocated to the corresponding grid spatio-temporal domain module; A grid semantic calculation unit, configured to, for each of the grid spatio-temporal domain modules, extract the text semantic vectors corresponding to the text UGC data in the grid spatio-temporal domain module, and calculate the corresponding grid text semantic center vector; A hot spot grid screening unit, configured to calculate the semantic drift amount of the target area grid according to the first grid text semantic center vector corresponding to the first historical time window and the second grid text semantic center vector corresponding to the second historical time window, so as to screen at least one candidate hot spot area grid from each of the target area grids; A cross-modal alignment verification unit, configured to, for each of the candidate hot spot area grids, sample a preset number of key frames from each of the visual UGC data corresponding to the second historical time window of the candidate hot spot area grid, so as to calculate the corresponding grid visual semantic center vector, and calculate the cross-modal semantic similarity in combination with the second grid text semantic center vector of the candidate hot spot area grid; A micro-destination recommendation unit, configured to generate a cultural and tourism destination recommendation list according to at least one target hot spot area grid whose corresponding cross-modal semantic similarity exceeds a preset similarity threshold.

Citation Information

Patent Citations

  • Deep semantic understanding method based on cross-modal model

    CN116680578A

  • Speech Recognition Accuracy with Natural-Language Understanding based Meta-Speech Systems for Assistant Systems

    US20210118440A1

Cited By

  • AIGC-based cultural feature accurate identification method and system

    CN121303147A

  • A method and system for accurate identification of cultural features based on AIGC

    CN121303147B