Group intelligent driven multi-modal travel knowledge graph retrieval method and system

By collecting multimodal data and performing intelligent analysis and dynamic updates, a multimodal cultural tourism knowledge graph is constructed, which solves the problems of insufficient semantic analysis and weak personalized retrieval capabilities in existing technologies, and realizes more accurate and personalized intelligent retrieval.

CN121833966APending Publication Date: 2026-04-10GUANGDONG URBAN & RURAL PLANNING & DESIGN INST
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing cultural and tourism knowledge graphs lack semantic analysis, multimodal integration, dynamic adaptability, and personalized intelligent retrieval capabilities, resulting in unsatisfactory intelligent retrieval results.

Method used

Collect multimodal historical data, analyze and perform triplet structuring and alignment fusion through various artificial intelligence models, construct a dynamic index framework, and combine group co-creation and dynamic self-learning mechanisms to achieve multimodal cultural tourism knowledge graph retrieval.

Benefits of technology

It enhances the expressiveness and interpretability of the map, provides more accurate and personalized search results, has strong dynamic update capabilities, and supports efficient fusion of multimodal data and personalized recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833966A_ABST
    Figure CN121833966A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence graph analysis, in particular to a group intelligent driven multi-modal travel knowledge graph retrieval method and system.The method comprises the steps that firstly, multi-modal historical data of a target tourist group is collected, and then an artificial intelligence model is adopted to analyze the multi-modal historical data; performing triple structuring and aligned fusion on various analysis results to obtain a triple unit and a node multi-mode embedding vector; then, writing the triple unit and the node multi-mode embedded vector into a knowledge graph database, obtaining a high-dimensional index framework through graph fusion and redundancy resolution, inputting real-time contribution data of a target tourist group into a preset dynamic evolution mechanism to drive the high-dimensional index framework to perform dynamic updating, obtaining a dynamic index framework, and obtaining a knowledge graph database; and finally, in response to triggering of a user retrieval instruction, inputting the user retrieval instruction and personal information authorized by the user into a preset intelligent application interface so as to call the dynamic index framework for retrieval to obtain a retrieval result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence graph analysis. More particularly, the present application relates to a multi-modal cultural tourism knowledge graph retrieval method and system driven by swarm intelligence. BACKGROUND

[0002] The scenic spot distribution map is generally placed at the scenic spot, and the self-driving tourists need to actually arrive at the destination to make more specific travel plans according to the map content, which has certain regional limitations. Moreover, such a map is generally static and does not have dynamic audio-visual experience, and the actual content presented is limited.

[0003] With the rise of artificial intelligence technology, a cultural tourism knowledge graph has emerged to replace the above-mentioned scenic spot distribution map. It can be loaded in a mobile cultural tourism app application, and is mainly used to associate cultural entities (such as the Forbidden City, the Great Wild Goose Pagoda, or the Lijiang Ancient City, etc.) with cultural attributes (such as geographical location, development time change, historical time, and surrounding industry, etc.) to provide intelligent retrieval, travel planning, and intelligent explanation that meet the preferences or expectations of self-driving tourists. Among them, the Chinese patent application file with publication number CN120542539A discloses a cross-modal cultural tourism knowledge graph retrieval system based on a large language model, which mainly uses an ARIMA model and a K-Means algorithm to realize the above-mentioned intelligent recommendation and travel planning functions. However, it has the following defects: Firstly, the semantic analysis dimension of the overall scheme is insufficient: a lot of visual, auditory, and behavioral information is contained in the cultural tourism scene, such as scenic spot photos, folk music, video images, tourist travel trajectories, etc. Using only text / picture information cannot fully depict the tourism experience, making the graph unable to cover multi-sensory semantics, and its expression ability is insufficient, lacking integration of multi-modal such as vision, audio, and text, resulting in unsatisfactory intelligent retrieval.

[0004] Secondly, poor dynamic adaptability: the ARIMA model is more suitable for stable time series prediction, and has weak response capability for real-time dynamic data (such as sudden public opinion or temporary activities), resulting in unsatisfactory intelligent retrieval effect.

[0005] Thirdly, poor individualized intelligent retrieval capability: the K-Means algorithm mainly divides users based on group characteristics, and the recommended results are biased towards "massification", making it difficult to meet the fine-grained individualized needs.

[0006] Therefore, the existing knowledge graph technology in the cultural tourism field has the problem of poor intelligent retrieval capability. SUMMARY

[0007] To solve the above-mentioned technical problem of poor intelligent retrieval capability, the present application discloses a multi-modal cultural tourism knowledge graph retrieval method and system driven by swarm intelligence.

[0008] In a first aspect, the method of the present application discloses a swarm intelligence driven multi-modal cultural tourism knowledge graph retrieval method, comprising: Collecting multi-modal historical data of a target tourist group; wherein the multi-modal historical data includes text / speech data, image / video data, audio data, and trajectory positioning data; Using multiple pre-set artificial intelligence models to analyze corresponding types of multi-modal historical data, and performing triple structure and alignment fusion on multiple analysis results to obtain triple units and node multi-modal embedding vectors, respectively; Writing the triple units and node multi-modal embedding vectors into a pre-set knowledge graph database, and performing knowledge graph fusion and redundancy resolution to obtain a high-dimensional index framework; Inputting real-time contribution data of the target tourist group into a pre-set dynamic evolution mechanism to drive the high-dimensional index framework to update dynamically and obtain a dynamic index framework; In response to the triggering of a user retrieval instruction, inputting the user retrieval instruction and the user's authorized personal information into a pre-set intelligent application interface to call the dynamic index framework for retrieval and obtain a retrieval result.

[0009] Beneficial effects: The present method first introduces a joint representation method of five modalities including text, image, audio, video, and trajectory in the cultural tourism knowledge graph. Through multi-modal vector fusion, the semantic representation of the graph nodes not only contains encyclopedic textual knowledge, but also fuses visual perception, auditory clues, and tourist behavior preferences. This expands the knowledge dimension of the graph, enabling it to more comprehensively depict cultural tourism experiences. Compared with existing technologies, the expressiveness and interpretability of the present method have been significantly enhanced, which helps to improve user understanding and satisfaction with the graph retrieval results. On this basis, the present method also structures and vectorizes the above multi-modal historical data, and constructs a dynamic index framework supporting user retrieval through fusion, group co-creation, dynamic self-learning, and other strategies, making the retrieval results obtained by users more accurate and personalized.

[0010] Preferably, the artificial intelligence models include a natural language processing model, a computer vision model, an audio feature extraction model, and a time-space clustering model; and the corresponding types of multi-modal historical data are analyzed using multiple pre-set artificial intelligence models, specifically: Inputting the text / speech data into the natural language processing model for text feature recognition and analysis to obtain a first analysis result; Inputting the image / video data into the computer vision model for visual feature recognition and analysis to obtain a second analysis result; Inputting the audio data into the audio feature extraction model for music feature recognition and analysis to obtain a third analysis result; The trajectory positioning data is input into the space-time clustering model for trajectory feature recognition and analysis, to obtain a fourth analysis result.

[0011] Preferably, the plurality of analysis results are aligned and fused, specifically: The first analysis result, the second analysis result, the third analysis result, and the fourth analysis result are aligned using a preset multi-modal alignment and fusion model. The vectors corresponding to the aligned first analysis result, the second analysis result, the third analysis result, and the fourth analysis result are spliced to obtain node multi-modal embedding vectors.

[0012] Preferably, the triple unit includes entity name, key connection relationship, and cultural attribute.

[0013] Preferably, the knowledge graph database includes a graph database and a vector database; the triple unit and the node multi-modal embedding vector are written into the preset knowledge graph database, including: The triple unit is stored in the graph database; The node multi-modal embedding vector is stored in the vector database; The triple unit and the corresponding node multi-modal embedding vector are associated through a node unique ID.

[0014] Preferably, knowledge graph fusion and redundancy resolution are performed, including: The cosine similarity between node multi-modal embedding vectors of two different sources in the vector database is calculated, and if the cosine similarity is higher than a similarity threshold, the node multi-modal embedding vectors of the two different sources are merged; The key connection relationships of all triple units in the graph database are semantically normalized.

[0015] Preferably, real-time contribution data of the target tourist group is input into a preset dynamic evolution mechanism to drive dynamic updating of the high-dimensional index framework, including: Defining data entry conditions for real-time contribution data; Performing credibility operation on real-time contribution data that meets the data entry conditions to obtain credibility; If the credibility is higher than a credibility threshold, the real-time contribution data is written into the knowledge graph database; According to the real-time contribution data, the node multi-modal embedding vector written into the knowledge graph database is fine-tuned.

[0016] Preferably, the user retrieval instruction includes a structured retrieval instruction and a semantic retrieval instruction.

[0017] Preferably, the intelligent application interface includes RESTful API and SDK.

[0018] In a second aspect, the present application further discloses a multi-modal cultural and travel knowledge graph retrieval system driven by swarm intelligence, comprising a processor and a memory, and the memory stores computer program instructions, which realize the multi-modal cultural and travel knowledge graph retrieval method driven by swarm intelligence as recorded in the first aspect when executed by the processor.

[0019] The present application has the following advantages: (1) Compared with the prior art, the expressiveness and interpretability of the method of the present application are significantly enhanced, which helps to improve the user's understanding and satisfaction with the graph retrieval results. On this basis, the method of the present application further structures and vectorizes the above-mentioned multi-modal historical data, and constructs a dynamic index framework supporting user retrieval through fusion, group co-creation, dynamic self-learning and other strategies, so that the retrieval results obtained by the user are more accurate and personalized.

[0020] (2) Unlike the traditional knowledge base constructed by experts in a closed manner, the method of the present application fully utilizes the wisdom of the crowd to generate and continuously optimize the knowledge graph database. When the credibility is higher than the credibility threshold, the real-time contribution data is written into the knowledge graph database to realize decentralized knowledge co-construction and review. This technical solution is introduced into the field of cultural and travel graph, and combined with credit points to prevent low-quality contributions, to ensure that the knowledge source is transparent and credible.

[0021] (3) Compared with the prior art, the method of the present application can meet the user's precise definition of retrieval intention, while also taking into account the needs of semantic expansion and fuzzy matching. BRIEF DESCRIPTION OF DRAWINGS

[0022] The above and other objects, features and advantages of the exemplary embodiments of the present application will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which several embodiments of the present application are shown by way of example, and in which the same or corresponding elements are referred to by the same or corresponding reference numerals, in which: Figure 1 is a flow chart of the multi-modal cultural and travel knowledge graph retrieval method driven by swarm intelligence in the first embodiment of the present application; Figure 2 is a structural schematic diagram of the multi-modal cultural and travel knowledge graph retrieval system driven by swarm intelligence in the second embodiment of the present application. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the present application.

[0024] The specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0025] Embodiment one As shown in the figure, the embodiment discloses a multi-modal travel knowledge graph retrieval method driven by swarm intelligence, which comprises: Figure 1 S10: Collecting multi-modal historical data of the target tourist group.

[0026] Among them, the above-mentioned multi-modal historical data is mainly divided into text / speech data, image / video data, audio data and trajectory positioning data. It needs to be declared that the collection of the above-mentioned multi-modal historical data needs to be authorized by the user's mobile phone APP and opened the corresponding storage access right if it is for the user personally; if it is for the network community / platform, it can be collected through the multi-source content uploaded in the network community / platform.

[0027] More specifically, the above-mentioned text / speech data can be collected from the content of the tourists' travel notes, scenic spot comments, voice explanation, etc., mainly collecting the text and its metadata of the above-mentioned content. Among them, the voice can be converted into text through the existing voice recognition model. Metadata can include uploader ID, time, geographic location and device.

[0028] More specifically, the above-mentioned image / video data can be collected from the visual content such as scenic spot photos and short video clips shared by users, and the image / video data contains associated scenic spots, event tags, shooting locations and time data. For example, a tourist uploads the video and photo of the Miao New Year's celebration. More specifically, the above-mentioned audio data refers to the folk music clips and explanation audio recorded by the target tourist group during the travel process, which is collected in combination with its text description or metadata (such as music name and type).

[0029] More specifically, the above-mentioned trajectory positioning data can be collected by starting the positioning sharing of the travel route GPS trajectory of the tourist, which is a time series of geographic coordinate points used to record the location and stay time of the tourist during the journey.

[0030] It should be noted that when collecting the above-mentioned multi-modal historical data, the associated geographic location (latitude and longitude or place name), timestamp (content generation or upload time), user ID and data source / device metadata need to be recorded uniformly to assist subsequent knowledge fusion.

[0031] S20: Using multiple preset artificial intelligence models to analyze the corresponding types of multi-modal historical data, and performing triple structure and alignment fusion on the multiple analysis results to obtain triple units and node multi-modal embedding vectors respectively. ​

[0032] In the present embodiment, the artificial intelligence model includes a natural language processing model, a computer vision model, an audio feature extraction model, and a time-space clustering model.

[0033] In the step S20, the multi-modal historical data of each type is analyzed by using a plurality of preset artificial intelligence models, specifically: S21: input the text / speech data into the natural language processing model to identify and analyze the text features, and obtain a first analysis result.

[0034] For text / speech data analysis, the present embodiment can use the BERT (Bidirectional Encoder Representations from Transformers) model in the natural language processing model to perform named entity recognition and relationship extraction on the text content. For speech content, first perform speech recognition to convert it into text, and then apply the same NLP (Natural Language Processing) pipeline for named entity recognition and relationship extraction.

[0035] For example, from a tourist comment: “During the Miao New Year Festival, I enjoyed the X tribe silverware making performance in Southeast Guizhou, which was very impressive”, the entities “Miao New Year Festival”, “Southeast Guizhou” and “X tribe silverware making” can be identified, as well as the relationship “the relationship between the three”. NLP analyzes this as a triple (X tribe silverware making, belongs to, Intangible Cultural Heritage project) and (Miao New Year Festival, location, Southeast Guizhou, Guizhou), etc.

[0036] S22: input the image / video data into the computer vision model to identify and analyze the visual features, and obtain a second analysis result.

[0037] For image / video data analysis, the present embodiment uses the CLIP (Contrastive Language-Image Pre-training) model or the visual Transformer ViT model in the computer vision model to perform image recognition on pictures and videos frame by frame, and extract cultural symbols, scene elements and entities contained therein.

[0038] For example, through image recognition, the “X tribe silverware making” handicraft scene or X tribe costume characters appearing in the photo, and the “Miao New Year Festival drum dance performance” appearing in the video, etc. The visual model outputs the labels of the recognized entities and possible relationships, such as “detecting that a building belongs to a scenic spot” as the second analysis result.

[0039] S23: input the audio data into the audio feature extraction model to identify and analyze the music features, and obtain a third analysis result.

[0040] For audio data analysis, the method inputs audio data into a VGGish or MFCC (Mel Frequency Cepstral Coefficient) model in the audio feature extraction model to identify key information in the audio.

[0041] For example, the model identifies a segment of audio as a segment of X-tribe big song, extracts the non-heritage music entity "X-tribe big song", and associates the village where it is located as the third analysis result.

[0042] S24: Input the trajectory positioning data into the time-space clustering model to identify and analyze the trajectory features, and obtain the fourth analysis result.

[0043] For trajectory positioning data analysis, the method mainly inputs GPS trajectory data into a time-space clustering algorithm (DBSCAN) to detect the stay points and movement paths of tourists. Specifically, positions with a speed lower than a threshold and a stay time longer than a certain time are identified as stay points (Stay Point), and are connected to form a behavior path. For example, through trajectory clustering, it is found that a tourist stayed near the coordinates of Miao Village for 2 hours, and the stay point corresponds to the entity "West River Thousand Miao Village", and the travel route of the tourist is obtained by connecting multiple stay points. These stay points and routes are also important knowledge, and can extract relationships such as (tourist X, visited West River Thousand Miao Village at time T) or (route Y, contains, sequence of scenic spots).

[0044] Through the above steps S21-S24, the method calls multiple artificial models to realize the specific analysis of multi-modal historical data related to tourism, and provides more reliable and comprehensive data sources for subsequent steps.

[0045] It should be explained that the above triple unit includes entity name, key relationship and cultural attribute. For example, the specific expression of the above triple unit can be: (X-tribe silverware making belongs to non-heritage project), which indicates that "X-tribe silverware making" is a national non-heritage project. (Miao New Year Festival is held in Guizhou Qianzhong South), which indicates that "Miao New Year Festival" is held in Guizhou Qianzhong South. In addition, appropriate key relationships are also given to the entities identified by images and audio, such as (X-tribe silverware making technique demonstration in West River Thousand Miao Village Museum), etc. Each triple unit should also be accompanied by content source data (text sentence or picture ID) as evidence.

[0046] Further, for the alignment and fusion of multiple analysis results in step S20, specifically: S25: Use a preset multimodal alignment and fusion model to align the first, second, third, and fourth parsing results.

[0047] In order to represent the parsing results from the different modal sources in a unified semantic space, this embodiment adopts a multimodal alignment fusion model to map the feature vectors of text, vision, audio, trajectory and other features into a shared vector representation.

[0048] Specifically, the information of each modality is first encoded into vectors: text uses CLS vectors from a pre-trained language model to represent text semantics; images use the last layer feature vectors of a visual model (such as ResNet or ViT) to represent image content; audio features are extracted from corresponding dimensions; and trajectories are extracted from statistical feature vectors of the stop point sequence. Alignment mainly refers to unifying the feature lengths of the above feature vectors.

[0049] S26: Concatenate the vectors corresponding to the aligned first, second, third, and fourth parsing results to obtain the node multimodal embedding vector.

[0050] Specifically, the above vectors are aligned and fused using a multimodal alignment and fusion model to generate a unified node multimodal embedding vector. In this embodiment, the fusion strategy mainly employs nonlinear mapping data concatenation; in other embodiments, weighted summation or attention fusion may also be used.

[0051] For example, the 768-dimensional features of text and images are concatenated and mapped back to 768 dimensions through a fully connected network to obtain a fused semantic representation. The fused vector contains both the structural semantics of the nodes and features from the perceptual levels such as vision and hearing, giving the graph nodes a hybrid representation of "structure + vector".

[0052] Through the above technical solution, each entity node, in addition to relational connections, i.e. triple units, also stores a multimodal embedding vector, i.e. node multimodal embedding vector, which can support vectorized semantic retrieval and computation.

[0053] S30: Write the triplet units and node multimodal embedding vectors into the preset knowledge graph database, and perform knowledge graph fusion and redundancy resolution to obtain a high-dimensional index framework.

[0054] The aforementioned knowledge graph database includes a graph database and a vector database, and step S30 includes: S31: Write the triplet cells into the graph database for storage.

[0055] Exemplarily, the structured triple units are stored in a graph database such as Neo4j to support efficient relationship queries.

[0056] S32: Write the node multi-modal embedding vector into the vector database for storage.

[0057] Step S32 is performed simultaneously with step S31 described above. Exemplarily, the node multi-modal embedding vector can be stored in a vector database such as Milvus to support embedding-based nearest neighbor search.

[0058] S33: Associate the triple unit with the corresponding node multi-modal embedding vector through the node unique ID.

[0059] Specifically, the above two are associated through the node unique ID, realizing the representation of the node with both graph structure and semantic vector. Through the above steps S31-S33, the knowledge graph database is converted into a hybrid architecture of graph database + vector database.

[0060] It should be noted that, in order to facilitate subsequent simple description, the subsequent "graph" or "knowledge graph" refers to the above-mentioned knowledge graph database. The "high-dimensional index framework" and "dynamic index framework" referred to in the context are essentially knowledge graph databases, and the two designations are mainly used to distinguish the two different states of the knowledge graph database.

[0061] Further, in order to maintain the consistency of the data in the graph, it is necessary to handle the same entities and equivalent relationships that may occur from different sources, therefore, the knowledge graph fusion and redundancy resolution in step S30 described above, specifically: S34: Calculate the cosine similarity between the node multi-modal embedding vectors of two different sources in the vector database. If the cosine similarity is higher than the similarity threshold, merge the node multi-modal embedding vectors of the two different sources.

[0062] Specifically, entity alignment is performed according to the semantic similarity between node multi-modal embedding vectors. By calculating the cosine similarity between the newly added node and the existing node vector in the graph, if the similarity is higher than the preset threshold (which can be 0.85-0.9), it is judged that the entities represented are the same or highly similar. At this time, node merging is performed: the original node ID is retained, and the new data is written as a new alias or additional attribute of the entity, thereby avoiding duplicate nodes.

[0063] Exemplarily, a new "Guizhou Kaili" node is identified, and the vector similarity with the existing "Kaili City" node is 0.92, indicating that the two refer to the same city, so they are merged into one node.

[0064] S35: Semantically normalize the key connection relationship of all triple units in the graph database.

[0065] It should be explained that the above step S35 is mainly used for semantic normalization of the key connection relationship. Different users may use different expressions to describe the same relationship, such as "located" and "located" semantics. These equivalent relationships are judged by a rule library or a vector representation and are unified as a graph specification relationship. For the same pair of entities, only one semantic repeated relationship is retained. Through the above steps S34-S35, the above process is equivalent to injecting semantic similar connections in the knowledge graph, and the similar nodes calculated by the node multi-modal embedding vector are connected by edges, so as to correct the redundancy and divergence of the original graph. Through the above technical solution, the vector similarity is used as an additional connection to improve the accuracy and reasoning ability of the graph representation.

[0066] S40: input the real-time contribution data of the target tourist group into the preset dynamic evolution mechanism to drive the high-dimensional index framework to update dynamically and obtain a dynamic index framework.

[0067] The above step S40 includes: S41: define the data entry conditions of the real-time contribution data.

[0068] It should be explained that the evolution of the knowledge graph is driven and evolved by the participation of a large number of users. The various interactive behaviors of tourists or community users on the platform are regarded as real-time contribution data of the knowledge graph by the method of the embodiment. These real-time contribution data include but are not limited to: (1) new tourism content on the platform, such as travel guides, travel notes, photos and videos, etc., which are used to provide new knowledge sources for the graph.

[0069] (2) entities or relationships in the graph, such as correcting scenic spot information, supplementing cultural anecdote descriptions.

[0070] (3) evaluation and feedback on existing knowledge content, such as likes, dislikes, collections of certain knowledge, or comments on scenic spot descriptions.

[0071] (4) data traffic of referencing or sharing graph content related content.

[0072] (5) user-submitted error reports or correction suggestions pointing out omissions or errors in the graph. The real-time contribution data meeting the above conditions are recorded by the system and associated with the affected knowledge objects to form the basic data for group co-creation or correction of the knowledge graph.

[0073] S42: perform a credibility operation on the real-time contribution data meeting the data entry conditions to obtain a credibility.

[0074] Specifically, for real-time contribution data of user contribution, the system introduces a group consensus evaluation mechanism to determine the credibility of the data being officially adopted by the knowledge graph. In this embodiment, the calculation formula of the credibility is:

[0075] In the formula, is the credibility; is a weighting coefficient, which is adjusted according to system experience or reinforcement learning process, so that the weighted sum of the three is 1; is the group voting approval rate, that is, the ratio between the likes and dislikes of other users to the content. is the historical reputation value of the contributor user, which is mainly calculated according to the correctness and the number of times of being adopted of the content contributed by the user in the past. The user with high reputation value is more reliable in submission, so is high; is the data source quality score, which gives a quality weight according to the source type. Specifically, official authoritative data > verified UGC (user-generated content) > ordinary network crawled data. For example, the data source quality score from the official website is the highest, and the network crawler data without correction has a lower data source quality score.

[0076] Preferably, in the initial state, to emphasize the public consensus while taking into account user reputation and source reliability.

[0077] S43: If the credibility is higher than the credibility threshold, write the real-time contribution data into the knowledge graph database.

[0078] More specifically, reflects the degree to which the knowledge is recognized by the group. If exceeds the preset credibility threshold, it is considered that the knowledge has reached group consensus and the credibility is sufficient, and it is automatically confirmed to be included in the graph. Otherwise, the content is placed in the "verification pool" and does not affect the graph temporarily. The content in the "verification pool" is waiting for data or manual review.

[0079] S44: Fine-tune the multi-modal embedding vector of the node written into the knowledge graph database according to the real-time contribution data.

[0080] Specifically, for the new knowledge content (i.e. real-time contribution data) written through the above step S43, the method of the embodiment introduces a node embedding online updating mechanism at the graph level to realize the adaptive evolution of semantic representation. The specific fine-tuning algorithm is:

[0081] In the formula, is the new content vector after fine-tuning, is the original node multimodal embedding vector of the entity node; is the vector representation calculated by the new content (incremental knowledge), for example, a newly uploaded description text about a scenic spot, a new text vector of the scenic spot is obtained through BERT, or a newly uploaded scenic spot photo is obtained through CLIP image vector; is an adaptive learning rate, which can be dynamically adjusted with , for example, , the higher , the greater the value is taken to fully approach the new content.

[0082] The above algorithm embodies the strategy of high consensus reinforcement and low consensus weakening: When the group highly approves, that is, Close to 1, the new content vector occupies a large weight, so as to obviously pull the node semantic vector to move towards the new knowledge direction; if the content is controversial, the new content only makes a small amount of fine-tuning or even negligible. In this way, while continuously absorbing new knowledge, it can avoid excessive interference of noise on the existing semantics.

[0083] Exemplarily, the original embedding vector of a node "Miao New Year Festival" mainly comes from the past text description. When a large number of users upload live photos and videos and obtain a high number of likes, these visual contents will generate new vectors , which are integrated into by high weighting, so that the vector of "Miao New Year Festival" node is optimized towards the direction containing visual features. On the contrary, if a description without approval tries to update the node, the original vector will hardly be changed.

[0084] Further, considering that the embedding update of large-scale knowledge graph cannot be retrained globally every time, the present application adopts a small batch continuous learning strategy to optimize the model parameters regularly, specifically: Every fixed period, a batch of newly added triples and node vectors are accumulated as an incremental training set, and fine-tuning is performed on the basis of the original model. In the training process, the representation strategy of the original important knowledge or knowledge distillation is protected to prevent catastrophic forgetting.

[0085] After the incremental fine-tuning is completed, the embedding vectors of the relevant nodes in the graph are updated, and the consistency of the overall semantic space is re-evaluated. If it is found that some old embedding vectors are significantly deviated from the association with the newly added knowledge, local embedding vector retraining or neighborhood updating is triggered, which is similar to the embedding aggregator inducing the neighborhood information of the new node to embed the new entity without retraining the entire model. This continuously evolving embedding learning framework ensures that old knowledge is preserved and new knowledge is gradually integrated, and the vector space of the graph can evolve in synchronization with the latest knowledge content.

[0086] Compared with traditional static one-time training, continuous learning enables the model to adapt to knowledge increments, and endows the graph with certain self-organizing and self-evolution capabilities. Through the introduction of swarm intelligence and dynamic evolution mechanisms, the knowledge graph of the method of the embodiment can be continuously improved in the continuous contribution of users, and automatically adjusts its internal representation to adapt to knowledge changes, so as to always maintain high credibility and high semantic accuracy.

[0087] S50: In response to the triggering of the user retrieval instruction, input the user retrieval instruction and the user-authorized personal information into the preset intelligent application interface to call the dynamic indexing framework for retrieval to obtain a retrieval result.

[0088] In the embodiment, the user retrieval instruction includes a structured retrieval instruction and a semantic retrieval instruction. That is, the method of the embodiment supports both structured retrieval and vectorized semantic retrieval.

[0089] For structured retrieval, the method of the embodiment can use the above-mentioned graph database to perform accurate structured queries on knowledge. Users or applications can propose entity relationship-based queries through a graph database query language such as Cypher or SPARQL. For example, the query of a specific relationship path: "What are the non-heritage festival activities in Guizhou Qianzhong and Southeast?" can be translated into a graph query to locate the nodes in the Guizhou Qianzhong and Southeast region, traverse the relationship of the venue to find all festival entities, and return a list of nodes that meet the conditions. Based on the reasoning of the graph path: "What non-heritage skills are related to the Miao New Year Festival?" The graph database can search along the relationship edge from the "Miao New Year Festival" node to find associated performance and skill entities, such as "Miao Silver Jewelry Making" and "Lusheng Dance" nodes. Structured retrieval has the advantages of logical accuracy and interpretability, and is suitable for scenarios that require strict matching and rule-based reasoning.

[0090] For vectorized semantic retrieval, the method of the embodiment can support semantic similarity retrieval by means of the embedding vector representation of nodes in the graph and the vector database. For a natural language query or example content of a user, the query or example content is first converted into a corresponding vector representation, and then approximate nearest neighbor (ANN) retrieval is performed in the vector database to find nodes or content that are semantically similar to the query or example content. For example, a user inputs a picture of a Miao ethnic festival, the system extracts the image embedding, searches for the nearest neighbor in the vector database, and returns the knowledge nodes and corresponding scenic spots that are most semantically similar to the picture. The user asks, “What are the minority celebrations similar to the Miao New Year Festival?” At this time, the question is converted into an embedding, and other festival nodes that are close to the “Miao New Year Festival” node vector are found in the vector space, such as the “Dong New Year Festival” and the “Qiang New Year”, to recommend similar cultural festivals. Compared with structured retrieval, vectorized retrieval can go beyond literal matching and achieve fuzzy matching at the semantic level, and is very effective for natural language, pictures, audio, and other queries.

[0091] On the basis of the two basic retrieval methods described above, the method of the embodiment further proposes a structure-semantic hybrid retrieval strategy. Specifically, the related retrieval fusion scoring algorithm of the strategy is as follows:

[0092] In the formula, represents the matching degree based on the graph structure, for example, whether there is a matching relationship path, the number of matching results, and the like represents the similarity score based on the embedding, and the two are weighted , The comprehensive score is obtained by weighted summation. According to the different properties of the query task, the weights of the two retrievals need to be dynamically adjusted.

[0093] Specifically, when the query has a clear entity relationship constraint, the proportion of is increased to ensure that the results meet the accurate conditions. For example, problems involving numerical filtering or multi-hop relationship reasoning mainly rely on structured queries, and semantics only performs auxiliary sorting. When the query is open or described using natural language, the weight of is increased to let the semantic similarity dominate the result sorting, while ensuring basic type matching. For example, the user asks, “Recommend some Miao festivals with song and dance performances”, first filter out the nodes of the type festival (structural constraint), and then determine which festivals are semantically related to “song and dance performances” according to the embedding to sort and recommend.

[0094] Through the hybrid retrieval strategy described above, the user can obtain query results that meet both logical conditions and semantic relevance, thereby improving the search experience.

[0095] Further, the intelligent application interface includes a RESTful API and an SDK.

[0096] It should be explained that based on the above retrieval framework, the method of the embodiment provides a RESTful API interface and an SDK to facilitate various smart cultural tourism applications to call the knowledge graph to achieve intelligent retrieval and other functions. The following provides more specific exemplary explanations: For similar cultural content recommendation: the user inputs a cultural entity (such as a non-heritage project or a scenic spot), and the interface retrieves and returns other cultural nodes with high relevance from the knowledge graph, thereby achieving personalized recommendation. For example, other metal technology non-heritage projects can be recommended based on "X clan silverware making". Trajectory-based intelligent route planning: according to the past travel preferences of tourists (such as preference for natural scenery or folk activities), combined with the relationship and vector similarity of scenic spots provided by the knowledge graph, personalized tourism routes are planned. The intelligent application interface receives user preferences and the current location, retrieves and returns several recommended route node sequences from the knowledge graph. Multi-modal cultural tourism content retrieval: the user uploads a photo or a recording, and the API returns the corresponding scenic spot / activity name and introduction according to the graph matching, realizing functions such as "picture search" and "sound recognition", so that tourists can obtain background knowledge of what they see and hear. RAG retrieval enhancement generation service: for cultural tourism question answering or content generation scenarios, first find relevant knowledge from the knowledge graph using hybrid retrieval, then inject knowledge into a large language model to generate more accurate answers or descriptions. For example, a tourist asks "What is the origin of the Miao New Year Festival?", by retrieving the historical legend nodes and related chronological data of the Miao New Year Festival in the knowledge graph, the LLM is assisted to generate authoritative and culturally based explanations. The open API makes the knowledge graph built by the present application easily used by upper-layer applications, such as intelligent tour guide App, cultural heritage digital museum platform, tourism chat robot, etc. The design of these interfaces follows the REST specification, supports on-demand query and batch data acquisition, and considers caching and permission control, ensuring good performance and security while providing high-quality knowledge services. Through the steps S10-S50, the knowledge graph constructed by the method of the embodiment realizes multi-modal data fusion, group wisdom co-construction, dynamic self-evolution, and structure-semantic hybrid retrieval. Compared with the prior art, the method of the embodiment improves in many aspects such as data source, database construction method, updating mechanism, user participation, and intelligent retrieval. Among them, multi-modal data fusion makes the graph information more rich; group intelligence driven improves the breadth and quality of knowledge acquisition; dynamic graph updating maintains the timeliness of knowledge; user consensus mechanism guarantees the credibility of knowledge; structure-semantic hybrid retrieval improves the user query experience. These improvements solve the pain points of the prior art, give the knowledge graph stronger adaptability and intelligence, and greatly enhance the retrieval ability.

[0097] Embodiment Two As shown in Figure 2 The embodiment discloses a group intelligence driven multi-modal cultural and tourism knowledge graph retrieval system, which comprises a processor and a memory. The memory stores computer program instructions. When the computer program instructions are executed by the processor, the group intelligence driven multi-modal cultural and tourism knowledge graph retrieval method recorded in embodiment one is realized.

[0098] The system of the embodiment also comprises other components such as communication interfaces, which are well known to those skilled in the art, and their settings and functions are known in the art, so they will not be described here.

[0099] In the present application, the aforementioned memory can be any tangible medium containing or storing a program, which can be used or combined with an instruction execution system, device or instrument. For example, the computer readable storage medium can be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory RRAM (Resistive Random Access Memory), dynamic random access memory DRAM (Dynamic Random Access Memory), static random access memory SRAM (Static Random-Access Memory), enhanced dynamic random access memory EDRAM (Enhanced Dynamic Random Access Memory), high bandwidth memory HBM (High-Bandwidth Memory), hybrid memory cube HMC (Hybrid Memory Cube), etc., or any other medium that can be used to store the required information and can be accessed by an application, module or both. Any such computer storage medium can be part of the device or accessible or connectable to the device. Any application or module described in the present application can be implemented using computer readable / executable instructions that can be stored or otherwise held by such computer readable medium.

[0100] In the description of the present specification, the meaning of "a plurality of" is at least two, for example, two, three or more, and the like, unless explicitly specifically limited.

[0101] While the present specification has shown and described several embodiments of the present application, it is to be understood that the same are presented by way of example only and not limitation. As such, numerous changes, modifications and substitutions can be made by one having ordinary skill in the art without departing from the spirit and scope of the present application. It is to be understood that in some instances, several of the described embodiments can be combined.

Claims

1. A swarm intelligence driven multi-modal tourism knowledge graph retrieval method, characterized in that, The method comprises the following steps: Collecting multi-modal historical data of a target tourist group, wherein the multi-modal historical data includes text / speech data, image / video data, audio data, and trajectory positioning data; Using multiple pre-set artificial intelligence models to analyze the corresponding types of multi-modal historical data, and performing triple structure and alignment fusion on the multiple analysis results to obtain triple units and node multi-modal embedding vectors, respectively; Writing the triple units and the node multi-modal embedding vectors into a pre-set knowledge graph database, and performing knowledge graph fusion and redundancy elimination to obtain a high-dimensional index framework; Inputting real-time contribution data of the target tourist group into a pre-set dynamic evolution mechanism to drive the high-dimensional index framework to update dynamically and obtain a dynamic index framework; In response to the triggering of a user search instruction, inputting the user search instruction and the user's authorized personal information into a pre-set intelligent application interface to call the dynamic index framework for searching and obtaining a search result.

2. The swarm intelligence driven multi-modal tourism knowledge graph retrieval method according to claim 1, characterized in that, The artificial intelligence models include natural language processing models, computer vision models, audio feature extraction models, and time-space clustering models; using multiple pre-set artificial intelligence models to analyze the corresponding types of multi-modal historical data, specifically: Inputting the text / speech data into the natural language processing model for text feature recognition and analysis to obtain a first analysis result; Inputting the image / video data into the computer vision model for visual feature recognition and analysis to obtain a second analysis result; Inputting the audio data into the audio feature extraction model for music feature recognition and analysis to obtain a third analysis result; Inputting the trajectory positioning data into the time-space clustering model for trajectory feature recognition and analysis to obtain a fourth analysis result.

3. The swarm intelligence driven multi-modal tourism knowledge graph retrieval method according to claim 2, characterized in that, Performing alignment fusion on the multiple analysis results, specifically: Using a pre-set multi-modal alignment fusion model to align the first analysis result, the second analysis result, the third analysis result, and the fourth analysis result; Concatenating the vectors corresponding to the aligned first analysis result, the second analysis result, the third analysis result, and the fourth analysis result to obtain the node multi-modal embedding vector.

4. The swarm intelligence driven multi-modal tourism knowledge graph retrieval method according to claim 1, characterized in that, The triple unit includes entity name, key relationship, and cultural attribute.

5. The swarm intelligence driven multi-modal tourism knowledge graph retrieval method according to claim 4, characterized in that, The knowledge graph database includes a graph database and a vector database; writing the triple unit and the node multi-modal embedding vector into the pre-set knowledge graph database includes: Storing the triple unit in the graph database; Storing the node multi-modal embedding vector in the vector database; Associating the triple unit with the corresponding node multi-modal embedding vector through a node unique ID.

6. The swarm intelligence driven multi-modal tourism knowledge graph retrieval method according to claim 5, characterized in that, Performing knowledge graph fusion and redundancy elimination includes: Calculating the cosine similarity between two node multi-modal embedding vectors of different sources in the vector database; if the cosine similarity is higher than a similarity threshold, merging the two node multi-modal embedding vectors of different sources; Performing semantic normalization on the key relationships of all triple units in the graph database.

7. The swarm intelligence driven multi-modal tourism knowledge graph retrieval method according to claim 1, characterized in that, Input real-time contribution data of a target tourist group into a preset dynamic evolution mechanism to drive the high-dimensional index framework to dynamically update, including: Defining data entry conditions of the real-time contribution data; Performing credibility operation on the real-time contribution data meeting the data entry conditions to obtain a credibility; If the credibility is higher than a credibility threshold, writing the real-time contribution data into the knowledge graph database; According to the real-time contribution data, fine-tuning node multi-modal embedding vectors written into the knowledge graph database.

8. The swarm intelligence driven multi-modal tourism knowledge graph retrieval method according to claim 1, characterized in that, The user retrieval instruction includes a structured retrieval instruction and a semantic retrieval instruction.

9. The swarm intelligence driven multi-modal tourism knowledge graph retrieval method according to claim 1, characterized in that, The intelligent application interface includes a RESTful API and an SDK.

10. A swarm intelligence driven multi-modal tourism knowledge graph retrieval system, characterized in that, A processor and a memory are included, and the memory stores computer program instructions which, when executed by the processor, implement the group intelligence driven multi-modal cultural and tourism knowledge graph retrieval method of any one of claims 1-9.

Citation Information

Patent Citations

  • Cross-modal travel knowledge graph construction system based on large language model

    CN120542539A

  • Multi-modal data distributed retrieval method and system based on mapping knowledge domain and vector matching

    CN118551086A

  • Trout sentiment analysis response method and system based on multi-modal fusion and incremental learning

    CN120256563A

  • Policy knowledge graph construction method and system based on digital human interaction data analysis

    CN120471160A

  • Building elevator detection, diagnosis and decision-making method based on graph retrieval enhanced agent

    CN120929785A