Index selection method for cross-domain multi-dimensional query features
By building a hierarchical and navigable small-world ontology graph indexing algorithm and user feedback optimization, the redundancy and inefficiency of index structures in cross-domain queries are solved, and efficient and accurate knowledge retrieval and personalized recommendation are achieved.
Patent Information
- Application Number
- CN202510231477.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-11
AI Technical Summary
The traditional single index structure is difficult to cope with the diversity and complexity of cross-domain queries, resulting in system redundancy and inefficiency. How to achieve semantic correlation and dynamic mapping between indexes in different domains while maintaining the index structure in a streamlined manner is an urgent problem.
Through user query understanding and knowledge construction, query intent recognition, feature extraction and optimization, dynamically adjust the index structure, and build a hierarchical and navigable small world ontology graph indexing algorithm, and combine user feedback behavior optimization indexing strategies to provide personalized knowledge recommendation services.
It improves the efficiency and accuracy of knowledge retrieval, optimizes index performance and adaptability, enhances user experience, and meets the needs of different users.
Smart Images

Figure CN120296207A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and particularly to an index selection method for cross - domain multi - dimensional query features. Background Art
[0002] In the selection of cross - domain multi - dimensional query features index, there is a complex technical problem. When facing query requests from different domains, how to effectively select and combine appropriate index structures to meet the specific concerns and query requirements of each domain has become an urgent problem to be solved. Traditional single - index structures are difficult to cope with the diversity and complexity of cross - domain queries, and simply establishing independent indexes for each domain will lead to system redundancy and low efficiency. The key challenge lies in how to achieve semantic association and dynamic mapping between indexes of different domains while keeping the index structure concise. This requires an index selection mechanism that can be adaptively adjusted, taking into account both the multi - dimensional features of the query and the specific concerns of different domains. At the same time, the index structure also needs to have good scalability and navigability to support seamless linking and integration of cross - domain knowledge. In addition, how to optimize the index selection strategy in real - time in a dynamically changing query environment, balancing query efficiency and system resource consumption, is also a thorny problem. This requires the index structure to have the ability of self - organization and self - optimization, and be able to dynamically adjust its internal structure and association relationships according to changes in query patterns. Summary of the Invention
[0003] The present invention provides an index selection method for cross - domain multi - dimensional query features, mainly including:
[0004] According to the historical query behavior data of users, by classifying query topics and identifying query intents, judge the query domain and purpose of users, and obtain the domain label and intent label of the query;
[0005] For different query domain labels and intent labels, construct a domain ontology knowledge base, and generate an intent template library based on the analysis of large - scale query logs, perform semantic parsing and feature extraction on the query, and obtain multi - dimensional attribute information including keywords, entities, time, and location in the query;
[0006] According to the weight calculation of query terms and the semantic similarity between query terms, optimize the combination of the extracted query multi - dimensional attributes to form a query feature vector representation, and use query feature clustering to mine potential query patterns, and optimize the node layering and edge connection strength of the ontology knowledge base;
[0007] In the domain ontology knowledge base, a graph embedding representation learning method is adopted to map ontology concepts into a low-dimensional vector space. The correlation degree between different concepts is calculated through the semantic similarity between vectors, and the mapping weights of concepts are dynamically adjusted using query patterns to construct a hierarchical and navigable small-world ontology graph indexing algorithm. Among them, the tree-like topology of the hierarchical structure serves as the backbone of the small-world indexing algorithm network;
[0008] In the ontology graph index, the hierarchical division of concept nodes is optimized to meet the average shortest path characteristics of the small-world indexing algorithm network, control the hop count and query complexity of the navigation path, and dynamically adjust the edge weights and node connection patterns through the network mapping of user query feedback behavior;
[0009] According to the similarity between the user query feature vector and the ontology concept vector, the most relevant ontology subgraph is dynamically selected as the candidate index. If the number of candidate indexes exceeds the threshold, the candidate index range is further narrowed according to the domain label and pattern of the query;
[0010] In the candidate index subgraph, starting from the user query feature, a breadth-first search strategy based on semantic similarity is adopted to discover the concept nodes related to the query and their semantic relationships, forming cross-concept knowledge links. During the search process, the importance of concepts in the hierarchical structure and the small-world indexing algorithm network is comprehensively considered, nodes with a high degree of match with the query intent are preferentially expanded, and the navigation path between nodes is recorded;
[0011] The navigation path of the knowledge link is optimized, comprehensively considering the semantic relevance between nodes and the characteristics of the small-world indexing algorithm, to recommend the optimal knowledge navigation path for the user, and dynamically adjust the optimization strategy according to the user's query history behavior and feedback;
[0012] The knowledge navigation results are presented through a visual interaction method, providing multi-dimensional attribute filtering and sorting functions to support users' exploratory browsing and in-depth mining. Then, according to the user's query scenario and preferences, a personalized knowledge recommendation service is provided.
[0013] The technical solution provided by the embodiment of the present invention may include the following beneficial effects:
[0014] The present invention discloses an index selection method for cross - domain multi - dimensional query features. Through user query understanding, knowledge construction and extraction, it can accurately understand the user's query domain and purpose, extract multi - dimensional attribute information of the query, taking into account both the multi - dimensional features during query and the specific concerns of different domains. At the same time, through index construction and optimization, query optimization and pattern mining, it not only improves the accuracy and efficiency of the query, but also optimizes the structure and connection strength of the ontology knowledge base, enhances the performance and adaptability of the index, and supports seamless link and integration of cross - domain knowledge. In addition, by dynamically adjusting the hierarchical division, edge weights and node connection patterns of the index according to user feedback and query behavior, it forms knowledge links and searches, and optimizes the navigation path, making the index more in line with user needs, realizing dynamic adjustment of strategies according to user behavior, enhancing the user experience, and meeting the needs of different users.
[0015] Generally speaking, through efficient query intention recognition, precise feature extraction and optimization, dynamic ontology knowledge base management and optimization, as well as efficient knowledge navigation and personalized recommendation, the present invention significantly improves the efficiency and accuracy of knowledge retrieval, not only optimizing the knowledge management and retrieval process, but also enhancing the interactivity and satisfaction between users and the knowledge base. Brief Description of the Drawings
[0016] Figure 1 It is a flowchart of an index selection method for cross - domain multi - dimensional query features of the present invention.
[0017] Figure 2 It is a schematic diagram of an index selection method for cross - domain multi - dimensional query features of the present invention.
[0018] Figure 3 It is another schematic diagram of an index selection method for cross - domain multi - dimensional query features of the present invention. Detailed Embodiment
[0019] Next, the technical solutions of the present invention will be described clearly and completely in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work belong to the scope of protection of the present invention.
[0020] As Figures 1-3 , a specific index selection method for cross - domain multi - dimensional query features in this embodiment may specifically include:
[0021] S101. According to the historical query behavior data of the user, through query topic classification and query intention recognition, judge the user's query domain and purpose, and obtain the domain label and intention label of the query.
[0022] Obtain the user's historical query records, extract the query text, timestamp, and clicked link from the user's historical query records; perform word segmentation on the query text to obtain a set of query keywords. Conduct text clustering on a large amount of historical query data, extract keywords for each category from the text clustering to form a topic classification dictionary; according to the topic classification dictionary, perform topic vectorization on the set of query keywords, calculate the similarity between the query text and each topic category using cosine similarity, and select the category with the highest similarity as the query domain label. Use a dependency syntax analysis tree to extract the subject-verb-object relationship in the query text, and combine named entity recognition to identify entities and attributes in the query text; extract the semantic features of the query based on the subject-verb-object relationship, entities, and attributes; based on the semantic features, use a support vector machine to judge the intent type of the query to obtain a query intent label. Divide the user's historical query data according to a fixed time window, calculate the occurrence frequency of each query domain and intent within each time window; construct a user interest vector based on the occurrence frequency. Adopt a sliding time window method, combine the query domain label and query intent label, and calculate the cosine similarity with the user interest vector; if the cosine similarity is higher than a preset threshold, it is determined as the user's main query domain and purpose.
[0023] Exemplarily, user behavior data such as query text, timestamp, and clicked link are extracted from the user's historical query records, and the query text is segmented to obtain a set of query keywords. Through text clustering of a large amount of historical query data, keywords of various categories are extracted to form a topic classification dictionary. According to the topic classification dictionary, the query keyword set is vectorized for topics, and the cosine similarity is used to calculate the similarity between the query text and each topic category. The category with the highest similarity is selected as the query domain label. The subject-verb-object relationship in the query is extracted using a dependency syntax analysis tree, and entities and attributes in the query are identified in combination with named entity recognition to extract the semantic features of the query. Based on the extracted semantic features, a support vector machine is used to judge the intent type of the query to obtain a query intent label. The user's historical query data is divided by a fixed time window, and the occurrence frequencies of each query domain and intent within each time window are calculated to construct a user interest vector. The sliding time window method is adopted, and in combination with the domain label and intent label of the current query, the cosine similarity with the historical interest vector is calculated. According to the calculated similarity value, a threshold is set to judge the query domain and purpose of the user. If the similarity is higher than the threshold, it is determined as the user's main query domain and purpose. Data is extracted from the user's historical query records, including query text such as smartphone recommendation, a timestamp of 2024-07-15 10:30:25, and a clicked link of www.example.com / smartphones. The query text is segmented to obtain keywords such as smart, phone, and recommendation. Through text clustering of 1 million historical query data, 20 topic categories such as electronic products, household items, and tourist attractions are extracted, with each category containing 100 keywords, forming a topic classification dictionary. The query keywords and topic categories are vectorized, and the cosine similarity is calculated. For example, the similarity between smartphone recommendation and the electronic product category is 0.85, which is the highest. Therefore, the query domain label is determined as electronic products. The subject-verb-object relationship in the query is extracted using a dependency syntax analysis tree. For example, recommendation is the predicate and phone is the object. Through named entity recognition, it is identified that smartphone is a product entity. Based on the extracted semantic features, a support vector machine is used to judge the intent type of the query. For example, the query is judged as a product recommendation intent. The user's historical query data for the past 30 days is divided by a 1-day window, and the occurrence frequencies of each query domain and intent within each window are calculated. For example, the occurrence frequency of the electronic product domain on a certain day is 0.6, and the occurrence frequency of the product recommendation intent is 0.4, to construct a user interest vector. The 7-day sliding time window is adopted, and in combination with the current query's electronic product, domain label, and product recommendation intent label, the cosine similarity with the historical interest vector is calculated, obtaining a similarity value of 0.75. A judgment threshold of 0.7 is set. Since 0.75 is greater than 0.7, it is determined that electronic products and product recommendation are the user's main query domain and purpose.
[0024] S102. For different query domain tags and intent tags, construct a domain ontology knowledge base, and generate an intent template library based on the analysis of a large-scale query log. Perform semantic parsing and feature extraction on the query to obtain multi-dimensional attribute information including keywords, entities, time, and location in the query.
[0025] According to the preset query domain tags and intent tags, extract relevant entities, attributes, and relationships from the knowledge graph to construct the domain ontology knowledge base. At the same time, use the Apriori algorithm for frequent pattern mining based on the query log to generate the intent template library. Segment the input query text, perform part-of-speech tagging using the hidden Markov model to obtain the tagged word sequence. Identify the named entities in the tagged word sequence through the conditional random field model, extract the keywords, entities, time, and location information contained in the query text, and obtain the feature vector of the query. Use dependency syntax analysis to obtain the syntactic structure of the query text, and combine the entity relationships in the domain ontology knowledge base to perform semantic parsing on the query text to identify the semantic roles and relationships in the query text. Calculate the cosine similarity between the feature vector of the query and the templates in the intent template library, and select the template with the highest similarity as the query intent. If the time and location attributes are identified, use the entity types and attribute definitions in the ontology knowledge base to map the identified entities and keywords to the corresponding attribute dimensions to obtain the multi-dimensional attribute information including keywords, entities, time, and location in the query text.
[0026] Exemplarily, according to the preset query domain tags and intent tags, relevant entities, attributes, and relationships are extracted from DBpedia or a self-built domain knowledge graph to construct a domain ontology knowledge base. At the same time, the Apriori algorithm is used for frequent pattern mining based on a large-scale query log to generate a corresponding intent template library. The input query text is segmented, and the Hidden Markov Model is used for part-of-speech tagging. Named entities are identified through the Conditional Random Field Model, and keywords, entities, time, and location information included in the query are extracted to generate a feature vector of the query. Dependency syntactic analysis is used to obtain the syntactic structure of the query, and combined with the entity relationships in the domain ontology knowledge base, semantic parsing of the query is performed to identify semantic roles and relationships in the query. The cosine similarity between the feature vector of the query and the templates in the intent template library is calculated, and the template with the highest similarity is selected as the query intent. Combining attributes such as time and location identified by the Conditional Random Field Model, and using the entity types and attribute definitions in the ontology knowledge base, the identified entities and keywords are mapped to the corresponding attribute dimensions to obtain multi-dimensional attribute information including keywords, entities, time, and location in the query. 10,000 entities, 100 attributes, and 500 relationships in the electronics product domain are extracted from DBpedia to construct a domain ontology knowledge base. The Apriori algorithm is used for frequent pattern mining on 1 million query logs, with the minimum support set to 0.1%, generating 1,000 intent templates. The input query "recommend smartphones with high cost performance" is segmented to obtain words such as "recommend", "cost performance", "high", and "smartphone". The Hidden Markov Model is used for part-of-speech tagging, and the tagging results are "recommend / v", "cost performance / n", "high / a", "smartphone / n". Named entities are identified through the Conditional Random Field Model, and "smartphone" is identified as a product entity. A query feature vector with a dimension of 100 is generated, including information such as word frequency, part of speech, and entity type. Dependency syntactic analysis is used to obtain the query structure, determining that "recommend" is the core predicate and "smartphone" is the object. Combining the attribute definition of "smartphone" in the ontology knowledge base, "cost performance" is identified as a product attribute. The cosine similarity between the query feature vector and the 1,000 templates in the intent template library is calculated, and the template with the highest similarity of 0.85, "[recommend][product][attribute]", is selected as the query intent. Finally, the identified entity "smartphone" is mapped to the product dimension, and "cost performance" is mapped to the attribute dimension to form multi-dimensional attribute information: product - smartphone, attribute - cost performance, attribute value - high.
[0027] S103. According to the weight calculation of the query terms and the semantic similarity between the query terms, the extracted query multi-dimensional attributes are combined and optimized to form a query feature vector representation, and potential query patterns are mined through query feature clustering to optimize the node layering and edge connection strength of the ontology knowledge base.
[0028] Calculate the weight of the query terms according to the TF-IDF algorithm, where TF is the number of occurrences of a term in the current query, and IDF is the inverse document frequency of the term in the entire query set; measure the semantic similarity between query terms through cosine similarity to obtain the query feature vector representation. Use the k-means++ algorithm to select the initial cluster centers. If the selection is completed, use the k-means algorithm to cluster the query feature vectors, and set the number of clusters to k; iteratively calculate the cluster centers until the cluster centers are stable or the maximum number of iterations is reached to obtain k query pattern clusters. For the k query pattern clusters, set the support threshold; calculate the occurrence frequency of the attribute combination in the query pattern cluster. If the occurrence frequency exceeds the support threshold, it is determined as a frequent attribute combination. Extract the frequent attribute combinations to construct potential query patterns; update the attribute set of the corresponding concept nodes in the ontology knowledge base according to the potential query patterns. Calculate the similarity between concept nodes based on the common attributes of the samples within the query pattern cluster; determine whether the similarity is higher than the preset threshold. If it is higher than the preset threshold, adjust the concept nodes to a parent-child relationship or a sibling relationship; update the hierarchical relationship between concept nodes in the ontology knowledge base according to the adjusted relationship.
[0029] Exemplarily, the weights of query terms are calculated according to the TF-IDF algorithm, where TF is the number of occurrences of a term in the current query, and IDF is the inverse document frequency of the term in the entire query set. The semantic similarity between query terms is measured by cosine similarity, and the extracted multi-dimensional attributes of the query are combined and optimized to form a query feature vector representation. The k-means++ algorithm is used to select the initial class centers, and then the k-means algorithm is used to cluster the query feature vectors. The number of clusters is set to k, and the class centers are iteratively calculated until the class centers are stable or the maximum number of iterations is reached, obtaining k query pattern clusters. For each query pattern cluster, a support threshold is set, and the occurrence frequency of the attribute combination in the query pattern cluster is calculated. If it exceeds the threshold, it is considered a frequent attribute combination, and these frequent attribute combinations are extracted to construct potential query patterns, and the attribute sets of the corresponding concept nodes in the ontology knowledge base are updated. Based on the common attributes of the samples within the query pattern cluster, the similarity between concept nodes is calculated, and nodes with high similarity are adjusted to a parent-child relationship or a sibling relationship, thereby adjusting the hierarchical relationship between concept nodes in the ontology knowledge base. The co-occurrence frequency between nodes is calculated, and the connection strength of the edges is updated. For querying the ranking of the cost performance of smartphones, the TF-IDF algorithm is used to calculate the term weights. The TF of smartphone is 1, the IDF is 0.5, and the weight is 0.5; the TF of cost performance is 1, the IDF is 0.8, and the weight is 0.8; the TF of ranking is 1, the IDF is 0.6, and the weight is 0.6. The word vectors are calculated by the Word2Vec model, and the cosine similarity is used to measure the semantic similarity between words. The similarity between smartphone and cost performance is 0.7, and the similarity between cost performance and ranking is 0.6. After combined optimization, a 100-dimensional query feature vector is formed. The k-means++ is used to select 5 initial class centers, and the k-means algorithm is applied to cluster 10 million query feature vectors. After 50 iterations, 5 query pattern clusters are obtained. In the electronic product evaluation cluster, the support threshold is set to 0.3, and it is found that the occurrence frequency of the brand-price-performance attribute combination is 0.35, exceeding the threshold, and it is extracted as a frequent attribute combination. The attribute set of the smartphone node in the ontology knowledge base is updated, and the cost performance attribute is added. The similarity between the smartphone and tablet computer nodes is calculated to be 0.8 and adjusted to a sibling relationship. The co-occurrence frequency between the smartphone and electronic product nodes is 0.9, and the connection strength is updated to 0.9.
[0030] S104. In the domain ontology knowledge base, the graph embedding representation learning method is adopted to map ontology concepts into a low-dimensional vector space, calculate the association degree between different concepts through the semantic similarity between vectors, and dynamically adjust the mapping weights of concepts using query patterns to construct a hierarchical and navigable small-world ontology graph indexing algorithm. Among them, the tree-like topology of the hierarchical structure serves as the backbone of the small-world indexing algorithm network.
[0031] Use the Word2Vec algorithm to perform graph embedding representation learning on the domain ontology knowledge base to obtain the low-dimensional vector representation of ontology concepts in a 128-dimensional vector space; calculate the cosine similarity between different concept vectors according to the low-dimensional vector representation, and if the cosine similarity is greater than the preset similarity threshold, establish a semantic association edge between the two concepts; use the PageRank algorithm to calculate the importance score of each concept node, and the importance score is used to construct a tree-like topological structure as the backbone of the index network; obtain the user query log, count the occurrence frequency of each concept in different query patterns, calculate the conditional probability of the concept in each query pattern according to the occurrence frequency, and the conditional probability is used as the weight of this concept in the query pattern; based on the tree-like topological structure, use the Kleinberg algorithm to add long-range connections, and the establishment probability of the long-range connections is proportional to the semantic similarity and importance score between nodes, to obtain a hierarchical and navigable small-world ontology graph index.
[0032] Exemplarily, the Word2Vec algorithm is used to perform graph embedding representation learning on the domain ontology knowledge base, mapping ontology concepts into a 128-dimensional vector space to obtain the low-dimensional vector representation of each concept node. The semantic similarity between different concept vectors is calculated through cosine similarity. A similarity threshold of 0.8 is set. If the similarity between two concepts is greater than the threshold, a semantic association edge is established between the two concepts, and these edges will be used for the subsequent establishment of long-range connections to improve the navigation efficiency of the small-world network. Using the PageRank algorithm, the concept co-occurrence relationship is regarded as a directed graph, and the PageRank value of each concept node is calculated as its importance score to construct a tree-like topological structure as the backbone of the index network. According to the user query log, the occurrence frequency of each concept in different query modes is counted, and the conditional probability of the concept in each query mode is calculated as the weight of the concept in this query mode, and these weights are updated regularly to achieve dynamic adjustment. Based on the tree-like topology, the Kleinberg algorithm is used to add long-range connections, where the establishment probability of long-range connections is proportional to the semantic similarity and importance score between nodes, forming a hierarchical and navigable small-world ontology graph index. In the ontology knowledge base in the e-commerce field, the Word2Vec algorithm is used to perform graph embedding representation learning on 1 million concept nodes, mapping each concept into a 128-dimensional vector space. For example, the vector representation of the smartphone concept is [-0.23, 0.45,..., 0.12]. The cosine similarity between concepts is calculated, and the similarity between smartphone and tablet computer is 0.85, exceeding the set threshold of 0.8, so a semantic association edge is established. The PageRank algorithm is used to calculate the concept importance. After 50 iterations, the PageRank value of the smartphone is 0.0075, and a tree-like topological structure is constructed accordingly. By analyzing 100 million user query logs, it is found that the conditional probability of the smartphone in the price comparison query mode is 0.35, which is used as the weight of this concept in this query mode. The weights are updated once a week to achieve dynamic adjustment. Based on the tree-like topology, the Kleinberg algorithm is used to add long-range connections, setting parameters p = 2 and q = 3 to generate a small-world network. For example, the probability of establishing a long-range connection between the smartphone node and the processor node is 0.015, and the probability value is determined by the product of the semantic similarity of 0.75 and the importance scores of 0.0075 and 0.0025 of the two nodes. The finally formed small-world ontology graph index contains 1 million nodes and 5 million edges, with an average path length of 6.2 and a clustering coefficient of 0.4, reflecting the characteristics of the small-world network.
[0033] S105. In the ontology graph index, optimize the hierarchical division of concept nodes to make it meet the average shortest path characteristics of the small-world index algorithm network, control the number of hops of the navigation path and the query complexity, and dynamically adjust the edge weights and node connection modes through the network mapping of user query feedback behavior.
[0034] The AGNES agglomerative hierarchical clustering algorithm is used to hierarchically partition the concept nodes in the ontology graph index to obtain an initial hierarchical structure. The average shortest path length of the ontology graph index is calculated according to the initial hierarchical structure. If the average shortest path length exceeds a preset threshold, cross-layer connection edges are added based on the semantic similarity and hierarchical difference between nodes until the small-world network characteristics are satisfied. For the graph structure that satisfies the small-world network characteristics, a maximum hop count threshold is set, the paths exceeding the maximum hop count threshold are pruned, and the A* search algorithm is used to optimize the query path. The user query logs are obtained, the access frequency between nodes is statistically counted according to the user query logs, and the access frequency is used as the initial weight of the edge, and a weighted graph structure is constructed based on the optimized network structure. The Q-learning algorithm is used to process the weighted graph structure, and the user query feedback behavior is mapped into a reward signal according to the query success rate and query time, and the edge weights and node connection modes are dynamically adjusted through the reward signal to achieve adaptive optimization of the hierarchical structure.
[0035] Exemplarily, the AGNES agglomerative hierarchical clustering algorithm is used to hierarchically partition the concept nodes in the ontology graph index, a clustering threshold is set to control the number of nodes in each layer, and an initial hierarchical structure is obtained. The average shortest path length of the ontology graph index is calculated. If the length exceeds the preset threshold, based on the semantic similarity and hierarchical difference between nodes, node pairs with high similarity and large hierarchical difference are selected to add cross-layer connection edges until the small-world network characteristics are satisfied. A maximum hop threshold is set to prune the paths exceeding the threshold, and the A* search algorithm is used to optimize the query path to control the hop count and query complexity of the navigation path. According to the user query logs, the access frequency between nodes is statistically analyzed, and the frequency is used as the initial weight of the edge. A weighted graph structure is constructed based on the optimized network structure. Using the Q-learning algorithm, the user query feedback behavior is mapped to a reward signal according to the query success rate and query time. Paths with high success rate and short query time obtain high rewards, and the edge weights and node connection patterns are dynamically adjusted to achieve the adaptive optimization of the hierarchical structure. In the ontology graph index in the e-commerce field, the AGNES agglomerative hierarchical clustering algorithm is used to hierarchically partition 1 million concept nodes, the clustering threshold is set to 0.6, and the number of nodes in each layer is controlled not to exceed 10,000, obtaining a 5-layer initial hierarchical structure. The average shortest path length of the ontology graph index is calculated to be 8.5, exceeding the preset threshold of 6. Therefore, cross-layer connection edges are added based on the semantic similarity and hierarchical difference between nodes. Node pairs with similarity greater than 0.8 and hierarchical difference greater than 2 are selected, such as smartphones and processors, and cross-layer connection edges are added until the average shortest path length drops to 5.8, satisfying the small-world network characteristics. The maximum hop threshold is set to 10, paths exceeding the threshold are pruned, and the A* search algorithm is used to optimize the query path, reducing the average query time from 200 ms to 150 ms. 100 million user query logs are analyzed, the access frequency between nodes is statistically analyzed, and the frequency is normalized and used as the initial weight of the edge. A weighted graph structure is constructed based on the optimized network structure. Using the Q-learning algorithm, the learning rate α = 0.1 and the discount factor γ = 0.9 are set, and the user query feedback behavior is mapped to a reward signal. The reward increases by 1 point for every 10% increase in the query success rate and 0.5 points for every 10 ms reduction in the query time. After 10,000 iterations of learning, the edge weights and node connection patterns are dynamically adjusted, increasing the average query success rate from 85% to 92% and further reducing the average query time to 120 ms, achieving the adaptive optimization of the hierarchical structure.
[0036] S106. Dynamically select the most relevant ontology subgraph as the candidate index according to the similarity between the user query feature vector and the ontology concept vector. If the number of candidate indexes exceeds the threshold, further narrow the candidate index range according to the domain label and pattern of the query.
[0037] Obtain the user query, and convert the user query into a feature vector according to the pre-established Word2Vec model; at the same time, obtain the preset ontology concept vector representation. Calculate the similarity between the user query feature vector and the ontology concept vector using cosine similarity; if the similarity is greater than the preset threshold, determine the corresponding ontology concept as the seed node. For the seed node, expand three layers deep outward through the breadth-first search algorithm, and extract the subgraph containing the seed node; use the subgraph as the candidate index. Use a support vector machine classifier to perform domain classification on the user query; obtain the domain label of the user query from the output of the support vector machine classifier. Combine the domain label and the query pattern in the pre-stored historical query log, and use a weighted scoring method to filter the candidate index; calculate the comprehensive score according to the domain relevance and query pattern matching degree; determine the subgraph with the highest score as the final candidate index.
[0038] Exemplarily, the user query is converted into a feature vector through the Word2Vec model, and at the same time, the vector representations of ontology concepts are obtained. The cosine similarity is used to calculate the similarity between the user query vector and the ontology concept vector. According to the similarity ranking result, the ontology concepts with similarity greater than the preset threshold are selected as seed nodes. Centered on the seed nodes, the breadth-first search algorithm is used to expand three layers deep outward, and the subgraph containing these nodes is extracted as the candidate index. According to the complexity of the current query and the system load, the similarity threshold is dynamically adjusted to control the number of candidate indexes. The number of nodes contained in the extracted subgraph is counted. If the number of candidate indexes exceeds the preset threshold, the support vector machine classifier is used to perform domain classification on the user query to obtain the domain label of the query. Combining the query domain label and the query patterns in the historical query log, the weighted scoring method is used to filter and reorder the candidate indexes, and the comprehensive score is calculated according to the domain relevance and query pattern matching degree, and the subgraph with the highest score is retained as the final candidate index. In the e-commerce field, the pre-trained Word2Vec model is used to convert the user query "smartphone cost performance" into a 300-dimensional feature vector [-0.2, 0.5,..., 0.3], and at the same time, the vector representations of ontology concepts such as "mobile phone", "price", and "performance" are obtained. Through cosine similarity calculation, the similarity between the query vector and the "mobile phone" concept is 0.85, the "price" is 0.72, and the "performance" is 0.68. The initial similarity threshold is set to 0.7, and "mobile phone" and "price" are selected as seed nodes. Centered on these two nodes, the breadth-first search algorithm is used to expand 3 layers deep outward, and the subgraph containing 200 nodes is extracted as the candidate index. The current system load is 65%, and the query complexity score is 0.8. Accordingly, the similarity threshold is dynamically adjusted to 0.75, and the seed nodes are re-screened to obtain a candidate index containing 150 nodes. Since 150 exceeds the preset threshold of 100, the support vector machine classifier is used to perform domain classification on the query, and the domain label of "electronic product" is obtained with a confidence of 0.92. Combining the query patterns in the historical query log, it is found that the price comparison pattern has a matching degree of 0.88 with the current query. Using the weighted scoring method, the domain relevance weight is set to 0.6, and the query pattern matching degree weight is set to 0.4 to calculate the comprehensive score of each candidate index. Finally, the subgraph with the highest score and containing 120 nodes is retained as the final candidate index, and this index covers key concept nodes such as "smartphone", "price", "performance", and "brand".
[0039] S107. In the candidate index subgraph, starting from the user query features, a breadth-first search strategy based on semantic similarity is adopted to discover the concept nodes related to the query and their semantic relationships, forming cross-concept knowledge links. During the search process, the importance of the concept in the hierarchical structure and the small-world index algorithm network is comprehensively considered, nodes with a high matching degree with the query intention are preferentially expanded, and the navigation paths between the nodes are recorded.
[0040] Obtain the user query feature vector, calculate the semantic similarity between the user query feature vector and all concept nodes in the candidate index subgraph, and select the top three nodes with the highest similarity as the search starting points according to the semantic similarity; adopt the breadth-first search algorithm based on the priority queue to expand the search starting points, and calculate the matching degree between the nodes to be expanded and the query intention; according to the matching degree, use the PageRank algorithm to calculate the importance score of the nodes in the small-world index network, and combine the hierarchical information of the nodes in the hierarchical structure to obtain the comprehensive importance score of the nodes through weighted average; determine the expansion priority of the nodes according to the comprehensive importance score, traverse the adjacent nodes layer by layer, record the semantic relationships between the nodes, and the semantic relationships include types such as "is a", "belongs to", and "contains", and store the semantic relationships in the form of triples; if the formed knowledge links reach the preset node quantity threshold, then use the Kruskal minimum spanning tree algorithm to prune and optimize the knowledge links, remove redundant nodes and relationships, and obtain the final cross-concept knowledge links.
[0041] Exemplarily, taking the user query feature vector as the starting node, calculate its semantic similarity with all concept nodes in the candidate index subgraph, and select the top three nodes with the highest similarity as the search starting points. If the similarities of multiple nodes are similar, comprehensively consider the degree centrality of the nodes and select the node with the highest degree centrality. Adopt the breadth-first search algorithm based on the priority queue, and calculate the matching degree between the node to be expanded and the query intention during each expansion. Use the PageRank algorithm to calculate the importance score of the node in the small-world index network, combine the hierarchical information of the node in the hierarchical structure, and calculate the comprehensive importance score of the node in a weighted average manner to determine the expansion priority of the node. According to the expansion priority, traverse the adjacent nodes layer by layer, record the semantic relationships between the nodes, including types such as "is a", "belongs to", and "contains", and store them in the form of triples. At the same time, record the navigation path. If the similarity between the current node and the query feature is higher than the preset threshold, add it to the knowledge link. Set the maximum search depth and the threshold for the number of knowledge link nodes. When the threshold is reached or the search is completed, use the Kruskal minimum spanning tree algorithm to prune and optimize the formed knowledge link, remove redundant nodes and relationships, and retain the most important connections to obtain the final cross-concept knowledge link. In the e-commerce field, for the user query of high-cost-performance smartphones, calculate the cosine similarity between its feature vector and 1000 concept nodes in the candidate index subgraph. Select the three nodes with the highest similarity: smartphone with 0.95, cost performance with 0.88, and price with 0.85 as the search starting points. Adopt the breadth-first search algorithm based on the priority queue, and the initial queue contains these three nodes. Use the PageRank algorithm to calculate the node importance. After 50 iterations, the PageRank value of the smartphone node is 0.025. Combine the hierarchical information of the node in the 4-layer structure (the smartphone is located in the 2nd layer with a weight of 0.8), and calculate the comprehensive importance score of the node in the way of 0.6PageRank + 0.4 hierarchical weight to be 0.037. According to the matching degree of 0.75 between the node and the query intention and the comprehensive importance score, determine the expansion priority of the smartphone node to be 0.028. Traverse the adjacent nodes and record the semantic relationships as triples such as "smartphone, is a, electronic product", "cost performance, belongs to, smartphone", etc. Set the maximum search depth to 5 and the threshold for the number of knowledge link nodes to 50. After the search is completed, obtain an initial knowledge link containing 65 nodes. Use the Kruskal algorithm for pruning, and finally retain 48 nodes and 52 edges to form an optimized cross-concept knowledge link, including key concepts such as smartphones, cost performance, processors, brands, and their relationships.
[0042] S108. Optimize the navigation path of the knowledge link, comprehensively consider the semantic relevance between nodes and the characteristics of the small-world index algorithm, recommend the optimal knowledge navigation path for the user, and dynamically adjust the optimization strategy according to the user's query history behavior and feedback.
[0043] Calculate the semantic relevance score for node pairs in the knowledge link, and construct a weighted graph structure based on the semantic relevance score, degree centrality, clustering coefficient, and average path length of nodes in the small-world index algorithm; wherein, the weighted graph structure uses the weighted sum of the semantic relevance score, degree centrality, clustering coefficient, and average path length as the comprehensive weight of the edge. According to the weighted graph structure, use the multi-objective A* algorithm to calculate the set of optimal navigation paths from the starting node to multiple target nodes to obtain the initial recommended path; the initial recommended path consists of multiple nodes. Obtain the user interest characteristics from the user's historical query records, use the LDA topic model to extract the topics and preferences that the user is interested in, and assign personalized weights to the nodes in the initial recommended path according to the user interest characteristics. Perform weighted fusion on the personalized weight and the weight of the initial recommended path to obtain the fusion weight; the fusion weight is determined by the weighted sum of the personalized weight and the weight of the initial recommended path. Use the Q-learning algorithm to process the user's click behavior and stay time, and set the reward function; wherein, the reward function is determined by the sum of the product of the click weight and whether to click and the product of the stay time weight and the normalized stay time; update the Q-value table according to the reward function, and update the weights of the nodes and edges according to the Q-value table to optimize the navigation path personalized.
[0044] Exemplarily, calculate the semantic relevance score for node pairs in the knowledge link, and combine the degree centrality, clustering coefficient, and average path length of nodes in the small-world index algorithm to construct a weighted graph structure. Calculate the comprehensive weight of edges in the way of 0.6 semantic relevance + 0.2 degree centrality + 0.1 clustering coefficient + 0.1 average path length. Use the multi-objective A* algorithm to calculate the set of optimal navigation paths from the starting node to multiple target nodes on the weighted graph to obtain the initial recommended paths. According to the user's historical query records, use the LDA topic model to extract the topics and preferences that the user is interested in, and assign personalized weights to the nodes in the recommended paths. Perform weighted fusion on the personalized weights and the initial path weights to obtain the personalized weight = 0.7 * initial path weight + 0.3 * user preference weight. Use the Q-learning algorithm, take the user's click behavior and dwell time as feedback signals, set the reward function = click weight * click behavior + dwell time weight * dwell time, and dynamically adjust the path optimization strategy. Update the Q-value table, and update the weights of nodes and edges according to the new Q-values to achieve continuous optimization of the personalized navigation path. In the knowledge link in the e-commerce field, the calculated semantic relevance score for the pair of nodes "smartphone" and "cost performance" is 0.85. Combining the characteristics of the small-world index algorithm, the degree centrality of the "smartphone" node is 0.72, the clustering coefficient is 0.68, and the average path length is 3.2. Apply the weighted formula 0.6×0.85 + 0.2×0.72 + 0.1×0.68 + 0.1×(1 / 3.2), and the obtained comprehensive weight of the edge is 0.766. Use the multi-objective A algorithm to calculate the set of optimal navigation paths from "smartphone" to the three target nodes of "brand", "price", and "performance" to obtain the initial recommended paths. Analyze the user's nearly 100 query records, and use the LDA topic model to extract 5 topics, among which the weight of the "cost performance" topic is 0.4. Assign a personalized weight of 0.8 to the "cost performance" node in the recommended path. Fuse the initial path weight of 0.766 and the user preference weight of 0.8 to obtain the personalized weight 0.7×0.766 + 0.3×0.8 = 0.7762. Apply the Q-learning algorithm, set the click weight to 0.6 and the dwell time weight to 0.4. The user clicks on the "cost performance" node, isClicked = 1, and the dwell time of 30 seconds is normalized to 0.8. Calculate the reward value 0.6×1 + 0.4×0.8 = 0.92. Update the Q-value table, and increase the Q-value of the "cost performance" node from the initial 0.5 to 0.68. According to the new Q-value, increase the weight of the "cost performance" node in the navigation path to 0.82 to achieve dynamic optimization of the personalized navigation path.
[0045] S109. Present the knowledge navigation results through a visual interaction method, provide multi-dimensional attribute filtering and sorting functions, support the user's exploratory browsing and in-depth mining, and then provide personalized knowledge recommendation services according to the user's query scenarios and preferences.
[0046] The force-directed graph algorithm is used to transform the knowledge navigation result into a visual network structure, and the visual network structure includes nodes representing concepts and edges representing the relationships between concepts; a multi-dimensional attribute index is constructed according to the visual network structure, and the concept nodes in the knowledge graph are classified and labeled by using an inverted index structure; the Node2Vec algorithm is used to perform graph embedding on the knowledge graph to obtain embedding vectors that capture the graph structure information; a user-concept interaction matrix is constructed according to the embedding vectors and the user historical behavior data, and the similarities between users and the similarities between concepts are calculated; a recommendation algorithm based on the attention mechanism is used to dynamically adjust and sort the recommendation list according to the similarities between users and the similarities between concepts; through a real-time data processing tool, the user interaction operation is mapped into a dynamic query request of the knowledge graph; the knowledge graph is updated according to the dynamic query request to realize the exploratory browsing and in-depth mining functions with real-time response; if the user interaction operation includes a multi-dimensional attribute filtering instruction, the knowledge graph is dynamically filtered and sorted according to the multi-dimensional attribute filtering instruction.
[0047] Exemplarily, the force-directed graph algorithm is used to transform the knowledge navigation results into a visual network structure, where nodes represent concepts, edges represent relationships between concepts, node size and color coding represent concept importance and category, and edge thickness represents relationship strength. A multidimensional attribute index is constructed, and the inverted index structure is used to classify and annotate the concept nodes in the knowledge graph to achieve fast attribute-based retrieval and filtering. The Node2Vec algorithm is used to embed the knowledge graph to capture the structural information of the graph, and then the user-concept interaction matrix is constructed based on the user's historical behavior data. The similarity between users and the similarity between concepts are calculated in combination with the graph embedding results to generate a personalized concept recommendation list. The recommendation algorithm based on the attention mechanism is used to dynamically adjust and sort the recommendation list according to the user's query scenario and preference to provide more accurate personalized recommendations. Through real-time data processing tools, user interaction operations are mapped to dynamic queries and updates of the knowledge graph, realizing real-time responsive exploratory browsing and deep mining functions, and supporting dynamic screening and sorting of multidimensional attributes. In the knowledge navigation system in the field of e-commerce, the force-directed graph algorithm is used to visualize the knowledge graph containing 5,000 concept nodes and 10,000 edges. The node size range is set to 10-50 pixels, 10 different colors are used to distinguish concept categories, and the edge thickness range is 1-5 pixels. A multi-dimensional attribute index is constructed, and an inverted index structure is used to index 20 attributes of concept nodes, such as price, brand, performance, etc. The Node2Vec algorithm is applied, and the window size is set to 10 and the walk length is set to 80 to generate a 128-dimensional node embedding vector. Based on the historical behavior data of 1 million users, a user-concept interaction matrix in the form of a sparse matrix is constructed. Combined with the graph embedding results, cosine similarity is used to calculate the similarity between users and concepts, and the threshold is set to 0.7 to generate a top-50 personalized concept recommendation list for each user. A multi-head attention mechanism is used, and the number of heads is set to 8. The recommendation list is dynamically adjusted and sorted according to the user's most recent 10 query scenarios and browsing preferences in the past 30 days. Using a stream processing tool, a processing delay of 100ms is set to map user interactions such as clicks, slides, and zooms to graph query and update operations in real time. It supports combined filtering of up to 5 attributes and 3 sorting methods: relevance, popularity, and time. It achieves millisecond-level response speed and provides users with smooth exploratory browsing and deep mining experience.
[0048] It should be noted that the above examples are only some specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments, and there are many variations. All variations that can be directly derived or associated with the content disclosed by a person skilled in the art should be considered as the protection scope of the present invention.
Claims
1. An index selection method for cross - domain multi - dimensional query features, characterized in that, The method includes: Based on the user's historical query behavior data, through query topic classification and query intention recognition, judge the user's query field and purpose, and obtain the domain label and intention label of the query; For different query domain labels and intention labels, construct a domain ontology knowledge base, and generate an intention template library based on the analysis of a large-scale query log. Perform semantic parsing and feature extraction on the query to obtain multi-dimensional attribute information including keywords, entities, time, and location in the query; According to the weight calculation of the query terms and the semantic similarity between the query terms, optimize the combination of the extracted query multi-dimensional attributes to form a query feature vector representation, and use query feature clustering to mine potential query patterns, and optimize the node stratification and edge connection strength of the ontology knowledge base; In the domain ontology knowledge base, adopt the graph embedding representation learning method to map the ontology concepts to a low-dimensional vector space, calculate the correlation degree between different concepts through the semantic similarity between vectors, and dynamically adjust the mapping weight of the concepts using the query pattern to construct a hierarchical and navigable small-world ontology graph indexing algorithm, where the tree-like topology of the hierarchical structure serves as the backbone of the small-world indexing algorithm network; In the ontology graph index, optimize the hierarchical division of the concept nodes to make it meet the average shortest path characteristics of the small-world indexing algorithm network, control the hop count and query complexity of the navigation path, and dynamically adjust the edge weight and node connection mode through the network mapping of the user's query feedback behavior; According to the similarity between the user's query feature vector and the ontology concept vector, dynamically select the most relevant ontology subgraph as the candidate index. If the number of candidate indexes exceeds the threshold, further narrow the candidate index range according to the domain label and pattern of the query; In the candidate index subgraph, starting from the user's query features, adopt a breadth-first search strategy based on semantic similarity to discover the concept nodes related to the query and their semantic relationships, form cross-concept knowledge links, and comprehensively consider the importance of the concepts in the hierarchical structure and the small-world indexing algorithm network during the search process, and preferentially expand the nodes with a high degree of match with the query intention, and record the navigation path between the nodes; Optimize the navigation path of the knowledge link, comprehensively consider the semantic relevance between nodes and the characteristics of the small-world indexing algorithm, recommend the optimal knowledge navigation path for the user, and dynamically adjust the optimization strategy according to the user's query history behavior and feedback; Present the knowledge navigation results through a visual interaction method, provide multi-dimensional attribute filtering and sorting functions, support the user's exploratory browsing and in-depth mining, and then provide personalized knowledge recommendation services according to the user's query scenario and preferences.
2. The method according to claim 1, wherein The step of judging the user's query field and purpose, and obtaining the domain label and intention label of the query by query topic classification and query intention recognition based on the user's historical query behavior data includes: Obtain the user's historical query records, and extract the query text, timestamp, and click link from the user's historical query records; Perform word segmentation processing on the query text to obtain a set of query keywords; Perform text clustering on a large amount of historical query data, extract keywords of each category from the text clustering to form a topic classification dictionary; According to the above-mentioned topic classification dictionary, the query keyword set is subjected to topic vectorization, the cosine similarity is used to calculate the similarity between the query text and each topic category, and the category with the highest similarity is selected as the query domain label; The subject-verb-object relationship in the query text is extracted by using the dependency syntactic analysis tree, and the entities and attributes in the query text are identified in combination with named entity recognition; The semantic features of the query are extracted according to the subject-verb-object relationship and entities and attributes; Based on the semantic features, a support vector machine is used to judge the intention type of the query to obtain the query intention label; The user historical query data is divided according to a fixed time window, and the occurrence frequencies of each query domain and intention in each time window are calculated; A user interest vector is constructed according to the occurrence frequencies; The sliding time window method is adopted, and in combination with the query domain label and the query intention label, the cosine similarity with the user interest vector is calculated; If the cosine similarity is higher than the preset threshold, it is determined as the main query domain and purpose of the user.
3. The method according to claim 1, wherein, For different query domain labels and intention labels, a domain ontology knowledge base is constructed, and an intention template library is generated based on the analysis of a large-scale query log. Semantic parsing and feature extraction are performed on the query to obtain multi-dimensional attribute information including keywords, entities, time, and location in the query, including: According to the preset query domain labels and intention labels, relevant entities, attributes, and relationships are extracted from the knowledge graph to construct the domain ontology knowledge base. At the same time, the Apriori algorithm is used for frequent pattern mining based on the query log to generate the intention template library; The input query text is segmented, and the hidden Markov model is used for part-of-speech tagging to obtain the tagged word sequence; Named entities in the tagged word sequence are identified through the conditional random field model, and the keyword, entity, time, and location information included in the query text is extracted to obtain the feature vector of the query; The syntactic structure of the query text is obtained by using dependency syntactic analysis, and in combination with the entity relationships in the domain ontology knowledge base, semantic parsing of the query text is performed to identify the semantic roles and relationships in the query text; The cosine similarity between the feature vector of the query and the templates in the intention template library is calculated, and the template with the highest similarity is selected as the query intention; If the time and location attributes are identified, the entity type and attribute definition in the ontology knowledge base are used to map the identified entities and keywords to the corresponding attribute dimensions to obtain the multi-dimensional attribute information including keywords, entities, time, and location in the query text.
4. The method according to claim 1, wherein According to the weight calculation of the query words and the semantic similarity between the query words, the extracted query multi-dimensional attributes are combined and optimized to form a query feature vector representation, and potential query patterns are mined by using query feature clustering to optimize the node layering and edge connection strength of the ontology knowledge base, including: The weight of the query word is calculated according to the TF-IDF algorithm, where TF is the number of occurrences of the word in the current query, and IDF is the inverse document frequency of the word in the entire query set; The semantic similarity between the query words is measured by the cosine similarity to obtain the query feature vector representation; The k-means++ algorithm is used to select the initial class centers. If the selection is completed, the k-means algorithm is used to cluster the query feature vectors, and the number of clusters is set to k; Iteratively calculate the class centers until the class centers are stable or the maximum number of iterations is reached, and obtain k query pattern clusters; For the k query pattern clusters, set the support threshold; Calculate the occurrence frequency of the attribute combination in the query pattern cluster. If the occurrence frequency exceeds the support threshold, it is determined as a frequent attribute combination; Extract the frequent attribute combinations to construct potential query patterns; Update the attribute set of the corresponding concept nodes in the ontology knowledge base according to the potential query patterns; Based on the common attributes of the samples within the query pattern cluster, calculate the similarity between concept nodes; Judge whether the similarity is higher than the preset threshold. If it is higher than the preset threshold, adjust the concept nodes to a parent-child relationship or a sibling relationship; Update the hierarchical relationship between concept nodes in the ontology knowledge base according to the adjusted relationship.
5. The method according to claim 1, wherein In the domain ontology knowledge base, the graph embedding representation learning method is adopted to map ontology concepts to a low-dimensional vector space, calculate the association degree between different concepts through the semantic similarity between vectors, and dynamically adjust the mapping weights of concepts by using query patterns, and construct a hierarchical and navigable small-world ontology graph indexing algorithm. Among them, the tree-like topology of the hierarchical structure is used as the backbone of the small-world indexing algorithm network, including: Use the Word2Vec algorithm to perform graph embedding representation learning on the domain ontology knowledge base to obtain the low-dimensional vector representation of ontology concepts in a 128-dimensional vector space; Calculate the cosine similarity between different concept vectors according to the low-dimensional vector representation. If the cosine similarity is greater than the preset similarity threshold, establish a semantic association edge between the two concepts; Use the PageRank algorithm to calculate the importance score of each concept node, and the importance score is used to construct a tree-like topological structure as the backbone of the indexing network; Obtain the user query log, count the occurrence frequency of each concept under different query patterns, calculate the conditional probability of the concept under each query pattern according to the occurrence frequency, and the conditional probability is used as the weight of this concept under the query pattern; Based on the tree-like topological structure, use the Kleinberg algorithm to add long-range connections. The establishment probability of the long-range connections is proportional to the semantic similarity and importance score between nodes, and a hierarchical and navigable small-world ontology graph index is obtained.
6. The method according to claim 1, wherein In the ontology graph index, optimize the hierarchical division of concept nodes to make it meet the average shortest path characteristics of the small-world indexing algorithm network, control the hop count and query complexity of the navigation path, and dynamically adjust the edge weights and node connection patterns through the network mapping of user query feedback behavior, including: Use the AGNES agglomerative hierarchical clustering algorithm to perform hierarchical division on the concept nodes in the ontology graph index to obtain the initial hierarchical structure; Calculate the average shortest path length of the ontology graph index according to the initial hierarchical structure. If the average shortest path length exceeds the preset threshold, add cross-layer connection edges based on the semantic similarity and hierarchical difference between nodes until the small-world network characteristics are met; For a graph structure that satisfies the characteristics of a small-world network, set a maximum hop count threshold, prune paths that exceed the maximum hop count threshold, and optimize the query path using the A* search algorithm; Obtain user query logs, count the access frequencies between nodes according to the user query logs, use the access frequencies as the initial weights of the edges, and construct a weighted graph structure based on the optimized network structure; Process the weighted graph structure using the Q-learning algorithm, map the user query feedback behavior to a reward signal according to the query success rate and query time, and dynamically adjust the edge weights and node connection patterns through the reward signal to achieve adaptive optimization of the hierarchical structure.
7. The method according to claim 1, wherein Based on the similarity between the user query feature vector and the ontology concept vector, dynamically select the most relevant ontology subgraph as a candidate index. If the number of candidate indexes exceeds the threshold, further narrow the range of candidate indexes according to the domain label and pattern of the query, including: Obtain a user query, and convert the user query into a feature vector according to a pre-established Word2Vec model; At the same time, obtain the preset ontology concept vector representation; Calculate the similarity between the user query feature vector and the ontology concept vector using cosine similarity; If the similarity is greater than the preset threshold, determine the corresponding ontology concept as a seed node; For the seed node, expand three layers deep outward through the breadth-first search algorithm, and extract the subgraph containing the seed node; Use the subgraph as a candidate index; Use a support vector machine classifier to perform domain classification on the user query; Obtain the domain label of the user query from the output of the support vector machine classifier; Combine the domain label and the query pattern in the pre-stored historical query logs, and filter the candidate indexes using a weighted scoring method; Calculate the comprehensive score according to the domain relevance and query pattern matching degree; Determine the subgraph with the highest score as the final candidate index.
8. The method according to claim 1, wherein In the candidate index subgraph, starting from the user query feature, adopt a breadth-first search strategy based on semantic similarity to discover concept nodes related to the query and their semantic relationships, form cross-concept knowledge links, and comprehensively consider the importance of concepts in the hierarchical structure and the small-world index algorithm network during the search process, preferentially expand nodes with a high degree of match with the query intent, and record the navigation path between nodes, including: Obtain the user query feature vector, calculate the semantic similarity between the user query feature vector and all concept nodes in the candidate index subgraph, and select the top three nodes with the highest similarity as the search starting points according to the semantic similarity; Adopt a breadth-first search algorithm based on a priority queue to expand the search starting points, and calculate the matching degree between the nodes to be expanded and the query intent; According to the matching degree, use the PageRank algorithm to calculate the importance score of the nodes in the small-world index network, and combine the hierarchical information of the nodes in the hierarchical structure to obtain the comprehensive importance score of the nodes through a weighted average method; Determine the expansion priority of nodes according to the comprehensive importance score, traverse adjacent nodes layer by layer, record the semantic relationships between nodes, where the semantic relationships include types such as "is a kind of", "belongs to", and "contains", and store the semantic relationships in the form of triples; If the formed knowledge links reach the preset node quantity threshold, then use the Kruskal minimum spanning tree algorithm to prune and optimize the knowledge links, remove redundant nodes and relationships, and obtain the final cross-concept knowledge links.
9. The method according to claim 1, wherein, The optimization of the navigation path of the knowledge link comprehensively considers the semantic relevance between nodes and the characteristics of the small-world index algorithm, recommends the optimal knowledge navigation path for the user, and dynamically adjusts the optimization strategy according to the user's query history behavior and feedback, including: Calculate the semantic relevance score for node pairs in the knowledge link, and construct a weighted graph structure according to the semantic relevance score and the degree centrality, clustering coefficient, and average path length of nodes in the small-world index algorithm; Among them, the weighted graph structure uses the weighted sum of the semantic relevance score, degree centrality, clustering coefficient, and average path length as the comprehensive weight of the edge; According to the weighted graph structure, use the multi-objective A* algorithm to calculate the set of optimal navigation paths from the starting node to multiple target nodes to obtain the initial recommended path; The initial recommended path consists of multiple nodes; Obtain the user interest characteristics from the user's historical query records, use the LDA topic model to extract the topics and preferences that the user is interested in, and assign personalized weights to the nodes in the initial recommended path according to the user interest characteristics; Perform weighted fusion on the personalized weight and the weight of the initial recommended path to obtain the fusion weight; The fusion weight is determined by the weighted sum of the personalized weight and the initial recommended path weight; Use the Q-learning algorithm to process the user's click behavior and residence time, and set the reward function; Among them, the reward function is determined by the sum of the product of the click weight and whether it is clicked and the product of the residence time weight and the normalized residence time; Update the Q-value table according to the reward function, and update the weights of nodes and edges according to the Q-value table to optimize the navigation path personalized.
10. The method according to claim 1, wherein, The presentation of the knowledge navigation result through a visual interaction method provides multi-dimensional attribute filtering and sorting functions, supports exploratory browsing and in-depth mining by users, and then provides personalized knowledge recommendation services according to the user's query scenario and preferences, including: Use the force-directed graph algorithm to transform the knowledge navigation result into a visual network structure, where the visual network structure includes nodes representing concepts and edges representing relationships between concepts; Construct a multi-dimensional attribute index according to the visual network structure, and classify and label the concept nodes in the knowledge graph through the use of an inverted index structure; Use the Node2Vec algorithm to perform graph embedding on the knowledge graph to obtain embedding vectors that capture the graph structure information; Construct a user-concept interaction matrix according to the embedding vectors and the user's historical behavior data, and calculate the similarity between users and the similarity between concepts; Use a recommendation algorithm based on the attention mechanism to dynamically adjust and sort the recommendation list according to the similarity between users and the similarity between concepts; Map user interaction operations into dynamic query requests for the knowledge graph through real-time data processing tools; Update the knowledge graph according to the dynamic query request to achieve exploratory browsing and in-depth mining functions with real-time response; If the user interaction operation contains multi-dimensional attribute filtering instructions, dynamically filter and sort the knowledge graph according to the multi-dimensional attribute filtering instructions.
Citation Information
Cited By
Intelligent intention judgment method based on template content automatic preferential selection
CN120745654A
A Smart Intent Judgment Method Based on Automatic Template Content Optimization
CN120745654B
Manual navigation method and system based on knowledge graph
CN120975202A
Manual navigation method and system based on knowledge graph
CN120975202B
Dynamic index generation method and system driven by business form fields
CN120994670A