A method for automatically recommending standard terms

By constructing a knowledge graph and using a gated attention fusion mechanism, the problem of messy recommendation results in traditional term recommendation methods is solved, and efficient and accurate standard term recommendations are achieved.

CN119938902BActive Publication Date: 2025-07-11CHINA STANDARD TECH DEV CORP

Patent Information

Application Number
CN202510431046.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The traditional term recommendation method lacks advanced screening and accurate matching mechanisms, resulting in messy recommendation results, making it difficult to quickly find professional data that meets user needs.

Method used

By constructing a knowledge graph, discrete semantic alignment using entity vectors and path coding, combined with the gating attention fusion mechanism and graph sampling algorithm, efficient matching of user intentions and accurate recommendations of candidate terms are achieved.

Benefits of technology

Fast vector similarity retrieval and structure matching are achieved, improving the accuracy and efficiency of recommendations, and ensuring that the recommendation results always hit user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938902B_ABST
    Figure CN119938902B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for automatically recommending standard terms, including: collecting industry data to be evaluated, constructing a knowledge graph, discretizing semantic alignment between entity vectors and path encodings, constructing a matrix-graph joint index, using a gated attention fusion mechanism to process user input data to obtain input intentions, searching for candidate recommended terms associated with the input intentions in the knowledge graph based on the matrix-graph joint index and a graph sampling algorithm and screening to obtain recommended standard terms, clustering the recommended standard terms, using a rime algorithm to aggregate the standard terms within the cluster, obtaining representative terms for each cluster and incrementally updating them to the matrix-graph joint index, and if no recommended standard terms are searched, recommending synonyms. This method improves the information retrieval and screening efficiency through the matrix-graph joint index, and searches the knowledge graph through the graph sampling algorithm, making the recommended results conform to the user's intentions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automatic recommendation of standard terms, and particularly to a method for automatically recommending standard terms. Background Art

[0002] With the vigorous development of modern technology and the continuous growth of personal data and industrial data volume, more and more data has been accumulated in various fields, and there are repeated expressions of various professional term concepts. People have higher and more accurate requirements for data classification, especially for the association and recommendation function of professional data.

[0003] However, the problems of traditional term recommendation are significant, and the update of domain information lags behind. Due to the lack of advanced screening and precise matching mechanisms, ordinary recommendation results are often presented in a disorderly manner to users, including many contents that are not relevant or have a low degree of association with people's needs, making it difficult to quickly find professional data recommendation content that truly meets their own needs. Therefore, a method for automatically recommending standard terms is needed. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for automatically recommending standard terms.

[0005] To achieve the above object, the present invention is implemented according to the following technical solution:

[0006] In a first aspect of the present invention, a method for automatically recommending standard terms is provided, including: collecting industry data to be evaluated, preprocessing the industry data to be evaluated to obtain standard term data, and performing semantic annotation on the entities of the standard term data;

[0007] Constructing a knowledge graph based on the entities and semantic annotations, and using the semantic association relationships of the entities in the knowledge graph to find synonyms to obtain a synonym library;

[0008] Calculating the cosine similarity between the entities in the knowledge graph, constructing a semantic association sparse matrix, marking the matrix index position for each entity to form a semantic association matrix index, traversing the knowledge graph to extract entity association paths and encoding them, constructing a graph structure index, and mapping the entity vectors and path encodings to a shared codebook through vector quantization for discrete semantic alignment to construct a matrix-graph joint index supporting dynamic expansion;

[0009] Using a gated attention fusion mechanism to divide the user input data to obtain feature text and input intention, searching for candidate recommended terms associated with the input intention in the knowledge graph based on the matrix-graph joint index and graph sampling algorithm, and uniformly extracting the candidate recommended terms to obtain recommended standard terms;

[0010] If recommended standard terms are obtained through search, cluster the recommended standard terms, use the rime algorithm to aggregate the standard terms within the cluster, obtain the representative terms of each cluster and incrementally update them to the matrix-graph joint index;

[0011] If no recommended standard terms are found, search for synonyms of the user input data in the thesaurus to obtain recommended synonyms.

[0012] As a further method, the step of constructing a knowledge graph based on the entity and semantic annotation includes: extracting entity attributes and constraint conditions based on the entity and semantic annotation, constructing a standard term ontology model, and using the standard term ontology model to inject the entities in the term library into the knowledge graph according to the mapping rule of entity-attribute-constraint to form an entity association network.

[0013] As a further method, the step of discretely aligning the entity vector and the path encoding through vector quantization mapping to a shared codebook to construct a matrix-graph joint index that supports dynamic expansion includes: training the entities in the knowledge graph based on the BERT model to obtain entity vectors; calculating the cosine similarity between entity vectors, constructing an n-order semantic association matrix based on the number n of entity vectors, and using the matrix elements to represent the semantic similarity between entity vectors i and j; sparsifying the semantic association matrix, retaining the top 5% of non-zero elements with high similarity, setting the remaining elements to zero values, assigning an identifier to each entity vector, and establishing a bidirectional hash mapping between the entity identifier and the matrix row and column indexes to form a semantic association matrix index;

[0014] Traverse the association paths between entities in the knowledge graph, use the message passing mechanism of the graph neural network to aggregate the features of nodes and edges on the path, and generate path encoding; map the path encoding to a semantic space with the same dimension as the entity vector through a graph embedding algorithm to form a graph structure index;

[0015] Discretely align the entity vector and the path encoding through vector quantization mapping to a shared codebook to construct a matrix-graph joint index that supports dynamic expansion.

[0016] As a further method, the step of using a gated attention fusion mechanism to partition the user input data to obtain feature text and input intent includes: obtaining the entity type and annotation information of the user input data, where the annotation information includes the start and end position information of the entity in the original text, splicing the annotation information with the user input text according to the start and end position information, and through the gated attention fusion mechanism, dynamically allocate the weight of the annotation information in the splicing according to the semantic relevance between the annotation information and the original text. The gated attention fusion formula is:

[0017] ,

[0018] Among them represents the final vector after fusing gated attention and residual connections, is the vector obtained after encoding the original text, is the vector obtained after encoding the annotation information, is the parameter matrix, e is the natural constant, represents the weight of the annotation information during the concatenation process;

[0019] Input the concatenated feature text into the intent recognition model to obtain the user input intent vector.

[0020] As a further method, the step of searching for candidate recommended terms associated with the input intent in the knowledge graph based on the matrix-graph joint indexing and graph sampling algorithm includes: based on the entity type of the user input data and the keywords associated with the annotation information, using the keywords as query points, locating the key entities of the keywords in the knowledge graph through the matrix-graph joint indexing, and using the graph structure index to mine the potential entities related to the key entities through the distance-aware subgraph sampling algorithm. The formula of the distance-aware subgraph sampling algorithm is:

[0021] ,

[0022] Among them represents the set of potential entities obtained after screening by the formula, e represents a single entity in the knowledge graph subgraph, represents the subgraph of the knowledge graph, h represents the head entity, r represents the relationship, and t represents the tail entity, represents the head entity h in the embedding vector in the hyperbolic space, represents the embedding vector of the relationship r in the hyperbolic space, represents the embedding vector of the head entity t in the hyperbolic space, represents the screening threshold, which is the attention weight coefficient;

[0023] For the located key entities and potential entities, find their synonyms in the thesaurus, use these synonyms as new search keywords, search again in the knowledge graph, and integrate all the entities obtained from the two searches to remove duplicate data to form candidate recommended terms.

[0024] As a further method, the step of uniformly extracting the candidate recommended terms to obtain the recommended standard terms includes: converting the entity vectors of the terms in the candidate recommended term set and the user input intent vector into hyperbolic space vectors, and calculating the semantic similarity between the candidate recommended term vector and the user intent vector using the hyperbolic distance formula. The hyperbolic distance formula is:

[0025] ,

[0026] where S represents semantic similarity, and the closer the value is to 1, the more similar. represents the difference vector between the candidate recommended term vector x and the user input intention vector y. x is the candidate recommended term vector, and y is the square of the norm of the user input intention vector y. is used to adjust the influence degree of the function on the final similarity. i.e., the inverse hyperbolic cosine function;

[0027] Arrange the candidate standard term set in descending order according to semantic similarity, extract the candidate recommended terms with odd orders, and extract repeatedly until 15% of the candidate standard term set remains to obtain the recommended labeled terms.

[0028] As a further method, the step of clustering the recommended standard terms and using the rime algorithm to aggregate the standard terms within the cluster to obtain the representative term of each cluster includes: searching to obtain the recommended standard terms, clustering the recommended standard terms, optimizing the parameters of the DBSCAN clustering algorithm using the rime algorithm, setting the optimization objective function as minimizing the inter-cluster similarity and maximizing the within-cluster weight variance. The within-cluster weight value is the weighted sum of the TF-IDF scores of the standard terms in the recommended standard term set and the centrality of the knowledge graph of the standard terms. The centrality of the knowledge graph is calculated by the PageRank algorithm. The formula for minimizing the inter-cluster similarity is:

[0029] ,

[0030] where is the neighborhood radius, adjusted according to the distribution of terms in the vector space, and the calculation formula is , where is the mean of the distances between all terms, is the standard deviation; is the minimum sample number, set as the logarithmic function of the domain term density, , N is the total number of current domain terms, is used to control the importance of the objective in the comprehensive optimization. k represents the number of clusters obtained after clustering. i ranges from 1 to k - 1, and j ranges from i + 1 to k, used to traverse different cluster pairs. is the number of terms in cluster ; is the number of terms in cluster ; is the vector of the m-th term in cluster ; is the number of terms in cluster ;

[0031] The formula for maximizing the within-cluster weight variance is as follows:

[0032] ,

[0033] where is the neighborhood radius, which is adjusted according to the distribution of terms in the vector space, and the calculation formula is , where is the mean of the distances between all terms, is the standard deviation; is the minimum number of samples, which is set as the logarithmic function of the domain term density, , N is the total number of current domain terms, is used to control the importance of the target in the comprehensive optimization. k represents the number of clusters obtained after clustering. i ranges from 1 to k and is used to traverse different cluster pairs, is the number of terms in cluster s, represents the weight set of all terms in the i-th cluster, represents the i-th cluster the weight set of terms within;

[0034] According to the frequency of occurrence of the standard terms within the cluster, select the standard term with the highest frequency of occurrence or the first occurrence in the term library as the representative standard term of the current cluster, and incrementally update the representative term to the matrix-graph joint index.

[0035] As a further method, the incrementally updating the representative term to the matrix-graph joint index includes: adopting a delayed merging strategy, accumulating the representative standard terms to reach 10% of the total historical update data volume, comparing the change rate of the representative standard term with the local density kurtosis value of the recent historical update data. If the change rate is less than 20%, use the representative standard term to search for associated nodes in the knowledge graph, use the centrality of the representative term in the knowledge graph as the entity attribute of the associated node to incrementally update the knowledge graph, dynamically expand the shared codebook based on the representative term through a vector algorithm and update the matrix-graph joint index. The calculation formula for the local density kurtosis value is:

[0036] ,

[0037] where CDI represents the local density kurtosis value, m represents the number of samples, and T represents for each term data point , when its distance from the point z to be evaluated is less than or equal to 1, its corresponding subscript i will be included in the set T, is the mean of m , z represents the target vector of the deviation degree of the existing term data distribution, w represents the target vector of the deviation degree of the existing term data distribution, and k represents the dimension of the term vector.

[0038] As a further method, the step of obtaining recommended synonyms includes: when the candidate recommended term set is empty, based on the knowledge graph, keywords associated with the user input text are obtained according to the user input intention vector, the keywords are input into the standard term synonym library to obtain synonyms, and the recommended synonyms are obtained.

[0039] Compared with the prior art, the embodiments of the present invention at least have the following advantages or beneficial effects:

[0040] (1) The present invention obtains standard term data through the induction and collation of industry data to be evaluated, constructs a knowledge graph, and performs discrete semantic alignment through vector quantization mapping of the entity vectors and path encodings in the knowledge graph to a shared codebook, constructs a matrix-graph joint index supporting dynamic expansion, the matrix index supports fast vector similarity retrieval, and at the same time the graph index supports structure-based exact matching.

[0041] (2) The present invention divides the user input data through the gated attention fusion mechanism to obtain feature text and input intention, and searches for candidate recommended terms associated with the input intention in the knowledge graph based on the matrix-graph joint index and the distance-aware subgraph sampling algorithm, realizing efficient matching of user intention data.

[0042] (3) The present invention performs clustering analysis on the obtained recommended standard terms, so that the hot words, high-frequency words, and new words searched by people can enter the matrix-graph joint index, further improving the recommendation efficiency of search hot terms, and through the recommendation of synonyms for unknown keywords, the recommendation requirements of people can always be hit. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a flowchart of the steps of a method for automatically recommending standard terms in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] Referring to Figure 1 as shown, the present invention provides a method for automatically recommending standard terms, including:

[0046] Step S1: Collect industry data to be evaluated, preprocess the industry data to be evaluated to obtain standard term data, and perform semantic annotation on the entities of the standard term data;

[0047] In this embodiment, power industry data is collected to obtain industry data in the order of 100,000, including power equipment maintenance records, power grid fault reports, power industry standard documents, technical specifications, etc.; the power industry data is cleaned and de-duplicated, entities are identified through the BERT-BiLSTM-CRF model, power terms are matched using regular expressions, and term disambiguation is performed in combination with the domain dictionary "Nomenclature for Electric Power Engineering" to obtain standard term data, and the entities of the standard term data are semantically annotated, annotating equipment entities such as "main transformer" and "110 kV busbar", annotating fault types such as "insulation breakdown" and "poor contact", and annotating parameter entities such as "rated current 1250 A" and "power frequency withstand voltage 42 kV"

[0048] Step S2 constructs a knowledge graph based on the entities and semantic annotations, and uses the semantic association relationships of the entities in the knowledge graph to find synonyms to obtain a synonym library;

[0049] It should be explained that the step of constructing a knowledge graph based on the entities and semantic annotations includes extracting entity attributes and constraint conditions based on the entities and semantic annotations, constructing a standard term ontology model, and injecting the entities in the term library into the knowledge graph according to the entity-attribute-constraint mapping rules of the standard term ontology model to form an entity association network; industry terms are complex, and the ontology model can clearly define the concepts and classify the terms, clearly define the attributes and relationships of each entity, can construct a hierarchical and structured knowledge system, organize the scattered terms and data in the industry, make the knowledge graph more systematic and comprehensive, provide an accurate semantic basis for the knowledge graph, and improve the practicality of the knowledge graph;

[0050] In this embodiment, based on the entities and semantic annotations of the standard term data, entities such as "110 kV XX Substation #1 Main Transformer" are imported from the equipment ledger, associated technical parameters such as "capacity 50 MVA" and "connection group YN,d11" are extracted, entity attributes and constraint conditions are obtained, the relationship of "series connection" is established between "circuit breaker" and "disconnecting switch", and the relationship of "causing a fault" is established between "insulator" and "flashover". Based on the extracted attributes and constraint relationships, attributes such as "voltage level", "protection level", and "arc extinguishing medium" are set for entities such as "circuit breaker", and constraint conditions such as "action time coordination" and "temperature rise limit" are set. A standard term ontology model is constructed, and the entities in the term library are injected into the knowledge graph according to the entity-attribute-constraint mapping rules of the standard term ontology model to form an entity association network; by analyzing the nodes and edges related to entities such as "circuit breaker" and "relay protection" in the knowledge graph, combined with the standard term habits in the power industry, synonyms of "circuit breaker" are selected: "switch", etc., and synonyms of "relay protection" are selected: "relay protection", etc., to obtain a standard synonym library.

[0051] Step S3 calculates the cosine similarity between entities in the knowledge graph, constructs a semantic association sparse matrix, assigns a matrix index position to each entity to form a semantic association matrix index, traverses the knowledge graph to extract entity association paths and encodes them, constructs a graph structure index, and maps the entity vectors and path encodings to a shared codebook through vector quantization for discrete semantic alignment, thus constructing a matrix-graph joint index that supports dynamic expansion;

[0052] In this embodiment, the cosine similarity between entities in the knowledge graph is calculated. For example, the cosine similarity between "circuit breaker" and "disconnector" is 0.85. The top 5% of the high-similarity edges are retained, such as equipment-fault associations, to construct a semantic association sparse matrix. The pairwise similarities of 1000 power terms such as "circuit breaker", "disconnector", and "relay protection" are calculated to generate a 1000-order semantic matrix M, and the matrix elements store the similarity values between terms i and j, count 499500 similarity values, and take the top 5% as high-similarity edges to obtain 24975 edges);

[0053] In addition to "circuit breaker - disconnector (0.85)", edges such as "circuit breaker - short-circuit breaking (0.92)" (equipment-function association), "disconnector - electrical isolation (0.88)" (equipment-function association), and "circuit breaker - fault - contact ablation (0.81)" (equipment - fault association) are also retained to construct a semantic association sparse matrix. The matrix is filled with high-similarity edges. For example, the historical association probability between the circuit breaker (index 1) and the contact ablation fault (index 501) is 0.7, and it is filled into , and only the index positions and association values of non-zero elements are stored to form a semantic association matrix index;

[0054] Use depth-first search to traverse the knowledge graph to extract entity association paths and encode them. For example, the path "substation → main transformer → winding → overheating" is encoded as [0.3, 0.6, 0.2, 0.7]. Use the TransE knowledge graph embedding model to convert the path into a vector and construct a graph structure index;

[0055] Use the BERT model to train power terms such as "circuit breaker" to generate 768-dimensional vectors: , and adopt the K-Means clustering algorithm to cluster the 768-dimensional "circuit breaker" vectors and the vectors of their related entities into the codebook entry C003, and cluster the path encoding vectors into C127 to achieve the semantic alignment of vectors and path encodings, establish a shared codebook, store discrete semantic units, and construct a matrix-graph joint index that supports dynamic expansion.

[0056] Step S4 uses a gated attention fusion mechanism to partition the user input data to obtain a feature text and an input intention, searches for candidate recommended terms associated with the input intention in the knowledge graph based on the matrix-graph joint index and the graph sampling algorithm, and uniformly extracts the candidate recommended terms to obtain recommended standard terms;

[0057] In this embodiment, the user inputs "Find the common causes of 220kV line tripping". The BERT-BiLSTM-CRF model is used to identify the entity "220kV line" and the corresponding annotation information (equipment) [3,8], "tripping" and the corresponding annotation information (fault type) [9,11]. The annotation information is spliced with the user input text according to the start and end position information. Through the gated attention fusion mechanism, the weight of this annotation information in the splicing is obtained according to the semantic relevance between the annotation information and the original text. The weight is 1.02, exceeding the benchmark value of 1, indicating that the annotation information spliced this time has a high relevance to the original text. The spliced feature text is input into the intention recognition model to obtain the user input intention vector. and the input intention "220kV line fault analysis";

[0058] Based on the entity type of the user input data and the initial keywords associated with the annotation information: "220kV line", "tripping", using the keywords as query points, the key entities of the keywords are located in the knowledge graph through the matrix-graph joint index. The graph structure index is used to mine potential entities related to the key entities through the distance-aware subgraph sampling algorithm of the knowledge graph, and the attention weight coefficient is set to 0.7, and the set of potential entities such as "maloperation of relay protection" and "insulator pollution" less than 0.7 is filtered out. For the located key entities and potential entities, their synonyms are searched in the thesaurus of synonyms. The synonyms such as "220 kV transmission line" and "220kV transmission line" are used as new search keywords, and the search is performed again in the knowledge graph. All the entities obtained from the two searches are integrated, and duplicate data is removed to form candidate recommended terms "220kV line, tripping, maloperation of relay protection, insulator pollution, 220 kV transmission line, switch tripping, misoperation of protection device, lightning protection measures for transmission line, incorrect protection setting value, pollution flashover fault". The entity vectors of the terms in the candidate recommended term set and the user input intention vector are converted into hyperbolic space vectors, and the semantic similarity between the candidate recommended term vector and the user intention vector is calculated using the hyperbolic distance formula. Let the adjustment parameter Then, the semantic similarity S of "maloperation of relay protection" is 0.656, the semantic similarity S of "insulator contamination" is 0.521, and the semantic similarity S of "line lightning protection design" is 0.789. According to the semantic similarity, the candidate standard terms are sorted in descending order, and the candidate recommended terms in odd orders are extracted. The extraction is repeated until 15% of the candidate standard term set remains, and the recommended annotation term "insulator contamination" is obtained.

[0059] It should be explained that when the selected recommended term is empty, to prevent the user from not getting any recommendations, keywords associated with the user input text are directly obtained based on the knowledge graph according to the user input intention vector. These keywords are input into the standard term synonym library to obtain synonyms for recommendation. This can make up for the matching time consumption of the previous two searches, avoid the user waiting, and directly recommending synonyms can enable the user to refine the further search, improving the user's recommendation experience and ensuring that the recommendation target can always hit the user's needs.

[0060] In this embodiment, when the user inputs "abnormal heating of GIS equipment" and there are no matching terms, the keywords "GIS equipment" are extracted, and synonyms such as "GIS combined electrical apparatus", "fully enclosed combined electrical apparatus", "partial discharge of GIS equipment", and "temperature rise test of combined electrical apparatus" are obtained by searching the synonym library and used for recommendation.

[0061] If recommended standard terms are obtained in step S5, cluster the recommended standard terms, use the rime algorithm to aggregate the standard terms within the cluster, obtain the representative term of each cluster, and incrementally update it to the matrix-graph joint index.

[0062] It should be explained that since industry data is constantly updated and the actual content searched by users is also constantly changing, for actual hot words and new terms that exist, they need to be incorporated into the term library and knowledge graph to optimize the recommendation efficiency. By performing cluster analysis on the recommendation results of the user's previous searches, more refined hot word features and new word impressions are obtained. Through the incremental update method, the overall calculation of the index is avoided, reducing the update time consumption and resources. By judging the local kurtosis value of the updated data and index, each updated data has a good embedding degree for the index.

[0063] In this embodiment, the candidate terms "maloperation of relay protection", "insulator contamination", "line overload", and "CT saturation" are obtained by searching. Cluster the recommended standard terms. The rime algorithm is used to optimize the parameters of the DBSCAN clustering algorithm to improve the clustering accuracy. The optimization objective function is set to minimize the inter-cluster similarity and maximize the intra-cluster weight variance. Calculate according to the formula The neighborhood radius is obtained as 0.386, which means that if the distance between two term vectors is less than 0.386, they are considered to be in the same neighborhood, and the minimum sample number is obtained is 5.39, indicating that when clustering, a neighborhood contains at least 5 terms to form a valid cluster. The intra-cluster weight value is the weighted sum of the TF-IDF score of the standard term in the recommended standard term set and the knowledge graph centrality of the standard term. For example, the document frequency of "relay protection" is low, and the TF-IDF value is 0.7. In the knowledge graph, "relay protection" is closely related to entities such as "circuit breaker" and "fault analysis". The value calculated by the PageRank algorithm is 0.6, and the intra-cluster weight value is 0.66;

[0064] Entities such as "relay protection", "circuit breaker", and "fault analysis" are divided into clusters based on the frequency of occurrence of standard terms in the cluster. ; "Insulator Dirty" is divided into clusters , according to the frequency of occurrence of standard terms in the cluster, select "circuit breaker" in the cluster The higher the frequency of occurrence, the more likely it is to be a cluster. Representative standard terms for "insulator pollution" are clustered representative standard terms; using the delayed merging strategy of SPFresh, after the cumulative representative standard terms reach the update threshold of 100, the change rate of the local density kurtosis value of the representative standard terms and the most recent historical update data is compared. If the change rate is less than 20%, the representative standard terms are used to find related nodes in the knowledge graph. After calculation, the local density kurtosis value CDI is 0.16, the most recent historical CDI is 0.14, and the change rate is 14%, which is less than 20%, and the CDI is 0.16 close to zero, which means that the index data updated this time has little impact on the overall index structure, and the overall data smoothness is high. The collection of representative standard terms such as "circuit breaker" and "insulator pollution" is dynamically expanded through the vector algorithm to share the code book and update the matrix-graph joint index for subsequent search.

[0065] In this embodiment, the step of using the gated attention fusion mechanism to divide the user input data to obtain feature text and input intent includes:

[0066] The entity type and annotation information of the user input data are obtained. The annotation information includes the start and end position information of the entity in the original text. The annotation information is spliced ​​with the user input text according to the start and end position information. Through the gated attention fusion mechanism, the weight of the annotation information in the splicing is dynamically allocated according to the semantic relevance between the annotation information and the original text. The gated attention fusion formula is:

[0067] ,

[0068] in represents the final vector after fusing gated attention and residual connection, is the vector obtained after encoding the original text, is the vector after the labeled information is encoded, is the parameter matrix, and e is the natural constant, represents the weight of the labeled information during the splicing process;

[0069] Input the spliced feature text into the intent recognition model to obtain the user input intent vector.

[0070] In this embodiment, the step of searching for candidate recommended terms associated with the input intent in the knowledge graph based on the matrix-graph joint index and graph sampling algorithm includes:

[0071] Based on the entity type of the user input data and the keywords associated with the labeled information, using the keywords as the query points, locate the key entities of the keywords in the knowledge graph through the matrix-graph joint index, and use the graph structure index to mine the potential entities related to the key entities through the distance-aware subgraph sampling algorithm. The formula of the distance-aware subgraph sampling algorithm is:

[0072] ,

[0073] where represents the set of potential entities obtained after screening by the formula, e represents a single entity in the knowledge graph subgraph, represents the subgraph of the knowledge graph, h represents the head entity, r represents the relationship, and t represents the tail entity, represents the head entity h 's embedding vector in the hyperbolic space, represents the embedding vector of the relationship r in the hyperbolic space, represents the embedding vector of the head entity t in the hyperbolic space, represents the screening threshold and is the attention weight coefficient;

[0074] For the located key entities and potential entities, find their synonyms in the thesaurus, use these synonyms as new search keywords, search again in the knowledge graph, and integrate all the entities obtained from the two searches to remove duplicate data to form candidate recommended terms.

[0075] In this embodiment, the step of uniformly extracting candidate recommended terms to obtain recommended standard terms includes:

[0076] Convert the entity vectors of the terms in the candidate recommended term set and the user input intent vector into hyperbolic space vectors, and calculate the semantic similarity between the candidate recommended term vector and the user intent vector using the hyperbolic distance formula. The hyperbolic distance formula is:

[0077] ,

[0078] Among them, S represents semantic similarity. The closer the value is to 1, the more similar. It represents the difference vector between the candidate recommended term vector x and the user input intention vector y. x is the candidate recommended term vector, and y is the square of the norm of the user input intention vector y. Used to adjust The influence degree of the function on the final similarity. That is, the inverse hyperbolic cosine function;

[0079] Arrange the candidate standard term set in descending order according to semantic similarity, extract the candidate recommended terms with odd order, and extract repeatedly until 15% of the candidate standard term set remains to obtain the recommended annotation terms.

[0080] In this embodiment, the step of clustering the recommended standard terms and using the rime algorithm to aggregate the standard terms within the cluster to obtain the representative terms of each cluster includes:

[0081] Search to obtain the recommended standard terms, cluster the recommended standard terms, optimize the parameters of the DBSCAN clustering algorithm using the rime algorithm, set the optimization objective function to minimize the inter-cluster similarity and maximize the intra-cluster weight variance. The intra-cluster weight value is the weighted sum of the TF-IDF scores of the standard terms in the recommended standard term set and the centrality of the knowledge graph of the standard terms. The centrality of the knowledge graph is calculated by the PageRank algorithm;

[0082] The formula for minimizing the inter-cluster similarity is:

[0083] ,

[0084] Among them Is the neighborhood radius, adjusted according to the distribution of terms in the vector space, and the calculation formula is , among which Is the mean of the distances between all terms, Is the standard deviation; Is the minimum sample number, set as the logarithmic function of the domain term density, , N is the total number of current domain terms, Used to control the importance of the target in the comprehensive optimization. k represents the number of clusters obtained after clustering. i ranges from 1 to k - 1, and j ranges from i + 1 to k, used to traverse different cluster pairs. Is the cluster The number of terms within, Is the cluster The number of terms within, Is the cluster The vector of the m-th term in, Is the cluster The vector of the nth term in;

[0085] The formula for maximizing the within-cluster weight variance is:

[0086] ,

[0087] where is the neighborhood radius, adjusted according to the distribution of terms in the vector space, and the calculation formula is , where is the mean of the distances between all terms, is the standard deviation; is the minimum number of samples, set as the logarithmic function of the domain term density, , N is the total number of current domain terms, is used to control the importance of the target in the comprehensive optimization, k represents the number of clusters obtained after clustering, and i ranges from 1 to k, used to traverse different cluster pairs, is the number of terms in cluster s, represents the set of weights of all terms in the ith cluster, represents the ith cluster the set of term weights within;

[0088] According to the frequency of occurrence of the standard terms within the cluster, select the standard term with the highest frequency of occurrence or the first occurrence in the term library as the representative standard term of the current cluster, and incrementally update the representative term to the matrix-graph joint index.

[0089] In this embodiment, the incrementally updating the representative term to the matrix-graph joint index includes:

[0090] Adopt a delayed merging strategy. When the cumulative representative standard terms reach 10% of the total historical update data volume, compare the change rate of the local density kurtosis value between the representative standard terms and the most recent historical update data. If the change rate is less than 20%, use the representative standard terms to search for associated nodes in the knowledge graph, use the centrality of the representative terms in the knowledge graph as the entity attribute of the associated nodes to incrementally update the knowledge graph, and dynamically expand the shared codebook and update the matrix-graph joint index based on the representative terms. The calculation formula for the local density kurtosis value is:

[0091] ,

[0092] where CDI represents the local density kurtosis value, m represents the number of samples, and T represents for each term data point , when its distance from the point z to be evaluated is less than or equal to 1, its corresponding subscript i will be included in the set T, is m The mean value, z represents the target vector of the deviation degree from the distribution of existing term data, w represents the target vector of the deviation degree from the distribution of existing term data, and k represents the dimension of the term vector.

[0093] In summary, the embodiments of the present application aim to efficiently meet user needs by constructing an industry knowledge graph and a matrix-graph joint index, realize a standard term recommendation method that conforms to the timeliness of user needs through the hyperbolic space recommendation ability, and update the index in a timely manner to improve the search efficiency of hot words, so as to create a more professional and accurate recommendation option for people.

[0094] The above content is only an example and illustration of the structure of the present invention. Those skilled in the art of the present technology can make various modifications or supplements to the described specific embodiments or use similar methods to replace them. As long as they do not deviate from the structure of the invention or exceed the scope defined by this claims, they should all fall within the protection scope of the present invention.

Claims

1. A method for automatically recommending standard terms, characterized in that, Including the following steps: Collect industry data to be evaluated, preprocess the industry data to be evaluated to obtain standard term data, and perform semantic annotation on the entities of the standard term data; Construct a knowledge graph based on the entities and semantic annotations, use the semantic association relationships of the entities in the knowledge graph to find synonyms, and obtain a synonym library; Calculate the cosine similarity between the entities in the knowledge graph, construct a semantic association sparse matrix, identify the matrix index position for each entity to form a semantic association matrix index, traverse the knowledge graph to extract entity association paths and encode them, construct a graph structure index, map the entity vectors and path encodings to a shared codebook through vector quantization for discrete semantic alignment, and construct a matrix-graph joint index that supports dynamic expansion; Use a gated attention fusion mechanism to partition the user input data to obtain feature text and input intent, search for candidate recommended terms associated with the input intent in the knowledge graph based on the matrix-graph joint index and graph sampling algorithm, and uniformly extract the candidate recommended terms to obtain recommended standard terms; If recommended standard terms are obtained through the search, cluster the recommended standard terms, use the rime algorithm to aggregate the standard terms within the cluster, obtain the representative term for each cluster, and incrementally update it to the matrix-graph joint index; If no recommended standard terms are found through the search, search for synonyms of the user input data in the synonym library to obtain recommended synonyms.

2. The automatic recommendation method for standard terms according to claim 1, wherein The step of constructing a knowledge graph based on the entities and semantic annotations includes: Based on the entities and semantic annotations, extract entity attributes and constraint conditions, construct a standard term ontology model, and use the standard term ontology model to inject the entities in the term library into the knowledge graph according to the entity-attribute-constraint mapping rule to form an entity association network.

3. The automatic recommendation method for standard terms according to claim 1, wherein The step of mapping the entity vectors and path encodings to a shared codebook through vector quantization for discrete semantic alignment and constructing a matrix-graph joint index that supports dynamic expansion includes: Train the entities in the knowledge graph based on the BERT model to obtain entity vectors; calculate the cosine similarity between entity vectors, construct an n-order semantic association matrix based on the number n of entity vectors, and use the matrix elements to represent the semantic similarity between entity vectors i and j; perform sparsification processing on the semantic association matrix, retain the top 5% of non-zero elements with high similarity, set the remaining elements to zero values, assign an identifier to each entity vector, establish a two-way hash mapping between the entity identifier and the matrix row and column indexes, and form a semantic association matrix index; Traverse the association paths between the entities in the knowledge graph, use the message passing mechanism of the graph neural network to aggregate the features of the nodes and edges on the path to generate path encodings; map the path encodings to a semantic space with the same dimension as the entity vectors through a graph embedding algorithm to form a graph structure index; Map the entity vectors and path encodings to a shared codebook through vector quantization for discrete semantic alignment, and construct a matrix-graph joint index that supports dynamic expansion.

4. A method for automatically recommending standard terms according to claim 1, characterized in that, The step of using a gated attention fusion mechanism to partition the user input data to obtain feature text and input intent includes: Obtain the entity type and annotation information of the user input data, where the annotation information includes the start and end position information of the entity in the original text, splice the annotation information and the user input text according to the start and end position information, and through the gated attention fusion mechanism, dynamically allocate the weight of the annotation information in the splicing according to the semantic relevance between the annotation information and the original text. The gated attention fusion formula is: , where represents the final vector after fusing gated attention and residual connections, is the vector obtained after encoding the original text, is the vector obtained after encoding the annotation information, is the parameter matrix, e is the natural constant, represents the weight of the annotation information in the concatenation process; Input the spliced feature text into an intent recognition model to obtain a user input intent vector.

5. The automatic recommendation method for standard terms according to claim 1, characterized in that The step of searching for candidate recommended terms associated with the input intent in the knowledge graph based on the matrix-graph joint index and graph sampling algorithm includes: Based on the keywords associated with the entity type and annotation information of the user input data, using the keywords as query points, locate the key entities of the keywords in the knowledge graph through matrix-graph joint indexing. Utilize the graph structure indexing to mine potential entities related to the key entities through the distance-aware subgraph sampling algorithm. The formula of the distance-aware subgraph sampling algorithm is as follows: , Among them represents the set of potential entities obtained by formula screening, where e represents a single entity in the knowledge graph subgraph. represents the subgraph of the knowledge graph. h represents the head entity, r represents the relationship, and t represents the tail entity. represents the head entity h embedding vector in the hyperbolic space. represents the embedding vector of the relationship r in the hyperbolic space. represents the embedding vector of the head entity t in the hyperbolic space. represents the screening threshold, which is the attention weight coefficient. For the located key entities and potential entities, search for their synonyms in the thesaurus, use these synonyms as new search keywords, search again in the knowledge graph, and integrate all the entities obtained from the two searches to remove duplicate data to form candidate recommended terms.

6. The automatic recommendation method for standard terms according to claim 1, wherein The step of uniformly extracting the candidate recommended terms to obtain the recommended standard terms includes: Convert the entity vectors of the terms in the candidate recommended term set and the user input intention vector into hyperbolic space vectors, and calculate the semantic similarity between the candidate recommended term vector and the user intention vector using the hyperbolic distance formula. The hyperbolic distance formula is as follows: , Among them, S represents semantic similarity. The closer the value is to 1, the more similar it is. It represents the difference vector between the candidate recommended term vector x and the user input intention vector y. x is the candidate recommended term vector, and y is the square of the norm of the user input intention vector y. Used to adjust The influence degree of the function on the final similarity. That is, the inverse hyperbolic cosine function; Sort the candidate standard terms in descending order according to the semantic similarity, extract the candidate recommended terms with odd indices, and repeatedly extract until 15% of the candidate standard term set remains to obtain the recommended annotation terms.

7. The automatic recommendation method for standard terms according to claim 1, characterized in that The step of clustering the recommended standard terms and using the rime algorithm to aggregate the standard terms within the cluster to obtain the representative term of each cluster includes: Search to obtain the recommended standard terms, cluster the recommended standard terms, use the rime algorithm to optimize the parameters of the DBSCAN clustering algorithm, set the optimization objective function to minimize the inter-cluster similarity and maximize the intra-cluster weight variance. The intra-cluster weight value is the weighted sum of the TF-IDF scores of the standard terms in the recommended standard term set and the centrality of the knowledge graph of the standard terms. The centrality of the knowledge graph is calculated by the PageRank algorithm; The formula for minimizing the inter-cluster similarity is as follows: , Among them is the neighborhood radius, which is adjusted according to the distribution of terms in the vector space. The calculation formula is , where is the mean of the distances between all terms, is the standard deviation; is the minimum number of samples, which is set as a logarithmic function of the density of domain terms, , N is the total number of current domain terms, is used to control the importance of the target in the comprehensive optimization. k represents the number of clusters obtained after clustering. i ranges from 1 to k - 1, and j ranges from i + 1 to k, which are used to traverse different pairs of clusters, is the number of terms in cluster , is the number of terms in cluster , is the vector of the m-th term in cluster , is the vector of the n-th term in cluster ; The formula for maximizing the intra-cluster weight variance is as follows: , Among them is the neighborhood radius, which is adjusted according to the distribution of terms in the vector space. The calculation formula is , where is the mean of the distances between all terms, is the standard deviation; is the minimum number of samples, which is set as the logarithmic function of the density of domain terms, , N is the total number of current domain terms, is used to control the importance of the target in the comprehensive optimization. k represents the number of clusters obtained after clustering. i ranges from 1 to k and is used to traverse different cluster pairs, is the number of terms in cluster s, represents the set of weights of all terms in the i-th cluster, represents the i-th cluster the set of term weights within; Based on the frequency of occurrence of the standard terms within the cluster, select the standard term with the highest frequency of occurrence or the first occurrence in the thesaurus as the representative standard term of the current cluster, and incrementally update the representative term to the matrix-graph joint indexing.

8. The automatic recommendation method for standard terms according to claim 7, wherein The step of incrementally updating the representative term to the matrix-graph joint indexing includes: Adopt a delayed merging strategy. When the cumulative representative standard terms reach 10% of the total historical update data volume, compare the change rate of the local density kurtosis value of the representative standard terms and the most recent historical update data. If the change rate is less than 20%, use the representative standard terms to search for associated nodes in the knowledge graph, use the centrality of the knowledge graph of the representative terms as the entity attribute increment of the associated nodes to update the knowledge graph, and dynamically expand the shared codebook and update the matrix-graph joint indexing based on the representative terms through the vector algorithm. The formula for the local density kurtosis value is as follows: , where CDI represents the local density kurtosis value, m represents the number of samples, and T represents for each term data point , when its distance from the z point to be evaluated is less than or equal to 1, its corresponding subscript i will be included in the set T, is the mean of m , z represents the target vector of the deviation degree from the existing term data distribution, w represents the target vector of the deviation degree from the existing term data distribution, and k represents the dimension of the term vector.

9. The automatic recommendation method for standard terms according to claim 1, characterized in that The step of obtaining the recommended synonyms includes: When the candidate recommended terms are empty, based on the knowledge graph, obtain the keywords associated with the user input text according to the user input intention vector, input the keywords into the standard term thesaurus to obtain synonyms, and obtain the recommended synonyms.

Citation Information

Patent Citations

  • Aspect-level sentiment classification method based on gated convolutional neural network

    CN112784043A

  • Document classification method based on multi-element hypergraph gating attention network

    CN118394936A

Cited By

  • Self-evolution process entry routing closed-loop method and system based on user feedback

    CN122332500A

  • A self-evolving process entry routing closed-loop method and system based on user feedback

    CN122332500B

  • External information recommendation method and system based on electric power construction industry

    CN122594481A