Automatic standard term recommendation method
By constructing a knowledge graph and using matrix-graph joint index, combined with the gating attention fusion mechanism, the problems of information update lag and lack of accurate matching in the traditional term recommendation method are solved, and efficient and accurate term recommendation is achieved.
Patent Information
- Application Number
- CN202510431046.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The traditional term recommendation method has lag in information updates and lacks accurate matching mechanisms, resulting in messy recommendation results and it is difficult to quickly find professional data that meets the needs.
By collecting and preprocessing industry data, building a knowledge graph, using matrix-graph joint indexing and gated attention fusion mechanisms to match users' intentions, achieving efficient search and clustering analysis of candidate recommendation terms.
Improve the accuracy and efficiency of recommendation results, ensure that users can quickly find professional data that meets their needs, and that the recommendation system can be dynamically updated to adapt to industry changes.
Smart Images

Figure CN119938902A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automatic recommendation of standard terms, and in particular to a method for automatic recommendation of standard terms. Background Art
[0002] With the vigorous development of modern science and technology, the continuous growth of personal and industry data, the accumulation of more and more data in various fields, and the repeated expression of various professional terms and concepts, people have higher and more accurate requirements for data classification, especially for the associative recommendation function of professional data.
[0003] However, traditional terminology recommendation has significant problems, and the update of field information lags behind. Due to the lack of advanced screening and precise matching mechanisms, ordinary recommendation results are often presented to users in a disorderly manner, mixed with many contents that are irrelevant or have low relevance to people's needs, making it difficult to quickly find professional data recommendation content that truly meets their needs. Therefore, a standard terminology automatic recommendation method is needed. Summary of the invention
[0004] The object of the present invention is to provide a method for automatically recommending standard terms.
[0005] To achieve the above object, the present invention is implemented according to the following technical solutions: A first aspect of the present invention provides a method for automatically recommending standard terminology, comprising: collecting industry data to be evaluated, preprocessing the industry data to be evaluated to obtain standard terminology data, and semantically annotating entities in the standard terminology data; Building a knowledge graph based on the entities and semantic annotations, searching for synonyms using the semantic associations of the entities in the knowledge graph, and obtaining a synonym library; Calculate the cosine similarity between knowledge graph entities, construct a semantic association sparse matrix, identify the matrix index position of each entity, form a semantic association matrix index, traverse the knowledge graph to extract the entity association path and encode it, build a graph structure index, map the entity vector and path encoding to the shared codebook through vector quantization for discretized semantic alignment, and build a matrix-graph joint index that supports dynamic expansion; Use the gated attention fusion mechanism to divide the user input data to obtain feature text and input intent. Based on the matrix-graph joint index and graph sampling algorithm, search for candidate recommendation terms associated with the input intent in the knowledge graph, and evenly extract the candidate recommendation terms to obtain the recommended standard terms. If the search obtains recommended standard terms, cluster the recommended standard terms, use the rime algorithm to cluster the standard terms within the cluster, obtain the representative terms of each cluster and incrementally update them to the matrix-graph joint index; If no recommended standard term is found, the synonyms of the user input data are searched in the synonym database to obtain recommended synonyms.
[0006] As a further method, the step of constructing a knowledge graph based on the entity and semantic annotations includes: based on the entity and semantic annotations, extracting entity attributes and constraints, constructing a standard terminology ontology model, and using the standard terminology ontology model to inject entities in the terminology library into the knowledge graph according to the entity-attribute-constraint mapping rules to form an entity association network.
[0007] As a further method, the step of mapping the entity vector and the path encoding to the shared codebook through vector quantization for discretized semantic alignment and constructing a matrix-graph joint index that supports dynamic expansion includes: training the entities in the knowledge graph based on the BERT model to obtain the entity vectors; calculating the cosine similarity between the entity vectors, constructing an n-order semantic association matrix based on the number of entity vectors n, and using the matrix elements Characterize the semantic similarity between entity vectors i and j; perform sparse processing on the semantic association matrix, retain the top 5% of non-zero elements with high similarity, set the remaining elements to zero, assign an identifier to each entity vector, establish a bidirectional hash mapping between the entity identifier and the matrix row and column index, and form a semantic association matrix index; Traverse the association paths between entities in the knowledge graph, use the message passing mechanism of the graph neural network to aggregate the features of nodes and edges on the path, and generate path encoding; map the path encoding to the semantic space of the same dimension as the entity vector through the graph embedding algorithm to form a graph structure index; The entity vector and path encoding are mapped to a shared codebook through vector quantization for discretized semantic alignment, and a matrix-graph joint index that supports dynamic expansion is constructed.
[0008] As a further method, the step of using the gated attention fusion mechanism to divide the user input data to obtain feature text and input intention includes: obtaining the entity type and annotation information of the user input data, the annotation information includes the start and end position information of the entity in the original text, splicing the annotation information with the user input text according to the start and end position information, and dynamically allocating the weight of the annotation information in the splicing according to the semantic relevance between the annotation information and the original text through the gated attention fusion mechanism. The gated attention fusion formula is: , in represents the final vector after fusing gated attention and residual connection, is the vector obtained after encoding the original text, is the vector after the label information is encoded, is the parameter matrix, e is the natural constant, Represents the weight of the annotation information in the splicing process; The concatenated feature text is input into the intent recognition model to obtain the user input intent vector.
[0009] As a further method, the step of searching for candidate recommendation terms associated with the input intent in the knowledge graph based on the matrix-graph joint index and the graph sampling algorithm includes: based on the entity type of the user input data and the keyword associated with the annotation information, taking the keyword as the query point, locating the key entity of the keyword in the knowledge graph through the matrix-graph joint index, and using the graph structure index to mine potential entities related to the key entity through the distance-aware subgraph sampling algorithm based on the association path of the knowledge graph. The formula of the distance-aware subgraph sampling algorithm is: , in represents the potential entity set obtained after screening by the formula, e represents a single entity in the knowledge graph subgraph, represents a subgraph of the knowledge graph, h represents the head entity, r represents the relationship, and t represents the tail entity. Represents the header entity h The embedding vector in hyperbolic space, represents the embedding vector of relation r in hyperbolic space, represents the embedding vector of the head entity t in the hyperbolic space, represents the screening threshold, which is the attention weight coefficient; For the located key entities and potential entities, their synonyms are searched in the synonym library, and these synonyms are used as new search keywords to search again in the knowledge graph. All entities obtained from the two searches are integrated, duplicate data is removed, and candidate recommendation terms are formed.
[0010] As a further method, the step of uniformly extracting the candidate recommendation terms to obtain the recommended standard terms includes: converting the entity vectors of the terms in the candidate recommendation term set and the user input intention vector into hyperbolic space vectors, and calculating the semantic similarity between the candidate recommendation term vector and the user intention vector using a hyperbolic distance formula, wherein the hyperbolic distance formula is: , Where S represents semantic similarity. The closer the value is to 1, the more similar it is. Represents the difference vector between the candidate recommendation term vector x and the user input intention vector y, where x is the candidate recommendation term vector and y is the square of the norm of the user input intention vector y. For adjustment The degree of influence of the function on the final similarity, That is, the inverse hyperbolic cosine function; The candidate standard term set is sorted in descending order according to semantic similarity, and the odd-numbered candidate recommendation terms are extracted. The extraction is repeated until 15% of the candidate standard term set remains, and the recommended annotation terms are obtained.
[0011] As a further method, the step of clustering the recommended standard terms, clustering the standard terms within the cluster using the rime algorithm, and obtaining representative terms for each cluster includes: searching for recommended standard terms, clustering the recommended standard terms, optimizing the parameters of the DBSCAN clustering algorithm using the rime algorithm, setting the optimization objective function to minimize the similarity between clusters and maximize the variance of the weight within the cluster, the weight value within the cluster is the weighted sum of the TF-IDF score of the standard term in the recommended standard term set and the knowledge graph centrality of the standard term, the knowledge graph centrality is calculated by the PageRank algorithm, and the formula for minimizing the similarity between clusters is: , in is the neighborhood radius, which is adjusted according to the distribution of terms in the vector space. The calculation formula is: ,in is the mean distance between all terms, is the standard deviation; is the minimum number of samples, set to the logarithmic function of the domain term density, , N is the total number of terms in the current field, It is used to control the importance of the target in the comprehensive optimization. k represents the number of clusters obtained after clustering. i ranges from 1 to k-1, and j ranges from i+1 to k. It is used to traverse different cluster pairs. It is a cluster The number of internal terms, It is a cluster The number of internal terms, It is a cluster The vector of the mth term in , It is a cluster The vector of the nth term in ; The formula for maximizing the intra-cluster weight variance is: , in is the neighborhood radius, which is adjusted according to the distribution of terms in the vector space. The calculation formula is: ,in is the mean distance between all terms, is the standard deviation; is the minimum number of samples, set to the logarithmic function of the domain term density, , N is the total number of terms in the current field, It is used to control the importance of the target in the comprehensive optimization. k represents the number of clusters obtained after clustering. i ranges from 1 to k and is used to traverse different cluster pairs. is the number of terms in cluster s, represents the weight set of all terms in the i-th cluster, represents the i-th cluster The set of internal term weights; According to the frequency of occurrence of standard terms in the cluster, the standard terms with the highest frequency or the first appearance in the terminology library are selected as the representative standard terms of the current cluster, and the representative terms are incrementally updated to the matrix-graph joint index.
[0012] As a further method, the representative terms are incrementally updated to the matrix-graph joint index, including: adopting a delayed merging strategy, accumulating representative standard terms to reach 10% of the total amount of historical update data, comparing the change rate of the local density kurtosis value of the representative standard terms and the most recent historical update data, and using the representative standard terms to find the associated nodes in the knowledge graph if the change rate is less than 20%, using the knowledge graph centrality of the representative terms as the entity attribute of the associated nodes to incrementally update the knowledge graph, dynamically expanding the shared codebook based on the representative terms through a vector algorithm and updating the matrix-graph joint index, and the local density kurtosis value calculation formula is: , Where CDI represents the local density kurtosis value, m represents the number of samples, and T represents the data point for each term. , when the distance between it and the point z to be evaluated is less than or equal to 1, its corresponding subscript i will be included in the set T, There are m , z represents the target vector of the degree of deviation from the existing term data distribution, w represents the target vector of the degree of deviation from the existing term data distribution, and k represents the dimension of the term vector.
[0013] As a further method, the step of obtaining recommended synonyms includes: when the candidate recommendation term set is empty, obtaining keywords associated with the user input text based on the knowledge graph according to the user input intention vector, inputting the keywords into the standard term synonym library to obtain synonyms, and obtaining recommended synonyms.
[0014] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects: (1) The present invention obtains standard terminology data by summarizing and organizing the industry data to be evaluated, constructs a knowledge graph, and performs discretized semantic alignment based on the entity vectors and path encodings in the knowledge graph mapped to a shared codebook through vector quantization. A matrix-graph joint index that supports dynamic expansion is constructed. The matrix index supports fast vector similarity retrieval, while the graph index supports structure-based precise matching.
[0015] (2) The present invention uses a gated attention fusion mechanism to divide the user input data to obtain feature text and input intent, and searches for candidate recommendation terms associated with the input intent in the knowledge graph based on a matrix-graph joint index and a distance-aware subgraph sampling algorithm, thereby achieving efficient matching of user intent data.
[0016] (3) The present invention performs cluster analysis on the recommended standard terms, so that the hot words, high-frequency words and new words that people search for can be included in the matrix-graph joint index, so that the recommendation efficiency of the search hot words is further improved. By recommending synonyms for unknown keywords, people's recommendation needs can always be met. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 The figure is a flowchart of the steps of a method for automatically recommending standard terms in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0019] Reference Figure 1 As shown, the present invention provides a method for automatically recommending standard terms, comprising: Step S1 collects industry data to be evaluated, pre-processes the industry data to be evaluated to obtain standard terminology data, and semantically annotates entities in the standard terminology data; In this embodiment, power industry data is collected to obtain 100,000-level industry data, including power equipment maintenance records, power grid fault reports, power industry standard documents, technical specifications, etc.; the power industry data is cleaned and deduplicated, and entities are identified through the BERT-BiLSTM-CRF model, regular expressions are used to match power terms, and term disambiguation is performed in combination with the domain dictionary "Power Engineering Terminology" to obtain standard term data, and semantic annotation is performed on the entities of the standard term data, annotating equipment entities such as "main transformer" and "110kV busbar", annotating fault types such as "insulation breakdown" and "poor contact", and annotating parameter entities such as "rated current 1250A" and "power frequency withstand voltage 42kV". Step S2 constructs a knowledge graph based on the entities and semantic annotations, and uses the semantic associations of the entities in the knowledge graph to search for synonyms and obtain a synonym library; It should be explained that the steps of constructing the knowledge graph based on the entity and semantic annotations include extracting entity attributes and constraints based on the entity and semantic annotations, constructing a standard terminology ontology model, and using the standard terminology ontology model to inject entities in the terminology library into the knowledge graph according to the entity-attribute-constraint mapping rules to form an entity association network; industry terminology is complex, and the ontology model can clearly define and classify the terms, clearly define the attributes and relationships of each entity, and construct a hierarchical and structured knowledge system, organize the scattered terms and data in the industry, make the knowledge graph more systematic and comprehensive, provide an accurate semantic basis for the knowledge graph, and improve the practicality of the knowledge graph; In this embodiment, based on the entity and semantic annotation of standard terminology data, entities such as "110kVXX substation #1 main transformer" are imported from the equipment ledger, technical parameters such as "capacity 50MVA" and "connection group YN, d11" are associated, entity attributes and constraints are extracted, and the "circuit breaker" and "disconnector" establish a "series connection" relationship, and the "insulator" and "pollution flashover" establish a "fault causing" relationship. Based on the extracted attributes and constraint relationships, attributes such as "voltage level", "protection level" and "arc extinguishing medium" are set for entities such as "circuit breaker", and constraints such as "action time coordination" and "temperature rise limit" are set. A standard terminology ontology model is constructed, and the entities in the terminology library are injected into the knowledge graph according to the mapping rules of entity-attribute-constraint using the standard terminology ontology model to form an entity association network; by analyzing the nodes and edges related to entities such as "circuit breaker" and "relay protection" in the knowledge graph, combined with the standard terminology habits of the power industry, the synonyms of "circuit breaker" are screened out: "switch" and "relay protection" Synonyms of: "Relay protection", etc., to obtain the standard synonym database.
[0020] Step S3 calculates the cosine similarity between knowledge graph entities, constructs a semantic association sparse matrix, identifies the matrix index position of each entity, forms a semantic association matrix index, traverses the knowledge graph to extract entity association paths and encodes them, constructs a graph structure index, maps entity vectors and path encodings to a shared codebook through vector quantization for discretized semantic alignment, and constructs a matrix-graph joint index that supports dynamic expansion; In this embodiment, the cosine similarity between knowledge graph entities is calculated. For example, the cosine similarity between "circuit breaker" and "disconnector" is 0.85. The top 5% high similarity edges, such as equipment-fault associations, are retained. A semantic association sparse matrix is constructed. The pairwise similarity of 1,000 power terms such as "circuit breaker", "disconnector", and "relay protection" is calculated to generate a 1,000-order semantic matrix M. The matrix elements are Store the similarity values of terms i and j, count 499,500 similarity values, take the top 5% as high-similarity edges, and obtain 24,975 edges). In addition to "circuit breaker-disconnector (0.85)", we also retain edges such as "circuit breaker-short circuit disconnection (0.92)" (equipment-function association), "disconnector-electrical isolation (0.88)" (equipment-function association), and "circuit breaker-fault-contact burnout (0.81)" (equipment-fault association) to build a semantic association sparse matrix. We use edges with high similarity to fill the matrix. For example, the historical association probability between circuit breaker (index 1) and contact burnout fault (index 501) is 0.7, and the historical association probability between circuit breaker (index 1) and contact burnout fault (index 501) is 0.7. , only the index position and associated value of non-zero elements are stored to form the semantic association matrix index; Use depth-first search to traverse the knowledge graph to extract entity association paths and encode them. For example, the path "substation → main transformer → winding → overheating" is encoded as [0.3, 0.6, 0.2, 0.7]. Use the TransE knowledge graph embedding model to convert the path into a vector and build a graph structure index. Use the BERT model to train electricity terms such as "circuit breaker" to generate 768-dimensional vectors: , the K-Means clustering algorithm is used to cluster the 768-dimensional "circuit breaker" vector and the vectors of its related entities into the codebook entry C003, and the path encoding vector is clustered into C127, to achieve the semantic alignment of vector and path encoding, establish a shared codebook, store discretized semantic units, and construct a matrix-graph joint index that supports dynamic expansion.
[0021] Step S4 uses the gated attention fusion mechanism to divide the user input data to obtain feature text and input intent, searches for candidate recommendation terms associated with the input intent in the knowledge graph based on the matrix-graph joint index and graph sampling algorithm, and evenly extracts the candidate recommendation terms to obtain recommended standard terms; In this embodiment, the user inputs "find the common causes of 220kV line tripping", and the BERT-BiLSTM-CRF model recognizes the entity "220kV line" and the corresponding annotation information (equipment) [3,8], "tripping" and the corresponding annotation information (fault type) [9,11]. The annotation information is concatenated with the user input text based on the start and end position information. The gated attention fusion mechanism is used to obtain the weight of the annotation information in the concatenation based on the semantic relevance between the annotation information and the original text. The value is 1.02, which exceeds the benchmark value of 1, indicating that the annotation information of this splicing is highly correlated with the original text. The spliced feature text is input into the intent recognition model to obtain the user input intent vector. and input intention "220kV line fault analysis"; Based on the initial keywords associated with the entity type and annotation information of the user input data: "220kV line", "tripping", take the keywords as query points, locate the key entities of the keywords in the knowledge graph through the matrix-graph joint index, use the graph structure index based on the association path of the knowledge graph through the distance-aware subgraph sampling algorithm to mine potential entities related to the key entities, and the attention weight coefficient Set to 0.7, filter out The set of potential entities such as "relay protection malfunction" and "insulator contamination" with a probability less than 0.7 is used. For the located key entities and potential entities, their synonyms are searched in the synonym database. The synonyms such as "220 kV transmission line" and "220kV transmission line" are used as new search keywords and searched again in the knowledge graph. All entities obtained from the two searches are integrated, and duplicate data is removed to form candidate recommendation terms "220kV line, tripping, relay protection malfunction, insulator contamination, 220 kV transmission line, switch tripping, protection device malfunction, transmission line lightning protection measures, protection setting error, flashover fault". The entity vectors of the terms in the candidate recommendation term set and the user input intention vector are converted into hyperbolic space vectors. The semantic similarity between the candidate recommendation term vector and the user intention vector is calculated using the hyperbolic distance formula. The adjustment parameter , the semantic similarity S of "relay protection malfunction" is 0.656, the semantic similarity S of "insulator contamination" is 0.521, and the semantic similarity S of "line lightning protection design" is 0.789. The candidate standard terms are arranged in descending order according to the semantic similarity, and the odd-numbered candidate recommended terms are extracted. The extraction is repeated until 15% of the candidate standard term set remains, and the recommended annotation term "insulator contamination" is obtained.
[0022] It should be explained that when the selected recommended term is empty, in order to avoid the user being unable to obtain recommendations, the keywords associated with the user input text are directly obtained based on the knowledge graph according to the user input intention vector, and the keywords are input into the standard term synonym library to obtain synonyms for recommendation. This can make up for the matching time of the previous two searches and avoid users waiting. Directly recommending synonyms can enable users to refine further searches, improve users' recommendation experience, and ensure that the recommendation target can always hit user needs; In this embodiment, when the user inputs "GIS equipment abnormal heating" and there is no matching term, the keyword "GIS equipment" is extracted, and the synonym library is searched to obtain synonyms such as "GIS combination electrical appliance", "fully enclosed combination electrical appliance", "GIS equipment partial discharge", and "combination electrical appliance temperature rise test", and synonyms are used for recommendation.
[0023] Step S5: If the recommended standard terms are obtained through the search, cluster the recommended standard terms, use the rime algorithm to cluster the standard terms in the cluster, obtain the representative terms of each cluster and incrementally update them to the matrix-graph joint index; It should be explained that, since industry data is constantly being updated, the actual content of user searches is also constantly changing. For existing hot words and new terms, they need to be included in the terminology library and knowledge graph to optimize recommendation performance. By clustering the recommendation results of previous user searches, more refined hot word features and new word impressions can be obtained. By incrementally updating, the overall calculation of the index can be avoided, and the update time and resources can be reduced. By judging the local kurtosis value of the updated data and the index, the data updated each time has a better embedding degree in the index. In this embodiment, the candidate terms "relay protection malfunction", "insulator contamination", "line overload", and "CT saturation" are searched and obtained, and the recommended standard terms are clustered. The parameters of the DBSCAN clustering algorithm are optimized by the rime algorithm to improve the clustering accuracy. The optimization objective function is set to minimize the similarity between clusters and maximize the variance of the weights within the cluster. The formula is used to calculate The neighborhood radius is 0.386, which means that if the distance between two term vectors is less than 0.386, they are considered to be in the same neighborhood and the minimum number of samples is obtained. is 5.39, indicating that when clustering, a neighborhood contains at least 5 terms to form a valid cluster. The intra-cluster weight value is the weighted sum of the TF-IDF score of the standard term in the recommended standard term set and the knowledge graph centrality of the standard term. For example, the document frequency of "relay protection" is low, and the TF-IDF value is 0.7. In the knowledge graph, "relay protection" is closely related to entities such as "circuit breaker" and "fault analysis". The value calculated by the PageRank algorithm is 0.6, and the intra-cluster weight value is 0.66; Entities such as "relay protection", "circuit breaker", and "fault analysis" are divided into clusters based on the frequency of occurrence of standard terms in the cluster. ; "Insulator pollution" is divided into clusters , according to the frequency of occurrence of standard terms in the cluster, select "circuit breaker" in the cluster The frequency of occurrence is high, so it is considered a cluster Representative standard terms for "insulator pollution" are clustered representative standard terms; using the delayed merging strategy of SPFresh, after the cumulative representative standard terms reach the update threshold of 100, the change rate of the local density kurtosis value of the representative standard terms and the most recent historical update data is compared. If the change rate is less than 20%, the representative standard terms are used to find related nodes in the knowledge graph. After calculation, the local density kurtosis value CDI is 0.16, the most recent historical CDI is 0.14, and the change rate is 14%, which is less than 20%, and the CDI is 0.16 close to zero, which means that the index data updated this time has little impact on the overall index structure, and the overall data smoothness is high. The collection of representative standard terms such as "circuit breaker" and "insulator pollution" is dynamically expanded through the vector algorithm to share the code book and update the matrix-graph joint index for subsequent search.
[0024] In this embodiment, the step of using the gated attention fusion mechanism to divide the user input data to obtain feature text and input intent includes: The entity type and annotation information of the user input data are obtained. The annotation information includes the start and end position information of the entity in the original text. The annotation information is spliced with the user input text according to the start and end position information. Through the gated attention fusion mechanism, the weight of the annotation information in the splicing is dynamically allocated according to the semantic relevance between the annotation information and the original text. The gated attention fusion formula is: , in represents the final vector after fusing gated attention and residual connection, is the vector obtained after encoding the original text, is the vector after the label information is encoded, is the parameter matrix, e is the natural constant, Represents the weight of the annotation information in the splicing process; The concatenated feature text is input into the intent recognition model to obtain the user input intent vector.
[0025] In this embodiment, the step of searching the knowledge graph for candidate recommendation terms associated with the input intent based on the matrix-graph joint index and graph sampling algorithm includes: Based on the entity type of the user input data and the keywords associated with the annotation information, the keywords are used as query points, and the key entities of the keywords are located in the knowledge graph through the matrix-graph joint index. The graph structure index is used to mine potential entities related to the key entities through the distance-aware subgraph sampling algorithm based on the association path of the knowledge graph. The formula of the distance-aware subgraph sampling algorithm is as follows: , in represents the potential entity set obtained after screening by the formula, e represents a single entity in the knowledge graph subgraph, represents a subgraph of the knowledge graph, h represents the head entity, r represents the relationship, and t represents the tail entity. Represents the header entity h The embedding vector in hyperbolic space, represents the embedding vector of relation r in hyperbolic space, represents the embedding vector of the head entity t in the hyperbolic space, represents the screening threshold, which is the attention weight coefficient; For the located key entities and potential entities, their synonyms are searched in the synonym library, and these synonyms are used as new search keywords to search again in the knowledge graph. All entities obtained from the two searches are integrated, duplicate data is removed, and candidate recommendation terms are formed.
[0026] In this embodiment, the step of uniformly extracting the candidate recommendation terms to obtain the recommended standard terms includes: The entity vectors of the terms in the candidate recommendation term set and the user input intention vector are converted into hyperbolic space vectors, and the semantic similarity between the candidate recommendation term vector and the user intention vector is calculated using the hyperbolic distance formula, and the hyperbolic distance formula is: , Where S represents semantic similarity. The closer the value is to 1, the more similar it is. Represents the difference vector between the candidate recommendation term vector x and the user input intention vector y, where x is the candidate recommendation term vector and y is the square of the norm of the user input intention vector y. For adjustment The degree of influence of the function on the final similarity, That is, the inverse hyperbolic cosine function; The candidate standard term set is sorted in descending order according to semantic similarity, and the odd-numbered candidate recommendation terms are extracted. The extraction is repeated until 15% of the candidate standard term set remains, and the recommended annotation terms are obtained.
[0027] In this embodiment, the step of clustering the recommended standard terms, clustering the standard terms in the cluster using the rime algorithm, and obtaining representative terms of each cluster includes: Searching for recommended standard terms, clustering the recommended standard terms, optimizing the parameters of the DBSCAN clustering algorithm using the rime algorithm, setting the optimization objective function to minimize the similarity between clusters and maximize the variance of the weight within the cluster, and the weight within the cluster is the weighted sum of the TF-IDF score of the standard term in the recommended standard term set and the knowledge graph centrality of the standard term, and the knowledge graph centrality is calculated by the PageRank algorithm; The formula for minimizing the inter-cluster similarity is: , in is the neighborhood radius, which is adjusted according to the distribution of terms in the vector space. The calculation formula is: ,in is the mean distance between all terms, is the standard deviation; is the minimum number of samples, set to the logarithmic function of the domain term density, , N is the total number of terms in the current field, It is used to control the importance of the target in the comprehensive optimization. k represents the number of clusters obtained after clustering. i ranges from 1 to k-1, and j ranges from i+1 to k. It is used to traverse different cluster pairs. It is a cluster The number of internal terms, It is a cluster The number of internal terms, It is a cluster The vector of the mth term in , It is a cluster The vector of the nth term in ; The formula for maximizing the intra-cluster weight variance is: , in is the neighborhood radius, which is adjusted according to the distribution of terms in the vector space. The calculation formula is: ,in is the mean distance between all terms, is the standard deviation; is the minimum number of samples, set to the logarithmic function of the domain term density, , N is the total number of terms in the current field, It is used to control the importance of the target in the comprehensive optimization. k represents the number of clusters obtained after clustering. i ranges from 1 to k and is used to traverse different cluster pairs. is the number of terms in cluster s, represents the weight set of all terms in the i-th cluster, represents the i-th cluster The set of internal term weights; According to the frequency of occurrence of standard terms in the cluster, the standard terms with the highest frequency or the first appearance in the terminology library are selected as the representative standard terms of the current cluster, and the representative terms are incrementally updated to the matrix-graph joint index.
[0028] In this embodiment, the incremental updating of representative terms to the matrix-graph joint index includes: A delayed merging strategy is adopted, and the cumulative representative standard terms reach 10% of the total amount of historical update data. The change rate of the local density kurtosis value of the representative standard terms and the most recent historical update data is compared. If the change rate is less than 20%, the representative standard terms are used to find the associated nodes in the knowledge graph, and the knowledge graph centrality of the representative terms is used as the entity attribute increment of the associated nodes to update the knowledge graph. Based on the representative terms, the shared code book is dynamically expanded through the vector algorithm and the matrix-graph joint index is updated. The local density kurtosis value calculation formula is: , Where CDI represents the local density kurtosis value, m represents the number of samples, and T represents the data point for each term. , when the distance between it and the point z to be evaluated is less than or equal to 1, its corresponding subscript i will be included in the set T, There are m , z represents the target vector of the degree of deviation from the existing term data distribution, w represents the target vector of the degree of deviation from the existing term data distribution, and k represents the dimension of the term vector.
[0029] In summary, the embodiments of the present application aim to efficiently meet user needs by constructing an industry knowledge graph and a matrix-graph joint index, implement a standard terminology recommendation method that meets the timeliness of user needs through hyperbolic space recommendation capabilities, and update the index in a timely manner to improve the search efficiency of hot words, creating more professional and accurate recommendation options for people.
[0030] The above contents are merely examples and explanations of the structure of the present invention. The technicians in this technical field may make various modifications or additions to the specific embodiments described or replace them in a similar manner. As long as they do not deviate from the structure of the invention or exceed the scope defined by the claims, they should all fall within the protection scope of the present invention.
Claims
1. A method for automatically recommending standard terms, characterized in that: The following steps are involved: Collect the industry data to be evaluated, pre-process the industry data to be evaluated to obtain standard terminology data, and semantically annotate the entities of the standard terminology data; Building a knowledge graph based on the entities and semantic annotations, searching for synonyms using the semantic associations of the entities in the knowledge graph, and obtaining a synonym library; Calculate the cosine similarity between knowledge graph entities, construct a semantic association sparse matrix, identify the matrix index position of each entity, form a semantic association matrix index, traverse the knowledge graph to extract the entity association path and encode it, build a graph structure index, map the entity vector and path encoding to the shared codebook through vector quantization for discretized semantic alignment, and build a matrix-graph joint index that supports dynamic expansion; Use the gated attention fusion mechanism to divide the user input data to obtain feature text and input intent. Based on the matrix-graph joint index and graph sampling algorithm, search for candidate recommendation terms associated with the input intent in the knowledge graph, and evenly extract the candidate recommendation terms to obtain the recommended standard terms. If the search obtains recommended standard terms, cluster the recommended standard terms, use the rime algorithm to cluster the standard terms within the cluster, obtain the representative terms of each cluster and incrementally update them to the matrix-graph joint index; If no recommended standard term is found, the synonyms of the user input data are searched in the synonym database to obtain recommended synonyms.
2. The method for automatically recommending standard terms according to claim 1, characterized in that: The step of constructing a knowledge graph based on the entities and semantic annotations includes: Based on the entities and semantic annotations, entity attributes and constraints are extracted, and a standard terminology ontology model is constructed. The standard terminology ontology model is used to inject entities in the terminology library into the knowledge graph according to the entity-attribute-constraint mapping rules to form an entity association network.
3. The method for automatically recommending standard terms according to claim 1, characterized in that: The step of mapping the entity vector and the path encoding to a shared codebook through vector quantization for discretization semantic alignment and constructing a matrix-graph joint index supporting dynamic expansion includes: Based on the BERT model, the entities in the knowledge graph are trained to obtain entity vectors; the cosine similarity between entity vectors is calculated, and an n-order semantic association matrix is constructed based on the number of entity vectors n. The matrix elements are used Characterize the semantic similarity between entity vectors i and j; perform sparse processing on the semantic association matrix, retain the top 5% of non-zero elements with high similarity, set the remaining elements to zero, assign an identifier to each entity vector, establish a bidirectional hash mapping between the entity identifier and the matrix row and column index, and form a semantic association matrix index; Traverse the association paths between entities in the knowledge graph, use the message passing mechanism of the graph neural network to aggregate the features of nodes and edges on the path, and generate path encoding; map the path encoding to the semantic space of the same dimension as the entity vector through the graph embedding algorithm to form a graph structure index; The entity vector and path encoding are mapped to a shared codebook through vector quantization for discretized semantic alignment, and a matrix-graph joint index that supports dynamic expansion is constructed.
4. The method for automatically recommending standard terms according to claim 1, characterized in that: The step of using the gated attention fusion mechanism to divide the user input data to obtain feature text and input intent includes: The entity type and annotation information of the user input data are obtained. The annotation information includes the start and end position information of the entity in the original text. The annotation information is spliced with the user input text according to the start and end position information. Through the gated attention fusion mechanism, the weight of the annotation information in the splicing is dynamically allocated according to the semantic relevance between the annotation information and the original text. The gated attention fusion formula is: , in represents the final vector after fusing gated attention and residual connection, is the vector obtained after encoding the original text, is the vector after the label information is encoded, is the parameter matrix, e is the natural constant, Represents the weight of the annotation information in the splicing process; The concatenated feature text is input into the intent recognition model to obtain the user input intent vector.
5. The method for automatically recommending standard terms according to claim 1, characterized in that: The step of searching the knowledge graph for candidate recommendation terms associated with the input intent based on the matrix-graph joint index and graph sampling algorithm includes: Based on the entity type of the user input data and the keywords associated with the annotation information, the keywords are used as query points, and the key entities of the keywords are located in the knowledge graph through the matrix-graph joint index. The graph structure index is used to mine potential entities related to the key entities through the distance-aware subgraph sampling algorithm based on the association path of the knowledge graph. The formula of the distance-aware subgraph sampling algorithm is as follows: , in represents the potential entity set obtained after screening by the formula, e represents a single entity in the knowledge graph subgraph, represents a subgraph of the knowledge graph, h represents the head entity, r represents the relationship, and t represents the tail entity. Represents the header entity h The embedding vector in hyperbolic space, represents the embedding vector of relation r in hyperbolic space, represents the embedding vector of the head entity t in the hyperbolic space, represents the screening threshold, which is the attention weight coefficient; For the located key entities and potential entities, their synonyms are searched in the synonym library, and these synonyms are used as new search keywords to search again in the knowledge graph. All entities obtained from the two searches are integrated, duplicate data is removed, and candidate recommendation terms are formed.
6. The method for automatically recommending standard terms according to claim 1, characterized in that: The step of uniformly extracting the candidate recommended terms to obtain the recommended standard terms includes: The entity vectors of the terms in the candidate recommendation term set and the user input intention vector are converted into hyperbolic space vectors, and the semantic similarity between the candidate recommendation term vector and the user intention vector is calculated using the hyperbolic distance formula, and the hyperbolic distance formula is: , Where S represents semantic similarity. The closer the value is to 1, the more similar it is. Represents the difference vector between the candidate recommendation term vector x and the user input intention vector y, where x is the candidate recommendation term vector and y is the square of the norm of the user input intention vector y. For adjustment The degree of influence of the function on the final similarity, That is, the inverse hyperbolic cosine function; The candidate standard terms are arranged in descending order according to semantic similarity, and the odd-numbered candidate recommended terms are extracted. The extraction is repeated until 15% of the candidate standard term set remains, and the recommended annotation terms are obtained.
7. The method for automatically recommending standard terms according to claim 1, characterized in that: The step of clustering the recommended standard terms, clustering the standard terms in the cluster using the rime algorithm, and obtaining representative terms of each cluster includes: Searching for recommended standard terms, clustering the recommended standard terms, optimizing the parameters of the DBSCAN clustering algorithm using the rime algorithm, setting the optimization objective function to minimize the similarity between clusters and maximize the variance of the weight within the cluster, and the weight within the cluster is the weighted sum of the TF-IDF score of the standard term in the recommended standard term set and the knowledge graph centrality of the standard term, and the knowledge graph centrality is calculated by the PageRank algorithm; The formula for minimizing the inter-cluster similarity is: , in is the neighborhood radius, which is adjusted according to the distribution of terms in the vector space. The calculation formula is: ,in is the mean distance between all terms, is the standard deviation; is the minimum number of samples, set to the logarithmic function of the domain term density, , N is the total number of terms in the current field, It is used to control the importance of the target in the comprehensive optimization. k represents the number of clusters obtained after clustering. i ranges from 1 to k-1, and j ranges from i+1 to k. It is used to traverse different cluster pairs. It is a cluster The number of internal terms, It is a cluster The number of internal terms, It is a cluster The vector of the mth term in , It is a cluster The vector of the nth term in ; The formula for maximizing the intra-cluster weight variance is: , in is the neighborhood radius, which is adjusted according to the distribution of terms in the vector space. The calculation formula is: ,in is the mean distance between all terms, is the standard deviation; is the minimum number of samples, set to the logarithmic function of the domain term density, , N is the total number of terms in the current field, It is used to control the importance of the target in the comprehensive optimization. k represents the number of clusters obtained after clustering. i ranges from 1 to k and is used to traverse different cluster pairs. is the number of terms in cluster s, represents the weight set of all terms in the i-th cluster, represents the i-th cluster The internal term weight set; According to the frequency of occurrence of standard terms in the cluster, the standard terms with the highest frequency or the first appearance in the terminology library are selected as the representative standard terms of the current cluster, and the representative terms are incrementally updated to the matrix-graph joint index.
8. The method for automatically recommending standard terms according to claim 7, characterized in that: The step of incrementally updating the representative terms to the matrix-graph joint index comprises: A delayed merging strategy is adopted, and the cumulative representative standard terms reach 10% of the total amount of historical update data. The change rate of the local density kurtosis value of the representative standard terms and the most recent historical update data is compared. If the change rate is less than 20%, the representative standard terms are used to find the associated nodes in the knowledge graph, and the knowledge graph centrality of the representative terms is used as the entity attribute increment of the associated nodes to update the knowledge graph. Based on the representative terms, the shared code book is dynamically expanded through the vector algorithm and the matrix-graph joint index is updated. The local density kurtosis value calculation formula is: , Where CDI represents the local density kurtosis value, m represents the number of samples, and T represents the data point for each term. , when the distance between it and the point z to be evaluated is less than or equal to 1, its corresponding subscript i will be included in the set T, There are m , z represents the target vector of the degree of deviation from the existing term data distribution, w represents the target vector of the degree of deviation from the existing term data distribution, and k represents the dimension of the term vector.
9. The method for automatically recommending standard terms according to claim 1, characterized in that: The step of obtaining the recommended synonyms includes: When the candidate recommendation terms are empty, the keywords associated with the user input text are obtained based on the knowledge graph according to the user input intention vector, and the keywords are input into the standard term synonym library to obtain synonyms, and the recommended synonyms are obtained.
Citation Information
Patent Citations
Aspect-level sentiment classification method based on gated convolutional neural network
CN112784043A
Document classification method based on multi-element hypergraph gating attention network
CN118394936A
Construction method and device of knowledge base question-answering system, equipment and storage medium
CN119293164A
Potential relation reasoning-based medical knowledge graph retrieval system and method
CN119739867A
Hierarchical multi-task term embedding learning for synonym prediction
US20210012215A1
Cited By
File disassembling method based on multi-dimensional file association analysis and automatic interpretation
CN120631847A
A file decomposition method based on multi-dimensional file association analysis and automated interpretation
CN120631847B
Tube feeding nursing complication retrieval recommendation processing method and system
CN120929583A
Retrieval method and device applied to technical manager platform
CN121092647A
A search method and device applied to a technical manager platform
CN121092647B