Word vector optimization equipment name intelligent matching method based on equipment knowledge graph
By building equipment knowledge graph and word vector model, combined with multimodal similarity fusion technology, we can solve the problems of irregular naming of distribution network equipment and errors in OCR identification, realize high-precision device name matching, and improve the intelligence and data consistency of power grid management.
Patent Information
- Application Number
- CN202510278458.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-04
AI Technical Summary
The matching accuracy caused by irregular naming of distribution network equipment, OCR identification errors and semantic ambiguity, affecting data consistency and intelligence requirements.
Build a device knowledge graph, train the semantic features of the device name through the Word2Vec word vector model, calculate character similarity based on editing distance and pinyin similarity, calculate semantic similarity using cosine similarity, and weight the fusion character and semantic similarity to judge the device name matching relationship.
Significantly improve the accuracy of equipment name matching to 97%, reduce the risk of mismatch, support cross-system data integration, and improve the intelligence level of power grid scheduling and maintenance management.
Smart Images

Figure CN120256642A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system dispatching and maintenance management, and particularly to an intelligent matching method for optimizing device names with word vectors based on a device knowledge graph. Background Art
[0002] In the operation and management of the distribution network, the standardized matching of device names is a core link to ensure the execution of dispatching instructions, the formulation of maintenance plans, and the reliability of operation data. However, in practical applications, the diversity and non-standardization of device names lead to insufficient matching accuracy, which are specifically manifested as the following technical problems to be solved urgently:
[0003] 1. Matching ambiguity caused by the lack of device naming specifications:
[0004] Due to the lack of a unified naming standard, the same device has significant differences in different systems, documents, or units. For example, standard names such as "10kV #1 transformer" coexist with regional variants such as "10 kV No. 1 transformer" and "10kV No. 1 transformer"; redundant information such as "10kV #1 incoming line switch" simplified to "#1 incoming line" is mixed with non-standard abbreviations such as "circuit breaker" abbreviated as "switch". Such naming differences directly result in the inability to accurately associate device names, affecting cross-system data integration and dispatching decisions.
[0005] 2. OCR recognition errors and character noise interference:
[0006] When digitizing paper maintenance forms or handwritten records through OCR technology, device names are easily affected by the following problems; character form distortion: for example, "10kV #1 circuit breaker" is misrecognized as "10KV No. 1 section breaker"; confusion of similar characters: for example, "circuit breaker" and "section breaker" are misjudged due to similar glyphs. Such errors will introduce false device names, reduce data credibility, and even pose a risk of dispatching misoperation.
[0007] 3. Lack of semantic association and limitations of rule matching:
[0008] Traditional matching methods based on string comparison or regular expressions have the following defects; insufficient semantic understanding ability: unable to recognize the synonymous relationship between "10kV #1 transformer" and "10kV No. 1 transformer"; limited rule coverage: poor adaptability to undefined naming variants, such as new abbreviations and cross-unit naming habits; lack of context association: ignoring the physical connection or attribution relationship between devices, such as the "circuit breaker - busbar" relationship and the "transformer - substation" hierarchy, resulting in a contradiction between the matching result and the actual topological structure.
[0009] In summary, the above problems together lead to the following technical bottlenecks: low matching accuracy: traditional methods, due to relying on manual rules and local features, are difficult to handle complex naming variations, and the matching accuracy rate is usually lower than 85%; insufficient intelligence level: lacking in-depth mining of the semantic associations and context relationships of devices, it cannot support the automation and intelligence requirements of power grid dispatching; poor data consistency: naming differences result in data fragmentation among the dispatching system, maintenance records, and asset ledgers, affecting fault location and operation and maintenance efficiency. Summary of the Invention
[0010] To solve the above problems of the prior art, the present invention provides an intelligent matching method for optimizing device names based on word vectors of a device knowledge graph, aiming to solve the problems of non-standard naming, OCR noise interference, and semantic ambiguity, and improve the standardization and intelligence level of distribution network device management.
[0011] The technical solution adopted by the present invention is as follows:
[0012] An intelligent matching method for optimizing device names based on word vectors of a device knowledge graph, the intelligent matching method for optimizing device names based on word vectors of a device knowledge graph includes the following steps:
[0013] Step 1 Knowledge graph construction: Extract the standard names, hierarchical relationships, connection relationships, and synonymous relationships of devices from the power grid asset management system, dispatching system database, and historical maintenance orders, and construct a device knowledge graph containing a triple structure; where the triple relationships include: belongs to, connects, and is synonymous;
[0014] Step 2 Word vector training: Construct a training corpus based on the device knowledge graph, and use the Word2Vec word vector model to perform semantic feature training on device names;
[0015] Step 3 Character similarity calculation: Calculate the character similarity between device names using the edit distance similarity + pinyin similarity;
[0016] Step 4 Semantic similarity calculation: Based on the trained device word vectors, use the cosine similarity to calculate the semantic matching degree of device names;
[0017] Step 5 Comprehensive decision-making: Combine the character similarity and semantic similarity to calculate the final matching score by weighting, and judge the matching relationship of device names.
[0018] Further, in the triple relationships of Step 1, "belongs to" represents the attribution relationship between a device and a substation, "connects" represents the physical connection relationship between switches, circuit breakers, and busbars, and "is synonymous" represents different naming methods of the same device in different units or systems.
[0019] Further, in step 2, triples are extracted from the knowledge graph and a text sequence is constructed as the training corpus; the training corpus includes: The NkV#X transformer belongs to the NkV substation X, the NkV#X circuit breaker is connected to the NkV#X bus, and the NkVX transformer is synonymous with the NkV#X transformer.
[0020] Further, in step 2, a word vector model is trained using the Skip-gram architecture of Word2Vec, and its parameter settings are: the dimension of the word vector is 128, the window size is 5, and the number of training epochs is 100;
[0021] The output after semantic feature training is to represent the training corpus through word vectors, so as to compare semantic similarity and judge the actual relationship.
[0022] Further, in step 3, the edit distance calculation uses the dynamic programming algorithm Levenshtein to calculate the minimum number of edit operations between two strings, so as to detect OCR misrecognition or spelling mistakes;
[0023] The calculation formula is:
[0024]
[0025] In the formula, D(i,j) represents the edit distance between the first i characters and j characters of the string; δ(S i ,t j ) represents the string replacement cost. If S i = t j , the cost is 0, otherwise it is 1;
[0026] The edit distance similarity, the calculation formula is:
[0027]
[0028] In the formula, S edit is the edit distance similarity, which measures the similarity between two strings and ranges from [0,1]; the edit distance is the minimum number of edit operations, that is, the number of insert, delete, and replace operations required to convert string A to string B; length A is the number of characters of string A, that is, the length of the original device name; length B is the number of characters of string B, that is, the length of the device name to be matched.
[0029] Further, in step 3, the pinyin similarity calculation is to convert Chinese characters into pinyin, and through the pinyin similarity, that is, the Levenshtein distance, it is determined whether there is a mistake of similar-shaped characters;
[0030] If the pinyin similarity ≥ 90%, it is regarded as a potential match;
[0031] The pinyin similarity calculation, its calculation formula is:
[0032]
[0033] Wherein, S pinyin is the pinyin similarity, which is used to measure the similarity after converting two Chinese character strings into pinyin, and the value range is [0,1]; the pinyin edit distance is the minimum number of operations for pinyin conversion, that is, the number of insert, delete, and replace operations required to convert pinyin A into pinyin B; the pinyin length A is the number of characters in the pinyin string A, that is, the length after converting the original device name into pinyin; the pinyin length B is the number of characters in the pinyin string B, that is, the length after converting the device name to be matched into pinyin.
[0034] Finally, perform the comprehensive character similarity calculation:
[0035] The calculation formula is:
[0036]
[0037] Wherein, S char is the comprehensive character similarity.
[0038] Furthermore, in step 4, the calculation formula for cosine similarity is as follows:
[0039]
[0040] Wherein, S sem (A,B) is the semantic similarity score, which is used to measure the semantic relevance of the names of device A and device B, and the value range is [-1,1]; V A and V B represent the word vectors of device A and device B, which are 128-dimensional semantic feature vectors obtained through knowledge graph training; V A *V B represents the dot product of vectors, which is used to calculate the similarity of two vectors in direction; ||V A || and ||V B || represent the modulus length of the vector, which is the square root of the sum of the squares of each dimension of the vector.
[0041] Furthermore, in step 5, the weighted calculation formula for character similarity and semantic similarity is as follows:
[0042] S final = αS char + βS sem
[0043] Wherein, S final is the weighted score; S char is the character similarity, that is, the edit distance similarity + pinyin similarity; S sem is the semantic similarity, that is, the cosine similarity; α is the character weight; β is the semantic weight.
[0044] Furthermore, in step 5, the matching relationship of the device name is judged based on a set matching threshold.
[0045] If S final > 0.90, the device name matching is successful, and the standardized name is output.
[0046] If 0.75 < S_final < 0.90, manual review is entered, and the device name is confirmed by the grid dispatcher.
[0047] If S final < 0.75, the matching fails.
[0048] Furthermore, based on the judged matching relationship of the device name, the successfully matched device names are stored in the dispatching system or the maintenance management system.
[0049] Advantages of the present invention:
[0050] The intelligent matching method for optimizing device names based on word vectors of the device knowledge graph proposed by the present invention solves the core pain points in device name matching through multi-dimensional technology integration. Its technical advantages are specifically reflected in the following aspects:
[0051] 1. Significantly improve the matching accuracy: Through multi-modal similarity fusion, use character similarity to correct OCR errors and spelling variants, covering character-level differences; capture the deep associations of naming variants through semantic similarity, and solve the problems of synonyms and abbreviations. Through experiments, the matching accuracy rate can be increased from 85% of the traditional method to 97%, significantly reducing the risk of mis-matching.
[0052] 2. Strengthen the adaptability to complex naming problems: For the robustness of OCR noise, identify near-homophone errors through pinyin similarity, and reduce matching failures caused by character distortion. For cross-system naming compatibility, the "synonym" relationship of the knowledge graph and the semantic association of word vectors unify the naming habits of different units and systems.
[0053] 3. Improve the level of intelligent decision-making and automation: Through dynamic threshold determination, high-confidence matches are automatically associated with standardized names, reducing manual intervention; low-confidence matches trigger the review process to balance efficiency and security. For semantic reasoning ability, combine the "belong to" and "connect" relationships of the knowledge graph to avoid conflicts between matching results and the actual structure of the power grid.
[0054] 4. Support the full-scenario data governance of the power grid: For cross-system data integration, unify the device names in the dispatching system, maintenance orders, and asset ledgers, eliminate data islands, and improve the efficiency of fault location and operation and maintenance. The matching results are directly written into the dispatching system or the maintenance management system to ensure the reliability of subsequent instruction execution and data analysis.
[0055] In summary, the intelligent matching method for optimizing device names based on the device knowledge graph's word vectors realizes high-precision, strong robustness, and intelligence in device name matching through the technical closed-loop of knowledge graph + word vectors + multi-similarity fusion, providing core data support for distribution network dispatching, maintenance, and asset management, and promoting the digital transformation and intelligent upgrade of the power grid. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0057] Figure 1 It is a schematic flowchart of the intelligent matching method for optimizing device names based on the device knowledge graph's word vectors of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] Aiming at the problem of low matching accuracy of distribution network device names caused by non-standard naming, OCR recognition errors, and semantic ambiguities, this embodiment provides an intelligent matching method for optimizing device names based on the device knowledge graph's word vectors; as Figure 1 shown, the intelligent matching method for optimizing device names based on the device knowledge graph's word vectors includes the following steps:
[0060] Step 1 Knowledge graph construction:
[0061] The purpose of knowledge graph construction is to establish a standardized association network of device names through multi-source data integration, providing context basis for semantic matching. The device knowledge graph is used to store and manage the hierarchical structure, connection relationship, and synonymous relationship of distribution network devices, providing background information for device name matching.
[0062] First, data collection is carried out:
[0063] Extract the device standard names, hierarchical relationships, connection relationships, and synonymous relationships from the power grid asset management system, dispatching system database, and historical maintenance orders. Specifically, extract the device standard names, categories, and hierarchical relationships from the power grid asset management system; extract the dispatching instructions and device call information from the dispatching system database; extract the possible variants of device names, such as abbreviations and aliases, from the historical maintenance orders.
[0064] Then, perform relationship modeling:
[0065] Adopt the triple structure - subject, predicate, object - to represent the relationships between devices, such as:
[0066] ("110kV #1 Transformer", "belongs to", "110kV Substation A")
[0067] ("10kV #1 Circuit Breaker", "is connected to", "10kV #1 Busbar")
[0068] ("10kV No. 1 Transformer", "is synonymous with", "10kV #1 Transformer")
[0069] Among them, the triple relationships include: belongs to, is connected to, and is synonymous with;
[0070] Belongs to represents the attribution relationship between the device and the substation;
[0071] Is connected to represents the physical connection relationship between switches, circuit breakers, and busbars;
[0072] Is synonymous with represents different naming methods of the same device in different units or systems.
[0073] Finally, store the device knowledge graph:
[0074] Use a graph database, such as Neo4j, to store the constructed device knowledge graph, supporting fast retrieval of device association relationships and semantic reasoning.
[0075] Step 2: Word vector training:
[0076] The purpose of word vector training is to: learn the semantic features of device names through the context of the knowledge graph and solve the semantic association problem of naming variants. To obtain the semantic features of device names, in this embodiment, a word vector model based on the device knowledge graph is used for training. Specifically:
[0077] First, generate the corpus:
[0078] Convert the triples in the knowledge graph into text sequences to form the training corpus; for example:
[0079] 110kV #1 Transformer belongs to 110kV Substation A
[0080] The 10kV #1 circuit breaker is connected to the 10kV #1 busbar
[0081] 10kV No. 1 transformer is synonymous with 10kV #1 transformer
[0082] This corpus is used to train the word vectors of the device.
[0083] Then, model training is carried out:
[0084] In this embodiment, the Skip-gram architecture of the Word2Vec word vector model is used to train the word vector model for semantic feature training of device names; the word vector training parameters are: the word vector dimension is 128, the window size is 5, and the number of training epochs is 100.
[0085] The output after semantic feature training is to represent the training corpus through word vectors, so as to compare semantic similarities and judge actual relationships; the results are as follows:
[0086] "10kV #1 transformer" → [0.12, 0.05, …, -0.43];
[0087] "10kV No. 1 transformer" → [0.11, 0.83, …, -0.41]; Highly similar to the standard name vector.
[0088] Step 3 Character similarity calculation:
[0089] The purpose of character similarity calculation is: Character similarity calculation is for problems such as possible spelling mistakes, OCR misidentifications, and abbreviations in device names; through character-level and pinyin-level similarities, correct OCR errors, spelling variants, and abbreviation problems to ensure that similar device names can be correctly matched.
[0090] Character similarity calculation mainly uses the edit distance + pinyin similarity calculation method to calculate the character similarity between device names; specifically as follows:
[0091] First, edit distance calculation is carried out:
[0092] Edit distance is used to measure the minimum number of edit operations between two strings, such as: insertion, deletion, replacement, so as to detect OCR misidentifications or spelling mistakes. Edit distance calculation uses the dynamic programming algorithm Levenshtein to calculate the minimum number of edit operations between two strings, and the calculation formula is:
[0093]
[0094] In the formula, D(i,j) represents the edit distance between the first i characters and j characters of the string; δ(S i ,t j ) represents the string replacement cost. If Si = t j If it is, the cost is 0; otherwise, it is 1.
[0095] For example, "10kV #1 Circuit Breaker" vs. "10kV Section 1 Circuit Breaker" → Edit Distance = 2.
[0096] The edit distance similarity, the calculation formula is:
[0097]
[0098] In the formula, S edit is the edit distance similarity, which measures the similarity between two strings, and the value range is [0, 1]; the edit distance is the minimum number of edit operations, that is, the number of insert, delete, and replace operations required to convert string A to string B; Length A is the number of characters in string A, that is, the length of the original device name; Length B is the number of characters in string B, that is, the length of the device name to be matched.
[0099] Then, perform the pinyin similarity calculation:
[0100] The pinyin similarity calculation is to convert Chinese characters into pinyin, and determine whether they are similar in form through the pinyin similarity, that is, the Levenshtein distance. Specifically:
[0101] Since OCR may cause misrecognition of similar-looking characters, such as:
[0102] "10kV #1 Circuit Breaker" is misrecognized as "10kV #1 Section Circuit Breaker"
[0103] Calculate the pinyin similarity by comparing the conversion of Chinese characters to pinyin:
[0104] "duàn lùqì" vs. "duàn lùqì" → Approximately the same;
[0105] If the pinyin similarity is higher than 90%, it can be determined as a possible matching item.
[0106] The pinyin similarity calculation, its calculation formula is:
[0107]
[0108] In the formula, S pinyin is the pinyin similarity, which is used to measure the similarity between two Chinese character strings after being converted into pinyin, and the value range is [0, 1]; the pinyin edit distance is the minimum number of operations for pinyin conversion, that is, the number of insert, delete, and replace operations required to convert pinyin A to pinyin B; Pinyin Length A is the number of characters in pinyin string A, that is, the length of the original device name after being converted into pinyin; Pinyin Length B is the number of characters in pinyin string B, that is, the length of the device name to be matched after being converted into pinyin.
[0109] Finally, perform comprehensive character similarity calculation:
[0110] The calculation formula is:
[0111]
[0112] In the formula, S char is the comprehensive character similarity.
[0113] Step 4 Semantic similarity calculation:
[0114] The purpose of semantic similarity calculation is to capture the deep semantic associations of device names based on word vectors and solve semantic ambiguity problems such as synonyms and abbreviations.
[0115] To match variants of device names in different units and systems, such as "10kV #1 transformer" vs "10 kV No. 1 transformer", the cosine similarity calculation of word vectors based on the device knowledge graph is adopted.
[0116] The calculation formula for cosine similarity is as follows:
[0117]
[0118] In the formula, S sem (A, B) is the semantic similarity score, used to measure the semantic relevance of the names of device A and device B, and the value range is [-1, 1]; V A and V B represent the word vectors of device A and device B, which are 128-dimensional semantic feature vectors obtained through knowledge graph training; V A *V B represents the dot product of vectors, which is used to calculate the similarity of the directions of two vectors; ||V A || and ||V B || represent the norms of the vectors, which are the square roots of the sum of the squares of each dimension of the vectors.
[0119] Assume: The word vector of device A, 10kV #1 transformer: V A = [0.12, 0.05,..., -0.43]; The word vector of device B, 10 kV No. 1 transformer: V B = [0.11, 0.83,..., -0.41];
[0120] Through dot product calculation: V A *V B = (0.12 × 0.11) + (0.05 × 0.83) +... + (-0.43 × -0.41) ≈ 12.5;
[0121] Through norm calculation: ||VB || ≈ 1.3;
[0122] Then the similarity score is: S sem (A, B) = 12.5 / (1.2 × 1.3) ≈ 0.98;
[0123] The semantic similarity is as high as 0.98, and it can be determined as different naming variants of the same device.
[0124] Similarly, for 10kV #1 circuit breaker vs 10kV #1 switch, the semantic similarity is 0.95, and it can be determined as approximate devices. For 10kV No. 1 transformer vs 35kV #1 transformer, the semantic similarity is 0.65, and it can be determined as having low similarity.
[0125] Step 5 Comprehensive decision-making:
[0126] The purpose of comprehensive decision-making is to achieve high-precision device name matching through multi-dimensional similarity fusion and dynamic threshold determination. In this step, the character similarity and semantic similarity are combined for weighted calculation of the final matching score, the matching relationship of the device names is judged, and the final matching judgment is formed.
[0127] First, calculate the weighted score:
[0128] The weighted calculation formula for character similarity and semantic similarity is as follows:
[0129] S final = αS char + βS sem
[0130] In the formula, S final is the weighted score; S char is the character similarity, that is, the edit distance similarity + pinyin similarity; S sem is the semantic similarity, that is, the cosine similarity; α is the character weight; β is the semantic weight.
[0131] Then, set the matching threshold:
[0132] In this embodiment, the set matching thresholds are 0.75 and 0.9;
[0133] If S final > 0.90, it is determined as the same device;
[0134] If 0.75 < S final < 0.90, manual review is recommended;
[0135] If S final < 0.75, the matching fails.
[0136] Based on the matching relationship of the device names, store the successfully matched device names in the scheduling system or the maintenance management system.
[0137] Hypothesis: The standard name is: 10kV#1 Transformer, and the name to be matched is: 10 kV No. 1 Transformer;
[0138] First, the edit distance:
[0139] Original string: 10kV#1 Transformer, its length = 10 vs. 10 kV No. 1 Transformer, its length = 9;
[0140] The minimum number of edit operations = 2, where, replace "#" with "No.", and delete "Transformer".
[0141] Edit distance similarity:
[0142] Then, process the pinyin similarity:
[0143] Convert Chinese characters to pinyin: 10kV#1 Transformer → shi kV#1bian ya qi; 10 kV No. 1 Transformer → shi qian fu1hao bian; The difference is: kV vs. qian fu, #1 vs. 1hao, missing "ya qi"; Pinyin edit distance = 3 + 2 + 2 = 7;
[0144] Pinyin similarity:
[0145] Then, calculate the comprehensive character similarity:
[0146]
[0147] Then, calculate the semantic similarity:
[0148] Cosine similarity of word vectors: Word vector of device A (10kV#1 Transformer): [0.12, 0.05,..., -0.43]; Word vector of device B (10 kV No. 1 Transformer): [0.11, 0.83,..., -0.41]; Then, cosine similarity: S sem ≈0.98;
[0149] Finally, calculate the final score and matching determination:
[0150] Weighted calculation: S final = αS char + βS sem = 0.6×0.4 + 0.4×0.98 = 0.632
[0151] According to the threshold rule: If S final > 0.90, then it is determined as the same device; If 0.75 < Sfinal <0.90, it is recommended for manual review; if S final <0.75, the matching fails. The weighted calculation result of this case is 0.632, so the matching fails, an error prompt is given, and the manual review mechanism is triggered; the device name with failed matching is marked as "to be confirmed" and is not written into the scheduling system or the maintenance management system for the time being.
[0152] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent replacements or changes, and should be covered within the protection scope of the present invention.
Claims
1. A method for optimizing the intelligent matching of device names with word vectors based on a device knowledge graph, characterized in that: The intelligent matching method for optimizing device names based on the device knowledge graph includes the following steps: Step 1: Knowledge graph construction: Extract the device standard names, hierarchical relationships, connection relationships, and synonymous relationships from the power grid asset management system, dispatching system database, and historical maintenance orders to construct a device knowledge graph containing a triple structure; among them, the triple relationships include: belongs to, connection, and synonymy; Step 2: Word vector training: Construct a training corpus based on the device knowledge graph and perform semantic feature training on the device names using the Word2Vec word vector model; Step 3: Character similarity calculation: Calculate the character similarity between device names using the edit distance + pinyin similarity; Step 4: Semantic similarity calculation: Based on the trained device word vectors, calculate the semantic matching degree of device names using the cosine similarity; Step 5: Comprehensive decision-making: Combine the character similarity and semantic similarity to calculate the final matching score weighted and judge the matching relationship of device names.
2. The method for intelligent matching of device names with optimized word vectors based on a device knowledge graph according to claim 1, wherein: In the triple relationship of Step 1, "belongs to" represents the attribution relationship between the device and the substation, "connection" represents the physical connection relationship between switches, circuit breakers, and busbars, and "synonymy" represents different naming methods of the same device in different units or systems.
3. The method for intelligent matching of device names with optimized word vectors based on the device knowledge graph according to claim 1, wherein: In Step 2, extract triples from the knowledge graph and construct a text sequence as the training corpus; The training corpus includes: Transformer NkV#X belongs to Substation NkV X, Circuit Breaker NkV#X is connected to Busbar NkV#X, and Transformer NkV X is synonymous with Transformer NkV#X.
4. The method for intelligent matching of optimized device names with word vectors based on a device knowledge graph according to claim 3, characterized in that: In Step 2, use the Skip-gram architecture of Word2Vec to train the word vector model, and its parameter settings are: the dimension of the word vector is 128, the window size is 5, and the number of training epochs is 100; The output after semantic feature training is to represent the training corpus through word vectors, so as to compare the semantic similarity and judge the actual relationship.
5. The method for intelligent matching of device names with optimized word vectors based on a device knowledge graph according to claim 1, characterized in that: In Step 3, the edit distance calculation uses the dynamic programming algorithm Levenshtein to calculate the minimum number of edit operations between two strings to detect OCR misrecognition or spelling mistakes; The calculation formula is: Where D(i, j) represents the edit distance between the first i characters and the first j characters of the string; δ(S i , t j ) represents the string replacement cost. If S i = t j , the cost is 0; otherwise, the cost is 1. The edit distance similarity, the calculation formula is: Where S edit is the edit distance similarity, which measures the similarity between two strings and ranges from [0, 1]; the edit distance is the minimum number of edit operations, that is, the number of insert, delete, and replace operations required to convert string A into string B; length A is the number of characters in string A, that is, the length of the original device name; length B is the number of characters in string B, that is, the length of the device name to be matched.
6. The method for intelligent matching of device names with optimized word vectors based on a device knowledge graph according to claim 5, wherein: In Step 3, the pinyin similarity calculation is to convert Chinese characters into pinyin, and through the pinyin similarity, that is, the Levenshtein distance, judge whether it is a mistake of similar-shaped characters; If the pinyin similarity ≥ 90%, it is regarded as a potential match; The pinyin similarity calculation, its calculation formula is: where S pinyin is the pinyin similarity, which is used to measure the similarity after converting two Chinese character strings into pinyin, and the value range is [0, 1]; the pinyin edit distance is the minimum number of operations for pinyin conversion, that is, the number of insertion, deletion, and replacement operations required to convert pinyin A into pinyin B; The pinyin length A is the number of characters in the pinyin string A, that is, the length after converting the original device name into pinyin; The pinyin length B is the number of characters in the pinyin string B, that is, the length after converting the device name to be matched into pinyin. Finally, perform the comprehensive character similarity calculation: The calculation formula is: Where S char is the comprehensive character similarity.
7. The method for intelligent matching of device names with optimized word vectors based on a device knowledge graph according to claim 1, characterized in that: In Step 4, the calculation formula of the cosine similarity is as follows: Where S sem (A, B) is the semantic similarity score, which is used to measure the semantic relevance of the names of device A and device B, and the value range is [-1, 1]; V A and V B represent the word vectors of device A and device B, which are 128-dimensional semantic feature vectors obtained through knowledge graph training; V A *V B represents the dot product of vectors, which is used to calculate the similarity of two vectors in direction; ‖‖V A ‖‖ and ‖‖V B ‖‖ represent the modulus length of the vector, which is the square root of the sum of the squares of each dimension of the vector.
8. The method for intelligent matching of device names with optimized word vectors based on a device knowledge graph according to claim 1, characterized in that: In Step 5, the weighted calculation formula of the character similarity and semantic similarity is as follows: S final = αS char + βS sem Where S final is the weighted score; S char is the character similarity, i.e., the edit distance similarity + the pinyin similarity; S sem is the semantic similarity, i.e., the cosine similarity; α is the character weight; β is the semantic weight.
9. The method for intelligent matching of optimized device names with word vectors based on a device knowledge graph according to claim 8, characterized in that: In Step 5, judge the matching relationship of device names based on the set matching threshold; If S final > 0.90, the device name match is successful, and the standardized name is output; If 0.75 < S_final < 0.90, enter the manual review, and the power grid dispatcher confirms the device name; If S final <0.75, then the match fails.
10. The method for intelligent matching of device names with optimized word vectors based on a device knowledge graph according to claim 1, characterized in that: The intelligent matching method for device names with optimized word vectors based on the device knowledge graph stores the successfully matched device names into the scheduling system or the maintenance management system based on judging the matching relationship of the device names.
Citation Information
Cited By
Method for assisting commodity matching and code matching based on artificial intelligence vectorization technology
CN121093004A