Multi-source power dispatching data fusion method and device, equipment and storage medium
Through multi-dimensional similarity analysis and synonym knowledge graph, the problem of low accuracy of power scheduling data fusion in the existing technology is solved, and more efficient and reliable multi-source data fusion is achieved.
Patent Information
- Application Number
- CN202411984087.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
AI Technical Summary
The existing power scheduling data fusion methods cannot provide reasonable similarity results when processing highly specialized terms, resulting in low accuracy of multi-source data fusion.
A multi-dimensional similarity analysis method is adopted, including character similarity, sequence similarity and semantic similarity. By constructing a scheduling data synonym knowledge graph and synonym reasoning engine, the string similarity between the power scheduling data attributes is determined, and data fusion processing is carried out.
It improves the accuracy of string similarity between power scheduling data attributes, enhances the reliability of attribute matching and the credibility of data fusion, is more adaptable, and can more accurately fusion of multi-source power scheduling data.
Smart Images

Figure CN119989259A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of electric power technology, and in particular to a multi-source electric power dispatching data fusion method, device, computer equipment, computer-readable storage medium and computer program product. Background Art
[0002] With the development of new power systems, the types of elements such as equipment and software have increased significantly. The grid dispatching department needs to integrate massive data from different sources. Data fusion is crucial to ensure the safe and stable operation of the power system.
[0003] At present, power dispatching data fusion is implemented based on the existing string similarity calculation method. However, the existing string similarity calculation method is mainly designed for the common language on the Internet. When processing highly specialized terms related to power grid dispatching business, it often cannot provide reasonable similarity results, resulting in low accuracy of multi-source data fusion.
[0004] It can be seen that the current power system data fusion has the problem of low accuracy. Summary of the invention
[0005] Based on this, it is necessary to provide a multi-source power dispatching data fusion method, device, computer equipment, computer-readable storage medium and computer program product that can improve the accuracy of multi-source data fusion of power dispatching data in response to the above technical problems.
[0006] In a first aspect, the present application provides a multi-source power dispatching data fusion method, comprising:
[0007] Obtain power dispatch data from different data sources and integrate the power dispatch data belonging to the same entity into one entity record;
[0008] For the power dispatch data recorded by each entity, the following processing is performed:
[0009] Preprocessing of power dispatching data;
[0010] determining character similarity, sequence similarity, and semantic similarity between attributes of the preprocessed power dispatch data;
[0011] Determine the string similarity between attributes of power dispatch data based on character similarity, sequence similarity, and semantic similarity;
[0012] Based on the string similarity between the attributes of the power dispatching data, the power dispatching data is fused.
[0013] In one embodiment, determining character similarity between attributes of power dispatch data includes:
[0014] Taking two attributes as a group, the two attributes are processed as follows to obtain the character similarity between the attributes of the power dispatching data:
[0015] Determine the string length of each attribute, and the minimum edit distance between two attributes;
[0016] Determines the character similarity between two attributes based on the maximum of the minimum edit distance and the string length.
[0017] In one embodiment, determining the sequence similarity between attributes of power dispatch data includes:
[0018] Taking two attributes as a group, the two attributes are processed as follows to obtain the sequence similarity between the attributes of the power dispatching data:
[0019] Determine the longest common subsequence between two attributes;
[0020] Determines the sequence similarity between two attributes based on the maximum of the longest common subsequence and string length.
[0021] In one embodiment, determining semantic similarity between attributes of power dispatch data includes:
[0022] For each attribute of the power dispatching data, the synonym attribute set of the attribute is determined based on the established dispatching data synonym knowledge graph and synonym inference engine;
[0023] Taking two attributes as a group, the two attributes are processed as follows to obtain the semantic similarity between the attributes of the power dispatching data:
[0024] Determine the similarity between synonymous attributes in the synonymous attribute sets of two attributes;
[0025] The highest synonymous attribute similarity is determined as the semantic similarity between the two attributes.
[0026] In one embodiment, for each attribute of the power dispatching data, based on the established dispatching data synonym knowledge graph and synonym inference engine, a synonym attribute set of the attribute is determined, including:
[0027] For each attribute of the power dispatching data, perform word segmentation on the attribute to obtain a word set of the attribute;
[0028] For each word in the word set, the synonyms of the word are filtered out from the scheduling data synonym knowledge graph to obtain the synonym set of the word;
[0029] For each attribute, a synonym is extracted from the synonym set of each word of the attribute to perform word concatenation to obtain a synonym attribute of the attribute. The synonym attribute set of the attribute contains multiple synonym attributes.
[0030] In one embodiment, determining the string similarity between attributes of the power dispatching data based on character similarity, sequence similarity, and semantic similarity includes:
[0031] Determine the string type of each attribute of the power dispatching data respectively;
[0032] Taking two attributes as a group, the two attributes are processed as follows to obtain the string similarity between the attributes of the power dispatching data:
[0033] Based on the string types of the two attributes, weights corresponding to the character similarity, sequence similarity and semantic similarity between the two attributes are determined respectively;
[0034] Based on the corresponding weights of character similarity, sequence similarity and semantic similarity, the character similarity, sequence similarity and semantic similarity are weightedly summed to determine the string similarity between two attributes.
[0035] In one embodiment, after determining the string similarity between the two attributes, the method further includes:
[0036] When the string similarity between the two attributes is higher than the preset string similarity threshold, the following processing is performed for each of the two attributes:
[0037] Sampling the attribute value of the attribute to obtain the attribute value sample of the attribute;
[0038] Determining attribute value similarity between attribute value samples of two attributes;
[0039] Determine the attributes whose attribute value similarity is higher than a preset attribute value similarity threshold as candidate fusion attributes;
[0040] The candidate fusion attribute with the highest attribute value similarity is fused with the attribute.
[0041] In one embodiment, the attribute value sample includes a plurality of attribute values; determining the attribute value similarity between the attribute value samples of two attributes includes:
[0042] Determine the similarity between each attribute value in two attribute value samples to obtain multiple attribute value similarities;
[0043] An average value of the plurality of attribute value similarities is determined as the attribute value similarity between the attribute value samples of the two attributes.
[0044] In a second aspect, the present application also provides a multi-source power dispatching data fusion device, comprising:
[0045] A data acquisition module, used to acquire power dispatch data from different data sources and integrate the power dispatch data belonging to the same entity into one entity record;
[0046] A data preprocessing module, used for preprocessing power dispatching data;
[0047] An attribute similarity determination module is used to determine the character similarity, sequence similarity and semantic similarity between the attributes of the preprocessed power dispatching data; based on the character similarity, sequence similarity and semantic similarity, determine the string similarity between the attributes of the power dispatching data;
[0048] The fusion processing module is used to perform fusion processing on the power dispatching data based on the string similarity between the attributes of the power dispatching data.
[0049] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps in any one of the above-mentioned multi-source power dispatching data fusion method embodiments are implemented.
[0050] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps in any one of the above-mentioned multi-source power dispatching data fusion method embodiments are implemented.
[0051] In a fifth aspect, the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps in any one of the above-mentioned multi-source power dispatching data fusion method embodiments.
[0052] The above-mentioned multi-source power dispatching data fusion method, device, computer equipment, computer-readable storage medium and computer program product are different from the traditional data fusion method based on general string similarity. Considering that most of the power dispatching data are highly specialized terms related to the power grid dispatching business, this application proposes a power dispatching data fusion process based on multi-dimensional similarity to improve the accuracy of multi-source data fusion of power dispatching data. Similarity analysis is performed from multiple dimensions of characters, sequences and semantics of power dispatching data attributes. According to the similarity of characters, sequences and semantic dimensions between the attributes of power dispatching data, the similarity between the attributes of power dispatching data is comprehensively analyzed to obtain string similarity. On the one hand, it can capture attributes with different names but the same meaning, improve the accuracy of string similarity between power dispatching data attributes, and help improve the accuracy of attribute matching based on string similarity between attributes. On the other hand, compared with the similarity of attributes in a single dimension, multi-dimensional similarity calculation is conducive to reducing the impact of errors in a single dimension, improving the reliability of attribute matching, and thus improving the credibility of data fusion. On the other hand, the adaptability of multi-source data fusion in the field of power dispatching containing rich professional terms is improved. Therefore, multi-source data fusion processing is performed according to the similarity calculation results of the attributes of the power dispatching data, which improves the fusion accuracy of the multi-source power dispatching data, is beneficial to improving the reliability of power data management, and further helps to ensure the stability of power system operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0054] Figure 1 An application environment diagram of a multi-source power dispatching data fusion method in one embodiment;
[0055] Figure 2 A schematic diagram of a flow chart of a multi-source power dispatching data fusion method in one embodiment;
[0056] Figure 3 A schematic diagram of a flow chart of a multi-source power dispatching data fusion method in another embodiment;
[0057] Figure 4 A schematic diagram of a flow chart of a multi-source power dispatching data fusion method in yet another embodiment;
[0058] Figure 5A schematic diagram of a flow chart of a multi-source power dispatching data fusion method in yet another embodiment;
[0059] Figure 6 A schematic diagram of a flow chart of a multi-source power dispatching data fusion method in another embodiment;
[0060] Figure 7 It is a structural block diagram of a multi-source power dispatching data fusion device in one embodiment;
[0061] Figure 8 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0063] The multi-source power dispatching data fusion method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the terminal 102 communicates with the server 104 through a network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers.
[0064] Specifically, the operator may upload the collected power dispatching data from different data sources to the server 104 through the terminal 102, and then send a data fusion message to the server 104 through the terminal 102. The server 104 obtains the power dispatching data from different data sources, and integrates the power dispatching data belonging to the same entity into one entity record. Secondly, for the power dispatching data under each entity record, the power dispatching data is pre-processed, and then the character similarity, sequence similarity and semantic similarity between the attributes of the pre-processed power dispatching data are determined. Based on the character similarity, sequence similarity and semantic similarity, the string similarity between the attributes of the power dispatching data is determined. Finally, based on the string similarity between the attributes of the power dispatching data, the power dispatching data is fused.
[0065] The terminal 102 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart TVs, smart car devices, projection devices, etc. The portable wearable devices may be smart watches, smart bracelets, etc. The server 104 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0066] In an exemplary embodiment, Figure 2 As shown in the figure, a multi-source power dispatching data fusion method is provided, which is applied to Figure 1 The server 104 in the example is used for explanation, and includes the following S100 to S500. Among them:
[0067] S100, acquiring power dispatching data from different data sources, and integrating the power dispatching data belonging to the same entity into one entity record.
[0068] The power dispatching data from different data sources include various types of power dispatching data from different entities and systems. The different data sources may be power plants, substation systems, and the like.
[0069] In practical applications, real-time and historical operating data of power generation facilities and substation facilities can be obtained through SCADA systems, energy management systems, distribution management systems, power plant control systems, smart meters, etc. Determine the unique identifier of each entity, such as power generation equipment ID, substation equipment ID, etc., and use the unique identifier of the entity for matching to identify the same entity. Integrate the power dispatch data belonging to the same entity into an entity record and store it in the data warehouse to ensure that there is only one complete record for an entity. Each entity contains multiple attributes. For example, power generation equipment includes multiple attributes such as equipment type, rated power, real-time output, and reference voltage.
[0070] For the power dispatch data recorded by each entity, the following processing is performed:
[0071] S200, pre-processing the power dispatching data.
[0072] In practical applications, data preprocessing such as data cleaning, data format standardization, and data alignment can be performed on power dispatch data. In particular, for each attribute of the power dispatch data, the string is segmented into multiple words according to the preset word segmentation rules, and all characters are uniformly converted to lowercase. For each string, a connector (such as an underscore "_") is added between the segmented words to connect them to update each attribute.
[0073] S300, determining character similarity, sequence similarity and semantic similarity between attributes of pre-processed power dispatching data.
[0074] Among them, character similarity mainly focuses on the degree of character-level matching between attributes. The character similarity between attributes of power dispatching data can be determined by methods such as Jaro-Winkler distance or Hamming distance. Sequence similarity mainly focuses on the internal structure and order between attributes. The sequence similarity between attributes can be determined by methods such as Smith-Waterman algorithm or Needleman-Wunsch algorithm. Semantic similarity focuses on the meaning and context between attributes. The semantic similarity between attributes can be determined by pre-trained language models (such as BERT model), TF-IDF (term frequency-inverse document frequency) and cosine similarity.
[0075] Considering that in the power dispatching business, there may be inconsistencies in the naming specifications of entity attributes in different scenarios. In order to identify and match attributes that point to the same attribute but have different names, identification and matching are performed through the multi-dimensional similarity calculation results between attributes.
[0076] In practical applications, for any two attributes, the Hamming distance can be used to determine the number of characters required to replace one attribute to convert it into another attribute, and the character similarity between the two attributes can be determined based on the number of characters required to be replaced; the Smith-Waterman algorithm can be used to find the fragments with high similarity between the two attributes, and the sequence similarity between the two attributes can be determined based on the fragments with high similarity; the deep learning ability of the BERT model is used to capture the semantic information between attributes and the contextual dependencies, and the semantic similarity between the attributes can be determined.
[0077] S400, determining the string similarity between the attributes of the power dispatching data based on the character similarity, the sequence similarity and the semantic similarity.
[0078] Among them, string similarity represents the similarity of multi-dimensional comprehensive evaluation between attributes.
[0079] In practical applications, weights can be assigned to character similarity, sequence similarity, and semantic similarity respectively. For any two attributes, the string similarity between the attributes is obtained by weighted summation based on the character similarity, sequence similarity, and semantic similarity between the attributes and the corresponding weights.
[0080] S500, performing fusion processing on the power dispatching data based on the string similarity between the attributes of the power dispatching data.
[0081] In practical applications, a threshold may be pre-set to determine whether two attributes point to the same attribute. For any two attributes, the string similarity between the attributes is compared with the preset threshold. When the string similarity is higher than the preset threshold, the two attributes are determined to point to the same attribute. Otherwise, they do not point to the same attribute. For attributes pointing to the same attribute, data fusion processing may be performed through data consistency check, data merging, etc. Integrate the power dispatching data under each attribute of the same entity to integrate data from different sources, types and time scales together to build a unified, complete and consistent data set to support the dispatching and operation decisions of the power system.
[0082] In the above-mentioned multi-source power dispatching data fusion method, different from the traditional data fusion method based on general string similarity, this application takes into account that most of the power dispatching data are highly specialized terms related to the power grid dispatching business. In order to improve the accuracy of multi-source data fusion of power dispatching data, a power dispatching data fusion process based on multi-dimensional similarity is proposed. Similarity analysis is performed from multiple dimensions of characters, sequences and semantics of power dispatching data attributes. According to the similarity of characters, sequences and semantic dimensions between the attributes of power dispatching data, the similarity between the attributes of power dispatching data is comprehensively analyzed to obtain string similarity. On the one hand, it can capture attributes with different names but the same meaning, improve the accuracy of string similarity between power dispatching data attributes, and help improve the accuracy of attribute matching based on string similarity between attributes. On the other hand, compared with the similarity of attributes in a single dimension, multi-dimensional similarity calculation is conducive to reducing the impact of errors in a single dimension, improving the reliability of attribute matching, and thus improving the credibility of data fusion. On the other hand, the adaptability of multi-source data fusion in the field of power dispatching containing rich professional terms is improved. Therefore, multi-source data fusion processing is performed according to the similarity calculation results of the attributes of the power dispatching data, which improves the fusion accuracy of the multi-source power dispatching data, is beneficial to improving the reliability of power data management, and further helps to ensure the stability of power system operation.
[0083] In an exemplary embodiment, determining the character similarity between the attributes of the power dispatching data includes S310 to S320. Among them:
[0084] Taking two attributes as a group, the two attributes are processed as follows to obtain the character similarity between the attributes of the power dispatching data:
[0085] S310: Determine the character string length of each attribute and the minimum edit distance between two attributes.
[0086] S320 , determining the character similarity between the two attributes based on the minimum edit distance and the maximum value of the character string length.
[0087] The minimum edit distance is the minimum number of edit operations required to transform one string into another string. The allowed edit operations include: replacing one character with another, inserting a character, and deleting a character.
[0088] Considering that in the power dispatching business, the attributes of the power dispatching data often use English abbreviations and a mixture of Chinese and English names, the similarity of the character dimension between the attributes is calculated by the minimum edit distance to identify the relationship between the full spelling and the abbreviation.
[0089] In practical applications, Levenshtein Distance (LD) is used to calculate the character similarity between attributes. For example, the character similarity between attribute ta and attribute tb is determined as follows:
[0090] set up Representation attributes and properties The minimum edit distance between express Before Characters and Before The minimum edit distance between characters. Thus, by recursively calculating the minimum edit distance between characters, the minimum edit distance between ta and tb is obtained. :
[0091]
[0092] In the formula, yes No. characters, yes No. characters, is an indicator function that is equal to 1 when p is true and 0 when it is false.
[0093] The initial state of the above formula is:
[0094]
[0095] In the formula, and Respectively represent attributes and properties The length of the string.
[0096] After obtaining the string lengths of the two attributes and the minimum edit distance between the two attributes, the character similarity between the two attributes can be determined based on the maximum value of the minimum edit distance and the string length:
[0097]
[0098] In the formula, Representation attributes and properties The similarity between characters.
[0099] In this embodiment, considering that the naming conventions for power dispatching data attributes may be different in different power dispatching scenarios, in order to improve the accuracy of multi-source data fusion, the character similarity between attributes is determined by the minimum edit distance between attributes, which is conducive to further evaluating the comprehensive similarity between attributes based on the character similarity, thereby facilitating improving the accuracy of determining the comprehensive similarity between attributes, and further facilitating multi-source data fusion based on the comprehensive similarity between attributes to improve data fusion accuracy.
[0100] In an exemplary embodiment, determining the sequence similarity between the attributes of the power dispatching data includes S330 to S340. Among them:
[0101] Taking two attributes as a group, the two attributes are processed as follows to obtain the sequence similarity between the attributes of the power dispatching data:
[0102] S330, determining the longest common subsequence between two attributes.
[0103] S340 , determining the sequence similarity between the two attributes based on the longest common subsequence and the maximum value of the string length.
[0104] Among them, sequence similarity is used to characterize the similarity of the sequence dimension between attributes.
[0105] In practical applications, by determining the longest common subsequence between attributes, the sequence features in the string containing abbreviations or a mixture of Chinese and English are identified to determine the similarity of the sequence dimension between the attributes. and The longest common subsequence between , by recursively calculating the properties and The common subsequence length between them is:
[0106]
[0107] In the formula, c[u][v] represents The first u characters of The length of the similarity between the first v characters.
[0108] After obtaining the longest common subsequence between the attributes, the sequence similarity between the two attributes can be determined based on the maximum value of the longest common subsequence and the string length:
[0109]
[0110] In the formula, Representation attributes and The sequence similarity between .
[0111] In this embodiment, by determining the longest common subsequence between attributes and calculating the sequence similarity between the attributes, it is beneficial to evaluate the comprehensive similarity between the attributes based on the sequence similarity, thereby facilitating improving the accuracy of the comprehensive similarity evaluation.
[0112] In an exemplary embodiment, Figure 3 As shown, determining the semantic similarity between the attributes of the power dispatching data includes S350 to S370. Among them:
[0113] S350, for each attribute of the power dispatching data, based on the constructed dispatching data synonym knowledge graph and synonym inference engine, determine the synonym attribute set of the attribute.
[0114] The dispatch data synonym knowledge graph includes professional terms in the dispatch business and synonyms corresponding to the professional terms. The synonym attribute set includes one or more synonym attributes of the attribute, and the synonym attribute is used to represent the same or similar meaning as the attribute.
[0115] Considering that the data attributes in the power dispatching business involve professional terminology, directly calculating the semantic similarity between attributes may have the problem of low accuracy. Therefore, through the constructed dispatching data synonymous knowledge graph and synonym inference engine, the professional terminology in the attributes of the power dispatching data is converted into common vocabulary, and the similarity calculation of the semantic dimension is performed based on the attributes after synonymous conversion to improve the accuracy and effectiveness of the semantic dimension similarity calculation.
[0116] In practical applications, a synonym knowledge graph for dispatching data can be pre-built. Specifically, professional terms can be obtained by extracting terms from technical manuals, operating procedures, fault reports and other documents within the power grid company, or by obtaining professional terms from industry standards and technical specifications. Use word segmentation tools to segment words and identify synonyms of professional terms. For example, synonym matching rules can be created based on common sense and term definitions in the power grid industry to identify synonyms of professional terms. It can also be through machine learning methods, training classifiers or using clustering algorithms to automatically discover potential synonym relationships and perform synonym identification. In combination with the characteristics of power dispatching data, the data structure of the dispatching data synonym knowledge graph (DKG) is defined as follows:
[0117]
[0118] in: Represents the entity set in DKG, including professional words and corresponding synonyms; R represents the synonym relationship set between professional words and corresponding synonyms in DKG; It is a collection of entity types, including systems and devices; Represents a set of synonymous relationship types; and are functions that map each synonym relation to its corresponding type. There are three main types of relationships:
[0119] (1) Direct association : If the synonym relationship type is ,but can always be considered A synonym for , which indicates that there is a clear and direct synonym relationship between two attributes, without considering additional contextual information:
[0120]
[0121] (2) Contextual association of system constraints : This synonymous relationship type includes properties such as "system". Indicates the property value of "system" gather, Inference The system context at that time is:
[0122]
[0123] For example, = "voltage", = "pressure", in the general context, these two words are not synonymous, but in a specific power system scenario, they may express the same meaning, that is, only when the current system context (sa) belongs to the predefined system constraint set (Sab), Are synonyms.
[0124] (3) Device type context association :This type of synonymous relationship includes attributes such as "device type". represents a set of values for the "Device Type" attribute, and For its elements, we have:
[0125]
[0126] For example, = "terminal", = "connection point" , in general scenarios, may not be synonymous, but for a specific type of equipment (da), such as "transformer", they may mean the same thing, that is, only when the current equipment type (da) belongs to a predefined equipment type set (Dab), Are synonyms.
[0127] Build a synonym inference engine to identify all synonyms of professional words in a specific scenario. The synonym inference engine can include identifying the system type and device type related to the professional word, searching for entities corresponding to the professional word, and the conditions and logic of the synonym search. Exemplarily, the synonym inference engine is used to identify professional words from DKG. All synonyms of Represents a set of recognized synonyms:
[0128] a) Identify the system type ss and device type ds related to professional terms;
[0129] b) Based on Search for the corresponding entity ea in the domain knowledge graph;
[0130] c) Perform a recursive forward search starting from ea until it reaches a synonym that can no longer be considered Synonyms of , all identified synonyms form a set , which includes itself.
[0131] In practical applications, for each attribute, each word in the attribute can be dispatched by scheduling the data synonym knowledge graph and the synonym string inference engine to obtain multiple synonyms for each word. According to the order of the words in the attribute, the synonyms of each word are combined to obtain multiple synonymous attributes of the attribute, thereby constructing a synonym attribute set for the attribute.
[0132] Taking two attributes as a group, the two attributes are processed as follows to obtain the semantic similarity between the attributes of the power dispatching data:
[0133] S360: Determine the similarity between the synonymous attributes in the synonymous attribute sets of the two attributes.
[0134] In practical applications, after obtaining the synonymous attribute sets of the attributes respectively, for any two attributes, the similarity between the synonymous attributes in the synonymous attribute sets of the two attributes can be determined by calculating the similarity between each synonymous attribute in the synonymous attribute set of one attribute and each synonymous attribute in the synonymous attribute set of the other attribute. Specifically, for each synonymous attribute, all the words in the synonymous attribute are extracted, each word is converted into a word vector through the GloVe model, and the average value of the word vectors of each word in the attribute is determined as the vector of the synonymous attribute:
[0135]
[0136] in, A vector representing synonymous attributes, is the number of words in the synonym attribute, It is the first of the synonymous attributes The word vector of each word.
[0137] Take two synonymous attributes as a group and process them as follows to get the similarity between them:
[0138] The similarity between synonymous attributes is calculated by cosine similarity:
[0139]
[0140] in, and Synonymous attributes and Vector.
[0141] S370: Determine the highest similarity of the synonymous attributes as the semantic similarity between the two attributes.
[0142] In practical applications, for any two attributes, after obtaining the similarities between the synonymous attributes in the synonymous attribute sets of the two attributes, the highest similarity of the synonymous attributes is determined as the semantic similarity between the two attributes. , can be constructed synonym attributes; for attributes , can be constructed Synonymous attributes. The similarity calculation combination of synonymous attributes is used to determine the highest similarity as the semantic similarity between the attributes.
[0143] In this embodiment, the limitations of semantic understanding of professional terms in multi-source power dispatching data fusion are taken into consideration. Through the constructed dispatching data synonym knowledge graph and synonym inference engine, synonyms and professional terminology variants in the power field are effectively identified, and the professional terms in the attributes of the dispatching data are converted into common vocabulary to obtain a set of synonymous attributes of the attributes. The semantic similarity between the attributes is determined based on the similarity between the synonymous attribute sets of the attributes, which improves the accuracy of the semantic similarity, thereby improving the accuracy of identifying data pointing to the same attribute based on semantic similarity, which is beneficial to improving the accuracy of multi-source power dispatching data fusion.
[0144] In order to improve the accuracy of semantic similarity between attributes, in an exemplary embodiment, Figure 4 As shown, S350 includes S352 to S356. Among them:
[0145] S352, for each attribute of the electric power dispatching data, perform word segmentation processing on the attribute to obtain a word set of the attribute.
[0146] S354, for each word in the word set, filter out synonyms of the word from the scheduling data synonym knowledge graph to obtain a synonym set of the word.
[0147] S356, for each attribute, extract a synonym from the synonym set of each word of the attribute to perform word concatenation to obtain a synonym attribute of the attribute, and the synonym attribute set of the attribute includes multiple synonym attributes.
[0148] In practical applications, for each attribute, word segmentation is performed according to the connector in the attribute to obtain a word set of the attribute. and , extract all the words in the attribute through word segmentation processing, and get the word set of the attribute and , containing n and m words respectively.
[0149] After obtaining the word set of the attribute, for each word in the word set, the synonyms of the word are filtered out from the scheduling data synonym knowledge graph to obtain the synonym set of the word. The implementation process of using the synonym inference engine to identify all synonyms of professional words in the scheduling data synonym knowledge graph in the above embodiment is referred to and will not be repeated here.
[0150] So, for the set Each word in , based on the synonym knowledge graph of scheduling data and the synonym inference engine, all synonyms are identified and the word Synonyms for For the collection Each word in , get the word Synonyms for .
[0151] After obtaining the synonym set of each word of the attribute, a synonym is extracted from the synonym set of each word of the attribute for word concatenation to obtain the synonym attribute of the attribute. According to different extracted synonyms, multiple synonym attributes of the attribute are obtained to construct the synonym attribute set of the attribute.
[0152] For example, for the attribute , a set of synonyms for each word from the attribute Extract a synonym from the attribute and compose the attribute according to the position of the word in the attribute. Synonymous combination of . Different combinations:
[0153]
[0154] For each synonym combination, the synonyms are concatenated to obtain the synonymous attributes of the attribute.
[0155] For attributes , from each set Extract a synonym from the , and form a synonym combination of attribute tb. There are Nb different combinations:
[0156]
[0157] For each synonym combination, the synonyms are concatenated to obtain the synonymous attributes of the attribute.
[0158] In this embodiment, by performing word segmentation processing on the attributes of the power dispatching data, a synonym set of each word in the attribute is searched from the dispatching data synonym knowledge graph, and for each attribute, a synonym attribute set of the attribute is obtained based on the synonym combination of each word in the attribute. This can solve the limitations of existing general methods in processing professional terms and is conducive to improving the accuracy of determining the semantic similarity between attributes.
[0159] To improve the accuracy of determining the string similarity between attributes, in an exemplary embodiment, Figure 5 As shown, S400 includes S420 to S460. Among them:
[0160] S420, respectively determine the character string type of each attribute of the power dispatching data.
[0161] Among them, the string type of the attribute can include abbreviation splicing type and complete phrase type. Among them, the abbreviation splicing type attribute is formed by splicing abbreviations, usually does not contain complete words, for example, the abbreviation splicing type attribute can include the d-axis transient reactance "Xdpp" and q-axis transient time constant "Tq0pp" of the synchronous generator; the complete phrase type attribute consists of one or more complete words, such as the name of the synchronous generator "Name" or the reference voltage "BaseVoltage".
[0162] In practical applications, for each attribute, the attribute vector is determined, the semantic similarity between the attribute and the standard vocabulary is determined, and the semantic similarity between the attribute and the standard vocabulary is compared with a preset semantic similarity threshold. If the semantic similarity is higher than the preset semantic similarity threshold, the semantic similarity between the attribute and the standard vocabulary is high, and the attribute is determined to be a complete phrase type; conversely, if the semantic similarity is not higher than the preset semantic similarity threshold, the semantic similarity between the attribute and the standard vocabulary is low, and the attribute is determined to be an abbreviation connection type. The standard vocabulary can be set according to industry standards.
[0163] Taking two attributes as a group, the two attributes are processed as follows to obtain the string similarity between the attributes of the power dispatching data:
[0164] S440 , based on the character string types of the two attributes, respectively determine weights corresponding to the character similarity, sequence similarity, and semantic similarity between the two attributes.
[0165] S460 , based on the weights corresponding to the character similarity, the sequence similarity and the semantic similarity, weighted summation is performed on the character similarity, the sequence similarity and the semantic similarity to determine the string similarity between the two attributes.
[0166] When determining the string similarity between attributes, the weights corresponding to the multi-dimensional similarities are dynamically adjusted according to the string type of the attribute. Specifically, considering that the semantic correlation between the attributes of the abbreviation splicing type is low, the characters and sequences between the two can be compared in detail, that is, the weights of character similarity and sequence similarity are increased, and the weight of semantic similarity is reduced. Exemplarily, the weight corresponding to character similarity is 0.4, the weight corresponding to sequence similarity is 0.4, and the weight of semantic similarity is 0.2. For complete phrase-type attributes, the weight of semantic similarity can be increased. Exemplarily, the weight corresponding to character similarity is 0.2, the weight corresponding to sequence similarity is 0.3, and the weight of semantic similarity is 0.5.
[0167] After obtaining the corresponding weights of character similarity, sequence similarity, and semantic similarity between attributes, dynamic weighted multidimensional string similarity calculation is implemented through the three dimensions of character, sequence, and semantics to obtain the string similarity S between attributes:
[0168]
[0169] in, Represent character similarity, sequence similarity and semantic similarity respectively. are weights, all greater than zero and sum to 1.
[0170] In this embodiment, based on the characteristics of the power dispatching data attributes, the attributes are classified into character strings, and weights of similarities in different dimensions are assigned according to the character string types, so as to obtain the character string similarities between the attributes, thereby improving the accuracy of measuring the character similarities between the attributes.
[0171] In order to improve the accuracy of attribute matching, in an exemplary embodiment, Figure 6 As shown, after S460, the multi-source power dispatching data fusion method further includes S620 to S680. Among them:
[0172] When the string similarity between the two attributes is higher than the preset string similarity threshold, the following processing is performed for each of the two attributes:
[0173] S620: Sample the attribute value of the attribute to obtain an attribute value sample of the attribute.
[0174] S640: Determine the attribute value similarity between the attribute value samples of two attributes.
[0175] S660: Determine the attributes whose attribute value similarity is higher than a preset attribute value similarity threshold as candidate fusion attributes.
[0176] S680: Fuse the candidate fusion attribute with the highest attribute value similarity with the attribute.
[0177] In practical applications, when the string similarity between two attributes is higher than a preset string similarity threshold, it indicates that the two attributes may point to the same attribute. In order to further verify whether the two attributes point to the same attribute, the similarity between their attribute values is used for verification. Specifically, the attribute values of the attributes are sampled to obtain attribute value samples of the attributes, and the attribute value samples may include multiple attribute values.
[0178] After obtaining the attribute value samples of the attributes, the types of the attribute values of the two attributes are determined, and the types of the attribute values include strings, integers, and floating-point numbers. For string types, the attribute value similarity between the attribute value samples can be determined by calculating the intersection and union of the two strings based on each attribute value sample as a set of non-repeating strings, calculating the Jaccard coefficient between the two strings based on the intersection and the union, and determining the Jaccard coefficient as the attribute value similarity between the attribute value samples of the two attributes. For integer and floating-point type attribute values, the attribute value similarity between the attribute value samples can be determined based on the Euclidean distance.
[0179] After obtaining the attribute value similarity between the attribute value samples of the two attributes, compare the attribute value similarity with the preset attribute value similarity threshold. If the attribute value similarity is higher than the preset attribute value similarity threshold, it indicates that the probability that the two attributes point to the same attribute is high. The attribute value whose attribute value similarity is higher than the preset attribute value similarity threshold is determined as a candidate fusion attribute.
[0180] After all candidate fusion attributes corresponding to the attribute are obtained, the candidate fusion attribute with the highest attribute value similarity is determined as the attribute requiring data fusion, and data fusion processing is performed on the data under the attribute.
[0181] In this embodiment, when the string similarity between attributes is higher than a preset string similarity threshold, the attribute values under the attributes are sampled to determine the attribute value similarity between the attribute value samples, and the attributes corresponding to the attributes that need to be fused are further determined, thereby improving the accuracy of attribute matching, as well as the accuracy and reliability of multi-source data fusion.
[0182] In an exemplary embodiment, the attribute value sample includes multiple attribute values, and S640 includes S642 to S644. Among them:
[0183] S642, determining the similarity between the attribute values in the two attribute value samples, and obtaining a plurality of attribute value similarities.
[0184] S644: Determine an average value of the plurality of attribute value similarities as the attribute value similarity between the attribute value samples of the two attributes.
[0185] In practical applications, for the similarity calculation between the attribute values in the attribute value samples between two attributes, each attribute value in the attribute value sample of one of the attributes can be similarly calculated with each attribute value in the attribute value sample of the other attribute. If the attribute value type of the two attributes is a string type, the character similarity, sequence similarity and semantic similarity between the attribute values can be determined, and the attribute value similarity between the attribute values is determined based on the character similarity, sequence similarity and semantic similarity. Specifically, referring to the method for determining the character similarity, sequence similarity and semantic similarity between attributes in the above-mentioned embodiment, the method for determining the string similarity between attributes based on the character similarity, sequence similarity and semantic similarity is not repeated here.
[0186] If the attribute values of two attributes are of integer or floating-point type, the similarity between the attribute values can be determined by the following formula: and Represents attribute values of the same entity from different data sources, and The similarity of attribute values between:
[0187]
[0188] After obtaining the attribute value similarities between multiple attribute values, the average value of the multiple attribute value similarities is determined as the attribute value similarity between the attribute value samples of the two attributes:
[0189]
[0190] in Indicates the number of attribute value samples.
[0191] In this embodiment, by calculating the similarity of multiple groups of attribute values and determining the attribute value similarity between attribute value samples, it is possible to reduce misjudgment caused by a single abnormal attribute value or noise data, which is beneficial to improving the accuracy and reliability of attribute verification.
[0192] In order to make a clearer description of the multi-source power dispatching data fusion method provided by the present application, a specific embodiment is described below, and the specific embodiment includes the following steps:
[0193] S1, obtains power dispatch data from different data sources and integrates the power dispatch data belonging to the same entity into one entity record.
[0194] For the power dispatching data recorded by each entity, execute S2 to S8:
[0195] S2, preprocessing the power dispatching data.
[0196] For the attributes of the preprocessed power dispatching data, execute S3 to S6:
[0197] S3, taking two attributes as a group, performs: determining the string length of each attribute and the minimum edit distance between the two attributes, and determining the character similarity between the two attributes based on the maximum value of the minimum edit distance and the string length.
[0198] S4, taking two attributes as a group, performs: determining the longest common subsequence between the two attributes, and determining the sequence similarity between the two attributes based on the maximum value of the longest common subsequence and the string length.
[0199] S5, for each attribute of the power dispatching data, perform word segmentation on the attribute to obtain a word set of the attribute, for each word in the word set, filter out synonyms of the word from the dispatching data synonym knowledge graph to obtain a synonym set of the word, for each attribute, extract a synonym from the synonym set of each word of the attribute to perform word concatenation to obtain a synonym attribute of the attribute, and the synonym attribute set of the attribute contains multiple synonym attributes.
[0200] S6, taking two attributes as a group, executing: determining the similarity between synonymous attributes in a synonymous attribute set of the two attributes, and determining the highest similarity of the synonymous attributes as the semantic similarity between the two attributes.
[0201] S7, respectively determine the string type of each attribute of the power dispatching data, take two attributes as a group, and execute: based on the string types of the two attributes, respectively determine the weights corresponding to the character similarity, sequence similarity and semantic similarity between the two attributes; based on the weights corresponding to the character similarity, sequence similarity and semantic similarity, perform weighted summation of the character similarity, sequence similarity and semantic similarity to determine the string similarity between the two attributes.
[0202] S8, when the string similarity between the two attributes is higher than a preset string similarity threshold, for each of the two attributes, perform the following steps: sampling the attribute values of the attributes to obtain attribute value samples of the attributes, determining the similarity between the attribute values in the two attribute value samples to obtain multiple attribute value similarities, determining the average value of the multiple attribute value similarities as the attribute value similarity between the attribute value samples of the two attributes, determining the attributes whose attribute value similarity is higher than a preset attribute value similarity threshold as candidate fusion attributes of the attributes, and fusing the candidate fusion attributes with the highest attribute value similarity with the attributes.
[0203] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0204] In an exemplary embodiment, Figure 7 As shown, a multi-source power dispatching data fusion device 600 is provided, comprising: a data acquisition module 610, a data preprocessing module 620, an attribute similarity determination module 630 and a fusion processing module 640, wherein:
[0205] The data acquisition module 610 is used to acquire power dispatching data from different data sources and integrate the power dispatching data belonging to the same entity into one entity record.
[0206] The data preprocessing module 620 is used to preprocess the power dispatching data.
[0207] The attribute similarity determination module 630 is used to determine the character similarity, sequence similarity and semantic similarity between the attributes of the preprocessed power dispatching data; based on the character similarity, sequence similarity and semantic similarity, determine the string similarity between the attributes of the power dispatching data.
[0208] The fusion processing module 640 is used to perform fusion processing on the power dispatching data based on the string similarity between the attributes of the power dispatching data.
[0209] In an exemplary embodiment, the attribute similarity determination module 630 is also used to group two attributes and perform the following processing on the two attributes to obtain the character similarity between the attributes of the power dispatching data: determine the string length of each attribute and the minimum edit distance between the two attributes; based on the minimum edit distance and the maximum value of the string length, determine the character similarity between the two attributes.
[0210] In an exemplary embodiment, the attribute similarity determination module 630 is also used to group two attributes and perform the following processing on the two attributes to obtain the sequence similarity between the attributes of the power dispatching data: determine the longest common subsequence between the two attributes; and determine the sequence similarity between the two attributes based on the longest common subsequence and the maximum value of the string length.
[0211] In an exemplary embodiment, the multi-source power dispatching data fusion device 600 further includes a synonymous attribute set determination module 650 for determining a synonymous attribute set of each attribute of the power dispatching data based on a constructed dispatching data synonym knowledge graph and a synonym inference engine.
[0212] The attribute similarity determination module 630 is also used to group two attributes and perform the following processing on the two attributes to obtain the semantic similarity between the attributes of the power dispatching data: determine the similarity between the synonymous attributes in the synonymous attribute set of the two attributes; and determine the highest similarity of the synonymous attributes as the semantic similarity between the two attributes.
[0213] In an exemplary embodiment, the synonymous attribute set determination module 650 is also used to perform word segmentation processing on each attribute of the power dispatching data to obtain a word set of the attribute; for each word in the word set, the synonyms of the word are screened out from the dispatching data synonym knowledge graph to obtain a synonym set of the word; for each attribute, a synonym is extracted from the synonym set of each word of the attribute to perform word splicing to obtain a synonymous attribute of the attribute, and the synonymous attribute set of the attribute includes multiple synonymous attributes.
[0214] In an exemplary embodiment, the attribute similarity determination module 630 is also used to determine the string type of each attribute of the power dispatching data respectively; taking two attributes as a group, the two attributes are processed as follows to obtain the string similarity between the attributes of the power dispatching data: based on the string type of the two attributes, the weights corresponding to the character similarity, sequence similarity and semantic similarity between the two attributes are determined respectively; based on the weights corresponding to the character similarity, sequence similarity and semantic similarity, the character similarity, sequence similarity and semantic similarity are weightedly summed to determine the string similarity between the two attributes.
[0215] In an exemplary embodiment, the multi-source power dispatching data fusion device 600 also includes a candidate fusion attribute determination module 660, which is used to, when the string similarity between two attributes is higher than a preset string similarity threshold, execute, for each of the two attributes: sampling the attribute value of the attribute to obtain an attribute value sample of the attribute; determining the attribute value similarity between the attribute value samples of the two attributes; and determining the attribute value samples whose attribute value similarity is higher than the preset attribute value similarity threshold as candidate fusion attributes of the attribute.
[0216] The fusion processing module 640 is further used to fuse the candidate fusion attribute with the highest attribute value similarity with the attribute.
[0217] In an exemplary embodiment, the candidate fusion attribute determination module 660 is further used to determine the similarity between each attribute value in two attribute value samples to obtain multiple attribute value similarities; and determine the average value of the multiple attribute value similarities as the attribute value similarity between the attribute value samples of the two attributes.
[0218] Each module in the multi-source power dispatching data fusion device 600 can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.
[0219] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 8 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a multi-source power dispatching data fusion method is implemented.
[0220] Those skilled in the art will understand that Figure 8The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0221] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps in any one of the above-mentioned multi-source power dispatching data fusion method embodiments are implemented.
[0222] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in any one of the above-mentioned multi-source power dispatching data fusion method embodiments are implemented.
[0223] In one embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the steps in any one of the above-mentioned multi-source power dispatching data fusion method embodiments.
[0224] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0225] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0226] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0227] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the present application. It should be noted that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A multi-source power dispatching data fusion method, characterized in that: The method comprises: Obtain power dispatch data from different data sources and integrate the power dispatch data belonging to the same entity into one entity record; For the power dispatch data recorded by each entity, the following processing is performed: Preprocessing the power dispatching data; Determining character similarity, sequence similarity, and semantic similarity between attributes of the preprocessed power dispatching data; Determining the string similarity between the attributes of the power dispatching data based on the character similarity, the sequence similarity and the semantic similarity; Based on the string similarity between the attributes of the power dispatching data, the power dispatching data is fused.
2. The method according to claim 1, characterized in that Determining the character similarity between the attributes of the power dispatching data includes: Taking two attributes as a group, the two attributes are processed as follows to obtain the character similarity between the attributes of the power dispatching data: Determining the string length of each of the attributes, and the minimum edit distance between two attributes; Based on the minimum edit distance and the maximum value of the character string length, the character similarity between the two attributes is determined.
3. The method according to claim 2, characterized in that Determining the sequence similarity between the attributes of the power dispatching data includes: Taking two attributes as a group, the two attributes are processed as follows to obtain the sequence similarity between the attributes of the power dispatching data: determining a longest common subsequence between the two attributes; The sequence similarity between the two attributes is determined based on the longest common subsequence and the maximum value of the string length.
4. The method according to claim 3, characterized in that Determining the semantic similarity between the attributes of the power dispatching data includes: For each attribute of the power dispatching data, based on the constructed dispatching data synonym knowledge graph and synonym inference engine, determine a synonym attribute set of the attribute; Taking two attributes as a group, the two attributes are processed as follows to obtain the semantic similarity between the attributes of the power dispatching data: Determining similarities between synonymous attributes in a set of synonymous attributes of the two attributes; The highest similarity of the synonymous attributes is determined as the semantic similarity between the two attributes.
5. The method according to claim 4, characterized in that For each attribute of the power dispatching data, based on the established dispatching data synonym knowledge graph and synonym inference engine, a synonym attribute set of the attribute is determined, including: For each attribute of the electric power dispatching data, performing word segmentation processing on the attribute to obtain a word set of the attribute; For each word in the word set, the synonyms of the word are screened out from the scheduling data synonym knowledge graph to obtain a synonym set of the word; For each of the attributes, a synonym is extracted from the synonym set of each word of the attribute to perform word concatenation to obtain a synonym attribute of the attribute, and the synonym attribute set of the attribute includes multiple synonym attributes.
6. The method according to claim 2, characterized in that The determining, based on the character similarity, the sequence similarity and the semantic similarity, the string similarity between the attributes of the power dispatching data comprises: Determine the character string type of each attribute of the power dispatching data respectively; Taking two attributes as a group, the two attributes are processed as follows to obtain the string similarity between the attributes of the power dispatching data: Based on the character string types of the two attributes, respectively determining weights corresponding to the character similarity, the sequence similarity, and the semantic similarity between the two attributes; Based on the weights corresponding to the character similarity, the sequence similarity and the semantic similarity, the character similarity, the sequence similarity and the semantic similarity are weightedly summed to determine the string similarity between the two attributes.
7. The method according to claim 4, characterized in that After determining the string similarity between the two attributes, the method further includes: When the string similarity between the two attributes is higher than a preset string similarity threshold, the following processing is performed for each of the two attributes: Sampling the attribute value of the attribute to obtain an attribute value sample of the attribute; Determining attribute value similarity between attribute value samples of the two attributes; Determine the attributes whose attribute value similarity is higher than a preset attribute value similarity threshold as candidate fusion attributes of the attribute; The candidate fusion attribute with the highest attribute value similarity is fused with the attribute.
8. The method according to claim 7, characterized in that The attribute value sample includes a plurality of attribute values; and determining the attribute value similarity between the attribute value samples of the two attributes includes: Determine the similarity between each attribute value in two attribute value samples to obtain multiple attribute value similarities; An average value of the plurality of attribute value similarities is determined as the attribute value similarity between the attribute value samples of the two attributes.
9. A multi-source power dispatching data fusion device, characterized in that: The device comprises: A data acquisition module, used to acquire power dispatch data from different data sources and integrate the power dispatch data belonging to the same entity into one entity record; A data preprocessing module, used for preprocessing the power dispatching data; An attribute similarity determination module, used to determine the character similarity, sequence similarity and semantic similarity between the attributes of the preprocessed power dispatching data; based on the character similarity, the sequence similarity and the semantic similarity, determine the string similarity between the attributes of the power dispatching data; The fusion processing module is used to perform fusion processing on the power dispatching data based on the string similarity between the attributes of the power dispatching data.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Trusted data sharing collaboration platform
CN120995464A