Power grid data processing method, electronic equipment, storage medium and program product

By classifying and clustering keywords in power grid data, and optimizing the classification results using multiple preset classification models, the problem of low accuracy caused by imbalance in keyword counts is solved, and the accuracy of categories with small keyword counts is improved.

CN120296413APending Publication Date: 2025-07-11HEYUAN POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510312789.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When processing power grid data, the existing classification model is prone to ignore important information about the classification with relatively small keywords when facing the imbalance of keywords under different classifications, resulting in a low classification accuracy.

Method used

By extracting keywords from the power grid data, using preset classification algorithms and clustering algorithms, first classify keywords to obtain the first category, then cluster the keywords of each category, and finally use the preset classification model to determine the target second category of the clustering results to ensure that the accuracy of categories with relatively small keywords is improved.

Benefits of technology

In the case of unbalanced keyword counts, the accuracy of categories with relatively small numbers of keywords determined is improved, ensuring the comprehensiveness and accuracy of the information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296413A_ABST
    Figure CN120296413A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a power grid data processing method, electronic equipment, a storage medium and a program product. The method comprises the following steps: extracting a plurality of keywords from power grid data; based on a first preset classification algorithm, classifying the plurality of keywords to obtain a first category to which each keyword belongs; for each first category, based on a preset clustering algorithm, clustering the plurality of keywords belonging to the first category according to the association relationship between the keywords to obtain at least one clustering result; and for each clustering result, based on a second preset classification algorithm and according to the features of the clustering result, determining the category of the clustering result from a plurality of known second categories, and taking the category as a target second category to which each keyword in the clustering result belongs. The method is used for achieving the effect of improving the accuracy of the determined category with relatively small keyword quantity when the quantity of the keywords under different classifications is unbalanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of power information technology and natural language processing, and in particular, to a method for processing power grid data, an electronic device, a storage medium, and a program product. Background Art

[0002] With the rapid development of the digital economy, it is necessary to retrieve valuable information related to keywords belonging to a classification from a large amount of power grid data (such as web news, log files, configuration files, etc. related to the power industry). Therefore, it is necessary to determine the classification to which the keywords belong.

[0003] In the related art, multiple keywords can be input into a classification model for classification to obtain the classification of each keyword.

[0004] However, when the classification model classifies multiple keywords, if the number of keywords in different classifications is unbalanced, the classification model may ignore the important information of the classification with relatively few keywords, resulting in low accuracy of these classifications. Summary of the Invention

[0005] Embodiments of the present application provide a method for processing power grid data, an electronic device, a storage medium, and a program product, so as to achieve the effect of improving the accuracy of the classification with relatively few keywords determined when the number of keywords in different classifications is unbalanced.

[0006] In a first aspect, an embodiment of the present application provides a method for processing power grid data, including:

[0007] Extracting multiple keywords from the power grid data;

[0008] Determining a first classification to which each of the multiple keywords belongs; wherein, the first classifications to which the multiple keywords belong are not completely the same; each of the first classifications includes at least one of the keywords;

[0009] For each of the first classifications, clustering the multiple keywords belonging to the first classification according to the keywords and the association relationships between the keywords to obtain at least one clustering result; wherein, each of the clustering results includes one or more of the keywords;

[0010] For each of the clustering results, determining a target second classification to which each of the keywords in the clustering result belongs from multiple known second classifications.

[0011] In a possible implementation manner, the determining the first classification to which each of the multiple keywords belongs includes:

[0012] Input the multiple keywords into a pre-trained first preset classification model. The first preset classification model classifies the multiple keywords according to the multiple keywords and the association relationships between the keywords to obtain the first category to which each of the multiple keywords belongs.

[0013] In a possible implementation manner, each of the first categories includes multiple first category labels; the first category labels within the same first category are not contradictory; the first category labels between different first categories do not overlap.

[0014] In a possible implementation manner, for each of the clustering results, determining the target second category to which each of the keywords in the clustering result belongs from multiple known second categories includes:

[0015] Input the multiple clustering results into a pre-trained second preset classification model. The second preset classification model determines a first candidate second category corresponding to each clustering result, and the first candidate second category belongs to the multiple known second categories;

[0016] For each clustering result, input each keyword in the clustering result into a pre-trained third preset classification model. The third preset classification model outputs a second candidate second category to which the keyword belongs, and the second candidate second category belongs to the multiple known second categories;

[0017] In response to the first candidate second category and the second candidate second category being consistent, take the first candidate second category as the target second category;

[0018] In response to the first candidate second category and the second candidate second category being inconsistent, replace the clustering result to which the keyword belongs with the clustering result corresponding to the second candidate second category, and re-determine the first candidate second category and the second candidate second category of the keyword.

[0019] In a possible implementation manner, each of the target second categories includes multiple second category labels; the second category labels within the same second category are not contradictory; the second category labels between different target second categories do not overlap.

[0020] In a possible implementation manner, it further includes:

[0021] Extract multiple training data from the training power grid data. Each training data corresponds to a keyword extracted from the training power grid data and the association relationship between this keyword and other keywords;

[0022] Input multiple pieces of the training data into a preset algorithm based on grounded theory, and determine initial categories for the multiple pieces of training data by the preset algorithm;

[0023] For each of the initial categories, receive a label for the initial category, and apply the label to the multiple pieces of training data belonging to the initial category;

[0024] Train an initial first preset classification model using the multiple pieces of training data to obtain the pre-trained first preset classification model.

[0025] In a second aspect, an embodiment of the present application provides a processing device for power grid data, including:

[0026] An extraction module, configured to extract multiple keywords from the power grid data;

[0027] A first determination module, configured to determine a first category to which each of the multiple keywords belongs; wherein, the first categories to which the multiple keywords belong are not completely the same; each of the first categories includes at least one of the keywords;

[0028] A clustering module, configured to, for each of the first categories, cluster the multiple keywords belonging to the first category according to the keywords and the association relationship between the keywords to obtain at least one clustering result; wherein, each of the clustering results includes one or more of the keywords;

[0029] A second determination module, configured to, for each of the clustering results, determine a target second category to which each of the keywords in the clustering result belongs from multiple known second categories.

[0030] In a possible implementation manner, the first determination module is specifically configured to:

[0031] Input the multiple keywords into a pre-trained first preset classification model, and classify the multiple keywords by the first preset classification model according to the multiple keywords and the association relationship between the keywords to obtain the first categories to which the multiple keywords belong.

[0032] In a possible implementation manner, each of the first categories includes multiple first category labels; the first category labels of the same first category are not contradictory; the first category labels between different first categories do not overlap.

[0033] In a possible implementation manner, the second determination module is specifically configured to:

[0034] Input multiple of the clustering results into a pre-trained second preset classification model, and use the second preset classification model to determine a first candidate second category corresponding to each clustering result, where the first candidate second category belongs to the multiple known second categories;

[0035] For each clustering result, input each keyword in the clustering result into a pre-trained third preset classification model, and use the third preset classification model to output a second candidate second category to which the keyword belongs, where the second candidate second category belongs to the multiple known second categories;

[0036] In response to the first candidate second category and the second candidate second category being the same, use the first candidate second category as the target second category;

[0037] In response to the first candidate second category and the second candidate second category being different, replace the clustering result to which the keyword belongs with the clustering result corresponding to the second candidate second category, and re-determine the first candidate second category and the second candidate second category of the keyword.

[0038] In a possible implementation manner, each of the target second categories includes multiple second category labels; the second category labels of the same second category are not contradictory; the second category labels of different target second categories do not overlap.

[0039] In a possible implementation manner, the apparatus further includes:

[0040] Extract multiple training data from the training power grid data, where each training data corresponds to a keyword extracted from the training power grid data and the association relationship between the keyword and other keywords;

[0041] Input the multiple training data into a preset algorithm based on grounded theory, and use the preset algorithm to determine initial categories for the multiple training data;

[0042] For each of the initial categories, receive a label for the initial category and apply the label to the multiple training data belonging to the initial category;

[0043] Use the multiple training data to train an initial first preset classification model to obtain the pre-trained first preset classification model.

[0044] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0045] The memory stores computer execution instructions;

[0046] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementation manners of the first aspect.

[0047] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the above first aspect and / or various possible implementation manners of the first aspect.

[0048] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above first aspect and / or various possible implementation manners of the first aspect.

[0049] The method for processing grid data, electronic device, storage medium, and program product provided by the embodiments of the present application first extracts a plurality of keywords from the grid data; secondly, based on a first preset classification algorithm, classifies the plurality of keywords to obtain the first category to which each keyword belongs; wherein, the first categories to which the plurality of keywords belong are not completely the same; each first category includes at least one keyword; then, for each first category, based on a preset clustering algorithm, clusters the plurality of keywords belonging to the first category according to the keywords and the association relationships between the keywords to obtain at least one clustering result; wherein, each clustering result includes one or more keywords; for each clustering result, based on a second preset classification algorithm, determines the category of the clustering result from a plurality of known second categories according to the characteristics of the clustering result, and uses it as the target second category to which each keyword in the clustering result belongs, and can implement determining the target second category to which the keywords in each clustering result belong. Therefore, when the number of keywords in different classifications is unbalanced, the accuracy of the category with a relatively small number of determined keywords is improved. Description of the Drawings

[0050] The drawings here are incorporated into the description and constitute a part of this description, showing embodiments consistent with the present application, and are used together with the description to explain the principles of the present application.

[0051] Figure 1 It is a schematic flowchart of a method for processing grid data provided by the present application;

[0052] Figure 2 It is a schematic flowchart of another method for processing grid data provided by the present application;

[0053] Figure 3 It is a schematic structural diagram of a device for processing grid data provided by the present application;

[0054] Figure 4Schematic diagram of the structure of the electronic device provided by this application.

[0055] Through the above-mentioned drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Specific Embodiments

[0056] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description involves the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.

[0057] With the rapid development of the digital economy, the amount of power grid data has grown explosively, and this data includes web news, log files, configuration files, etc. related to the power industry. In order to retrieve valuable information related to keywords belonging to a specific category from this vast amount of data, it is necessary to first determine the category to which the keywords belong.

[0058] In one example, multiple keywords can be classified based on a classification model to obtain the classification of each keyword.

[0059] However, when the classification model classifies multiple keywords, if the number of keywords under different classifications is unbalanced, it may be biased towards the characteristics of the classification with a large number of keywords and ignore the important information of the classification with a relatively small number of keywords, resulting in a low accuracy rate for these classifications.

[0060] This application provides a method for processing power grid data, an electronic device, a storage medium, and a program product to solve the above technical problems.

[0061] The following uses specific embodiments to detail the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the drawings.

[0062] Figure 1 Schematic flow diagram of a method for processing power grid data provided by this application, as Figure 1 shown, the method includes:

[0063] S101. Extract multiple keywords from the power grid data.

[0064] Exemplarily, the execution subject of this embodiment can be any device such as a server, a distributed system, a terminal device, other electronic devices / computer devices, and other devices that can implement the solution of this application, and there is no limitation thereto. Among them, the server can be an independent server or a server cluster, such as any form including but not limited to cloud servers, distributed servers, blockchain servers, etc.

[0065] This embodiment will be introduced with the execution subject being a server.

[0066] Exemplarily, power grid data can be obtained through multiple channels. For example, power-related web news data, social media data, etc. can be obtained from the Internet through data mining, or log files, configuration files, emails, etc. can be obtained from the power system.

[0067] In one example, the power grid data can be segmented, stop words can be removed, and then multiple keywords can be extracted. Among them, the keywords can be one or more of text, numbers, time, etc.

[0068] S102. Determine the first category to which each of the multiple keywords belongs; among them, the first categories to which the multiple keywords belong are not completely the same; each first category includes at least one keyword.

[0069] Exemplarily, after extracting multiple keywords from the power grid data, based on the first preset classification algorithm, the multiple keywords can be classified to obtain the first category to which each keyword belongs. Among them, the first preset classification algorithm can be any algorithm for classifying keywords, such as a classification model, regular expression matching, etc.

[0070] S103. For each first category, cluster the multiple keywords belonging to the first category according to the keywords and the association relationship between the keywords to obtain at least one clustering result; among them, each clustering result includes one or more keywords.

[0071] Exemplarily, after obtaining the first category to which each of the multiple keywords belongs, for each first category, based on the preset clustering algorithm, cluster the multiple keywords belonging to the first category according to the keywords and the association relationship between the keywords to obtain at least one clustering result. Among them, the association relationship between the keywords can be the context association relationship between the keyword and other keywords in the power grid data.

[0072] S104. For each clustering result, determine the target second category to which each keyword in the clustering result belongs from multiple known second categories.

[0073] Exemplarily, a plurality of known second categories can be predefined. After obtaining at least one clustering result for a plurality of keywords belonging to each first category, for each clustering result, based on a second preset classification algorithm, according to the characteristics of the clustering result, the category of the clustering result is determined from the plurality of known second categories, and it is used as the target second category to which each keyword in the clustering result belongs.

[0074] Among them, the second preset classification algorithm can be any algorithm used to classify the characteristics of the clustering result, such as a classification model, regular expression matching, etc. The characteristics of the clustering result characterize the common attributes or patterns of each cluster during the clustering process, such as the centroid of each cluster, high-frequency keywords, etc.

[0075] The method for processing power grid data provided in the embodiments of the present application extracts a plurality of keywords from the power grid data first; secondly, classifies the plurality of keywords based on a first preset classification algorithm to obtain the first category to which each keyword belongs; among them, the first categories to which the plurality of keywords belong are not completely the same; each first category includes at least one keyword; then, for each first category, based on a preset clustering algorithm, the plurality of keywords belonging to the first category are clustered according to the keywords and the association relationship between the keywords to obtain at least one clustering result; among them, each clustering result includes one or more keywords; for each clustering result, based on a second preset classification algorithm, according to the characteristics of the clustering result, the category of the clustering result is determined from the plurality of known second categories, and it is used as the target second category to which each keyword in the clustering result belongs, which can realize determining the target second category to which the keywords in each clustering result belong. Therefore, when the number of keywords in different classifications is unbalanced, the accuracy of the category with relatively few determined keywords is improved.

[0076] Figure 2 It is a schematic flowchart of another method for processing power grid data provided by the present application, as Figure 2 shown, the method includes:

[0077] S201. Extract a plurality of keywords from the power grid data.

[0078] Exemplarily, the execution subject of this embodiment can be any device such as a server, a distributed system, a terminal device, other electronic devices / computer devices, and other devices that can implement the solution of the present application, and there is no limitation thereto. Among them, the server can be an independent server or a server cluster, such as any form including but not limited to cloud servers, distributed servers, blockchain servers, etc.

[0079] This embodiment is introduced with the execution subject being a server.

[0080] Exemplarily, the implementation of step S201 is similar to that of step S101. For details, reference can be made to the description in step S101, which will not be elaborated here.

[0081] S202. Construct a pre-trained first preset classification model.

[0082] In one example, step S202 includes the following process:

[0083] Extract multiple training data from the training power grid data. Each training data corresponds to a keyword extracted from the training power grid data and the association relationship between this keyword and other keywords;

[0084] Input the multiple training data into a preset algorithm based on grounded theory, and the preset algorithm determines the initial categories for the multiple training data;

[0085] For each initial category, receive the label for this initial category and apply the label to the multiple training data belonging to this initial category;

[0086] Use the multiple training data to train the initial first preset classification model to obtain a pre-trained first preset classification model.

[0087] Exemplarily, the acquisition channels of the training power grid data are similar to those of the power grid data introduced in step S101 and will not be elaborated here. The label of each training data can be understood as the correct category to which each keyword belongs.

[0088] In one example, an initial first preset classification model can be constructed, and then the multiple training data and the label of each training data are input into the initial first preset classification model in batches for iterative training. In each iteration, the initial first preset classification model classifies the multiple training data in this batch to obtain the predicted category of each training data. Then, the predicted category of each training data and the label of this training data are input into the loss function to obtain a loss value. Through the backpropagation of the loss value in the initial first preset classification model, the parameters of the initial first preset classification model are adjusted until the model converges, and the initial first preset classification model with adjusted parameters is used as the pre-trained first preset classification model.

[0089] By inputting multiple training data into a preset algorithm based on grounded theory, the preset algorithm determines initial categories for the multiple training data. Then, for each initial category, a label for the initial category is received and applied to the multiple training data belonging to the initial category, so that more accurate labels for each training data can be obtained. Furthermore, based on the more accurate labels for each training data, the initial first preset classification model is trained and learned, and the parameters of the model can be adjusted more accurately. Thus, a pre-trained first preset classification model for more accurate classification can be obtained.

[0090] S203. Input multiple keywords into the pre-trained first preset classification model. The first preset classification model classifies the multiple keywords according to the multiple keywords and the association relationships between the keywords, and obtains the first category to which each keyword belongs.

[0091] Among them, the first categories to which the multiple keywords belong are not completely the same; each first category includes at least one keyword.

[0092] In one example, each first category includes multiple first category labels; the first category labels within the same first category are not contradictory; the first category labels between different first categories do not overlap.

[0093] Exemplarily, after extracting multiple keywords from grid data and constructing a pre-trained first preset classification model, the multiple keywords can be input into the pre-trained first preset classification model. The first preset classification model classifies the multiple keywords according to the multiple keywords and the context relationships between each keyword and other keywords in the grid data, and obtains the first category to which each keyword belongs.

[0094] In one example, after classifying the keyword "photovoltaic power generation", the first category to which "photovoltaic power generation" belongs includes the first category labels "renewable energy" and "clean energy technology".

[0095] By classifying multiple keywords based on the first preset classification model and obtaining multiple first category labels to which each keyword belongs, it is possible to perform retrievals respectively based on different first category labels to which the same keyword belongs, and valuable information related to the keyword can be retrieved in both cases.

[0096] S204. For each first category, cluster the multiple keywords belonging to the first category according to the keywords and the association relationships between the keywords, and obtain at least one clustering result; among them, each clustering result includes one or more keywords.

[0097] Exemplarily, the implementation manner of step S204 is similar to that of step S103. For details, reference can be made to the description in step S103, which will not be elaborated here.

[0098] By clustering multiple keywords belonging to each first category to obtain at least one clustering result, for the keywords under the first category, regardless of the number of keywords belonging to a category with a finer granularity than the first category, keywords with high semantic or feature similarity will form a separate clustering result.

[0099] S205. Input the multiple clustering results into a pre-trained second preset classification model, and the second preset classification model determines a first candidate second category corresponding to each clustering result. The first candidate second category belongs to multiple known second categories.

[0100] Exemplarily, different multiple known second categories can be predefined for different first categories. The multiple known second categories corresponding to a first category can be understood as a finer-grained classification of each keyword under the first category. For example, keyword A, keyword B, and keyword C can all belong to the first category "dynamic operation monitoring data". Keyword A can belong to the known second category "time reference parameter", keyword B can belong to the known second category "thermal power generation monitoring parameter", keyword C can belong to the known second category "clean energy power generation parameter", etc.

[0101] Different second preset classification models can be pre-trained for the multiple clustering results under different first categories, that is, based on the multiple known second categories corresponding to each first category, the initial second preset classification model corresponding to the first category is trained and learned to obtain the second preset classification model corresponding to the first category.

[0102] Among them, when training and learning the initial second preset classification model, the sample data input into the initial second preset classification model can be the sample data of multiple evenly distributed clustering results under the first category. Then, the initial second preset classification model can learn the characteristics of the sample data of each clustering result during training, and then adjust the model parameters to achieve more accurate classification of the sample data of each clustering result.

[0103] Then, for each first category, input the multiple clustering results under the first category into the second preset classification model corresponding to the first category. The second preset classification model determines the first candidate second category corresponding to the clustering result from the multiple known second categories corresponding to the first category according to the characteristics of each clustering result, and uses the first candidate second category as the first candidate second category to which the keyword in the clustering result belongs.

[0104] By pre - defining multiple known second categories for different first categories respectively, and for multiple clustering results under different first categories, based on the multiple known second categories corresponding to the first category, pre - training respective second preset classification models, the accuracy rate of determining the first candidate second category corresponding to each clustering result under different first categories can be improved.

[0105] S206. For each clustering result, input each keyword in the clustering result into the pre - trained third preset classification model, and the third preset classification model outputs the second candidate second category to which the keyword belongs, and the second candidate second category belongs to multiple known second categories.

[0106] Exemplarily, different third preset classification models can be pre - trained for keywords under different known second categories corresponding to each first category, that is, for each known second category corresponding to each first category, based on the multiple known second categories corresponding to the first category, train and learn the initial third preset classification model corresponding to the known second category to obtain the third preset classification model corresponding to the known second category.

[0107] Among them, when training and learning the initial third preset classification model, the sample data input into the initial third preset classification model can be sample data containing the majority of the known second category and a small number of sample data of other known second categories. Thus, the initial third preset classification model can learn more features of the sample data of the known second category during training, and then can adjust the model parameters to achieve more accurate classification of the sample data of the known second category.

[0108] Then, for each clustering result under each category, after obtaining the first candidate second category corresponding to each clustering result, the known second category corresponding to the clustering result can be determined according to the first candidate second category, and then the corresponding third preset classification model can be found according to the known second category. Furthermore, each keyword in the clustering result is input into the third preset classification model, and the third preset classification model determines the second candidate second category of the keyword from the multiple known second categories corresponding to the first category.

[0109] By pre - training respective third preset classification models for keywords under different known second categories corresponding to each first category, the accuracy rate of determining the second candidate second category to which the keywords under different known second categories corresponding to each first category belong can be improved.

[0110] S207. In response to the inconsistency between the first candidate second category and the second candidate second category, replace the clustering result to which the keyword belongs with the clustering result corresponding to the second candidate second category, and re - determine the first candidate second category and the second candidate second category of the keyword.

[0111] Exemplarily, the first candidate second category can be understood as the classification result corresponding to the feature of the clustering result where the keyword is located; the second candidate second category can be understood as the classification result of classifying the keyword based on the corresponding third preset classification model; if there is a keyword whose first candidate second category and second candidate second category are inconsistent, then replace the clustering result to which the keyword belongs with the clustering result corresponding to the second candidate second category, and re - execute according to step S205 - step S207.

[0112] By replacing the clustering result to which the keyword belongs with the clustering result corresponding to the second candidate second category if there is a keyword whose first candidate second category and second candidate second category are inconsistent, it is possible to optimize the clustering result based on the classification result of the keyword, thereby improving the accuracy of the clustering result.

[0113] S208. In response to the first candidate second category and the second candidate second category being consistent, take the first candidate second category as the target second category.

[0114] In one example, each target second category includes multiple second - category labels; the second - category labels of the same second category are not contradictory; the second - category labels between different target second categories do not overlap.

[0115] Exemplarily, if the first candidate second category and the second candidate second category to which each keyword belongs are both consistent, then take the first candidate second category to which each keyword belongs as the target second category.

[0116] By taking the first candidate second category to which each keyword belongs as the target second category if the first candidate second category and the second candidate second category to which each keyword belongs are both consistent, it is possible to verify the accuracy of the classification of the keyword from two dimensions of the clustering result and the keyword, thereby improving the accuracy of the target second category to which the determined keyword belongs.

[0117] Figure 3 It is a schematic structural diagram of the power - grid data processing device provided by the present application, as Figure 3 shown, the power - grid data processing device 30 provided in this embodiment includes:

[0118] An extraction module 301, configured to extract multiple keywords from the power - grid data;

[0119] A first determination module 302, configured to determine the first category to which each of the multiple keywords belongs; wherein, the first categories to which the multiple keywords belong are not completely the same; each first category includes at least one keyword;

[0120] The clustering module 303 is configured to, for each first category, cluster multiple keywords belonging to the first category according to the keywords and the association relationships between the keywords, so as to obtain at least one clustering result; wherein each clustering result includes one or more keywords.

[0121] The second determination module 304 is configured to, for each clustering result, determine the target second category to which each keyword in the clustering result belongs from multiple known second categories.

[0122] In a possible implementation manner, the first determination module 302 is specifically configured to:

[0123] Input the multiple keywords into a pre-trained first preset classification model, and the first preset classification model classifies the multiple keywords according to the multiple keywords and the association relationships between the keywords, so as to obtain the first category to which each keyword belongs.

[0124] In a possible implementation manner, each first category includes multiple first category labels; the first category labels of the same first category are not contradictory; the first category labels of different first categories do not overlap.

[0125] In a possible implementation manner, the second determination module 304 is specifically configured to:

[0126] Input the multiple clustering results into a pre-trained second preset classification model, and the second preset classification model determines a first candidate second category corresponding to each clustering result, and the first candidate second category belongs to the multiple known second categories;

[0127] For each clustering result, input each keyword in the clustering result into a pre-trained third preset classification model, and the third preset classification model outputs a second candidate second category to which the keyword belongs, and the second candidate second category belongs to the multiple known second categories;

[0128] In response to the first candidate second category and the second candidate second category being consistent, use the first candidate second category as the target second category;

[0129] In response to the first candidate second category and the second candidate second category being inconsistent, replace the clustering result to which the keyword belongs with the clustering result corresponding to the second candidate second category, and re-determine the first candidate second category and the second candidate second category of the keyword.

[0130] In a possible implementation manner, each target second category includes multiple second category labels; the second category labels of the same second category are not contradictory; the second category labels of different target second categories do not overlap.

[0131] In a possible implementation, the device 30 further includes:

[0132] Extract a plurality of training data from the training power grid data, each training data corresponding to a keyword extracted from the training power grid data, and the association relationship between the keyword and other keywords;

[0133] Input the plurality of training data into a preset algorithm based on grounded theory, and the preset algorithm determines the initial categories for the plurality of training data;

[0134] For each initial category, receive the label for the initial category and apply the label to the plurality of training data belonging to the initial category;

[0135] Train the initial first preset classification model using the plurality of training data to obtain a pre-trained first preset classification model.

[0136] The processing device for power grid data provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0137] Figure 4 It is a schematic structural diagram of the electronic device provided in this application. As Figure 4 shown, the electronic device 40 provided in this embodiment includes: at least one processor 401 and a memory 402. Optionally, the device 40 further includes a communication component 403. Among them, the processor 401, the memory 402, and the communication component 403 are connected through a bus 404.

[0138] In a specific implementation process, at least one processor 401 executes the computer execution instructions stored in the memory 402, so that at least one processor 401 executes the above method.

[0139] The specific implementation process of the processor 401 can refer to the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0140] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU for short), or may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.

[0141] The memory may include a high-speed random access memory (RAM), and may also include non-volatile memory (NVM), such as at least one disk memory.

[0142] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0143] This application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0144] This application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the above method is implemented.

[0145] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disk. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0146] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0147] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed between each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can be in electrical, mechanical, or other forms.

[0148] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0149] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0150] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0151] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks or optical discs that can store program codes.

[0152] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A method for processing power grid data, characterized in that Including: Extracting a plurality of keywords from the grid data; Determining a first category to which each of the plurality of keywords belongs; wherein, the first categories to which the plurality of keywords belong are not completely the same; each of the first categories includes at least one of the keywords; For each of the first categories, clustering the plurality of keywords belonging to the first category according to the keywords and the association relationships between the keywords to obtain at least one clustering result; wherein, each of the clustering results includes one or more of the keywords; For each of the clustering results, determining a target second category to which each of the keywords in the clustering result belongs from a plurality of known second categories.

2. The method according to claim 1, wherein The determining the first category to which each of the plurality of keywords belongs includes: Inputting the plurality of keywords into a pre-trained first preset classification model, and classifying the plurality of keywords by the first preset classification model according to the plurality of keywords and the association relationships between the keywords to obtain the first categories to which the plurality of keywords belong.

3. The method according to claim 1, characterized in that, Each of the first categories includes a plurality of first category labels; the first category labels of the same first category are not contradictory; the first category labels between different first categories do not overlap.

4. The method according to claim 1, wherein The for each of the clustering results, determining a target second category to which each of the keywords in the clustering result belongs from a plurality of known second categories includes: Inputting the plurality of clustering results into a pre-trained second preset classification model, and determining a first candidate second category corresponding to each clustering result by the second preset classification model, the first candidate second category belonging to the plurality of known second categories; For each clustering result, inputting each keyword in the clustering result into a pre-trained third preset classification model, and outputting a second candidate second category to which the keyword belongs by the third preset classification model, the second candidate second category belonging to the plurality of known second categories; In response to the first candidate second category and the second candidate second category being consistent, taking the first candidate second category as the target second category; In response to the first candidate second category and the second candidate second category being inconsistent, replacing the clustering result to which the keyword belongs with the clustering result corresponding to the second candidate second category, and re-determining the first candidate second category and the second candidate second category of the keyword.

5. The method according to any one of claims 1-4, characterized in that, Each of the target second categories includes a plurality of second category labels; the second category labels of the same second category are not contradictory; the second category labels between different target second categories do not overlap.

6. The method according to claim 2, characterized in that, Further including: Extracting a plurality of training data from the training grid data, each training data corresponding to a keyword extracted from the training grid data and the association relationship between the keyword and other keywords; Inputting the plurality of training data into a preset algorithm based on grounded theory, and determining initial categories for the plurality of training data by the preset algorithm; For each of the initial categories, receiving a label for the initial category and applying the label to the plurality of training data belonging to the initial category; Training the initial first preset classification model using multiple pieces of the training data to obtain the pre-trained first preset classification model.

7. A processing device for power grid data, characterized in that Including: An extraction module, configured to extract multiple keywords from the power grid data; A first determination module, configured to determine the first category to which each of the multiple keywords belongs; wherein, the first categories to which the multiple keywords belong are not completely the same; each of the first categories includes at least one of the keywords; A clustering module, configured to, for each of the first categories, cluster the multiple keywords belonging to the first category according to the keywords and the association relationships between the keywords to obtain at least one clustering result; wherein, each of the clustering results includes one or more of the keywords; A second determination module, configured to, for each of the clustering results, determine the target second category to which each of the keywords in the clustering result belongs from multiple known second categories.

8. An electronic device, characterized in that, Including: A memory and a processor; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory, so that the processor executes the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the method according to any one of claims 1-6.

10. A computer program product, characterized in that, Including a computer program, which when executed by a processor implements the method according to any one of claims 1-6.