Power standard data mining method based on graph clustering

Through the data mining method based on graph clustering, node and attribute construction of power standard documents, node relationship scores are calculated, and information clustering is carried out, which solves the problems of complexity and scenario adaptability of power standard documents, and achieves efficient standard document clustering and usage efficiency improvement.

CN120011420APending Publication Date: 2025-05-16YUNNAN POWER GRID CO LTD ELECTRIC POWER RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510004770.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The complexity and diversity of power standard documents lead to difficulties for practitioners in selecting and understanding standards, increasing learning and use costs, and it is difficult to quickly and accurately obtain standards that match specific scenarios when multiple standards are applied in parallel.

Method used

Using a data mining method based on graph clustering, the nodes and attributes are constructed on the power standard documents, the edges between nodes are defined, and the similarity model is constructed for the shared encoder training, the vector encoding model is obtained, the node relationship score is calculated, and information clustering is finally performed based on the aggregated node characterization vector.

Benefits of technology

It effectively improves the clustering accuracy of power standard documents, simplifies the understanding and application process of standard documents, reduces the learning and use costs of practitioners, and improves the efficiency of standard use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011420A_ABST
    Figure CN120011420A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an electric power standard data mining method based on graph clustering, and the method comprises the steps: carrying out the construction of nodes and attributes based on an electric power standard document, defining an attribute value as a text description, and if a plurality of attribute values exist, defining an attribute value list, each element in the attribute value list is a text description of the attribute value; according to the attributes, defining and constructing the edges of the nodes; determining a neighbor matrix through the relationship of edges, and carrying out vector coding on attributes of all nodes to obtain a feature matrix; carrying out matrix multiplication on the neighbor matrix and the feature matrix, dividing by the number of nodes, and averaging to obtain an aggregated node representation vector; performing node information clustering based on the aggregated node representation vectors to obtain a clustering result; the accuracy of power standard document clustering can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a data mining method for power standards based on graph clustering. Background Art

[0002] Electric power standard documents are not only the cornerstone of industry standardization and standardized operation, but also provide important guidance and support in many key areas. First, it ensures the unified standards of power system design, construction, operation, maintenance, etc., and promotes standardized management of the industry and efficient allocation of resources. Secondly, power standards provide direction for technological development, especially in the application and promotion of new technologies, through standardized definitions and requirements, it promotes technological progress and innovation. In addition, the maintenance and updating of power equipment are also inseparable from the support of standards to ensure the stability and safety of equipment operation. More importantly, in the process of power production, transmission and use, safety assurance is always the top priority, and standard documents provide a scientific and systematic safety management framework to help practitioners reduce operational risks and ensure system safety. Therefore, power standard documents are not only an important tool for technical management, but also the core support for ensuring the long-term sustainable development of the power industry.

[0003] In the power industry, the complexity and diversity of standard documents are common and growing problems. Power standards cover multiple dimensions such as safety, technology, management, and environmental protection, and each standard document often provides detailed regulations for specific application scenarios or operating procedures. Although this multi-level and multi-dimensional standard system has played an important role in improving the overall standardized management and technical support of the industry, it has also led to practitioners facing the dual challenges of selection and understanding in actual work. On the one hand, different working environments, equipment types, and operating conditions require practitioners to accurately select suitable standard documents and apply them flexibly according to actual conditions; on the other hand, the standards themselves are highly technical and professional, and are difficult to understand. In addition, the interrelationships and differences between standard documents make learning and mastering more difficult. With the rapid development of the power industry, new technologies, equipment, and management models continue to emerge, and relevant standards are also updated and increased. This ever-changing standard system further increases the learning and use costs of practitioners. In particular, when facing the parallel application of multiple standards, how to quickly and accurately obtain standards that match specific scenarios has become a problem that needs to be solved urgently. If these standards cannot be selected and applied efficiently, it may not only affect the operating efficiency and safety of the power system, but also increase the operational risks and compliance issues of enterprises. Therefore, how to simplify the understanding and application process of standard documents through innovative technical means and improve the efficiency of standard use by practitioners has become a key issue that the industry needs to solve urgently. Summary of the invention

[0004] The main purpose of the present invention is to provide a data mining method for power standards based on graph clustering to solve the problems of complexity and scenario adaptability of standard documents, and to obtain relatively high-quality clustering results through large model composition and information aggregation.

[0005] To achieve the above objectives, the present application provides, in a first aspect, a data mining method for power standards based on graph clustering, the method comprising:

[0006] Nodes and attributes are constructed based on the power standard document, where the attribute value is defined as a text description. If there are multiple attribute values, an attribute value list is defined, and each element in the attribute value list is a text description of the attribute value;

[0007] According to the attributes, define and construct the edges of the nodes;

[0008] Determine the neighbor matrix through the relationship of the edges, perform vector encoding on the attributes of all nodes, and obtain a feature matrix; perform matrix multiplication on the neighbor matrix and the feature matrix, and divide by the number of nodes for averaging to obtain an aggregated node representation vector;

[0009] Node information is clustered based on the aggregated node representation vectors to obtain a clustering result.

[0010] Optionally, the node and attribute construction based on the electric power standard document includes:

[0011] defining each of the electric power standard documents as a node, and determining the attributes of the node;

[0012] Using the content of the electric power standard document as user prompt words and the attribute extraction prompt words as system prompt words, a prompt word project is constructed;

[0013] The determined prompt words are used as input of the large model, and the output of the large model is a text description of the attribute value of each attribute.

[0014] Optionally, the defining and constructing the edge of the node according to the attribute includes:

[0015] Evaluate the importance of the attributes to obtain the attribute weight of each attribute that contributes to the document relevance;

[0016] defining the matching degree of the document attributes as the edge of the node;

[0017] Construct a dual-tower model of a shared encoder, use the attribute value as a positive sample, and use the batches in the batch as random negative samples to train a similarity model and obtain a vector encoding model;

[0018] The similarity of the attribute values ​​of the nodes is calculated, and the attribute weights are fused through vector dot multiplication. Then, edge screening is performed by setting boundary values ​​to obtain the node relationship score.

[0019] Optionally, the similarity calculation of the attribute values ​​of the nodes is performed, and the attribute weights are integrated by vector dot multiplication, and then edge screening is performed by setting a boundary value to obtain a node relationship score, including:

[0020] The node relationship score is calculated using the following formula:

[0021]

[0022] in, represents the node relationship score, the SimModel is the attribute value similarity calculation model, the Represents node n i and node n j The similarity set of the attributes of the position, represents the transpose of the attribute weights.

[0023] Optionally, the neighbor matrix is ​​determined by edge relationships, and the attributes of all nodes are vector-encoded to obtain a feature matrix, including:

[0024] The nodes connected by edges are taken as the neighbor nodes. The nodes without edges are defined as 0, and the node relationship of the node itself is 1, and the neighbor matrix is ​​obtained.

[0025] All attribute values ​​of each node are vector-encoded using SimModel to obtain the feature vectors of all attributes of each node to obtain the feature matrix F {k,z} , where the first dimension represents the number of nodes k, the second dimension represents the number of attribute feature vectors z, and the hidden third dimension is the attribute feature vector dimension m.

[0026] Optionally, clustering the node information based on the aggregated node representation vector to obtain a clustering result includes:

[0027] Step 1: When the first node arrives, initialize the cluster and obtain the cluster center vector;

[0028] Step 2: When the second node arrives, the Euclidean distance between the new node and the cluster center is calculated. If the distance is less than the threshold, the new node is integrated into the cluster, otherwise a new cluster is established;

[0029] Step 3: When the clustering result is updated, find the point with the smallest average distance from each node to any other node in the cluster and set it as the new cluster center;

[0030] Step 4: Repeat steps 2 and 3 until all nodes are calculated to obtain the clustering result.

[0031] Optionally, finding the point with the smallest average distance from each node to any other node in the cluster and setting it as a new cluster center includes:

[0032] The following formula is used to calculate the node Become a cluster c i The probability of the center point is:

[0033]

[0034] D represents the Euclidean distance calculation, is the cluster c i The center vector of

[0035] The above formula is used to calculate the probability set of all nodes becoming the center point Take the node corresponding to the maximum probability as the cluster c i New center point

[0036] A second aspect of the present application provides a data mining device for power standards based on graph clustering, comprising:

[0037] A construction module is used to construct nodes and attributes based on the power standard document, wherein the attribute value is defined as a text description. If there are multiple attribute values, an attribute value list is defined, and each element in the attribute value list is a text description of the attribute value;

[0038] The construction module is further used to define and construct the edge of the node according to the attribute;

[0039] An information aggregation module is used to determine a neighbor matrix through edge relationships, perform vector encoding on the attributes of all nodes, and obtain a feature matrix; perform matrix multiplication on the neighbor matrix and the feature matrix, and divide the matrix by the number of nodes for averaging to obtain an aggregated node representation vector;

[0040] The clustering module is used to cluster the node information based on the aggregated node representation vector to obtain a clustering result.

[0041] A third aspect of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the first aspect and any possible implementation thereof.

[0042] A fourth aspect of the present application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes each step of the method described in the first aspect.

[0043] The present application provides a data mining method for electric power standards based on graph clustering, which constructs nodes and attributes based on electric power standard documents, wherein an attribute value is defined as a text description, and if there are multiple attribute values, an attribute value list is defined, and each element in the attribute value list is a text description of the attribute value; based on the attributes, the edges of the nodes are defined and constructed; the neighbor matrix is ​​determined through the relationship between the edges, and the attributes of all nodes are vector-encoded to obtain a feature matrix; matrix multiplication is performed on the neighbor matrix and the feature matrix, and the results are averaged by dividing by the number of nodes to obtain an aggregated node representation vector; node information is clustered based on the aggregated node representation vector to obtain a clustering result; the accuracy of clustering of electric power standard documents can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0045] in:

[0046] Figure 1 A schematic diagram of a flow chart of a data mining method for power standards based on graph clustering provided in an embodiment of the present application;

[0047] Figure 2 A schematic diagram of the structure of a data mining device for power standards based on graph clustering provided in an embodiment of the present application;

[0048] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.

[0050] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.

[0051] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0052] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.

[0053] See also Figure 1 , is a flow chart of a data mining method for power standards based on graph clustering provided in an embodiment of the present application, such as Figure 1 As shown, the method includes:

[0054] 101. Nodes and attributes are constructed based on the power standard document, wherein the attribute value is defined as a text description. If there are multiple attribute values, an attribute value list is defined, and each element in the attribute value list is a text description of the attribute value.

[0055] First, the nodes and attributes of the power standard document are constructed. Since the object to be mined is the power standard document, each standard document can be directly used as a node. Then, by observing the content of the power standard document, the document attributes are redefined, and the characteristic values ​​of usage conditions, requirements, etc. are selected as node attributes, and the prompt words and the input and output of the large model are set. This can be done by performing a prompt word project, assembling attribute extraction prompt words, using the content of the power standard document as the user prompt word, and the attribute extraction prompt word as the system prompt word. The two use the message list mode as the input of the large model, and the output is the attribute value text of each attribute.

[0056] In an optional implementation, the above step 101 includes:

[0057] Define each of the above power standard documents as a node, and determine the attributes of the above nodes;

[0058] The content of the above-mentioned power standard document is used as the user prompt word, and the attribute extraction prompt word is used as the system prompt word to build a prompt word project;

[0059] The determined prompt words are used as input of the large model, and the output of the large model is a text description of the attribute value of each attribute.

[0060] Specifically, by observing the content of the power standard documents, it can be found that the descriptions of the documents differ greatly. In order to better describe the possible associations between the documents, the attributes of the documents are redefined in the embodiment of the present application, and characteristic values ​​of usage conditions, requirements, etc. are selected as node attributes.

[0061]

[0062] Table 1

[0063] Table 1 is a node attribute definition table provided in an embodiment of the present application. It is necessary to define the range of attribute values ​​for the attributes. Since the descriptions of each attribute in the power standard document are different and difficult to standardize, in order to better extract the attribute values, the attribute values ​​are defined as a text description in an embodiment of the present application. If there are multiple attribute values, they are defined as an attribute value list, and each element is a text description of the attribute value. The advantage of this approach is that the attribute extraction work can be converted into a large model question-and-answer method for processing. By utilizing the powerful text comprehension ability of the large model, the attribute values ​​of each different attribute can be well summarized.

[0064] Specifically, we can carry out prompt word engineering, assemble the attribute extraction prompt, use the content of the power standard document as the user prompt word user, and the attribute extraction prompt as the system prompt word system. Both use the message list mode as the input of the large model, and the output is the attribute value text of each attribute.

[0065] 102. Based on the above attributes, define and construct the edges of the above nodes.

[0066] After the above processing, the nodes and attributes of the power standard document can be obtained, and the relationship between the nodes needs to be processed later.

[0067] In an optional implementation, the above step 102 includes:

[0068] Evaluate the importance of the above attributes and obtain the attribute weight of each attribute's contribution to the document relevance;

[0069] Define the matching degree of document attributes as the edges of the above nodes;

[0070] Build a dual-tower model with a shared encoder, use the above attribute values ​​as positive samples, and use them as random negative samples in batches to train the similarity model and obtain a vector encoding model.

[0071] The attribute values ​​of the above nodes are similarly calculated, and the attribute weights are integrated through vector dot multiplication. Then, edge screening is performed by setting boundary values ​​to obtain the node relationship score.

[0072] First, an importance evaluation is performed on all attributes to obtain the attribute weight of each attribute's contribution to the document's relevance, and the matching degree of the document's attributes is defined as the edge of the node. Then, a dual-tower model of a shared encoder is constructed, with attribute values ​​as positive samples and batches within a batch as random negative samples to train a similarity model and obtain a vector encoding model. Finally, similarity calculations are performed on node attribute values, and attribute weights are fused through vector dot multiplication. Edge screening is then performed by setting boundary values ​​to obtain the final node relationship score.

[0073] Specifically: 1) First, the embodiment of the present application evaluates the importance of all attributes to obtain the importance of each attribute in contributing to the document relevance, which can be marked as:

[0074] Weight attr represents a weight list, marking the weight value of each attribute, and w attr It represents the weight value of a specific attribute. Based on this, the application can measure the relevance of each group of attributes separately, and then add the weight coefficient to combine them, so as to effectively measure the relationship between nodes.

[0075] 2) Define the matching degree of document attributes as the edge of the node.

[0076] 3) Since the attribute values ​​were extracted into a text description in previous work, measuring feature relevance can be regarded as a similarity calculation task. The more similar the attribute descriptions are, the higher the similarity. This approach is based on the fact that the description forms of different attributes vary greatly, while the descriptions of the same attributes vary less. The main focus is on the changes in keywords, so the similarity model can well express the differences in attribute values. The specific method is to build a dual-tower model with a shared encoder, use attribute values ​​as positive samples, and use batches within a batch as random negative samples to train the similarity model.

[0077] Through the above simple training, the present invention can obtain a model SimModel that effectively measures the similarity of attribute values. After calculating the similarity of attribute values, combined with Weight attr By combining them, we can get a specific score that measures the relationship between nodes.

[0078] Further optionally, similarity calculation is performed on the attribute values ​​of the above nodes, and the attribute weights are integrated by vector dot multiplication, and then edge screening is performed by setting a boundary value to obtain a node relationship score, including:

[0079] The above node relationship score is calculated using the following formula:

[0080]

[0081] in, represents the above node relationship score, the above SimModel is the attribute value similarity calculation model, the above Represents node n i and node n j The similarity set of the attributes of the position, represents the transpose of the above attribute weights.

[0082] Specifically, in the above formula, CoS represents cosine calculation. By pre-calculating the attribute vector obtained by SimModel, the similarity of each attribute can be obtained. Represents node n i and node n j A similarity set of attributes. Represents the transpose of the attribute weights, and finally merges them through vector dot multiplication to obtain the final node relationship score

[0083] The embodiment of the present application can effectively measure the relevance between each document through the above method and regard it as the edge between nodes.

[0084] Finally, the edge is trimmed. Here, in the embodiment of the present application, a hyperparameter margin score is defined to trim the edge. The edge that exceeds the margin score is retained, and the edge that is lower than the margin score is set to 0.

[0085] 103. Determine the neighbor matrix through the relationship of the edges, perform vector encoding on the attributes of all nodes, and obtain a feature matrix; perform matrix multiplication on the neighbor matrix and the feature matrix, and divide by the number of nodes for averaging to obtain an aggregated node representation vector.

[0086] This step mainly aggregates the information of neighbor nodes. First, we can obtain the neighbor matrix through edge relationships, then vectorize the attributes of all nodes to obtain the feature matrix, and finally perform matrix multiplication on the neighbor matrix and the feature matrix, and divide it by the number of nodes to average, so as to obtain the attribute vector of each node that integrates the neighbor node relationship.

[0087] In an optional implementation, the neighbor matrix is ​​determined by edge relationships, and the attributes of all nodes are vector-encoded to obtain a feature matrix, including:

[0088] Take the nodes connected by edges as the above neighbor nodes, define the node relationship of nodes without edges as 0, and the node relationship of the node itself as 1, and get the neighbor matrix

[0089] All attribute values ​​of each node are vectorized using SimModel to obtain the feature vectors of all attributes of each node to obtain the feature matrix F {k,z} , where the first dimension represents the number of nodes k, the second dimension represents the number of attribute feature vectors z, and the hidden third dimension is the attribute feature vector dimension m.

[0090] Specifically, by averaging the attribute vectors of all nodes in the node cluster and weighting them according to the size of the edges, the information of neighboring nodes can be effectively aggregated.

[0091]

[0092] in, Represents the eigenvector of attribute j in node i. It can be seen that the elements in the A matrix are scalars, the i-th row represents the node relationship set between the i-th node and all other nodes, and the elements of the F matrix are vectors, and the j-th column represents the j-th attribute of all nodes. It can be seen that the multiplication of the rows and columns of A and F is exactly the multiplication and addition of the node relationship scalar and the corresponding attribute vector. Therefore, in the embodiment of the present application, it is only necessary to define the matrix multiplication operation of A and F, and then divide it by the node dimension k for averaging to obtain the attribute vector of the fused neighbor node relationship of each node. Finally, in the embodiment of the present application, all the attribute vectors of each node are combined with the attribute weight Weight attr Perform weighted summation to obtain the aggregated node representation vector

[0093] At this step, the embodiment of the present application has completed the related work of neighbor node aggregation and obtained the node representation vector, on which the node information clustering can be performed subsequently.

[0094] 104. Cluster the node information based on the above aggregated node representation vectors to obtain a clustering result.

[0095] Node information clustering in the embodiment of the present application refers to the process of dividing nodes into different clusters or classes according to the characteristics and structural relationships of the nodes in graph structure data.

[0096] Specifically, we can first initialize the cluster and obtain the cluster center vector, then calculate the Euclidean distance between the new node and the cluster center. If the distance is less than the threshold, the node is integrated into the cluster, otherwise a new cluster is established. Finally, the distance center is updated and this step is repeated to obtain the final clustering result.

[0097] In an optional implementation, clustering the node information based on the aggregated node representation vector to obtain a clustering result includes:

[0098] Step 1: When the first node arrives, initialize the cluster and obtain the cluster center vector;

[0099] Step 2: When the second node arrives, the Euclidean distance between the new node and the cluster center is calculated. If the distance is less than the threshold, the new node is integrated into the above cluster, otherwise a new cluster is established;

[0100] Step 3: When the clustering results are updated, find the point with the smallest average distance from each node to any other node in the above cluster and set it as the new cluster center;

[0101] Step 4: Repeat steps 2 and 3 until all nodes are calculated and the above clustering results are obtained.

[0102] In order to explain the above clustering method more clearly, the above steps are described in detail as follows:

[0103] 1) Select a node n1 as the initial cluster Assume that this application has obtained a node representation vector with rich information And As the center vector of cluster c1, denoted as

[0104] 2) The Euclidean distance calculation is recorded as D, and the distance threshold is recorded as d megin When the second node n2 arrives, this application only needs to update the node vector and Calculate Euclidean distance Less than d megin , then update the cluster c1 = {n1, n2}, otherwise n2 is used as the new cluster As the C2 cluster center vector

[0105] 3) Node The basis for becoming a cluster center is that this node The average distance of other nodes in is the smallest, assuming There are x nodes in total, so the following formula can be used to calculate the number of nodes n that form cluster c i The probability of the center point.

[0106]

[0107] Among them, g n Represents the probability of node n becoming the center of the cluster. This formula can be used to calculate the probability set of all nodes becoming the center. Take the node corresponding to the maximum value as the cluster c i New center point

[0108] 4) Repeat steps 2) and 3) until all nodes are calculated, and then we can get the cluster set C = {c1, c2, ..., c i ,…}, and the center points of all clusters, denoted as the set When the new document node n k+1 When entering, the Euclidean distance is calculated for the center points of all clusters If there is a cluster whose size is less than the threshold, the node will be added to the cluster and the cluster center will be updated, otherwise it will become a new cluster. Since then, all the work has been completed and the graph clustering results for the power standard document have been obtained.

[0109] In the embodiment of the present application, the electric power standard document is taken as a node, and the various elements in the electric power standard document (such as clauses, requirements, application scenarios, etc.) are taken as node attributes, and the association between nodes is represented by edges; then a large model is introduced to assist in graph construction, which overcomes the shortcomings of traditional graph construction methods in node definition and attribute association, so that the constructed graph can better reflect the key information of the electric power standard document; at the same time, by aggregating neighbor node information, the document node obtains implicit features with attribute preferences; finally, the incremental clustering algorithm is used to realize the clustering of the electric power standard document, which effectively improves the accuracy of the clustering of the electric power standard document.

[0110] In order to better demonstrate the effect of the method in the embodiments of the present application, the following is an example of experimental verification based on the above method.

[0111] 1) This application uses 121 Chinese documents and 29 English documents related to power standards, a total of 150 documents, involving a total of 1392 power standard terms. Table 2 is an analysis table of a power standard document data set provided in an embodiment of this application.

[0112]

[0113] Table 2

[0114] 2) This application manually observes the document structure and content, and finally defines a total of 76 attributes, and ensures that these attributes can have attribute values ​​in multiple documents. Among them, there are about 20 attributes with high commonality and about 56 attributes with low commonality. Based on commonality, the importance is evaluated and the corresponding weight values ​​are given.

[0115] 3) This application uses Qwen-200B as the large model base and sets the temperature to 0.8; in addition, the tiny-Bert open source pre-trained model is used as the encoder when training SimModel, and the maximum length of the attribute value is set to 100 characters; the learning rate is set to 0.01; the batch size is set to 64; the training rounds are set to 20; and finally, the marginscore is selected as 0.6 in information aggregation. The evaluation indicators of this application mainly use F1-score (F1 value) and Overlapping Normalized Mutual Information (ONMI) to evaluate the clustering effect.

[0116] 4) This application mainly uses composition to enhance the clustering effect for power standard documents, so it is necessary to pay more attention to the impact of composition effect on clustering results when conducting comparative experiments. Four classic clustering models are mainly used as baseline models for comparison in the embodiments of this application. The baseline models are as follows:

[0117] Incremental K-means: a classic clustering method based on iterative solution, using raw data as input;

[0118] AE+K-means: clustering using the data potential representation obtained by the autoencoder and the K-means algorithm;

[0119] DEC: Based on AE+K-means, a self-training clustering loss is designed to optimize the model;

[0120] DCN: Based on AE+K-means, the objective function of the K-means algorithm is used to optimize the model.

[0121] In order to make the comparative experiment more fair, the baseline model was simply modified in the embodiment of the present application, and the range division principle was changed from selecting the nearest range to being smaller than the preset range, which is consistent with the method in the embodiment of the present application.

[0122] 5) When conducting experiments, the other four algorithms do not perform composition, but directly characterize the document content and then perform clustering. Table 3 is a table of comparative experimental results of a baseline model provided in the embodiment of the present application. The experimental results are as follows:

[0123] As shown in Table 3:

[0124]

[0125] Table 3

[0126] According to the results in Table 3, it can be seen that the K-means model using the original data has the worst effect, and the improved clustering models have all been relatively improved, but the performance is still poor. This is because the content of the power standard document itself is quite different, and its complexity and diversity are the root causes of the poor performance of these clustering algorithms. The method of this application does not use the content of the document itself in the early modeling process, but adds a large number of attribute features, and obtains a comprehensive representation of the document nodes through the graph structure, and finally achieves a relatively significant improvement. Compared with the best performing DEC model in the baseline model, F1 has increased by 13.8%, and ONMI has increased by 20.8%. This also strongly illustrates the effectiveness of the method in the embodiment of this application in using the composition strategy on the power standard document.

[0127] 6) In order to verify the effectiveness of the key modules in the method of this application. The experimental results are shown in Table 4. Special note, "(-) Attribute definition" means that attribute construction is not used, and the large model is directly used to summarize the document content to express the document content, and then the node relationship is constructed by training the similarity model. Since there is no attribute representation vector, the neighbor matrix and the node matrix can be directly multiplied to fuse the neighbor node information; "(-) Attribute evaluation" means that when defining attributes, no importance evaluation is performed on the attributes, and the weight values ​​are all the same by default, and the others remain unchanged; "(-) SimModel" means that no similarity model training is performed, and the pre-trained model tiny-BERT is used directly; "(-) Information aggregation" means that no composition is performed, and the extracted attributes are directly used to characterize the nodes. Table 4 is an ablation experiment result table provided in an embodiment of the present application. The experimental results are as follows

[0128] As shown in Table 4.

[0129]

[0130] As shown in the results of Table 4, in the case of "(-) attribute definition", the evaluation result decreased by about 7.2% of the F1 value compared with the model of this application, and ONMI decreased by 12.1%, which shows that attribute definition is crucial to structured power standard text. The structured power standard text has largely alleviated the problems of complexity and diversity, and maintained uniformity in format. It also shows that the method of this application makes full use of attribute information; in the case of "(-) attribute evaluation", the evaluation result also decreased compared with the text model, which shows that for the power standard text, distinguishing and utilizing the importance of attributes can more accurately obtain the representation information of the power standard document, and it also shows that the method of the present invention effectively uses the key information of attribute weight; in the case of "(-) SimModel", the evaluation result also decreased, which shows that comparative learning and training for text characteristics can better measure the relevance of attributes; finally, in the case of "(-) information aggregation", it can be seen that the evaluation result of this application has also suffered a large loss, the F1 value decreased by about 5%, and ONMI decreased by 4.7%, which also effectively illustrates that the information aggregation link of the method of this application makes full use of neighbor node information, and also illustrates the importance of composition for understanding power standard text.

[0131] 7) In order to verify the influence of the large model capability on the performance of the whole method, this application used large models of different types and sizes on the market for experiments, including GPT-4o, Qwen series, GLM-3 and GLM-4. In addition, this application also designed three different prompt styles to compare the effects. One is the simpleprompt that reduces the content of the attribute definition, the other is the middle prompt used in this application, which includes the normal attribute definition, and the last is the complex prompt that adds many detailed requirements and conditions to the attribute definition, such as output style, word count, etc. Table 5 is a large model impact analysis experimental result table provided in the embodiment of this application, and the specific experimental results are shown in Table 5.

[0132]

[0133] Table 5

[0134] As shown in the results of Table 5, first of all, in the case of middle prompt, this application compares the performance of the Qwen series models and finds that after 72B, the change in the scale of the large model has basically negligible impact on the final performance. By comparing the results of GPT-4o, GLM-4 and Qwen-200b, it can be seen that the performance difference between the three is within 0.2, which means that the impact of different types of models on the composition quality is very small. From this, it is not difficult for this application to draw the conclusion that the performance of the large model does not require a particularly large parameter scale for the composition method and data set of this application. It only needs to maintain performance above 32B to achieve the desired effect of this application. Of course, this is also because the power standard document for this application is a standard document with a very clear structure, and this application is a property definition that caters to the document structure, so this phenomenon is also expected.

[0135] In addition, the present invention compares three sets of data of simple prompt, meddle prompt and complex prompt. It can be seen that meddle prompt achieves the best effect in both sets of evaluation indicators. This simple phenomenon can explain that the more complex the prompt is, the better it is. Reasonable prompt design and attribute definition play a key role in improving the performance of the entire method.

[0136] Based on the description of the above method embodiment, in one embodiment of the present application, a data mining device for power standards based on graph clustering is also proposed, see Figure 2 , Figure 2 A schematic diagram of a data mining device for power standards based on graph clustering provided in an embodiment of the present application. Figure 2 As shown, the data mining device for power standards based on graph clustering includes:

[0137] A construction module 210 is used to construct nodes and attributes based on the power standard document, wherein the attribute value is defined as a text description. If there are multiple attribute values, an attribute value list is defined, and each element in the attribute value list is a text description of the attribute value;

[0138] The construction module 210 is further used to define and construct the edge of the node according to the attribute;

[0139] The information aggregation module 220 is used to determine the neighbor matrix through the relationship of the edges, perform vector encoding on the attributes of all nodes, and obtain a feature matrix; perform matrix multiplication on the neighbor matrix and the feature matrix, and divide the matrix by the number of nodes to obtain an average, so as to obtain an aggregated node representation vector;

[0140] The clustering module 230 is used to cluster the node information based on the aggregated node representation vectors to obtain a clustering result.

[0141] Based on the description of the above method embodiment, in one embodiment of the present application, an electronic device is also provided. Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 3 As shown, the electronic device 300 includes a processor 301 and a memory 302, wherein the memory 302 stores a computer program. When the computer program is executed by the processor 301, the following operations are performed: Figure 1 Any step in the method embodiment shown. The electronic device 300 may also include an input / output device, etc. In a specific implementation, the electronic device may be a terminal device, etc.

[0142] In one embodiment, a computer-readable storage medium is further provided. The computer-readable storage medium stores a computer program. When the computer program is executed by the processor 301, the processor 301 executes any step in the above method embodiment.

[0143] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0144] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0145] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A data mining method for power standards based on graph clustering, characterized in that: The method comprises: Nodes and attributes are constructed based on the power standard document, where the attribute value is defined as a text description. If there are multiple attribute values, an attribute value list is defined, and each element in the attribute value list is a text description of the attribute value; According to the attributes, define and construct the edges of the nodes; Determine the neighbor matrix through the relationship of the edges, perform vector encoding on the attributes of all nodes, and obtain a feature matrix; perform matrix multiplication on the neighbor matrix and the feature matrix, and divide by the number of nodes for averaging to obtain an aggregated node representation vector; Node information is clustered based on the aggregated node representation vectors to obtain a clustering result.

2. The data mining method for power standards based on graph clustering according to claim 1 is characterized in that: The node and attribute construction based on the electric power standard document includes: defining each of the electric power standard documents as a node, and determining the attributes of the node; Using the content of the electric power standard document as user prompt words and the attribute extraction prompt words as system prompt words, a prompt word project is constructed; The determined prompt words are used as input of the large model, and the output of the large model is a text description of the attribute value of each attribute.

3. The data mining method for power standards based on graph clustering according to claim 2 is characterized in that: Defining and constructing the edge of the node according to the attribute includes: Evaluate the importance of the attributes to obtain the attribute weight of each attribute that contributes to the document relevance; defining the matching degree of the document attributes as the edge of the node; Construct a dual-tower model of a shared encoder, use the attribute value as a positive sample, and use the batches in the batch as random negative samples to train a similarity model to obtain a vector encoding model; The similarity of the attribute values ​​of the nodes is calculated, and the attribute weights are fused through vector dot multiplication. Then, edge screening is performed by setting boundary values ​​to obtain the node relationship score.

4. The data mining method for power standards based on graph clustering according to claim 3 is characterized in that: The similarity calculation of the attribute values ​​of the nodes is performed, and the attribute weights are integrated by vector dot multiplication, and then edge screening is performed by setting a boundary value to obtain a node relationship score, including: The node relationship score is calculated using the following formula: in, represents the node relationship score, the SimModel is the attribute value similarity calculation model, the Represents node n i and node n j The similarity set of the attributes of the position, represents the transpose of the attribute weights.

5. The data mining method for power standards based on graph clustering according to claim 4 is characterized in that: The neighbor matrix is ​​determined by edge relationships, and the attributes of all nodes are vector-encoded to obtain a feature matrix, including: The nodes connected by edges are taken as the neighbor nodes. The nodes without edges are defined as 0, and the node relationship of the node itself is 1, and the neighbor matrix is ​​obtained. All attribute values ​​of each node are vector-encoded using SimModel to obtain the feature vectors of all attributes of each node to obtain the feature matrix F {k,z} , where the first dimension represents the number of nodes k, the second dimension represents the number of attribute feature vectors z, and the hidden third dimension is the attribute feature vector dimension m.

6. The data mining method for power standards based on graph clustering according to claim 5 is characterized in that: The clustering of node information based on the aggregated node representation vector to obtain a clustering result includes: Step 1: When the first node arrives, initialize the cluster and obtain the cluster center vector; Step 2: When the second node arrives, the Euclidean distance between the new node and the cluster center is calculated. If the distance is less than the threshold, the new node is integrated into the cluster, otherwise a new cluster is established; Step 3: When the clustering result is updated, find the point with the smallest average distance from each node to any other node in the cluster and set it as the new cluster center; Step 4: Repeat steps 2 and 3 until all nodes are calculated to obtain the clustering result.

7. The data mining method for power standards based on graph clustering according to claim 6 is characterized in that: The step of finding the point with the smallest average distance from each node to any other node in the cluster and setting it as a new cluster center includes: The following formula is used to calculate the node Become a cluster c i The probability of the center point is: D represents the Euclidean distance calculation, is the cluster c i The center vector of The above formula is used to calculate the probability set of all nodes becoming the center point Take the node corresponding to the maximum probability as the cluster c i New center point 8. A data mining device for power standards based on graph clustering, characterized in that: include: A construction module is used to construct nodes and attributes based on the power standard document, wherein the attribute value is defined as a text description. If there are multiple attribute values, an attribute value list is defined, and each element in the attribute value list is a text description of the attribute value; The construction module is further used to define and construct the edge of the node according to the attribute; The information aggregation module is used to determine the neighbor matrix through the relationship of the edges, vectorize the attributes of all nodes, and obtain the feature matrix; Performing matrix multiplication on the neighbor matrix and the feature matrix, and dividing by the number of nodes for averaging, to obtain an aggregated node representation vector; The clustering module is used to cluster the node information based on the aggregated node representation vector to obtain a clustering result.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 7.