A system and method for classifying and sorting power data
Through semantic fusion, entity recognition and association evaluation, the power archive knowledge graph is built, and the intelligent classification model is trained, which solves the problem of inefficient classification and sorting of multi-source power data, and realizes efficient and accurate power archive data classification.
Patent Information
- Application Number
- CN202411126343.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-08-15
AI Technical Summary
The prior art has inaccurate identification of entity of multi-source power data and poor accuracy of association relationships, resulting in low efficiency in classification and sorting of power files.
The semantic fusion module is used to perform semantic matching and fusion of multi-source power archive data, and the historical target entity is extracted through the entity recognition module, the association relationship acquisition module is used to evaluate and correct the association relationship, build a power archive knowledge graph, and train the power archive data intelligent classification model for classification and sorting.
It improves the accuracy and efficiency of power file data classification and improves the automation level of power file classification and sorting.
Smart Images

Figure CN119226842B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of power data management, and particularly to a system and method for classifying and sorting power data. Background Art
[0002] With the development of smart grid and Internet of Things technologies, a vast amount of data is generated during the operation of the power system, and this data needs to be effectively stored, processed, and analyzed. Currently, existing power archive management systems mainly rely on traditional manual classification and retrieval methods, and are unable to handle the massive and multi-source power archive data effectively, resulting in low efficiency. At the same time, due to the lack of effective data fusion and standardization means, there are significant differences and inconsistencies between power archive data from different sources.
[0003] In summary, due to the inaccurate entity recognition of multi-source data and the poor accuracy of the association relationship in the prior art, the efficiency of power archive classification and sorting is low. Summary of the Invention
[0004] The purpose of this application is to provide a system and method for classifying and sorting power data, so as to solve the problem that the efficiency of power archive classification and sorting is low due to the inaccurate entity recognition of multi-source data and the poor accuracy of the association relationship in the prior art.
[0005] In view of the above problems, this application provides a system and method for classifying and sorting power data.
[0006] In a first aspect, the present application provides a system for classifying and organizing power data. The system for classifying and organizing power data includes: a semantic fusion module for collecting multi-source power archive data, performing semantic matching on the multi-source power archive data, and performing semantic fusion according to the semantic matching result to generate a target power archive data set to be classified; an entity recognition module for extracting the classified power archive data set from the power data archive data system, performing entity recognition on the classified power archive data set based on a predetermined classification entity through an archive text recognition model, and extracting a plurality of historical target entities, where the predetermined classification entity includes grid connection dispatching protocols, power business licenses, and project approvals, and the plurality of historical target entities include historical grid connection dispatching protocols, historical power business licenses, and historical project approvals; an association relationship acquisition module for identifying a first association relationship set among the historical grid connection dispatching protocols, historical power business licenses, and historical project approvals, evaluating the accuracy of the first association relationship set, and correcting the first association relationship set according to the evaluation result to generate a historical target association relationship set; a knowledge graph construction module for constructing a power archive knowledge graph according to the historical grid connection dispatching protocols, historical power business licenses, historical project approvals, and the historical target association relationship set; a classification model acquisition module for extracting classification features of the power archive knowledge graph and training an intelligent classification model for power archive data; and a classification and organization module for classifying and organizing the target power archive data set to be classified according to the intelligent classification model for power archive data.
[0007] Second aspect, the present application provides a method for classifying and organizing power data, and the method for classifying and organizing power data is implemented through a system for classifying and organizing power data. Among them, the method for classifying and organizing power data includes: collecting multi-source power archive data, performing semantic matching on the multi-source power archive data, performing semantic fusion according to the semantic matching results, and generating a target power archive data set to be classified; extracting the classified power archive data set from the power data archive data system, performing entity recognition on the classified power archive data set through an archive text recognition model based on a predetermined classification entity, and extracting multiple historical target entities. Among them, the predetermined classification entity includes grid connection dispatching agreement, power business license, and project approval, and the multiple historical target entities include historical grid connection dispatching agreement, historical power business license, and historical project approval; identifying the first association relationship set among the historical grid connection dispatching agreement, historical power business license, and historical project approval, performing accuracy evaluation on the first association relationship set, and correcting the first association relationship set according to the evaluation results to generate a historical target association relationship set; constructing a power archive knowledge graph according to the historical grid connection dispatching agreement, historical power business license, historical project approval, and the historical target association relationship set; extracting the classification features of the power archive knowledge graph, and training an intelligent classification model for power archive data; classifying and organizing the target power archive data set to be classified according to the intelligent classification model for power archive data.
[0008] One or more technical solutions provided in the present application have at least the following technical effects or advantages:
[0009] A semantic fusion module is used to collect multi-source power archive data, perform semantic matching on the multi-source power archive data, perform semantic fusion according to the semantic matching results, and generate a target power archive data set to be classified; an entity recognition module is used to extract the classified power archive data set from the power data archive data system, and perform entity recognition on the classified power archive data set through an archive text recognition model based on predetermined classification entities, and extract multiple historical target entities, where the predetermined classification entities include grid connection dispatching agreements, power business licenses, and project approvals, and the multiple historical target entities include historical grid connection dispatching agreements, historical power business licenses, and historical project approvals; a correlation relationship acquisition module is used to identify a first correlation relationship set among the historical grid connection dispatching agreements, historical power business licenses, and historical project approvals, evaluate the accuracy of the first correlation relationship set, and correct the first correlation relationship set according to the evaluation results to generate a historical target correlation relationship set; a knowledge graph construction module is used to construct a power archive knowledge graph according to the historical grid connection dispatching agreements, historical power business licenses, historical project approvals, and the historical target correlation relationship set; a classification model acquisition module is used to extract classification features of the power archive knowledge graph and train an intelligent classification model for power archive data; a classification and sorting module is used to classify and sort the target power archive data set to be classified according to the intelligent classification model for power archive data, effectively solving the problem that the efficiency of power archive classification and sorting is low in the prior art due to inaccurate entity recognition of multi-source data and poor accuracy of correlation relationships, improving the accuracy and efficiency of power archive data classification, and enhancing the automation level of power archive classification and sorting.
[0010] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are given below. It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0012] Figure 1 It is a schematic structural diagram of a system for classifying and sorting power data according to the present application;
[0013] Figure 2 This is a flowchart of a method for classifying and organizing power data in this application.
[0014] Explanation of the reference numerals in the drawings:
[0015] Semantic fusion module 11, entity recognition module 12, association relationship acquisition module 13, knowledge graph construction module 14, classification model acquisition module 15, classification and organization module 16. Specific implementation manners
[0016] By providing a system and method for classifying and organizing power data, this application solves the problem that in the prior art, due to inaccurate entity recognition of multi-source data and poor accuracy of association relationships, the efficiency of classifying and organizing power archives is low, improves the accuracy and efficiency of classifying power archive data, and enhances the automation level of classifying and organizing power archives.
[0017] Next, the technical solutions in this application will be described clearly and completely with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments of this application. It should be understood that this application is not limited by the exemplary embodiments described herein. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the scope of protection of this application. Additionally, it should be noted that for the sake of description, only the parts related to this application are shown in the drawings rather than all of them.
[0018] Embodiment 1, this application provides a system for classifying and organizing power data. Please refer to the attached Figure 1 , the system for classifying and organizing power data includes:
[0019] The semantic fusion module 11 is used to collect multi-source power archive data, perform semantic matching on the multi-source power archive data, and perform semantic fusion according to the semantic matching result to generate a target power archive data set to be classified.
[0020] Specifically, determine the data sources to be collected, including but not limited to the file management system within the power company, public data on the data sharing platform, etc. For structured data, such as data tables in a database, database query statements, such as SQL, can be used for direct extraction. For unstructured or semi-structured data, such as text files, PDF documents, pictures, etc., web crawler technology, OCR recognition technology or document parsing tools are used for extraction. Clean the collected data to remove duplicate, incorrect or irrelevant data, and standardize the data, including unifying data formats, encodings, units, etc. Use natural language processing technology to process text data, such as word segmentation, part-of-speech tagging, named entity recognition, etc. Apply semantic similarity algorithms, such as cosine similarity, Jaccard similarity, similarity calculation based on word vectors, etc., to evaluate the semantic similarity between data items in different data sources. Cross-validate the matching results to ensure the accuracy of the matching, and adjust the semantic matching strategy or parameters according to the evaluation results to optimize the matching effect. According to the matching results and data characteristics, formulate appropriate fusion strategies. For example, for highly similar data items, merging can be selected; for data items with differences, decisions are made according to specific business rules. Use data fusion tools to implement the actual data fusion operation, including steps such as merging similar data items, handling data conflicts, and filling in missing values. Conduct quality checks on the fused data, including verification in terms of integrity, consistency, accuracy, etc. Based on the fused data, construct a power file dataset to be classified. This dataset contains all power file data items to be classified, and each data item should be accurate, complete and consistent.
[0021] The entity recognition module 12 is used to extract the classified power file dataset from the power data file data system, and perform entity recognition on the classified power file dataset through the file text recognition model based on predetermined classification entities, and extract multiple historical target entities, where the predetermined classification entities include grid connection dispatching agreements, power business licenses, project approvals, and the multiple historical target entities include historical grid connection dispatching agreements, historical power business licenses, historical project approvals.
[0022] Specifically, extract the classified power file dataset from the power data file storage database. The classified power file dataset has been classified and stored according to certain rules or standards. Next, use the file text recognition model to perform entity recognition on the extracted dataset. The model has the ability to recognize predetermined classified entities. In this process, the predetermined classified entities refer to the specific information that needs to be extracted from the power file data, such as grid connection dispatching agreements, power business licenses, project approvals, etc. Through the processing of the file text recognition model, multiple historical target entities can be extracted from the power file data. Use SQL query statements or database access interfaces to extract the classified power file dataset from the power data file storage database. Perform text preprocessing on the extracted dataset, including cleaning data, removing irrelevant information, unifying text formats, etc. Use the labeled dataset for training, and configure the model to recognize predetermined classified entities such as historical grid connection dispatching agreements, historical power business licenses, historical project approvals, etc. Input the preprocessed text data into the file text recognition model for entity recognition. The model will output the recognized entities and their corresponding text segments. Extract multiple historical target entities from the output of the model, including historical grid connection dispatching agreements, historical power business licenses, historical project approvals, etc.
[0023] The association relationship acquisition module 13 is used to identify the first association relationship set among the historical grid connection dispatching agreements, historical power business licenses, and historical project approvals, evaluate the accuracy of the first association relationship set, and correct the first association relationship set according to the evaluation results to generate a historical target association relationship set.
[0024] Specifically, use text mining or natural language processing techniques to extract the association relationships among the historical grid connection dispatching agreements, historical power business licenses, and historical project approvals from the power file data, which is achieved through methods such as keyword matching, dependency syntax analysis, and semantic role annotation. Organize the extracted association relationships into a set, that is, the first association relationship set, which contains the association relationships among all recognized entities. Evaluate the first association relationship set based on a machine learning model. The evaluation is to determine the reliability and accuracy of the association relationships. According to the evaluation results, identify the errors or inconsistencies in the first association relationship set, and correct the incorrect or inconsistent association relationships to ensure their accuracy and consistency, including deleting incorrect association relationships, adding missing association relationships, modifying inaccurate association relationships, etc. Integrate the corrected association relationships into a new set, that is, generate a historical target association relationship set.
[0025] The knowledge graph construction module 14 is used to construct a power file knowledge graph based on the historical grid connection dispatching agreements, historical power business licenses, historical project approvals, and the historical target association relationship set.
[0026] Specifically, ensure that historical grid-connected dispatching agreements, historical power business licenses, historical project approvals, etc. are correctly identified and extracted, and entities are deduplicated and standardized to ensure the uniqueness of each entity in the knowledge graph. Sort out the historical target association set, ensure that all associations are accurate and consistent, and classify and standardize the associations. In the knowledge graph, each historical target entity will be represented as a node, and the associations between entities will be represented as edges between nodes. Design attributes for each node, such as the name, type, description, etc. of the entity, and design attributes for each edge, such as the type, strength, direction, etc. of the association. Create nodes in the knowledge graph database, assign a unique identifier to each node, and add the attribute information of the entity to the corresponding node. Create edges in the knowledge graph based on the historical target association set, connect related nodes, add attribute information to each edge, describe the relationship between nodes, and perform data verification on the constructed knowledge graph to ensure the correctness of nodes and edges.
[0027] The classification model acquisition module 15 is used to extract the classification features of the power archive knowledge graph and train the power archive data intelligent classification model.
[0028] Specifically, extract features from the nodes of the power archive knowledge graph, such as the type, name, description, etc. of the entity. Extract the features of the edges in the knowledge graph, such as the type, strength, direction, etc. of the association relationship, to represent the relationship and connectivity between nodes. Extract the graph structure features of the knowledge graph, such as the neighbors, paths, subgraphs, etc. of the nodes, which can provide information about the position and importance of the nodes in the graph. Encode the extracted features, and the features can be encoded using methods such as one-hot encoding and embedding representation. Prepare a labeled power archive dataset, which contains classified power archive samples, and each sample has a corresponding label to indicate the category to which it belongs. Based on the extracted features, construct a feature matrix, in which each row represents a sample and each column represents a feature. Divide the labeled dataset into a training set, a validation set, and a test set. Select a machine learning model suitable for the classification of power archive data, such as a support vector machine, a random forest, a deep learning model, etc. Use the training set to train the selected model, adjust the model parameters to optimize the classification performance, use the validation set to validate the trained model, evaluate its classification accuracy and generalization ability, and tune the model based on the validation results, such as adjusting hyperparameters and optimizing feature selection. The classification features of the power archive knowledge graph can be extracted, and an intelligent classification model for power archive data can be trained. This model will be able to automatically classify new power archive data and improve the efficiency and accuracy of power archive management.
[0029] A classification and sorting module 16, configured to classify and sort the target power archive dataset to be classified according to the intelligent classification model of the power archive data.
[0030] Specifically, clean the target power archive dataset to be classified, remove irrelevant information, fill in missing values, correct incorrect data, etc., load the trained intelligent classification model of the power archive data, input the processed power archive dataset to be classified into the model, and the model performs classification prediction on the input power archive data and assigns a class label to each sample. Output the classification result of the model to obtain a classification list containing the sample ID and the corresponding class label, and sort the power archive dataset to be classified into subsets of different classes according to the classification result.
[0031] Furthermore, the entity recognition module 12 is also used for:
[0032] Randomly annotate the classified power archive dataset according to the grid connection scheduling agreement, power business license, and project approval to obtain an unannotated power archive dataset and an annotated power archive dataset; based on self-training and generative adversarial networks, train the archive text recognition model according to the unannotated power archive dataset and the annotated power archive dataset.
[0033] Specifically, randomly select a part of the samples from the classified power archive dataset as the annotation objects. Conduct detailed annotation on the selected samples, including text content, key information, category to which they belong, etc., to generate an annotated power archive dataset, and use the remaining unannotated samples as the unannotated power archive dataset. Preprocess the annotated and unannotated power archive datasets, such as text cleaning, word segmentation, removing stop words, etc. Extract text features, such as bag-of-words model, TF-IDF, word embedding, etc., for model training, and combine self-training and generative adversarial networks to construct a training framework for the archive text recognition model. Train an initial archive text recognition model using the annotated power archive dataset, use the initial model to predict the unannotated power archive dataset to obtain pseudo-annotations, add the pseudo-annotation samples with high confidence to the annotated dataset, and retrain the model using the updated annotated dataset, and iterate this process. Construct a generator and a discriminator, where the generator generates pseudo-samples and the discriminator distinguishes between real samples and pseudo-samples. Through adversarial training, improve the generalization ability and recognition accuracy of the model, and optimize the model according to the performance of the validation set, such as adjusting hyperparameters, optimizing the model structure, etc.
[0034] Furthermore, the entity recognition module 12 is also used for:
[0035] Load the pre-trained text recognition model, perform supervised learning on the labeled power archive dataset to obtain an initial text recognition model; use the initial text recognition model to predict the unlabeled power archive dataset, generate pseudo-labels, screen the pseudo-labels to obtain target pseudo-labels; merge the target pseudo-labels and the labeled power archive dataset to obtain a target training dataset, and train the text recognition model to obtain the archive text recognition model.
[0036] Specifically, use the labeled power archive dataset to further train the pre-trained model, adjust the model parameters through supervised learning to make it more suitable for the recognition task of power archive texts, thereby obtaining an initial text recognition model. Use the initial text recognition model to predict the unlabeled power archive dataset, and assign a prediction label, that is, a pseudo-label, to each unlabeled sample. Evaluate the quality of the generated pseudo-labels, and screen out the pseudo-labels with high confidence and accurate prediction as target pseudo-labels. Merge the filtered target pseudo-labels with the original labeled power archive dataset to obtain a new and larger target training dataset. Use the target training dataset to further train the initial text recognition model, and improve the performance of the model in the power archive text recognition task by iteratively optimizing the model parameters, and finally obtain the archive text recognition model.
[0037] Furthermore, the entity recognition module 12 is further configured to:
[0038] Construct a dual generator and a multi-level target training data discriminator for the target training data. The dual generator of the target training data includes a first generator and a second generator, and the first generator and the second generator have the same structure; use the first generator to generate initial synthetic data, evaluate the authenticity of the initial synthetic data, and adjust the parameters of the second generator according to the evaluation results to obtain a second optimized generator; optimize the initial synthetic data according to the second optimized generator to obtain target synthetic data; evaluate the target synthetic data through the multi-level target training data discriminator, and add the target synthetic data into the target training dataset according to the evaluation results.
[0039] Specifically, the goal of the first generator is to generate initial synthetic data. New power file text data similar to the labeled power file dataset can be generated according to the characteristics and distribution of the labeled power file dataset. The first generator and the second generator have the same structure, which is a multi-layer neural network that maps the input random noise to the data space through multi-layer non-linear transformations. Using the first generator, a batch of initial synthetic data is generated according to the characteristics and distribution of the labeled power file dataset. Then, the reward function is used to calculate the reward value of the initial synthetic data. With the goal of maximizing the reward, the gradient descent method is used for optimization, and the negative reward value is used as the loss function. The goal of the generator is to minimize the negative reward value, that is, to maximize the reward. The gradient of the loss function with respect to the generator parameters is calculated through the backpropagation algorithm. The gradient reflects the direction and degree of parameter adjustment. According to the calculated gradient, the parameters of the second generator are updated to obtain the second optimized generator. The initial synthetic data is input into the second optimized generator for further optimization and adjustment to obtain the target synthetic data with higher quality. The discriminator is a multi-level network structure that can capture the complex characteristics of power file text data, including multiple convolutional layers, pooling layers, fully connected layers, etc., to achieve accurate evaluation of synthetic data. Using the labeled power file dataset as positive samples and the initial synthetic data as negative samples, the discriminator is trained to enable the discriminator to accurately distinguish real data from synthetic data. The trained multi-level target training data discriminator is used to evaluate the target synthetic data, including the confidence level output by the discriminator, the similarity between the synthetic data and the real data, etc. According to the evaluation results, the target synthetic data with higher quality is selected and added to the target training dataset, which can expand the scale of the training dataset and improve the generalization ability and performance of the file text recognition model.
[0040] Further, the association relationship acquisition module 13 is further configured to:
[0041] Read a predetermined feature picking strategy, and pick features from the historical grid connection scheduling protocol, historical power business license, and historical project approval based on the predetermined feature picking strategy to obtain historical protocol feature information, historical license feature information, and historical approval feature information respectively; perform association analysis based on the historical protocol feature information, historical license feature information, and historical approval feature information to obtain the first association relationship; introduce an association relationship evaluation function to evaluate the accuracy of the first association relationship to obtain the first accuracy; when the first accuracy does not meet the predetermined constraint, correct the first association relationship to obtain the historical target association relationship.
[0042] Specifically, for historical grid connection dispatching agreements, key features are extracted according to the feature picking strategy, such as agreement number, signing date, grid connection capacity, dispatching method, etc., to obtain historical agreement feature information. For historical power business licenses, features are also extracted according to the strategy, such as license number, issuing agency, validity period, business scope, etc., to obtain historical license feature information. For historical project approvals, features such as approval number, approval date, project name, investment scale, etc. are extracted to obtain historical approval feature information. The extracted historical agreement feature information, historical license feature information, and historical approval feature information are subjected to correlation analysis. The internal connections and rules between these feature information are found. For example, it can be analyzed whether a certain grid connection dispatching agreement corresponds to a specific power business license, or whether a certain project approval is associated with a specific grid connection dispatching agreement. An association relationship evaluation function is introduced to evaluate the accuracy of the first association relationship obtained from the analysis. This evaluation function is based on a machine learning model or can also be based on statistical methods. The purpose of the evaluation is to determine the reliability and accuracy of the association relationship. If the evaluation result shows that the accuracy of the first association relationship does not meet the predetermined constraints, for example, the accuracy is lower than the accuracy preset threshold, and the accuracy preset threshold is set by those skilled in the art, the association relationship is corrected, including re-analyzing the feature information, adjusting the association rules, introducing more feature information, etc. The corrected association relationship is called the historical target association relationship.
[0043] Further, the association relationship obtaining module 13 is also used for:
[0044] Convert the historical agreement feature information, historical license feature information, and historical approval feature information into an association analysis data set, where each data in the association analysis data set represents a feature information; traverse the association analysis data set, calculate the occurrence frequency of the first data, and when the occurrence frequency is greater than the preset minimum frequency threshold, add the first data to the descending frequent association data set; sort the association analysis data set according to the descending frequent association data set to obtain a descending association analysis data set; construct a frequent association tree according to the descending association analysis data set, and at the same time construct a frequent head node; randomly extract data from the descending association analysis data set and add it to the frequent head node; mine the frequent association tree according to the frequent head node to generate multiple conditional association pattern bases; construct multiple conditional frequent association trees according to the multiple conditional association pattern bases, and recursively mine the multiple conditional frequent association trees respectively to generate the first association relationship set.
[0045] Specifically, convert historical protocol feature information, historical license feature information, and historical approval feature information into an association analysis data set. In this data set, each data item represents a feature information, such as a protocol number, a license issuing agency, or a project approval date, etc. Traverse the association analysis data set and calculate the occurrence frequency of each feature information. When the occurrence frequency of a certain feature information is greater than the preset minimum frequency threshold, add this feature information to the descending frequent association data set, and sort the association analysis data set according to the descending frequent association data set to obtain a descending association analysis data set. Construct a frequent association tree based on the descending association analysis data set. This tree will be used to store the association relationships between feature information, and at the same time construct frequent head nodes, which will be used as the starting points for mining association patterns. Randomly extract data from the descending association analysis data set and add these data to the frequent head nodes. Mine the frequent association tree according to the frequent head nodes to generate multiple conditional association pattern bases, which represent the potential association relationships between feature information. Further analyze and process each conditional association pattern base to extract useful association rules. Construct multiple conditional frequent association trees according to multiple association pattern bases to store more specific association relationships. Recursively mine each conditional frequent association tree to generate a first association relationship set, which contains all useful association relationships mined from historical protocol feature information, historical license feature information, and historical approval feature information. Finally, output the first association relationship set as the final result.
[0046] Furthermore, the classification and sorting module 16 is also used for:
[0047] Collect user feedback data in real time; according to the user feedback data, identify misclassified archive data, and calculate the error frequency of the misclassified archive data; adjust and optimize the intelligent classification model of the power archive data according to the error frequency.
[0048] Specifically, collect user feedback on the classification results of power archive data in real time. Conduct in-depth analysis on the collected user feedback data to identify misclassified archive data, which is achieved by comparing user feedback with model classification results. Calculate the error frequency of each misclassified archive data, that is, the ratio of the number of times the archive data is misclassified to the total number of classifications. According to the error frequency of the misclassified archive data, adjust the weights of the intelligent classification model of the power archive data. Increase the weight of the misclassified archive data in model training to reduce the possibility of its misclassification. Retrain the model, and use an extended data set containing misclassified archive data to improve the classification performance of the model, which can be achieved through incremental learning or online learning techniques in machine learning algorithms. Verify the performance of the optimized model. Use an independent test data set to evaluate the classification accuracy of the model to ensure that the optimization measures effectively improve the performance of the model.
[0049] Example 2. Please refer to the appendix Figure 2 In this application, a method for classifying and sorting power data is provided. Among them, the method for classifying and sorting power data is applied to a system for classifying and sorting power data. The method for classifying and sorting power data specifically includes the following steps:
[0050] S1: Collect multi-source power archive data, perform semantic matching on the multi-source power archive data, and perform semantic fusion according to the semantic matching results to generate a target power archive data set to be classified.
[0051] S2: Extract the classified power archive data set from the power data archive data system, perform entity recognition on the classified power archive data set through an archive text recognition model based on a predetermined classification entity, and extract multiple historical target entities. Among them, the predetermined classification entity includes grid connection dispatching agreements, power business licenses, and project approvals, and the multiple historical target entities include historical grid connection dispatching agreements, historical power business licenses, and historical project approvals.
[0052] S3: Identify the first association relationship set among the historical grid connection dispatching agreements, historical power business licenses, and historical project approvals, evaluate the accuracy of the first association relationship set, and correct the first association relationship set according to the evaluation results to generate a historical target association relationship set.
[0053] S4: Construct a power archive knowledge graph according to the historical grid connection dispatching agreements, historical power business licenses, historical project approvals, and the historical target association relationship set.
[0054] S5: Extract the classification features of the power archive knowledge graph and train an intelligent classification model for power archive data.
[0055] S6: Classify and sort the target power archive data set to be classified according to the intelligent classification model for power archive data.
[0056] Furthermore, step S2 of this application further includes:
[0057] Perform random data annotation on the classified power archive data set according to the grid connection dispatching agreements, power business licenses, and project approvals to obtain an unannotated power archive data set and an annotated power archive data set; based on self-training and a generative adversarial network, train the archive text recognition model according to the unannotated power archive data set and the annotated power archive data set.
[0058] Furthermore, step S2 of this application further includes:
[0059] Load a pre-trained text recognition model, perform supervised learning on the labeled power file dataset to obtain an initial text recognition model; use the initial text recognition model to predict the unlabeled power file dataset, generate pseudo-labels, screen the pseudo-labels to obtain target pseudo-labels; merge the target pseudo-labels and the labeled power file dataset to obtain a target training dataset, and train the text recognition model to obtain the file text recognition model.
[0060] Further, step S2 of this application further includes:
[0061] Construct a dual generator and a multi-level target training data discriminator for the target training data. The dual generator of the target training data includes a first generator and a second generator, and the first generator and the second generator have the same structure; use the first generator to generate initial synthetic data, perform authenticity evaluation on the initial synthetic data, and adjust the parameters of the second generator according to the evaluation results to obtain a second optimized generator; optimize the initial synthetic data according to the second optimized generator to obtain target synthetic data; evaluate the target synthetic data through the multi-level target training data discriminator, and add the target synthetic data to the target training dataset according to the evaluation results.
[0062] Further, step S3 of this application further includes:
[0063] Read a predetermined feature picking strategy, and pick features from the multiple historical target entities based on the predetermined feature picking strategy. The multiple target entities include historical grid connection dispatching protocols, historical power business licenses, and historical project approvals, and obtain historical protocol feature information, historical license feature information, and historical approval feature information respectively; perform correlation analysis according to the historical protocol feature information, historical license feature information, and historical approval feature information to obtain the first correlation relationship; introduce a correlation relationship evaluation function to evaluate the accuracy of the first correlation relationship to obtain the first accuracy; when the first accuracy does not meet the predetermined constraints, correct the first correlation relationship to obtain the historical target correlation relationship.
[0064] Further, step S3 of this application further includes:
[0065] Convert the historical protocol feature information, historical license feature information, and historical approval feature information into an association analysis data set, where each data in the association analysis data set represents a feature information; traverse the association analysis data set, calculate the occurrence frequency of the first data, and when the occurrence frequency is greater than the preset minimum frequency threshold, add the first data to the descending frequent association data set; sort the association analysis data set according to the descending frequent association data set to obtain a descending association analysis data set; construct a frequent association tree based on the descending association analysis data set, and at the same time construct a frequent header node; randomly extract data from the descending association analysis data set and add it to the frequent header node; mine the frequent association tree according to the frequent header node to generate multiple conditional association pattern bases; construct multiple conditional frequent association trees based on the multiple conditional association pattern bases, and recursively mine the multiple conditional frequent association trees respectively to generate the first association relationship set.
[0066] Further, step S6 of this application further includes:
[0067] Collect user feedback data in real time; according to the user feedback data, identify misclassified archive data, and calculate the error frequency of the misclassified archive data; optimize the intelligent classification model of the power archive data according to the error frequency.
[0068] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0069] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of this application and its equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. A system for classifying and sorting power data, characterized in that, Including: A semantic fusion module, which is used to collect multi-source power archive data, perform semantic matching on the multi-source power archive data, perform semantic fusion according to the semantic matching result, and generate a target power archive data set to be classified; An entity recognition module, which is used to extract the classified power archive data set from the power data archive data system, perform entity recognition on the classified power archive data set through an archive text recognition model based on predetermined classification entities, and extract multiple historical target entities, where the predetermined classification entities include grid connection dispatching protocols, power business licenses, and project approvals, and the multiple historical target entities include historical grid connection dispatching protocols, historical power business licenses, and historical project approvals; An association relationship acquisition module, which is used to identify the first association relationship set among the historical grid connection dispatching protocols, historical power business licenses, and historical project approvals, evaluate the accuracy of the first association relationship set, and correct the first association relationship set according to the evaluation result to generate a historical target association relationship set; A knowledge graph construction module, which is used to construct a power archive knowledge graph according to the historical grid connection dispatching protocols, historical power business licenses, historical project approvals, and the historical target association relationship set; A classification model acquisition module, which is used to extract the classification features of the power archive knowledge graph and train an intelligent classification model for power archive data; A classification and sorting module, which is used to classify and sort the target power archive data set to be classified according to the intelligent classification model for power archive data; The entity recognition module is further used for: Constructing a dual generator and a multi-level target training data discriminator for target training data, where the dual generator of the target training data includes a first generator and a second generator, and the first generator and the second generator have the same structure; Using the first generator to generate initial synthetic data, performing authenticity evaluation on the initial synthetic data, and adjusting the parameters of the second generator according to the evaluation result to obtain a second optimized generator; Optimizing the initial synthetic data according to the second optimized generator to obtain target synthetic data; Evaluating the target synthetic data through the multi-level target training data discriminator, and adding the target synthetic data into the target training data set according to the evaluation result; The multi-level target training data discriminator is a multi-level network structure that can capture the complex features of power archive text data, including multiple convolutional layers, pooling layers, and fully connected layers, so as to achieve accurate evaluation of synthetic data.
2. The system for classifying and sorting power data according to claim 1, characterized in that, The entity recognition module is further used for: Randomly labeling the classified power archive data set according to the grid connection dispatching protocol, power business license, and project approval to obtain an unlabeled power archive data set and a labeled power archive data set; Training the archive text recognition model based on self-training and a generative adversarial network according to the unlabeled power archive data set and the labeled power archive data set.
3. A system for classifying and organizing power data according to claim 2, characterized in that, The entity recognition module is further used for: Loading a pre-trained text recognition model and performing supervised learning on the labeled power archive data set to obtain an initial text recognition model; Use the initial text recognition model to predict the unlabeled power file dataset, generate pseudo-labels, and screen the pseudo-labels to obtain target pseudo-labels; Merge the target pseudo-labels and the labeled power file dataset to obtain a target training dataset, and train the text recognition model to obtain the file text recognition model.
4. The system for classifying and sorting power data according to claim 1, wherein The association relationship acquisition module is further configured to: Read a predetermined feature picking strategy, and pick features from the historical grid connection scheduling protocol, historical power business license, and historical project approval based on the predetermined feature picking strategy to obtain historical protocol feature information, historical license feature information, and historical approval feature information respectively; Perform association analysis based on the historical protocol feature information, historical license feature information, and historical approval feature information to obtain the first association relationship; Introduce an association relationship evaluation function to evaluate the accuracy of the first association relationship to obtain a first accuracy; When the first accuracy does not meet the predetermined constraint, correct the first association relationship to obtain the historical target association relationship.
5. The system for classifying and sorting power data according to claim 4, characterized in that, The association relationship acquisition module is further configured to: Convert the historical protocol feature information, historical license feature information, and historical approval feature information into an association analysis dataset, where each data in the association analysis dataset represents a feature information; Traverse the association analysis dataset, calculate the occurrence frequency of the first data, and when the occurrence frequency is greater than the preset minimum frequency threshold, add the first data to the descending frequent association dataset; Sort the association analysis dataset according to the descending frequent association dataset to obtain a descending association analysis dataset; Construct a frequent association tree according to the descending association analysis dataset, and construct a frequent header node at the same time; Randomly extract data from the descending association analysis dataset and add it to the frequent header node; Mine the frequent association tree according to the frequent header node to generate multiple conditional association pattern bases; Construct multiple conditional frequent association trees according to the multiple conditional association pattern bases, and perform recursive mining on the multiple conditional frequent association trees respectively to generate the first association relationship set.
6. The system for classifying and sorting power data according to claim 1, characterized in that, The classification and sorting module is further configured to: Collect user feedback data in real time; Identify misclassified file data according to the user feedback data, and calculate the error frequency of the misclassified file data; Optimize the intelligent classification model of the power file data according to the error frequency.
7. A method for classifying and sorting power data, characterized in that, The method for classifying and sorting power data is implemented by the system for classifying and sorting power data according to any one of claims 1 to 6, wherein the method for classifying and sorting power data includes: Collect multi-source power file data, perform semantic matching on the multi-source power file data, and perform semantic fusion according to the semantic matching result to generate a target power file dataset to be classified; Extract the classified power archive data set from the power data archive data system, and perform entity recognition on the classified power archive data set through an archive text recognition model based on predetermined classification entities to extract multiple historical target entities, where the predetermined classification entities include grid connection dispatching agreements, power business licenses, and project approvals, and the multiple historical target entities include historical grid connection dispatching agreements, historical power business licenses, and historical project approvals; Identify the first association relationship set among the historical grid connection dispatching agreements, historical power business licenses, and historical project approvals, evaluate the accuracy of the first association relationship set, and correct the first association relationship set according to the evaluation results to generate a historical target association relationship set; Construct a power archive knowledge graph based on the historical grid connection dispatching agreements, historical power business licenses, historical project approvals, and the historical target association relationship set; Extract the classification features of the power archive knowledge graph and train a power archive data intelligent classification model; Classify and organize the target power archive data set to be classified according to the power archive data intelligent classification model.
Citation Information
Patent Citations
Knowledge graph-based airport data classification and grading method and system
CN117473431A
Intelligent management system for enterprise outsourcing
CN118052408A