Method and system for automatic labeling of power standard knowledge and storage medium
By designing a system for collecting, cleaning, expanding, and labeling power standards, and combining corrections by a power industry expert group with deep learning network training, the problem of low efficiency in traditional power standard knowledge labeling has been solved, enabling rapid, accurate, and automated labeling of power standard knowledge.
Patent Information
- Application Number
- CN202410829538.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-06-25
AI Technical Summary
Traditional methods of labeling electrical standards are inefficient, prone to errors, and difficult to guarantee consistency and accuracy.
We collected a power standard knowledge dataset, cleaned and standardized the data, expanded the text using a large language model, designed a power knowledge tagging system, and achieved automated tagging through correction by a power expert group and optimization by deep learning network training.
It enables rapid and accurate labeling of power standard knowledge, improves labeling efficiency and quality, and ensures the consistency and accuracy of labeling.
Smart Images

Figure CN118626594B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automated labeling technology, specifically to a method, system, and storage medium for automated labeling of electrical standard knowledge. Background Technology
[0002] With the rapid development of the power industry and the in-depth advancement of smart grid construction, the management and application of power standards knowledge has become increasingly important. Power standards knowledge not only covers the operating specifications and safety standards of power equipment, but also involves multiple aspects such as power quality and power market operation, and is the foundation for ensuring the safe, stable, and efficient operation of the power system.
[0003] However, with the explosive growth of power data and the continuous expansion of the power knowledge system, traditional power standard knowledge annotation methods are no longer sufficient to meet practical needs. Traditional methods often rely on manual knowledge annotation, which is not only inefficient and prone to errors, but also faces the problem of high labor costs. In addition, due to the complexity and diversity of power knowledge, manual annotation makes it difficult to guarantee the consistency and accuracy of annotation, which to some extent restricts the effective application of power standard knowledge. Summary of the Invention
[0004] This application provides a method, system, and storage medium for automated labeling of electrical standard knowledge, which addresses the technical problem that manual labeling in the prior art is difficult to guarantee in terms of consistency and accuracy.
[0005] In view of the above problems, this application provides a method, system and storage medium for automated labeling of electrical standard knowledge.
[0006] The first aspect of this application provides a method for automated labeling of electrical standard knowledge, the method comprising:
[0007] A source power standard knowledge dataset is collected, and the dataset is cleaned and standardized to obtain a standard power knowledge dataset. A large language model is used to augment the standard power knowledge dataset, generating a power standard knowledge text description dataset. Text association filtering rules are obtained, and the power standard knowledge text description dataset is subjected to multiple filtering processes based on these rules to obtain an augmented power standard knowledge text dataset. A power knowledge tagging system is designed according to power knowledge application needs, and the standard power knowledge dataset and the augmented power standard knowledge text dataset are labeled based on this system to obtain an labeled power standard knowledge text dataset. The labeled power standard knowledge text dataset is then corrected by a power domain knowledge expert group to obtain a labeled power standard knowledge text sample set. A deep learning network is used to train and optimize the labeled power standard knowledge text sample set to obtain a power standard knowledge labeling model for automated labeling of power standard knowledge.
[0008] Furthermore, the acquisition of the standard power knowledge dataset also includes:
[0009] A text knowledge cleaning program is obtained, which includes non-text content removal, special symbol removal, HTML tag removal, and text format unification. Based on the text knowledge cleaning program, the source power standard knowledge dataset is cleaned to obtain a usable power standard knowledge dataset. A custom power knowledge dictionary is established, and the usable power standard knowledge dataset is segmented using the custom power knowledge dictionary to obtain a segmented power standard knowledge dataset. A stop word list is constructed, and the segmented power standard knowledge dataset is removed based on the stop word list to obtain the standard power knowledge dataset.
[0010] Furthermore, the obtained power standard knowledge text augmented dataset also includes:
[0011] Based on the text association filtering rules, a content duplication filtering step, a relevance filtering step, and a semantic clarity filtering step are determined. Based on the content duplication filtering step, a text removal algorithm is used to initially filter duplicate content in the power standard knowledge text description dataset, resulting in a first power knowledge text dataset. Based on the relevance filtering step, a keyword matching algorithm is used to perform a secondary filtering of irrelevant content in the first power knowledge text dataset, resulting in a second power knowledge text dataset. Based on the semantic clarity filtering step, a natural language processing algorithm is used to filter semantically unclear content in the second power knowledge text dataset, resulting in the expanded power standard knowledge text dataset.
[0012] Furthermore, the design of the power knowledge tagging system based on the application needs of power knowledge also includes:
[0013] The application requirements for power knowledge are labeled with attributes to obtain a set of knowledge labeling requirements attributes, which includes the labeling purpose, application scenario, and users. The set of knowledge labeling requirements attributes is then analyzed for classification according to the power operation and maintenance knowledge architecture to determine the knowledge tag classification method. Based on the knowledge tag classification method, a tag hierarchy is designed for the power operation and maintenance knowledge architecture to obtain a knowledge tag classification hierarchy. Finally, based on the knowledge tag classification method and the knowledge tag classification hierarchy, the power knowledge tag system is determined.
[0014] Furthermore, the obtained power standard knowledge text annotation sample set also includes:
[0015] Based on the review and correction of the power standard knowledge text annotation dataset by each expert in the power field knowledge expert group, an expert annotation correction information set is obtained; by evaluating the modification index of each correction information in the expert annotation correction information set by each expert, an initial annotation information correction index set is obtained; based on the initial annotation information correction index set and the index weighting calculation of each correction information by the power field knowledge expert group, a weighted annotation information correction index set is determined; the annotation data in the weighted annotation information correction index set that reaches the preset correction index threshold is corrected to obtain the power standard knowledge text annotation sample set.
[0016] Furthermore, the determination of the weighted correction index set for the annotation information also includes:
[0017] Trustworthiness assessments are conducted on each expert in the power field knowledge expert group based on their knowledge and experience in the power industry, resulting in a trustworthiness set for the expert group. Trustworthiness correction factor information is generated based on the trustworthiness set of the expert group. The weighted calculation result of the initial annotation information correction index set and the trustworthiness correction factor information is used as the annotation information weighted correction index set for each correction information.
[0018] Furthermore, the power standard knowledge annotation model also includes:
[0019] A basic power knowledge annotation model is obtained by iteratively training the power standard knowledge text annotation sample set using a deep learning network; the performance of the basic power knowledge annotation model is verified and evaluated using a model loss function to obtain a set of annotation model performance evaluation parameters; based on the annotation model performance evaluation parameter set, the model loss function is minimized using a gradient descent algorithm to obtain model optimization parameters; based on the model optimization parameters, the model parameters of the basic power knowledge annotation model are updated and configured to obtain the power standard knowledge annotation model.
[0020] A second aspect of this application provides an automated labeling system for electrical standard knowledge, the system comprising:
[0021] The system includes: a power standard knowledge text description dataset generation module, which collects source power standard knowledge datasets, cleans and standardizes them to obtain standard power knowledge datasets, and uses a large language model to augment the text in the standard power knowledge datasets to generate a power standard knowledge text description dataset; a power standard knowledge text augmentation dataset acquisition module, which acquires text association filtering rules and performs multiple filtering processes on the power standard knowledge text description datasets based on these rules to obtain a power standard knowledge text augmentation dataset; and a power standard knowledge text annotation dataset acquisition module. The module includes a power standard knowledge text annotation dataset acquisition module, which designs a power knowledge tagging system based on the application needs of power knowledge, and annotates the standard power knowledge dataset and the power standard knowledge text extended dataset based on the power knowledge tagging system to obtain a power standard knowledge text annotation dataset; and a power standard knowledge tag automated annotation module, which corrects the power standard knowledge text annotation dataset through a power domain knowledge expert group to obtain a power standard knowledge text annotation sample set, and uses a deep learning network to train and optimize the power standard knowledge text annotation sample set to obtain a power standard knowledge annotation model for automated power standard knowledge tag annotation.
[0022] Thirdly, this application provides a computer-readable storage medium storing computer instructions that cause a processor to execute the steps of any of the methods described in the first aspect above.
[0023] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0024] This application collects a source power standard knowledge dataset, cleans and standardizes it to obtain a standard power knowledge dataset, and uses a large language model to augment the standard power knowledge dataset, generating a power standard knowledge text description dataset. It then obtains text association filtering rules and performs multiple filtering processes on the power standard knowledge text description dataset based on these rules, resulting in an augmented power standard knowledge text dataset. A power knowledge tagging system is designed according to the application needs of power knowledge, and the standard power knowledge dataset and the augmented power standard knowledge text dataset are labeled based on this system, resulting in an labeled power standard knowledge text dataset. The labeled power standard knowledge text dataset is then corrected by a power domain knowledge expert group to obtain a labeled power standard knowledge text sample set. A deep learning network is used to train and optimize the labeled power standard knowledge text sample set, resulting in a power standard knowledge labeling model for automated labeling of power standard knowledge. This invention solves the technical problem that manual labeling in existing technologies struggles to guarantee consistency and accuracy. Through an automated labeling method, it achieves rapid and accurate labeling of power standard knowledge, thereby improving labeling efficiency and quality. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 A schematic diagram of the automated labeling method for electricity standard knowledge provided in this application embodiment;
[0027] Figure 2 A schematic diagram of the structure of an automated labeling system for electricity standard knowledge provided in this application embodiment.
[0028] Figure labeling: Module 11 for generating power standard knowledge text description dataset, Module 12 for obtaining power standard knowledge text extended dataset, Module 13 for obtaining power standard knowledge text annotation dataset, and Module 14 for automatic annotation of power standard knowledge tags. Detailed Implementation
[0029] This application addresses the technical problem that manual labeling in existing technologies struggles to guarantee consistency and accuracy by providing an automated labeling method, system, and storage medium for electrical standard knowledge. Through automated labeling, it achieves rapid and accurate labeling of electrical standard knowledge, thereby improving labeling efficiency and quality.
[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.
[0032] Example 1
[0033] like Figure 1 As shown, this application provides an automated labeling method for electrical standard knowledge, the method comprising:
[0034] Step S1: Collect the source power standard knowledge dataset, perform data cleaning and standardization on the source power standard knowledge dataset to obtain the standard power knowledge dataset, and use a large language model to augment the text of the standard power knowledge dataset to generate a power standard knowledge text description dataset.
[0035] In this embodiment, the source power standard knowledge dataset is the starting point for the automated annotation process. The source power standard knowledge dataset refers to a collection of raw, unprocessed power standard knowledge data obtained from various sources. When collecting the source power standard knowledge dataset, web crawlers, API calls, or other methods are used to retrieve power standard knowledge data from professional databases within the power industry.
[0036] After data collection, the source power standard knowledge dataset undergoes data cleaning and standardization. Data cleaning removes noise, duplicates, and errors, improving data quality. Standardization unifies data format and units, ensuring consistency and comparability. These steps result in the standard power knowledge dataset.
[0037] Subsequently, large language models such as BERT and GPT were used to augment the text of the standard power knowledge dataset. During the augmentation process, the text from the standard power knowledge dataset was input into the large language model, which then output new text related to the input text. This new text serves as an explanation, expansion, or supplement to the original knowledge. Finally, the generated new text was integrated with the original dataset to form a power standard knowledge text description dataset.
[0038] Step S2: Obtain text association filtering rules, and perform multiple filtering processes on the power standard knowledge text description dataset based on the text association filtering rules to obtain the power standard knowledge text augmentation dataset;
[0039] In this embodiment, the text association filtering rules are formulated through expert review or historical data analysis. These rules include content relevance, semantic clarity, information value, and logical coherence. Content relevance means the text content should be directly related to power standard knowledge, covering various areas of power technology, such as power generation, transmission, and distribution. Semantic clarity means the text should be clearly expressed, avoiding the use of vague or ambiguous words and sentence structures. Information value means the text should contain valuable power standard knowledge, providing effective information for subsequent annotation and model training. Logical coherence means the text should have logical coherence, forming a complete and meaningful knowledge expression.
[0040] After obtaining the text association filtering rules, the power standard knowledge text description dataset underwent multiple filtering processes. First, based on content relevance rules, the original text description dataset was initially filtered to remove texts irrelevant or weakly relevant to power standard knowledge. Then, natural language processing techniques, such as lexical analysis and syntactic analysis, were used to perform semantic clarity analysis on the initially filtered texts. By calculating indicators such as semantic complexity and sentence structure rationality, texts that are semantically clear and easy to understand were further filtered. Combining the knowledge and experience of power industry experts, the information value of the texts filtered for semantic clarity was assessed. By determining whether the texts contain valuable power standard knowledge, and the importance and practicality of this knowledge, high-value texts were further filtered. Finally, for the texts filtered for information value, logical coherence analysis was performed. By examining the logical relationships between sentences and the connections between paragraphs, the filtered texts were ensured to be logically coherent and consistent.
[0041] Through the above multiple filtering processes, an expanded dataset of power standard knowledge text was obtained.
[0042] Step S3: Design an electric knowledge tagging system based on the application requirements of electric knowledge, and annotate the standard electric knowledge dataset and the extended dataset of electric standard knowledge text based on the electric knowledge tagging system to obtain an annotated dataset of electric standard knowledge text.
[0043] In this embodiment, the application requirements for power knowledge include multiple aspects such as power grid operation, equipment maintenance, energy management, and market analysis. A power knowledge tagging system is designed based on these application requirements, covering all aspects of the power sector, including but not limited to power generation technology, transmission technology, distribution technology, the power market, and energy policy.
[0044] When annotating the standard power knowledge dataset and the extended power standard knowledge text dataset based on the power knowledge tagging system, appropriate annotation tools, such as Label Studio, are selected according to the annotation needs and scale. During annotation, for each text, the relevant power field knowledge is analyzed to determine its tag category. Based on the power knowledge tagging system, corresponding tags are assigned to each text. After annotation, the results are cleaned and organized, including removing duplicate data, handling outliers, and formatting data to ensure the accuracy and consistency of the annotated dataset. Finally, the annotation results of the standard power knowledge dataset and the extended power standard knowledge text dataset are integrated to form the power standard knowledge text annotated dataset.
[0045] Step S4: Correct the power standard knowledge text annotation dataset to obtain a power standard knowledge text annotation sample set. Use a deep learning network to train and optimize the power standard knowledge text annotation sample set to obtain a power standard knowledge annotation model. Then, automatically annotate power standard knowledge tags based on the power standard knowledge annotation model.
[0046] In this embodiment of the application, the power sector knowledge expert group corrects the power standard knowledge text annotation dataset. By checking the accuracy, completeness, and consistency of the annotations, it ensures that each piece of text is correctly assigned the corresponding label. For any erroneous annotations discovered during the review process, the expert group corrects them one by one, including correcting incorrect labels, supplementing missing labels, and adjusting inaccurate labels.
[0047] Based on the characteristics and requirements of the power standard knowledge text annotation dataset, a suitable deep learning model was selected for training. The modified power standard knowledge text annotation dataset underwent preprocessing, including text cleaning, word segmentation, and vectorization, to enable the model to better understand and process text data. The deep learning model was then trained using the preprocessed dataset. During training, the model learned how to extract features from the text and predict corresponding labels. The model's performance was optimized by continuously adjusting its parameters and structure. The model's performance was evaluated using evaluation metrics, and necessary tuning operations were performed based on the evaluation results, such as adjusting the learning rate and increasing the number of training epochs, to further improve the model's annotation performance.
[0048] Finally, the trained power standard knowledge annotation model is used to automatically annotate power standard knowledge tags.
[0049] Furthermore, in step S1 of the method provided in the application embodiment, obtaining the standard power knowledge dataset includes:
[0050] S11: Obtain the text knowledge cleaning program, which includes non-text content removal, special symbol removal, HTML tag removal, and text format unification;
[0051] S12: Based on the text knowledge cleaning program, perform data cleaning on the source power standard knowledge dataset to obtain a usable power standard knowledge dataset;
[0052] S13: Establish a custom dictionary of power knowledge, and use the custom dictionary of power knowledge to perform word segmentation on the available power standard knowledge dataset to obtain a power standard knowledge word segmentation dataset;
[0053] S14: Construct a stop word list, and perform stop word removal processing on the power standard knowledge word segmentation dataset based on the stop word list to obtain the standard power knowledge dataset.
[0054] In this embodiment, a text knowledge cleaning procedure is obtained, which includes non-text content removal, special symbol removal, HTML tag removal, and text format unification. Non-text content removal involves identifying and removing non-text elements from the text, such as images, audio, video, and other multimedia content, retaining only plain text information. Special symbol removal involves deleting special symbols from the text, such as punctuation marks and emoticons. HTML tag removal involves removing HTML tags from text obtained from sources such as web pages; these tags are web page formatting marks. Text format unification standardizes the text format to a standard format, such as unifying capitalization and formatting dates and numbers into a uniform form.
[0055] Use the text knowledge cleaning program to clean the source power standard knowledge dataset, clean the imported data, remove non-text content, special symbols, and HTML tags, and unify the text format. During the cleaning process, conduct data quality checks to ensure that the quality of the cleaned data meets the requirements and there are no errors or omissions. Export the cleaned data to obtain an available power standard knowledge dataset.
[0056] Collect power knowledge from sources such as professional books, papers, and specifications in the power field, including professional terms, abbreviations, common expressions, etc. Organize the collected power knowledge into the form of a custom dictionary to ensure that each word has a clear definition and boundary. Use a word segmentation tool in combination with the custom dictionary of power knowledge to perform word segmentation on the available power standard knowledge dataset. Check the word segmentation results to ensure that the word segmentation is accurate and there are no incorrect segmentations or missing words. Through this step, obtain the power standard knowledge word segmentation dataset.
[0057] Stop words refer to words that frequently appear in the text but do not contribute substantially to the understanding of the text content, such as common function words like "de", "le", "zai", etc. When constructing the stop word list, collect common stop words from existing stop word libraries or according to the text characteristics of the power field. Organize the collected stop words into a stop word list.
[0058] Use the stop word list to remove the words that appear in the stop word list from the dataset by comparing the words after word segmentation, and perform stop word removal processing on the power standard knowledge word segmentation dataset. Check the processing results after removing the stop words to ensure that no important words are accidentally deleted. Through the above steps, obtain the standard power knowledge dataset.
[0059] Furthermore, the steps in the method provided by the application embodiment to obtain the power standard knowledge text expansion dataset in step S2 include:
[0060] S2l: Determine the content duplication screening step, relevance screening step, and semantic clarity screening step according to the text association screening rules;
[0061] S22: Based on the content duplication screening step, use a text removal algorithm to perform a preliminary screening of duplicate content on the power standard knowledge text description dataset to obtain the first power knowledge text dataset;
[0062] S23: Based on the relevance screening step, use a keyword matching algorithm to perform a secondary screening of irrelevant content on the first power knowledge text dataset to obtain the second power knowledge text dataset;
[0063] S24: Based on the semantic clarity filtering step, a natural language processing algorithm is used to filter semantically unclear content in the second power knowledge text dataset to obtain the power standard knowledge text extended dataset.
[0064] In this embodiment, a content duplication screening step is first determined based on text association filtering rules. This step aims to remove duplicate or highly similar text content from the dataset, avoiding redundant information in subsequent analysis. During the content duplication screening step, text similarity algorithms, such as cosine similarity and Jaccard similarity, are used to compare the similarity between texts, and an appropriate threshold is set to identify duplicate content. Next, a keyword matching algorithm is used to construct a keyword database in the power industry and match it with the text content to determine the relevance of the text. Finally, natural language processing algorithms, such as syntactic analysis and semantic role labeling, are used to parse the structure and semantic information of the text.
[0065] Based on the determined content duplication screening steps, a text removal algorithm is used to initially screen for duplicate content in the power standard knowledge text description dataset. First, the power standard knowledge text description dataset is loaded, and the necessary tools and libraries for the text removal algorithm are prepared. Then, each text record in the dataset is traversed, and a text similarity algorithm is used to calculate the similarity between each text and other texts. Next, duplicate content is identified based on the similarity calculation results. Specifically, a similarity threshold is set; when the similarity between two texts exceeds this threshold, they are considered duplicates. Then, redundant items in the duplicate texts are removed, retaining only one as a representative. After this step, the preliminary screening result of duplicate content removal is obtained, which is the first power knowledge text dataset.
[0066] Following the relevance screening steps, a keyword matching algorithm is used to perform a secondary screening of irrelevant content in the first power knowledge text dataset to further refine the dataset. During this secondary screening, a keyword database for the power industry is first constructed. This database contains professional terminology, commonly used vocabulary, and specific expressions related to power. These keywords are collected from professional books, papers, and standards in the power industry, and then organized and categorized. Next, a keyword matching algorithm is used to match each text in the first power knowledge text dataset. During the matching process, the text content is segmented and compared with keywords in the keyword database. If a text contains a certain number of keywords, and these keywords have a high weight in the text—for example, the keywords appear frequently or are located in important positions in the text—then the text is considered to have a high relevance to the power industry. Through keyword matching, texts highly relevant to the power industry are selected, while texts irrelevant or with low relevance are excluded. The above steps yield the second power knowledge text dataset.
[0067] Natural language processing algorithms are used to parse each text in the second power knowledge text dataset to obtain its structural and semantic information. Then, based on the parsing results, the semantic clarity is assessed by checking the completeness of the grammatical structure, the fluency of sentences, and the clarity of meaning. Specific judgment criteria are set according to the actual situation; for example, rules are established such as avoiding excessive grammatical errors and ambiguous sentences. Text with unclear semantics is removed from the dataset to ensure that the final power standard knowledge text augmented dataset has clear and accurate semantics. Through this processing step, the power standard knowledge text augmented dataset is obtained.
[0068] Furthermore, in step S3 of the method provided in the application embodiment, a power knowledge tagging system is designed based on the application needs of power knowledge, including:
[0069] S31: Extract the annotation attributes from the power knowledge application requirements to obtain a knowledge annotation requirement attribute set, which includes annotation purpose, application scenario and users;
[0070] S32: Analyze the classification method of the knowledge labeling requirement attribute set according to the power operation and maintenance knowledge architecture, and determine the knowledge label classification method;
[0071] S33: Based on the knowledge tag classification method, design the tag hierarchy of the power operation and maintenance knowledge architecture to obtain the knowledge tag classification hierarchy;
[0072] S34: Determine the power knowledge tag system based on the knowledge tag classification method and the knowledge tag classification hierarchy.
[0073] In this embodiment, the application requirements for power knowledge are first analyzed to understand the underlying business logic and actual needs. This includes acquiring specific workflows and scenarios related to power operation and maintenance, fault handling, and equipment maintenance. Next, key attributes related to knowledge annotation are extracted. These key attributes include annotation purpose, application scenario, and users. Annotation purpose includes improving knowledge retrieval efficiency and supporting decision analysis. Application scenarios include online learning and fault diagnosis. Users include operation and maintenance personnel and management personnel. Finally, the extracted annotation attributes are integrated into a set of knowledge annotation requirement attributes.
[0074] The power operation and maintenance knowledge architecture encompasses all aspects of the power system, including power generation, transmission, substation, and distribution, as well as related equipment, technologies, and management. The process involves aligning the knowledge annotation requirement attribute set with the power operation and maintenance knowledge architecture, matching the purpose, application scenarios, and user attributes in the annotation requirements with specific knowledge points in the knowledge architecture. This alignment process determines which knowledge points are the focus of annotation and which attributes are instructive for knowledge point classification. The results are then analyzed to determine the classification method for knowledge tags. Classification can be based on the nature of the knowledge points, such as categorizing them as technical, management, or safety-related; it can also be based on application scenarios, such as categorizing them as routine operation and maintenance, fault handling, and equipment repair; or it can be based on the user's role, such as categorizing knowledge points as essential for operation and maintenance personnel or as reference for management personnel.
[0075] After determining the classification method for knowledge tags, the next step is to design a hierarchical tag structure for the power operation and maintenance knowledge architecture. This creates a clear and logically structured knowledge tag system to better organize and present power operation and maintenance knowledge. The tag hierarchy design first divides power operation and maintenance knowledge into different levels based on the classification method, the granularity of knowledge points, and their importance, ensuring the hierarchy and logic of the tags. Next, corresponding tags are designed for each level. These tags accurately reflect the characteristics and content of the knowledge at that level while maintaining relevance and consistency with tags from other levels. Finally, all tags are integrated into a knowledge tag classification hierarchy.
[0076] After designing the knowledge tag classification methods and hierarchies, these methods and hierarchies were integrated to form a complete tag system framework. This framework comprehensively covers all aspects of power operation and maintenance knowledge while maintaining structural clarity and logical rigor. Next, the tag system was further optimized and improved by adjusting tag granularity, merging or splitting certain tags, and adding necessary metadata to ensure its accuracy and usability. Finally, the power knowledge tag system was obtained.
[0077] Furthermore, in step S4 of the method provided in the application embodiment, obtaining the power standard knowledge text annotation sample set includes:
[0078] S41: Based on the review and correction of the power standard knowledge text annotation dataset by each expert in the power field knowledge expert group, a set of expert annotation correction information is obtained;
[0079] S42: By evaluating the modification index of each correction information in the expert annotation correction information set by each expert, an initial annotation information correction index set is obtained;
[0080] S43: Based on the initial annotation information correction index set and the power field knowledge expert group, perform index weighting calculation on each correction information to determine the annotation information weighted correction index set;
[0081] S44: Correct the labeled data that reach the preset correction index threshold in the weighted correction index set of the labeled information to obtain the power standard knowledge text labeled sample set.
[0082] In this embodiment, the power standard knowledge text annotation dataset is submitted to a power field knowledge expert group for review and correction. Each expert, based on their professional knowledge and experience, reviews the annotation information in the dataset one by one, identifies errors or inaccuracies, and proposes corresponding correction suggestions. The experts' correction suggestions include modifying the annotation category, adjusting the annotation position, and adding missing annotations. Finally, all the experts' correction suggestions are compiled into an expert annotation correction information set.
[0083] After obtaining the set of expert-annotated correction information, I evaluated the modification index of each piece of information. The modification index is a quantitative indicator used to measure the importance and urgency of the correction information. During the evaluation process, each expert gave an evaluation value for each piece of information based on factors such as the nature of the correction information, its scope of impact, and the difficulty of correction. Finally, the evaluation values of all experts for the same piece of information were summarized, and the average or weighted average was calculated as the initial modification index for that piece of information.
[0084] After obtaining the initial set of correction indices for the annotation information, a weighted calculation is performed to determine the weighted correction indices for the annotation information. During the weighted calculation, information such as the qualifications, experience, and research achievements of experts is collected, and a weight coefficient is assigned to each expert. Then, the initial correction index for each piece of corrected information is multiplied by the corresponding expert's weight coefficient to obtain the weighted correction index. Finally, all weighted correction indices are aggregated into a set of weighted correction indices for the annotation information.
[0085] After obtaining the weighted correction index set for the annotation information, a correction index threshold is set to filter out the annotation data that needs priority correction. The correction index threshold is set based on a comprehensive consideration of the accuracy and importance of the annotation information. When the weighted correction index exceeds this threshold, it means that the corresponding annotation data has significant errors or uncertainties, and further correction is required. Next, the weighted correction index set for the annotation information is traversed, and the weighted correction index of each annotation data is compared with the preset threshold one by one. When the weighted correction index of a certain annotation data is found to exceed the threshold, it is marked as annotation data that needs correction. For the annotation data that needs correction, review and analysis are performed to check whether the annotation category is correct, whether the annotation position is accurate, and whether the annotation text is clear. Based on the review and analysis, correction processing is carried out. Correction processing includes changing the annotation category, adjusting the annotation position, and correcting errors in the annotation text. Appropriate correction measures are taken according to the specific situation to ensure that the corrected annotation data is more accurate and reliable. After the correction processing is completed, the power standard knowledge text annotation dataset is updated, and the original erroneous data is replaced with the corrected annotation data. Through the above process, a power standard knowledge text annotation sample set is obtained.
[0086] Furthermore, in step S43 of the method provided in the application embodiment, determining the set of weighted correction indices for the annotation information includes:
[0087] Based on their knowledge and experience in the power industry, the trustworthiness of each expert in the power field knowledge expert group was assessed to obtain the trustworthiness set of the expert group.
[0088] Based on the trustworthiness set of the expert group, trustworthiness correction factor information is generated;
[0089] The weighted calculation result of the initial annotation information correction index set and the trust correction factor information is used as the annotation information weighted correction index set of each correction information.
[0090] In this embodiment, the knowledge, experience, and skill level of experts in the power industry have a significant impact on the accuracy and reliability of the labeled dataset. Therefore, a trustworthiness assessment is conducted on each expert in the power industry knowledge expert group to quantify their professional competence and level of trust. During the assessment, firstly, basic information about each expert is collected, such as educational background, work experience, and professional field. Secondly, the experts' knowledge, experience, and practical achievements in the power industry are obtained. These include the expert's participation in power projects, published academic papers, patents obtained, and industry recognition. Next, other experts or relevant institutions in the industry evaluate the target expert group. The evaluation is based on the experts' professional skills, problem-solving abilities, and industry reputation. By collecting this evaluation information, the industry recognition and trustworthiness of the target expert group are further obtained. Then, based on the collected information, appropriate evaluation methods, such as scoring or ranking, are used to assess the trustworthiness of each expert. Finally, the evaluation results are expressed in numerical or graded form, forming a trustworthiness set for the expert group.
[0091] After obtaining the trustworthiness set of the expert group, corresponding trustworthiness adjustment factors are generated based on the trustworthiness information. When generating the trustworthiness adjustment factors, firstly, certain conversion rules or algorithms are set to convert the values or levels in the expert group's trustworthiness set into corresponding trustworthiness weights. For example, different weight coefficients are set according to the trustworthiness level. Secondly, for experts with higher expertise and trustworthiness in certain specific fields or issues, their annotation information is given greater weight. The trustworthiness weights are further adjusted based on factors such as the expert's professional field, experience background, and past performance. Finally, trustworthiness adjustment factor information is generated. This information includes the trustworthiness weight and adjustment coefficient for each expert.
[0092] After obtaining the initial set of annotation information correction indices and the trust level correction factor, these two are weighted to obtain a weighted set of annotation information correction indices. During the weighted calculation, each correction index in the initial set of annotation information correction indices is first multiplied by the corresponding expert's trust level correction factor. Then, the weighted correction indices are summarized or averaged to obtain the weighted correction index for each piece of annotation information. The weighted correction index combines the importance of the corrected information with the expert's trustworthiness, and can more accurately reflect the degree to which the annotation information needs correction. Finally, a weighted set of annotation information correction indices is obtained, containing the weighted correction index for each piece of annotation information.
[0093] Furthermore, in step S4 of the method provided in the application embodiment, obtaining the power standard knowledge annotation model includes:
[0094] S45: Use a deep learning network to iteratively train the power standard knowledge text annotation sample set to obtain a basic power knowledge annotation model;
[0095] S46: The performance of the basic power knowledge annotation model is verified and evaluated using the model loss function to obtain a set of annotation model performance evaluation parameters;
[0096] S47: Based on the set of performance evaluation parameters of the labeled model, the model loss function is minimized using the gradient descent algorithm to obtain the model optimization parameters;
[0097] S48: Based on the model optimization parameters, update and configure the model parameters of the basic power knowledge annotation model to obtain the power standard knowledge annotation model.
[0098] In this embodiment, a suitable deep learning network structure, such as a convolutional neural network or a recurrent neural network, is selected and customized according to the characteristics of the power standard knowledge text and the requirements of the annotation task. Then, the annotation sample set is divided into a training set, a validation set, and a test set. The training set is used to train the model, the validation set is used to adjust model parameters and monitor the training process, and the test set is used to evaluate the model's performance. During training, the backpropagation algorithm and gradient descent optimizer are used to update the network parameters. Specifically, samples from the training set are input into the network, the network output is calculated through forward propagation, and compared with the actual annotation information to calculate the value of the loss function. Then, the gradient of the loss function with respect to the network parameters is calculated using the backpropagation algorithm, and the gradient descent optimizer is used to update the network parameters to reduce the value of the loss function. This process is repeated until a preset number of training rounds is reached or other stopping conditions are met. Through iterative training, a basic power knowledge annotation model is obtained.
[0099] The basic electricity knowledge annotation model was then tested using a validation set. The samples in the validation set were not used for training, allowing for a more objective evaluation of the model's performance. The samples from the validation set were input into the model to obtain its annotation results, which were then compared with the actual annotation information. The model's loss function was then calculated, reflecting the degree of difference between the model's annotation results and the actual annotation information. A smaller loss function value indicates better model annotation performance. In addition to the loss function value, other performance evaluation metrics, such as precision, recall, and F1 score, were calculated to more comprehensively evaluate the model's performance. Finally, a set of performance evaluation parameters for the annotation model was obtained, including the loss function value and the values of other performance evaluation metrics.
[0100] After obtaining the set of labeled model performance evaluation parameters, the gradient descent algorithm is used to minimize the model's loss function to obtain the optimized model parameters. During the minimization process, the gradient direction of the loss function under the current model parameters is first determined based on the loss function values in the performance evaluation parameter set. The gradient direction indicates the direction in which the loss function value decreases, i.e., the direction of model optimization. Then, the gradient descent algorithm is used to update the model parameters. Specifically, an appropriate learning rate is chosen, and the model parameters are updated along the gradient direction to reduce the value of the loss function. This process is repeated until the loss function value converges to a small value or other stopping conditions are met. Through the optimization process of the gradient descent algorithm, the optimized model parameters are obtained.
[0101] After obtaining the optimized model parameters, they are used to update the parameter configuration of the basic power knowledge annotation model, thus obtaining the final power standard knowledge annotation model. Specifically, the optimized parameters are loaded into the model, replacing the original model parameters. In this way, the model performs annotation tasks under the new parameter configuration. The updated model not only inherits the learning ability of the basic model but also further improves the accuracy and reliability of annotation through the adjustment of optimized parameters. The final power standard knowledge annotation model is thus obtained.
[0102] In summary, the embodiments of this application have at least the following technical effects:
[0103] This application collects a source power standard knowledge dataset, cleans and standardizes it to obtain a standard power knowledge dataset, and uses a large language model to augment the standard power knowledge dataset, generating a power standard knowledge text description dataset. It then obtains text association filtering rules and performs multiple filtering processes on the power standard knowledge text description dataset based on these rules, resulting in an augmented power standard knowledge text dataset. A power knowledge tagging system is designed according to the application needs of power knowledge, and the standard power knowledge dataset and the augmented power standard knowledge text dataset are labeled based on this system, resulting in an labeled power standard knowledge text dataset. The labeled power standard knowledge text dataset is then corrected by a power domain knowledge expert group to obtain a labeled power standard knowledge text sample set. A deep learning network is used to train and optimize the labeled power standard knowledge text sample set, resulting in a power standard knowledge labeling model for automated labeling of power standard knowledge. This invention solves the technical problem that manual labeling in existing technologies struggles to guarantee consistency and accuracy. Through an automated labeling method, it achieves rapid and accurate labeling of power standard knowledge, thereby improving labeling efficiency and quality.
[0104] Example 2
[0105] Based on the same inventive concept as the automated labeling method for electricity standard knowledge in the foregoing embodiments, such as Figure 2 As shown, this application provides an automated labeling system for electricity standard knowledge. The system and method embodiments in this application are based on the same inventive concept. The system includes:
[0106] The power standard knowledge text description dataset generation module 11 is used to collect source power standard knowledge datasets, perform data cleaning and standardization on the source power standard knowledge datasets to obtain standard power knowledge datasets, and use a large language model to perform text expansion on the standard power knowledge datasets to generate power standard knowledge text description datasets.
[0107] The power standard knowledge text extended dataset acquisition module 12 is used to acquire text association filtering rules, and perform multiple filtering processes on the power standard knowledge text description dataset based on the text association filtering rules to obtain the power standard knowledge text extended dataset.
[0108] The power standard knowledge text annotation dataset acquisition module 13 is used to design a power knowledge tag system according to the power knowledge application requirements, and to perform data annotation on the standard power knowledge dataset and the power standard knowledge text extended dataset based on the power knowledge tag system to obtain the power standard knowledge text annotation dataset.
[0109] The power standard knowledge tag automated annotation module 14 is used to correct the power standard knowledge text annotation dataset, obtain a power standard knowledge text annotation sample set, train and optimize the power standard knowledge text annotation sample set using a deep learning network, obtain a power standard knowledge annotation model, and perform automated annotation of power standard knowledge tags based on the power standard knowledge annotation model.
[0110] Furthermore, the system is also used for:
[0111] Obtain a text knowledge cleaning program, which includes non-text content removal, special symbol removal, HTML tag removal, and text format unification;
[0112] The source power standard knowledge dataset is cleaned using the text knowledge cleaning program to obtain a usable power standard knowledge dataset.
[0113] A custom dictionary of power knowledge is established, and the available power standard knowledge dataset is segmented using the custom dictionary of power knowledge to obtain a segmented power standard knowledge dataset.
[0114] A stop word list is constructed, and the stop words are removed from the power standard knowledge word segmentation dataset based on the stop word list to obtain the standard power knowledge dataset.
[0115] Furthermore, the system is also used for:
[0116] Based on the text association filtering rules, determine the content duplication filtering step, the relevance filtering step, and the semantic clarity filtering step;
[0117] Based on the content duplication screening step, a text removal algorithm is used to perform preliminary screening of duplicate content in the power standard knowledge text description dataset to obtain the first power knowledge text dataset.
[0118] Based on the relevance screening step, the keyword matching algorithm is used to perform a second screening of irrelevant content on the first power knowledge text dataset to obtain the second power knowledge text dataset.
[0119] Based on the semantic clarity screening step, a natural language processing algorithm is used to screen the second power knowledge text dataset for semantically unclear content, thereby obtaining the power standard knowledge text augmentation dataset.
[0120] Furthermore, the system is also used for:
[0121] The power knowledge application requirements are labeled with attributes to obtain a set of knowledge labeling requirements attributes, which includes labeling purpose, application scenario and users.
[0122] Based on the power operation and maintenance knowledge architecture, the classification method of the knowledge labeling requirement attribute set is analyzed to determine the knowledge label classification method.
[0123] Based on the knowledge tag classification method, a tag hierarchy is designed for the power operation and maintenance knowledge architecture to obtain a knowledge tag classification hierarchy.
[0124] Based on the knowledge tag classification method and the knowledge tag classification hierarchy, the power knowledge tag system is determined.
[0125] Furthermore, the system is also used for:
[0126] Based on the review and correction of the power standard knowledge text annotation dataset by each expert in the power field knowledge expert group, a set of expert annotation correction information is obtained.
[0127] By evaluating the modification index of each correction information in the expert annotation correction information set by each expert, an initial annotation information correction index set is obtained;
[0128] Based on the initial annotation information correction index set and the power field knowledge expert group, the weighted index calculation of each correction information is performed to determine the annotation information weighted correction index set;
[0129] The labeled data that reach the preset correction index threshold in the weighted correction index set of the labeled information are corrected to obtain the power standard knowledge text labeled sample set.
[0130] Furthermore, the system is also used for:
[0131] Based on their knowledge and experience in the power industry, the trustworthiness of each expert in the power field knowledge expert group was assessed to obtain the trustworthiness set of the expert group.
[0132] Based on the trustworthiness set of the expert group, trustworthiness correction factor information is generated;
[0133] The weighted calculation result of the initial annotation information correction index set and the trust correction factor information is used as the annotation information weighted correction index set of each correction information.
[0134] Furthermore, the system is also used for:
[0135] A basic power knowledge annotation model is obtained by iteratively training the power standard knowledge text annotation sample set using a deep learning network.
[0136] The basic power knowledge annotation model is evaluated and its performance is verified using the model loss function to obtain a set of performance evaluation parameters for the annotation model.
[0137] Based on the set of performance evaluation parameters of the labeled model, the model loss function is minimized using the gradient descent algorithm to obtain the model optimization parameters;
[0138] Based on the model optimization parameters, the basic power knowledge annotation model is updated and configured to obtain the power standard knowledge annotation model.
[0139] This application also provides a computer-readable storage medium storing computer instructions for causing a processor to execute the automated labeling method for power standard knowledge described in the foregoing embodiments.
[0140] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0141] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0142] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. An automated labeling method for electrical standard knowledge, characterized in that, The method includes: A source power standard knowledge dataset is collected, and the source power standard knowledge dataset is cleaned and standardized to obtain a standard power knowledge dataset. A large language model is used to augment the standard power knowledge dataset with text to generate a power standard knowledge text description dataset. Obtain text association filtering rules, and perform multiple filtering processes on the power standard knowledge text description dataset based on the text association filtering rules to obtain the power standard knowledge text augmentation dataset; A power knowledge tagging system is designed based on the application needs of power knowledge. Based on the power knowledge tagging system, the standard power knowledge dataset and the power standard knowledge text augmentation dataset are annotated to obtain the power standard knowledge text annotation dataset. The power standard knowledge text annotation dataset is corrected to obtain a power standard knowledge text annotation sample set. A deep learning network is used to train and optimize the power standard knowledge text annotation sample set to obtain a power standard knowledge annotation model. Power standard knowledge labels are automatically annotated based on the power standard knowledge annotation model. The obtained power standard knowledge text augmented dataset includes: Based on the text association filtering rules, determine the content duplication filtering step, the relevance filtering step, and the semantic clarity filtering step; Based on the content duplication screening step, a text removal algorithm is used to perform preliminary screening of duplicate content in the power standard knowledge text description dataset to obtain the first power knowledge text dataset. Based on the relevance screening step, the keyword matching algorithm is used to perform a second screening of irrelevant content on the first power knowledge text dataset to obtain the second power knowledge text dataset. Based on the semantic clarity screening step, a natural language processing algorithm is used to screen the second power knowledge text dataset for semantically unclear content, thereby obtaining the power standard knowledge text augmentation dataset.
2. The method as described in claim 1, characterized in that, The obtained standard power knowledge dataset includes: Obtain a text knowledge cleaning program, which includes non-text content removal, special symbol removal, HTML tag removal, and text format unification; The source power standard knowledge dataset is cleaned using the text knowledge cleaning program to obtain a usable power standard knowledge dataset. A custom dictionary of power knowledge is established, and the available power standard knowledge dataset is segmented using the custom dictionary of power knowledge to obtain a segmented power standard knowledge dataset. A stop word list is constructed, and the stop words are removed from the power standard knowledge word segmentation dataset based on the stop word list to obtain the standard power knowledge dataset.
3. The method as described in claim 1, characterized in that, The design of the power knowledge tagging system based on the application needs of power knowledge includes: The power knowledge application requirements are labeled with attributes to obtain a set of knowledge labeling requirements attributes, which includes labeling purpose, application scenario and users. Based on the power operation and maintenance knowledge architecture, the classification method of the knowledge labeling requirement attribute set is analyzed to determine the knowledge label classification method. Based on the knowledge tag classification method, a tag hierarchy is designed for the power operation and maintenance knowledge architecture to obtain a knowledge tag classification hierarchy. Based on the knowledge tag classification method and the knowledge tag classification hierarchy, the power knowledge tag system is determined.
4. The method as described in claim 1, characterized in that, The obtained power standard knowledge text annotation sample set includes: Based on the review and correction of the power standard knowledge text annotation dataset by each expert in the power field knowledge expert group, a set of expert annotation correction information is obtained. By evaluating the modification index of each correction information in the expert annotation correction information set by each expert, an initial annotation information correction index set is obtained; Based on the initial annotation information correction index set and the power field knowledge expert group, the weighted index calculation of each correction information is performed to determine the annotation information weighted correction index set; The labeled data that reach the preset correction index threshold in the weighted correction index set of the labeled information are corrected to obtain the power standard knowledge text labeled sample set.
5. The method as described in claim 4, characterized in that, The set of weighted correction indices for determining the labeled information includes: Based on their knowledge and experience in the power industry, the trustworthiness of each expert in the power field knowledge expert group was assessed to obtain the trustworthiness set of the expert group. Based on the trustworthiness set of the expert group, trustworthiness correction factor information is generated; The weighted calculation result of the initial annotation information correction index set and the trust correction factor information is used as the annotation information weighted correction index set of each correction information.
6. The method as described in claim 1, characterized in that, The obtained power standard knowledge annotation model includes: A basic power knowledge annotation model is obtained by iteratively training the power standard knowledge text annotation sample set using a deep learning network. The basic power knowledge annotation model is evaluated and its performance is verified using the model loss function to obtain a set of performance evaluation parameters for the annotation model. Based on the set of performance evaluation parameters of the labeled model, the model loss function is minimized using the gradient descent algorithm to obtain the model optimization parameters; Based on the model optimization parameters, the basic power knowledge annotation model is updated and configured to obtain the power standard knowledge annotation model.
7. An automated labeling system for electrical standard knowledge, characterized in that, The system comprises: steps for implementing the method according to any one of claims 1 to 6; The power standard knowledge text description dataset generation module is used to collect source power standard knowledge datasets, perform data cleaning and standardization on the source power standard knowledge datasets to obtain standard power knowledge datasets, and use a large language model to perform text augmentation on the standard power knowledge datasets to generate power standard knowledge text description datasets. The power standard knowledge text augmentation dataset acquisition module is used to acquire text association filtering rules, and to perform multiple filtering processes on the power standard knowledge text description dataset based on the text association filtering rules to obtain the power standard knowledge text augmentation dataset. The module for obtaining the power standard knowledge text annotation dataset is used to design a power knowledge tagging system based on the application requirements of power knowledge, and to annotate the standard power knowledge dataset and the power standard knowledge text extended dataset based on the power knowledge tagging system to obtain the power standard knowledge text annotation dataset. An automated annotation module for power standard knowledge tags is used to correct the power standard knowledge text annotation dataset through a power field knowledge expert group to obtain a power standard knowledge text annotation sample set. The power standard knowledge text annotation sample set is then trained and optimized using a deep learning network to obtain a power standard knowledge annotation model for automated annotation of power standard knowledge tags.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps in the automated labeling method for electrical standard knowledge as described in any one of claims 1-6.
Citation Information
Patent Citations
Electric power industry standard provision searching method and system based on semantic understanding
CN118210869A