Document quality inspection method and device, computer equipment and readable storage medium
Through automated document quality inspection methods, using feature extraction and quality inspection models, efficient and accurate quality inspection of documents related to power emergency incidents is solved, and the problem of low efficiency in the existing technology is solved.
Patent Information
- Application Number
- CN202510108100.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
AI Technical Summary
Document quality inspection in power emergency events is inefficient, relies on manual review and is prone to errors.
Provide a document quality inspection method, by obtaining documents to be inspected associated with power emergency events, performing feature extraction and input to trained quality inspection models, and generating quality inspection results.
Automatic document quality inspection is realized, inspection efficiency and accuracy are improved, and the time and cost of manual review are reduced.
Smart Images

Figure CN119988939A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a document quality inspection method, apparatus, computer equipment, and computer-readable storage medium. Background Art
[0002] During the operation of the power system, various emergencies are often encountered, such as natural disasters, equipment failures, etc. These events often require rapid response and effective handling. During the emergency handling process, a large number of documents will be generated, including accident reports, operation records, emergency plans, etc. The quality of these documents directly affects the effect of emergency handling and subsequent investigation and analysis.
[0003] In traditional technology, document quality inspection under power emergency events usually relies on manual review. However, manual review is time-consuming, labor-intensive and error-prone, with low efficiency and accuracy.
[0004] Therefore, there is a problem of low efficiency in the current document quality inspection technology under power emergency events. Summary of the invention
[0005] Based on this, it is necessary to provide a document quality inspection method, device, computer equipment, computer-readable storage medium and computer program product that can improve efficiency in response to the above technical problems.
[0006] In a first aspect, the present application provides a document quality inspection method, comprising:
[0007] Acquire documents to be checked that are associated with power emergency events; the documents to be checked include at least one of accident reports, operation records, and emergency plans;
[0008] Extracting features of the document to be checked to obtain document features of the document to be checked;
[0009] The document features are input into a trained quality inspection model to obtain a quality inspection result of the document to be inspected.
[0010] In one embodiment, the document features include accident report features, operation record features, and emergency plan features; the inputting of the document features into a trained quality inspection model to obtain the quality inspection results of the document to be inspected includes:
[0011] Correcting the accident report feature according to the operation record feature to obtain a corrected accident report feature;
[0012] Correcting the emergency plan characteristics according to the corrected accident report characteristics to obtain corrected emergency plan characteristics;
[0013] The corrected emergency plan features are input into the trained quality inspection model to obtain the quality inspection results of the emergency plan.
[0014] In one of the embodiments, before extracting features from the document to be checked to obtain document features of the document to be checked, the method further includes:
[0015] The document to be checked is preprocessed; the preprocessing includes at least one of text cleaning, format standardization, word segmentation, tagging, deduplication, stop word filtering, and data enhancement.
[0016] In one embodiment, extracting features from the document to be checked to obtain document features of the document to be checked includes:
[0017] Extracting keywords from the document to be checked;
[0018] Performing named entity recognition on the keywords;
[0019] If the recognition is successful, extracting the syntactic features, semantic features and content structure features of the document to be checked;
[0020] The document feature is obtained according to the syntactic feature, the semantic feature and the content structure feature.
[0021] In one embodiment, the quality inspection model includes a rule engine model and a machine learning model; the step of inputting the document features into the trained quality inspection model to obtain the quality inspection result of the document to be inspected includes:
[0022] Inputting the document features into a preset rule engine model to obtain a first inspection result of the document to be inspected;
[0023] Inputting the document features into a trained machine learning model to obtain a second inspection result of the document to be inspected;
[0024] The first inspection result and the second inspection result are combined to obtain the quality inspection result.
[0025] In one embodiment, the step of inputting the document feature into a preset rule engine model to obtain a first inspection result of the document to be inspected includes:
[0026] Inputting the document features into the rule engine model to obtain the format check result, content integrity check result and terminology consistency check result of the document to be checked;
[0027] The first check result of the document to be checked is obtained according to the format check result, the content integrity check result and the terminology consistency check result.
[0028] In a second aspect, the present application also provides a document quality inspection device, comprising:
[0029] A document acquisition module, used to acquire documents to be checked that are associated with power emergency events; the documents to be checked include at least one of accident reports, operation records, and emergency plans;
[0030] A feature extraction module, used to extract features of the document to be checked to obtain document features of the document to be checked;
[0031] The quality inspection module is used to input the document features into the trained quality inspection model to obtain the quality inspection result of the document to be inspected.
[0032] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0033] Acquire documents to be checked that are associated with power emergency events; the documents to be checked include at least one of accident reports, operation records, and emergency plans;
[0034] Extracting features of the document to be checked to obtain document features of the document to be checked;
[0035] The document features are input into a trained quality inspection model to obtain a quality inspection result of the document to be inspected.
[0036] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0037] Acquire documents to be checked that are associated with power emergency events; the documents to be checked include at least one of accident reports, operation records, and emergency plans;
[0038] Extracting features of the document to be checked to obtain document features of the document to be checked;
[0039] The document features are input into a trained quality inspection model to obtain a quality inspection result of the document to be inspected.
[0040] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:
[0041] Acquire documents to be checked that are associated with power emergency events; the documents to be checked include at least one of accident reports, operation records, and emergency plans;
[0042] Extracting features of the document to be checked to obtain document features of the document to be checked;
[0043] The document features are input into a trained quality inspection model to obtain a quality inspection result of the document to be inspected.
[0044] The above-mentioned document quality inspection method, device, computer equipment, computer-readable storage medium and computer program product obtain documents to be inspected related to power emergency events, and the documents to be inspected include at least one of accident reports, operation records, and emergency plans. Feature extraction is performed on the documents to be inspected to obtain document features of the documents to be inspected, and the document features are input into a trained quality inspection model to obtain quality inspection results of the documents to be inspected. Quality inspection can be automatically performed on documents such as accident reports, operation records, emergency plans, etc. related to power emergency events, thereby improving the efficiency of document quality inspection. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0046] Figure 1 A schematic diagram of a flow chart of a document quality inspection method in one embodiment;
[0047] Figure 2 is a flowchart of a document quality inspection method in another embodiment;
[0048] Figure 3 is a structural block diagram of a document quality inspection device in one embodiment;
[0049] Figure 4 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0051] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0052] In an exemplary embodiment, Figure 1 As shown, a document quality inspection method is provided. This embodiment takes the method applied to a terminal as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0053] Step S102, obtaining documents to be checked that are associated with the power emergency event; the documents to be checked include at least one of an accident report, an operation record, and an emergency plan.
[0054] Among them, power emergency events refer to emergency events that occur suddenly and cause or may cause casualties, damage to power equipment, large-scale power outages in the power grid, environmental damage, etc., and emergency measures need to be taken to deal with them. The documents to be checked may be documents that need to be checked for format, content integrity, terminology consistency, logical coherence, etc. Accident reports may be report documents of accidents such as casualties, damage to power equipment, large-scale power outages in the power grid, and environmental damage. Operation records may be record documents of operations on power equipment. Emergency plans may be pre-formulated response plans for power emergency events. It is understandable that the documents to be checked also include but are not limited to emergency response cards, emergency analysis reports, emergency event summary reports, etc.
[0055] In a specific implementation, documents such as accident reports, operation records, emergency plans, etc. associated with power emergency events can be input into the terminal as documents to be checked.
[0056] Step S104: extract features of the document to be checked to obtain document features of the document to be checked.
[0057] The document feature may be useful information extracted from the document to be checked.
[0058] In the specific implementation, after obtaining the document to be checked, the terminal can first perform preprocessing such as text cleaning, format standardization, word segmentation, annotation, deduplication, stop word filtering, data enhancement, etc. on the document to be checked, and then perform feature extraction on the preprocessed document to be checked to obtain document features.
[0059] During the feature extraction process, the terminal can first extract keywords from the preprocessed document to be inspected, and then perform named entity recognition on the keywords to identify whether the keywords contain the proper nouns of the corresponding document. For example, whether the keywords extracted from the operation record contain the equipment name, operating procedure number, time and date, etc. If they do, syntactic analysis, semantic analysis, content structure analysis, etc. can be continued to extract the syntactic features, semantic features, content structure features, etc. of the document to be inspected. Otherwise, if the keywords do not contain the proper nouns of the corresponding document, it means that the document to be inspected is incorrect, and it is determined that the quality inspection of the document to be inspected has failed.
[0060] Step S106, inputting the document features into the trained quality inspection model to obtain the quality inspection result of the document to be inspected.
[0061] The quality inspection model may be a pre-trained model, which may be implemented through a rule engine, machine learning, deep learning, or multimodal fusion. The quality inspection result may be a result of whether the quality inspection of the document to be inspected has passed or failed.
[0062] In a specific implementation, a quality inspection model can be pre-trained, and the extracted document features can be input into the quality inspection model. The quality inspection model automatically performs quality inspection on the document to be inspected based on the document features to obtain corresponding quality inspection results.
[0063] The above-mentioned document quality inspection method obtains the documents to be inspected related to the power emergency event, and the documents to be inspected include at least one of accident reports, operation records, and emergency plans. Feature extraction is performed on the documents to be inspected to obtain document features of the documents to be inspected, and the document features are input into a trained quality inspection model to obtain quality inspection results of the documents to be inspected. Quality inspection can be automatically performed on documents such as accident reports, operation records, emergency plans, etc. related to power emergency events, thereby improving the efficiency of document quality inspection.
[0064] In an exemplary embodiment, the document features include accident report features, operation record features and emergency plan features; the above step S106 may specifically include: correcting the accident report features according to the operation record features to obtain corrected accident report features; correcting the emergency plan features according to the corrected accident report features to obtain corrected emergency plan features; inputting the corrected emergency plan features into the trained quality inspection model to obtain the quality inspection results of the emergency plan.
[0065] The accident report feature may be a feature extracted from the accident report, the operation record feature may be a feature extracted from the operation record, and the emergency plan feature may be a feature extracted from the emergency plan.
[0066] In the specific implementation, the terminal can extract accident report features from the accident report, extract operation record features from the operation record, extract emergency plan features from the emergency plan, use the operation record features to correct the accident report features, and use the corrected accident report features to correct the emergency plan features. Finally, the corrected emergency plan features are input into the trained quality inspection model to obtain the quality inspection results of the emergency plan.
[0067] In practical applications, the correlation between the operation record features and the accident report features can be identified based on a pre-trained model, and when the correlation is lower than a preset threshold, the accident report features are corrected until the correlation between the operation record features and the accident report features is no lower than the preset threshold. Alternatively, the accident report features corresponding to the operation record features can be predicted based on a pre-trained model, and the predicted accident report features can be used to correct the accident report features obtained by feature extraction.
[0068] Similarly, the correlation between the accident report feature and the emergency plan feature can also be identified based on the pre-trained model, and the emergency plan feature can be corrected when the correlation is lower than the preset threshold, until the correlation between the accident report feature and the emergency plan feature is not lower than the preset threshold. Alternatively, the emergency plan feature corresponding to the accident report feature can be predicted based on the pre-trained model, and the predicted emergency plan feature can be used to correct the emergency plan feature obtained by feature extraction.
[0069] It should be noted that the above-mentioned pre-trained models include but are not limited to machine learning, deep learning, neural network and other models.
[0070] It is understandable that the accident report can also be corrected according to the operation record to obtain a corrected accident report, the emergency plan can be corrected according to the corrected accident report to obtain a corrected emergency plan, features can be extracted from the corrected emergency plan to obtain emergency plan features, the emergency plan features can be input into the trained quality inspection model to obtain the quality inspection results of the emergency plan, and the correction method can be the same as the correction method of the accident report features and the emergency plan features.
[0071] In this embodiment, the accident report characteristics are corrected according to the operation record characteristics to obtain the corrected accident report characteristics, the emergency plan characteristics are corrected according to the corrected accident report characteristics to obtain the corrected emergency plan characteristics, and the corrected emergency plan characteristics are input into the trained quality inspection model to obtain the quality inspection results of the emergency plan. The emergency plan can be corrected according to the correlation between the accident report, operation record and emergency plan to improve the accuracy of the emergency plan and ensure the reliability of the power emergency response.
[0072] In an exemplary embodiment, before the above step S104, it can also specifically include: preprocessing the document to be checked; the preprocessing includes at least one of text cleaning, format standardization, word segmentation, tagging, deduplication, stop word filtering, and data enhancement.
[0073] In specific implementation, the terminal can perform text cleaning on the document to be inspected, including deleting non-essential content such as headers, footers, page numbers, watermarks, comments, etc. in the document to ensure that the retained text information is directly related to the power emergency event; it also includes repairing spelling errors, punctuation errors, garbled characters and other problems in the document to ensure the readability and accuracy of the text; and converting the document to a unified character encoding format (such as UTF-8) to avoid parsing errors caused by inconsistent encoding.
[0074] The terminal can also standardize the format of the documents to be inspected, including converting documents of different formats (such as PDF, Word, TXT, etc.) into a unified format (such as plain text or HTML) for subsequent processing; it also includes adjusting the document's layout format, such as unifying fonts, font sizes, line spacing, paragraph indents, etc., to ensure the visual consistency and standardization of the document; and organizing the document content into structured forms such as chapters, paragraphs, lists, etc., to facilitate subsequent feature extraction and analysis.
[0075] The terminal can also perform word segmentation and annotation on the documents to be inspected, including using word segmentation tools (such as jieba Chinese word segmentation, natural language processing toolkit NLTK, etc.) to divide the text in the document into words to ensure that each word can be processed separately; it also includes part-of-speech tagging of the words after word segmentation, such as nouns, verbs, adjectives, etc., to provide a basis for subsequent semantic analysis and named entity recognition; as well as identifying and annotating proper nouns in the document, such as equipment names, operating procedure numbers, time and date, etc., to ensure the accurate extraction of these key information.
[0076] The terminal can also perform deduplication processing on the documents to be checked, remove duplicate content in the documents, and avoid redundant information interfering with subsequent analysis.
[0077] The terminal can also perform stop-word filtering on the document to be inspected, removing common stop words (such as "de", "shi", "zai", etc.) to reduce noise and improve processing efficiency.
[0078] The terminal can also perform data augmentation on the document to be inspected. For some documents with less key information, data augmentation techniques (such as synonym replacement, sentence restructuring, etc.) are used to increase the diversity of training data and improve the generalization ability of the model.
[0079] In this embodiment, by preprocessing the document to be inspected, the document to be inspected can be converted into a processable format, facilitating automatic processing of the document to be inspected by a computer and improving the efficiency of document quality inspection.
[0080] In an exemplary embodiment, step S104 above may specifically include: extracting keywords from the document to be inspected; performing named entity recognition on the keywords; in the case of successful recognition, extracting syntactic features, semantic features, and content structure features of the document to be inspected; and obtaining document features based on the syntactic features, semantic features, and content structure features.
[0081] In specific implementation, the terminal can extract keywords from the document to be inspected and perform named entity recognition on the extracted keywords. If the recognition is successful, syntactic analysis, semantic analysis, and content structure analysis are sequentially performed on the document to be inspected to obtain syntactic features, semantic features, and content structure features respectively, and the syntactic features, semantic features, and content structure features are used as the document features of the document to be inspected. Otherwise, if the recognition fails, it indicates that the document to be inspected is incorrect, and it is determined that the quality inspection of the document to be inspected fails.
[0082] In practical applications, after preprocessing the document to be inspected, the Term Frequency-Inverse Document Frequency (TF-IDF) method can be used to extract features from it, and the TextRank method can also be used to sort the extracted keywords. Specific entity types are defined according to the characteristics of the power industry, and a custom Named Entity Recognition (NER) model is trained. The extracted keywords are input into the trained NER model to identify proper nouns in the document, such as equipment names, operating procedure numbers, dates, etc. If the proper nouns identified in the document to be inspected conform to the predefined entity types, syntactic analysis, semantic analysis, and content structure analysis are sequentially performed on the document to be inspected to obtain syntactic features, semantic features, and content structure features of the document to be inspected.
[0083] In this embodiment, by extracting keywords from the document to be checked, named entity recognition is performed on the keywords, and if the recognition is successful, the syntactic features, semantic features, and content structure features of the document to be checked are extracted; based on the syntactic features, semantic features, and content structure features, document features are obtained, and multiple features can be extracted from the document to be checked, thereby increasing the accuracy and reliability of document quality inspection.
[0084] In an exemplary embodiment, the quality inspection model includes a rule engine model and a machine learning model; the above step S106 may specifically include: inputting document features into a preset rule engine model to obtain a first inspection result of the document to be inspected; inputting document features into a trained machine learning model to obtain a second inspection result of the document to be inspected; and integrating the first inspection result with the second inspection result to obtain a quality inspection result.
[0085] The rule engine model may be a series of rules defined according to the standards and specifications of the power industry, and used to check the format, content integrity, terminology consistency, etc. of the document. The first check result may be the result obtained by checking according to the standards and specifications of the power industry. The second check result may be the result identified by machine learning.
[0086] In the specific implementation, a rule engine model can be built in advance, and a machine learning model can be pre-trained. The terminal inputs the document features into the rule engine model to obtain a first inspection result of the document to be inspected. The terminal can also input the document features into the trained machine learning model to obtain a second inspection result of the document to be inspected. The first inspection result is then integrated with the second inspection result to obtain the final quality inspection result.
[0087] In practical applications, a series of rules can be defined according to the standards and specifications of the power industry as a rule engine to check the format, content integrity, terminology consistency, etc. of the document. Specifically, the font, font size, paragraph spacing, list format and other standards of the document can be defined to ensure that the format of the document meets the requirements; the chapters and content that the document must contain can also be defined, such as event description, cause analysis, treatment measures, follow-up plans, etc., to ensure the content integrity of the document; standard expressions of commonly used terms in the power industry can also be defined to ensure the consistency and accuracy of the terms in the document; rule matching algorithms can also be developed to compare the pre-processed documents with the defined rules to check whether the documents meet the requirements of the rules; in addition, regular expressions can be used to match specific formats and content, and string matching algorithms can be used to check the consistency of terms; finally, the execution process of the rule engine is designed to ensure that each rule can be correctly applied and the first inspection result is generated.
[0088] Select appropriate machine learning models, such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Decision Tree (GBDT), etc. In addition, you can also select deep learning models suitable for text processing, such as Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), Transformer, etc., to identify document features and obtain the second inspection result. Then, you can combine the first inspection result and the second inspection result to obtain the final document quality. For example, if either the first inspection result or the second inspection result fails, the final quality inspection result is failed.
[0089] In this embodiment, by inputting document features into a pre-set rule engine model to obtain a first inspection result of the document to be inspected, the document features are input into a trained machine learning model to obtain a second inspection result of the document to be inspected, and the first inspection result is integrated with the second inspection result to obtain a quality inspection result, thereby improving the accuracy of the document quality inspection results.
[0090] In an exemplary embodiment, the above-mentioned step of inputting document features into a preset rule engine model to obtain a first inspection result of the document to be inspected may specifically include: inputting document features into a rule engine model to obtain a format check result, a content integrity check result and a terminology consistency check result of the document to be inspected; and obtaining a first inspection result of the document to be inspected based on the format check result, the content integrity check result and the terminology consistency check result.
[0091] The format check result refers to whether the document's font, font size, paragraph spacing, list format, etc. meet the requirements. The content integrity check result refers to whether the document contains the necessary chapters and content. The terminology consistency check result refers to whether the terminology in the document is consistent with the standard expressions of commonly used terminology in the power industry.
[0092] In a specific implementation, the terminal can define the format rules, content integrity rules and terminology consistency rules of the document to be checked through a rule engine model, input the document features into the rule engine model, check the document to be checked through the format rules to obtain a format check result, check the document to be checked through the content integrity rules to obtain a content integrity check result, check the document to be checked through the terminology consistency rules to obtain a terminology consistency check result, and use the format check result, content integrity check result and terminology consistency check result as the first check result of the document to be checked.
[0093] In this embodiment, by inputting document features into the rule engine model, the format check results, content integrity check results and terminology consistency check results of the document to be checked are obtained. Based on the format check results, content integrity check results and terminology consistency check results, the first inspection result of the document to be checked is obtained, and whether the document complies with the standards and specifications of the electric power industry can be checked, thereby improving the accuracy of document quality inspection.
[0094] In order to facilitate those skilled in the art to have a deeper understanding of the embodiments of the present application, a specific example will be described below.
[0095] This application combines Natural Language Processing (NLP) technology and professional knowledge of the power industry, aiming to use automated means to perform quality inspection and processing on a large number of documents generated in power emergency events. This method not only improves the efficiency and accuracy of document processing, but also provides strong support for emergency management and decision-making of power companies. Specifically, this application deeply integrates NLP technology with power emergency event management, and realizes efficient and accurate quality inspection of a large number of emergency documents in an automated manner. Compared with traditional manual review, this solution not only significantly improves the speed and accuracy of document processing, but also realizes intelligent analysis and evaluation of document content by introducing machine learning and deep learning models, thereby effectively improving the decision-making efficiency and management level of power companies in emergency response.
[0096] In one embodiment, the present application provides a method for quality inspection and processing of documents related to power emergency events based on NLP technology. The method uses automated means to conduct a comprehensive inspection of the format, content integrity, terminology consistency, and logical coherence of a large number of documents generated in power emergency events, and uses machine learning and deep learning models to improve the accuracy and efficiency of the inspection, and finally generates a detailed inspection report and improvement suggestions, significantly improving the document processing capabilities and decision-making level of power companies in emergency response. The method specifically includes the following steps:
[0097] Step S201: document preprocessing
[0098] (1) Text cleaning
[0099] Remove irrelevant information: Delete non-essential content such as headers, footers, page numbers, watermarks, comments, etc. in the document to ensure that the retained text information is directly related to the power emergency event.
[0100] Correct formatting errors: Fix spelling errors, punctuation errors, garbled characters, etc. in the document to ensure the readability and accuracy of the text.
[0101] Unified encoding: Convert the document into a unified character encoding format (such as UTF-8) to avoid parsing errors caused by inconsistent encodings.
[0102] (2) Format standardization
[0103] Unify document format: Convert documents in different formats (such as PDF, Word, TXT, etc.) into a unified format (such as plain text or HTML) for easier subsequent processing.
[0104] Normalize typesetting: Adjust the typesetting format of the document, such as unifying the font, font size, line spacing, paragraph indentation, etc., to ensure the visual consistency and standardization of the document.
[0105] Structured processing: Organize the content of the document in a structured form such as chapters, paragraphs, lists, etc., for easier subsequent feature extraction and analysis.
[0106] (3) Word segmentation and annotation
[0107] Word segmentation processing: Use word segmentation tools (such as jieba, NLTK, etc.) to segment the text in the document into words to ensure that each word can be processed separately.
[0108] Part-of-speech tagging: Tag the words after word segmentation with part-of-speech tags such as nouns, verbs, adjectives, etc., to provide a basis for subsequent semantic analysis and named entity recognition.
[0109] Named entity recognition: Identify and tag the proper nouns in the document, such as equipment names, operating procedure numbers, dates, etc., to ensure the accurate extraction of these key information.
[0110] (4) Data preprocessing
[0111] Duplicate removal processing: Remove the duplicate content in the document to avoid redundant information interfering with subsequent analysis.
[0112] Stop word filtering: Remove common stop words (such as "de", "shi", "zai", etc.) to reduce noise and improve processing efficiency.
[0113] Data augmentation: For some documents with less key information, data augmentation techniques (such as synonym replacement, sentence restructuring, etc.) can be used to increase the diversity of training data and improve the generalization ability of the model.
[0114] Through the above preprocessing steps, it can be ensured that the document data input into the subsequent model is clean, standardized, and structured, laying a solid foundation for high-quality document inspection and processing.
[0115] Step S202, Feature extraction
[0116] Feature extraction is a key step in document quality inspection. It provides a basis for subsequent quality assessment and inspection by extracting useful information from preprocessed documents. The following are the detailed feature extraction steps:
[0117] (1) Keyword extraction
[0118] TF-IDF algorithm: Use the TF-IDF algorithm to extract keywords from documents. The higher the TF-IDF value, the more important the word is in the document.
[0119] Calculate term frequency (TF): count the number of times each word appears in the document.
[0120] Calculate Inverse Document Frequency (IDF): A measure of the general importance of a term, usually by taking the logarithm of the inverse of the number of documents containing the term.
[0121] Calculate the TF-IDF value: Multiply TF and IDF to get the TF-IDF value of each word.
[0122] TextRank algorithm: Use the TextRank algorithm to extract keywords. This is a graph-based ranking model, similar to the PageRank algorithm, which evaluates the importance of words through the relationship between words.
[0123] Specifically, TF-IDF and TextRank can be combined. First, TF-IDF is used to preliminarily screen out candidate keywords, and then TextRank is used to perform secondary sorting on the candidate keywords.
[0124] (2) Named Entity Recognition
[0125] Pre-trained model: Use pre-trained NER models (such as BERT-NER, Spacy, etc.) to identify proper nouns in documents, such as equipment names, operating procedure numbers, time and date, etc.
[0126] Customized entity recognition: Based on the characteristics of the power industry, we define specific entity types and train customized NER models to more accurately identify industry-related proper nouns.
[0127] Context analysis: Combine context information to improve the accuracy of named entity recognition and avoid misidentification.
[0128] It should be noted that named entity recognition in the preprocessing stage is mainly used for text preprocessing and preliminary understanding. By identifying proper nouns, it can ensure the consistent expression of these entities in the entire document and avoid understanding deviations caused by different expressions. Named entity recognition in the feature extraction stage is mainly used to verify whether the document contains all necessary information. For example, the power emergency response document should include the name of the equipment involved, the operating procedure number, etc. The lack of this information may cause the document to be incomplete or invalid.
[0129] (3) Syntactic analysis
[0130] Dependency syntactic analysis: Use dependency syntactic analysis tools (such as Stanford Parser, Spacy, etc.) to analyze the dependency relationship between words in a sentence and build a dependency tree to help understand the structure and logical relationship of the sentence.
[0131] Component Syntax Analysis: Use component syntax analysis tools to break down sentences into different grammatical components (such as subject, predicate, object, etc.) and check the grammatical correctness and logical coherence of the sentences.
[0132] (4) Semantic analysis
[0133] Word vector representation: Use word embedding technology (such as Word2Vec, GloVe, etc.) to convert words into high-dimensional vectors to capture the semantic relationship between words.
[0134] Sentence vector representation: Aggregate the word vectors in the sentence (such as average, weighted average, etc.) to generate a vector representation of the sentence for subsequent semantic similarity calculation.
[0135] Semantic similarity calculation: Use methods such as cosine similarity to calculate the semantic similarity between different sentences in the document and detect duplicate content and logical inconsistencies.
[0136] (5) Content structure analysis
[0137] Paragraph division: Divide the document into multiple paragraphs, and treat each paragraph as an independent unit to facilitate subsequent analysis.
[0138] Title recognition: Identify titles and subtitles in a document, build a hierarchical structure for the document, and help understand the overall framework of the document.
[0139] Logical structure analysis: Analyze the logical structure of the document, such as event description, cause analysis, treatment measures, follow-up plans, etc., to ensure the integrity and logic of the document content.
[0140] Through the above feature extraction steps, the content and structure of the document can be comprehensively analyzed from multiple perspectives, providing rich information support for subsequent quality inspection and evaluation. These features will be used in subsequent model building and quality inspection steps to ensure the quality and compliance of the document.
[0141] Step S203: Model construction
[0142] Model building is the core part of this application. By building a rule engine and machine learning / deep learning models, we can achieve automated quality inspection of documents related to power emergency events. The following are the detailed model building steps:
[0143] (1) Rule engine construction
[0144] Rule definition: Based on the standards and specifications of the power industry, define a series of rules to check the format, content integrity, terminology consistency, etc. of the document.
[0145] Formatting rules: define the document's font, font size, paragraph spacing, list format and other standards to ensure that the document's format meets the requirements.
[0146] Content integrity rules: define the sections and content that the document must include, such as event description, cause analysis, treatment measures, follow-up plan, etc.
[0147] Terminology consistency rules: define the standard expressions of commonly used terms in the power industry. The terms in the specification documents must adopt unified standard expressions to ensure the consistency and accuracy of the terms in the documents.
[0148] For example, first, the equipment name in the document must be consistent with the ledger type of the equipment center, such as substation, main transformer, distribution transformer, etc.; secondly, the unit name in the document must be consistent with the unit organizational structure; finally, documents related to power emergencies must include designated chapters. For example, the wind and flood emergency plan must include general principles, purpose of preparation, basis for preparation, scope of application, working principles, relationship with other plans, risk and resource analysis, risk analysis, resource analysis, equipment accident classification, organizational structure and responsibilities, emergency command organization, responsibilities of the emergency command organization, emergency response, response classification and disposal subject, information reporting, emergency disposal, emergency response changes, emergency end and other chapters.
[0149] Rule matching: Develop a rule matching algorithm to compare the preprocessed documents with the defined rules to check whether the documents meet the rule requirements.
[0150] Regular expressions: Use regular expressions to match specific patterns and content.
[0151] String Matching: Check the consistency of terms using a string matching algorithm.
[0152] Rule execution: Design the execution process of the rule engine to ensure that each rule is applied correctly and generate detailed inspection results.
[0153] (2) Machine Learning Model Construction
[0154] Data preparation: Use preprocessed documents and manually annotated labels (such as quality ratings, question types, etc.) as training data.
[0155] Data cleaning: Remove outliers and noise to ensure data accuracy and consistency.
[0156] Data segmentation: Divide the dataset into training set, validation set, and test set for model training, tuning, and evaluation.
[0157] Feature selection: Select the most relevant features from the features extracted in the preprocessing step for model training.
[0158] Feature Importance Evaluation: Use feature selection algorithms such as recursive feature elimination, random forest feature importance, etc. to evaluate the importance of features.
[0159] Feature dimensionality reduction: Use dimensionality reduction techniques (such as principal component analysis (PCA), linear discriminant analysis (LDA), etc.) to reduce feature dimensions and improve model training efficiency.
[0160] Model selection: Choose an appropriate machine learning model, such as support vector machine (SVM), random forest, gradient boosted tree (GBDT), etc.
[0161] Model training: Use the training set to train the selected model and optimize the model parameters.
[0162] Model tuning: Use the validation set to tune the model and select the best hyperparameter combination.
[0163] Model evaluation: Use the test set to evaluate the performance of the model, such as accuracy, recall, F1 score, etc.
[0164] Model ensemble: Combine multiple models through ensemble learning methods (such as voting and stacking) to improve the overall prediction performance.
[0165] (3) Deep learning model construction
[0166] Data preparation: Similar to the machine learning model, the preprocessed documents and manually annotated labels are used as training data.
[0167] Model selection: Select a deep learning model suitable for text processing, such as convolutional neural network (CNN), long short-term memory network (LSTM), Transformer, etc.
[0168] Sequence model: Use sequence models such as LSTM or recurrent neural network GRU to process the text sequence of documents and capture long-term dependencies.
[0169] Attention Mechanism: Use attention mechanisms (such as Transformer) to highlight key information in the document and improve the interpretability of the model.
[0170] Model training: Use the training set to train the selected deep learning model and optimize the model parameters.
[0171] Loss function: Select a suitable loss function (such as cross entropy loss, mean square error loss, etc.) to guide the learning process of the model.
[0172] Optimization algorithm: Use optimization algorithms (such as adaptive moment estimation Adam, stochastic gradient descent SGD, etc.) to update model parameters and speed up convergence.
[0173] Model tuning: Use the validation set to tune the model and select the best hyperparameter combination.
[0174] Learning rate adjustment: Dynamically adjust the learning rate to avoid overfitting and underfitting.
[0175] Regularization: Use regularization techniques (such as L1, L2 regularization) to prevent overfitting.
[0176] Model evaluation: Use the test set to evaluate the performance of the model, such as accuracy, recall, F1 score, etc.
[0177] (4) Model fusion
[0178] Multimodal fusion: Combine the advantages of rule engines, machine learning models, and deep learning models to build a multimodal fusion model.
[0179] Rule engine + machine learning: The inspection results of the rule engine are input into the machine learning model as features to improve the robustness of the model.
[0180] Machine learning + deep learning: The output of the machine learning model is input into the deep learning model as features to improve the prediction accuracy of the model.
[0181] Ensemble learning: Combine multiple models through ensemble learning methods (such as voting method, stacking method, etc.) to further improve the overall prediction performance.
[0182] Through the above model building steps, an efficient and accurate document quality inspection system can be built to ensure the quality and compliance of documents related to power emergency events. These models will be used in subsequent quality inspection steps to generate detailed inspection reports and improvement suggestions.
[0183] Step S204: Quality inspection
[0184] Quality inspection is an important part of this application. By applying the built rule engine and machine learning / deep learning model, a comprehensive quality assessment and inspection of documents related to power emergency events is carried out. The following are the details:
[0185] (1) Format check
[0186] Font and size check: Verify whether the font and size in the document meet predefined standards, such as using bold 14-point font for titles and Song 12-point font for body text.
[0187] Paragraph format check: Check whether the paragraph indentation, line spacing, alignment, etc. meet the requirements to ensure the visual consistency of the document.
[0188] List and table check: verify whether the numbering and symbols of the list items are correct, and whether the format, borders, and alignment of the table are standardized.
[0189] Header and footer check: confirm whether the content of the header and footer meets the standards, such as page number position, company logo, etc.
[0190] (2) Content integrity check
[0191] Chapter completeness: Check whether the document contains all necessary chapters, such as event description, cause analysis, treatment measures, follow-up plan, etc.
[0192] Content completeness: Ensure that the content of each chapter is complete and no key information is omitted. For example, the event description section should include basic information such as time, location, and people involved.
[0193] Terminology consistency: Check whether the terminology used in the document is consistent to avoid using different expressions for the same concept in different places.
[0194] (3) Terminology consistency check
[0195] Termbase comparison: Compare the terms in the document with the predefined termbase to ensure the accuracy and consistency of the terms.
[0196] Context analysis: Combine context information to check whether the terminology is used appropriately and avoid ambiguity.
[0197] Proper noun recognition: Use named entity recognition technology to identify and verify proper nouns in documents, such as equipment names, operating procedure numbers, etc.
[0198] (4) Logical consistency check
[0199] Grammar check: Use grammar check tools (such as LanguageTool, Grammarly, etc.) to check for grammatical errors in the document and ensure that the sentence structure is correct.
[0200] Logical relationship check: Use dependency syntax analysis and component syntax analysis to check the logical relationship between sentences and ensure the logical coherence of the document content.
[0201] Consistency check: Verify the consistency of the previous and subsequent contents in the document, such as whether the event description matches the handling measures, and whether the cause analysis and subsequent plans are reasonable.
[0202] (5) Problem detection and marking
[0203] Problem classification: Classify detected problems into categories such as format problems, content problems, terminology problems, logic problems, etc.
[0204] Problem marking: Mark the specific problem location and type in the document to facilitate subsequent corrections.
[0205] Severity assessment: Assess the severity of the problem based on its nature and scope of impact, such as minor, moderate, severe, etc.
[0206] (6) Generate inspection report
[0207] Report content: Generate a detailed inspection report, including basic information of the document, inspection results, problem list, improvement suggestions, etc.
[0208] Visual display: The quality score and problem distribution of the document are intuitively displayed in the form of charts, etc., so that users can quickly understand the overall situation of the document.
[0209] Improvement suggestions: Provide specific improvement suggestions for detected problems to help users quickly correct documents.
[0210] (7) Automatic correction suggestions
[0211] Automatic correction: For some common formatting problems and grammatical errors, an automatic correction function is provided to repair problems in the document with one click.
[0212] Manual correction tips: For issues that require manual intervention, detailed correction tips are provided to guide users to perform manual corrections.
[0213] Through the above quality inspection steps, the quality of documents related to power emergency events can be comprehensively evaluated to ensure that the format of the documents is standardized, the content is complete, the terminology is consistent, and the logic is coherent. These inspection results will provide a reliable basis for subsequent decision-making and management, and improve the efficiency and level of power companies in emergency response.
[0214] Step S205: Result output and optimization
[0215] Result output and optimization is the last step of this application. By generating detailed inspection reports and improvement suggestions, and continuously optimizing the model based on user feedback, the accuracy and reliability of the system are ensured. The following are the details:
[0216] (1) Generate inspection report
[0217] Report Contents:
[0218] Basic document information: including the document title, author, creation date, version number, etc.
[0219] Overall score: Give the document an overall quality score, such as a full score of 100.
[0220] Problem Summary: Lists all detected problems, sorted by category (formatting issues, content issues, terminology issues, logic issues, etc.).
[0221] Problem Details: Provide a detailed description of each problem, including its location, type, severity, etc.
[0222] Improvement suggestions: Provide specific improvement suggestions for each issue to help users quickly correct the document.
[0223] Report Format:
[0224] Text Report: Generate detailed text reports for easy reading and archiving.
[0225] HTML report: Generates an interactive HTML report that allows you to jump directly to the problem location by clicking a link.
[0226] PDF Report: Generate reports in PDF format for easy printing and sharing.
[0227] Visual display:
[0228] Graphical display: Use bar charts, pie charts and other charts to display the distribution of problems, such as the percentage of each type of problem.
[0229] Heatmap: Generate a heatmap of the document to visually display the dense areas of problems.
[0230] Rating trend: If the same document is checked multiple times, a rating trend chart is generated to show the changes in document quality.
[0231] (2) Output inspection results
[0232] File Export:
[0233] Annotated file: Generates a copy of the document with problem marks, and users can make modifications directly in the annotated file.
[0234] CSV / Excel file: Export the list of issues and improvement suggestions for users to view and manage in a spreadsheet.
[0235] View online:
[0236] Web interface: Provides a web interface for viewing reports online. Users can view and download reports in a browser.
[0237] Application Programming API interface: Provides an API interface to facilitate integration with other systems and call inspection results.
[0238] (3) User feedback collection
[0239] Feedback channels:
[0240] Online form: A feedback form is provided on the web interface where users can fill in their opinions and suggestions on the inspection results.
[0241] Email feedback: Users can send feedback via email, and the system will automatically collect and organize the feedback information.
[0242] Telephone support: Provide telephone support, users can directly contact technical support staff to report problems.
[0243] Feedback content:
[0244] Problem accuracy: User feedback checks the accuracy of the results, pointing out issues with false positives and false negatives.
[0245] User experience: Users provide feedback on their experience during use and make suggestions for improvements.
[0246] New requirements: Users put forward new requirements and feature suggestions, such as adding new inspection items, optimizing report formats, etc.
[0247] (4) Model optimization
[0248] Data update:
[0249] New data: Based on user feedback, new document samples and annotated data are collected to enrich the training data set.
[0250] Data cleaning: Clean and preprocess the new data to ensure data quality.
[0251] Model tuning:
[0252] Parameter adjustment: Based on user feedback, adjust the model’s hyperparameters to optimize model performance.
[0253] Rule update: Update the rules in the rule engine, add new rules or modify existing rules to improve the coverage and accuracy of the rules.
[0254] Model retraining: Retrain the model using the updated dataset to improve the model’s generalization ability and accuracy.
[0255] Performance evaluation:
[0256] Internal testing: Test the optimized model in an internal testing environment to evaluate the model's performance indicators, such as accuracy, recall, F1 score, etc.
[0257] User testing: Invite some users to participate in the test, collect user feedback, and further optimize the model.
[0258] (5)Continuous iteration
[0259] Version management:
[0260] Version record: record the version number and changes after each optimization to facilitate tracking and backtracking.
[0261] Version release: New versions are released regularly to notify users to update the system.
[0262] User Training:
[0263] Training materials: Provide user training materials, including manuals, video tutorials, etc., to help users better use the system.
[0264] Training courses: Regularly hold online or offline training courses to answer users' questions and provide technical support.
[0265] Through the above steps, it is possible to ensure the continuous optimization and improvement of the document quality inspection system, continuously improve the accuracy of the system and user experience, and provide strong support for the document management of power companies in emergency response.
[0266] The above method is highly efficient and uses NLP technology to achieve automated quality inspection of documents related to power emergency events, significantly reducing the time and cost of manual review. Automated processing enables the system to complete the inspection of a large number of documents in a short period of time, greatly improving the efficiency of emergency response and ensuring that decisions can be made quickly in emergency situations.
[0267] The method also has accuracy, combining the rule engine and machine learning / deep learning models to improve the accuracy of document quality inspection through a multimodal fusion approach. The rule engine ensures the consistency of format and terminology, while the machine learning and deep learning models can identify complex semantic and logical problems to ensure the accuracy and completeness of the document content.
[0268] In addition, it is comprehensive, covering multiple aspects of the document, including format check, content integrity check, terminology consistency check and logical coherence check. Through multi-level and multi-angle checks, it ensures that the document complies with standards and specifications in all aspects, providing comprehensive support for emergency management and decision-making of power companies.
[0269] Finally, this method also has good scalability and continuous optimization capabilities. The system can continuously update the rule engine and training model based on user feedback and new requirements to improve the performance and adaptability of the system. In addition, through the continuous accumulation of data and iterative optimization of the model, the system can continuously improve the accuracy and efficiency of inspections to meet changing business needs.
[0270] This application achieves efficient and accurate quality inspection of documents related to power emergency events by combining NLP technology, rule engine and machine learning / deep learning models. Specific technical effects include: first, automated processing significantly reduces the time and cost of manual review and improves the efficiency of emergency response; second, the multimodal fusion method ensures the high quality of documents in terms of format, content integrity, terminology consistency and logical coherence, and improves the reliability and compliance of documents; third, the system can generate detailed inspection reports and improvement suggestions to help users quickly locate and correct problems, and improve the level of refinement of document management; finally, through continuous data accumulation and model optimization, the system has good scalability and adaptability, and can continuously adapt to new business needs and standard changes to ensure long-term efficient operation. These technical effects have jointly improved the decision-making efficiency and management level of power companies in emergency response.
[0271] In one embodiment, Figure 2 As shown, a document quality inspection method is provided, which is described by taking the method applied to a terminal as an example, and includes the following steps:
[0272] Step S301, obtaining documents to be checked that are associated with power emergency events; the documents to be checked include at least one of accident reports, operation records, and emergency plans;
[0273] Step S302, preprocessing the document to be checked; the preprocessing includes at least one of text cleaning, format standardization, word segmentation, tagging, deduplication, stop word filtering, and data enhancement;
[0274] Step S303, extracting keywords from the preprocessed document to be checked, performing named entity recognition on the keywords, and if the recognition passes, extracting syntactic features, semantic features, and content structure features of the preprocessed document to be checked, and obtaining document features based on the syntactic features, semantic features, and content structure features;
[0275] Step S304, input the document features into a preset rule engine model and a trained machine learning model respectively, obtain a first inspection result and a second inspection result of the document to be inspected, and merge the first inspection result with the second inspection result to obtain a quality inspection result.
[0276] In the specific implementation, accident reports, operation records, emergency plans and other documents to be inspected can be input into the terminal. The terminal performs preprocessing on the documents to be inspected, such as text cleaning, format standardization, word segmentation, annotation, duplication removal, stop word filtering, data enhancement, etc., and performs keyword extraction on the preprocessed documents to be inspected. Named entity recognition is performed on the extracted keywords. If the recognition passes, the syntactic features, semantic features and content structure features of the preprocessed documents to be inspected are extracted, and the syntactic features, semantic features and content structure features are respectively input into the rule engine model and the machine learning model to obtain the first inspection result and the second inspection result. The first inspection result and the second inspection result are combined to obtain the final quality inspection result of the document to be inspected.
[0277] The above-mentioned document quality inspection method obtains the documents to be inspected related to the power emergency event, preprocesses and extracts features of the documents to be inspected, and then inputs them into the rule engine model and the machine learning model respectively, and integrates the first inspection result with the second inspection result to obtain the quality inspection result; it can automatically perform quality inspection on documents such as accident reports, operation records, emergency plans, etc. related to the power emergency event, thereby improving the efficiency of document quality inspection.
[0278] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0279] Based on the same inventive concept, the embodiment of the present application also provides a document quality inspection device for implementing the document quality inspection method involved above. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above method, so the specific limitations in one or more document quality inspection device embodiments provided below can refer to the limitations of the document quality inspection method above, and will not be repeated here.
[0280] In an exemplary embodiment, Figure 3 As shown, a document quality inspection device is provided, including: a document acquisition module 402, a feature extraction module 404 and a quality inspection module 406, wherein:
[0281] The document acquisition module 402 is used to acquire documents to be checked that are associated with power emergency events; the documents to be checked include at least one of accident reports, operation records, and emergency plans;
[0282] The feature extraction module 404 is used to extract features of the document to be checked to obtain document features of the document to be checked;
[0283] The quality inspection module 406 is used to input the document features into the trained quality inspection model to obtain the quality inspection result of the document to be inspected.
[0284] In an exemplary embodiment, the above-mentioned quality inspection module 406 is also used to correct the accident report characteristics according to the operation record characteristics to obtain corrected accident report characteristics; correct the emergency plan characteristics according to the corrected accident report characteristics to obtain corrected emergency plan characteristics; input the corrected emergency plan characteristics into the trained quality inspection model to obtain the quality inspection results of the emergency plan.
[0285] In an exemplary embodiment, the document quality inspection device further includes a preprocessing module for preprocessing the document to be inspected; the preprocessing includes at least one of text cleaning, format standardization, word segmentation, tagging, deduplication, stop word filtering, and data enhancement.
[0286] In an exemplary embodiment, the feature extraction module 404 is also used to extract keywords from the document to be checked; perform named entity recognition on the keywords; if the recognition is successful, extract the syntactic features, semantic features and content structure features of the document to be checked; and obtain the document features based on the syntactic features, the semantic features and the content structure features.
[0287] In an exemplary embodiment, the above-mentioned quality inspection module 406 is also used to input the document features into a preset rule engine model to obtain a first inspection result of the document to be inspected; input the document features into a trained machine learning model to obtain a second inspection result of the document to be inspected; and merge the first inspection result with the second inspection result to obtain the quality inspection result.
[0288] In an exemplary embodiment, the above-mentioned quality inspection module 406 is also used to input the document features into the rule engine model to obtain the format inspection result, content integrity inspection result and terminology consistency inspection result of the document to be inspected; based on the format inspection result, the content integrity inspection result and the terminology consistency inspection result, the first inspection result of the document to be inspected is obtained.
[0289] Each module in the above-mentioned document quality inspection device can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each module above.
[0290] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 4 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a document quality inspection method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.
[0291] Those skilled in the art will understand that Figure 4The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0292] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0293] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0294] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0295] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0296] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0297] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0298] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the present application. It should be noted that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A document quality inspection method, characterized in that: The method comprises: Acquire documents to be checked that are associated with power emergency events; the documents to be checked include at least one of accident reports, operation records, and emergency plans; Extracting features of the document to be checked to obtain document features of the document to be checked; The document features are input into a trained quality inspection model to obtain a quality inspection result of the document to be inspected.
2. The method according to claim 1, characterized in that The document features include accident report features, operation record features and emergency plan features; The step of inputting the document features into a trained quality inspection model to obtain a quality inspection result of the document to be inspected includes: Correcting the accident report feature according to the operation record feature to obtain a corrected accident report feature; Correcting the emergency plan characteristics according to the corrected accident report characteristics to obtain corrected emergency plan characteristics; The corrected emergency plan features are input into the trained quality inspection model to obtain the quality inspection results of the emergency plan.
3. The method according to claim 1, characterized in that Before extracting features from the document to be checked to obtain document features of the document to be checked, the method further includes: The document to be checked is preprocessed; the preprocessing includes at least one of text cleaning, format standardization, word segmentation, tagging, deduplication, stop word filtering, and data enhancement.
4. The method according to claim 1, characterized in that The step of extracting features from the document to be checked to obtain document features of the document to be checked includes: Extracting keywords from the document to be checked; Performing named entity recognition on the keywords; If the recognition is successful, extracting the syntactic features, semantic features and content structure features of the document to be checked; The document feature is obtained according to the syntactic feature, the semantic feature and the content structure feature.
5. The method according to claim 1, characterized in that The quality inspection model includes a rule engine model and a machine learning model; the document features are input into the trained quality inspection model to obtain the quality inspection result of the document to be inspected, including: Inputting the document features into a preset rule engine model to obtain a first inspection result of the document to be inspected; Inputting the document features into a trained machine learning model to obtain a second inspection result of the document to be inspected; The first inspection result and the second inspection result are combined to obtain the quality inspection result.
6. The method according to claim 5, characterized in that The step of inputting the document feature into a preset rule engine model to obtain a first inspection result of the document to be inspected includes: Inputting the document features into the rule engine model to obtain the format check result, content integrity check result and terminology consistency check result of the document to be checked; The first check result of the document to be checked is obtained according to the format check result, the content integrity check result and the terminology consistency check result.
7. A document quality inspection device, characterized in that: The device comprises: A document acquisition module, used to acquire documents to be checked that are associated with power emergency events; the documents to be checked include at least one of accident reports, operation records, and emergency plans; A feature extraction module, used to extract features of the document to be checked to obtain document features of the document to be checked; The quality inspection module is used to input the document features into the trained quality inspection model to obtain the quality inspection result of the document to be inspected.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.