Power system relay protection setting value list information extraction method and related device
Through the combination of global layout analysis and multimodal model, the logical partitions and semantic entities in the power system relay protection fixed value list are extracted and identified, which solves the problem of information extraction in the existing technology, and realizes the generation of structured data and the intelligent processing of complex documents.
Patent Information
- Application Number
- CN202510054966.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively extract and process information in power system relay protection fixed value sheets in complex formats, especially when facing non-digital outputs, diversified document structures and lack of semantic entity recognition capabilities.
The logical partition is identified by global layout analysis, combined with text detection and recognition model, extract text content, perform semantic classification and entity recognition, and infer the logical relationship between semantic entities through graph neural network or Transformer model to generate structured data.
It realizes automated information extraction and structured data generation of complex layout documents, improves the processing capability and accuracy of non-standardized fixed value orders, and is suitable for a variety of input forms.
Smart Images

Figure CN120014664A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power information extraction, and more specifically, relates to a method for extracting relay protection setting value information of a power system and a related device. Background Art
[0002] The relay protection setting sheet in the power system is an important technical document that records the protection setting and operating parameters of the equipment. Large setting sheets are converted into PDF format by scanning photos or documents. The setting sheets of different sites have different formats and layouts, and contain many terms, complex characters, and manual signatures. The existing system cannot intelligently extract information in this format. In order to achieve unified management of these data, it is necessary to extract the data information in the image format file and convert it into unified structured data. However, the manual entry method is inefficient and error-prone, and cannot meet the needs of large-scale setting sheet processing.
[0003] The method of extracting information from fixed-value sheets based on optical character recognition (OCR) technology has made certain progress in the field of document processing. It mainly realizes the extraction of text information through two core steps: text detection and text recognition. Although OCR technology has achieved certain results in general document processing, these general models are usually trained for various common documents and cannot fully cope with the complex needs of the specific field of fixed-value sheets. It is mainly manifested in: 1. The text information in fixed-value sheets is not only distributed in different areas, but also involves the combination of various structured information such as tables. Existing OCR technology is difficult to fully process these complex documents, especially when extracting specific field information and its relationship, it is often unable to accurately distinguish and establish the association relationship between fields. 2. Fixed-value sheet information not only requires text recognition, but also involves the accurate extraction of semantic entities (such as "title", "serial number", "fixed value name", "fixed value", etc.). When introducing the necessary semantic entity recognition (SER) technology, the recognition process not only depends on the text content, but also needs to be combined with the spatial layout characteristics of the document. 3. There are also complex logical relationships between the fields in the fixed-value sheet.
[0004] Prior art document 1 (CN118132620A) discloses a method, device and equipment for automatically extracting and comparing the setting value information of a relay protection device. The method includes: obtaining a print data stream output by the relay protection device in the process of executing relay protection; the print data stream includes the relay protection setting value information generated by the relay protection device when the relay protection device performs relay protection on the power system; based on a preset information matching method, extracting the actual setting value information from the print data stream; based on the information matching method, matching and comparing the actual setting value information and the preset reference setting value information to obtain the corresponding comparison result; the matching and comparison is used to determine the matching degree between the actual setting value information and the preset reference setting value information; the information matching method includes a fuzzy matching method based on keywords and / or an accurate matching method based on keywords.
[0005] The shortcomings of prior art document 1 are that the technology relies on the printed data stream output by the relay protection device and is only applicable to digital output scenarios, and cannot effectively process non-digital fixed value sheets (such as scanned copies or paper documents); at the same time, its keyword-based fuzzy matching and exact matching methods are easily limited by the quality of keyword settings when faced with complex format fixed value sheets, and it is difficult to cope with diverse document structures; in addition, it lacks semantic entity recognition capabilities and can only extract simple text information from printed data, and cannot accurately extract specific semantic entities such as "fixed value name", "fixed value" and their associations; finally, the technology does not combine the spatial layout characteristics of the document, and it is difficult to infer the logical relationship between fields in complex documents.
[0006] Prior art document 2 (CN113934865A) discloses a processing method for relay protection information standardization based on a knowledge graph, which relates to the field of power grid relay protection technology. The processing method is implemented in the process of standardizing the relay protection information text on the basis of establishing a standardized model. The standardized model is divided into three levels according to the large type, small type and manufacturer of the relay protection setting information to form a tree topology structure, and defines the first extraction value of the large type and the second extraction value of the small type of the relay protection setting information. The process of processing the relay protection information text includes: a. Scanning and reading the relay protection setting information text of each manufacturer; b. Extracting the small type and large type to which all the setting information of each manufacturer belongs through the first extraction value and the second extraction value, forming a mapping relationship, and storing it in a standard database; c. Formatting the relay protection setting information after the mapping.
[0007] The shortcoming of the prior art document 2 is that the technology mainly relies on a manually defined standardized model, and realizes the standardized processing of relay protection information through a tree topology structure and extracted values. The method is limited to rule-driven static mapping, and it is difficult to cope with complex and diverse document formats and non-standard expressions; at the same time, it lacks the ability to intelligently identify semantic entities and infer relationships in the text, and can only complete simple type mapping and formatting processing, and has limited ability to extract complex logical relationships; in addition, the technology does not combine the spatial layout characteristics of the document, and cannot directly process the fixed value orders stored in the form of pictures or scans.
[0008] Prior art document 3 (CN116598990A) discloses an online verification method for relay protection settings. The method includes: S1. In a regional power grid with a device as a protection unit, determine the relationship between the setting item and the selectivity, sensitivity and speed. Each setting item consists of a setting value and an action time. The sensitivity of the protection depends on the setting value, the speed is determined by the action time, and the selectivity of the protection is affected by both the setting value and the action time; S2. Investigate whether the sensitivity, selectivity and speed of the protection setting and action time meet the requirements as the core idea, and use the layered idea to maintain the device model, setting item and protection principle contained in the protection device in the system in turn.
[0009] The shortcomings of prior art document 3 are that its method mainly verifies the sensitivity, selectivity and speed of set values and action times, relies on manual configuration and a rule-based hierarchical verification mechanism, and lacks automated and intelligent support for the maintenance of protection device models, setting items and protection principles, as well as the generation of verification results; in addition, this method is only applicable to set value data with a high degree of digitization and standardization, and is difficult to cope with complex and irregular paper set value sheets or scanned input forms; at the same time, its verification process only focuses on the timing logic matching within and between protection devices, and does not conduct in-depth analysis of the semantic entities and logical relationships of specific fields in the set value sheet, resulting in limited applicability and efficiency in non-standardized scenarios. Summary of the invention
[0010] In order to solve the deficiencies in the prior art, the present invention provides a method and a related device for extracting information from a relay protection setting sheet of a power system, which are used to realize OCR recognition and information extraction of a relay protection setting sheet of a power system.
[0011] The present invention adopts the following technical solution.
[0012] A first aspect of the present invention provides a method for extracting information of relay protection setting values in a power system, comprising the following steps:
[0013] Perform global layout analysis on the content of the relay protection setting list image, identify the logical partitions in the relay protection setting list image, including the title area, setting table area and signature area, extract the boundary information of the logical partitions, and generate a logical partition framework diagram;
[0014] According to the logical partition framework diagram, extract the text area of each logical partition, identify the text content in the area, and record the spatial position information of each character in the text; arrange the text content according to the spatial position information of the characters to construct the text sequence in the logical partition; combine the context information across the logical partitions, merge the text sequences in the logical partitions, and generate a continuous text stream;
[0015] According to the text flow and partition framework diagram, the text sequence is semantically classified to generate semantic entities of title, fixed value name, fixed value pair, serial number and date to construct the hierarchical relationship of semantic entities;
[0016] By utilizing the hierarchical relationship of semantic entities, a relationship extraction model is constructed to analyze the logical relationship between semantic entities, transform the content and logical relationship of semantic entities into structured data, and extract the information of relay protection setting list.
[0017] Preferably, the performing of global layout analysis on the content of the relay protection setting value single image to identify the logical partitions in the relay protection setting value single image includes:
[0018] Extract the relay protection setting single image, use the text detection model to process the relay protection setting single image, and automatically detect the text area in the relay protection setting single image; wherein, the text detection model adopts the DBNet detection and recognition neural network model in PaddleOCR, extracts features of the relay protection setting single image content through the convolutional neural network, and generates a probability map to refine the text boundary; outputs the bounding box containing the text area and the corresponding position information, and the text area includes at least one text box, which is used to identify the text area in the relay protection setting single image;
[0019] Analyze the text area in the relay protection setting value single image, detect the character dense area, table lines and relay protection setting value single image structure, and generate partition candidate areas.
[0020] Preferably, the step of extracting the boundary information of the partitions and generating a partition framework diagram includes:
[0021] Based on the partition candidate area, combined with the spatial distribution pattern and context information, the boundary points of the logical partition are extracted and the bounding box is generated, and the spatial coordinates and category labels of the partition frame are recorded;
[0022] Construct a partition framework diagram, which includes the boundary locations and category labels of logical partitions.
[0023] Preferably, extracting the text area of each logical partition includes:
[0024] The DBNet detection and recognition neural network model is used to detect the text content in the logical partition and output the bounding box containing the text area and the corresponding confidence level;
[0025] The detected bounding box area is cropped, and the text content in the cropped area is recognized through a convolutional recurrent neural network, and a character sequence is output;
[0026] Combining the confidence and detection position information of the characters, secondary detection and optimization are performed on the text areas with confidence below the set threshold.
[0027] Preferably, the recording of the spatial position information of each character in the text; arranging the text content according to the spatial position information of the characters, and constructing a text sequence within the logical partition includes:
[0028] According to the boundary coordinates of the logical partition framework diagram, the single image content of the relay protection setting value in each logical partition is cropped to extract the text area image corresponding to the logical partition;
[0029] The text recognition model is used to extract the bounding box coordinate information of each character in the relay protection setting single image, including the starting coordinates and occupied area of the character. The text recognition model uses a convolutional recurrent neural network to extract the image features of the text area, and combines the recurrent neural network and the CTC conditional random field loss function to analyze the time series relationship of the characters.
[0030] Arrange the character contents according to the spatial position information of the characters and construct the text sequence within each logical partition.
[0031] Preferably, the semantic classification of the text sequence according to the text flow and the partition framework diagram includes:
[0032] Based on the logical partition information in the partition framework diagram, partition association analysis is performed on the text sequence in the text stream, and each text sequence is matched with a partition category;
[0033] Based on the content and spatial position information of each text sequence in the text stream, the entity category labels are classified and generated, including semantic entities such as title, fixed value name, fixed value value pair, serial number and date;
[0034] Output semantic classification results, which include the category label, text content and spatial location information of each semantic entity.
[0035] Preferably, the step of constructing a hierarchical relationship of semantic entities includes:
[0036] Based on the generated semantic entities, the graph neural network is used to analyze the subordinate relationships between semantic entities by combining the category label, text content and spatial location information of each semantic entity.
[0037] According to the arrangement order and logical rules of semantic entities, a hierarchical relationship is constructed, including the subordinate relationship between the title and the text, the preliminary association relationship between the fixed value name and the fixed value pair, and the matching relationship between the serial number and the date; the hierarchical relationship of the semantic entities is output.
[0038] Preferably, the step of constructing a relationship extraction model by utilizing the hierarchical relationship of semantic entities includes:
[0039] The graph neural network model is used to logically model the hierarchical relationship, and the semantic entities and their associations are converted into a graph structure, where nodes represent semantic entities, node features include category labels, text content, and spatial location information, and edges represent logical associations between semantic entities. Edge features are calculated through the spatial distance, contextual relationship, and semantic relevance of semantic entities.
[0040] Based on the hierarchical relationship of semantic entities, the features of nodes and edges are aggregated through graph neural networks, including weighted combination of the feature vector of each node with the feature vectors of adjacent nodes and their edges to generate a new node feature representation. At the same time, the feature weights of different relationship categories are defined through a weighted learning mechanism of multiple relationship types to optimize the logical association of semantic entities.
[0041] Based on the aggregated node and edge features, the semantic entity pairs are classified and the classification results are generated according to the classification standards of subordinate relationship, corresponding relationship and no relationship.
[0042] According to the relationship classification results, a relationship extraction model is generated. By integrating the category labels, text content and spatial information of semantic entity nodes, as well as the logical associations between entities, structured data for single semantic relationship parsing of relay protection settings is output, including entity categories, entity contents and logical associations.
[0043] Preferably, converting the content and logical relationship of the semantic entity into structured data includes:
[0044] The output of structured data is in the following format:
[0045] Convert semantic entity content and logical relationships into JSON format files. The JSON format contains entity content, labels, and hierarchical relationships expressed in the form of key-value pairs.
[0046] For tabular display scenarios, the semantic entity content and association relationships are generated into an XML tree structure data file for structured representation of the semantic entity content and association relationships.
[0047] A second aspect of the present invention provides a device for extracting information of relay protection setting values in a power system, comprising:
[0048] A text detection module is used to perform text detection on an image containing a target relay protection setting list using a text detection model to obtain a text area; wherein the text area includes at least one text box and corresponding position information;
[0049] A text recognition module is used to perform text recognition on a text area using a text recognition model to obtain text content;
[0050] A semantic entity recognition module, used to identify semantic entities using a semantic entity recognition model according to the text content and the corresponding text area;
[0051] The relationship extraction module is used to infer the relationship between semantic entities and generate structured association data based on the relationship between semantic entities; wherein the format of the structured association data includes the category, content, spatial location information and inferred logical association of the semantic entity.
[0052] Compared with the prior art, the beneficial effects of the present invention include at least: the information extraction method and related device of the power system relay protection setting list provided by the present invention have the following advantages: the bounding box of all text areas in the setting list is identified through the DBNet detection and recognition neural network model in PaddleOCR. The convolutional recurrent neural network CRNN model in PaddleOCR is used to identify the text in the bounding box as a readable string. Through the LayoutXLM model, the text is classified and the specific entity type (such as "setting name", "setting", "date", etc.) is identified by combining visual and text information. Through the graph neural network or Transformer model, the relationship between different semantic entities is inferred to generate structured associated data. PaddleOCR's multimodal model LayoutXLM can combine visual features and language features, and is suitable for entity recognition of complex layout documents such as setting lists. Relationship extraction technology extracts deeper logical relationships from semantic entities through graph neural networks or Transformer models to reconstruct structured information in documents. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a flow chart of a method for extracting information of relay protection setting value list of a power system provided in accordance with an embodiment of the present invention;
[0054] Figure 2 It is a schematic diagram of a power system relay protection setting value single information extraction device provided according to an embodiment of the present invention;
[0055] Figure 3is a schematic diagram of an electronic device provided according to an embodiment of the present invention;
[0056] Figure 4 is a schematic diagram of a semantic entity recognition result provided according to an embodiment of the present invention;
[0057] Figure 5 It is a schematic diagram of a relationship extraction result provided according to an embodiment of the present invention. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical scheme and advantages of the present invention clearer, the technical scheme of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. The embodiments described in this application are only part of the embodiments of the present invention, not all of them. Based on the spirit of the present invention, other embodiments obtained by ordinary technicians in this field without creative work are all within the scope of protection of the present invention.
[0059] In the description of the present invention, "several" means more than one, "many" means more than two, "greater than", "less than", "exceed", etc. are understood to exclude the number itself, and "above", "below", "within", etc. are understood to include the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0060] In the description of the present invention, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0061] like Figure 1 As shown, the method for extracting information from a relay protection setting list of a power system provided by the present invention includes a series of steps. First, a text detection model is used to locate the text area in the relay protection setting list image to generate a bounding box and position information; then a text recognition model is used to convert the content in the text area into a character sequence; then, a semantic classification model is used to identify semantic entities such as titles, setting names, setting value pairs, serial numbers and dates; on this basis, a graph neural network model is used to infer the logical relationship between semantic entities to form a mapping relationship between setting names and setting value pairs and an association relationship between serial numbers and dates; finally, the extracted content and logical relationship are converted into structured data to support subsequent analysis and processing.
[0062] like Figure 2 As shown, the power system relay protection setting value single information extraction device of the present invention is composed of multiple modules, including a text detection module, a text recognition module, a semantic classification module and a relationship extraction module. The text detection module receives the input image and extracts the text area, the text recognition module is responsible for parsing the text area content to generate a character sequence, the semantic classification module semantically annotates the text through a model, extracts the semantic entity category and location information, and the relationship extraction module infers the logical relationship based on the semantic entity and generates structured data. These modules work together to achieve efficient information extraction and conversion in complex document scenarios.
[0063] like Figure 3 As shown, the method for extracting setting list information of relay protection in a power system of the present invention can be implemented by an electronic device. The electronic device includes a processor, a memory and an input-output interface. The processor loads the computer program in the memory and sequentially executes the steps of text detection, text recognition, semantic classification and logical modeling to complete the extraction of setting list information. The input-output interface is used to receive image input and output structured data results. The device design simplifies the process of extracting complex document information and can be applied to real-time data processing scenarios.
[0064] like Figure 4 As shown, the semantic entity recognition results in the present invention include classification and annotation of multiple text regions. By semantically classifying the extracted text content, semantic entities of categories such as title, fixed value name, fixed value pair, serial number and date are generated; at the same time, the spatial location information and contextual features of the text are combined to intuitively present the hierarchical structure and logical association of the semantic entity. The visualized semantic classification results provide a reliable foundation for subsequent logical modeling and data conversion.
[0065] Example 1 of the present invention provides a method for extracting information of relay protection setting values in a power system, comprising the following steps:
[0066] Perform global layout analysis on the content of the relay protection setting list image, identify the logical partitions in the relay protection setting list image, including the title area, setting table area and signature area, extract the boundary information of the logical partitions, and generate a partition framework diagram;
[0067] Preferably, the performing of global layout analysis on the content of the relay protection setting value single image to identify the logical partitions in the relay protection setting value single image includes:
[0068] Extract the relay protection setting single image, use the text detection model to process the relay protection setting single image, and automatically detect the text area in the image; wherein, the text detection model uses the DBNet detection and recognition neural network model in PaddleOCR, extracts features from the relay protection setting single image content through a convolutional neural network, and generates a probability map to refine the text boundary; outputs the bounding box containing the text area and the corresponding position information, and the text area includes at least one text box, which is used to identify the text area in the relay protection setting single image;
[0069] Analyze the text area in the relay protection setting single image, detect the character dense area, table lines and image structure, and generate partition candidate areas.
[0070] Preferably, the step of extracting the boundary information of the partitions and generating a partition framework diagram includes:
[0071] Based on the candidate partition regions, combined with the spatial distribution pattern and context information, the logical partition boundaries are determined, including the title region, the set value table region and the signature region;
[0072] Extract boundary points and generate boundary boxes for the boundary information of logical partitions, and record the spatial coordinates, category labels, and hierarchical relationships of the partition framework;
[0073] Construct a partition framework diagram, which includes the boundary locations of logical partitions, category labels, and logical associations between regions.
[0074] In some embodiments, the text detection model adopts a DBNet model.
[0075] The text detection module in PaddleOCR is used to automatically identify text areas from images. It extracts image features through a convolutional neural network (CNN) and uses a specific detection algorithm to identify the location of text.
[0076] The DBNet model in PaddleOCR is a commonly used text detection model, which is suitable for processing scenes with complex backgrounds, such as technical documents in fixed value orders. DBNet generates probability maps on feature maps to further refine text boundaries, thereby generating more accurate text area detection results.
[0077] In this embodiment, the workflow of the text detection model is as follows:
[0078] Input: Image of bill of value or ticket.
[0079] Output: Bounding box with text area.
[0080] Model: DBNet model provided by PaddleOCR for text detection.
[0081] PaddleOCR uses multi-scale training and feature pyramid networks to improve the detection accuracy of texts of different sizes, especially for small fonts and long texts in large-size documents such as A4 paper.
[0082] The significant difference between the present invention and the prior art is that, by introducing OCR technology, it supports text detection and recognition of complex fixed value orders in scanned or photo formats, expanding the input form; using the LayoutXLM model combined with the document space layout characteristics, it can accurately identify semantic entities such as "fixed value name", "fixed value", and "serial number"; at the same time, through the graph neural network (GNN) or Transformer model, relationship extraction is performed to infer the logical relationship between semantic entities, and a complete structured associated data is constructed, thereby effectively processing complex layout documents. The beneficial effect achieved by the present invention relative to the prior art is that it is suitable for non-digital fixed value orders in various forms such as paper documents and scanned copies, and can automatically identify specific fields and their associated relationships, significantly improving the accuracy and processing efficiency of information extraction of complex and irregular format fixed value orders, and providing an efficient and comprehensive solution for the automated processing of large-scale fixed value orders.
[0083] The beneficial effect of the present invention compared to the prior art is that it not only supports the extraction of setting value information in complex and irregular formats, but also can automatically verify the logical relationship between the setting value and the action time in combination with the spatial layout and semantic characteristics of the setting value list; by inferring the relationship between semantic entities through a deep model, it reduces the reliance on manual configuration rules, improves the efficiency and accuracy of the verification process, and provides a more intelligent and comprehensive solution for the online verification of relay protection settings in complex scenarios.
[0084] According to the partition framework diagram, extract the text area of each logical partition, identify the text content in the area, and record the spatial position information of each character in the text; arrange the text content according to the spatial position information of the characters to construct the text sequence in the logical partition; combine the context information across the logical partitions, merge the text sequences in the logical partitions, and generate a continuous text stream;
[0085] Preferably, the step of constructing a text sequence within a logical partition and combining the context information across logical partitions to merge the text sequence within the logical partition to generate a continuous text stream includes:
[0086] According to the boundary coordinates of the partition frame diagram, the image content in each logical partition is cropped to extract the text area image corresponding to the logical partition;
[0087] The extracted text area image is recognized using a text recognition model, character content is extracted, and the spatial position information of each character is recorded, including the starting coordinates and occupied area of the character. The text recognition model uses a convolutional recurrent neural network to extract image features of the text area, and combines the convolutional recurrent neural network and the CTC conditional random field loss function to analyze the time series relationship of the characters.
[0088] Arrange character contents according to the spatial position information of characters and construct text sequences within each logical partition;
[0089] Combined with the contents of multiple logical partitions in the partition framework diagram, the related character contents between the partitions are merged according to the contextual logic and spatial position of the text sequence to generate a continuous text flow.
[0090] In some embodiments, the text recognition model adopts a CRNN model; the CTC (Connectionist Temporal Classification) loss function is used when training the CRNN model. The CRNN model is not only suitable for Chinese character recognition in the value list, but also supports multi-language recognition. The CTC loss function effectively solves the sequence-to-sequence mapping problem, and is particularly suitable for the case where the character spacing in the value list is uneven.
[0091] The task of text recognition is to extract the content in the detected text area and convert it into machine-readable text.
[0092] The text recognition model in PaddleOCR is usually based on CRNN (Convolutional Recurrent Neural Network) and CTC (Connectionist Temporal Classification) loss function. This combined model can efficiently recognize continuous characters in fixed value orders and adapt to different character lengths and complex font styles.
[0093] In this embodiment, the workflow of the text recognition model is as follows:
[0094] Input: Text region obtained by text detection.
[0095] Output: recognized text content.
[0096] Model: The CRNN model provided by PaddleOCR is used to recognize text content from text area images.
[0097] According to the text flow and partition framework diagram, the text sequence is semantically classified to generate semantic entities of title, fixed value name, fixed value pair, serial number and date, and the hierarchical relationship of semantic entities is constructed;
[0098] Preferably, the semantic classification of the text sequence according to the text flow and the partition framework diagram includes:
[0099] Based on the logical partition information in the partition framework diagram, partition association analysis is performed on the text sequence in the text stream, and each text sequence is matched with a partition category;
[0100] Based on the content and spatial position information of each text sequence in the text stream, the entity category labels are classified and generated, including semantic entities such as title, fixed value name, fixed value value pair, serial number and date;
[0101] Output semantic classification results, which include the category label, text content and spatial location information of each semantic entity.
[0102] Preferably, the step of constructing a hierarchical relationship of semantic entities includes:
[0103] Based on the generated semantic entities, the graph neural network is used to analyze the subordinate relationships between semantic entities by combining the category label, text content and spatial location information of each semantic entity.
[0104] According to the arrangement order and logical rules of semantic entities, a hierarchical relationship is constructed, including the subordinate relationship between the title and the text, the preliminary association relationship between the fixed value name and the fixed value pair, and the matching relationship between the serial number and the date; the hierarchical relationship of the semantic entities is output.
[0105] In some embodiments, the semantic entity recognition model adopts the LayoutXLM model. In some embodiments, the types of the semantic entities include: "title", "serial number", "fixed value name", "fixed value", and "date".
[0106] Semantic Entity Recognition (SER) is to extract fields or entities with specific meanings from the text recognized by OCR, such as "serial number", "fixed value name", "fixed value", etc. in the fixed value list.
[0107] LayoutXLM provided by PaddleOCR is a pre-trained model based on multimodality, combining the features of text and images. Through the combination of vision and language, the LayoutXLM model can not only extract text from images, but also understand the spatial layout relationship of text. Therefore, it is very suitable for processing files with complex layouts such as tables and forms.
[0108] In this embodiment, the workflow of the semantic entity recognition model is as follows:
[0109] Input: text content and location information obtained through text recognition.
[0110] Output: text with entity labels, such as "serial number", "setting name", "setting value", etc.
[0111] Model: The LayoutXLM model in PaddleOCR combines text content and spatial location information for entity recognition.
[0112] Utilizing the hierarchical relationship of semantic entities, the logical relationship between semantic entities is constructed through graph neural network, the content and logical relationship of semantic entities are converted into structured data, and the information of relay protection setting list is extracted.
[0113] Preferably, the step of constructing the logical relationship between semantic entities through a graph neural network and converting the content and logical relationship of the semantic entities into structured data includes:
[0114] The graph neural network model is used to logically model the hierarchical relationship, and the semantic entities and their associations are converted into a graph structure, where nodes represent semantic entities and edges represent logical associations between entities.
[0115] Combining the context information of the semantic entity and the connection relationship in the graph structure, the mapping relationship between the fixed value name and the fixed value pair, as well as the association relationship between the serial number and the date, is inferred;
[0116] Output structured data, including the category, content, spatial location information and inferred logical association of semantic entities, for information storage and processing of relay protection setting sheets.
[0117] Relation extraction (RE) is to infer the relationship between the identified semantic entities. The relationship between the semantic entities in the fixed value list includes the correspondence between "fixed value name" and "fixed value", or the association between "serial number" and other fields. The relation extraction model uses spatial layout and context information to infer the logical relationship between entities. Especially in table documents, relation extraction can help reconstruct the logical structure of data.
[0118] The significant difference between the present invention and the prior art is that the text detection and recognition of scanned and picture documents is realized through OCR technology, and multiple input forms are supported; the semantic entity recognition model is used to automatically extract semantic entities such as "fixed value name", "fixed value", and "serial number" in combination with visual and text features; and the relationship between semantic entities is inferred through a graph neural network (GNN) or a Transformer model to generate structured associated data, thereby breaking through the dependence of traditional rule-driven methods on classification systems and document formats. The beneficial effect of the present invention relative to the prior art is that, through semantic entity recognition and relationship extraction technology, complex and non-standardized fixed value orders can be intelligently processed, and unstructured input forms such as scanned copies and pictures can be supported; combined with spatial layout features, not only can the field content be extracted, but also the logical relationship between fields can be reconstructed, providing a more efficient, intelligent and flexible solution for large-scale and diversified fixed value order processing, while avoiding dependence on manually predefined classification systems, and significantly improving the generalization ability and efficiency of processing.
[0119] In some embodiments, the relationship extraction model adopts a graph neural network GNN or a Transformer model.
[0120] PaddleOCR models the relationship between semantic entities by using Graph Neural Network (GNN) or Transformer structure. GNN can effectively capture the relationship between entities, especially suitable for documents in the form of tables in value orders.
[0121] In this embodiment, the workflow of the relationship extraction model is as follows:
[0122] Input: Identified semantic entities (such as "serial number", "fixed value name") and their location information.
[0123] Output: Relationships between entities (such as the association between “fixed value name” and “fixed value”).
[0124] Model: PaddleOCR provides a relation extraction model based on GNN or Transformer, which predicts relations by learning the dependencies between entities.
[0125] Specific application examples:
[0126] Data preparation
[0127] 1.1PDF Conversion
[0128] 1.1.1 Separate the images in the PDF into single jpg images;
[0129] 1.1.2 Unify file names and divide the images into training and validation sets to obtain an unlabeled dataset.
[0130] 1.2 Custom Labels
[0131] First, customize the labels according to the actual needs in the setting list. Divide all elements into the following categories of labels:
[0132] HEADER, QUESTION, ANSWER, INDEX, DATE, OTHER; These tags may be used in the value sheet to identify text areas such as header information, question, answer, serial number, date, etc.
[0133] Save the labels in text format and place them in the root directory of the dataset.
[0134] 1.3 Data annotation:
[0135] Use PPOCRLabel provided by PaddlePaddle to annotate the data to prepare the dataset with labels and relationship information required for model training;
[0136] 1.3.1 Import data and use automatic annotation to mark out the text in the value list;
[0137] Open PPOCRLabel to label the data, and use the built-in labeling function to label the text in the fixed value list;
[0138] 1.3.2RE Relationship Definition
[0139] When using the PPOCRLabel tool to annotate text, since it does not support relationship annotation (linking) itself, we need to use some techniques to indirectly implement the definition of these relationships. Specifically, we can assign specific serial numbers (such as INDEX1, QUESTION1, ANSWER1) to different key text blocks during the annotation stage, and then use code to automatically add id and linking fields according to the same serial number in the exported annotation data to establish the relationship between text blocks.
[0140] Here are the detailed steps:
[0141] 1.PPOCRLabel labeling stage
[0142] When annotating, we cannot directly define the relationship between text blocks, so we need to use the same sequence number for key text blocks to identify the relationship between them. For example, for a question and its corresponding answer and index item, give them the same number:
[0143] INDEX1: Indicates that this text block is the index of sequence number 1.
[0144] QUESTION1: Indicates that this text block is question number 1.
[0145] ANSWER1: Indicates that this text block is the answer to number 1.
[0146] Similarly, each group of related text blocks is marked with the same number (such as 1, 2, etc.).
[0147] 2. Export annotation data
[0148] The annotation file generated by PPOCRLabel does not contain relationships (linking), and only exports the basic information of each text block, such as text content, location coordinates, and labels.
[0149] 3. Add id and linking fields through code;
[0150] After exporting the annotation data, you need to generate unique IDs for these text blocks through code and establish linking relationships based on the same number (such as 1). The specific steps are as follows:
[0151] Generate unique id for each text block: Iterate through each text block in the dataset and assign them a unique id in sequence.
[0152] Generate linking based on the number in the label: traverse the text blocks in the dataset and find the text blocks with the same number (such as QUESTION1, ANSWER1, INDEX1). Connect the IDs of these text blocks to form a linking relationship.
[0153] 4. After processing all the data, save the results back to the file and use it to train the RE model. In this process, by manually labeling the text blocks with the same number, and then automatically generating IDs and linking through the code, the relationship annotation required for the RE task can be effectively achieved.
[0154] Finally, through annotation and code processing, the complete dataset structure is obtained.
[0155] Text Detection Model
[0156] The text detection model is the first step in the OCR process. Its main function is to detect the area containing text from the input image, that is, to identify all the text boxes in the image. These text boxes mark the areas in the image that may contain text for use by the subsequent text recognition model.
[0157] The text detection model locates these text areas by analyzing the visual features in the image and returns their location information (usually expressed in the form of a rectangular box or quadrilateral). For complex documents, bills, and forms, the text detection model can effectively extract text content from the background. It can not only detect neatly arranged text, but also handle irregular situations such as handwritten text and tilted text.
[0158] The trained text detection model can be used to detect text regions in new images during the inference phase. The output of the model is the coordinates of each text region, which will be used for text recognition in the next step.
[0159] Text recognition model
[0160] The main function of the text recognition model is to identify the specific text content from the text area (i.e. text box) output by the text detection model. In simple terms, the task of the text recognition model is to convert the text in these images into text information that can be understood by the machine. It can handle various types of text, including printed text, handwritten text, text in different fonts, and even rotated, tilted, or distorted text areas.
[0161] In the OCR process, the text recognition model is a key step after text detection, which directly affects the final text recognition accuracy. Through the output of the recognition model, the text information in the document can be extracted and used for further processing and analysis.
[0162] The dataset for the text recognition model is different from the previously prepared dataset. The specific format is as follows
[0163] "Image file name image annotation information"
[0164] The training process of the text recognition model includes:
[0165] Load pre-trained model: Directly use the pre-trained text recognition model provided by PaddleOCR and fine-tune it on a specific dataset. This can significantly shorten the training time and improve the model performance.
[0166] Loss function: The commonly used loss function for text recognition models is CTC (Connectionist Temporal Classification) loss, which can handle output sequences of variable length and is very suitable for text recognition tasks.
[0167] Optimizer settings: Use the optimizer to update model parameters during training. Set appropriate learning rate, regularization parameters, etc. to control the stability and convergence speed of the training process.
[0168] Model Inference
[0169] After training, the text recognition model can be used to perform inference on new images. The specific steps are:
[0170] Enter the text box image to be recognized.
[0171] The model outputs the corresponding text sequence.
[0172] Post-process the output, such as deleting invalid characters and performing character correction, to obtain the final recognition result.
[0173] The text recognition model is a key part of the OCR process. It converts the text area image output by the text detection model into actual text information. By using the text recognition model in PaddleOCR and combining data preprocessing, model training, reasoning and post-processing techniques, the text of the fixed value order can be effectively recognized.
[0174] SE (Semantic Entity) Model
[0175] The SE (Semantic Entity) model is used to identify specific types of semantic entities from documents. In the fixed-value single OCR task, the SE model can identify key content in the document, such as header, index, question, answer and other entities. These entities often have clear semantics and can be classified and extracted by the SE model.
[0176] The role of the SE model is to further analyze and classify the text that has been recognized by OCR, so that the model can understand the structure and content of the document, providing a basis for subsequent relationship extraction and information extraction.
[0177] The following information is automatically printed in the log:
[0178]
[0179] The prediction results are as follows Figure 5 As shown:
[0180] Through the prediction of the RE model, different entities in the value sheet are connected by straight lines, which represent that the relationship between entities has been successfully identified. Through this relationship extraction, the originally independent text boxes are orderly associated to construct a semantic network between entities. Combined with the entity label information previously obtained through the SE (Semantic Entity Recognition) model, the key information in the entire value sheet, such as Header, Index, Question and Answer, are all structured and extracted and associated together. Finally, the combination of SE and RE models realizes the complete extraction of the data required for the value sheet, and converts it into a consistent format for subsequent processing and application. This unified data format can better support tasks such as automated analysis and information retrieval.
[0181] Data normalization process
[0182] Through the SE (semantic entity recognition) and RE (relation extraction) models, we have extracted and identified the key information in the value list, but in order to facilitate subsequent analysis and application, the results of the SE and RE models need to be normalized and converted into a unified data format. The core of data normalization is to convert different entities and their relationships into a structured and easy-to-process format. The following are the specific normalization steps:
[0183] Step 1: Collect the output of the SE model
[0184] The output of the SE model is the identified text entities and their category information, such as Header, Index, Question, Answer, etc. These entities are often marked in the form of "label + serial number", such as Question1, Answer1, in order to provide a basis for subsequent relationship matching.
[0185] Step 2: Collect the output of the RE model
[0186] The output of the RE model is the relationship between the identified entities. The RE model identifies the logical relationship between these entities by connecting them with straight lines, such as the corresponding relationship between Header and Index, or the mapping relationship between Question and Answer.
[0187] Step 3: Create a unified data structure
[0188] Convert the output of SE and RE models into a unified structured data format, which usually contains the following fields:
[0189] Entity ID: A unique identifier for each entity.
[0190] Entity tags: such as Header, Index, Question, Answer, etc.
[0191] Text content: The actual text corresponding to the entity.
[0192] Example 2 of the present invention provides a device for extracting information of relay protection setting values in a power system, comprising:
[0193] A text detection module is used to perform text detection on an image containing a target relay protection setting list using a text detection model to obtain a text area; wherein the text area includes at least one text box and corresponding position information;
[0194] A text recognition module is used to perform text recognition on a text area using a text recognition model to obtain text content;
[0195] A semantic entity recognition module, used to identify semantic entities using a semantic entity recognition model according to the text content and the corresponding text area;
[0196] The relationship extraction module is used to infer the relationship between semantic entities and generate structured association data based on the relationship between semantic entities; wherein the format of the structured association data includes the category, content, spatial location information and inferred logical association of the semantic entity.
[0197] Example 3 of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method for extracting information from a relay protection setting value sheet of a power system is implemented.
[0198] Example 4 of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the method for extracting information from a relay protection setting value sheet of a power system when executing the computer program.
[0199] Compared with the prior art, the beneficial effects of the present invention include at least: the information extraction method and related device of the power system relay protection setting list provided by the present invention have the following advantages: the bounding box of all text areas in the setting list is identified through the DBNet detection and recognition neural network model in PaddleOCR. The convolutional recurrent neural network CRNN model in PaddleOCR is used to identify the text in the bounding box as a readable string. Through the LayoutXLM model, the text is classified and the specific entity type (such as "setting name", "setting", "date", etc.) is identified by combining visual and text information. Through the graph neural network or Transformer model, the relationship between different semantic entities is inferred to generate structured associated data. PaddleOCR's multimodal model LayoutXLM can combine visual features and language features, and is suitable for entity recognition of complex layout documents such as setting lists. Relationship extraction technology extracts deeper logical relationships from semantic entities through graph neural networks or Transformer models to reconstruct structured information in documents.
[0200] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0201] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0202] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1A function specified in one or more boxes.
[0203] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0204] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0205] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for extracting information of relay protection setting values in a power system, characterized in that: The following steps are involved: Perform global layout analysis on the content of the relay protection setting list image, identify the logical partitions in the relay protection setting list image, including the title area, setting table area and signature area, extract the boundary information of the logical partitions, and generate a logical partition framework diagram; According to the logical partition framework diagram, extract the text area of each logical partition, identify the text content in the area, and record the spatial position information of each character in the text; arrange the text content according to the spatial position information of the characters to construct the text sequence in the logical partition; combine the context information across the logical partitions, merge the text sequences in the logical partitions, and generate a continuous text stream; According to the text flow and partition framework diagram, the text sequence is semantically classified to generate semantic entities of title, fixed value name, fixed value pair, serial number and date to construct the hierarchical relationship of semantic entities; By utilizing the hierarchical relationship of semantic entities, a relationship extraction model is constructed to analyze the logical relationship between semantic entities, transform the content and logical relationship of semantic entities into structured data, and extract the information of relay protection setting list.
2. The method for extracting information from relay protection setting lists of a power system according to claim 1, characterized in that: The global layout analysis of the relay protection setting value single image content is performed to identify the logical partitions in the relay protection setting value single image, including: Extract the relay protection setting single image, use the text detection model to process the relay protection setting single image, and automatically detect the text area in the relay protection setting single image; wherein, the text detection model adopts the DBNet detection and recognition neural network model in PaddleOCR, extracts features of the relay protection setting single image content through the convolutional neural network, and generates a probability map to refine the text boundary; outputs the bounding box containing the text area and the corresponding position information, and the text area includes at least one text box, which is used to identify the text area in the relay protection setting single image; Analyze the text area in the relay protection setting value single image, detect the character dense area, table lines and relay protection setting value single image structure, and generate partition candidate areas.
3. The method for extracting information from relay protection setting values of a power system according to claim 2, characterized in that: The step of extracting the boundary information of the partition and generating the partition framework diagram includes: Based on the partition candidate area, combined with the spatial distribution pattern and context information, the boundary points of the logical partition are extracted and the bounding box is generated, and the spatial coordinates and category labels of the partition frame are recorded; Construct a partition framework diagram, which includes the boundary locations and category labels of logical partitions.
4. A method for extracting relay protection setting value information of a power system according to claim 1 or 2, characterized in that: The extracting of the text area of each logical partition comprises: The DBNet detection and recognition neural network model is used to detect the text content in the logical partition and output the bounding box containing the text area and the corresponding confidence score; The detected bounding box area is cropped, and the text content in the cropped area is recognized through a convolutional recurrent neural network, and a character sequence is output; Combining the confidence and detection position information of the characters, secondary detection and optimization are performed on the text areas with confidence below the set threshold.
5. The method for extracting information from relay protection setting lists of a power system according to claim 1, characterized in that: The step of recording the spatial position information of each character in the text; arranging the text content according to the spatial position information of the characters, and constructing a text sequence within a logical partition includes: According to the boundary coordinates of the logical partition framework diagram, the single image content of the relay protection setting value in each logical partition is cropped to extract the text area image corresponding to the logical partition; The text recognition model is used to extract the bounding box coordinate information of each character in the relay protection setting single image, including the starting coordinates and occupied area of the character. The text recognition model uses a convolutional recurrent neural network to extract the image features of the text area, and combines the recurrent neural network and the CTC conditional random field loss function to analyze the time series relationship of the characters. Arrange the character contents according to the spatial position information of the characters and construct the text sequence within each logical partition.
6. A method for extracting relay protection setting value information of a power system according to claim 1, characterized in that: The semantic classification of the text sequence according to the text flow and the partition framework diagram includes: Based on the logical partition information in the partition framework diagram, partition association analysis is performed on the text sequence in the text stream, and each text sequence is matched with a partition category; Based on the content and spatial position information of each text sequence in the text stream, the entity category labels are classified and generated, including semantic entities such as title, fixed value name, fixed value value pair, serial number and date; Output semantic classification results, which include the category label, text content and spatial location information of each semantic entity.
7. A method for extracting relay protection setting value information of a power system according to claim 1, characterized in that: The step of constructing a hierarchical relationship of semantic entities includes: Based on the generated semantic entities, the graph neural network is used to analyze the subordinate relationships between semantic entities by combining the category label, text content and spatial location information of each semantic entity. According to the arrangement order and logical rules of semantic entities, a hierarchical relationship is constructed, including the subordinate relationship between the title and the text, the preliminary association relationship between the fixed value name and the fixed value pair, and the matching relationship between the serial number and the date; the hierarchical relationship of the semantic entities is output.
8. The method for extracting relay protection setting value information of a power system according to claim 1, characterized in that: The method of constructing a relation extraction model by utilizing the hierarchical relationship of semantic entities includes: The graph neural network model is used to logically model the hierarchical relationship, and the semantic entities and their associations are converted into a graph structure, where nodes represent semantic entities, node features include category labels, text content, and spatial location information, and edges represent logical associations between semantic entities. Edge features are calculated through the spatial distance, contextual relationship, and semantic relevance of semantic entities. Based on the hierarchical relationship of semantic entities, the features of nodes and edges are aggregated through graph neural networks, including weighted combination of the feature vector of each node with the feature vectors of adjacent nodes and their edges to generate a new node feature representation. At the same time, the feature weights of different relationship categories are defined through a weighted learning mechanism of multiple relationship types to optimize the logical association of semantic entities. Based on the aggregated node and edge features, the semantic entity pairs are classified and the classification results are generated according to the classification standards of subordinate relationship, corresponding relationship and no relationship. According to the relationship classification results, a relationship extraction model is generated. By integrating the category labels, text content and spatial information of semantic entity nodes, as well as the logical associations between entities, structured data for single semantic relationship parsing of relay protection settings is output, including entity categories, entity contents and logical associations.
9. A method for extracting relay protection setting value information of a power system according to claim 8, characterized in that: The converting the content and logical relationship of the semantic entity into structured data includes: The output of structured data is in the following format: Convert semantic entity content and logical relationships into JSON format files. The JSON format contains entity content, labels, and hierarchical relationships expressed in the form of key-value pairs. For tabular display scenarios, the semantic entity content and association relationships are generated into an XML tree structure data file for structured representation of the semantic entity content and association relationships.
10. A device for extracting information from relay protection setting values in a power system, characterized in that: include: A text detection module is used to perform text detection on an image containing a target relay protection setting list using a text detection model to obtain a text area; wherein the text area includes at least one text box and corresponding position information; A text recognition module is used to perform text recognition on a text area using a text recognition model to obtain text content; A semantic entity recognition module, used to identify semantic entities using a semantic entity recognition model according to the text content and the corresponding text area; The relation extraction module is used to infer the relationship between semantic entities and generate structured associated data based on the relationship between semantic entities; The format of the structured association data includes the category, content, spatial location information and inferred logical association of the semantic entity.
Citation Information
Patent Citations
Relay protection information standardization processing method based on knowledge graph
CN113934865A
Relay protection setting value online checking method
CN116598990A
Method, device and equipment for automatically extracting and comparing constant value information of relay protection device
CN118132620A
Cited By
Visual setting calculation system and method thereof
CN120494072A
A visual tuning calculation system and method thereof
CN120494072B
Relay protection setting value analysis and verification method and system
CN120611316A
A method and system for relay protection setting value analysis and verification
CN120611316B
Power plant relay protection setting value checking method and system fused with deep learning
CN120851903A