Unstructured file identification method based on primitive identification
Through the unstructured file recognition method based on element recognition, the problems of element and label text recognition and semantic understanding in the electrical wiring diagram are solved, and the accurate positioning and maintenance of electrical equipment are achieved.
Patent Information
- Application Number
- CN202510217549.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to effectively identify the essays and label texts in electrical wiring diagrams, especially label texts with different fonts and styles, and it is difficult to understand semantic information, resulting in difficulty in positioning and maintaining electrical equipment.
Unstructured file recognition methods based on element recognition are adopted, including preprocessing, element recognition, text recognition and electrical element association analysis, and feature extraction and text area detection are used to combine template matching and regular expressions for electrical element association.
It realizes accurate identification and semantic understanding of the elements and texts in the electrical wiring diagram, avoids overlap and interference of labeled text, and can directly correspond to actual electrical equipment, making it convenient for positioning and maintenance.
Smart Images

Figure CN120260065A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electrical unstructured document recognition, and in particular to an unstructured document recognition method based on graphic element recognition. Background Art
[0002] Electrical unstructured document recognition generally constructs a benchmark library based on SVG graphics, converts vector graphics and bitmaps such as PDF / CAD and Visio into SVG graphics, and realizes the intelligent recognition of graphic power components and spatial topological structures. For an electrical wiring diagram, which is a representation of the layout of electrical equipment in a substation, to vectorize the electrical wiring diagram, it is first necessary to accurately identify and locate the electrical equipment contained in the drawing. Electrical equipment exists in the form of graphic element symbols in the electrical wiring drawing, and they are composed of some simple geometric elements, so it can be completed by the method of object detection. The marked text information in engineering drawings is an important part of drawing reading and understanding. In electrical wiring drawings, the marked text is generally divided into two categories: one is used to describe the detailed information of electrical graphic elements such as numbers, so that the graphic element symbols can be directly corresponding to the electrical equipment in actual work, facilitating positioning and maintenance. The font size of this type of marked text is usually small; the other is to describe the global information of the drawing, such as the drafter of the drawing, the location area, and the bus voltage level, etc. The font size of this type of marked text is usually large. The content of the marked text is composed of Chinese characters, letters, and numbers, and may also contain a small number of Greek letters. Electrical wiring drawings in different places may use texts with different fonts and styles for marking.
[0003] Unstructured document recognition is to detect and extract electrical graphic elements and electrical marked text in electrical wiring drawings. However, in the process of vectorizing electrical wiring drawings, it is not enough to simply identify the electrical graphic elements and marked text contained therein, because they are all independent individuals and do not contain any available semantic information, while semantic information is the most crucial part of understanding a drawing. Semantic information includes two parts: the association relationship between graphic elements and marked text, and the connection relationship between graphic elements. The traditional method is to directly search and match the surrounding marked text or connection line segments according to the positional relationship between electrical elements and reproduce them in the vector drawing. This method is relatively rough and time-consuming, and it is difficult to achieve good results.
[0004] There is an urgent need for an unstructured document recognition method that can obtain the graphic elements and text of an electrical wiring diagram through graphic element recognition and text recognition, and can solve the problem of different fonts and styles of marked text, understand semantic information, so that the graphic element symbols can be directly corresponding to the electrical equipment in actual work, facilitating positioning and maintenance. Summary of the Invention
[0005] To overcome the above-mentioned disadvantages of the prior art, the present invention proposes an unstructured document recognition method based on graphic element recognition, which can obtain graphic elements and text of an electrical wiring diagram through graphic element recognition and text recognition, and can solve the problem of different fonts and styles of marked text, understand semantic information, and directly correspond graphic element symbols to electrical equipment in actual work, facilitating positioning and maintenance.
[0006] An unstructured document recognition method based on graphic element recognition includes the following steps:
[0007] (1) Preprocessing, which is used to crop, scale an electrical wiring drawing to a required size and perform grayscale processing;
[0008] (2) Graphic element recognition, which uses a YOLO electrical graphic element detection algorithm integrating an attention mechanism to construct a feature extraction network of an electrical graphic element detection model and extract graphic element features;
[0009] (3) Text recognition, including the following steps: removing graphic elements, detecting marked text regions, constructing text candidate boxes, segmenting and adjusting the boundaries of marked text, recognizing the content of marked text, and filtering the results of electrical text recognition;
[0010] Detecting marked text regions is used to obtain feature maps of different scales, and texture information of text of different sizes in the input drawing can be analyzed and obtained from these feature maps;
[0011] Constructing text candidate boxes directly uses a vertical anchor box regression mechanism in the feature map to detect a series of text candidate regions of small scales. This vertical anchor box regression mechanism forms several candidate solutions through several types of anchor boxes, and outputs many text candidate boxes with a fixed width and different heights by jointly predicting the text or non-text scores of each candidate solution; through a text line construction algorithm, these text candidate regions are connected in sequence to obtain text lines; through a merging and deduplication algorithm, text regions with relatively independent semantics are obtained;
[0012] Segmenting and adjusting the boundaries of marked text, specifically using a vertical projection method to segment connected marked text to obtain marked text boxes;
[0013] Recognizing the content of marked text, cropping text regions one by one and sending them into a CRNN model to recognize their marked content and identify the text content;
[0014] Filtering the results of electrical text recognition: designing an electrical text recognition result cleaning algorithm through analyzing and summarizing the characteristics of electrical marked text itself, performing filtering and cleaning operations, removing illegal characters or even the entire marked text in the marked text, and finally obtaining correct text recognition results.
[0015] (4) Analysis of the association relationship of electrical elements, including the following steps: extraction of frame elements, division of primitive tuple regions, template matching, and association between electrical primitives and annotation texts;
[0016] The extraction of frame elements is to divide the drawing into several sub - drawings according to the positions of the frame elements; the frame elements refer to the key elements that can affect the overall layout of the electrical wiring drawing, and the frame elements include busbars, transformers, and instrument transformers;
[0017] The division of primitive tuple regions is to design an algorithm to obtain the primitive tuple regions inside the sub - drawing according to these sub - drawings and combined with the results of primitive recognition;
[0018] Template matching is to match predefined templates by analyzing the composition and distribution characteristics of each primitive tuple region, and parse the sub - drawing into a combination of several templates;
[0019] Association between electrical primitives and annotation texts: obtain the set of annotation texts around the primitive tuple through range search, and then design a regular expression in combination with the inherent characteristics of the numbers of core electrical primitives to screen out the numbers of core primitives from the set of annotation texts, fill the placeholders of the primitive numbers in the primitive tuple template file according to their distribution, and generate the numbers of auxiliary equipment at the same time to realize the association between electrical primitives and annotation texts.
[0020] Furthermore, in the primitive recognition step, the fusion attention mechanism specifically uses the channel attention mechanism to implement SE - Net, and the core part of SE - Net consists of three operations: squeeze, excitation, and attention.
[0021] Furthermore, the electrical primitive detection model is trained using a multi - task loss function, and the loss function consists of confidence loss, classification loss, and bounding box regression loss.
[0022] Furthermore, the CTPN model is used for the detection of the annotation text region. This model uses VGG - 16 to extract the image features of the input image; each convolutional model of VGG - 16 includes a convolutional kernel, and the convolutional kernel is connected to a max - pooling layer. The feature map obtained by the max - pooling layer is sent to a bidirectional LSTM to continue learning the sequence features of the input image. The bidirectional LSTM is connected to a fully - connected layer FC on the previous layer to output the parameters to be predicted, and the parameters include foreground score, background score, position prediction information, and horizontal correction amount.
[0023] Furthermore, the anchor box has a fixed width, which is 16 pixels by default, and 12 different heights are set to form 12 candidate solutions.
[0024] Further, the text line construction algorithm is as follows: Calculate the horizontal distance Horiziontal_Gap, the vertical overlap threshold MIN_V_OVERLAPS, and the text candidate regions in the vertical direction between two text candidate regions; if the horizontal distance between two text boxes is less than Horiziontal_Gap and the IoU in the vertical direction is greater than MIN_V_OVERLAPS, then it is considered that these two text boxes can be connected.
[0025] Further, for the merging and deduplication algorithm of two text candidate regions, if there is an intersection between the two text candidate regions and the IoU in the vertical direction is greater than a certain threshold, it is considered that the two text candidate regions can be merged. Then, the original two text candidate regions are deleted, and a new text box generated by merging these two text candidate regions is added, and this operation is recursively repeated until no text box can be merged.
[0026] Further, the vertical projection method is as follows: Convert the image into a binary image according to a threshold. For this binary image, calculate the number of white pixels in each column of pixels of the image to form a statistical histogram. According to this histogram, the regions with characters and the blank regions without characters in the image can be obtained. Then, based on a certain threshold, if the length of a certain continuous blank region exceeds this threshold, it is considered that segmentation is necessary here.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] The present invention uses character recognition to eliminate graphic elements, avoiding interference from electrical graphic symbols in detection. Then, it successively performs detection of labeled text regions, construction of text candidate boxes, segmentation and boundary adjustment of labeled text, and recognition of labeled text content, avoiding detection overlap and interference from other types of symbols or connection lines, obtaining labeled text boxes without redundant regions, and finally obtaining text regions with relatively independent semantics. Due to the segmentation and boundary adjustment of labeled text, the problem that multiple independent labeled texts may be processed into a complete region, affecting subsequent processing of labeled information, is avoided. Filtering the electrical text recognition results can avoid the problem that a small number of labeled text detection results are affected by other symbols or connection lines, resulting in the directly recognized labeled content being doped with redundant characters.
[0029] In view of the modularization of power system diagrams by these "fixed combinations", the present invention can adopt a vectorization mapping strategy based on template matching. It pre-summarizes and defines common modular combinations in electrical wiring diagrams, defines these combinations as "templates", and represents in detail the association relationships between various electrical elements in the templates to form a structured file. The file contains the relative positions of all electrical elements that make up the electrical element group, the connecting line segments between the elements, the annotation text placeholders corresponding to the elements, and the connection point information of the element group with other structures. In this way, the problem of analyzing the complex association relationships between elements in electrical wiring diagrams is transformed into a module matching problem for local parts of the diagram. Subsequently, by analyzing the connection relationships between modules, the analysis of the association relationships of all electrical elements within the entire diagram can be completed. Brief Description of the Drawings
[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments.
[0031] Figure 1 It is the overall flowchart of a method for identifying unstructured documents based on graphic element recognition according to an embodiment of the present invention.
[0032] Figure 2 It is the flowchart of detecting the annotation text area for a method for identifying unstructured documents based on graphic element recognition according to an embodiment of the present invention.
[0033] Figure 3 It is a comparison diagram before and after implementing the text line construction algorithm for a method for identifying unstructured documents based on graphic element recognition according to an embodiment of the present invention. (a) is a display diagram with text candidate areas before implementing the text line construction algorithm, and (b) is a display diagram with text candidate areas after implementing the text line construction algorithm.
[0034] Figure 4 It is a state diagram of combining sub-diagrams into several templates for a method for identifying unstructured documents based on graphic element recognition according to an embodiment of the present invention. Detailed Embodiment
[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0036] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0037] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations. In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0038] As Figure 1 shown, an unstructured document recognition method based on primitive recognition in an embodiment of the present invention includes the following steps:
[0039] (1) Preprocessing, which is used to crop, scale an electrical wiring diagram to the required size and perform grayscale processing;
[0040] (2) Primitive recognition. A feature extraction network of an electrical primitive detection model is constructed by using a YOLO electrical primitive detection algorithm integrating an attention mechanism to extract primitive features. The integration of the attention mechanism specifically refers to implementing SE-Net with a channel attention mechanism. The core part of SE-Net consists of three operations: squeezing, excitation, and attention. The electrical primitive detection model is trained using a multi-task loss function, and the loss function is composed of a confidence loss, a classification loss, and a bounding box regression loss.
[0041] (3) Text recognition, including the following steps: labeled text area detection, text candidate box construction, labeled text segmentation and boundary adjustment, labeled text content recognition, and filtering of electrical text recognition results.
[0042] Labeled text region detection is used to obtain feature maps of different scales, in which the texture information of texts of different sizes in the input drawing can be analyzed and obtained. The CTPN model is adopted for labeled text region detection. This model uses VGG-16 to extract the image features of the input picture. Each convolutional model of VGG-16 includes a convolutional kernel, and the convolutional kernel is connected to a max pooling layer. The feature map obtained by the max pooling layer is sent to a bidirectional LSTM to continue learning the sequence features of the input picture. The bidirectional LSTM is connected to a fully connected layer FC on the previous layer to output the parameters to be predicted, and the parameters include foreground score, background score, position prediction information, and horizontal correction amount.
[0043] As Figure 2 shown, the preprocessed input electrical wiring drawing in step (1) is scaled to 1920*1920 pixels as the input picture and input into the model. First, the first four convolutional layers of VGG-16 are used to obtain a feature map with a width and height of 1 / 8 of the input picture size, that is, 240*240. Subsequently, three max pooling layers with sizes of 2*2, 3*3, and 4*4 are used, and the feature map is pooled with strides of 2, 3, and 4 respectively. Finally, feature maps with widths and heights of 120*120, 80*80, and 60*60 can be obtained. After obtaining the three feature maps, the subsequent operations are the same as those of the general CTPN. The features are sent to a bidirectional LSTM to continue learning the sequence features of the picture, and finally, a fully connected layer FC is connected to output the parameters we want to predict, and the parameters include foreground score, background score, position prediction information, and horizontal correction amount.
[0044] Using a three-layer feature pyramid structure, the model needs to predict the desired output on the three obtained feature maps. 12 small anchor boxes with a fixed width of 16 pixels and different heights are used to let the model predict the text regions in the electrical wiring drawing.
[0045] Among them, for the feature map with a larger size, its corresponding downsampling multiple is relatively low, which can be used to locate relatively small text regions. Therefore, anchor boxes with heights of 11, 16, 22, and 33 pixels are preset for it; for the feature map with a medium size, which can be used to detect medium-sized text regions, anchor boxes with heights of 48, 68, 87, and 97 pixels are preset for it; for the feature map with the smallest size, which is used to detect larger text regions in the drawing, anchor boxes with heights of 111, 139, 198, and 283 pixels are preset for it.
[0046] Finally, the CTPN model integrates the prediction results of the three feature maps as the output, which respectively include the foreground score, background score, position prediction information of these 12 anchor boxes, and the horizontal boundary correction parameter of the anchor box, i.e., the horizontal correction amount. When training the CTPN model, the following loss function is used to optimize the parameters. The loss consists of three parts, namely the classification loss Ls for judging foreground and background, the regression loss Lv for predicting the position of the anchor box, and the regression loss Lo for predicting the horizontal boundary correction amount of the anchor box. λ1 and λ2 are used to balance the losses of each task.
[0047]
[0048] Among them, the classification loss uses Cross Entropy Loss, and the regression loss uses Smooth L1 Loss. Assuming that the position information corresponding to the anchor box is [x a , y a , w a , h a , indicating its center point coordinates and width and height, and the text box information corresponding to the labeled GroundTruth is [x, y, w, h], then the prediction data [t x , t y , t w , t h of the CTPN model for each anchor box and its true value label [t * , t * , t * , t * are calculated by the formula.
[0049]
[0050] According to the predicted value and the true value, the corresponding functions in are used to calculate the loss. In Ls, s i represents the probability that the predicted anchor box contains text, and s * i represents its true value label, which is obtained by calculating the IoU between the anchor box and the labeled text box. When the IoU value is greater than a certain threshold (such as 0.7), the true value label takes the value of 1, otherwise it takes the value of 0; in Lv, v = [ty, th] represents the vertical position information of the anchor box predicted by the network, and v* = [t*, t*] represents its true value label; in Lo, o = tx represents the horizontal position information of the anchor box predicted by the network, and o* = t*x represents its true value label.
[0051] Compared with the standard CTPN model, the text region detection method used in the embodiments of the present invention aims to obtain feature maps of different scales through pooling operations. In these feature maps, texture information of texts of different sizes in the input drawing can be analyzed and obtained, which helps to subsequently obtain text regions of any size in the picture. The effective extraction of text regions will provide a very good basis for the subsequent recognition of the labeled text content.
[0052] The text candidate box is constructed to directly use the vertical anchor box regression mechanism in the feature map to detect a series of text candidate regions of small scales. The vertical anchor box regression mechanism forms several candidate solutions through several types of anchor boxes, and outputs many small text candidate boxes with a fixed width and different heights by jointly predicting the text or non-text scores of each candidate solution; through the text line construction algorithm, these text candidate regions are connected in sequence to obtain text lines; and semantic relatively independent text regions are obtained through the merging and deduplication algorithm.
[0053] The text line construction algorithm is as follows: calculate the horizontal distance Horiziontal_Gap, the vertical overlap threshold MIN_V_OVERLAPS, and the text candidate regions in the vertical direction between two text candidate regions; if the horizontal distance between two text boxes is less than Horiziontal_Gap and the IoU in the vertical direction is greater than MIN_V_OVERLAPS, then it is considered that these two text boxes can be connected. As Figure 3 shown, (a) is a display diagram with text candidate regions before implementing the text line construction algorithm, and (b) is a display diagram with text candidate regions after implementing the text line construction algorithm.
[0054] For the merging and deduplication algorithm, for two text candidate regions, if there is an overlapping part (i.e., an intersection) between the two text candidate regions and the IoU in the vertical direction is greater than a certain threshold (here, 0.7 is selected), then it is considered that the two text candidate regions can be merged (Connectable). Thus, the original two text candidate regions are deleted, and a new text box generated by merging these two text candidate regions is added. This operation is recursively repeated until no text box can be merged.
[0055] Labeled text segmentation and boundary adjustment are specifically to use the vertical projection method to segment the connected labeled text to obtain labeled text boxes; the vertical projection method is: convert the image into a binary image according to a threshold. For this binary image, calculate the number of white pixels in each column of pixels of the image to form a statistical histogram. According to this histogram, the regions with characters and the blank regions without characters in the image can be obtained. Then, based on a certain threshold (1 / 3 of the text box height is used as the segmentation threshold), if the length of a certain continuous blank region exceeds this threshold, it is considered that segmentation is necessary here.
[0056] For the recognition of the marked text content, the text area is cut one by one and sent into the CRNN model to recognize its marked content and identify the text content.
[0057] Filtering of the electrical text recognition result: By analyzing and summarizing the characteristics of the electrical marked text itself, an algorithm for cleaning the electrical text recognition result is designed to perform filtering and cleaning operations, removing illegal characters or even the entire mark in the marked text, and finally obtaining the correct text recognition result. The algorithm for cleaning the electrical text recognition result can be set according to needs and experience. For example, it can be set that the marked text is always no less than two characters and does not contain many symbols such as "|", "]", "→", etc.
[0058] (4) Analysis of the association relationship of electrical elements, including the following steps: extraction of frame elements, division of graphic element group areas, template matching, and association between electrical graphic elements and marked text.
[0059] The extraction of frame elements divides the drawing into several sub-drawings according to the positions of the frame elements; the frame elements refer to the key elements that can affect the overall layout of the electrical wiring drawing, and the frame elements include busbars, transformers, and instrument transformers.
[0060] The division of graphic element group areas designs an algorithm to obtain the graphic element group areas inside the sub-drawing according to these sub-drawings and in combination with the results of graphic element recognition.
[0061] Template matching analyzes the composition and distribution characteristics of each graphic element group area, matches predefined templates, and parses the sub-drawing into a combination of several templates; as Figure 4 shown.
[0062] Association between electrical graphic elements and marked text: Obtain the set of marked text around the graphic element group through range search, and then design a regular expression in combination with the inherent characteristics of the numbers of the core electrical graphic elements to screen out the numbers of the core graphic elements from the set of marked text, fill the placeholders of the graphic element numbers in the graphic element group template file according to their distribution, and generate the numbers of auxiliary equipment at the same time to realize the association between electrical graphic elements and marked text.
[0063] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for identifying unstructured documents based on graphic element recognition, characterized in that, Including the following steps: (1) Preprocessing, which is used to crop, scale the electrical wiring drawing to the required size and perform grayscale processing; (2) Graphic element recognition, using the YOLO electrical graphic element detection algorithm with a fusion attention mechanism to construct the feature extraction network of the electrical graphic element detection model and extract graphic element features; (3) Text recognition, including the following steps: removing graphic elements, detecting labeled text regions, constructing text candidate boxes, segmenting and adjusting the boundaries of labeled text, recognizing the content of labeled text, and filtering the results of electrical text recognition; Detecting the labeled text region is used to obtain feature maps of different scales, and the texture information of texts of different sizes in the input drawing can be analyzed and obtained from these feature maps; Constructing text candidate boxes is to directly use the vertical anchor box regression mechanism in the feature map to detect a series of text candidate regions of small scales. This vertical anchor box regression mechanism forms several candidate solutions through several types of anchor boxes, and outputs many text candidate boxes with a fixed width and different heights by jointly predicting the text or non-text scores of each candidate solution; through the text line construction algorithm, these text candidate regions are connected in sequence to obtain text lines; through the merging and deduplication algorithm, text regions with relatively independent semantics are obtained; Segmenting and adjusting the boundaries of labeled text, specifically using the vertical projection method to segment the connected labeled text to obtain labeled text boxes; Recognizing the content of labeled text, cropping the text region one by one and sending it into the CRNN model to recognize its labeled content and identify the text content; Filtering the results of electrical text recognition: designing an electrical text recognition result cleaning algorithm through analyzing and summarizing the characteristics of electrical labeled text, performing filtering and cleaning operations, removing illegal characters or even the entire label in the labeled text, and finally obtaining the correct text recognition result. (4) Analyzing the association relationship of electrical elements, including the following steps: extracting frame elements, dividing the graphic element group region, template matching, and associating electrical graphic elements with labeled text; Extracting frame elements is to divide the drawing into several sub-drawings according to the positions of the frame elements; the frame elements refer to the key elements that can affect the overall layout of the electrical wiring drawing, and the frame elements include busbars, transformers, and instrument transformers; Dividing the graphic element group region is to design an algorithm to obtain the graphic element group region inside the sub-drawing based on these sub-drawings and combined with the results of graphic element recognition; Template matching is to match the predefined templates by analyzing the composition and distribution characteristics of each graphic element group region and parse the sub-drawing into a combination of several templates; Associating electrical graphic elements with labeled text: obtaining the set of labeled text around the graphic element group through range search, and then designing a regular expression in combination with the inherent characteristics of the numbers of core electrical graphic elements to screen out the numbers of core graphic elements from the set of labeled text, filling the placeholders of the graphic element numbers in the graphic element group template file according to their distribution, and generating the numbers of auxiliary equipment at the same time to realize the association between electrical graphic elements and labeled text.
2. The unstructured document recognition method based on primitive recognition according to claim 1, wherein In the graphic element recognition step, the specific implementation of the fusion attention mechanism is the channel attention mechanism to implement SE-Net. The core part of SE-Net consists of three operations: squeezing, excitation, and attention.
3. The unstructured document recognition method based on graphic primitive recognition according to claim 1, characterized in that The electrical graphic element detection model is trained using a multi-task loss function, which consists of a confidence loss, a classification loss, and a bounding box regression loss.
4. A method for identifying unstructured documents based on primitive recognition according to claim 1, characterized in that The labeled text region detection uses the CTPN model, which uses VGG-16 to extract the image features of the input image; Each convolutional model of VGG-16 includes a convolutional kernel, which is connected to a max pooling layer. The feature map obtained by the max pooling layer is fed into a bidirectional LSTM to continue learning the sequence features of the input image. The bidirectional LSTM is connected to a fully connected layer FC at the previous layer to output the parameters to be predicted, which include foreground scores, background scores, position prediction information, and horizontal correction amounts.
5. A method for recognizing unstructured documents based on primitive recognition according to claim 1, characterized in that The anchor box is preset with a fixed width, which is 16 pixels by default, and 12 different heights are set to form 12 candidate solutions.
6. The unstructured document recognition method based on primitive recognition according to claim 1, characterized in that, The text line construction algorithm is as follows: Calculate the horizontal distance Horiziontal_Gap, the vertical overlap threshold MIN_V_OVERLAPS, and the text candidate regions in the vertical direction between two text candidate regions. If the horizontal distance between two text boxes is less than Horiziontal_Gap and the IoU in the vertical direction is greater than MIN_V_OVERLAPS, then these two text boxes are considered connectable.
7. A method for identifying unstructured documents based on primitive recognition according to claim 1, characterized in that For the merging and deduplication algorithm of two text candidate regions, if there is an intersection between the two text candidate regions and the IoU in the vertical direction is greater than a certain threshold, then the two text candidate regions are considered mergable. Then, the original two text candidate regions are deleted, and a new text box generated by merging these two text candidate regions is added. This operation is recursively repeated until no text boxes can be merged.
8. A method for identifying unstructured documents based on primitive recognition according to claim 1, characterized in that, The vertical projection method is as follows: Convert the image into a binary image according to a threshold. For this binary image, calculate the number of white pixels in each column of pixels of the image to form a statistical histogram. Based on this histogram, the regions with characters and the blank regions without characters in the image can be obtained. Then, based on a certain threshold, if the length of a certain continuous blank region exceeds the threshold, it is considered necessary to perform segmentation here.
Citation Information
Cited By
Optical character recognition method and system for mixing handwritten form and printed form
CN120976928A
Power grid wiring diagram composite primitive construction method, system, equipment and medium
CN121074417A
Power grid wiring diagram composite graph element construction method, system, device and medium
CN121074417B
An unstructured CAD table extraction method and system
CN122531053A