A structure perception and visual repair method for power operation tickets

CN122821575APending Publication Date: 2026-09-25YANTAI HAIYI SOFTWARE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611010501.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

随着信息技术的发展,利用AI技术实现电力作业票据的智能识别已经成为趋势,然而在实际生产中要可靠应用仍面临诸多问题

Benefits of technology

(1)本发明提出一种基于OCR文本、版面结构和bbox坐标的候选证据构建方法,结合OCR文本、版面块类型、阅读顺序、字段上下文和bbox坐标,对多类型电力作业票据中的时间、开关操作、设备名称、线路名称和杆号位置等信息进行整理和筛选,并保留其来源页码、空间位置和所属版面区域,为后续字段关系判断和复核区域确定提供依据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821575A_ABST
    Figure CN122821575A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of bill repair, and particularly relates to a structure perception and visual repair method for electric power operation bills. The method takes multiple types of electric power operation bills such as work tickets, switching operation tickets and repair tickets as processing objects, and realizes accurate identification of key information of the electric power operation bills based on a cheap AI infrastructure composed of OCR and a low-parameter multi-modal large model, and by comprehensively using technologies such as OCR analysis, layout structure identification, table structure identification, spatial position modeling, local area cropping and page cutting review.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of invoice repair technology, specifically relating to a structural perception and visual repair method for power operation invoices. Background Technology

[0002] Power operations, including on-site repairs and switching operations, generate a large number of different types of work tickets. The key business information recorded on these tickets, such as the start and end times, execution times, and switch operation details, is crucial for safety management, process traceability, statistical analysis, and automated archiving. With the development of information technology, using AI to achieve intelligent recognition of power work tickets has become a trend; however, reliable application in actual production still faces many challenges.

[0003] First, power operation tickets are diverse in type, with significant differences in layout, field names, key information locations, and information presentation formats, making it difficult to achieve unified parsing using a fixed template. Furthermore, tickets often contain handwritten content, low-quality scans, cross-line information, partial occlusion, and table line interference. Consequently, while traditional OCR (Optical Character Recognition) technology can identify some content, it cannot reliably and accurately identify key information in power operation tickets. Second, although multimodal large-scale models can effectively recognize images and handwritten content, due to security considerations and computing power limitations, the actual production solution typically uses a privately deployed, quantized, low-parameter version of the multimodal large-scale model, whose performance cannot meet the requirement of accurately recognizing complete power operation tickets. Therefore, achieving accurate power operation ticket recognition based on locally deployed low-parameter multimodal large-scale models and OCR under limited computing power is one of the core requirements for the practical application of AI in power business.

[0004] To address the aforementioned problems, this invention proposes a structure perception and visual repair method for power operation tickets. Summary of the Invention

[0005] To overcome the problems in the prior art, this invention proposes a structural perception and visual repair method for power operation tickets.

[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: This invention provides a structure perception and visual repair method for power operation tickets, comprising the following steps: Step 100: Receive the power work ticket documents and standardize them; Step 200: Perform OCR recognition on the standardized power operation ticket document to obtain the OCR recognition result, and set extraction rules to clean the OCR recognition result and organize candidate evidence; Step 300: Input the OCR recognition results, candidate evidence and extraction rules into the multimodal large model for structured extraction, perform integrity checks, and perform visual repair based on the integrity check results.

[0007] Furthermore, in step 100, the power work ticket document includes the original document of the power work ticket to be identified and a list of all switching equipment related to the power work ticket.

[0008] Furthermore, in step 200, the extraction rules are set to clean the OCR recognition results and organize candidate evidence, including: Clean the OCR recognition results, that is, remove the outer fields and model configuration fields that are not involved in business extraction; After cleaning, the OCR recognition results are compressed, retaining only the page and text block information that is valuable for business extraction, forming OCR evidence; the page and text block information that is valuable for business extraction includes page width, page height, text block content, text block type, text block coordinates, text block number, and reading order; Based on OCR evidence, each text block is evaluated for its candidate value according to business semantics, and candidate evidence is organized. The business semantics include time source semantics, action word semantics, device keyword semantics, and disabled template sentences.

[0009] Furthermore, in step 300, the integrity check includes: whether there are start time and end time fields, whether a valid switch operation event is extracted, and whether the preliminary extraction result is empty.

[0010] Furthermore, step 300 also includes: When the initial extraction results are empty, or key fields are missing and no valid candidates can be formed, then structure-aware page segmentation and local re-identification are performed. When the page contains a table area with a preset ratio, and the initial extraction results are missing key fields, the horizontal lines of the complex table are weakened and the OCR recognition is performed again. If the initial extraction results include fields that can be further processed, but the field content is incomplete, then field region positioning, partial cropping, and visual verification will be performed.

[0011] Furthermore, missing key fields include: the initial extraction results are missing any key field from the start time, end time, or switch operation event, i.e., the corresponding key field does not exist or is empty.

[0012] Furthermore, structure-aware page segmentation and local re-identification are performed, including: The entire page image is horizontally divided into image blocks for local recognition. The following principles are followed when dividing the image: First, keep the page width unchanged, and divide the image blocks horizontally along the page height so that each image block retains complete horizontal layout information. After each segmentation, the OCR model is called again, and the OCR recognition results are processed, candidate evidence is constructed, and multimodal large model structure extraction is performed again.

[0013] Furthermore, the weakening of horizontal lines in complex tables and re-OCR recognition include: By breaking and weakening the horizontal lines, the interference of the horizontal table lines on text recognition is reduced. After processing, OCR recognition is performed again and information is extracted.

[0014] Step 400: Field region location, partial cropping, and visual verification, including: Combining OCR layout blocks, table areas, and cell candidate information, a multimodal large model is used to match and determine the source of fields. The source of fields includes ordinary text blocks, table blocks, and specific cells. The table area and cell candidate information are obtained through a table recognition model. For fields not sourced from tables, the corresponding text block's bounding box is retained as the candidate region; for fields sourced from tables, fine-grained cell bounding boxes are used as the candidate region. Candidate cropping regions are generated based on candidate regions, and the images of the local cropping regions are used as inputs to the visual kernel of the multimodal large model. If more complete year, month, day, hour, and minute information can be identified in the cropped area, the initial extraction result is corrected using the visual verification result; if only part of the time can be identified in the cropped image, only the identifiable content is retained. The field that has undergone local visual verification and correction records the corresponding repair status, indicating that the time field has been verified by cropping the image.

[0015] The process of generating a local cropping region based on the candidate region includes: If the time field originates from a non-table range, the candidate range is used as the candidate clipping range. If the time field originates from a table range and contains fine-grained cell bounding boxes, then the candidate range is used as the candidate clipping range. If the time field originates from a table range but no reliable cell bbox exists, then either the table block bbox or the field block_bbox is selected as the candidate clipping region.

[0016] Compared with the prior art, the present invention has the following technical effects: (1) This invention proposes a candidate evidence construction method based on OCR text, page structure and bbox coordinates. By combining OCR text, page block type, reading order, field context and bbox coordinates, information such as time, switch operation, equipment name, line name and pole number position in multiple types of power operation tickets is sorted and filtered, and the source page number, spatial position and page area are retained, providing a basis for subsequent field relationship judgment and review area determination.

[0017] (2) This invention proposes a structure-aware field location method that is triggered on demand. It determines whether a field is located in a table area based on the candidate field bbox, the layout block type, and the field source region. For fields within a table, the table structure is parsed on demand to obtain cell boundaries, row and column relationships, and relationships between adjacent cells. For non-table fields, the verification area is determined directly based on the OCR bbox, layout block, and context, thus balancing processing efficiency and field location accuracy.

[0018] (3) This invention proposes a segmented visual repair method for OCR missing data and table interference, which triggers the repair process when fields are missing, recognition is incomplete, confidence is low, or spatial relationships are abnormal. For fields with clear locations, local cropping is performed to focus the multimodal large model on the target area; for fields with unclear locations, dense page text, or poor quality, page segmentation is performed; for complex table interference scenarios, horizontal table lines are further weakened and broken, and OCR recognition, structured extraction, and visual verification are re-executed to supplement the time fields, handwritten content, device information, and switch operation content that were missed in the recognition. Attached Figure Description

[0019] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram of the process of the present invention; Figure 2 The first page of the original invoice; Figure 3 This is the second page of the original invoice; Figure 4 This is a partial OCR recognition result of the second page of the original invoice; Figure 5 A clear image of the OCR recognition result for the second page of the original invoice; Figure 6 This is a partial cropped image of the end time on the second page of the original invoice; Figure 7 To output structured extraction results based on OCR evidence for multimodal large models; Figure 8 Image of the original invoice; Figure 9 This is a partial OCR result image of the original document; Figure 10 This is a table showing the results of processing the original invoices. Figure 11 This is a partial OCR result image after the horizontal lines in the table have been processed. Figure 12 Examples of field positioning, time clipping, and result verification; Figure 13 Extracting OCR result Prompt templates for large multimodal models; Figure 14 Match the Prompt template to the switch information; Figure 15 A Prompt template that matches the layout analysis results with the coordinate frame; Figure 16 For structure-aware segmentation Prompt templates; Figure 17 This is a Prompt template for visual repair. Detailed Implementation

[0021] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the specific implementation methods, structures, features, and effects of the technical solutions proposed according to the present invention are described in detail below with reference to the accompanying drawings and preferred embodiments. Specific features, structures, or characteristics in one or more embodiments may be combined in any suitable form. Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0022] This embodiment takes various types of power operation tickets, such as work orders, switching operation tickets, and emergency repair tickets, as the processing objects. Based on OCR and a low-cost AI infrastructure composed of low-parameter multimodal large models, it comprehensively uses technologies such as OCR parsing, page structure recognition, table structure recognition, spatial location modeling, local area cropping, and page segmentation verification to achieve accurate recognition of key information in power operation tickets.

[0023] Reference Figures 1-17This invention first utilizes OCR to identify page text, layout blocks, table regions, field coordinates, and contextual relationships to construct OCR evidence, and then performs preliminary structured extraction using a multimodal large model. Subsequently, the extraction results undergo a completeness check, and corresponding repair paths are selected based on the number of missing key fields, the availability of candidate evidence, and the table structure. For cases with missing key fields, insufficient candidate evidence, or severe table interference, further page segmentation or table line breaking and weakening are performed before re-performing OCR recognition and structured extraction. For the extracted time field, structure-aware localization is performed by combining OCR layout blocks, table regions, cell coordinates, and field source information. Local areas are cropped based on candidate bounding boxes, and the multimodal large model is used to visually verify the time content in the cropped image, thereby improving the completeness and accuracy of the extracted time field and other key information.

[0024] In this embodiment, a method for structure perception and visual repair of power work tickets is provided, including the following steps: Step 100: Receive the power work ticket documents and standardize them; Step 200: Perform OCR recognition on the standardized power operation ticket document to obtain the OCR recognition result, and set extraction rules to clean the OCR recognition result and organize candidate evidence; Step 300: Input the OCR recognition results, candidate evidence and extraction rules into the multimodal large model for structured extraction, perform integrity checks, and perform visual repair based on the integrity check results.

[0025] The following is a detailed explanation of each of the above steps: Step 100: Receive the power work ticket documents and standardize them.

[0026] Existing power business invoices come from diverse sources and formats, with variations in page structure, resolution, page numbering, and coordinate systems across different file types such as images, PDFs, and Word documents. This leads to inconsistencies in subsequent OCR recognition, field bounding box (bounding box) localization, partial cropping, page segmentation, and result backtracking. Therefore, this embodiment aims to address the issues of standardized processing of multi-format power operation invoices and cross-stage coordinate consistency, providing a unified foundation for subsequent field localization and visual verification.

[0027] As an example, step 100 specifically includes: Step 110: Receive the power work ticket document, which includes the original document of the power work ticket to be identified and a list of all switchgear related to the power work ticket.

[0028] In this embodiment, the power work ticket file consists of two parts: the original document of the power work ticket to be identified, and a list of all switchgear related to the power work ticket. The power work ticket file can be in formats such as Word, PDF, or image files, and requires recognition using OCR and a multimodal large model. The switchgear list serves as a standard set of equipment names in the subsequent switch matching stage, matching the switch operation descriptions extracted from the power work ticket into standardized switch names.

[0029] For example, the OCR in the power work invoice document may misidentify "Tianyi Mingzhu Opening Station 101 Switch" as "Tianyi Mingsuo Opening Station 101 Switch", where "Zhu" is misidentified as "Luo". By matching it with "Tianyi Mingzhu Opening Station 101 Switch" in the switch equipment list, the final output is the standardized result "Tianyi Mingzhu Opening Station 101 Switch".

[0030] Step 120: Convert the power work ticket files into fixed-resolution image files, i.e., standardized power work ticket files.

[0031] This embodiment first converts the input power operation ticket files into image files with a fixed resolution (e.g., 300 DPI), so that subsequent OCR recognition, page segmentation, field cropping, and bbox backfilling are all based on the same type of file.

[0032] If the power work invoice is a PDF file, convert it to an image file; if the power work invoice is a Word document, convert it to a PDF file first, and then to an image file.

[0033] For multi-page power work order documents, after each page has undergone the format conversion as described above, a mapping relationship between the global page number and the original power work order document is established.

[0034] Step 200: Perform OCR recognition on the standardized power operation ticket document to obtain the OCR recognition result, and set extraction rules to clean the OCR recognition result and organize candidate evidence.

[0035] Figures 2-3 For a two-page "Emergency Repair Order for Power Distribution Fault" document that was actually processed, the following is a combination of... Figures 2-3 Explanation: Step 210: Perform OCR recognition on the standardized power operation ticket document to obtain the OCR recognition result.

[0036] The OCR model with layout recognition capability (such as PaddleOCR-VL) is called to recognize the page image. The OCR model outputs the OCR recognition results, such as page size, layout blocks, text block content, reading order, text block number, and corresponding bbox coordinates.

[0037] Figure 3 Some OCR recognition results are as follows Figure 4 As shown. Comparison Figure 3 and Figure 4 It can be seen that the OCR model incorrectly extracted the emergency repair end time from the operation ticket as "January 8, 2025" ( Figure 4 The underlined part needs to be repaired in the subsequent process.

[0038] Step 220: Clean the OCR recognition results and organize candidate evidence.

[0039] Given the varying formats of different power work invoices, it's impossible to accurately extract key information directly from OCR recognition results using a fixed format. This embodiment first cleans, compresses, and organizes candidate evidence in the OCR recognition results, with the following extraction rules: (1) Clean the OCR recognition results, that is, remove the outer fields and model configuration fields that do not participate in business extraction.

[0040] For example: `page_count` (total number of pages, unrelated to invoice business), `markdown_ignore_labels` (Markdown ignored labels, belonging to layout control information), `return_layout_polygon_points` (returns the coordinates of the layout polygons, belonging to geometric metadata), and the `layout ParsingResults` field in the OCR configuration of `model_settings` (outer wrapper field of the OCR layout parsing results). These fields can be retained in the original OCR recognition result file for traceability, but they will not be included as candidate evidence in the business text, nor will they be submitted to the multimodal large model as the basis for judging the content of invoices.

[0041] (2) After cleaning, the OCR recognition results are compressed, and only the page and text block information that is valuable for business extraction is retained to form OCR evidence.

[0042] The page and text block information that is useful for business extraction includes: Page width (page_width) and page height (page_height) are used for subsequent slicing, cropping, and coordinate boundary processing. The text block content (block_content) is used to construct the OCR full text and candidate text. The text block type `block_label` is used to determine whether it belongs to plain text, a heading, or a table area. The text block coordinates, block_bbox, are used for subsequent field positioning and cropping verification; The text block number (block_id) and reading order (block_order) are used to maintain the traceability of evidence and the order of context.

[0043] (3) Based on OCR evidence, make a candidate value judgment on each text block according to the business semantics and complete the candidate evidence sorting.

[0044] Business semantics includes time source semantics, action word semantics, equipment keyword semantics, and disabled template sentences. Among them, time source semantics is used to distinguish different time types such as authorized emergency repair time, work completion time, and filling time; action word semantics is used to identify actual operation actions such as "pull open," "close," and "disconnect"; equipment keyword semantics is used to identify equipment objects such as switches, knife switches, and disconnect switches; disabled template sentences are used to exclude signature instructions, template prompts, and non-actual operation content to avoid these fixed texts interfering with structured extraction.

[0045] Specifically, for the candidate value judgment of each text block, four types of business semantics are used for screening: time source semantics are used to distinguish different types such as authorized repair time, work completion time, and filling time, and time candidates related to the actual start and end of the work are selected as the basis for preliminary start / end time extraction; action word semantics are used to identify actual operation actions such as "pull open", "close", and "disconnect", and equipment keyword semantics are used to identify equipment objects such as switches, knife switches, and disconnect switches. The two are combined to generate switch operation event candidates. At the same time, the specific action type (such as "pull open" corresponding to opening the switch, "close" corresponding to closing the switch) is further mapped to the operation nature field to distinguish between power outage and live operation; and prohibited template sentences are used to exclude non-actual operation content such as signature instructions and template prompts to avoid fixed text interfering with the extraction. The output of the four types of semantic judgments and the text block structure information retained after cleaning together constitute candidate evidence, which is used by the multimodal large model to complete structured extraction under the constraint of OCR evidence.

[0046] Figure 4 The OCR recognition results shown are the results after cleaning and candidate evidence preparation. Figure 5 As shown.

[0047] Step 300: Input the OCR evidence, candidate evidence and extraction rules into the multimodal large model for structured extraction, obtain preliminary extraction results, perform integrity checks, and perform visual repair based on the integrity check results.

[0048] Input the OCR evidence, candidate evidence, and extraction rules obtained in step 220 into the multimodal large model to obtain preliminary extraction results. When the preliminary extraction results have missing time, insufficient switch information, or incomplete structure, perform page segmentation and weaken the horizontal lines of complex tables according to the page size, layout structure, table area ratio, and field missing status, and obtain complete information through multiple operations.

[0049] As an example, step 300 specifically includes: Step 310: Input the OCR evidence, candidate evidence and extraction rules into the multimodal large model for structured extraction to obtain preliminary extraction results. The preliminary extraction results include invoice type, business nature, time field and switch operation event.

[0050] The OCR evidence obtained in steps (1) and (2) of 220, the candidate evidence obtained through business screening in step (3) of 220, and the extraction rules are used as input. The multimodal large model performs structured extraction under the constraints of OCR evidence. The OCR structured extraction prompts used in this stage are as follows: Figure 13 As shown. This stage does not directly rely on fixed field positions, but rather, based on the complete OCR text, it combines evidence such as ticket type candidates, time candidates, device candidates, and switch operation candidates to comprehensively determine the ticket type, business nature, time field, and switch operation event.

[0051] In the process of information extraction from multimodal large models, it is necessary to combine it with power business logic. For example, for time in the context of start-type business, it should be given priority as the start time source; for time in the context of end-type business, it should be given priority as the end time source; switch operation events need to be judged by combining action words and specific equipment information to avoid misidentifying template text as valid switch operation events.

[0052] After extraction is completed, the multimodal large model generates preliminary extraction results, including fields such as ticket type, business nature (whether it is an emergency repair order, operation nature), time field (start time and end time), and switch operation event.

[0053] Because OCR recognition results may exhibit issues such as localized recognition instability, field cross-regional issues, separation of time text and labels, and low confidence levels for time segments, the time fields in the initial extraction results may still be missing or incomplete. Therefore, this initial extraction result serves as the basis for subsequent field localization, region cropping, and visual verification, and is used to further refine key time fields.

[0054] Step 320: When the electricity bill document contains multiple pages, the OCR evidence, candidate evidence and preliminary extraction results of each page need to be merged into a task-level result.

[0055] During the merging process, it is necessary to retain the page source, textual evidence, and location information for each field, and to uniformly categorize them according to field semantics. For example, start-type time sources, end-type time sources, switch operation candidates, and prohibited time fields will be categorized into their respective candidate sets. The merged task-level results can simultaneously support cross-page time judgment, cross-page field verification, and subsequent region clipping correction, avoiding the omission of key business information due to analyzing only single-page content.

[0056] The final multimodal large model outputs structured extraction results based on OCR evidence, as shown in the example below. Figure 7 As shown.

[0057] Step 330: Perform an integrity check on the preliminary extraction results.

[0058] After the initial extraction is completed, a completeness check needs to be performed on the preliminary extraction results to determine whether the results meet the basic conditions for proceeding to subsequent processing steps. The completeness check mainly includes: whether the start time and end time fields exist, whether valid switch operation events have been extracted, and whether the preliminary extraction results are empty. It should be noted that "completeness" at this stage primarily means that the key fields exist or have traceable candidate sources, and does not equate to the field content being completely accurate or formatted correctly.

[0059] When the initial extraction result is empty (some OCR systems will directly recognize large or compact tables as empty), or when key fields are missing and cannot form valid candidates, it is necessary to trigger a page segmentation retry to perform structure-aware page segmentation and local re-identification, as detailed in step 340. Missing key fields and inability to form valid candidates include: first, the initial extraction result lacks any key field from the start time, end time, or on / off operation event, meaning the corresponding field is missing or has an empty value; second, in the OCR processing stage, valid time candidates or on / off operation candidates are not formed, meaning the candidate evidence contains neither business-related text content nor on / off operation text containing device keywords and action words, resulting in the inability to form reliable structured extraction results subsequently.

[0060] When a page contains a complex table structure that causes missing information in the OCR recognition results (i.e., the page has a large table area, and the initial extraction still lacks key fields), the complex table horizontal line processing flow is initiated. This involves weakening the complex table horizontal lines and performing re-OCR recognition, as detailed in step 350. Specifically, based on the OCR layout analysis results, a table structure is considered complex if the table area occupies more than 50% of the total page area, or the table width occupies more than 80% of the page width, or the table height occupies more than 50% of the page height.

[0061] If the initial extraction results already contain fields that can be processed further, but the field content is still incomplete, the process will not proceed to page-level retry. Instead, it will proceed to the field region location, partial cropping, and visual review process to further correct the incomplete fields. "Initial extraction results already contain fields that can be processed further" means that key fields such as start time, end time, or on / off operation events have been extracted, indicating that the fields are not completely missing. "Field content is still incomplete" means that the field exists, but the recognition result is missing some content. For example, a time field may only recognize partial time information (such as "January 2025"), or there may be missing characters. The field itself has been located, but it cannot meet subsequent business requirements.

[0062] Step 340: Structure-aware page segmentation and local re-identification.

[0063] When there are obvious missing results in the initial extraction, the structure-aware page segmentation is performed based on the page size, field missing status, and page layout distribution, splitting the entire page image horizontally into image blocks that are more suitable for local recognition.

[0064] The following principles are followed when segmenting: First, keep the page width unchanged and divide the page horizontally along the height direction so that each image block retains complete horizontal layout information; dynamically determine the number of blocks based on the page size and the number of retries, with smaller pages using fewer blocks and regular pages using different numbers of blocks for each retrieval according to the configuration; allow a certain amount of overlap between blocks to prevent fields from being truncated when they are located at the block boundaries.

[0065] After each segmentation, the OCR model is called again, and the OCR recognition results are processed, candidate evidence is constructed, and multimodal large model structure extraction is performed again.

[0066] Step 350: Weakening of horizontal lines in complex tables and re-OCR recognition.

[0067] When a page contains a large table area, and the initial extraction results still lack key fields, this invention performs complex table horizontal line weakening processing. This processing reduces the interference of horizontal table lines on text recognition by breaking and weakening the horizontal lines. After processing, OCR recognition and structured extraction are performed again.

[0068] like Figures 8-11 As shown, a complex table weakening example is presented. It can be seen that before the weakening process, the OCR recognition result only has a small amount of information, while after the process, the OCR recognizes a lot of useful information such as the start and end times.

[0069] Step 400: Field location, partial cropping, visual repair, and result output.

[0070] After initial extraction under OCR evidence constraints, the multimodal large model identifies specific switching operation events from these candidates and outputs corresponding text content, such as "open 10kV switch A", which is the switching operation text. The extracted switching operation text is then matched against an externally provided list of switching equipment to output standardized switch recognition results.

[0071] For time information, the system combines OCR layout blocks, table areas, cell coordinates, and field source information to perform structure-aware localization of the time field source area. Then, it performs local cropping based on candidate bounding boxes. Finally, a multimodal large model identifies and visually repairs the time content in the cropped area, and outputs key information such as start time, end time, and the name of the matched switch.

[0072] Step 410: Match the extracted switch operation text with the externally provided list of switch devices and output standardized switch recognition results.

[0073] To ensure the accuracy of the extraction results, this invention further matches the switch names in the extracted switch operation text with an external list of switchgear. By splitting and comparing information such as action words, line names, equipment names, pole numbers, serial numbers, switch types, and equipment locations in the switch operation text, it is determined whether they can reliably correspond to a standard device in the external equipment list. Finally, the switch names in the switch operation text are mapped to standard device names.

[0074] During the matching process, action words are mainly used to preserve operational attributes and do not directly participate in the standard equipment name judgment; equipment name matching focuses on comparing entity content such as line name, pole number, equipment type, and location information. For candidate equipment names, it is necessary to comprehensively compare their similarity relationship with the withdrawing switch operation text, and prioritize the results that are consistent in line, number, and equipment type.

[0075] If multiple candidates are similar but cannot be reliably distinguished, or if there is no clear corresponding item in the externally provided list of switchgear, a forced match will not be made.

[0076] Step 420: Conservative selection of ambiguous candidates in a multimodal large model.

[0077] When multiple similar candidates or candidates with small differences exist in the current matching step, the current match, related OCR text, candidate device name, action word, and matching information need to be input into the multimodal large model. The multimodal large model then makes a conservative selection within the candidate set, using prompt words such as... Figure 14 As shown.

[0078] Step 430: Locate the time information field area to obtain the candidate area.

[0079] All time fields were re-identified through partial image cropping to ensure completeness and accuracy. Specifically, combining OCR layout blocks, table areas, and cell candidate information, a multimodal large model was used to match and determine the source of the field, i.e., whether the field should be obtained from a plain text block, a table block, or a specific cell. The time candidate matching prompts used are attached. Figure 15 As shown, the purpose is to clarify the structural attribution of fields before cropping. The table region and cell candidate information can be obtained through a table recognition model that can output table structure and cell coordinate information, such as the Paddle general table recognition model. This model outputs the table region bbox, cell bbox and their corresponding relationship, providing structural information support for subsequent field source determination and candidate region location.

[0080] After determining the source region, a differentiated candidate region strategy is adopted based on different structures. For fields not originating from tables, the corresponding text block bounding boxes are preferentially retained as candidate regions. For fields originating from tables, finer-grained cell bounding boxes are preferentially used as candidate regions. If reliable cell-level positioning results cannot be obtained, the table block bounding boxes are used as candidate regions. This hierarchical strategy effectively distinguishes time fields in ordinary text regions and table regions, avoiding the mixing of fields from different structures.

[0081] After determining the field region, candidate cropping regions are generated based on the candidate regions, and the cropped image is used as input for visual verification of the multimodal large model. If the field source region is determined to be a table but lacks cell-level positioning results, the table block range is directly used for cropping.

[0082] Step 440: Select a local cropping region based on the candidate region and perform a visual review.

[0083] This stage begins with the multimodal large model selecting the most suitable clipping region for the review time field from the candidate regions. The field clipping region selection prompts used are attached. Figure 16 As shown.

[0084] The selection of candidate clipping regions follows a structure-aware rule: if the time field originates from a non-table area, the corresponding text block bbox is selected first, i.e., the corresponding candidate area is the candidate clipping region; if the time field originates from a table area and fine-grained cell bboxes exist, the cell bboxes are selected first, i.e., the corresponding candidate area is the candidate clipping region; if the time field originates from a table area but no reliable cell bboxes exist, the table block bbox or the field block_bbox is selected as the candidate clipping region.

[0085] To avoid information loss after cropping due to the written content being too close to the boundary, the candidate cropping area is appropriately expanded outward, and then a local cropping image is generated from the page image, such as... Figure 6 As shown. After cropping, the multimodal large model is called to perform a visual review of the visible time content in the local cropped image. The time field cropping review prompts used are as follows: Figure 17 As shown.

[0086] It is important to emphasize that this step does not involve artificially filling in the time information based on context or existing OCR results. Instead, it directly re-identifies the time field based on the visible content in the cropped image. If more complete year, month, day, hour, and minute information can be identified in the cropped area, the initial extraction result is corrected using the visual verification results. If only part of the time can be identified in the cropped image, only the identifiable content is retained, and no inferences are made about the invisible parts. Finally, the field corrected through local visual verification will record its corresponding repair status to indicate that the time field has been verified using the cropped image.

[0087] Figure 12 This demonstrates the intermediate results of switch matching, field location, cropping verification, and the final output. It can be seen that the initial OCR recognition result ended on "January 8, 2025"; after field location and partial cropping verification, the final output was corrected to "January 8, 2025, 6:30 AM". Figure 12 (Underlined). Since no reliable match was confirmed in the external device list, both matched_switch and switch_id are empty, but the action word, ticket type, and time fields are retained.

[0088] Step 450: Assembly of interface results and return of status.

[0089] After completing the above steps, the OCR recognition results, structured extraction results, switch matching results, field location results, cropping results, and time verification results are assembled in a unified manner to form the final pipeline result.

[0090] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for structure perception and visual repair of power operation tickets, characterized in that, Includes the following steps: Step 100: Receive the power work ticket documents and standardize them; Step 200: Perform OCR recognition on the standardized power operation ticket document to obtain the OCR recognition result, and set extraction rules to clean the OCR recognition result and organize candidate evidence; Step 300: Input the OCR recognition results, candidate evidence and extraction rules into the multimodal large model for structured extraction, perform integrity checks, and perform visual repair based on the integrity check results.

2. The method for structure perception and visual repair of power operation tickets according to claim 1, characterized in that, In step 100, the power work ticket document includes the original document of the power work ticket to be identified and a list of all switchgear related to the power work ticket.

3. The method for structure perception and visual repair of power operation tickets according to claim 1, characterized in that, In step 200, extraction rules are set to clean the OCR recognition results and organize candidate evidence, including: Clean the OCR recognition results, that is, remove the outer fields and model configuration fields that are not involved in business extraction; After cleaning, the OCR recognition results are compressed, retaining only the page and text block information that is valuable for business extraction, forming OCR evidence; the page and text block information that is valuable for business extraction includes page width, page height, text block content, text block type, text block coordinates, text block number, and reading order; Based on OCR evidence, each text block is evaluated for its candidate value according to business semantics, and candidate evidence is organized. The business semantics include time source semantics, action word semantics, device keyword semantics, and disabled template sentences.

4. The method for structure perception and visual repair of power operation tickets according to claim 3, characterized in that, In step 300, the integrity check includes: whether there are start time and end time fields, whether a valid switch operation event is extracted, and whether the preliminary extraction result is empty.

5. A method for structure perception and visual repair of power operation tickets according to claim 4, characterized in that, Step 300 further includes: When the initial extraction results are empty, or key fields are missing and no valid candidates can be formed, then structure-aware page segmentation and local re-identification are performed. When the page contains a table area with a preset ratio, and the initial extraction results are missing key fields, the horizontal lines of the complex table are weakened and the OCR recognition is performed again. If the initial extraction results include fields that can be further processed, but the field content is incomplete, then field region positioning, partial cropping, and visual verification will be performed.

6. The method for structure perception and visual repair of power operation tickets according to claim 5, characterized in that, Missing key fields include: the initial extraction results are missing any key field from the start time, end time, or switch operation event, meaning the corresponding key field is missing or empty.

7. A method for structure perception and visual repair of power operation tickets according to claim 5, characterized in that, Perform structure-aware page segmentation and local re-identification, including: The entire page image is horizontally divided into image blocks for local recognition. The following principles are followed when dividing the image: First, keep the page width unchanged, and divide the image blocks horizontally along the page height so that each image block retains complete horizontal layout information. After each segmentation, the OCR model is called again, and the OCR recognition results are processed, candidate evidence is constructed, and multimodal large model structure extraction is performed again.

8. A method for structure perception and visual repair of power operation tickets according to claim 5, characterized in that, Complex table horizontal line weakening and re-OCR recognition, including: By breaking and weakening the horizontal lines, the interference of the horizontal table lines on text recognition is reduced. After processing, OCR recognition is performed again and information is extracted.

9. A method for structure perception and visual repair of power operation tickets according to claim 5, characterized in that, Field region positioning, partial cropping, and visual verification include: Combining OCR layout blocks, table areas, and cell candidate information, a multimodal large model is used to match and determine the source of fields. The source of fields includes ordinary text blocks, table blocks, and specific cells; wherein, the table area and cell candidate information are obtained through a table recognition model. For fields not sourced from tables, the corresponding text block's bounding box is retained as the candidate region; for fields sourced from tables, fine-grained cell bounding boxes are used as the candidate region. Candidate cropping regions are generated based on candidate regions, and the images of the local cropping regions are used as inputs to the visual kernel of the multimodal large model. If complete year, month, day, hour, and minute information can be identified in the cropped area, the initial extraction result is corrected using the visual verification result; if only part of the time can be identified in the cropped image, only the identifiable content is retained. The field that has undergone local visual verification and correction records the corresponding repair status, indicating that the time field has been verified by cropping the image.

10. A method for structure perception and visual repair of power operation tickets according to claim 9, characterized in that, Generate local cropping regions based on candidate regions, including: If the time field originates from a non-table range, the candidate range is used as the candidate clipping range. If the time field originates from a table range and contains fine-grained cell bounding boxes, then the candidate range is used as the candidate clipping range. If the time field originates from a table range but no reliable cell bbox exists, then either the table block bbox or the field block_bbox is selected as the candidate clipping region.