Nuclear power system file layout identification and structured information extraction method and system
Through a method of document layout recognition and structured information extraction for nuclear power system, deep learning and OCR models are used, combined with rule engines and regular expressions, chapters, tables and management elements in the file are automatically identified and extracted, which solves the problems of low classification efficiency, high error rate and insufficient complex layout processing capabilities in the existing technology, and realizes efficient and accurate file classification and structured information extraction.
Patent Information
- Application Number
- CN202411687501.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-05-02
AI Technical Summary
The existing classification methods of nuclear power system documents rely on manual or simple rules, making it difficult to cope with the complex situation of irregular file header information and diversified formats, resulting in low classification efficiency and high error rate. The existing OCR tools have limited processing capabilities for complex layouts, and have low accuracy of recognition text, especially in complex layouts. The existing methods cannot automatically extract the chapter tree structure, and require manual analysis and labeling, which is inefficient and error-prone.
Provide a method for identifying and extracting structured information of nuclear power system files, including data classification, preprocessing, layout analysis and data restoration, text vectorization and semantic analysis, and generate structured output. This method uses deep learning object detection model and OCR model, combined with rule engines and regular expressions to automatically identify and extract chapters, tables and management elements in files.
It improves the accuracy and efficiency of file classification and significantly reduces the need for manual correction. The processing capability of scanning files has been enhanced, and the stability in complex layout processing has also been significantly enhanced. Automatic restoration of file hierarchy is realized, greatly reducing manual intervention time. Efficient recognition and restoration of table content, efficient text semantic extraction and clustering, and high-quality structured output is achieved.
Smart Images

Figure CN119919955A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document knowledge management, and in particular to a method and system for document layout recognition and structured information extraction in a nuclear power system. Background Art
[0002] Nuclear power system documents are quality control and management system documents to ensure the compliance, consistency and efficiency of nuclear power enterprise organization and operation, usually including: upstream laws, regulations, standards, industry specifications, technical guidelines, and superior enterprise management guidelines, etc.; enterprise internal management control management outline, field-level management procedures, business-level management procedures and operating instructions, etc. System documents usually cover the management areas of nuclear power enterprises such as business management, production safety, engineering construction, and market development.
[0003] The characteristics of system documents include:
[0004] There are many system documents, usually involving more than 2,000 upstream laws, regulations, standards, specifications and technical documents; more than 3,000 documents involving the superior guidelines of nuclear power enterprises, the enterprise's own management procedures and systems;
[0005] The system file formats are diverse or unformatted. Usually, the upstream files in the system files are mostly scanned PDF files. The system files in the enterprise are mostly in Word format, but some files are not formatted for chapters or numbers, for example: all text content is in the "text" format, and chapter information cannot be extracted, and the numbering information is not formatted, so the hierarchical relationship cannot be directly extracted;
[0006] The system document contains many management elements. For nuclear power management, it at least includes: scope of application, reference definition, reference, process, process regulations / requirements, positions / responsibilities, forms / records, etc. However, different nuclear power companies have different descriptions of management elements; the system document itself also contains some file management meta-information, including at least: file title, code, version, status, upgrade instructions, approval information, etc.
[0007] As the nuclear power industry continues to deepen its information and digital construction, the need to format, structure, model and standardize system documents has become increasingly prominent, and has become an important issue that nuclear power companies urgently need to solve. Summary of the invention
[0008] In view of the above-mentioned problems, the present invention is proposed.
[0009] Therefore, the technical problem solved by the present invention is that the existing classification method of nuclear power system files relies on manual or simple rules, which is difficult to cope with the complex situation of non-standard file header information and diversified formats, resulting in low classification efficiency and high error rate. About 60% of nuclear power files are scanned copies, which often contain interference information such as watermarks, seals, and direction offsets. Existing OCR tools have limited processing capabilities for these problems, and the accuracy of text recognition is low, especially in complex layouts (such as tables and nested pictures), the recognition failure rate is high. The chapter title formats of nuclear power files are diverse, and the paragraph styles are not unified. The existing methods cannot automatically extract the chapter tree structure, and manual analysis and annotation are required, which is inefficient and error-prone. The files output by the existing system are mostly simple plain text or picture formats, lacking structured information such as chapters, tables, and management elements, and it is difficult to meet the needs of intelligent management.
[0010] In order to solve the above technical problems, the present invention provides the following technical solutions: a method for nuclear power system document layout recognition and structured information extraction, comprising: collecting first object data and classifying the first object data.
[0011] The first preprocessing is performed on the classification results, and the layout analysis and data restoration are performed on the preprocessing results.
[0012] Text vectorization and semantic analysis are performed on the restored data to obtain structured output.
[0013] As a preferred solution of the method for nuclear power system document layout recognition and structured information extraction of the present invention, the classification of the first object data includes three classification methods. The three classification methods include classification by first object content, classification by first object medium, and automatic classification based on first object header information.
[0014] As a preferred solution of the method for nuclear power system document layout recognition and structured information extraction of the present invention, the classification result includes scanned files and editable files. The first preprocessing of the classification result includes respectively performing the first preprocessing on the scanned files and the editable files.
[0015] As a preferred solution of the method for nuclear power system document layout recognition and structured information extraction described in the present invention, the layout parsing and data restoration of the preprocessing results includes layout element recognition and analysis and OCR text recognition and restoration.
[0016] Layout element recognition uses the target detection model to detect and annotate layout elements for scanned documents, and parses the document object model for editable files.
[0017] Layout element analysis includes processing of text, pictures, and table blocks.
[0018] As a preferred solution of the method for nuclear power system document layout recognition and structured information extraction described in the present invention, the OCR text recognition includes using an OCR model to recognize text content.
[0019] Restoration includes identifying and merging text lines of the same paragraph, using regular expressions to identify chapter titles, generating a chapter tree structure, and forming a hierarchical document.
[0020] As a preferred solution of the method for nuclear power system document layout recognition and structured information extraction described in the present invention, wherein: the first object includes but is not limited to the nuclear power work system, and the first object data includes the nuclear power system document.
[0021] The first preprocessing includes scanned file preprocessing and editable file preprocessing. Scanned file preprocessing includes grayscale and binarization, watermark and stamp removal, size and direction normalization. Editable file preprocessing includes chapter number normalization and paragraph style processing formatting.
[0022] As a preferred solution of the method for nuclear power system document layout recognition and structured information extraction described in the present invention, wherein: text vectorization and semantic analysis are performed based on the restored data to obtain a structured output including text vectorization including text fragmentation, semantic vector generation and cross-chapter semantic clustering.
[0023] Structured output includes generating document files in Markdown format, including chapter titles, text content, pictures and tables, and outputting a document structure tree.
[0024] A nuclear power system document layout recognition and structured information extraction system, characterized by: including:
[0025] The classification module collects first object data and classifies the first object data.
[0026] The analysis module performs a first preprocessing on the classification results, and performs layout analysis and data restoration on the preprocessing results.
[0027] The output module performs text vectorization and semantic analysis based on the restored data to obtain structured output.
[0028] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0029] A computer-readable storage medium stores a computer program, which implements the steps of the method described above when executed by a processor.
[0030] The beneficial effects of the present invention are: improving the accuracy and efficiency of file classification, while significantly reducing the need for manual correction. The processing capability of scanned files is enhanced, and the stability in complex layout processing is also significantly enhanced. Automatic restoration of file hierarchical structure is achieved, greatly reducing the time of manual intervention. Efficient recognition and restoration of table content, efficient text semantic extraction and clustering are achieved, and high-quality structured output is generated. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0032] Figure 1 An overall flow chart of a method and system for nuclear power system document layout recognition and structured information extraction provided for the first embodiment of the present invention.
[0033] Figure 2 A document template diagram for constructing a method and system for nuclear power system document layout recognition and structured information extraction provided in the first embodiment of the present invention.
[0034] Figure 3 A bubbling algorithm flow chart of a method and system for nuclear power system document layout recognition and structured information extraction provided in the first embodiment of the present invention. DETAILED DESCRIPTION
[0035] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.
[0036] Example 1, reference Figure 1 to Figure 3 , which is an embodiment of the present invention, provides a method for nuclear power system document layout recognition and structured information extraction, comprising:
[0037] S1: Collect first object data and classify the first object data.
[0038] In the present invention, the first object is a nuclear power work system, the first object data is a nuclear power system file, and the system collects nuclear power system files from a variety of sources, including internal and external uploads of the enterprise, automatic database synchronization, and manual retrieval.
[0039] The file formats mainly include PDF (scanned or editable version), Word / WPS documents and a small number of other formats (TXT).
[0040] The classification includes the following situations:
[0041] Classification by document content: laws and regulations, technical guidelines, management systems, operating instructions, etc.
[0042] Classification by file medium: scanned PDF, editable PDF and Word / WPS files.
[0043] File classification is done automatically using a rule engine based on file header information (title, format), with an option for manual correction.
[0044] The first object includes but is not limited to nuclear power work systems, and may also be aviation industry operation systems and financial industry risk control systems. The first object data includes but is not limited to nuclear power system files, and may also be aviation industry system files and financial industry risk control system files.
[0045] In an optional embodiment of the present invention, the first object is an aviation industry operating system, and the first object data is an aviation industry system file. The aviation industry operating system file covers the technical and management documents involved in aviation manufacturing and operation, and its sources are diverse, including aviation enterprise uploads, industry database sharing, and manual uploads. Common file formats include scanned PDF, editable PDF, Word / WPS documents, and other formats (such as TXT). The specific content includes the following types:
[0046] Technical documents: technical specifications related to aviation design, production, maintenance, etc., such as engine testing standards, fuselage material inspection procedures, etc.
[0047] Regulations and standards: Domestic and international regulations and international standards for the aviation industry, such as the ICAO Regulations and Aircraft Airworthiness Standards.
[0048] Operating Manual: Aircraft operating procedures, including pilot manual, maintenance operation manual, etc.
[0049] Management documents: enterprise management specifications, quality control procedures, risk assessment reports, etc.
[0050] The aviation industry system document classification method is as follows:
[0051] Classification by file content category:
[0052] Regulations and standards: such as national aviation laws and regulations, and international aviation standards.
[0053] Technical documents: technical standards and operating guidelines for aircraft design, manufacturing, maintenance, etc.
[0054] Operations Manual: Guide to the operation and maintenance of an aircraft.
[0055] Management documents: such as quality control plans, risk assessment reports, etc.
[0056] Classification by file medium:
[0057] Scanned PDF: For example, regulatory documents digitized from paper archives.
[0058] Editable PDF: such as design specifications and technical standards.
[0059] Word / WPS files: such as the latest published corporate operating procedures.
[0060] Automatic classification based on file header information:
[0061] Automatically categorize files based on file title and format (such as regulation number and manual type) through the rule engine, such as:
[0062] Documents whose titles contain keywords such as "standard" and "regulation" can be classified as technical documents.
[0063] Documents containing the words "guide" or "manual" are classified as operating manuals.
[0064] In an optional embodiment of the present invention, the first object is the financial industry risk control system, and the first object data is the financial risk control system file. The financial risk control system file involves risk control files of financial institutions such as banks, securities, and insurance, covering internal control management, risk assessment, and emergency response. Data sources include internal uploads from financial institutions, releases from regulatory authorities, and industry shared databases. Common formats include PDF (scanned or editable version), Word / WPS documents, and EXCEL files. Specific contents include:
[0065] Regulations and supervisory documents: Risk management regulations and policy documents issued by central banks and financial regulatory authorities.
[0066] Internal control management system: financial institutions’ risk management strategies, internal audit processes, and data security management regulations.
[0067] Risk assessment report: quantitative and qualitative risk analysis report, including market risk, credit risk, liquidity risk, etc.
[0068] Emergency response plan: a response plan for financial crises or emergencies.
[0069] The classification of financial risk control system documents is as follows:
[0070] Classification by file content category:
[0071] Regulations and supervisory documents: such as the Commercial Bank Capital Management Measures.
[0072] Internal control management system: such as risk management system and anti-money laundering management measures.
[0073] Risk assessment reports: such as credit risk assessment and market risk control strategies.
[0074] Emergency response plan: such as the emergency response plan for financial crisis.
[0075] Classification by file medium:
[0076] Scanned PDF: Scanned documents such as those issued by regulatory agencies.
[0077] Editable PDF: such as reports produced internally by financial institutions.
[0078] Word / WPS files: such as daily risk control documents and policy updates.
[0079] Automatic classification based on file header information:
[0080] The rule engine reads the file title, annotation symbol, document structure, etc. for classification, for example:
[0081] Documents containing the keywords “management methods” and “policies” are classified as laws and regulations and regulatory documents.
[0082] Documents containing the keywords “plan” and “emergency” are classified as emergency response plans.
[0083] It should be noted that it is also necessary to use file templates to build models for file structure and formatting annotation and extraction. Through users uploading standard file templates, templates for upstream laws and regulations, technical documents, procedures, etc. are built, as well as templates for enterprise management procedures. Different types of templates are suitable for different types of document files. Template files are formatted structure files such as Word / WPS. The system automatically extracts the chapters, numbers and other information and rules of the file according to the file format. The extraction algorithm is a bubbling algorithm. The results are displayed on the front-end page. Users can adjust and modify the document structure outline according to actual needs, fine-tune the automatically extracted content, and add some labels, chapter numbers, indents, and text content constraints in chapters. For example Figure 2 shown.
[0084] S2: Perform a first preprocessing on the classification results, and perform layout analysis and data restoration on the preprocessing results.
[0085] In the present invention, the first preprocessing includes scanned file preprocessing and editable file preprocessing, wherein the scanned file preprocessing includes graying and binarization, watermark and seal removal, size and direction normalization, and the editable file preprocessing includes chapter number normalization and paragraph style processing formatting.
[0086] Grayscale and binarization: Grayscale each page of the PDF scanned file to extract the core information, and then use adaptive threshold segmentation to generate a binary image to reduce background interference.
[0087] Watermark and seal removal uses RGB channel separation technology to extract the color information of the red and blue channels respectively. Morphological processing (such as dilation and corrosion operations) is used to detect the edges of the seal or watermark and remove its area.
[0088] Size and orientation normalization: The images are uniformly adjusted to the A4 standard size (2480×3508 pixels). The rotation angle is corrected using the affine transformation matrix to ensure the image orientation is consistent.
[0089] Chapter number standardization checks the chapter number format (such as "1.1", "Chapter 1", etc.), and automatically corrects problems such as duplication, missing numbers, and skipped numbers.
[0090] Paragraph style processing detects whether the body content does not use the title style. If it is not formatted, it is divided into blocks and marked according to segmentation rules (such as continuous line breaks or punctuation marks).
[0091] Layout parsing and data restoration include layout element recognition, layout element analysis, OCR recognition, and paragraph and chapter restoration.
[0092] Layout element recognition uses deep learning-based target detection models (such as Faster R-CNN or YOLO) to detect the following layout elements on scanned documents: text area, image area, and table area. Through the training model, the envelope coordinates of text blocks, image blocks, and table blocks are identified, and the element type is marked.
[0093] Layout element recognition directly parses the Document Object Model (DOM) of editable files to extract the logical location and content of paragraphs, images and tables.
[0094] Layout element analysis includes the following steps:
[0095] Text block processing: Extract and mark the boundary information of each text block to ensure accurate subsequent OCR recognition.
[0096] Image block processing: record the coordinates and size information of the image block, and retain the original image for output.
[0097] Table block processing:
[0098] Detect table boundaries and annotate cell ranges.
[0099] Use the NMS (non-maximum suppression) algorithm to merge adjacent table elements and remove redundant tags.
[0100] OCR recognition uses a common Chinese and English OCR model (such as PaddleOCR or Tesseract) to recognize the text content in the text block. Perform independent OCR processing on the text in the table cells to ensure the accuracy of the table content.
[0101] Furthermore, for unstructured text recognized by OCR, and some of the aforementioned Word files that do not contain document structures, the bubble algorithm can be used to infer the document's chapters, paragraphs, chapter numbers and other structural information, and a chapter outline view of the document can be constructed. Figure 3 shown.
[0102] The algorithm process is as follows: each chapter in the document is distributed in a tree-like manner. In order to meet the database's tabular storage standards, each chapter is first assigned a unique identifier (ID), and the parent chapter ID (i.e. PID) of the chapter is also assigned. The chapter ID, chapter name, and chapter PID become the data structure of the chapter entity, which can be stored in the database in a tabular manner. When it is necessary to obtain the chapter tree path upward, a recursive query can be performed through a linked list, bubbling upward until the chapter without PID; when it is necessary to list the left and right descendant chapters downward, a reverse recursive query can also be performed through a linked list.
[0103] Paragraph merging identifies and merges text lines in the same paragraph, using line break detection or start character alignment rules to determine paragraph ranges.
[0104] Chapter title detection uses regular expressions to match chapter title formats (such as "1.", "1.1", "Chapter 1", etc.). Automatically generate a chapter tree structure, annotate chapter IDs and parent chapter IDs, and form a hierarchical document structure.
[0105] S3: Perform text vectorization and semantic analysis based on the restored data to obtain structured output.
[0106] Text fragmentation divides the content into chapters. If the chapter is too long, it is divided into paragraphs. If the paragraph is still too long, it is divided into sentences to ensure semantic continuity.
[0107] Semantic vector generation uses the BERT model to convert each text into a fixed-length dense vector. The vector length is set to 1024 dimensions to ensure the accuracy of semantic representation.
[0108] Cross-chapter semantic clustering uses K-Means algorithm to analyze similar semantic content in different chapters and generate a set of dual vectors for semantic matching and conflict detection.
[0109] Generates a document file in Markdown format, including chapter titles, text content, pictures, and tables to output a document structure tree for subsequent document visualization or retrieval.
[0110] Based on the association relationships and management elements between files, a multi-level knowledge graph is constructed, including a file-file association graph and a file content-management element association graph.
[0111] The extracted information is stored in a database, supporting full-text search, version control, and conflict detection.
[0112] The computer device may be a server. The computer device includes a processor, a memory, an input / output interface (I / O for short) and a communication interface. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data cluster data of a power monitoring system. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for document layout recognition and structured information extraction of a nuclear power system is implemented.
[0113] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0114] Example 2 is an embodiment of the present invention, which provides a method and system for document layout recognition and structured information extraction in a nuclear power system. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through simulation experiments.
[0115] Experimental preparation and implementation process
[0116] In order to verify the technical advantages of the present invention, by comparing the traditional manual analysis method and some automatic recognition tools in the prior art (hereinafter referred to as "the prior art"), 50 typical documents in the standard system document library of nuclear power enterprises were selected for the experiment, including scanned PDF and editable Word / WPS documents, among which:
[0117] Scanned PDFs account for 60%, including laws and regulations, technical specifications, and internal management manuals;
[0118] Editable files account for 40%, covering work instructions, quality control plans, etc.
[0119] Prior art steps:
[0120] Scanned documents: Use OCR tools (such as ABBYY FineReader) to recognize text content and manually correct it. The analysis of the document layout structure is completely manual (for example, manually marking chapter titles, paragraph levels, and table content).
[0121] Editable files: Direct manual analysis of file chapter content and layout information without the assistance of automated tools.
[0122] Output: Only basic plain text files or simple Excel spreadsheet records can be generated.
[0123] The present invention implements the steps:
[0124] Data classification: Automatically classify files by content and media through the classification module.
[0125] Preprocessing:
[0126] Scanned documents: complete grayscale and binarization, remove watermarks and stamps, and unify image size and orientation.
[0127] Editable files: Automatically normalize chapter numbers and paragraph styles to ensure consistent document formatting.
[0128] Layout analysis and text restoration:
[0129] Scanned documents: Use the deep learning object detection model (YOLO) to detect layout elements (such as text blocks, image blocks, and table blocks), and recognize text content through OCR.
[0130] Editable files: Parse the Document Object Model (DOM) and directly extract logical hierarchical information.
[0131] Structured Information Extraction:
[0132] Extract chapter trees, paragraph contents, and administrative elements.
[0133] Build Markdown format files and generate structured output (including associated information and semantic vectorization results).
[0134] Output: Store file contents in the form of structured data tables (Markdown and knowledge graphs) for easy downstream analysis.
[0135] The experimental results are shown in Table 1.
[0136] Table 1 Experimental results
[0137]
[0138]
[0139] The experimental results show that the present invention is superior to the prior art in many key indicators, and the specific analysis is as follows:
[0140] The present invention adopts automatic classification technology based on rule engine and file header information, with an accuracy rate of 98%, which is much higher than the 65% manual classification result of the prior art. The prior art classification has subjective errors, especially when the file title is unclear, the classification error rate is high.
[0141] In the method of the present invention, the OCR recognition accuracy rate reaches 95%, which is 17 percentage points higher than the existing technology. The combined application of image preprocessing and deep learning OCR model effectively reduces the probability of noise interference and misrecognition, while the general OCR tools in the existing technology are insufficient in processing complex layouts.
[0142] The existing technology completely relies on manual construction of chapter levels, which takes an average of 600 seconds per file. The present invention achieves automatic restoration through regular expression detection and tree structure algorithm, shortening the time to less than 120 seconds, and improving efficiency by about 5 times.
[0143] The present invention uses a technology combining table boundary detection and cell OCR, with an accuracy rate of 92%, which is 37 percentage points higher than the prior art. The prior art is prone to misalignment or omission when manually restoring a table.
[0144] The data integrity of the present invention reaches 99%, ensuring the comprehensiveness of document information. In the prior art, chapters, annotations and management elements are easily lost during the multi-step manual processing process.
[0145] The present invention only takes 12 seconds to process each page, which is 1 / 8 of the 90 seconds of the prior art, greatly improving the document parsing efficiency.
[0146] The prior art can only extract explicit association information (such as reference marks in file titles), while the present invention can extract implicit association relationships through semantic analysis, and the extraction rate is significantly improved to 93%.
[0147] Embodiment 3 is an embodiment of the present invention, including a nuclear power system document layout recognition and structured information extraction system, specifically:
[0148] a classification module, collecting first object data and classifying the first object data;
[0149] The analysis module performs a first preprocessing on the classification results, and performs layout analysis and data restoration on the preprocessing results;
[0150] The output module performs text vectorization and semantic analysis based on the restored data to obtain structured output.
[0151] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for nuclear power system document layout recognition and structured information extraction, characterized in that: include: collecting first object data, and classifying the first object data; Performing a first preprocessing on the classification results, and performing layout analysis and data restoration on the preprocessing results; Text vectorization and semantic analysis are performed on the restored data to obtain structured output.
2. The method for nuclear power system document layout recognition and structured information extraction according to claim 1, characterized in that: The classification of the first object data includes three classification methods, and the three classification methods include classification according to the first object content, classification according to the first object medium, and automatic classification based on the first object header information.
3. The method for nuclear power system document layout recognition and structured information extraction as claimed in claim 2, characterized in that: The classification result includes a scanned file and an editable file; and the first preprocessing on the classification result includes performing the first preprocessing on the scanned file and the editable file respectively.
4. The method for nuclear power system document layout recognition and structured information extraction as claimed in claim 3, characterized in that: The layout analysis and data restoration of the preprocessing results include layout element recognition and analysis and OCR text recognition and restoration; Layout element recognition uses the target detection model to detect and annotate layout elements in scanned documents, and parses the document object model for editable documents; Layout element analysis includes processing of text, pictures, and table blocks.
5. The method for nuclear power system document layout recognition and structured information extraction as claimed in claim 4, characterized in that: The OCR text recognition includes using an OCR model to recognize text content; Restoration includes identifying and merging text lines of the same paragraph, using regular expressions to identify chapter titles, generating a chapter tree structure, and forming a hierarchical document.
6. The method for nuclear power system document layout recognition and structured information extraction according to claim 5, characterized in that: The first object includes but is not limited to a nuclear power work system, and the first object data includes a nuclear power system file; The first preprocessing includes scanned file preprocessing and editable file preprocessing, and the scanned file preprocessing includes graying and binarization, watermark and seal removal, and size and direction normalization; Editable file preprocessing includes chapter number normalization and paragraph style processing formatting.
7. The method for nuclear power system document layout recognition and structured information extraction according to claim 6, characterized in that: The text vectorization and semantic analysis are performed on the restored data to obtain structured output, wherein the text vectorization includes text fragmentation, semantic vector generation, and cross-chapter semantic clustering; Structured output includes generating document files in Markdown format, including chapter titles, text content, pictures and tables, and outputting a document structure tree.
8. A nuclear power system document layout recognition and structured information extraction system using the method according to any one of claims 1 to 7, characterized in that: a classification module, collecting first object data and classifying the first object data; The analysis module performs a first preprocessing on the classification results, and performs layout analysis and data restoration on the preprocessing results; The output module performs text vectorization and semantic analysis based on the restored data to obtain structured output.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Multi-element document analysis method and system
CN120874816A
Nuclear power safety report data extraction method and system based on multi-modal feature fusion
CN121074928A