Automatic extraction and entry method for nuclear security supervision document data

By constructing an automatic data extraction model for nuclear safety regulatory documents and utilizing text recognition and regular expression matching technologies, the problem of fragmented data management in nuclear safety regulation has been solved, achieving efficient and accurate data entry and improving regulatory efficiency and data quality.

CN121706745APending Publication Date: 2026-03-20RES INST OF NUCLEAR POWER OPERATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511635875.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Nuclear safety regulation suffers from low efficiency and human error due to the fragmentation of data and knowledge and reliance on experience in management. Existing technologies make it difficult to efficiently input nuclear safety regulatory data.

Method used

By employing text recognition, file structure parsing, and regular expression matching technologies, an automatic data extraction model for nuclear safety regulatory documents is constructed to automatically extract key information and input it into the database.

Benefits of technology

It has improved the efficiency and accuracy of nuclear safety regulatory data acquisition, reduced labor costs, and achieved an accuracy rate of over 90% for automatic extraction and data entry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121706745A_ABST
    Figure CN121706745A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of nuclear security supervision, and particularly relates to an automatic extraction and entry method for nuclear security supervision document data. Comprising the following steps: step 1, sorting a nuclear security supervision document template and key fields; 2, constructing a nuclear security supervision document intelligent extraction model; and step 3, automatically inputting data and configuring items. The method for automatically extracting and inputting the nuclear security supervision document data has the beneficial effects that text recognition can be carried out on the nuclear security supervision documents of various formats and templates, key information fields are automatically extracted from the documents according to requirements and are filled into a fixed form, and the efficiency is improved. And automatic extraction and entry of nuclear safety supervision data are realized. By developing a corresponding intelligent system, the automatic extraction and entry accuracy of the nuclear safety supervision data can reach 90% or above, the actual application requirements are met, the labor cost can be remarkably reduced, and the obtaining efficiency and accuracy of the nuclear safety supervision data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of nuclear safety regulatory technology, specifically relating to a method for automatically extracting and inputting nuclear safety regulatory document data. Background Technology

[0002] With the continuous upgrading of nuclear science and technology and the increasing demands for nuclear safety regulation in my country, the current nuclear safety field faces shortcomings such as fragmented data and knowledge, reliance on experience in management, resulting in insufficient regulatory efficiency and inadequate quality of regulation. These issues urgently require solutions through next-generation digital and artificial intelligence technologies. Regarding the construction of nuclear safety regulatory databases, due to the diversity of various supervisory information and regulatory report formats, the entry of nuclear safety regulatory data is primarily done manually by searching through various regulatory reports, compiling key information fields, filling data forms, and storing them in the database. This method is inefficient and susceptible to human error. Summary of the Invention

[0003] The purpose of this invention is to provide a method for automatically extracting and entering nuclear safety regulatory document data. By combining technologies such as text recognition, file structure parsing, and regular expression matching with nuclear safety regulatory data services, the method can automatically extract key information from nuclear safety regulatory documents and enter it into a database form, thereby improving the efficiency of nuclear safety regulatory data acquisition.

[0004] The technical solution of the present invention is as follows: A method for automatically extracting and entering nuclear safety regulatory document data, comprising the following steps:

[0005] Step 1: Reviewing nuclear safety regulatory document templates and key fields;

[0006] Step 2: Construction of an intelligent extraction model for nuclear safety regulatory documents;

[0007] Based on the nuclear safety regulatory document template and extraction field requirements defined in step 1, an automated information extraction model is constructed using file structure parsing and regular expression matching methods.

[0008] Step 3: Automated data entry and configuration item matching;

[0009] The structured extraction results output by the model in step 2 are automatically populated into the corresponding form fields of the nuclear safety regulatory database. For fields involving predefined configuration items, a complete configuration item mapping rule library is pre-organized and established. During data entry, the extracted raw text is intelligently matched and mapped with the configuration item rule library, and the standardized configuration item results that are successfully matched are filled into the database form.

[0010] Step 1 includes:

[0011] Step 11: Based on the core business specifications and data requirements of nuclear safety supervision, identify the target document types that need to achieve automated information extraction and entry;

[0012] Step 12: For each document type, define its standard template structure and the key information fields to be extracted.

[0013] The key information fields in step 12 include license number, equipment model, inspection date, responsible person, and limit parameters.

[0014] Step 2 includes:

[0015] Step 21: User upload

[0016] Users upload attachments to nuclear safety regulatory documents;

[0017] Step 22: File Type Identification and Text Conversion

[0018] The document type is identified based on the file extension. For PDF files, the OCR engine is called to convert the content into a processable text format.

[0019] Step 23: Document Preprocessing and Structure Analysis

[0020] Uploaded files undergo cleaning and standardization preprocessing;

[0021] Step 24: Target text content localization. Based on the parsed structured information, locate the text segment containing the required key information.

[0022] Step 25: Field matching based on rules and regular expressions;

[0023] Step 26: Output the matching results. Output all successfully extracted key fields and their corresponding values.

[0024] The structural parsing in step 23 includes, for DOCX files, directly parsing their internal XML structure, identifying and extracting the structured elements and content of paragraphs, heading levels, tables, and lists in the document; for PDF, DOC, and WPS format files, structural parsing is performed based on page number order, table features, and specific text style rules to restore the document's logical hierarchy.

[0025] In step 25, for each key field to be extracted, a text pattern recognition rule is designed, and the target field value is extracted from the located text fragment using regular expressions. The state is changed according to the input order and the current state, and there will be one or more outputs in each state.

[0026] In the regular expression a|ab, the circle represents the initial state, and the ring represents the termination state. 1 is the initial state. After inputting 'a', the input can transition to two states, 2 and 3. If the input is 'b', the input transitions to state 4, which is the termination state, meaning the string 'ab' is matched. If the input is 'b', the input transitions to state 3, which is the termination state, meaning the string 'a' is matched. There can be one or more termination states. Starting from 'a', each part is checked, and the current text is checked to see if it matches the current part of the expression. If it does, the input continues to the next part of the expression. This continues until all parts of the expression match, meaning the entire expression is successfully matched.

[0027] The beneficial effects of this invention are as follows: The automatic extraction and entry method for nuclear safety regulatory document data constructed by this invention can perform text recognition on nuclear safety regulatory documents of various formats and templates, and automatically extract key information fields from the documents according to requirements, and fill them into fixed forms, thereby realizing the automatic extraction and entry of nuclear safety regulatory data. Through the development of a corresponding intelligent system, the accuracy rate of automatic extraction and entry of nuclear safety regulatory data can reach over 90%, meeting practical application needs, significantly reducing labor costs, and improving the efficiency and accuracy of nuclear safety regulatory data acquisition. Attached Figure Description

[0028] Figure 1 A schematic diagram of the automatic data extraction process for nuclear safety regulatory documents;

[0029] Figure 2 To automatically extract and match the schematic diagram;

[0030] Figure 3 The regular expression is a|ab. Detailed Implementation

[0031] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] This invention provides a method for automatically extracting and inputting nuclear safety regulatory document data. Based on cutting-edge technologies such as Optical Character Recognition (OCR) and Natural Language Processing (NLP), it uses various paper documents as input. Through document scanning and text recognition technologies, it performs text recognition on unstructured paper documents. Based on this, it organizes the format templates of documents requiring intelligent input, including various document types under multiple business modules such as nuclear safety licensing, nuclear safety review, and nuclear safety supervision. A customized information extraction model is developed for each document format template. Metadata annotation and knowledge tag extraction are performed on the recognized document content. The extracted electronic text information is mapped to fields of corresponding data modules in the nuclear safety database, ultimately achieving the digitization and intelligent input of paper documents. This includes the following:

[0033] (1) Organize the types and templates of nuclear safety regulatory documents, organize the fields of nuclear safety database forms, and determine the names of the fields that need to be extracted;

[0034] (2) Combining nuclear safety regulatory document templates, and using technologies such as OCR text recognition, file structure parsing, and regular expression matching, an automatic data extraction model for nuclear safety regulatory documents is constructed to realize document data recognition, data extraction, and matching;

[0035] (3) Develop an automatic data extraction and entry function for nuclear safety regulatory documents and integrate it into the database management system. Based on the fields of the nuclear safety database form, the extracted field information will be automatically filled into the form to realize the automatic entry of nuclear safety regulatory data.

[0036] A method for automatically extracting and entering nuclear safety regulatory document data includes the following steps:

[0037] Step 1: Reviewing Nuclear Safety Regulatory Document Templates and Key Fields

[0038] Including the following:

[0039] Step 11: Based on the core business specifications and data requirements of nuclear safety supervision, the system identifies the target document types that need to achieve automated information extraction and entry.

[0040] Step 12: For each document type, define its standard template structure and the key information fields to be extracted (e.g., license number, device model, inspection date, responsible person, limit parameters, etc.). See Table 1 for a detailed list of the document types covered and their corresponding key fields to be extracted. This step lays the foundation for subsequent automated processing, ensuring the extraction targets are accurate and clearly defined.

[0041] Table 1. Nuclear Safety Regulatory Document Types and Extracted Fields

[0042]

[0043]

[0044] Step 2: Construction of Intelligent Extraction Model for Nuclear Safety Regulatory Documents

[0045] Based on the document template and extraction field requirements defined in Step 1, an automated information extraction model is constructed by comprehensively utilizing core technologies such as file structure parsing and regular expression matching. The core processing flow of the model is as follows: Figure 1 As shown, it includes the following key steps:

[0046] Step 21: User upload: The user uploads attachments to the nuclear safety regulatory document through the relevant database system.

[0047] Step 22: File Type Recognition and Text Conversion: Identify the document type based on the file extension (supported: PDF, DOCX, DOC, WPS). For PDF files, use an OCR (Optical Character Recognition) engine to convert their content into a processable text format.

[0048] Step 23: Document Preprocessing and Structure Analysis: Perform necessary cleaning and standardization preprocessing on the uploaded file (and the OCR-converted PDF text). The deep document structure analysis process is as follows:

[0049] For DOCX files (based on the Office Open XML standard), the internal XML structure is directly parsed to accurately identify and extract structured elements such as paragraphs, heading levels, tables, and lists, as well as their content.

[0050] For PDF (after OCR), DOC, WPS and other file formats, the structure is parsed based on heuristic rules such as page number order, table features (such as borders and header patterns), and specific text styles (such as bold and heading fonts) to restore the document's logical hierarchy.

[0051] Step 24: Target text content location: Based on the parsed structured information, accurately locate the text fragment containing the required key information.

[0052] Step 25: Field Matching Based on Rules and Regular Expressions: For each key field to be extracted, design specific text pattern recognition rules (detailed recognition rules are shown in Table 2). Utilize the powerful pattern matching capabilities of regular expressions (RegEx), combined with the principles of Finite State Automatons (FSA) (matching principles are discussed in...). Figure 2 This method precisely extracts target field values ​​from located text fragments. The matching principle involves progressively changing the state based on the input order and the current state, with one or more outputs at each state. A finite state automaton contains a finite number of states, one of which is the initial state. Depending on the type of input, a set of transition rules determines the next state to switch to.

[0053] Table 2. Rules for Text Pattern Recognition in Nuclear Safety Regulatory Documents

[0054]

[0055]

[0056]

[0057]

[0058] Using the regular expression a|ab( Figure 3 For example:

[0059] The circle represents the initial state, and the ring represents the final state. 1 is the initial state, and after inputting 'a', it can transition to states 2 and 3.

[0060] The input transitions to state 2. Then, inputting 'b' transitions to state 4, which is the termination state. This means the string "ab" has been matched.

[0061] The program transitions to state 3, which is the termination state. This means the string 'a' has been matched.

[0062] There can be one or more termination states. It starts from "a" and checks a portion of the expression at a time (the regular expression engine checks a part of the expression). At the same time, it checks whether the current text matches the current part of the expression. If it does, it continues to the next part of the expression, and so on, until all parts of the expression match, that is, the entire expression matches successfully.

[0063] Step 26: Output matching results: Output all successfully extracted key fields and their corresponding values.

[0064] Step 3: Automated data entry and configuration item matching

[0065] The structured extraction results from the model in step 2 are automatically populated into the corresponding form fields of the nuclear safety regulatory database. For fields involving predefined configuration items (such as equipment type codes, status codes, standard clause numbers, etc.), a complete configuration item mapping rule base is pre-organized and established. During data entry, the system automatically performs intelligent matching and mapping between the extracted raw text and the configuration item rule base. The successfully matched standardized configuration item results (not the raw text) are accurately filled into the database form, ensuring the standardization, consistency, and analyzability of the data, significantly improving data quality and subsequent regulatory efficiency.

Claims

1. A method for automatically extracting and inputting nuclear safety regulatory document data, characterized in that, Includes the following steps: Step 1: Reviewing nuclear safety regulatory document templates and key fields; Step 2: Construction of an intelligent extraction model for nuclear safety regulatory documents; Based on the nuclear safety regulatory document template and extraction field requirements defined in step 1, an automated information extraction model is constructed using file structure parsing and regular expression matching methods. Step 3: Automated data entry and configuration item matching; The structured extraction results output by the model in step 2 are automatically populated into the corresponding form fields of the nuclear safety regulatory database. For fields involving predefined configuration items, a complete configuration item mapping rule library is pre-organized and established. During data entry, the extracted raw text is intelligently matched and mapped with the configuration item rule library, and the standardized configuration item results that are successfully matched are filled into the database form.

2. The method for automatically extracting and inputting nuclear safety regulatory document data as described in claim 1, characterized in that, Step 1 includes: Step 11: Based on the core business specifications and data requirements of nuclear safety supervision, identify the target document types that need to achieve automated information extraction and entry; Step 12: For each document type, define its standard template structure and the key information fields to be extracted.

3. The method for automatically extracting and entering nuclear safety regulatory document data as described in claim 2, characterized in that: The key information fields in step 12 include license number, equipment model, inspection date, responsible person, and limit parameters.

4. The method for automatically extracting and inputting nuclear safety regulatory document data as described in claim 1, characterized in that, Step 2 includes: Step 21: User upload Users upload attachments to nuclear safety regulatory documents; Step 22: File Type Identification and Text Conversion The document type is identified based on the file extension. For PDF files, the OCR engine is called to convert the content into a processable text format. Step 23: Document Preprocessing and Structure Analysis Uploaded files undergo cleaning and standardization preprocessing; Step 24: Target text content localization. Based on the parsed structured information, locate the text segment containing the required key information. Step 25: Field matching based on rules and regular expressions; Step 26: Output the matching results. Output all successfully extracted key fields and their corresponding values.

5. The method for automatically extracting and entering nuclear safety regulatory document data as described in claim 4, characterized in that, The structural parsing in step 23 includes, for DOCX files, directly parsing their internal XML structure, identifying and extracting the structured elements and content of paragraphs, heading levels, tables, and lists in the document; for PDF, DOC, and WPS format files, structural parsing is performed based on page number order, table features, and specific text style rules to restore the document's logical hierarchy.

6. The method for automatically extracting and inputting nuclear safety regulatory document data as described in claim 4, characterized in that, In step 25, for each key field to be extracted, a text pattern recognition rule is designed, and the target field value is extracted from the located text fragment using regular expressions. The state is changed according to the input order and the current state, and there will be one or more outputs in each state.

7. The method for automatically extracting and entering nuclear safety regulatory document data as described in claim 6, characterized in that: In the regular expression a|ab, the circle represents the initial state, and the ring represents the termination state. 1 is the initial state. After inputting 'a', the input can transition to two states, 2 and 3. If the input is 'b', the input transitions to state 4, which is the termination state, meaning the string 'ab' is matched. If the input is 'b', the input transitions to state 3, which is the termination state, meaning the string 'a' is matched. There can be one or more termination states. Starting from 'a', each part is checked, and the current text is checked to see if it matches the current part of the expression. If it does, the input continues to the next part of the expression. This continues until all parts of the expression match, meaning the entire expression is successfully matched.