Method and system for intelligently extracting data elements for national data standard file
Through the large-model Prompt technology combined with text chunking and natural language processing, the problems of inefficient and insufficient accuracy of information extraction of national data standard files are solved, and fast and accurate data element and rule extraction is achieved, which improves data processing efficiency and quality.
Patent Information
- Application Number
- CN202510488705.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the information extraction of national data standard documents mainly relies on manual interpretation and rule-based methods, and there are problems of inefficiency and insufficient accuracy, making it difficult to deal with diverse document formats and language expressions.
The large-scale Prompt technology is adopted, combining text chunking and natural language processing to realize the automated processing of data standard files, including document content reading, chunking, data element recognition and regular expression generation, improving extraction efficiency and accuracy.
It realizes the rapid and accurate extraction of data elements and data rules, reduces labor and time costs, supports data modeling and quality inspection, and improves the efficiency and quality of data processing.
Smart Images

Figure CN120336507A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of application of artificial intelligence large models, and specifically to a method and system for intelligently extracting data elements from national data standard documents. Background Technique
[0002] With the rapid development of information technology, national data standard documents play a crucial role in promoting information sharing, ensuring data quality and consistency. These documents define data elements (i.e., the smallest data unit definitions, such as format, value range, semantics) and data rules (such as regular expressions, etc.) within a specific domain, providing a unified basis for data exchange between government agencies or enterprises. However, in practical applications, since data standard documents are usually written in natural language and involve complex business logics and technical terms, this poses a huge challenge to automated processing.
[0003] Currently, the information extraction from national data standard documents mainly relies on two methods: (1) Manual interpretation: Professional personnel manually read and analyze the document content, and then enter the key information into a database or information system. Although this method can ensure high accuracy, it is inefficient, error-prone, and especially when faced with a large number of documents, the workload is huge and time-consuming. (2) Traditional rule-based method: By predefining a series of regular expressions or other pattern matching rules to automatically identify specific structured information in the text. Although this method can improve work efficiency to a certain extent, due to its lack of flexibility, it is difficult to handle diverse document formats and changing language expressions, resulting in low accuracy and high maintenance costs. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for intelligently extracting data elements from national data standard documents to solve the problems of low efficiency and insufficient accuracy in the existing manual extraction method. Through the present invention, it is possible to quickly and accurately extract data elements and data rules from national data standard documents, improve the efficiency and quality of data processing, and provide strong support for the intelligent processing of national data standard documents.
[0005] To achieve the above object, the present invention provides the following technical solution: A method for intelligently extracting data elements from national data standard documents, including the following steps:
[0006] The first step: Read the content of the data standard document, including WORD documents and PDF documents, and perform optical character recognition on the pictures in the document to extract the text content;
[0007] The second step: Chunk the read document content according to preset rules and identify each text chunk in turn;
[0008] Step 3: Compile the large model Prompt, and submit the chunked text to the large model for recognition to determine whether it is the content of the data element to be extracted; if so, continue to compile the data element extraction Prompt and submit it to the large model for data element extraction, otherwise discard the text chunk;
[0009] Step 4: Compile the large model Prompt, and input the extracted data element results into the large model to generate regular expressions for the data elements through the large model;
[0010] Step 5: Persistently store the data element and regular expression extraction results in the database for subsequent data modeling and data quality inspection processes.
[0011] Preferably, the specific steps for reading the content of the data standard document in the first step include: using a document processing library to read the content of the WORD document and extract the text in paragraphs and tables; using a PDF parsing library to read the content of the PDF document, extract the text and tables, and perform OCR character recognition on the pictures in the PDF; integrating the extracted text content into a unified text format and retaining the original structure information of the document.
[0012] Preferably, the specific steps for chunking the text content in the second step include: chunking the document content according to paragraphs, tables, headings, and picture text; each chunk contains chunk type, chunk content, and chunk location information; output the chunked text content as a chunk list for subsequent steps to process.
[0013] Preferably, the specific steps for data element extraction in the third step include: compiling a judgment Prompt and submitting it to the large model to determine whether the text chunk contains the content of the data element to be extracted; if the judgment is yes, compile a data element extraction Prompt and submit it to the large model to extract the name, definition, data type, and value range of the data element; store the extracted data element information as structured data for subsequent steps to use.
[0014] Preferably, the specific steps for regular expression generation in the fourth step include: compiling a generation Prompt and inputting the extracted data element information into the large model; generating regular expressions corresponding to the data element value range through the large model; verifying the generated regular expressions to ensure that they can correctly match the data element value range;
[0015] The specific steps for persisting the data standard extraction results in the fifth step include: designing the database table structure, including the data element table and the regular expression table; inserting the extracted data element information and the generated regular expressions into the database; supporting subsequent data modeling and data quality inspection processes through the query interface of the database.
[0016] A system for an intelligent data element extraction method for national data standard documents, comprising:
[0017] A document reading and recognition module, which is used to read the content of data standard documents, including WORD documents and PDF documents, perform optical character recognition on the pictures in the documents, and extract the text content;
[0018] A text chunking module, which is used to chunk the read document content according to preset rules and sequentially recognize each text chunk;
[0019] A data element recognition and extraction module, which is used to write a large model Prompt, hand over the chunked text to the large model for recognition, and judge whether it is the content of data elements to be extracted; if so, continue to write a data element extraction Prompt and hand it over to the large model for data element extraction, otherwise discard the text chunk;
[0020] A regular expression generation module, which is used to write a large model Prompt, input the extracted data element results into the large model, and generate a regular expression for the data elements through the large model;
[0021] A data persistent storage module, which is used to persistently store the data elements and the regular expression extraction results in a database for subsequent data modeling and data quality inspection processes.
[0022] Preferably, the document reading and recognition module is specifically used for: using a document processing library to read the content of WORD documents and extract the text in paragraphs and tables; using a PDF parsing library to read the content of PDF documents, extract the text and tables, and perform OCR optical character recognition on the pictures in the PDF; integrating the extracted text content into a unified text format and retaining the original structure information of the document.
[0023] Preferably, the document content is chunked according to paragraphs, tables, headings, and picture text; each chunk contains chunk type, chunk content, and chunk location information; the chunked text content is output as a chunk list for subsequent processing steps.
[0024] Preferably, the data element recognition and extraction module is specifically used for: writing a judgment Prompt and handing it over to the large model to judge whether the text chunk contains the content of data elements to be extracted; if the judgment is yes, writing a data element extraction Prompt and handing it over to the large model to extract the name, definition, data type, and value range of the data elements; storing the extracted data element information as structured data for subsequent steps to use.
[0025] Preferably, the regular expression generation module is specifically used for: writing a generation Prompt, inputting the extracted data element information into the large model; generating a regular expression corresponding to the value range of the data elements through the large model; verifying the generated regular expression to ensure that it can correctly match the value range of the data elements;
[0026] The data persistent storage module is specifically used for: designing database table structures, including data element tables and regular expression tables; inserting the extracted data element information and generated regular expressions into the database; and supporting subsequent data modeling and data quality inspection processes through the query interface of the database.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] The method and system for intelligently extracting data elements from national data standard documents proposed by the present invention perform in-depth semantic parsing of national standard documents by debugging the large model prompt, and use a combination of text segmentation and natural language processing techniques to achieve accurate identification and extraction of data elements in the documents. The present invention can quickly and accurately obtain key data information from national data standard documents, standardize the data structures of the extracted data elements and data rule information, and can be directly used for subsequent data modeling and data quality inspection, greatly saving manpower and time costs, and is of great significance for improving data processing efficiency and data quality. It is applicable to various fields involving the processing of national data standard documents and provides strong support for the intelligent processing of national data standard documents. Description of the Drawings
[0029] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments
[0030] In order to clearly and completely describe the objectives, technical solutions of the present invention, and make the advantages more clear, the following further details the embodiments of the present invention with reference to the drawings. It should be understood that the specific embodiments described herein are part of the embodiments of the present invention, rather than all of the embodiments, and are only used to explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0031] Embodiment 1, please refer to Figure 1 , the present invention provides a technical solution: a method for intelligently extracting data elements from national data standard documents, including the following steps:
[0032] Step 1: Read the content of the data standard document.
[0033] Data standard documents usually exist in various formats, such as WORD documents (.doc,.docx) and PDF documents (.pdf). Therefore, the core task of this step is to support the reading of multiple document formats and ensure the integrity and accuracy of the document content;
[0034] (1) For WORD documents, use a document processing library (the python-docx library in Python) to read them. Traverse the paragraphs and tables in the document to extract the text content. For the content in the tables, extract it in the order of rows and columns and retain the table structure information.
[0035] (2) For PDF documents, use a PDF parsing library (PyPDF2 in Python) to load the PDF document and extract the PDF text content by page number to ensure that the text is extracted in the correct order;
[0036] (3) For optical character recognition of images, whether it is a WORD document or a PDF document, the document may contain text content in the form of images. Use the OCR tool Tesseract to perform optical character recognition on the extracted images. First, preprocess the images, such as binarization and denoising, and then recognize the text content in the images;
[0037] (4) Text integration: Integrate the text content recognized by OCR with the text content in the document to ensure that all text content is correctly extracted.
[0038] Step 2: Block the text content.
[0039] In the first step, the document content has been read and integrated into a unified text format. However, standard documents usually contain a large amount of information, and directly processing the entire document will result in high computational complexity and low processing efficiency. Therefore, the core task of the second step is to block the document content according to certain rules so that subsequent steps can process it block by block, improving the efficiency and accuracy of data element and rule extraction.
[0040] (1) Title blocking: Traverse the document content to identify the start and end positions of each title. Titles are usually marked by specific title tags (such as <h1>、< / h1> <h2>label) or formatting features such as bold font, increased font size, etc.;
[0041] (2) Paragraph chunking: Based on the text chunks after title chunking in the first step, perform paragraph chunking in sequence to identify the start and end positions of each paragraph. Paragraphs are usually separated by line breaks or specific paragraph markers (such as Labels) are separated. Generate finer-grained paragraph text blocks;
[0042] (3) Table chunking: Traverse the content of the paragraph text blocks to identify the start and end positions of each table. Tables are usually marked by specific table tags (such as
[0043]
[0044]
[0045]
[0046]
[0047]
[0048]
[0049]
[0050]
[0051]
[0052]
[0053]
[0054]
[0055]
[0056]
[0057]
[0058]
[0059]
[0060]
[0061]
[0062]
[0063]
[0064] (4) After the above text chunking steps, the entire content of the data standard document is split into a list of text chunks with fine granularity. Step 3: Data element extraction. The data element extraction process mainly consists of two steps: First, determine whether the text chunk contains data element extraction content. If so, further extract the detailed information of the data element. The specific process is as follows: (1) Determine whether the text chunk contains data element extraction content a) Write a prompt for determining whether the text chunk contains data element extraction content. This prompt contains clear instructions and examples to guide the large model to make an accurate judgment. Inject the chunked text into the above prompt to call the large model for judgment. According to the output result of the large model, decide whether to continue extracting data element information. b) For text chunks determined to contain data element extraction content, design a prompt for extracting detailed data element information. Data elements usually include attributes such as data element name, English name, data element description, data type, data element definition, data format, etc. This prompt has a clear description of the data element extraction task, provides example guidance, and clearly requires the large model to output the extraction result in a specific format in the prompt, for subsequent processing and storage; Inject the text of the data element to be extracted into the above prompt, call the large model for data element extraction, and return the extraction result in a fixed format for subsequent result processing and storage. Step 4: Regular expression generation. The standard format and value range of data elements usually need to be verified through regular expressions to ensure data standardization and consistency. The core task of the fourth step is to intelligently generate corresponding regular expressions based on the extracted data element information for subsequent data quality inspection and verification processes. Through the natural language understanding and generation capabilities of the large model, regular expressions that conform to data element specifications can be generated efficiently and accurately. The implementation solution is as follows: (1) Organize the data element information extracted in step 3 into a structured input format and design a prompt for generating regular expressions for data elements. This prompt contains clear instructions and examples to guide the large model to generate accurately. (2) Inject the data element information extracted in step 3 into the above prompt and call the large model to output regular expressions that conform to data element specifications. Step 5: Persistence of data standard extraction results. After the processing of the first 4 steps, data elements and regular expressions have been extracted and generated from the national data standard document. The core task of the fifth step is to parse and persist these extraction results into the database so that subsequent data modeling and data quality inspection processes can directly reference the data standard. Regarding the parsing of the data element extraction results and regular expression generation results, call the response examples given in the large model prompt. The large model will return according to the example format. By writing regular expression parsing, design the database table structure and persist the parsing results into the database.Embodiment 2, based on Embodiment 1, proposes a system for an intelligent method of extracting data elements from national data standard documents, including: A document reading and recognition module, which is used to read the content of data standard documents, including WORD documents and PDF documents, and perform character recognition on the pictures in the documents to extract the text content; specifically: use a document processing library to read the content of WORD documents and extract the text in paragraphs and tables; use a PDF parsing library to read the content of PDF documents, extract text and tables, and perform OCR character recognition on the pictures in the PDF; integrate the extracted text content into a unified text format and retain the original structure information of the document. A text chunking module, which is used to chunk the read document content according to preset rules and sequentially recognize each text chunk; specifically: chunk the document content according to paragraphs, tables, headings, and picture text; each chunk contains chunk type, chunk content, and chunk location information; output the chunked text content as a chunk list for subsequent steps to process. A data element recognition and extraction module, which is used to write a large model Prompt, hand over the chunked text to the large model for recognition to determine whether it is the content of data elements to be extracted; if so, continue to write a data element extraction Prompt and hand it over to the large model for data element extraction, otherwise discard the text chunk; specifically: write a judgment Prompt and hand it over to the large model to judge whether the text chunk contains the content of data elements to be extracted; if the judgment is yes, write a data element extraction Prompt and hand it over to the large model to extract the name, definition, data type, and value range of the data element; store the extracted data element information as structured data for subsequent steps to use. A regular expression generation module, which is used to write a large model Prompt, input the extracted data element results into the large model, and generate a regular expression for the data element through the large model; specifically: write a generation Prompt and input the extracted data element information into the large model; generate a regular expression corresponding to the value range of the data element through the large model; verify the generated regular expression to ensure that it can correctly match the value range of the data element. A data persistent storage module, which is used to persistently store the data element and regular expression extraction results in a database for subsequent data modeling and data quality inspection processes; specifically: design the database table structure, including a data element table and a regular expression table; insert the extracted data element information and the generated regular expression into the database; support subsequent data modeling and data quality inspection processes through the query interface of the database. Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents. < / h2>
Claims
1. A method for intelligently extracting data elements from national data standard documents, characterized in that: It includes the following steps: The first step: Read the content of the data standard documents, including WORD documents and PDF documents, perform optical character recognition on the pictures in the documents, and extract the text content; The second step: Chunk the read document content according to preset rules and identify each text chunk in sequence; The third step: Compile the large model Prompt, submit the chunked text to the large model for recognition to determine whether it is the content to be extracted for data elements; if so, continue to compile the data element extraction Prompt and submit it to the large model for data element extraction, otherwise discard the text chunk; The fourth step: Compile the large model Prompt, input the extracted data element results into the large model, and generate a regular expression for the data element through the large model; The fifth step: Persistently store the data element and regular expression extraction results in the database for subsequent data modeling and data quality inspection processes.
2. The method for intelligently extracting data elements from a national data standard document according to claim 1, characterized in that: The specific steps for reading the content of the data standard documents in the first step include: Use a document processing library to read the content of the WORD document and extract the text in paragraphs and tables; Use a PDF parsing library to read the content of the PDF document, extract the text and tables, and perform OCR character recognition on the pictures in the PDF; Integrate the extracted text content into a unified text format and retain the original structure information of the document.
3. The method for intelligently extracting data elements from a national data standard document according to claim 2, characterized in that: The specific steps for chunking the text content in the second step include: Chunk the document content according to paragraphs, tables, headings, and picture text; Each chunk contains chunk type, chunk content, and chunk location information; Output the chunked text content as a chunk list for subsequent steps to process.
4. A method for intelligently extracting data elements from a national data standard document according to claim 3, characterized in that: The specific steps for data element extraction in the third step include: Compile a judgment Prompt and submit it to the large model to determine whether the text chunk contains the content to be extracted for data elements; If the judgment is yes, compile the data element extraction Prompt and submit it to the large model to extract the name, definition, data type, and value range of the data element; Store the extracted data element information as structured data for subsequent steps to use.
5. A method for intelligently extracting data elements from a national data standard document according to claim 4, characterized in that: The specific steps for regular expression generation in the fourth step include: Compile a generation Prompt and input the extracted data element information into the large model; Generate a regular expression corresponding to the value range of the data element through the large model; Verify the generated regular expression to ensure that it can correctly match the value range of the data element; The specific steps for persisting the data standard extraction results in the fifth step include: Design the database table structure, including the data element table and the regular expression table; Insert the extracted data element information and the generated regular expression into the database; Support subsequent data modeling and data quality inspection processes through the query interface of the database.
6. A system for the method of intelligently extracting data elements for a national data standard document according to claim 5, characterized in that: It includes: A document reading and recognition module for reading the content of data standard documents, including WORD documents and PDF documents, performing optical character recognition on the pictures in the documents, and extracting the text content; A text chunking module for chunking the read document content according to preset rules and identifying each text chunk in sequence; The data element identification and extraction module is used to write the large model Prompt, hand over the chunked text to the large model for identification, and judge whether it is the content of the data element to be extracted; if so, continue to write the data element extraction Prompt and hand it over to the large model for data element extraction, otherwise discard the text chunk; The regular expression generation module is used to write the large model Prompt, input the extracted data element results into the large model, and generate regular expressions for the data elements through the large model; The data persistent storage module is used to persistently store the data element and regular expression extraction results in the database for subsequent data modeling and data quality inspection processes.
7. A system according to claim 6, characterized in that: The document reading and recognition module is specifically used for: using the document processing library to read the content of the WORD document and extract the text in paragraphs and tables; using the PDF parsing library to read the content of the PDF document, extract the text and tables, and perform OCR character recognition on the pictures in the PDF; integrating the extracted text content into a unified text format and retaining the original structure information of the document.
8. A system according to claim 7, characterized in that: The text chunking module is specifically used for: chunking the document content according to paragraphs, tables, headings, and picture text; each chunk contains chunk type, chunk content, and chunk location information; outputting the chunked text content as a chunk list for subsequent steps to process.
9. A system according to claim 8, characterized in that: The data element identification and extraction module is specifically used for: writing the judgment Prompt and handing it over to the large model to judge whether the text chunk contains the content of the data element to be extracted; if the judgment is yes, write the data element extraction Prompt and hand it over to the large model to extract the name, definition, data type, and value range of the data element; store the extracted data element information as structured data for subsequent steps to use.
10. A system according to claim 9, wherein: The regular expression generation module is specifically used for: writing the generation Prompt, inputting the extracted data element information into the large model; generating regular expressions corresponding to the data element value range through the large model; verifying the generated regular expressions to ensure that they can correctly match the data element value range; The data persistent storage module is specifically used for: designing the database table structure, including the data element table and the regular expression table; inserting the extracted data element information and the generated regular expressions into the database; Support subsequent data modeling and data quality inspection processes through the query interface of the database.
Citation Information
Patent Citations
Data standardization method based on large model
CN119003583A
BIM rule specification extraction method based on large model
CN119337992A
Electric power standard structuring method and system based on image recognition technology
CN119625748A
Cited By
Multivariate data structured processing method, system and program product in automobile field
CN121188140A