A docx document business processing, data utilization system and method
Through parsing and mapping encoding technology, the docx document data is automatically processed, and the problem of repeated entry of multiple systems is solved, efficient data utilization and electronic signature are realized, and work efficiency and data utilization are improved.
Patent Information
- Application Number
- CN202111310456.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-04
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-11-04
AI Technical Summary
In the prior art, docx documents cannot directly use their data, and need to be manually copied or memorized to extract data, resulting in low efficiency when multiple documents are repeatedly entered, low data utilization rate, and they need to fill in and print signatures and seals, which increases work complexity and paper waste.
The file parsing unit, mapping rule configuration unit and data entry unit are used to parse docx documents through Python, identify differentiated fields and identify mapping encoding, automatically configure mapping relationships, and use Selenium units to realize automatic data entry and electronic signature, solving the problem of repeated data entry among multiple systems.
It realizes automatic extraction and entry of docx document data, improves work efficiency, reduces manual operations, reduces data error rate and paper consumption, and improves data utilization rate.
Smart Images

Figure CN114186549B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer information processing, and particularly to a text file processing technology, a method for writing, extracting and utilizing docx file data. Background Art
[0002] The traditional way of processing document services is to edit text in a docx document, download and print it, use it after signing and stamping, and then manually input the paper document data into the system. The data in the docx document cannot be directly utilized, which greatly reduces work efficiency and increases the data error rate, making the entire business process complex and lengthy.
[0003] Chinese Patent Application for Invention CN110083843A, a CAD drawing translation method discloses extracting the text content on a CAD file through Python parsing objects for manual or machine translation, and then backfilling and entering it into the CAD file. This method effectively solves the learning cost problem of translators for CAD. This application only performs parsing processing and utilization on one file type, CAD, and does not make corresponding solutions for other file types. Moreover, the parsed data is finally backfilled into the CAD file without further beneficial applications.
[0004] Chinese Invention Patent CN107797978A, a method and system for an input area in a document for a handwriting device. A server generates a form identifier to identify a page or an input area of a document; generates a position and a field type for the input area of the document; associates the position and the field type with the form identifier; and uses an identifier represented in a graphic converted from the form identifier to copy a second document. The position, the field type, and the form identifier are stored in the metadata of the document. A client device obtains the form identifier converted from the identifier represented in the graphic from the handwriting device. The form identifier is associated with the position and the field type for the input area of the second document. The form identifier, the position, and the field type are stored in the metadata of the second document. The client device obtains a position signal of the handwriting from the handwriting device, and associates the position signal with the input area based on the form identifier, the position, and the field type.
[0005] Chinese Patent Application CN108959626A, a method for efficiently and automatically generating cross-platform heterogeneous data briefs, discloses a method for efficiently and automatically generating cross-platform heterogeneous data briefs, which centrally manages data using the SX404DB key-value database; the SX404DB key-value database is a key-value NoSQL database based on the inverted index technology; the brief content is dynamically generated through the DocumentScript script control system; and it is completed by injecting content into a format template based on Office OpenXML and compressing it into a DOCX format document. This method has the characteristics of supporting massive heterogeneous data, flexible and extensible content generation methods, stable brief formats and high compatibility, and has good stability, operability and scalability.
[0006] The technologies disclosed in the above patent documents only complete the internal processing and association of document data, and the obtained data is applicable for internal use within the system, unable to achieve effective interaction and utilization with external systems. They are standard format electronic forms formed for specific business scenarios and belong to a one-to-one parsing and utilization method. If interaction with external systems is required, it still needs to be manually completed. DOCX documents cannot be utilized. When submitting a DOCX document to the system, the data in the required areas of the document cannot be extracted and utilized. It can only be extracted and reused through manual copying or memorization, reducing work efficiency and increasing data. For the data entry of multiple duplicate DOCX documents, for multiple DOCX documents with different contents, when duplicate data needs to be filled in, the traditional method can only extract and reuse it through manual copying or memorization, reducing work efficiency and increasing data. Staff need to complete the duplicate entry of data in multiple systems. When a piece of material needs to be entered into multiple systems, a large number of basic data information needs to be repeatedly entered. Summary of the Invention
[0007] In view of the problems in the prior art that when multiple documents need to be repeatedly entered into multiple systems, a large number of basic data need to be repeatedly entered, the processing speed is slow, and the data utilization rate is low, the present invention proposes a method for obtaining, entering and utilizing text file data. The present invention uses a text file processing technology, and the method for writing, extracting and utilizing DOCX file data can effectively solve the problems of internal and external data utilization. It can be combined with electronic signature and electronic seal technologies to solve the problem that electronic documents must be filled in, printed, signed and sealed, and automatically complete the repeated entry of key data, improve work efficiency, reduce paper waste, and increase data utilization efficiency.
[0008] The technical solution of the present invention to solve the above technical problems is a document service processing and data utilization system, including: a file parsing unit, a mapping rule configuration unit, a data entry unit, and a selenium unit. The file parsing unit parses the uploaded docx document blank template and parsing template, adds read position example data to the parsing template corresponding to the blank template, identifies different fields, and marks mapping codes; the mapping rule configuration unit uses the mapping codes to configure the mapping relationship with the data to be filled address in the blank template, selects the fields to be used from the parsing template, and marks unique codes; the data entry unit extracts the required structured data from the specified address of the parsing template through the marked unique code, and transmits the structured data with the mapping relationship to the selenium unit. The Selenium unit uses the configured data to be filled address, calls the browser to automatically open the data usage address, and through the configured mapping relationship, the data is automatically entered into the form input box to generate an electronic document, and the selenium unit uploads the completed electronic document to the corresponding path.
[0009] Traverse the paragraphs and tables between the blank template and the parsing template, compare the text content differences of each corresponding paragraph or table in the two template documents, determine the content to be filled and its position, and determine the type of content to be filled through the context relationship of the different content; parse the parsing template and the blank template to obtain the parsing document, mark the mapping code for the different content of the two document documents through the parsing document, and then match the mapping code with the address identification ID number of the content to be filled in the blank template.
[0010] Process the blank template table and the corresponding parsing template table filled with any content through an automated configuration file generation script, automatically obtain the configuration file of the table, and verify the configuration file.
[0011] A configuration file contains several lines of configuration information, and each line of configuration information contains three fields: extraction position, extraction method, and storage label. Information extraction for any word document that conforms to the specification is achieved through the cooperation of the above three fields with the configuration information. The extraction methods include: the diff mode for extracting different information in the extraction position label of the two documents, the full mode for extracting all information in the extraction position label, the ldiff mode for extracting from the first different character in the two documents in the position label, the cbox mode for extracting information in the checkbox in the extraction position label, and the lines mode for extracting the entire line and the following lines in the extraction position label.
[0012] When using the diff mode: According to the mapping code marked for the different content of the two document documents provided by the parsing document, extract relevant content from the corresponding position of the parsing template according to the different content. The different content can be divided into multiple paragraphs, and the array length of the text content extracted from the parsing template is not fixed.
[0013] When using the full mode: Obtain the positions in the parsing template where content needs to be filled, extract all the information at the corresponding marked positions according to a fixed length and store it in an array. The length of the returned array is a fixed length. When using the ldiff mode: According to the different positions in the parsing template, starting from the first different position of the blank template and the parsing template, extract all the subsequent information and return it. The length of the array returned by this mode is fixed. When using the cbox mode: According to the marked different positions in the parsing document, select and extract the information selected by the check box, and load the extracted content into an array and return it. The length of the array returned by this mode is not fixed. When using the lines mode: Configure the lines mode at the starting position of information extraction to extract the information in the variable-length form.
[0014] According to the mapping code identified in the parsing document and the address identification ID number of the matching content, extract fixed-length or variable-length content from the corresponding positions in the parsing template according to different extraction methods set for the extracted objects of text, table, and check box, and store it in an array.
[0015] The present invention also proposes a method for docx document service processing and data utilization, including: The file parsing unit parses the uploaded docx document blank template and parsing template, adds example data for reading positions in the parsing template corresponding to the blank template, identifies differential fields and identification mapping codes; The mapping rule configuration unit uses the mapping code to configure the mapping relationship with the data to be filled address in the blank template, selects the fields to be utilized from the parsing template, and identifies unique codes; The data entry unit extracts the required structured data from the specified address in the parsing template through the identified unique code, transmits the structured data with the mapping relationship to the selenium unit. The Selenium unit uses the configured data to be filled address, calls the browser to automatically open the data usage address, and through the configured mapping relationship, the data is automatically entered into the form input box to generate an electronic document. The selenium unit uploads the completed electronic document to the corresponding path.
[0016] The data entry unit uses online editing to automatically evoke the local word document editing unit, complete the data entry in the document blank template that needs to be filled, save and refresh, transmit the data back to the file parsing unit, parse the completed file, and fill it to the corresponding position in the blank template according to the configured mapping relationship for front-end storage in the form of structured data.
[0017] Traverse the paragraphs and tables between the blank template and the parsing template, compare the text content differences of each corresponding paragraph or table in the two template documents, determine the content to be filled and its location, and determine the type of content to be filled through the context relationship of the different content. Parse the parsing template and the blank template to obtain the parsing document. Identify and map the encoding of the different content in the two document templates through the parsing document, and then match the mapped encoding with the address identification ID number of the content to be filled in the blank template. Process the blank template and the parsing template table filled with any content through an automated configuration file generation script to automatically obtain the configuration file of the table and verify the configuration file. According to the mapped encoding identified in the parsing document and the address identification ID number of the matched content, extract the content of fixed length or variable length from the corresponding position in the parsing template according to different extraction methods set for text, table, and checkbox, and store it in an array.
[0018] A configuration file contains several lines of configuration information. Each line of configuration information includes three fields: extraction location, extraction method, and storage label. The information extraction of any word document is realized through the cooperation of the above three fields and the configuration information. The extraction methods include: the diff mode for extracting different information in the two documents in the extraction location label, the full mode for extracting all information in the extraction location label, the ldiff mode for extracting from the first different character in the two documents in the location label, the cbox mode for extracting information in the checkbox in the extraction location label, and the lines mode for extracting the entire line content from the current line and below in the extraction location label. When using the diff mode: According to the mapped encoding of the different content in the two documents, extract the different content with variable array length from the corresponding position in the parsing template; when using the full mode: Extract all the information at the corresponding position label according to the fixed array length and store it in the array; when using the ldiff mode: Start from the first different position of the two templates and extract all the subsequent information from the parsing template according to the fixed array length; when using the cbox mode: Extract the information selected in the checkbox according to the marked different position in the parsing document; when using the lines mode: Configure the lines mode at the starting position of information extraction to extract the information in the form.
[0019] The present invention parses and extracts the application data of docx document files through Python. Compared with the XML method, it parses and extracts the mainstream text document format (docx) files, can customize the parsing area, has better compatibility, solves the parsing of personalized templates, and supports one-to-many form filling; it can not only extract relevant data, but also write the extracted relevant data for reuse, and direct the data to be automatically filled into the corresponding positions of other doc documents at the specified path, improving work efficiency and reducing costs. Description of the Drawings
[0020] Figure 1 This is the flowchart of the text document processing method of the present invention;
[0021] Figure 2 It is the mapping relationship corresponding to the mapping code and the target expression;
[0022] Figure 3 It is the business process flowchart of mapping implementation;
[0023] Figure 4 It is the schematic diagram of the document extraction position mark;
[0024] Figure 5 It is the flowchart of document automatic configuration;
[0025] Figure 6 It is the document information extraction process;
[0026] Figure 7 This is the schematic diagram of the document business processing system of the present invention. Specific embodiments
[0027] Next, the technical solution of the present invention will be clearly and completely described in conjunction with specific examples and the accompanying drawings. This embodiment is a part of the embodiments of the present invention, and the protection scope of the present invention is not limited to the following embodiments.
[0028] As Figure 1 shown is the flowchart of the text document processing method of the present invention.
[0029] Upload a blank document docx (text format of Microsoft Word) template document without filled content as a blank template, and store the docx file with the filled content as the corresponding parsing template in the system Python-docx library (document configuration library, a database for creating or updating Word files, providing computing functions for creating or updating.doc\.docx and other files, adding and configuring paragraphs, pictures, tables, texts, etc.). Traverse the paragraphs and tables between the blank template and the parsing template, compare the text content differences of each corresponding paragraph or table in the two files, determine the content to be filled and its position, and determine the type of the filled content through the context relationship of the different content. Parse the parsing template and the blank template to obtain the parsing document, mark the mapping code for the document difference content through the parsing document, and then match the mapping code with the address identification ID number of the content to be filled in the blank template.
[0030] As Figure 2 shown is the mapping relationship corresponding to the mapping code and the target expression. The parsing rule configuration module fills the mapping code in the parsing template and the target expression to be filled into a group to form the corresponding mapping relationship.
[0031] For the array data stored in the file parsing unit, it will be returned to the selenium unit of the application module so that the selenium unit can transfer the corresponding data to the specified address and system.
[0032] In the background configuration system module, through the mapping rule configuration unit, the mapping encoding and the target expression can be customarily associated and configured, that is, the corresponding document data is selected for the target position to be filled. By parsing the mapping encoding of each structured data in the document, for example, T-0-0-1 represents the applicant unit name field in the document, which can be associated with the corresponding target filling form position. If the form position id = "orgName", the value filled in T-0-0-1 in the document can be extracted as the structured data "applicant unit name" and passed to the selenium unit, and then transmitted to the position with id = "orgName" of the specified address and system. At the same time, the selenium unit will submit the file to the corresponding list in the mapping rule configuration unit.
[0033] Such as Figure 3 As shown in the mapping implementation business flow chart. It mainly includes the following processes: obtaining structured data, obtaining mapping relationships, automatically matching mapping addresses, finding and filling in the correct form input positions through mapping relationships, and automatically submitting electronic documents to the corresponding paths.
[0034] Obtain structured data through the configuration process and parsing logic.
[0035] In the mapping rule configuration unit, the configuration file with mapping relationship information is set and stored in csv format. A configuration file contains several lines of configuration information, and each line of configuration information contains three fields: extraction position, extraction method, and storage label (which can be represented by three fields separated by English commas). Through the above three fields and the configuration information, information extraction can be realized for any word document that meets the specifications. The extraction methods of the three fields are introduced separately below.
[0036] For any template document, according to the mapping encoding that identifies the difference content between the two files in the parsed document, the content extraction position is determined by setting marks for the docx / doc file format of the parsed template. Two formats including paragraph (P) and table (T) marks can be set. Therefore, the "extraction position" starts with P or T. For the identifier starting with P, two numbers separated by the symbol "-" are set behind it. The former represents the position of the paragraph, and the latter represents the information of the nth blank in the extracted paragraph.
[0037] Such as Figure 4The following is a schematic diagram of the extraction position marker. If you want to extract the information on the space after "There is", you need to give the position of this space. Since this space is in a paragraph, the first digit of the configuration file is P; since it is in the 0th paragraph, the second digit is 0; since it is the 0th space to be filled with information in this paragraph, the third digit is 0. Then the "extraction position" of this position configuration file is P-0-0.
[0038] Next is the configuration of the "extraction method". Through the flexible configuration of this field, the parsing algorithm can flexibly extract information in various formats. As shown in the following table, there are five extraction methods.
[0039] Table 1: Extraction Methods
[0040]
[0041] According to the mapping code marked in the parsing document, match the address identification ID number of the content. According to the different extraction methods set for the extraction objects of text, table, and checkbox, extract fixed-length or variable-length content from the corresponding positions in the parsing template and store it in an array. The extraction methods include: extracting the different information between the two documents in the extraction position marker, extracting all the information in the extraction position marker, starting to extract from the first different character in the two documents in the position marker, extracting the information in the checkbox in the extraction position marker, and extracting the entire line and the following lines in the extraction position marker.
[0042] A configuration file contains several lines of configuration information. Each line of configuration information includes three fields: extraction position, extraction method, and storage label. Through the cooperation of the above three fields and the configuration information, information extraction can be realized for any word document that conforms to the specification. The extraction methods include: the diff mode for extracting the different information between the two documents in the extraction position marker, the full mode for extracting all the information in the extraction position marker, the ldiff mode for starting to extract from the first different character in the two documents in the position marker, the cbox mode for extracting the information in the checkbox in the extraction position marker, and the lines mode for extracting the entire line and the following lines in the extraction position marker.
[0043] The diff (difference extraction) mode: It is used to extract the different content at the same position in the input document and the template document. For example, in the above table, s1 is a blank template table, that is, an empty table without filled content, and s2 is the parsing template table. It is necessary to extract relevant information from the table s2 and fill it in the corresponding position of the table s1.
[0044] When extracting in diff mode, according to the mapping code of the document difference content identification provided by the parsed document, relevant content is extracted from the corresponding position of the parsing template according to the difference content and stored in the corresponding array of the document parsing unit. The extracted array can be divided into multiple segments according to the content (as shown in Table 2 above). Therefore, the length of the array for extracting text content from the parsing template through this mode is not fixed.
[0045] Full (full extraction) mode: It will not compare the input content with the template content. Instead, it obtains the positions where content needs to be filled from the parsing template and directly extracts all the information at the corresponding extraction position markers in a fixed length and stores it in an array. Therefore, the length of the array returned by this mode is a fixed length, which is 1 or 0.
[0046] Ldiff (difference matching) mode: It performs lazy matching on the input content and the output content. According to the difference positions in the parsing template, starting from the first different position of the blank template and the parsing template, all the subsequent information is extracted and returned. Therefore, the length of the array returned by this mode is also fixed, which is 1 or 0.
[0047] Cbox (checkbox extraction) mode: It is used to extract the content of checkboxes. According to the marked difference positions in the parsed document, the selected information of the checkboxes is selected and extracted, and the extracted content is loaded into an array and returned. The length of the array returned by this mode is also not fixed.
[0048] Lines (row extraction) mode: It is used to process table content with an indefinite length. For example, in the above table, it is impossible to determine how many rows of information the user will fill in the actual business. In response to this situation, the extraction method of the lines mode is configured at the starting position of information extraction. By traversing all the rows filled in by the user and removing the text content of each row, the information in the variable-length form can be extracted.
[0049] A word document may have a dozen or even dozens of pieces of information to be extracted. For this reason, an automated configuration file generation script is provided. By using the automated configuration file generation script to process the blank template table and the parsed template table filled with any content, the configuration file of this table is automatically obtained and the configuration file is verified.
[0050] As Figure 5 shown is the document automatic configuration flowchart. The formats of the blank template table (template table) and the parsed template table (comparison table) are the same. The parsed template table contains complete information, and the content of adjacent cells in the same row is not completely the same. The configuration generation module analyzes the different parts of the blank template table (empty table) of the docx file and the parsed template table (comparison table) filled with information to allocate the relative index positions of the information to be extracted.
[0051] Generate a visual docx file for automated form filling configuration in the mapping rule configuration unit, determine the position information corresponding to the fields to be filled, generate a configuration file, and each line of the configuration file contains three fields: extraction position, extraction method, and storage label. Generate a configuration table according to the configuration file.
[0052] As Figure 6 shown is the document information extraction process. According to the configuration table, extract the key information in the docx file that needs to be filled in the corresponding position of the blank template from the parsing template. For example, the extracted information can be in the format {position identifier: [information label, [infol, info2,...]]}, which contains the data information included in the corresponding position. For example: {T-2-1-3: ['ID number', ['510502199711111111']]}. Save the extracted information in the dictionary format Dict().
[0053] Figure 7 Shown is a schematic diagram of the document business processing system of the present invention.
[0054] The document business processing system of the present invention includes: a background configuration system module and an application program module. Among them, the background configuration system module includes: a file parsing unit and a mapping rule configuration unit; the application program module includes: a data entry unit and a selenium unit.
[0055] The data entry unit of the application program module uses online editing to automatically activate the local word document editing unit, and can use methods such as manual entry, ID card reading entry, OCR recognition entry, etc. to complete the entry of the data that needs to be filled in the document template and save and refresh it.
[0056] The data entry unit transmits the data back to the file parsing unit. After parsing the completed file, it then transmits the structured data with mapping relationships to the selenium (web automation tool) unit, and stores it front-end in the application program module in the form of structured data through the mapping relationships.
[0057] The Selenium unit uses the filled address configured by the mapping rule configuration unit to automatically open the data usage address by calling the browser, and then completes the automatic entry of data into the form input box to generate an electronic document through the configured mapping relationship. The Selenium unit uploads the completed electronic document to the corresponding path.
Claims
1. A document service processing and data utilization system, characterized in that, Including: A file parsing unit, a mapping rule configuration unit, a data entry unit, and an automated tool Selenium unit. The file parsing unit parses the uploaded blank template and parsing template of the docx document. The blank template is a blank docx template document without filled content, and the parsing template is a docx file with filled content; Add read position example data to the parsing template corresponding to the blank template, identify different fields, and mark mapping codes. The mapping rule configuration unit uses the mapping codes to configure the mapping relationship with the data filling address in the blank template, selects the fields to be used from the parsing template, and marks unique codes; The data entry unit extracts the required structured data from the specified address of the parsing template through the marked unique code, transfers the structured data with the mapping relationship to the Selenium unit. The Selenium unit uses the configured data filling address, automatically opens the data usage address by calling the browser, and through the configured mapping relationship, the data is automatically entered into the form input box to generate an electronic document. The Selenium unit uploads the completed electronic document to the corresponding path.
2. The system according to claim 1, wherein The data entry unit uses online editing to automatically evoke the local word document editing unit, enters and saves the data to be filled in the document blank template, refreshes it, and transmits the data back to the file parsing unit. The completed file is parsed, and filled in the corresponding position of the blank template according to the configured mapping relationship, and stored in the front end in the form of structured data.
3. The system according to claim 1, wherein Process the blank template and the corresponding parsing template filled with any content through an automated configuration file generation script to automatically obtain the configuration file and verify the configuration file.
4. The system according to claim 3, wherein A configuration file contains several lines of configuration information. Each line of configuration information includes three fields: extraction position, extraction method, and storage label. The information extraction of any word document is realized through the cooperation of the above three fields. The extraction methods include: diff mode for extracting different information between the two documents in the extraction position label, full mode for extracting all information in the extraction position label, ldiff mode for extracting from the first different character between the two documents in the position label, cbox mode for extracting information in the checkbox in the extraction position label, and lines mode for extracting the entire line content from the current line and below in the extraction position label.
5. The system according to any one of claims 1-4, characterized in that According to the mapping code marked in the parsed document and the address identification ID number of the matching content, different extraction methods are set according to whether the extraction object is text, table, or checkbox, and content with a fixed length or an unfixed length is extracted from the corresponding position in the parsing template and stored in an array.
6. A method for document service processing and data utilization, characterized in that, Including: Take the blank docx template document without filled content as the blank template, and the docx file with filled content as the parsing template. The file parsing unit parses the uploaded blank template and parsing template of the docx document, adds read position example data to the parsing template corresponding to the blank template, identifies different fields, and marks mapping codes; The mapping rule configuration unit uses mapping encoding to configure the mapping relationship with the data filling addresses in the blank template, selects the fields to be used from the parsed template, and identifies unique encodings. The data entry unit extracts the required structured data from the specified address of the parsed template through the identified unique encoding, transfers the structured data with the mapping relationship to the automated tool selenium unit. The selenium unit uses the configured data filling addresses, automatically opens the data usage address by calling the browser, and through the configured mapping relationship, the data is automatically entered into the form input box to generate an electronic document. The selenium unit uploads the completed electronic document to the corresponding path.
7. The method according to claim 6, wherein The data entry unit uses online editing to automatically invoke the local word document editing unit, completes the data entry for the data to be filled in the document blank template, saves and refreshes, and transfers the data back to the file parsing unit. The completed file is parsed and filled in the corresponding positions of the blank template according to the configured mapping relationship, and stored in the front end in the form of structured data.
8. The method according to claim 6, characterized in that, Traverse the paragraphs and tables between the blank template and the parsed template, compare the text content differences of each corresponding paragraph or table in the two template documents, determine the content to be filled and its position, and determine the type of content to be filled through the context relationship of the different content. Parse the parsed template and the blank template to obtain the parsed document, identify the mapping encoding for the different content of the two document templates through the parsed document, and then match the mapping encoding with the address identification ID number of the content to be filled in the blank template.
9. The method according to claim 6, characterized in that, Process the blank template and the parsed template table filled with any content through the automated configuration file generation script, automatically obtain the configuration file of the table, and verify the configuration file.
10. The method according to claim 9, wherein According to the mapping encoding identified in the parsed document and the address identification ID number of the matched content, extract the content of fixed length or variable length from the corresponding position in the parsed template into an array according to different extraction methods set for the extracted object being text, table, or checkbox.
11. The method according to claim 9 or 10, characterized in that, A configuration file contains several lines of configuration information. Each line of configuration information includes three fields: extraction position, extraction method, and storage label. Through the cooperation of the above three fields with the configuration information, information extraction from any word document is achieved. The extraction methods include: the diff mode for the different information in the two documents in the extraction position indication, the full mode for all the information in the extraction position indication, the ldiff mode for extracting from the first different character in the two documents in the position indication, the cbox mode for the information in the checkbox in the extraction position indication, and the lines mode for the entire line and the following lines in the extraction position indication.
12. The method according to claim 11, characterized in that, When using the diff mode: According to the mapping code identified by the differences between the two documents, extract the difference content with an unfixed array length from the corresponding position of the parsing template; when using the full mode: Extract all the information at the marked position according to the fixed array length and store it in an array; when using the ldiff mode: Start from the first different position of the two templates, and extract all the subsequent information from the parsing template according to the fixed array length; when using the cbox mode: Extract the information selected by the check box according to the marked difference position in the parsed document; when using the lines mode: By traversing all the lines filled in by the user, remove the text content of each line, and extract the information in the variable-length form.
Citation Information
Patent Citations
Method and system for input areas in documents for handwriting devices
CN107797978A
An efficient automatic generation method of cross-platform heterogeneous data briefing
CN108959626A
CAD drawing translation method
CN110083843A
Format form auto-filling method, device, compute device and storage medium
CN109308350A
Template type legal document information filling method and device
CN110096689A