Template-based data processing method and device and electronic equipment
Through predefined data analysis and conversion templates, unstructured data is automatically processed, which solves the problem of complex data analysis error in the existing technology, and realizes efficient and accurate data conversion and import.
Patent Information
- Application Number
- CN202510041815.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-26
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has analytical errors when processing complex or fuzzy unstructured data, especially when identifying text in handwritten or low-quality images, with low text accuracy.
Automatically identify and parse unstructured data in different formats through predefined data analysis templates and data conversion templates. The specific steps include obtaining the type of the target file, selecting the matching data analysis template for file content conversion, determining the target table object, and converting the data into structured data through the matching data conversion template.
It improves the efficiency and accuracy of data processing, reduces the time and labor cost of manual processing, ensures the consistency and standardization of data conversion, and improves the accuracy of data import.
Smart Images

Figure CN120067190A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a template-based data processing method, apparatus, and electronic device. Background Art
[0002] With the rapid development of big data and informatization, the demand for processing unstructured data is increasing day by day. Unstructured data mainly includes various formats of documents such as PDF, Word, Excel, CSV, TXT, HTML, and XML. These data widely exist in enterprises and organizations, but their diversity and complexity bring great challenges to data processing, analysis, and management. This requires data conversion of unstructured data.
[0003] In related technologies, optical character recognition technology can be used to process unstructured data. For example, for PDF and image files, optical character recognition technology can convert the text in the image into editable text. However, when processing complex or blurred data, there may still be parsing errors. For example, when optical character recognition technology recognizes the text in handwritten or low-quality images, the accuracy of the extracted text is relatively low. Summary of the Invention
[0004] The object of the present invention is to, through predefined data parsing templates and data conversion templates, for unstructured data of different formats, select appropriate data parsing templates and data conversion templates to parse and convert the unstructured data, so as to improve the efficiency and accuracy of data processing.
[0005] In a first aspect, an embodiment of the present invention provides a template-based data processing method, the method including:
[0006] Obtain the target file type of a target file, where the target file includes unstructured data;
[0007] Based on the target file type, determine a target data parsing template that matches the target file type among a plurality of predefined data parsing templates, where different file types correspond to different data parsing templates;
[0008] Convert the file content of the target file into multi-line data in a preset format through the target data parsing template;
[0009] For each line of data in the preset format, determine a target table object that matches the data in the preset format based on the data type of the data in the preset format, and store the data in the preset format in the target table object;
[0010] Among multiple predefined data conversion templates, determine a target data conversion template that matches the target table object, where the matching degree between the column labels in the target table object and the column labels of the data conversion template is greater than a preset matching degree;
[0011] Convert the data in the target table object into structured data through the target data conversion template.
[0012] Optionally, determining a target data parsing template that matches the target file type among multiple predefined data parsing templates based on the target file type includes:
[0013] For each predefined data parsing template, obtain the template feature of the data parsing template, where the template feature is used to characterize the file type of the file that the data parsing template can parse;
[0014] Match the target file type with the template feature of the data parsing template to obtain a matching result, where the matching result includes that the target file type matches the template feature of the data parsing template and that the target file type does not match the template feature of the data parsing template;
[0015] If the matching result is that the target file type matches the template feature of the data parsing template, determine the data parsing template as the target data parsing template that matches the target file type.
[0016] Optionally, for each row of data in a preset format, determining a target table object that matches the data in the preset format based on the data type of the data in the preset format includes:
[0017] For each row of data in a preset format, determine the data type of the data in the preset format, where the data type includes row data, column data, and extended data;
[0018] If the data type of the data in the preset format is row data, the page number corresponding to the data in the preset format is the same as the page number of the table object, and the data in the preset format matches the column data or row data of the predefined table object, determine that the data in the preset format matches the table object and determine the table object as the target table object.
[0019] Optionally, determining that the data in the preset format matches the column data or row data of the table object includes:
[0020] If the number of columns of the data in the preset format is the same as the number of columns of the column data in the predefined table object, or the row numbers of the data in the preset format are consecutive and the error in the number of columns is less than the first threshold compared to the column data in the predefined table object, it is determined that the data in the preset format matches the column data of the table object;
[0021] Or,
[0022] If the number of columns of the data in the preset format is the same as the number of columns of the row data in the predefined table object, or the row numbers of the data in the preset format are the same and the error in the number of columns is less than the second threshold compared to the row data in the predefined table object, it is determined that the data in the preset format matches the row data of the table object.
[0023] Optionally, the converting the data in the target table object into structured data through the target data conversion template includes:
[0024] Obtaining dictionary information required in the data processing process from the dictionary configuration node configured in the target data conversion template;
[0025] Obtaining table object processing logic from the table object conversion configuration node configured in the target data conversion template, and converting the data in the target table object based on the dictionary information and the table object processing logic to obtain structured data.
[0026] Optionally, before converting the data in the target table object into structured data through the target data conversion template, the method further includes:
[0027] Obtaining sample data from the data in the target table object;
[0028] Converting the sample data through the target data conversion template to obtain converted data;
[0029] If the conversion process of the sample data and the converted data meet the preset conditions, execute the step of converting the data in the target table object into structured data through the target data conversion template;
[0030] Wherein, the preset conditions include: no dictionary data is missing during the conversion process of the sample data, the data processing component of the target data conversion template is executable, there is no missing data field mapping and the data type is correct in the converted data.
[0031] Optionally, the method further includes:
[0032] Obtaining preset data quality configuration rules;
[0033] Perform data verification on each row of the structured data based on the data quality configuration rules to obtain the score corresponding to each row of data;
[0034] According to a preset data quality evaluation formula, calculate the scores corresponding to multiple rows of data of the structured data to obtain the score corresponding to the structured data, and the score corresponding to the structured data is used to evaluate the accuracy of the structured data.
[0035] In a second aspect, an embodiment of the present invention provides a template-based data processing device, and the device includes:
[0036] A file type acquisition module, configured to acquire the target file type of a target file, where the target file includes unstructured data;
[0037] A data parsing template determination module, configured to determine a target data parsing template that matches the target file type from a plurality of predefined data parsing templates based on the target file type, where different file types correspond to different data parsing templates;
[0038] A file content conversion module, configured to convert the file content of the target file into multiple rows of data in a preset format through the target data parsing template;
[0039] A target table object matching module, configured to, for each row of data in the preset format, determine a target table object that matches the data in the preset format based on the data type of the data in the preset format, and store the data in the preset format in the target table object;
[0040] A data conversion template determination module, configured to determine a target data conversion template that matches the target table object from a plurality of predefined data conversion templates, where the matching degree between the column labels in the target table object and the column labels of the data conversion template is greater than a preset matching degree;
[0041] A data conversion module, configured to convert the data in the target table object into structured data through the target data conversion template.
[0042] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0043] At least one processor;
[0044] A memory for storing executable instructions of the at least one processor;
[0045] Wherein, the at least one processor is configured to execute the instructions to implement the method described in the first aspect.
[0046] Fourthly, an embodiment of the present invention provides a computer-readable storage medium. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the method described in the first aspect.
[0047] For the technical solution provided by the embodiment of the present invention, the target file type of a target file including unstructured data is obtained. Based on the target file type, a target data parsing template matching the target file type is selected from predefined data parsing templates, and the file content of the target file is converted into data in a multi-line preset format through the target data parsing template. Then, a target table object is determined through the data type of each line of the preset format data, and the preset format data is stored in the target table object. The target table object is a data carrier for the multi-line preset format data. Finally, in a plurality of predefined data conversion templates, a target data conversion template matching the target table object is determined, and the data in the target table object is converted into structured data through the target data conversion template.
[0048] It can be seen that through predefined data parsing templates and data conversion templates, for unstructured data in different formats, appropriate data parsing templates and data conversion templates can be selected to parse and convert the unstructured data, that is, convert the unstructured data into structured data, thereby improving the efficiency and accuracy of data processing. Description of the Drawings
[0049] Figure 1 It is a flowchart of a template-based data processing method provided by an embodiment of the present invention;
[0050] Figure 2 is Figure 1 a flowchart of the specific implementation of S120 in
[0051] Figure 3 is Figure 1 a flowchart of the specific implementation of S140 in
[0052] Figure 4 is Figure 1 a flowchart of the specific implementation of S160 in
[0053] Figure 5 It is a flowchart of another template-based data processing method provided by an embodiment of the present invention;
[0054] Figure 6 It is a flowchart of another template-based data processing method provided by an embodiment of the present invention;
[0055] Figure 7 It is a structural schematic diagram of a template-based data processing device provided by an embodiment of the present invention;
[0056] Figure 8 Schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0057] The present invention will be described in detail below through embodiments.
[0058] With the rapid development of big data and informatization, the demand for processing unstructured data is increasing day by day. Unstructured data mainly includes various formats of documents such as PDF, Word, Excel, CSV, TXT, HTML, and XML. These data widely exist in enterprises and organizations, but their diversity and complexity bring great challenges to data processing, analysis, and management. This requires data conversion of unstructured data.
[0059] In related technologies, unstructured data can be processed through natural language processing technology and optical character recognition technology. Although natural language processing technology and optical character recognition technology have made great progress, there may still be parsing errors when processing complex or ambiguous data. For example, when optical character recognition technology recognizes text in handwritten or low-quality images, the accuracy of the extracted text is relatively low.
[0060] In view of the above technical problems existing in the prior art, an embodiment of the present invention provides a template-based data processing method. In the embodiment of the present invention, a large number of data parsing templates and data conversion templates are predefined to automatically identify and parse unstructured data in different formats. Specifically, for different unstructured data, appropriate data parsing templates can be selected based on the data type of the unstructured data to parse the unstructured data, and appropriate data conversion templates can be selected to convert the unstructured data, so as to convert the unstructured data into structured data and can be efficiently imported into the target data storage system, thereby improving the efficiency and accuracy of data processing.
[0061] Moreover, the predefined template library can include data parsing templates and data conversion templates in various formats, supporting flexible expansion and update. In addition, after converting unstructured data into structured data, the integrity, consistency, and correctness of the structured data can be verified to ensure the quality of the structured data. Also, it supports desensitization processing for key field information during data reading and preview.
[0062] A template-based data processing method, device, and electronic device provided by an embodiment of the present invention will be elaborated in detail below.
[0063] As Figure 1 shown, a template-based data processing method provided by an embodiment of the present invention may include the following steps:
[0064] S110, obtain the target file type of the target file.
[0065] Among them, the target file includes unstructured data.
[0066] Specifically, the target file type can be multiple file types such as PDF, Word, Excel, CSV, TXT, HTML, XML, etc. Of course, in practical applications, it can also be other file types, and the embodiments of the present invention do not make specific limitations on this.
[0067] The specific implementation manner of obtaining the target file type of the target file can be:
[0068] First step, read the file header, file content, and file suffix of the file. Among them, the file header can be the first few bytes of the file.
[0069] Second step, judge the file type according to the file header. Specifically, when judging the file type according to the file header, it can be judged according to the file format. For example, for a PDF format file, it can be judged whether the file header is "%PDF". If the file header is %PDF, then the file type can be judged as PDF. For a PNG format file, it can be judged whether the file header is 89 504E 47 0D 0A 1A 0A. If the file header is 89 50 4E 47 0D 0A 1A 0A, the file type can be judged as PNG. For a JPEG format file, it can be judged whether the file header is FF D8 FF. If the file header is FF D8 FF, then the file type can be judged as JPEG. For an XLS format file, it can be judged whether the file header is D0 CF 11E0 A1 B1 1AE1. If the file header is D0 CF 11E0 A1 B1 1A E1, the file type can be judged as XLS.
[0070] Third step, if the file type cannot be judged through the file header, the file type can be judged according to the file content. Specifically, the specific method of judging the file type according to the file content can be: match the content of the specified part according to the relevant format specification standard. For example, for an HTML format file, it can be matched whether the head and tail are wrapped by tags. If so, the file type is judged as HTML. For a CSV format file, it can be matched whether it is separated by line breaks and whether the first two lines are separated into equal columns by commas. If so, the file type is judged as CSV.
[0071] When the file header tag is "50 4B 03 04", this tag indicates a ZIP - formatted compressed file. It is possible to parse whether its file content contains "[Content_Types].xml". If it does, then it is determined that the file type is XLSX. And so on, other file types based on the ZIP format can be identified.
[0072] Step 4, if the file type cannot be determined from the file content, then determine the file type based on the file suffix. Specifically, to determine the file type based on the file suffix, it can be judged according to the file format specification. For example, if the file suffix is xml, then it can be determined that the file type is xml; if the file suffix is txt, then it can be determined that the file type is txt.
[0073] In the process of determining the file type through the above four steps, as long as the file type is determined, the process of determining the file type can be exited.
[0074] S120, based on the target file type, determine the target data parsing template that matches the target file type from multiple predefined data parsing templates.
[0075] Among them, different file types correspond to different data parsing templates.
[0076] Specifically, the parsing template is defined by an XML - formatted file, and the specific definition is as follows:
[0077] 1. Reader: Represents the root tag, containing an "impl" attribute pointing to the template implementation. For example: <Reader impl=”TxtReaderTemplate”>, which represents the parsing template configuration for a TXT format.
[0078] 2. Name: Represents the template name.
[0079] 3. Order: Represents the template priority.
[0080] 4. Feature: Represents the template feature, which consists of multiple defined sub - tags. Through the template feature, the file type can be matched with the data parsing template. Among them, the data parsing template can define different configuration tags according to different implementations. For example, the following two different configuration tags for data parsing templates.
[0081] The first data parsing template is used to parse the target file in CSV format. This data parsing template can define two configuration tags. The first configuration tag is: matching the file format with the CSV suffix, and the second configuration tag is: the column separator is ",".
[0082] The second data parsing template is used to parse the target file in XML format. This data parsing template can define multiple configuration tags. The first configuration tag can be: matching XML format files with the path " / Document / ... / EAA"; the second configuration tag can be: not parsing tag attributes and defaulting to using the short tag name; the third configuration tag can be: defining the parsing path; the fourth configuration tag can be: defining the escape of tag names. Among them, the definition of tag name escape means converting special characters in HTML tags into corresponding entity characters so that these characters can be correctly displayed on the web page instead of being parsed as HTML tags or entities.
[0083] And so on, data parsing templates for file types such as PDF, EXCEL, WORD, HTML, images, and audio can be defined and implemented. The template engine will perform loading, verification, and matching operations on the data parsing template defined in XML format based on the target file type of the target file.
[0084] In one implementation, S120, based on the target file type, determine the target data parsing template that matches the target file type among multiple predefined data parsing templates. As Figure 2 shown, it can include the following steps:
[0085] S121, for each predefined data parsing template, obtain the template features of the data parsing template.
[0086] Among them, the template features are used to characterize the file types of files that the data parsing template can parse.
[0087] S122, match the target file type with the template features of the data parsing template to obtain a matching result.
[0088] Among them, the matching result includes that the target file type matches the template features of the data parsing template and that the target file type does not match the template features of the data parsing template.
[0089] S123, if the matching result is that the target file type matches the template features of the data parsing template, determine the data parsing template as the target data parsing template that matches the target file type.
[0090] As can be seen from the above description, the template features of the data parsing template consist of multiple defined sub-tags. By matching the target file type with the multiple sub-tags included in the template features, the data parsing template whose template features match the target file type is determined as the target data parsing template.
[0091] S130, convert the file content of the target file into data in multiple preset formats through the target data parsing template.
[0092] Specifically, after determining the target data parsing template that matches the target file type, the file content of the target file can be converted into multi-line data in a specified format through the target data parsing template, and the specified format is a pre-set format. The specific format may include the following content:
[0093] 1. lineNumber. It represents the serial number of the data generation. For example, the serial number can be a positive integer such as 0, 1, 2, etc.
[0094] 2. lineString, which represents the conversion into a string according to the order of the data generation columns, that is, the column labels are concatenated into a string. Specifically, the column labels can be concatenated according to the tab character to obtain a string. The function of the tab character is to align the text vertically by column without using a table. Common applications include lists, simple lists, etc. For example, "lineString" can be: "Serial number\tName\tID number\tID type\tMobile phone number...".
[0095] 3. Sheet, which represents the virtual page number. For example, the page number can be a positive integer such as 1, 2, 3, etc.
[0096] 4. SheetName, which represents the virtual page name. For example, it can be the basic information of personnel.
[0097] Among them, the parsing logic of the target data parsing template for the file content of the target file can be converted according to different data formats and conversion result requirements. For example, an Excel format file can parse the above 1-4 structural contents according to the specification; a CSV format file needs to generate virtual page information. For an image format file, after image type recognition and feature extraction, the form data can be converted into multi-line data in a preset format. For an audio format file, after speech recognition, it can be parsed into multi-line data in a preset format such as time, object, content format, etc. Files of other formats can also perform file content conversion through their corresponding target data parsing templates, which will not be listed one by one here.
[0098] S140. For each line of data in the preset format, based on the data type of the data in the preset format, determine the target table object that matches the data in the preset format, and store the data in the preset format in the target table object.
[0099] Specifically, in the embodiment of the present invention, an abstract table object is defined as the data carrier of the multi-line data in the preset format in S130. Specifically, an abstract table object structure is defined as the data carrier for converting into the final target table structure. The table object may include the following content:
[0100] 1. File object. The file object may include the following content:
[0101] FileType: is the parsed file type; Charset: is the encoding format used by the file; FileHolder: is the physical information of the file. List <importfiletable>: It is a set of table data object collections.
[0102] 2. Table data object. The table data object may include the following content:
[0103] ImportFileIndex: It is the index information of the table in the file; ImportFileCol: It is the column header information of the table; ImportFileExtend: It is the extended information of the table, the data before the table header and not included in the previous table; Size: It is the size of the table data.
[0104] 3. Index object. The index object may include the following content:
[0105] startLine: It is the starting line; endLine: It is the ending line; sheet: It is the page number; sheetName: It is the page name
[0106] 4. Column header object. The column header object may include the following content:
[0107] ImportFileIndex: It is the index information of the column in the file; Cols: It is the column header array; Sign: It is the column label.
[0108] 5. Extended information object. The extended information object may include the following content:
[0109] ImportFileIndex: It is the index information of the extended data in the file; Row: It is the number of rows; maxCol: It is the maximum number of columns.
[0110] In one implementation, S140, for each row of data in the preset format, based on the data type of the data in the preset format, determine the target table object that matches the data in the preset format. As Figure 3 shown, it may include the following steps:
[0111] S141, for each row of data in the preset format, determine the data type of the data in the preset format.
[0112] Among them, the data types include row data, column data, and extended data.
[0113] Specifically, after parsing multiple rows of data in the preset format through S130, for each row of data in the preset format, the data type of the data in the preset format can be judged. Among them, the data types can be row data, column data, and extended data. The specific implementation manners of judging that the data in the preset format is row data, column data, and extended data will be elaborated in detail below.
[0114] Among them, the specific implementation manner of judging whether the data in the preset format of the current row is column data can be:
[0115] Define a keyword library. When at least 60% of the column labels match the data in the keyword library, it is determined as column data. When there is data in the column label that is not in the keyword library, it is supplemented to the keyword library. If the data in the preset format is an empty line, a single column, or consecutive N characters are numbers, it is determined not to be column data.
[0116] The specific implementation method for determining whether the data in the preset format of the current row is row data can be:
[0117] If the data in the preset format of the current row is not empty, not only one column has a value, not determined as column data, and the row data is not composed of key-value pair data, this data is determined as row data.
[0118] Among them, the method for judging key-value pair data can be judged according to the defined key-value regular expression. Those skilled in the art should be able to understand the specific judgment method of key-value pair data, which will not be elaborated here.
[0119] The specific implementation method for determining whether the data in the preset format of the current row is extended data is:
[0120] If the data in the preset format of the current row is not column data and not row data, then it can be determined that the data in the preset format of the current row is extended data.
[0121] In practical applications, column data and extended data can be converted into row data. By converting column data and extended data into row data, the data in the preset format can be stored in a table object. Specifically, in multi-column data, each column represents a specific attribute or variable, and each row represents a piece of data. By converting multi-column data into row data, it is more convenient to perform operations such as data screening, sorting, and aggregation. At the same time, retaining label information can ensure that each piece of data can still be accurately traced and identified after data conversion.
[0122] Converting extended data into row data means converting multi-column data in the original data table into row data and retaining the label information of the original data in the converted row data.
[0123] S142, if the data type of the data in the preset format is row data, the page number corresponding to the data in the preset format is the same as the page number of the table object, and the data in the preset format matches the column data or row data of the predefined table object, it is determined that the data in the preset format matches the table object, and the table object is determined as the target table object.
[0124] As an implementation manner of an embodiment of the present invention, determining that the data in the preset format matches the column data or row data of the table object may include the following steps:
[0125] If the number of columns of the data in the preset format is the same as the number of columns of the column data in the predefined table object, or the row numbers of the data in the preset format are consecutive and the difference in the number of columns is less than the first threshold compared to the column data in the predefined table object, it is determined that the data in the preset format matches the column data of the table object.
[0126] Or,
[0127] If the number of columns of the data in the preset format is the same as the number of columns of the row data in the predefined table object, or the row numbers of the data in the preset format are the same and the difference in the number of columns is less than the second threshold compared to the row data in the predefined table object, it is determined that the data in the preset format matches the row data of the table object.
[0128] Specifically, by comparing the data in the preset format with the column data and row data in the table object, it can be determined whether the data in the preset format matches the column data or the row data of the table object, and then the target table object that matches the data in the preset format can be accurately determined.
[0129] When receiving multiple lines of data in the preset format parsed in step S140, object conversion can be performed according to the following steps:
[0130] S150, among the multiple predefined data conversion templates, determine the target data conversion template that matches the target table object.
[0131] Among them, the matching degree between the column labels in the target table object and the column labels of the data conversion template is greater than the preset matching degree.
[0132] Specifically, the target table object is the data carrier of multiple lines of data in the preset format. After storing the multiple lines of data in the preset format into the target table object, the data conversion template is matched through the target table object. If the number of matching pairs between the column labels in the target table object and the column labels of a data conversion template is greater than the preset number of matching pairs, it indicates that the matching degree between the column labels of the target table object and the column labels of this data conversion template is relatively high, and the data in the target table object can be converted through this data conversion template. Therefore, this data conversion template can be determined as the target data conversion template that matches the template table object. Among them, the matching degree can be the similarity between the column labels in the standard table object and the column labels of a data conversion template, and the preset matching degree can be determined according to the actual situation. For example, the preset matching degree can be 90%, and the embodiments of the present invention do not make specific limitations on this.
[0133] For example, if the column labels of a data conversion template are exactly the same as the column labels of the shopping order table object, then this data conversion template is the data conversion template that matches the shopping order table object. The data conversion template can configure information such as the data source, target table, target field mapping, table data processor, and dictionary information for parsing dates.
[0134] S160. Convert the data in the target table object into structured data through the target data conversion template.
[0135] Specifically, after determining the target data conversion template, two methods can be set to automatically execute the target data conversion template operation or manually execute the target data conversion template operation. And through the target data conversion template, the data in the target table object is efficiently and accurately converted into structured data, that is, the unstructured data in the target file is efficiently and accurately converted into structured data.
[0136] In addition, in practical applications, if the target table object does not match the target data conversion template, the information of the failed data conversion template matching can be pushed to the user interface, and the user can add a new data conversion template and import it into the data conversion template library.
[0137] In one implementation, S160. Convert the data in the target table object into structured data through the target data conversion template, as Figure 4 shown, it can include the following steps:
[0138] S161. Obtain the dictionary information required during the data processing from the dictionary configuration node configured in the target data conversion template.
[0139] S162. Obtain the table object processing logic from the table object conversion configuration node configured in the target data conversion template, and convert the data in the target table object based on the dictionary information and the table object processing logic to obtain structured data.
[0140] Specifically, the target data conversion template has a dictionary configuration node and a table object conversion configuration node. The dictionary information required during the data processing can be obtained from the dictionary configuration node, the table object processing logic can be obtained from the table object conversion configuration node, and the data in the target table object is converted based on the dictionary information and the table object processing logic to obtain structured data.
[0141] The process of converting the data in the target table object into structured data can be:
[0142] 1. Load the target data conversion template;
[0143] 2. Since there are usually multiple target table objects, the multiple target table objects can be sorted according to the processing priority order of the multiple target table objects to process the multiple target table objects in sequence.
[0144] 3. There are usually multiple target data conversion templates, and the multiple target data conversion templates can be sorted according to the processor priority order of the multiple target data conversion templates.
[0145] 4. Load the data in the table object one by one and process the data according to the table object processing logic.
[0146] 5. Convert the data in the table object into the target table field structure according to the mapping relationship.
[0147] 6. Execute the data model according to the characteristics of the converted data, supplement the missing data in the conversion process, and obtain the converted structured data.
[0148] The technical solution provided by the embodiment of the present invention is to obtain the target file type of the target file including unstructured data, select the target data parsing template matching the target file type from the predefined data parsing templates based on the target file type, convert the file content of the target file into data in multiple preset formats through the target data parsing template, determine the target table object based on the data type of each line of data in the preset format, and store the data in the preset format in the target table object. The target table object is the data carrier of the data in multiple preset formats. Finally, among the predefined multiple data conversion templates, determine the target data conversion template matching the target table object, and convert the data in the target table object into structured data through the target data conversion template.
[0149] It can be seen that through the predefined data parsing template and data conversion template, for unstructured data in different formats, appropriate data parsing templates and data conversion templates can be selected to parse and convert the unstructured data, that is, convert the unstructured data into structured data, thereby improving the efficiency and accuracy of data processing.
[0150] Based on the embodiment shown in Figure 1 , before converting the data in the target table object into structured data through the target data conversion template, as shown in Figure 5 , this template-based data processing method further includes:
[0151] S160a. Obtain sample data from the data in the target table object;
[0152] S160b. Convert the sample data through the target data conversion template to obtain the converted data;
[0153] If the conversion process of the sample data and the converted data meet the preset conditions, execute S160, that is, the step of converting the data in the target table object into structured data through the target data conversion template.
[0154] Among them, the preset conditions include: no dictionary data is missing during the conversion process of the sample data, the data processing component of the target data conversion template is executable, there is no missing data field mapping in the converted data, and the data type is correct.
[0155] Specifically, to improve the accuracy of the target data conversion template, after determining the target data conversion template, sample data can be read from the data in the target table object to evaluate the target data conversion template. Among them, it can be evaluated whether there is a lack of dictionary data during the conversion process of the sample data, whether the data processing components of the target data conversion template are executable, whether there is a lack of data field mapping in the converted data, and whether the data type is correct.
[0156] It can be seen that by evaluating the target data conversion template, the accuracy and executability of the target data conversion template can be ensured to improve the accuracy of the converted structured data.
[0157] Based on the embodiment shown in Figure 5 as shown, the data processing method based on the template may further include the following steps: Figure 6 as shown in
[0158] S170, obtain the preset data quality configuration rules.
[0159] S180, perform data verification on each row of the structured data based on the data quality configuration rules to obtain the score corresponding to each row of data.
[0160] S190, calculate the scores corresponding to multiple rows of the structured data according to the preset data quality evaluation formula to obtain the score corresponding to the structured data.
[0161] Among them, the score corresponding to the structured data is used to evaluate the accuracy of the structured data.
[0162] Specifically, after the data conversion is completed and before storing the structured data in the database, the data quality can be verified according to the preset data quality configuration rules, that is, the accuracy of the data can be verified. The specific steps are as follows:
[0163] 1. Load the data quality configuration rules, which can score null values, non-dictionary data, data type errors, etc.
[0164] 2. Verify the data row by row and obtain the final score of each row of data.
[0165] 3. Calculate the scores corresponding to multiple rows of the structured data according to the preset data quality evaluation formula to obtain the final score of the entire structured data file.
[0166] Through the above steps, the accuracy of the structured data can be evaluated, and thus the accuracy of the converted structured data can be improved.
[0167] In addition, in the embodiments of the present invention, data can also be desensitized, that is, sensitive data in the data is hidden. Specifically, when editing and previewing a target file, if no data parsing template is hit, all field data is desensitized. If a specific parsing template and data conversion template are hit, the sensitive field values can be desensitized according to the configuration of the template. Among them, desensitization can select semi-desensitization and full-desensitization modes. Semi-desensitization can replace the middle information of the field with "*"; full-desensitization can replace the sensitive field with "*".
[0168] In summary, through the predefined data parsing template and data conversion template in the embodiments of the present invention, the system can automatically identify and parse unstructured data in different formats, reducing the time and labor costs of manual processing. It supports the processing and import of batch data, improving the efficiency of data processing, especially in the scenario of importing large-scale data sets. The data conversion template ensures the consistency and standardization of data conversion, reducing errors caused by human factors and improving the accuracy of data import. Integrity, consistency, and correctness verification can be performed during the data conversion process to ensure the quality of the converted data. It supports data parsing templates and data conversion templates in multiple formats, and the template library can be flexibly extended and updated according to needs to adapt to data conversion of different types of data.
[0169] Moreover, a template-based data processing method provided by the embodiments of the present invention has a wide range of application prospects. For example, it can be used in the following scenarios:
[0170] Enterprise data management: Enterprises can use the template-based data processing method to automatically convert document data such as contracts, reports, and statements in various formats into structured data and import them into enterprise resource planning (ERP) systems, customer relationship management (CRM) systems, etc.
[0171] Data analysis: Research institutions and data analysis companies can use this method to process a large amount of unstructured data, extract valuable information, and conduct in-depth data analysis and mining.
[0172] Government and public services: Government departments can use this method to process various official documents, statements, etc., improving the efficiency and accuracy of data processing.
[0173] In a second aspect, an embodiment of the present invention provides a template-based data processing device 70, as Figure 7 shown, the device includes:
[0174] A file type acquisition module 710, configured to acquire the target file type of a target file, where the target file includes unstructured data;
[0175] A data parsing template determination module 720, configured to determine a target data parsing template that matches the target file type from a plurality of predefined data parsing templates based on the target file type, where different file types correspond to different data parsing templates;
[0176] A file content conversion module 730, configured to convert the file content of the target file into data in a multi-line preset format through the target data parsing template;
[0177] A target table object matching module 740, configured to, for each line of data in the preset format, determine a target table object that matches the data in the preset format based on the data type of the data in the preset format, and store the data in the preset format in the target table object;
[0178] A data conversion template determination module 750, configured to determine a target data conversion template that matches the target table object from a plurality of predefined data conversion templates, where the matching degree between the column labels in the target table object and the column labels in the data conversion template is greater than a preset matching degree;
[0179] A data conversion module 760, configured to convert the data in the target table object into structured data through the target data conversion template.
[0180] The technical solution provided by the embodiment of the present invention obtains the target file type of a target file including unstructured data, selects a target data parsing template that matches the target file type from predefined data parsing templates based on the target file type, converts the file content of the target file into data in a multi-line preset format through the target data parsing template, determines a target table object through the data type of each line of data in the preset format, and stores the data in the preset format in the target table object, where the target table object is a data carrier for multi-line data in the preset format. Finally, in a plurality of predefined data conversion templates, a target data conversion template that matches the target table object is determined, and the data in the target table object is converted into structured data through the target data conversion template.
[0181] It can be seen that through predefined data parsing templates and data conversion templates, for unstructured data in different formats, appropriate data parsing templates and data conversion templates can be selected to parse and convert the unstructured data, that is, convert the unstructured data into structured data, thereby improving the efficiency and accuracy of data processing.
[0182] Optionally, the target table object matching module is specifically configured to:
[0183] For each line of data in the preset format, determine the data type of the data in the preset format, where the data type includes row data, column data, and extended data;
[0184] If the data type of the data in the preset format is row data, the page number corresponding to the data in the preset format is the same as the page number of the table object, and the data in the preset format matches the column data or row data of the predefined table object, it is determined that the data in the preset format matches the table object, and the table object is determined as the target table object.
[0185] Optionally, the target table object matching module is specifically configured to:
[0186] If the number of columns of the data in the preset format is the same as the number of columns of the column data in the predefined table object, or the row numbers of the column data in the predefined table object corresponding to the data in the preset format are consecutive and the column number error is less than a first threshold, it is determined that the data in the preset format matches the column data of the table object;
[0187] Or,
[0188] If the number of columns of the data in the preset format is the same as the number of columns of the row data in the predefined table object, or the row numbers of the row data in the predefined table object corresponding to the data in the preset format are the same and the column number error is less than a second threshold, it is determined that the data in the preset format matches the row data of the table object.
[0189] Optionally, the data conversion module is specifically configured to:
[0190] Obtain dictionary information required during the data processing from the dictionary configuration node configured by the target data conversion template;
[0191] Obtain the table object processing logic from the table object conversion configuration node configured by the target data conversion template, and convert the data in the target table object based on the dictionary information and the table object processing logic to obtain structured data.
[0192] Optionally, the apparatus further includes:
[0193] A sample data acquisition module, configured to obtain sample data from the data in the target table object before converting the data in the target table object into structured data through the target data conversion template;
[0194] A sample data conversion module, configured to convert the sample data through the target data conversion template to obtain converted data;
[0195] If the conversion process of the sample data and the converted data meet preset conditions, trigger the data conversion module to execute the step of converting the data in the target table object into structured data through the target data conversion template;
[0196] Among them, the preset conditions include: during the conversion process of the sample data, the dictionary data is not missing, the data processing component of the target data conversion template is executable, there is no missing data field mapping and the data type is correct after conversion.
[0197] Optionally, the device further includes:
[0198] A quality configuration rule acquisition module, configured to acquire a preset data quality configuration rule;
[0199] A first scoring module, configured to perform data verification on each row of the structured data based on the data quality configuration rule to obtain a score corresponding to each row of data;
[0200] A second scoring module, configured to calculate the scores corresponding to multiple rows of the structured data according to a preset data quality evaluation formula to obtain the score corresponding to the structured data, and the score corresponding to the structured data is used to evaluate the accuracy of the structured data.
[0201] In a third aspect, an embodiment of the present invention provides an electronic device 800, as Figure 8 shown, including:
[0202] At least one processor 801;
[0203] A memory 802 for storing instructions executable by the at least one processor;
[0204] Among them, the at least one processor is configured to execute the instructions to implement the method described in the first aspect.
[0205] The technical solution provided by the embodiment of the present invention obtains the target file type of the target file including unstructured data, selects a target data parsing template matching the target file type from predefined data parsing templates based on the target file type, and converts the file content of the target file into multiple rows of data in a preset format through the target data parsing template, and determines the target table object based on the data type of each row of data in the preset format, and stores the data in the preset format in the target table object, where the target table object is the data carrier of multiple rows of data in the preset format. Finally, among the predefined multiple data conversion templates, a target data conversion template matching the target table object is determined, and the data of the target table object is converted into structured data through the target data conversion template.
[0206] It can be seen that through predefined data parsing templates and data conversion templates, for unstructured data in different formats, appropriate data parsing templates and data conversion templates can be selected to parse and convert the unstructured data, that is, convert the unstructured data into structured data, thereby improving the efficiency and accuracy of data processing.
[0207] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method described in the first aspect.
[0208] The technical solution provided by the embodiment of the present invention is to obtain the target file type of a target file including unstructured data, select a target data parsing template matching the target file type from pre-defined data parsing templates based on the target file type, convert the file content of the target file into data in a multi-line preset format through the target data parsing template, determine a target table object based on the data type of each line of the preset format data, and store the preset format data in the target table object. The target table object is a data carrier for the multi-line preset format data. Finally, among a plurality of pre-defined data conversion templates, determine a target data conversion template matching the target table object, and convert the data in the target table object into structured data through the target data conversion template.
[0209] It can be seen that through pre-defined data parsing templates and data conversion templates, for unstructured data in different formats, appropriate data parsing templates and data conversion templates can be selected to parse and convert the unstructured data, that is, convert the unstructured data into structured data, thereby improving the efficiency and accuracy of data processing.
[0210] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.< / importfiletable>
Claims
1. A template-based data processing method, characterized in that: The method comprises: Acquire a target file type of a target file, wherein the target file includes unstructured data; Based on the target file type, determining a target data parsing template that matches the target file type from a plurality of predefined data parsing templates, wherein different file types correspond to different data parsing templates; Converting the file content of the target file into multiple lines of data in a preset format through the target data parsing template; For each row of data in a preset format, based on the data type of the data in the preset format, determining a target table object that matches the data in the preset format, and storing the data in the preset format in the target table object; Determining a target data conversion template that matches the target table object from among a plurality of predefined data conversion templates, wherein a matching degree between a column label in the target table object and a column label in the data conversion template is greater than a preset matching degree; The data in the target table object is converted into structured data through the target data conversion template.
2. The method according to claim 1, characterized in that The step of determining, based on the target file type, a target data parsing template that matches the target file type from a plurality of predefined data parsing templates comprises: For each predefined data parsing template, a template feature of the data parsing template is obtained, where the template feature is used to characterize the file type of the file that can be parsed by the data parsing template; Matching the target file type with the template feature of the data parsing template to obtain a matching result, wherein the matching result includes that the target file type matches the template feature of the data parsing template, and that the target file type does not match the template feature of the data parsing template; If the matching result is that the target file type matches the template feature of the data parsing template, the data parsing template is determined as a target data parsing template that matches the target file type.
3. The method according to claim 1, characterized in that The step of determining, for each row of data in a preset format, a target table object matching the data in the preset format based on the data type of the data in the preset format includes: For each row of data in a preset format, determining a data type of the data in the preset format, wherein the data type includes row data, column data, and extended data; If the data type of the data in the preset format is row data, the page number corresponding to the data in the preset format is the same as the page number of the table object, and the data in the preset format matches the column data or row data of the predefined table object, it is determined that the data in the preset format matches the table object, and the table object is determined as the target table object.
4. The method according to claim 3, characterized in that Determining that the data in the preset format matches the column data or row data of the table object includes: If the data in the preset format has the same number of columns as the column data in the predefined table object, or the row numbers of the data in the preset format and the column data in the predefined table object are continuous and the column number error is less than a first threshold, it is determined that the data in the preset format matches the column data in the table object; or, If the data in the preset format has the same number of columns as the row data in the predefined table object, or the data in the preset format has the same row number as the row data in the predefined table object and the column number error is less than a second threshold, it is determined that the data in the preset format matches the row data in the table object.
5. The method according to any one of claims 1 to 4, characterized in that: The step of converting the data in the target table object into structured data by using the target data conversion template includes: Obtaining dictionary information required in the data processing process from the dictionary configuration node configured in the target data conversion template; The table object processing logic is obtained from the table object conversion configuration node configured in the target data conversion template, and the data in the target table object is converted based on the dictionary information and the table object processing logic to obtain structured data.
6. The method according to any one of claims 1 to 4, characterized in that: Before converting the data in the target table object into structured data by using the target data conversion template, the method further includes: Obtain sample data from the data in the target table object; Convert the sample data using the target data conversion template to obtain converted data; If the conversion process of the sample data and the converted data meet the preset conditions, executing the step of converting the data in the target table object into structured data by using the target data conversion template; Among them, the preset conditions include: the dictionary data is not missing during the sample data conversion process, the data processing component of the target data conversion template is executable, there is no missing data field mapping in the converted data, and the data type is correct.
7. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Get the preset data quality configuration rules; Performing data verification on each row of the structured data based on the data quality configuration rule to obtain a score corresponding to each row of data; According to a preset data quality assessment formula, scores corresponding to multiple rows of data of the structured data are calculated to obtain scores corresponding to the structured data, and the scores corresponding to the structured data are used to evaluate the accuracy of the structured data.
8. A data processing device based on a template, characterized in that: The device comprises: A file type acquisition module, used to acquire a target file type of a target file, wherein the target file includes unstructured data; A data parsing template determining module, configured to determine, based on the target file type, a target data parsing template matching the target file type from among a plurality of predefined data parsing templates, wherein different file types correspond to different data parsing templates; A file content conversion module, used to convert the file content of the target file into multiple lines of data in a preset format through the target data parsing template; A target table object matching module is used to determine, for each row of data in a preset format, a target table object that matches the data in the preset format based on the data type of the data in the preset format, and store the data in the preset format in the target table object; A data conversion template determination module, used to determine a target data conversion template that matches the target table object from among a plurality of predefined data conversion templates, wherein a matching degree between a column label in the target table object and a column label in the data conversion template is greater than a preset matching degree; The data conversion module is used to convert the data in the target table object into structured data through the target data conversion template.
9. An electronic device, characterized in that: include: at least one processor; a memory for storing the at least one processor-executable instruction; The at least one processor is configured to execute the instructions to implement the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
Format conversion method and device for heterogeneous data and medium
CN121092506A