A method, apparatus, device and medium for table processing
By performing word segmentation processing and data type analysis on the initial table, combined with target relationship table comparison, the table header and column data are automatically identified, which solves the problem of low efficiency in table data processing of irregular table headers, and achieves efficient automatic processing.
Patent Information
- Application Number
- CN202310500574.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-05-06
AI Technical Summary
In the prior art, the table data processing of platform bills requires manual parsing into a system preset format, which leads to time-consuming and labor-intensive and cannot effectively deal with the problem of irregular table headers.
By obtaining the initial table, determining the target relationship table from the relationship table based on its category, performing header word segmentation processing and column data type analysis, using the comparison module to match the header data and column data types after word segmentation processing with the target relationship table, and automatically identifying the target table header and column data.
Automatic processing of irregular table headers is realized, table data processing efficiency is improved, time-consuming and labor-intensive manual analysis is avoided, and processing efficiency is improved.
Smart Images

Figure CN116702727B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly relates to a table processing method, apparatus, device and medium. Background Art
[0002] The general form of a platform bill is a data table, and usually the data in the table needs to be manually edited and sorted before it can be applied to actual business requirements. The traditional process is that the operator exports the corresponding bill from the platform's background and then manually processes the bill. With the development of business, manual processing is too time-consuming and laborious, so it is necessary to automatically strip the required content from the bill through the system.
[0003] Most of the existing excel (a spreadsheet software) processing logics export the positions of preset table headers and data columns through templates, and the system accurately identifies and extracts the subsequent required data content through the preset templates. This method can achieve automatic extraction of table data, but there is a drawback, that is, the data columns in the imported excel have been limited to which data must be placed, and the source data exported from the platform will not be able to fully match the preset template due to various factors such as bill type and bill time. It is still necessary to manually parse the source bill into the bill format preset by the system, which is time-consuming and laborious and reduces the table processing efficiency. Summary of the Invention
[0004] This application provides a table processing method, apparatus, device and medium to overcome the defects of the above-mentioned existing technologies, which can automatically process table data for irregular table headers and improve the table processing efficiency.
[0005] To solve the above technical problems, this application provides the following technical solutions:
[0006] According to the first aspect of the embodiments of this application, a table processing method is provided, including:
[0007] Obtain an initial table;
[0008] Based on the category of the initial table, determine a target relationship table from the relationship tables; the relationship tables are determined based on sample tables;
[0009] Extract the table header data in the initial table to obtain initial table header data;
[0010] Perform word segmentation processing on the initial table header data to obtain the word-segmented table header data;
[0011] Extract the column data corresponding to the initial table header data to obtain initial column data;
[0012] Analyze the data type of the initial column data to obtain the initial column data type;
[0013] Compare the tokenized header data and the initial column data types with the target relationship table to obtain a comparison result;
[0014] When the comparison result indicates a match, use the match result with the highest degree of matching in the sample header of the target relationship table as the target header, and use the initial column data as the target column data corresponding to the target header.
[0015] Further, the step of comparing the tokenized header data and the initial column data types with the target relationship table to obtain a comparison result; when the comparison result indicates a match, using the match result with the highest degree of matching in the sample header of the target relationship table as the target header, and using the initial column data as the target column data corresponding to the target header includes:
[0016] Compare the tokenized header data with the sample header in the target relationship table to obtain a first result;
[0017] When the first result indicates a match, compare the initial column data type with the data type of the sample column data to obtain a second result;
[0018] When the second result indicates a match, use the match result with the highest degree of matching in the sample header as the target header, and use the initial column data as the target column data.
[0019] Further, the method further includes:
[0020] When the first result indicates a mismatch, repeat the steps of determining the target relationship table from the relationship table based on the category of the initial table to analyzing the data type of the initial column data to obtain the initial column data type until the first result indicates a match.
[0021] Further, the method further includes:
[0022] When the first result indicates a match but the second result indicates a mismatch, repeat the steps of determining the target relationship table from the relationship table based on the category of the initial table to analyzing the data type of the initial column data to obtain the initial column data type until the second result indicates a match.
[0023] Further, the method further includes:
[0024] In the case where the first result characterization matches, sort the matching results in the first result according to the matching degree, retain the matching result with the highest matching degree in the sample header as the target header, and perform dematching processing on other matching results; wherein, the header data after word segmentation processing matches the word segmentation results of at least one of the sample headers.
[0025] According to a second aspect of an embodiment of the present application, there is provided a table processing device, the device includes:
[0026] An initial table acquisition module, configured to acquire an initial table;
[0027] A target relationship table determination module, configured to determine a target relationship table from the relationship table based on the category of the initial table; the relationship table is determined based on a sample table;
[0028] An initial header acquisition module, configured to extract the header data in the initial table to obtain initial header data;
[0029] A word segmentation processing module, configured to perform word segmentation processing on the initial header data to obtain header data after word segmentation processing;
[0030] An initial column acquisition module, configured to extract the column data corresponding to the initial header data to obtain initial column data;
[0031] An initial column analysis module, configured to analyze the data type of the initial column data to obtain an initial column data type;
[0032] A comparison module, configured to compare the header data after word segmentation processing and the initial column data type with the target relationship table to obtain a comparison result;
[0033] A target data acquisition module, configured to, in the case where the comparison result characterization matches, use the matching result with the highest matching degree in the sample header of the target relationship table as the target header, and use the initial column data as the target column data corresponding to the target header.
[0034] According to a third aspect of an embodiment of the present application, there is provided an electronic device, including a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement any one of the above-mentioned table processing methods.
[0035] According to a fourth aspect of an embodiment of the present application, there is provided a computer-readable storage medium, and at least one instruction or at least one program segment is stored in the storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement any one of the above-mentioned table processing methods.
[0036] With the above technical solution, the present application has the following beneficial effects:
[0037] A table processing method, device, equipment and medium provided by the present application perform word segmentation on the initial header data in the initial table, analyze the data types of the initial column data, and then compare the word-segmented header data and the initial column data types obtained with the target relationship table. When the comparison result indicates a match, the matching result with the highest degree of matching is found in the sample header of the target relationship table as the target header, and the initial column data is used as the target column data; this technical solution can automatically process table data for irregular headers, improve the table processing efficiency, and avoid the need for manual parsing of source bills into the bill formats preset by the system, which is time-consuming and laborious. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0039] Figure 1 It is a schematic flowchart of a table processing method provided by an embodiment of the present application;
[0040] Figure 2 It is a schematic diagram of the actual application of a table processing method provided by an embodiment of the present application;
[0041] Figure 3 It is a structural block diagram of a table processing device provided by an embodiment of the present application;
[0042] Figure 4 It is a hardware structural block diagram of an electronic device for running a table processing method provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present application belong to the scope of protection of the present application.
[0044] As used herein, "one embodiment" or "embodiment" means that it may be included in
[0045] Specific features, structures, or characteristics in at least one implementation of the present application. In the description of the embodiments of the present application, it should be understood that the orientation or positional relationships indicated by terms such as "upper", "lower", "top", "bottom", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present application. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. Moreover, the terms "first", "second", etc. are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein.
[0046] Please refer to Figure 1 , which shows a schematic flowchart of a table processing method provided by an embodiment of the present application. The table processing method includes:
[0047] Step S101: Obtain an initial table;
[0048] Step S102: Determine a target relationship table from the relationship table based on the category of the initial table; the relationship table is determined based on a sample table;
[0049] Step S103: Extract the header data in the initial table to obtain initial header data;
[0050] Step S104: Perform word segmentation processing on the initial header data to obtain the header data after word segmentation processing;
[0051] Step S105: Extract the column data corresponding to the initial header data to obtain initial column data;
[0052] Step S106: Analyze the data type of the initial column data to obtain the initial column data type;
[0053] Step S107: Compare the header data after word segmentation processing and the initial column data type with the target relationship table to obtain a comparison result;
[0054] Step S108: In the case where the comparison result indicates a match, use the matching result with the highest degree of match in the sample header of the target relationship table as the target header, and use the initial column data as the target column data corresponding to the target header.
[0055] In a specific embodiment, through step S101, an initial table is obtained; wherein, the initial table can be the table where the bill source data is located; through step S102, a target relationship table matching the type of the initial table is determined from the relationship tables; the specific steps may include: comparing the type information carried in the initial table with the type information carried in each relationship table, and determining the relationship table that matches the type information carried in the initial table as the target relationship table; wherein, the relationship table is determined based on a sample table, and the sample table is determined based on the sorted historical bills; the specific steps may include: inputting the sample table into a first neural network for training, repeatedly debugging the parameters of the first neural network according to the training results until the training results meet the first preset condition, wherein the parameters include departments, businesses, etc., the first preset condition is set according to the corresponding first neural network parameters, and the first neural network parameters meet the preset standard, and the preset standard can be obtained through industry standards, manual settings, etc.; through step S103, the header data in the initial table is extracted and de-duplicated to obtain the initial header data; through step S104, word segmentation processing is performed on the initial header data; the specific steps may include: determining a word segmentation library based on the sample table, matching the initial header data with the word segmentation library data, and determining the header data after word segmentation processing for the matching ones; optionally, the natural language processing library jieba of python (a computer programming language) is used to perform word segmentation on the header. Since there are many bill nouns in the bill and they do not exist in the Chinese word segmentation dictionary of the jieba library, some bill nouns need to be customarily added to the jieba dictionary. Importing the custom dictionary can make the word segmentation more accurate. Custom special character filtering does not participate in word segmentation storage. Such as:,,., 。; "“:、}]);Through step S105, the initial column data is extracted; through step S106, the data type of the initial column data is analyzed to obtain the initial column data type. Specifically, it may include: first obtaining the data type of the initial column data according to the initial column data, then determining the column data type library based on the sample table, and matching the data type of the initial column data with the column data type library. The matching one is the initial column data type. Optionally, if the data type is integer or floating point, these contents are marked with symbols (+-), and if the data type is character type, these contents are marked with the data length range, so that the initial column data and the corresponding initial table header data form a corresponding relationship table. Through step S107, the tokenized table header data, the initial column data type, and the target relationship table are compared to obtain a comparison result. Since the Chinese word segmentation will be in the form of a combination of multiple phrases, multiple matchable data will be found when doing the table header matching. Through step S108, when the comparison result indicates a match, the matching result with the highest matching degree in the sample table header of the target relationship table is used as the target table header, and the initial column data is used as the target column data corresponding to the target table header. Among them, the specific steps for obtaining the matching weight may include: inputting the sample table into the second neural network for training, repeatedly debugging the parameters of the second neural network according to the training results until the training results meet the second preset condition; or, performing fault-tolerant training and self-learning on the initial table header data type and the matching data type to obtain the matching weight. For example, the word segmentation result of the table header "customer return quantity" is "customer return" and "quantity". Two word segmentation results can be matched from the relationship table: "payable quantity: customer return quantity * ratio" and "customer return quantity: payable quantity * ratio". To improve the accuracy of table header recognition, first, keyword weights are set for the word segmentation in the relationship table. The keyword in "payable quantity: customer return quantity * ratio" is "payable quantity", and the keyword in "customer return quantity: payable quantity * ratio" is "customer return". Then, the corresponding matching weight is calculated through the algorithm of (matched phrase / total number of phrases) * keyword weight, and the matching weights are sorted, and the table header with the highest weight is used as the alternative column. For example, the word segmentation result in "payable quantity: customer return quantity * ratio" is ['payable quantity', 'customer return', 'quantity', 'ratio'], and the matching weight of the "customer return quantity" word segmentation is (2 / 4) * 1 = 50%; the word segmentation result in "customer return quantity: payable quantity * ratio" is ['customer return', 'quantity', 'payable quantity', 'ratio'], and the matching weight of the "customer return quantity" word segmentation is (2 / 4) * 200 = 100%. At this time, the match of "payable quantity: customer return quantity * ratio" will be removed, and "customer return quantity: payable quantity * ratio" is stored as the alternative column. Obtain the corresponding data content of the table header, judge the type, length, and range of the data content, and then compare it with the corresponding type annotation in the corresponding table. The comparison result that is successful is written into the corresponding column of the new sheet, and the comparison result that fails is deleted from the alternative column.
[0056] In an optional embodiment, the above steps S107 - S108 may include:
[0057] Compare the header data after word segmentation processing with the sample headers in the target relationship table to obtain a first result;
[0058] When the first result indicates a match, compare the initial column data type with the data types of the sample column data to obtain a second result;
[0059] When the second result indicates a match, use the matching result with the highest matching degree in the sample headers as the target header, and use the initial column data as the target column data.
[0060] In an optional embodiment, the above method may further include:
[0061] When the first result indicates a mismatch, repeat the steps of determining the target relationship table from the relationship table based on the category of the initial table to analyzing the data type of the initial column data to obtain the initial column data type until the first result indicates a match.
[0062] In an optional embodiment, the above method may further include:
[0063] When the first result indicates a match but the second result indicates a mismatch, repeat the steps of determining the target relationship table from the relationship table based on the category of the initial table to analyzing the data type of the initial column data to obtain the initial column data type until the second result indicates a match.
[0064] In an optional embodiment, the above method may further include:
[0065] When the first result indicates a match, sort the matching results in the first result by matching degree, retain the matching result with the highest matching degree in the sample headers as the target header, and perform dematching processing on other matching results; wherein, the header data after word segmentation processing matches the word segmentation results of at least one of the sample headers.
[0066] The execution code for dematching processing is as follows:
[0067]
[0068]
[0069]
[0070] Please refer toFigure 2 , which shows a schematic diagram of the actual application of a table processing method provided by an embodiment of the present application; in a specific application scenario, the method may include:
[0071] Step S201: Import the bill source data;
[0072] Step S202: Sequentially determine the bill category, perform word segmentation on the table header, extract the data content of the corresponding columns of the table header, and perform data type analysis;
[0073] Step S203: Search for the corresponding table header by word segmentation of the table header in the corresponding bill type in the relationship table;
[0074] Step S204: Determine whether there is a corresponding table header; if so, execute Step S205, if not, jump back to Step S202;
[0075] Step S205: Compare the data type in the relationship table with the data type in the source data;
[0076] Step S206: Determine whether there is a corresponding data type; if so, execute Step S207, if not, jump back to Step S202;
[0077] Step S207: Extract the data of this column in the source data, and form a new data table according to the column in the relationship table.
[0078] As can be seen from the above technical solution of the embodiment of the present application, in the embodiment of the present application, by performing word segmentation processing on the initial table header data in the initial table, performing data type analysis on the initial column data, and then comparing the word-segmented table header data and the initial column data type obtained with the target relationship table, and when the comparison result shows a match, finding the matching result with the highest matching degree in the sample table header of the target relationship table as the target table header, and using the initial column data as the target column data; this technical solution can realize automatic processing of table data for irregular table headers, improve the table processing efficiency, and avoid the need for manual parsing of source bills into the bill formats preset by the system, which is time-consuming and laborious.
[0079] Corresponding to the table processing method provided in the above embodiment, the embodiment of the present application also provides a table processing device. Since the table processing device provided in the embodiment of the present application corresponds to the table processing method provided in the above embodiment, the implementation manners of the foregoing table processing method are also applicable to the table processing device provided in this embodiment, and will not be described in detail in this embodiment.
[0080] Please refer to Figure 3 , which shows a structural block diagram of a table processing device provided by an embodiment of the present application; the device includes:
[0081] 001: Initial table acquisition module, used to acquire an initial table;
[0082] 002: Target relationship table determination module, used to determine a target relationship table from a relationship table based on the category of the initial table; the relationship table is determined based on a sample table;
[0083] 003: Initial table header acquisition module, used to extract the table header data in the initial table to obtain initial table header data;
[0084] 004: Word segmentation processing module, used to perform word segmentation processing on the initial table header data to obtain the word-segmented table header data;
[0085] 005: Initial column acquisition module, used to extract the column data corresponding to the initial table header data to obtain initial column data;
[0086] 006: Initial column analysis module, used to analyze the data type of the initial column data to obtain the initial column data type;
[0087] 007: Comparison module, used to compare the word-segmented table header data and the initial column data type with the target relationship table to obtain a comparison result;
[0088] 008: Target data acquisition module, used to, when the comparison result indicates a match, use the matching result with the highest degree of match in the sample table header of the target relationship table as the target table header, and use the initial column data as the target column data corresponding to the target table header.
[0089] In a specific embodiment, an initial table is obtained through an initial table acquisition module, where the initial table can be the table where the bill source data is located; a target relationship table is determined from the relationship tables through a target relationship table determination module, and the specific steps may include: comparing the type information carried in the initial table with the type information carried in each relationship table, and determining the relationship table that matches the type information carried in the initial table as the target relationship table; where the relationship table is determined based on a sample table, and the sample table is determined based on the sorted historical bills; the specific steps may include: inputting the sample table into a first neural network for training, repeatedly debugging the parameters of the first neural network according to the training results until the training results meet the first preset condition, where the parameters include departments, businesses, etc., the first preset condition is set according to the corresponding first neural network parameters, and the first neural network parameters meet the preset standards, and the preset standards can be obtained through industry standards, manual settings, etc.; through an initial table header acquisition module, the header data in the initial table is extracted and undergoes duplicate removal processing to obtain the initial header data; through a word segmentation processing module, word segmentation processing is performed on the initial header data; the specific steps may include: determining a word segmentation library based on the sample table, matching the initial header data with the word segmentation library data, and determining the header data after word segmentation processing for the matching ones; optionally, the natural language processing library jieba of python (a computer programming language) is used to perform word segmentation on the table header. Since there are many bill nouns in the bill and they do not exist in the Chinese word segmentation dictionary in the jieba library, some bill nouns need to be customarily added to the jieba dictionary. Importing the custom dictionary can make the word segmentation more accurate. Custom special character filtering does not participate in word segmentation storage. For example: ,,., 。; "“:、}]);Through the initial column acquisition module, extract the initial column data; through the initial column analysis module, analyze the data type of the initial column data to obtain the initial column data type. Specifically, it may include: first obtain the data type of the initial column data based on the initial column data, then determine the column data type library based on the sample table, and match the data type of the initial column data with the column data type library. The matching one is the initial column data type. Optionally, if the data type is integer or floating point, mark these contents with symbols (+-), and if the data type is character type, mark these contents with the data length range, so that the initial column data and the corresponding initial header data form a correspondence table; through the comparison module, compare the segmented header data, the initial column data type with the target relationship table to obtain a comparison result; because after Chinese word segmentation, it will be in the form of multiple phrases combined, so when doing header matching, multiple matchable data will be found; through the target data acquisition module, when the comparison result indicates a match, use the matching result with the highest matching degree in the sample header of the target relationship table as the target header, and use the initial column data as the target column data corresponding to the target header; among them, the specific steps to obtain the matching weight may include: input the sample table into the second neural network for training, and repeatedly debug the parameters of the second neural network according to the training results until the training results meet the second preset condition; or, perform fault tolerance training and self-learning on the initial header data type and the matching data type to obtain the matching weight; for example, the word segmentation result of the header customer return quantity is customer return, quantity. Two sets of word segmentation results can be matched from the relationship table: payable quantity: customer return quantity * ratio, customer return quantity: payable quantity * ratio. To improve the accuracy of header recognition, first set the keyword weights for the word segmentation in the relationship table. The keyword in payable quantity: customer return quantity * ratio is payable quantity, and the keyword in customer return quantity: payable quantity * ratio is customer return. Then calculate the corresponding matching weight through the algorithm of (matched phrase / total number of phrases) * keyword weight, sort the matching weights, and use the header with the highest weight as the alternative column. For example, the word segmentation result in payable quantity: customer return quantity * ratio is ['payable quantity', 'customer return', 'quantity', 'ratio'], and the matching weight of the customer return quantity word segmentation is (2 / 4) * 1 = 50%; the word segmentation result in customer return quantity: payable quantity * ratio is ['customer return', 'quantity', 'payable quantity', 'ratio'], and the matching weight of the customer return quantity word segmentation is (2 / 4) * 200 = 100%. At this time, the match of payable quantity: customer return quantity * ratio will be de-matched, and customer return quantity: payable quantity * ratio will be stored as the alternative column. Obtain the corresponding data content of the header, judge the type, length, and range of the data content, and then compare it with the corresponding type annotation in the corresponding table. The ones with successful comparison are written into the corresponding column of the new sheet, and the ones with failed comparison are deleted from the alternative column.
[0090] In an optional embodiment, the above device may further include:
[0091] A first result acquisition module, configured to compare the header data after word segmentation processing with the sample headers in the target relationship table to obtain a first result;
[0092] A second result acquisition module, configured to, when the first result indicates a match, compare the data type of the initial column data with the data type of the sample column data to obtain a second result;
[0093] A target result acquisition module, configured to, when the second result indicates a match, use the matching result with the highest matching degree in the sample headers as the target header and use the initial column data as the target column data.
[0094] In an optional embodiment, the above device may further include:
[0095] A first repetition module, configured to, when the first result indicates a mismatch, repeat the steps of determining the target relationship table from the relationship table based on the category of the initial table to analyzing the data type of the initial column data to obtain the initial column data type until the first result indicates a match.
[0096] In an optional embodiment, the above device may further include:
[0097] A second repetition module, configured to, when the first result indicates a match but the second result indicates a mismatch, repeat the steps of determining the target relationship table from the relationship table based on the category of the initial table to analyzing the data type of the initial column data to obtain the initial column data type until the second result indicates a match.
[0098] In an optional embodiment, the above device may further include:
[0099] A matching degree screening module, configured to, when the first result indicates a match, sort the matching results in the first result by matching degree, retain the matching result with the highest matching degree in the sample headers as the target header, and perform dematching processing on other matching results; wherein, the header data after word segmentation processing matches the word segmentation results of at least one of the sample headers.
[0100] The execution code for the dematching processing is as follows:
[0101]
[0102]
[0103]
[0104] It should be noted that when the device provided in the above embodiments realizes its functions, only the division of the above functional modules is used as an example for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.
[0105] The table processing device according to the embodiment of the present application performs word segmentation processing on the initial header data in the initial table, analyzes the data types of the initial column data, and then compares the header data after word segmentation processing and the initial column data types obtained with the target relationship table. When the comparison result indicates a match, the matching result with the highest matching degree is found in the sample header of the target relationship table as the target header, and the initial column data is used as the target column data; this technical solution can automatically process table data for irregular headers, improve the table processing efficiency, and avoid the need for manual parsing of source bills into the bill formats preset by the system, which is time-consuming and laborious.
[0106] The embodiment of the present application also provides an electronic device, including a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or at least one program segment is loaded and executed by the processor to implement the table processing method provided in the above method embodiments.
[0107] The memory can be used to store software programs and modules. The processor runs the software programs and modules stored in the memory to perform various functional applications and achieve high-level autonomous driving. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory can also include a memory controller to provide the processor with access to the memory.
[0108] The method embodiments provided by the embodiments of the present application can be executed on a computer terminal, a server, or a similar computing device, that is, the above electronic device can include a computer terminal, a server, or a similar computing device. Figure 4 is a hardware structure block diagram of an electronic device that runs a table processing method provided by an embodiment of the present application, as Figure 4As shown, the internal structure of the electronic device may include, but is not limited to: a processor, a network interface, and a memory. Among them, the processor, network interface, and memory in the electronic device may be connected by a bus or other means. In the embodiments of this specification, Figure 4 the connection through the bus is taken as an example.
[0109] Among them, the processor (or CPU (Central Processing Unit)) is the computing core and control core of the electronic device. The network interface may optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.). The memory is the memory device in the electronic device, used to store programs and data. It can be understood that the memory here can be a high-speed RAM storage device, or a non-volatile storage device, such as at least one disk storage device; optionally, it can also be at least one storage device located far from the aforementioned processor. The memory provides a storage space, and the operating system of the electronic device is stored in this storage space, which may include, but is not limited to: Windows system (an operating system), Linux (an operating system), Android (a mobile operating system) system, IOS (a mobile operating system) system, etc. This application does not make any limitations in this regard; and, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space, and these instructions can be one or more computer programs (including program codes). In the embodiments of this specification, the processor loads and executes one or more instructions stored in the memory to implement the table processing method provided in the above method embodiments.
[0110] The embodiments of this application also provide a computer-readable storage medium, in which at least one instruction or at least one segment of program is stored, and the at least one instruction or at least one segment of program is loaded and executed by the processor to implement the table processing method provided in the method embodiments.
[0111] Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.
[0112] It should be noted that the above sequence of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. Moreover, the above specific embodiments of this specification have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multi-small sample image classification and parallel processing are also possible or may be advantageous.
[0113] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0114] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.
[0115] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A table processing method, characterized in that, Including: Obtain an initial table; Determine a target relationship table from a relationship table based on the category of the initial table; The relationship table is determined based on a sample table; Extract the header data in the initial table to obtain initial header data; Perform word segmentation processing on the initial header data to obtain the header data after word segmentation processing; Extract the column data corresponding to the initial header data to obtain initial column data; Analyze the data type of the initial column data to obtain the initial column data type and perform corresponding annotation; Compare the header data after word segmentation processing and the initial column data type with the target relationship table to obtain a comparison result; specifically, compare the header data after word segmentation processing with the sample headers in the target relationship table, and compare the annotation of the initial column data type with the corresponding type annotation in the target relationship table; When the comparison result indicates a match, use the matching result with the highest matching degree in the sample headers of the target relationship table as the target header, and use the initial column data as the target column data corresponding to the target header.
2. The table processing method according to claim 1, characterized in that The step of comparing the header data after word segmentation processing and the initial column data type with the target relationship table to obtain a comparison result; when the comparison result indicates a match, using the matching result with the highest matching degree in the sample headers of the target relationship table as the target header, and using the initial column data as the target column data corresponding to the target header, includes: Compare the header data after word segmentation processing with the sample headers in the target relationship table to obtain a first result; When the first result indicates a match, compare the initial column data type with the data type of the sample column data to obtain a second result; When the second result indicates a match, use the matching result with the highest matching degree in the sample headers as the target header, and use the initial column data as the target column data.
3. The table processing method according to claim 2, wherein The method further includes: When the first result indicates a mismatch, repeat the steps of determining the target relationship table from the relationship table based on the category of the initial table to analyzing the data type of the initial column data to obtain the initial column data type until the first result indicates a match.
4. The table processing method according to claim 3, characterized in that The method further includes: When the first result indicates a match but the second result indicates a mismatch, repeat the steps of determining the target relationship table from the relationship table based on the category of the initial table to analyzing the data type of the initial column data to obtain the initial column data type until the second result indicates a match.
5. The table processing method according to any one of claims 2 to 4, characterized in that The method further includes: When the first result indicates a match, sort the matching results in the first result by matching degree, retain the matching result with the highest matching degree in the sample headers as the target header, and perform dematching processing on other matching results; wherein, the header data after word segmentation processing matches the word segmentation results of at least one of the sample headers.
6. A table processing device, characterized in that, The device includes: An initial table acquisition module for acquiring an initial table; A target relationship table determination module, configured to determine a target relationship table from a relationship table based on the category of the initial table; the relationship table is determined based on a sample table; An initial table header acquisition module, configured to extract the table header data in the initial table to obtain initial table header data; A word segmentation processing module, configured to perform word segmentation processing on the initial table header data to obtain the table header data after word segmentation processing; An initial column acquisition module, configured to extract the column data corresponding to the initial table header data to obtain initial column data; An initial column analysis module, configured to analyze the data type of the initial column data to obtain the initial column data type and perform corresponding annotation; A comparison module, configured to compare the table header data after word segmentation processing and the initial column data type with the target relationship table to obtain a comparison result; specifically, compare the table header data after word segmentation processing with the sample table headers in the target relationship table, and compare the annotation of the initial column data type with the corresponding type annotation in the target relationship table; A target data acquisition module, configured to, when the comparison result indicates a match, use the matching result with the highest degree of match in the sample table headers of the target relationship table as the target table header, and use the initial column data as the target column data corresponding to the target table header.
7. An electronic device, characterized in that, It includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the table processing method according to any one of claims 1 to 5.
8. A computer-readable storage medium, in which at least one instruction or at least one program segment is stored, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the table processing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Data processing method and device, storage medium and equipment
CN114969051A
Drug data acquisition and storage method and system and storage medium
CN115185947A