A data table classification method, device, equipment and storage medium

By using automated keyword matching and multi-dimensional chi-square value analysis, the problem of inaccurate manual data table partitioning was solved, and accurate partitioning of data tables into data warehouse layers was achieved.

CN115599975BActive Publication Date: 2025-10-21WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211208626.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-30
Publication Date
2025-10-21
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

In existing technologies, the partitioning of data tables into data warehouse layers relies on human experience, leading to inaccurate results.

Method used

By identifying target keywords that match the information in each dimension of the data table from multiple preset keywords, and based on the correspondence between classification labels and multi-dimensional chi-square values, the classification labels and classification results of the data table are automatically determined, reducing manual intervention.

Benefits of technology

Improves the accuracy of data table classification results, ensuring that data tables are accurately divided into corresponding data warehouse layers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115599975B_ABST
    Figure CN115599975B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of data table classification method, device, equipment and storage medium, it is related to big data processing technical field, the method comprises: from multiple preset keywords determine at least one target keyword matched with the dimension table information of data table, for any target keyword, based on classification label corresponding relationship, determine target classification label corresponding to target keyword, and the multidimensional chi-square value associated with target classification label and target keyword.Finally, based on the target classification label corresponding to each target keyword and the multidimensional chi-square value associated with target classification label and each target keyword, determine the classification result of data table, instead of relying on artificial experience, improve the accuracy of data table classification result, in turn guarantee the accuracy of data table division to corresponding data warehouse hierarchy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of big data processing technology, and in particular to a data table classification method, device, equipment and storage medium. Background Art

[0002] With the rapid development of the Internet, data has also exploded. To facilitate data storage, companies currently store data in the form of data tables in data warehouses.

[0003] Data warehouses are generally divided into five data warehouse layers: the Operation Data Store (ODS), the Dimension (DIM), the Data Warehouse Details (DWD), the Data Warehouse Service (DWS), and the Application Data Service (ADS). The Operation Data Store is an offline or near-real-time data access layer; the Dimension (DIM) stores data collected from multiple dimensions; the Data Details layer cleans and transforms the Data Operations layer; the Data Warehouse Details layer aggregates the Data Details layer; and the Application Data Service (ADS) aggregates and integrates the Data Warehouse Details layer to provide services such as subsequent business queries.

[0004] To more clearly manage data tables in a data warehouse, they need to be divided into different data warehouse tiers. Currently, data tables are typically analyzed manually and divided into corresponding data warehouse tiers based on the analysis results. This reliance on manual experience can easily lead to inaccurate data table division results. Summary of the Invention

[0005] The embodiments of the present application provide a data table classification method, apparatus, device, and storage medium for improving the accuracy of data table classification results.

[0006] In one aspect, an embodiment of the present application provides a data table classification method, the method comprising:

[0007] Determine at least one target keyword that matches information of each dimension table of the data table from a plurality of preset keywords;

[0008] For any target keyword, based on the classification label correspondence, determine the target classification label corresponding to the target keyword and the multi-dimensional chi-square value associated with the target classification label; the classification label correspondence is the classification relationship of each preset keyword determined based on multiple sample data tables; wherein the classification relationship of each preset keyword includes the classification label to which the preset keyword belongs and the multi-dimensional chi-square value associated with the preset keyword and the classification label to which it belongs; the multi-dimensional chi-square value is used to represent the correlation between the preset keyword and the classification label to which it belongs under multiple dimensional table information;

[0009] The classification result of the data table is determined based on the target classification label corresponding to each target keyword and the multi-dimensional chi-square value associated with each target keyword and the target classification label.

[0010] Optionally, the classification label correspondence is a classification relationship of each preset keyword determined according to multiple sample data tables, including:

[0011] For any preset keyword, based on the multiple sample data tables, respectively determine the multi-dimensional chi-square value of the preset keyword and each candidate classification label;

[0012] The largest multidimensional chi-square value is selected from multiple multidimensional chi-square values, and when the largest multidimensional chi-square value is greater than the critical value of the chi-square distribution, the candidate classification label corresponding to the largest multidimensional chi-square value is used as the classification label to which the preset keyword belongs, and the largest multidimensional chi-square value is used as the multidimensional chi-square value associated with the preset keyword and the classification label to which it belongs.

[0013] Optionally, for any preset keyword, determining the multi-dimensional chi-square value of the preset keyword and each candidate classification label based on the multiple sample data tables includes:

[0014] For any candidate classification label corresponding to any preset keyword, perform the following steps:

[0015] Based on the multiple sample data tables, respectively determine the confidence value corresponding to each dimension table information; the confidence value is used to characterize the relevance of each dimension information and the candidate classification label;

[0016] Based on the confidence value corresponding to each dimension table information and the single-dimensional chi-square value associated with the preset keyword and the candidate classification label in each dimension table information, the multi-dimensional chi-square value of the preset keyword and the candidate classification label is determined.

[0017] Optionally, respectively determining the confidence value corresponding to each dimension table information includes:

[0018] For any dimension table information, based on the multiple sample data tables, determining an association probability value between the preset keyword in the dimension table information and the candidate classification label;

[0019] Determining a weight factor for each dimension table information based on an association probability value between the preset keyword and the candidate classification label in each dimension table information;

[0020] For any dimension table information, the weight factor of the dimension table information is used to adjust the association probability value between the preset keyword and the candidate classification label in the dimension table information to obtain the confidence value of the dimension table information.

[0021] Optionally, determining the association probability values ​​between the preset keywords and the candidate classification labels in the dimension table information based on the multiple sample data tables includes:

[0022] Determining a first data table quantity containing the preset keyword in dimension table information of the plurality of sample data tables;

[0023] Determining that dimension table information of the plurality of sample data tables contains the preset keyword, and the plurality of sample data tables belong to a second data table quantity of the candidate classification label;

[0024] The ratio of the second data table amount to the first data table amount is used as the association probability value.

[0025] Optionally, determining the weight factor of each dimension table information based on the association probability value between the preset keyword and the candidate classification label in each dimension table information includes:

[0026] Determine the sum of the association probability values ​​of the preset keywords and the candidate classification labels in each dimension table information as the total association probability value;

[0027] For any dimension table information, the weight factor of the dimension information is determined based on the association probability value between the preset keyword and the candidate classification label in the dimension table information, and the total association probability value.

[0028] Optionally, determining the multi-dimensional chi-square value of the preset keyword and the candidate classification label based on the confidence value corresponding to each dimension table information and the single-dimensional chi-square value associated with the preset keyword and the candidate classification label in each dimension table information includes:

[0029] Sort by the confidence value corresponding to each dimension table information to obtain the sorted confidence value;

[0030] According to the preset matching relationship, the first confidence value and the second confidence value having a matching relationship are obtained from the sorted confidence values ​​in sequence;

[0031] For each first confidence value and second confidence value that have a matching relationship, determine a chi-square difference based on a single-dimensional chi-square value associated with the preset keyword and the candidate classification label in the dimension table information corresponding to the first confidence value, and a single-dimensional chi-square value associated with the preset keyword and the candidate classification label in the dimension table information corresponding to the second confidence value;

[0032] Based on the chi-square difference values ​​corresponding to the first confidence values ​​and the second confidence values ​​that have a matching relationship, a multi-dimensional chi-square value of the preset keyword and the candidate classification label is determined.

[0033] Optionally, determining the classification result of the data table based on the target classification label corresponding to each target keyword and the multi-dimensional chi-square value associated with each target keyword and the target classification label includes:

[0034] Determine a classification label group according to the target classification label corresponding to each target keyword; the classification label group corresponds to the target classification label one-to-one; the classification label group includes at least one target keyword;

[0035] The classification result of the data table is determined based on the number of tags corresponding to each classification tag group and the chi-square value of at least one target keyword in each classification tag group associated with the target classification tag.

[0036] Optionally, determining the classification result of the data table based on the number of tags corresponding to each classification tag group and the chi-square value associated with at least one target keyword in each classification tag group and the target classification tag includes:

[0037] If there are at least two classification label groups, and the number of labels in the at least two classification label groups is the largest and equal, the at least two classification label groups are used as reference label groups;

[0038] For any reference tag group, determining a reference chi-square value based on the chi-square value of at least one target keyword in the reference tag group being associated with the target classification tag;

[0039] The target classification label corresponding to the maximum reference chi-square value among the reference chi-square values ​​corresponding to each reference label group is taken as the classification result.

[0040] Optionally, the candidate classification labels include data operation class, public dimension class, data detail class, data intermediate class, and data application class.

[0041] Optionally, after determining the classification result of the data table, the method further includes:

[0042] Based on the dependency relationships among the data operation class, the common dimension class, the data detail class, the data intermediate class, and the data application class, the classification result of the data table is verified.

[0043] In one aspect, an embodiment of the present application provides a data table classification device, the device comprising:

[0044] A keyword determination module, configured to determine at least one target keyword that matches information in each dimension table of the data table from a plurality of preset keywords;

[0045] A classification label determination module is configured to determine, for any target keyword, the target classification label corresponding to the target keyword and the multidimensional chi-square value associated with the target classification label based on the classification label correspondence relationship; the classification label correspondence relationship is the classification relationship of each preset keyword determined based on multiple sample data tables; wherein the classification relationship of each preset keyword includes the classification label to which the preset keyword belongs and the multidimensional chi-square value associated with the preset keyword and the classification label to which it belongs; the multidimensional chi-square value is used to represent the correlation between the preset keyword and the classification label to which it belongs under multiple dimensional table information;

[0046] The classification result determination module is used to determine the classification result of the data table based on the target classification label corresponding to each target keyword and the multi-dimensional chi-square value associated with each target keyword and the target classification label.

[0047] Optionally, the classification label determination module is specifically configured to:

[0048] For any preset keyword, based on the multiple sample data tables, respectively determine the multi-dimensional chi-square value of the preset keyword and each candidate classification label;

[0049] The largest multidimensional chi-square value is selected from multiple multidimensional chi-square values, and when the largest multidimensional chi-square value is greater than the critical value of the chi-square distribution, the candidate classification label corresponding to the largest multidimensional chi-square value is used as the classification label to which the preset keyword belongs, and the largest multidimensional chi-square value is used as the multidimensional chi-square value associated with the preset keyword and the classification label to which it belongs.

[0050] Optionally, the classification label determination module is specifically configured to:

[0051] For any candidate classification label corresponding to any preset keyword, perform the following steps:

[0052] Based on the multiple sample data tables, respectively determine the confidence value corresponding to each dimension table information; the confidence value is used to characterize the relevance of each dimension information and the candidate classification label;

[0053] Based on the confidence value corresponding to each dimension table information and the single-dimensional chi-square value associated with the preset keyword and the candidate classification label in each dimension table information, the multi-dimensional chi-square value of the preset keyword and the candidate classification label is determined.

[0054] Optionally, the classification label determination module is specifically configured to:

[0055] For any dimension table information, based on the multiple sample data tables, determining an association probability value between the preset keyword in the dimension table information and the candidate classification label;

[0056] Determining a weight factor for each dimension table information based on an association probability value between the preset keyword and the candidate classification label in each dimension table information;

[0057] For any dimension table information, the weight factor of the dimension table information is used to adjust the association probability value between the preset keyword and the candidate classification label in the dimension table information to obtain the confidence value of the dimension table information.

[0058] Optionally, the classification label determination module is specifically configured to:

[0059] Determining a first data table quantity containing the preset keyword in dimension table information of the plurality of sample data tables;

[0060] Determining that dimension table information of the plurality of sample data tables contains the preset keyword, and the plurality of sample data tables belong to a second data table quantity of the candidate classification label;

[0061] The ratio of the second data table amount to the first data table amount is used as the association probability value.

[0062] Optionally, the classification label determination module is specifically configured to:

[0063] Determine the sum of the association probability values ​​of the preset keywords and the candidate classification labels in each dimension table information as the total association probability value;

[0064] For any dimension table information, the weight factor of the dimension information is determined based on the association probability value between the preset keyword and the candidate classification label in the dimension table information, and the total association probability value.

[0065] Optionally, the classification label determination module is specifically configured to:

[0066] Sort by the confidence value corresponding to each dimension table information to obtain the sorted confidence value;

[0067] According to the preset matching relationship, the first confidence value and the second confidence value having a matching relationship are obtained from the sorted confidence values ​​in sequence;

[0068] For each first confidence value and second confidence value that have a matching relationship, determine a chi-square difference based on a single-dimensional chi-square value associated with the preset keyword and the candidate classification label in the dimension table information corresponding to the first confidence value, and a single-dimensional chi-square value associated with the preset keyword and the candidate classification label in the dimension table information corresponding to the second confidence value;

[0069] Based on the chi-square difference values ​​corresponding to the first confidence values ​​and the second confidence values ​​that have a matching relationship, a multi-dimensional chi-square value of the preset keyword and the candidate classification label is determined.

[0070] Optionally, the classification result determination module is specifically configured to:

[0071] Determine a classification label group according to the target classification label corresponding to each target keyword; the classification label group corresponds to the target classification label one-to-one; the classification label group includes at least one target keyword;

[0072] The classification result of the data table is determined based on the number of tags corresponding to each classification tag group and the chi-square value of at least one target keyword in each classification tag group associated with the target classification tag.

[0073] Optionally, the classification result determination module is specifically configured to:

[0074] If there are at least two classification label groups, and the number of labels in the at least two classification label groups is the largest and equal, the at least two classification label groups are used as reference label groups;

[0075] For any reference tag group, determining a reference chi-square value based on the chi-square value of at least one target keyword in the reference tag group being associated with the target classification tag;

[0076] The target classification label corresponding to the maximum reference chi-square value among the reference chi-square values ​​corresponding to each reference label group is taken as the classification result.

[0077] Optionally, the candidate classification labels include data operation class, public dimension class, data detail class, data intermediate class, and data application class.

[0078] Optionally, a verification module is also included, specifically for:

[0079] After determining the classification result of the data table, the classification result of the data table is verified based on the dependency relationships among the data operation class, the common dimension class, the data detail class, the data intermediate class, and the data application class.

[0080] On the one hand, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned transaction parameter query method when executing the program.

[0081] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program that can be executed by a computer device. When the program is run on the computer device, the computer device executes the steps of the above-mentioned transaction parameter query method.

[0082] In an embodiment of the present application, at least one target keyword that matches the information in each dimension table of a data table is determined from multiple preset keywords. For each target keyword, the target classification label corresponding to the target keyword and the multi-dimensional chi-square value associated with the target keyword and the target classification label are determined based on the classification label correspondence. Finally, based on the target classification label corresponding to each target keyword and the multi-dimensional chi-square value associated with each target keyword and the target classification label, the classification result of the data table is determined, rather than relying on manual experience. This improves the accuracy of the data table classification result and further ensures the accuracy of the data table division into the corresponding data warehouse layers. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0084] Figure 1 A schematic diagram of a system architecture provided in an embodiment of the present application;

[0085] Figure 2 A flowchart of a data table classification method provided in an embodiment of the present application;

[0086] Figure 3 A schematic diagram of the blood relationship structure of a classification label provided in an embodiment of the present application;

[0087] Figure 4 A flowchart of a method for determining classification relationships provided in an embodiment of the present application;

[0088] Figure 5A flowchart of a method for determining a multi-dimensional chi-square value provided in an embodiment of the present application;

[0089] Figure 6 A flowchart of a method for determining a confidence value provided in an embodiment of the present application;

[0090] Figure 7 A flowchart of a method for determining a multi-dimensional chi-square value provided in an embodiment of the present application;

[0091] Figure 8 A schematic diagram of the structure of a data table classification device provided in an embodiment of the present application;

[0092] Figure 9 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0093] In order to make the purpose, technical solutions and beneficial effects of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0094] refer to Figure 1 , which is a data table classification system architecture diagram applicable to an embodiment of the present application, and the data table classification system architecture diagram at least includes a terminal device 101 and a data table classification system 102.

[0095] The terminal device 101 is installed with a target application for data table classification, which can be a pre-installed client, a web application, or a small program embedded in other applications, etc. The terminal device 101 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto.

[0096] Data table classification system 102 is the backend server of the target application, providing services for the target application. Data table classification system 102 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0097] The terminal device 101 and the data table classification system 102 can be directly or indirectly connected through wired or wireless communication, and this application does not impose any restrictions on this.

[0098] The terminal device 101 responds to the user's data table classification operation and sends a data table classification instruction to the data table classification system 102. The data table classification system 102 receives the data table classification instruction and determines at least one target keyword that matches the information of each dimension table of the data table from multiple preset keywords. For any target keyword, based on the classification label correspondence, the target classification label corresponding to the target keyword and the multi-dimensional chi-square value associated with the target keyword and the target classification label are determined; the classification label correspondence is the classification relationship of each preset keyword determined based on multiple sample data tables; wherein, the classification relationship of each preset keyword includes the classification label to which the preset keyword belongs and the multi-dimensional chi-square value associated with the preset keyword and the classification label to which it belongs; the multi-dimensional chi-square value is used to characterize the correlation between the preset keyword and the classification label to which it belongs under multiple dimension table information; based on the target classification label corresponding to each target keyword and the multi-dimensional chi-square value associated with each target keyword and the target classification label, the classification result of the data table is determined.

[0099] based on Figure 1 The system architecture diagram, the embodiment of the present application provides a process of a data table classification method, such as Figure 2 As shown, the process of this method is Figure 1 The data table classification system 102 shown is executed, including the following steps:

[0100] Step S201: Determine at least one target keyword that matches information of each dimension table of the data table from a plurality of preset keywords.

[0101] Specifically, a data table corresponds to different dimension table information at different dimensions. The dimension table information corresponding to the library dimension in a data table is the library name and library description. The dimension table information corresponding to the table dimension in a data table is the table name and table description. The dimension table information corresponding to the field dimension in a data table is the field name and field description.

[0102] Preset keywords generally include: ods, record, info, logs, etc. Preset keywords can be adjusted according to the business of different companies and are not limited here.

[0103] For any data table to be classified, determine the dimension table information of the data table, including library name, library description, table name, table description, field name and field description.

[0104] Each preset keyword is judged. If any dimension table information of the data table contains the preset keyword, the preset keyword is used as the target keyword. For example, if the preset keyword is ods and the database name of the data table contains the preset feature word "ods", the preset feature word ods is used as the target keyword.

[0105] Among them, in determining whether the dimension table information of the data table contains the preset keywords, methods such as word segmentation, filtering, and statistics can be used, which are not limited here.

[0106] Step S202 : for any target keyword, based on the category label correspondence, determine the target category label corresponding to the target keyword and the multi-dimensional chi-square value associated with the target keyword and the target category label.

[0107] Specifically, the classification label correspondence is a classification relationship of each preset keyword determined according to multiple sample data tables.

[0108] By performing statistics on multiple sample data tables, the classification label to which each preset keyword belongs, as well as the multi-dimensional chi-square value between each preset keyword and the classification label to which it belongs, are determined. The multi-dimensional chi-square value is used to represent the correlation between the preset keyword and the classification label under the information of multiple dimension tables. The multi-dimensional chi-square value is determined by the single-dimensional chi-square value of the preset keyword under the information of each dimension table.

[0109] The classification label described by each preset keyword and the multi-dimensional chi-square value associated with each preset keyword and the classification label to which it belongs are used as the classification relationship of each preset keyword, and the classification relationship of each preset keyword is used as the classification label correspondence relationship.

[0110] Among them, the classification labels include data operation category, public dimension category, data details category, data middle layer category, and data application category.

[0111] For example, the classification label correspondence is shown in Table 1. The preset keywords included in the classification label correspondence are ods, record, info, logs, ads, and application. Taking the record preset keyword as an example, the classification label to which the record preset keyword belongs is the public dimension class, and the multi-dimensional chi-square value associated with the record preset keyword and the public dimension class is 230.

[0112] For other preset keywords in Table 1, the correspondence between other preset keywords and classification labels, and the correspondence between them and multi-dimensional chi-square values ​​will not be further explained.

[0113] Table 1.

[0114] Preset keywords Category Tags Multidimensional chi-square value record Public dimension class 230 ods Data Operations 320 info Data details class 445 logs Data intermediate class 210 ads Data Application 108 application Data Application 200

[0115] Step S203 : determining the classification result of the data table based on the target classification label corresponding to each target keyword and the multi-dimensional chi-square value associated with each target keyword and the target classification label.

[0116] In one embodiment, a classification label group is determined based on the target classification label corresponding to each target keyword; wherein the classification label group corresponds to the target classification label one-to-one, and each classification label group includes at least one target keyword. A classification result of the data table is determined based on the number of labels corresponding to each classification label group and the chi-square value associated with the target classification label of at least one target keyword in each classification label group.

[0117] Specifically, there are two possibilities for the number of labels in each classification label group:

[0118] The first possibility: If there is only one classification label group with the largest number of labels, the target classification label corresponding to the classification label group is directly used as the classification result.

[0119] The second possibility: if there are at least two classification label groups, and the number of labels of the at least two classification label groups is the largest and equal, the at least two classification label groups are used as reference label groups.

[0120] For any reference tag group, a reference chi-square value is determined based on the chi-square values ​​of at least one target keyword associated with the target classification tag within the reference tag group. The reference chi-square value may be the sum of the chi-square values ​​associated with the at least one target keyword and the target classification tag, or the maximum chi-square value selected from the chi-square values ​​associated with the at least one target keyword and the target classification tag. Other methods may also be used to determine the reference chi-square value, which are not limited herein.

[0121] Finally, the target classification label corresponding to the maximum reference chi-square value among the reference chi-square values ​​corresponding to each reference label group is taken as the classification result.

[0122] In this application, after determining the classification result of the data table, the data table can be directly divided into the data warehouse layer corresponding to the classification result based on the classification result of the data table. For example, if the classification result of the data table is data operation type, the data table will be divided into the data operation layer.

[0123] After determining the classification results of the data table, the classification results of the data table can also be verified based on the dependency relationships among the data operation class, public dimension class, data detail class, data intermediate class, and data application class.

[0124] Specifically, the dependency relationship of each classification label is generally the blood relationship of each classification label. In the blood relationship, the classification label can generally be divided into outflow nodes, intermediate nodes, and inflow nodes. The blood relationship of data operation class, public dimension class, data detail class, data intermediate class, and data application class is as follows: Figure 3As shown in the figure, data operation and public dimension classes are outflow nodes, data detail and data intermediate classes are intermediate nodes, and data application classes are inflow nodes.

[0125] If a data table's classification label corresponds to an outflow node, the data table is the data provider and does not reference other data tables in terms of lineage. If a data table's classification label corresponds to an inflow node, the data table is not referenced by other data tables in terms of lineage. If a data table's classification label corresponds to an intermediate node, the data table references both other data tables with classification labels for outflow nodes and other data tables with classification labels for inflow nodes in terms of lineage.

[0126] Among them, the dependency relationship of each classification label is as follows:

[0127] The first dependency: When the classification label of a data table is data operation or public dimension, if the data table does not reference other data tables in the blood relationship, the classification result of the data table is accurate; otherwise, the classification result of the data table is inaccurate.

[0128] The second dependency relationship: When the classification label of a data table is data application class, if the data table is referenced by other data tables in a blood relationship, the classification result of the data table is inaccurate; otherwise, the classification result of the data table is accurate.

[0129] The third type of dependency: When the classification label of a data table is data detail class, if the data table references other data tables with classification labels of data operation class or public dimension class in lineage relationship, and the data table is referenced by other data tables with classification labels of data intermediate class in lineage relationship, then the classification result of the data table is accurate; otherwise, the classification result of the data table is inaccurate.

[0130] The fourth type of dependency: When the classification label of a data table is data intermediate class, if the data table references other data tables with the classification label of data detail class in lineage, and the data table is referenced by other data tables with the classification label of data application class in lineage, then the classification result of the data table is accurate; otherwise, the classification result of the data table is inaccurate.

[0131] After verifying the classification results using the dependency relationship of each classification label, if the classification results are inaccurate, the classification results can be modified, further improving the accuracy of the classification results.

[0132] In an embodiment of the present application, at least one target keyword that matches the information in each dimension table of a data table is determined from multiple preset keywords. For each target keyword, the target classification label corresponding to the target keyword and the multi-dimensional chi-square value associated with the target keyword and the target classification label are determined based on the classification label correspondence. Finally, based on the target classification label corresponding to each target keyword and the multi-dimensional chi-square value associated with each target keyword and the target classification label, the classification result of the data table is determined, rather than relying on manual experience. This improves the accuracy of the data table classification result and further ensures the accuracy of dividing the data table into the corresponding data warehouse layers.

[0133] Optionally, in the above step S202, the classification label correspondence is the classification relationship of each preset keyword determined according to multiple sample data tables, specifically including: Figure 4 The following steps are shown:

[0134] Step S401 : for any preset keyword, based on multiple sample data tables, respectively determine the multi-dimensional chi-square value of the preset keyword and each candidate classification label.

[0135] Specifically, the candidate classification labels include data operation class, public dimension class, data detail class, data intermediate class, and data application class.

[0136] For any preset keyword and any candidate classification label, before calculating the multi-dimensional chi-square value of the preset keyword and the candidate classification label, the following assumptions must be defined:

[0137] Null hypothesis H0: Assume that the preset keyword is unrelated to the candidate classification label.

[0138] Alternative hypothesis H1: Assume that the preset keyword is related to the candidate classification label.

[0139] For example, if the preset keyword is record, you need to define the hypothetical questions for the preset keyword record and the data operation class, public dimension class, data detail class, data intermediate class, and data application class. For example, the hypothetical questions for the preset keyword record and the data operation class are as follows:

[0140] Null hypothesis H0: Assume that the preset keyword record has nothing to do with data operation.

[0141] Alternative hypothesis H1: Assume that the preset keyword record is related to data operation.

[0142] The assumption problem between the preset keyword record and other candidate classification labels is similar to the assumption problem between the preset keyword record and data operation categories, and is not limited here.

[0143] Assume that through calculation, the multidimensional chi-square values ​​of the preset keyword record and each candidate classification label are shown in Table 2. The multidimensional chi-square values ​​of the preset keyword record and the data operation category, public dimension category, data detail category, data intermediate category, and data application category are 180, 230, 10, 30, and 50, respectively.

[0144] Table 2.

[0145] Preset keywords Category Tags Multidimensional chi-square value record Data Operations 180 record Public dimension class 230 record Data details class 10 record Data intermediate class 30 record Data Application 50

[0146] Step S402: Select the largest multidimensional chi-square value from multiple multidimensional chi-square values, and when the largest multidimensional chi-square value is greater than the critical value of the chi-square distribution, use the candidate classification label corresponding to the largest multidimensional chi-square value as the classification label to which the preset keyword belongs, and use the largest multidimensional chi-square value as the multidimensional chi-square value associated with the preset keyword and the classification label to which it belongs.

[0147] Specifically, the commonly used chi-square distribution table is shown in Table 3.

[0148] Table 3.

[0149]

[0150] The critical value of the chi-square distribution is generally set to 3.841. Generally speaking, when the multidimensional chi-square value is greater than the critical value of the chi-square distribution, it means that the probability that the null hypothesis H0 of the preset keyword and candidate classification label corresponding to the multidimensional chi-square value is true is 0.05. Since the probability of 0.05 is relatively small, it can be inferred that the null hypothesis H0 of the preset keyword and candidate classification label corresponding to the multidimensional chi-square value is not true, and the alternative hypothesis H1 of the preset keyword and candidate classification label corresponding to the multidimensional chi-square value is true.

[0151] Depending on different situations, the critical value of the chi-square distribution can also be set to 5.024, 6.635, etc., which is not limited here.

[0152] According to Table 2, the multidimensional chi-square value of 230 between the preset keyword record and the public dimension class is the largest. Since 230 is much larger than 3.841, the public dimension class is taken as the classification label to which the preset keyword belongs, and the maximum multidimensional chi-square value 230 is taken as the multidimensional chi-square value associated with the preset keyword record and the public dimension class.

[0153] In an embodiment of the present application, for any preset keyword, based on multiple sample data tables, the multidimensional chi-square value of the preset keyword and each candidate classification label is determined respectively, and the largest multidimensional chi-square value is selected from the multiple multidimensional chi-square values. When the largest multidimensional chi-square value is greater than the critical value of the chi-square distribution, the candidate classification label corresponding to the largest multidimensional chi-square value is used as the classification label to which the preset keyword belongs, and the largest multidimensional chi-square value is used as the multidimensional chi-square value associated with the preset keyword and the classification label to which it belongs. The classification label correspondence determined by the above method ensures the accuracy of the classification relationship of each preset keyword.

[0154] Optionally, in the above step S401, for any preset keyword, based on multiple sample data tables, the multi-dimensional chi-square value of the preset keyword and each candidate classification label is determined respectively. Specifically, for any candidate classification label corresponding to any preset keyword, the following is specifically performed: Figure 5 The following steps are shown:

[0155] Step S501: Based on multiple sample data tables, determine the confidence value corresponding to each dimension table information.

[0156] Among them, the confidence value is used to characterize the correlation between each dimension information and the candidate classification label.

[0157] The confidence value corresponding to each dimension table information is related to the preset keywords, candidate classification labels, and dimension table information.

[0158] For the same sample data table, when the preset keywords and candidate classification labels are the same, the confidence values ​​corresponding to different dimension table information are different; when the preset keywords are the same and the candidate classification labels are different, the confidence values ​​corresponding to the same dimension table information are different; when the candidate classification labels are the same and the preset keywords are different, the confidence values ​​corresponding to the same dimension table information are different.

[0159] Step S502 : determining a multi-dimensional chi-square value of the preset keyword and the candidate classification label based on the confidence value corresponding to each dimension table information and the single-dimensional chi-square value associated with the preset keyword and the candidate classification label in each dimension table information.

[0160] In an embodiment of the present application, the multidimensional chi-square value of the preset keyword and the candidate classification label is determined based on the confidence value corresponding to each dimension table information and the single-dimensional chi-square value associated with the preset keyword and the candidate classification label in each dimension table information, thereby ensuring that the multidimensional chi-square value fully considers the correlation between each dimension table information and each dimension table information, thereby increasing the accuracy of the multidimensional chi-square value.

[0161] In the above step S501, based on multiple sample data tables, the confidence value of each dimension table information corresponding to the preset keyword is determined, specifically including the following: Figure 6 The following steps are shown:

[0162] Step S601 : for any dimension table information, based on multiple sample data tables, determining the association probability value between the preset keywords in the dimension table information and the candidate classification labels.

[0163] Specifically, for any dimension table information, the method first determines the number of first data tables whose dimension table information contains the preset keyword, and then determines the number of second data tables whose dimension table information contains the preset keyword and whose sample data tables belong to candidate classification labels. Finally, the ratio of the second data table volume to the first data table volume is used as the association probability value.

[0164] Assuming the number of sample data tables is 765, when the dimension table information is the library name, the preset keyword is record, and the candidate classification label is the public dimension class, according to the content in Table 4: among the 765 sample data tables, 78 sample data tables have the preset keyword record in their library name, that is, the first data table quantity is 78; among the 765 sample data tables, 76 sample data tables have the preset keyword record in their library name and belong to the public dimension class, that is, the second data table quantity is 76. Therefore, the association probability value between the preset keyword record and the public dimension class is 76 / 78.

[0165] Table 4.

[0166]

[0167] When the dimension table information is library description, the preset keyword is record, and the candidate classification label is the public dimension class, Table 5 shows that: among the 765 sample data tables, 89 sample data tables contain the preset keyword record in the library description, that is, the first data table quantity is 89; among the 765 sample data tables, 66 sample data tables contain the preset keyword record in the library description and belong to the public dimension class, that is, the second data table quantity is 66. Therefore, the association probability value between the preset keyword record and the public dimension class is 66 / 89.

[0168] Table 5.

[0169]

[0170] When the dimension table information is table name, table description, field name, and field description, the association probability value between the preset keyword record and the public dimension class is no longer counted, and the statistical process is similar to the above process.

[0171] In an embodiment of the present application, the association probability is determined based on the preset keywords contained in the dimension table information of the sample data table and the relationship between the sample data table and the candidate classification label, which facilitates the subsequent calculation of the confidence value of each dimension table information.

[0172] Step S602 : determining a weight factor of each dimension table information based on an association probability value between a preset keyword and a candidate classification label in each dimension table information.

[0173] Specifically, the sum of the association probability values ​​of the preset keywords and the candidate classification labels in each dimension table information is first determined as the total association probability value; then, for any dimension table information, the weight factor of the dimension information is determined based on the association probability values ​​of the preset keywords and the candidate classification labels in the dimension table information, and the total association probability value.

[0174] Set the association probability values ​​of the preset keywords and candidate classification labels in each dimension table information to be P1, P2, ..., P n , where n is the number of dimension table information. The total value of the association probability is P1+…+P n .

[0175] Set the weight factor of the i-th dimension table information to W i , the association probability value P between the preset keywords and the candidate classification labels in the i-th dimension table information i The ratio of the total value of the association probability is used as the weight factor of the information in the i-th dimension table, as shown in formula (1):

[0176]

[0177] Step S603 : for any dimension table information, the weight factor of the dimension table information is used to adjust the association probability value between the preset keywords and the candidate classification labels in the dimension table information to obtain the confidence value of the dimension table information.

[0178] Specifically, for any dimension table information, the product of the weight factor of the dimension table information and the probability value of the association between the preset keywords and the candidate classification labels in the dimension table information is used as the adjusted probability value of the dimension table information. The ratio of the adjusted probability value of the dimension table information to the sum of the adjusted probability values ​​of all dimension table information is then used as the confidence value of the dimension table information.

[0179] Set the weight factors of n dimension table information to W1, W2, ..., W n , the association probability values ​​of the preset keywords and candidate classification labels in each dimension table information are P1, P2, ..., P n , the adjusted probability values ​​of each dimension table information are W1*P1, W2*P2, ..., Wn *P n . Confidence value CV for any dimension table information i , the corresponding calculation formula is shown in formula (2):

[0180]

[0181] In an embodiment of the present application, the confidence value of the dimension table information is determined based on the weight factor of the dimension table information and the association probability value between the preset keywords and the candidate classification labels in the dimension table information, which facilitates the subsequent calculation of the multi-dimensional chi-square value.

[0182] Optionally, in the above step S502, based on the confidence value corresponding to each dimension table information and the single-dimensional chi-square value associated with the preset keyword and the candidate classification label in each dimension table information, the multi-dimensional chi-square value of the preset keyword and the candidate classification label is determined, specifically including the following: Figure 7 The following steps are shown:

[0183] Step S701 : sorting the information in each dimension table according to the confidence value corresponding to the information to obtain the sorted confidence value.

[0184] Step S702 : According to the preset matching relationship, the first confidence value and the second confidence value having a matching relationship are sequentially obtained from the sorted confidence values.

[0185] Specifically, in order to meet the preset matching relationship, the number of dimension table information is first judged. If the number of dimension table information is odd, the dimension table information corresponding to the median of the sorted confidence value is deleted; if the number of dimension table information is even, it is not processed.

[0186] The preset matching relationship may be: taking the maximum confidence value as the first confidence value and the minimum confidence value as the second confidence value; then taking the second largest confidence value as the first confidence value and the second smallest confidence value as the second confidence value, and so on.

[0187] The preset matching relationship can also be: taking the maximum confidence value as the first confidence value and the confidence value at n / 2+1 as the second confidence value; then taking the second largest confidence value as the first confidence value and the confidence value at n / 2+2 as the second confidence value; where n is the number of dimension table information.

[0188] Step S703: For each first confidence value and second confidence value that have a matching relationship, determine the chi-square difference based on the single-dimensional chi-square value associated with the preset keyword and the candidate classification label in the dimension table information corresponding to the first confidence value, and the single-dimensional chi-square value associated with the preset keyword and the candidate classification label in the dimension table information corresponding to the second confidence value.

[0189] Specifically, a commonly used chi-square calculation method is used to determine the single-dimensional chi-square value associated with the preset keywords and the candidate classification labels in each dimension table information.

[0190] The number of sample data tables is set to 765. When the dimension table information is the library name, the preset keyword is record, and the candidate classification label is the public dimension class, the relationship between the preset keyword and the candidate classification label is shown in Table 4 above.

[0191] Based on the content in Table 4, the commonly used chi-square value determination method is used to determine whether the library name of the sample data table contains the preset keyword record, and whether the sample data table belongs to the public dimension class. The theoretical values ​​are shown in Table 6.

[0192] Table 6.

[0193]

[0194] Based on the contents in Table 4 and Table 6, the chi-square value calculation formula is used. Determine the single-dimension chi-square value associated with the preset keyword record in the library name and the public dimension class. represents the chi-square value, A represents the actual value of whether the sample data table's library name contains a record and whether the sample data table belongs to the public dimension class, and T represents the theoretical value of whether the sample data table's library name contains a record and whether the sample data table belongs to the public dimension class. For example, when A indicates that the sample data table's library name contains a record and the sample data table belongs to the public dimension class, A = 76; when A indicates that the sample data table's library name does not contain a record and the sample data table belongs to the public dimension class, A = 2. When T indicates that the sample data table's library name contains a record and the sample data table belongs to the public dimension class, T = 23.751; when T indicates that the sample data table's library name does not contain a record and the sample data table belongs to the public dimension class, T = 209.1915.

[0195] The square of the difference between the single-dimensional chi-square value associated with the preset keywords and the candidate classification labels in the dimension table information corresponding to the first confidence value and the single-dimensional chi-square value associated with the preset keywords and the candidate classification labels in the dimension table information corresponding to the second confidence value is used as the chi-square difference value.

[0196] Step S704 : determining a multi-dimensional chi-square value of the preset keyword and the candidate classification label based on the chi-square difference values ​​corresponding to the first confidence values ​​and the second confidence values ​​that have a matching relationship.

[0197] Specifically, the sum of multiple chi-square differences is averaged and then the square root is calculated to obtain the multi-dimensional chi-square value of the preset keyword and the candidate classification label. The specific formula is shown in formula (3):

[0198]

[0199] Among them, r represents the multi-dimensional chi-square value of the preset keywords and candidate classification labels, Indicates the single-dimensional chi-square value associated with the preset keywords and candidate classification labels in the n-dimensional table information, are the chi-square differences between the first and second confidence values ​​that have a matching relationship, and n is the number of dimension table information.

[0200] For example, the single-dimensional chi-square values ​​associated with the preset keyword record and the data operation category in each dimension table information are determined as shown in Table 7.

[0201] Table 7.

[0202] Preset keywords Candidate classification labels Dimension table information Single dimension chi-square value record Data Operations Library Name 183.98 record Data Operations Library Description 90.81 record Data Operations Table name 293.37 record Data Operations Table Description 181.62 record Data Operations Field Name 16.23 record Data Operations Field Description 211.62

[0203] When the preset keyword is record and the candidate classification label is data operation, the confidence values ​​of the information in each dimension table are shown in Table 8.

[0204] Table 8.

[0205]

[0206]

[0207] Sorted by confidence value, the confidence values ​​are 0.199334833, 0.189243599, 0.184665035, 0.178878552, 0.138426956, and 0.10945097. The order of dimension table information is table name, library name, table description, field description, field name, and library description. According to the preset matching relationship, the first confidence value and the second confidence value for determining the matching relationship are (confidence value of the table name, confidence value of the library description), (confidence value of the library name, confidence value of the field name), (confidence value of the table description, confidence value of the field description), i.e. (0.199334833, 0.10945097), (0.189243599, 0.138426956), (0.184665035, 0.178878552).

[0208] The single-dimensional chi-square values ​​of the first confidence value and the second confidence value with matching relationships are (293.37, 90.81), (183.98, 16.23), and (181.62, 211.62), respectively. The multi-dimensional chi-square value calculation formula is: Therefore, the multi-dimensional chi-square value of the preset keyword record and data operation category is 152.83.

[0209] By using the above process, the multi-dimensional chi-square values ​​of the preset keyword record and the common dimension class, data detail class, data intermediate class and data application class can be determined respectively.

[0210] For other preset keywords, the above process can be used to determine the multi-dimensional chi-square values ​​of other preset keywords and data operation category, public dimension category, data detail category, data intermediate category and data application category respectively.

[0211] In an embodiment of the present application, the multidimensional chi-square values ​​of the preset keywords and candidate classification labels are determined based on the confidence values ​​corresponding to each dimension table information, which fully reflects the weights of different dimension table information in the multidimensional chi-square values ​​and improves the accuracy of the multidimensional chi-square values.

[0212] Based on the same technical concept, the embodiment of the present application provides a data table classification device, such as Figure 8 As shown, the data table classification device 800 includes:

[0213] The keyword determination module 801 is configured to determine at least one target keyword that matches information in each dimension table of the data table from a plurality of preset keywords;

[0214] The classification label determination module 802 is used to determine, for any target keyword, the target classification label corresponding to the target keyword and the multi-dimensional chi-square value associated with the target classification label based on the classification label correspondence relationship; the classification label correspondence relationship is the classification relationship of each preset keyword determined based on multiple sample data tables; wherein the classification relationship of each preset keyword includes the classification label to which the preset keyword belongs and the multi-dimensional chi-square value associated with the preset keyword and the classification label to which it belongs; the multi-dimensional chi-square value is used to represent the correlation between the preset keyword and the classification label to which it belongs under multiple dimensional table information;

[0215] The classification result determination module 803 is configured to determine the classification result of the data table based on the target classification label corresponding to each target keyword and the multi-dimensional chi-square value associated with each target keyword and the target classification label.

[0216] Optionally, the classification label determination module 802 is specifically configured to:

[0217] For any preset keyword, based on the multiple sample data tables, respectively determine the multi-dimensional chi-square value of the preset keyword and each candidate classification label;

[0218] The largest multidimensional chi-square value is selected from multiple multidimensional chi-square values, and when the largest multidimensional chi-square value is greater than the critical value of the chi-square distribution, the candidate classification label corresponding to the largest multidimensional chi-square value is used as the classification label to which the preset keyword belongs, and the largest multidimensional chi-square value is used as the multidimensional chi-square value associated with the preset keyword and the classification label to which it belongs.

[0219] Optionally, the classification label determination module 802 is specifically configured to:

[0220] For any candidate classification label corresponding to any preset keyword, perform the following steps:

[0221] Based on the multiple sample data tables, respectively determine the confidence value corresponding to each dimension table information; the confidence value is used to characterize the relevance of each dimension information and the candidate classification label;

[0222] Based on the confidence value corresponding to each dimension table information and the single-dimensional chi-square value associated with the preset keyword and the candidate classification label in each dimension table information, the multi-dimensional chi-square value of the preset keyword and the candidate classification label is determined.

[0223] Optionally, the classification label determination module 802 is specifically configured to:

[0224] For any dimension table information, based on the multiple sample data tables, determining an association probability value between the preset keyword in the dimension table information and the candidate classification label;

[0225] Determining a weight factor for each dimension table information based on an association probability value between the preset keyword and the candidate classification label in each dimension table information;

[0226] For any dimension table information, the weight factor of the dimension table information is used to adjust the association probability value between the preset keyword and the candidate classification label in the dimension table information to obtain the confidence value of the dimension table information.

[0227] Optionally, the classification label determination module 802 is specifically configured to:

[0228] Determining a first data table quantity containing the preset keyword in dimension table information of the plurality of sample data tables;

[0229] Determining that dimension table information of the plurality of sample data tables contains the preset keyword, and the plurality of sample data tables belong to a second data table quantity of the candidate classification label;

[0230] The ratio of the second data table amount to the first data table amount is used as the association probability value.

[0231] Optionally, the classification label determination module 802 is specifically configured to:

[0232] Determine the sum of the association probability values ​​of the preset keywords and the candidate classification labels in each dimension table information as the total association probability value;

[0233] For any dimension table information, the weight factor of the dimension information is determined based on the association probability value between the preset keyword and the candidate classification label in the dimension table information, and the total association probability value.

[0234] Optionally, the classification label determination module 802 is specifically configured to:

[0235] Sort by the confidence value corresponding to each dimension table information to obtain the sorted confidence value;

[0236] According to the preset matching relationship, the first confidence value and the second confidence value having a matching relationship are obtained from the sorted confidence values ​​in sequence;

[0237] For each first confidence value and second confidence value that have a matching relationship, determine a chi-square difference based on a single-dimensional chi-square value associated with the preset keyword and the candidate classification label in the dimension table information corresponding to the first confidence value, and a single-dimensional chi-square value associated with the preset keyword and the candidate classification label in the dimension table information corresponding to the second confidence value;

[0238] Based on the chi-square difference values ​​corresponding to the first confidence values ​​and the second confidence values ​​that have a matching relationship, a multi-dimensional chi-square value of the preset keyword and the candidate classification label is determined.

[0239] Optionally, the classification result determination module 803 is specifically configured to:

[0240] Determine a classification label group according to the target classification label corresponding to each target keyword; the classification label group corresponds to the target classification label one-to-one; the classification label group includes at least one target keyword;

[0241] The classification result of the data table is determined based on the number of tags corresponding to each classification tag group and the chi-square value of at least one target keyword in each classification tag group associated with the target classification tag.

[0242] Optionally, the classification result determination module 803 is specifically configured to:

[0243] If there are at least two classification label groups, and the number of labels in the at least two classification label groups is the largest and equal, the at least two classification label groups are used as reference label groups;

[0244] For any reference tag group, determining a reference chi-square value based on the chi-square value of at least one target keyword in the reference tag group being associated with the target classification tag;

[0245] The target classification label corresponding to the maximum reference chi-square value among the reference chi-square values ​​corresponding to each reference label group is taken as the classification result.

[0246] Optionally, the candidate classification labels include data operation class, public dimension class, data detail class, data intermediate class, and data application class.

[0247] Optionally, a verification module 804 is further included, specifically configured to:

[0248] After determining the classification result of the data table, the classification result of the data table is verified based on the dependency relationships among the data operation class, the common dimension class, the data detail class, the data intermediate class, and the data application class.

[0249] Based on the same technical concept, the embodiment of the present application provides a computer device, which can be a terminal or a server, such as Figure 9 As shown, it includes at least one processor 901 and a memory 902 connected to the at least one processor. The specific connection medium between the processor 901 and the memory 902 is not limited in the embodiment of the present application. Figure 9 For example, the processor 901 and the memory 902 are connected via a bus. The bus can be divided into an address bus, a data bus, a control bus, and the like.

[0250] In an embodiment of the present application, the memory 902 stores instructions that can be executed by at least one processor 901. The at least one processor 901 can execute the steps included in the above-mentioned data table classification method by executing the instructions stored in the memory 902.

[0251] The processor 901 is the control center of the computer device. It can connect various parts of the computer device using various interfaces and lines, and perform data table classification by running or executing instructions stored in the memory 902 and calling data stored in the memory 902. Optionally, the processor 901 may include one or more processing units. The processor 901 may integrate an application processor and a modem processor. The application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the modem processor may not be integrated into the processor 901. In some embodiments, the processor 901 and the memory 902 may be implemented on the same chip. In some embodiments, they may also be implemented on separate chips.

[0252] The processor 901 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor.

[0253] The memory 902 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 902 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 902 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 902 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.

[0254] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program that can be executed by a computer device. When the program runs on the computer device, the computer device executes the steps of the above-mentioned data table classification method.

[0255] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0256] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0257] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0258] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0259] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A data table classification method, characterized in that: include: Determine at least one target keyword that matches information of each dimension table of the data table from a plurality of preset keywords; For any target keyword, based on the classification label correspondence, determine the target classification label corresponding to the target keyword and the multi-dimensional chi-square value associated with the target classification label; the classification label correspondence is the classification relationship of each preset keyword determined based on multiple sample data tables; wherein the classification relationship of each preset keyword includes the classification label to which the preset keyword belongs and the multi-dimensional chi-square value associated with the preset keyword and the classification label to which it belongs; the multi-dimensional chi-square value is used to represent the correlation between the preset keyword and the classification label to which it belongs under multiple dimensional table information; Determining a classification result of the data table based on a target classification label corresponding to each target keyword and a multi-dimensional chi-square value associated with each target keyword and the target classification label; The classification label correspondence is a classification relationship of each preset keyword determined based on multiple sample data tables, including: For any candidate classification label corresponding to any preset keyword, the following steps are performed: based on the multiple sample data tables, respectively determining the confidence value corresponding to each dimension table information; the confidence value is used to characterize the correlation between each dimension information and the candidate classification label; based on the confidence value corresponding to each dimension table information and the single-dimensional chi-square value associated with the preset keyword and the candidate classification label in each dimension table information, determining the multi-dimensional chi-square value between the preset keyword and the candidate classification label; The largest multidimensional chi-square value is selected from multiple multidimensional chi-square values, and when the largest multidimensional chi-square value is greater than the critical value of the chi-square distribution, the candidate classification label corresponding to the largest multidimensional chi-square value is used as the classification label to which the preset keyword belongs, and the largest multidimensional chi-square value is used as the multidimensional chi-square value associated with the preset keyword and the classification label to which it belongs.

2. The method according to claim 1, wherein Determining the confidence value corresponding to each dimension table information separately includes: For any dimension table information, based on the multiple sample data tables, determining an association probability value between the preset keyword in the dimension table information and the candidate classification label; Determining a weight factor for each dimension table information based on an association probability value between the preset keyword and the candidate classification label in each dimension table information; For any dimension table information, the weight factor of the dimension table information is used to adjust the association probability value between the preset keyword and the candidate classification label in the dimension table information to obtain the confidence value of the dimension table information.

3. The method according to claim 2, wherein The determining, based on the multiple sample data tables, the association probability values ​​between the preset keywords in the dimension table information and the candidate classification labels includes: Determining a first data table quantity containing the preset keyword in dimension table information of the plurality of sample data tables; Determining that dimension table information of the plurality of sample data tables contains the preset keyword, and the plurality of sample data tables belong to a second data table quantity of the candidate classification label; The ratio of the second data table amount to the first data table amount is used as the association probability value.

4. The method according to claim 2, wherein The determining of the weight factor of each dimension table information based on the association probability value between the preset keyword and the candidate classification label in each dimension table information includes: Determine the sum of the association probability values ​​of the preset keywords and the candidate classification labels in each dimension table information as the total association probability value; For any dimension table information, the weight factor of the dimension information is determined based on the association probability value between the preset keyword and the candidate classification label in the dimension table information, and the total association probability value.

5. The method according to claim 1, wherein The determining of the multi-dimensional chi-square value between the preset keyword and the candidate classification label based on the confidence value corresponding to each dimension table information and the single-dimensional chi-square value associated with the preset keyword and the candidate classification label in each dimension table information includes: Sort by the confidence value corresponding to each dimension table information to obtain the sorted confidence value; According to the preset matching relationship, the first confidence value and the second confidence value having a matching relationship are obtained from the sorted confidence values ​​in sequence; For each first confidence value and second confidence value that have a matching relationship, determine a chi-square difference based on a single-dimensional chi-square value associated with the preset keyword and the candidate classification label in the dimension table information corresponding to the first confidence value, and a single-dimensional chi-square value associated with the preset keyword and the candidate classification label in the dimension table information corresponding to the second confidence value; Based on the chi-square difference values ​​corresponding to the first confidence values ​​and the second confidence values ​​that have a matching relationship, a multi-dimensional chi-square value of the preset keyword and the candidate classification label is determined.

6. The method according to any one of claims 1 to 5, characterized in that The determining of the classification result of the data table based on the target classification label corresponding to each target keyword and the multi-dimensional chi-square value associated with each target keyword and the target classification label includes: Determine a classification label group according to the target classification label corresponding to each target keyword; the classification label group corresponds to the target classification label one-to-one; the classification label group includes at least one target keyword; The classification result of the data table is determined based on the number of tags corresponding to each classification tag group and the chi-square value of at least one target keyword in each classification tag group associated with the target classification tag.

7. The method according to claim 6, wherein The determining of the classification result of the data table based on the number of labels corresponding to each classification label group and the chi-square value associated with at least one target keyword in each classification label group and the target classification label includes: If there are at least two classification label groups, and the number of labels in the at least two classification label groups is the largest and equal, the at least two classification label groups are used as reference label groups; For any reference tag group, determining a reference chi-square value based on the chi-square value of at least one target keyword in the reference tag group being associated with the target classification tag; The target classification label corresponding to the maximum reference chi-square value among the reference chi-square values ​​corresponding to each reference label group is taken as the classification result.

8. The method according to any one of claims 1 to 5, characterized in that The candidate classification labels include data operation class, public dimension class, data detail class, data intermediate class, and data application class.

9. The method according to claim 8, wherein After determining the classification result of the data table, the method further includes: Based on the dependency relationships among the data operation class, the common dimension class, the data detail class, the data intermediate class, and the data application class, the classification result of the data table is verified.

10. A data table classification device, characterized in that: include: A keyword determination module, configured to determine at least one target keyword that matches information in each dimension table of the data table from a plurality of preset keywords; A classification label determination module is configured to determine, for any target keyword, the target classification label corresponding to the target keyword and the multidimensional chi-square value associated with the target classification label based on the classification label correspondence relationship; the classification label correspondence relationship is the classification relationship of each preset keyword determined based on multiple sample data tables; wherein the classification relationship of each preset keyword includes the classification label to which the preset keyword belongs and the multidimensional chi-square value associated with the preset keyword and the classification label to which it belongs; the multidimensional chi-square value is used to represent the correlation between the preset keyword and the classification label to which it belongs under multiple dimensional table information; a classification result determination module, configured to determine a classification result of the data table based on a target classification label corresponding to each target keyword and a multi-dimensional chi-square value associated with each target keyword and the target classification label; The classification label determination module is specifically configured to: for any candidate classification label corresponding to any preset keyword, perform the following steps: based on the multiple sample data tables, respectively determine the confidence value corresponding to each dimension table information; the confidence value is used to characterize the correlation between each dimension information and the candidate classification label; based on the confidence value corresponding to each dimension table information and the single-dimensional chi-square value associated with the preset keyword and the candidate classification label in each dimension table information, determine the multi-dimensional chi-square value between the preset keyword and the candidate classification label; The classification label determination module is specifically used to: select the largest multidimensional chi-square value from multiple multidimensional chi-square values, and when the largest multidimensional chi-square value is greater than the critical value of the chi-square distribution, use the candidate classification label corresponding to the largest multidimensional chi-square value as the classification label to which the preset keyword belongs, and use the largest multidimensional chi-square value as the multidimensional chi-square value associated with the preset keyword and the classification label to which it belongs.

11. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 9 are implemented.

12. A computer-readable storage medium, characterized in that It stores a computer program that can be executed by a computer device. When the program is run on the computer device, the computer device executes the steps of any one of the methods according to claims 1 to 9.

Citation Information

Patent Citations

  • Data resource obtaining method and device, storage medium and processor

    CN111125086A

  • Text multi-label classification method and device, equipment and storage medium

    CN112765965A