Data Table Field Relationship Recognition Method, Device, Electronic Device and Storage Medium
The method simplifies the identification of dimension field relationships in data tables by using field type determination and machine learning, improving accuracy and efficiency while reducing user complexity.
Patent Information
- Application Number
- CN202111223392.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-10-20
AI Technical Summary
In the prior art, determining the relationship between the dimension field and the dimension field requires a pivot table drag operation, resulting in cumbersome processing flow and reducing the data table processing efficiency and user experience.
By determining the field type in the data table, obtaining the set of cells corresponding to the dimension fields and their related information, building a cross-tab and using a pre-trained classification model to determine the relationship between dimension fields, including inclusion relationships, non-inclusion relationships and combination suggestions, simplifying the operation process.
It improves the accuracy and efficiency of field relationship recognition in data tables, provides a foundation for subsequent data table proofreading and field recommendations, and improves user experience.
Smart Images

Figure CN114003665B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method, apparatus, electronic device, and storage medium for identifying the relationship between data table fields. Background Art
[0002] In data standardization work, as new data tables are continuously connected to the database and with the rapid development of big data technology, the accuracy of data and the quality of identifying the relationship between data table fields are crucial for the value that data can produce.
[0003] A data table is composed of fields in the table and the data in the corresponding cells of each field. Among them, the fields in the data table are roughly divided into two categories, dimension fields and measure fields. A dimension field refers to a "classification field" used to describe what the attribute of the data in the cell is; a measure field is used to describe the quantity. In order to accurately analyze and judge the data in the data table, it is necessary to confirm the relationship between each dimension field.
[0004] In the prior art, the relationship between dimension fields is often determined only after constructing a pivot table. It is necessary to drag the corresponding two dimension fields to the rows and columns and then view them to determine the relationship between the two. This confirmation method is relatively complex, the processing flow is cumbersome, reducing the efficiency of data table processing and resulting in a poor user experience. Summary of the Invention
[0005] Based on the problems existing in the prior art, the present invention provides a method, apparatus, electronic device, and storage medium for identifying the relationship between data table fields, which realizes the determination of the relationship between each dimension field without a pivot table, improves the accuracy and efficiency of identifying the relationship between data table fields, provides a basis for subsequent data table proofreading and field recommendation, and has the advantages of improving the efficiency of data table processing and enhancing the user experience.
[0006] In a first aspect, the present invention provides a method for identifying the relationship between data table fields, including:
[0007] Determine the type of each field in the data table to be processed, and determine the dimension field and the set of cells corresponding to the dimension field according to the type of the field; wherein, the dimension field is a field used to describe the meaning represented by the data in the cell;
[0008] Obtain the relevant information of the set of cells corresponding to each dimension field, the relevant information of the constructed cross-table, and the information determined according to the type of the dimension field;
[0009] Determine the relationships between the respective dimensional fields based on the relevant information of the cell sets corresponding to the respective dimensional fields, the relevant information of the constructed cross-tabulation, and the information determined according to the types of the dimensional fields.
[0010] Furthermore, according to the method for identifying the relationships between the data table fields provided by the present invention, the obtaining of the relevant information of the cell sets corresponding to the respective dimensional fields, the relevant information of the constructed cross-tabulation, and the information determined according to the types of the dimensional fields includes:
[0011] Obtain the description information of the cell sets corresponding to the respective dimensional fields, the relationship information between any two cell sets, the description features and correlation features of the cross-tabulation constructed by any two cell sets, and the enumerated value types of the cell sets corresponding to the dimensional fields determined according to the types of any two dimensional fields;
[0012] Correspondingly, the determining of the relationships between the respective dimensional fields based on the relevant information of the cell sets corresponding to the respective dimensional fields, the relevant information of the constructed cross-tabulation, and the information determined according to the types of the dimensional fields includes:
[0013] Determine the relationships between the respective dimensional fields based on the description information of the cell sets corresponding to the respective dimensional fields, the relationship information between any two cell sets, the description features and correlation features of the cross-tabulation constructed by any two cell sets, and the enumerated value types of the cell sets corresponding to the dimensional fields determined according to the types of any two dimensional fields.
[0014] Furthermore, according to the method for identifying the relationships between the data table fields provided by the present invention, the obtaining of the description information of the cell sets corresponding to the respective dimensional fields, the relationship information between any two cell sets, the description features and correlation features of the cross-tabulation constructed by any two cell sets, and the enumerated value types of the corresponding cell sets determined according to the types of any two dimensional fields includes:
[0015] Obtain the description information of the first cell set, the description information of the second cell set, and the relationship information between the first cell set and the second cell set; wherein, the first cell set is the cell set corresponding to the first dimensional field, and the second cell set is the cell set corresponding to the second dimensional field; the first dimensional field is any one of the dimensional fields in the data table to be processed, and the second dimensional field is any one of the dimensional fields in the data table to be processed that is different from the first dimensional field;
[0016] Construct a cross-tabulation of the first cell set and the second cell set, and obtain the description features and correlation features of the cross-tabulation;
[0017] Determine the enumerated value types of the first cell set and the second cell set according to the types of the first dimension field and the second dimension field.
[0018] Further, according to the data table field relationship recognition method provided by the present invention, determining the relationship information between each dimension field according to the description information of the cell sets corresponding to each dimension field, the relationship information between any two cell sets, the description features and correlation features of the cross-table constructed by any two cell sets, and the enumerated value types of the cell sets corresponding to any two dimension fields includes:
[0019] Input the description information of the first cell set, the description information of the second cell set, the relationship information between the first cell set and the second cell set, the description features of the cross-table, the correlation features of the cross-table, and the enumerated value types of the first cell set and the second cell set into a pre-trained first classification model to obtain a first probability value indicating whether there is an inclusion relationship between the first dimension field and the second dimension field;
[0020] Input the description information of the first cell set, the description information of the second cell set, the relationship information between the first cell set and the second cell set, the description features of the cross-table, the correlation features of the cross-table, and the enumerated value types of the first cell set and the second cell set into a pre-trained second classification model to obtain a second probability value indicating whether combination is recommended when the first dimension field and the second dimension field are in a non-inclusion relationship;
[0021] Determine the relationship information between the first dimension field and the second dimension field according to the first probability value and the second probability value;
[0022] Wherein, the first classification model is trained based on the description information of the first sample cell set, the description information of the second sample cell set, the relationship information between the first sample cell set and the second sample cell set, the description features of the cross-table formed by the first sample cell set and the second sample cell set, the correlation features of the cross-table formed by the first sample cell set and the second sample cell set, the enumerated value types of the first sample cell set and the second sample cell set, and the label information indicating whether there is an inclusion relationship between the first sample cell set and the second sample cell set;
[0023] The second classification model is trained based on the description information of the first sample cell set, the description information of the second sample cell set, the relationship information between the first sample cell set and the second sample cell set, the description features of the contingency table formed by the first sample cell set and the second sample cell set, the correlation features of the contingency table formed by the first sample cell set and the second sample cell set, the enumerated value types of the first sample cell set and the second sample cell set, and the label information on whether combination is recommended under the non-inclusion relationship between the first sample cell set and the second sample cell set.
[0024] Further, according to the method for identifying data table field relationships provided by the present invention, determining the relationship information between the first dimension field and the second dimension field based on the first probability value and the second probability value includes:
[0025] When the first probability value is greater than the second probability value, there is an inclusion relationship between the first dimension field and the second dimension field;
[0026] When the first probability value is less than the second probability value and the second probability value is greater than or equal to a preset first threshold, there is a non-inclusion and recommended combination relationship between the first dimension field and the second dimension field;
[0027] When the first probability value is less than the second probability value and the second probability value is less than a preset first threshold, there is a non-inclusion and non-recommended combination relationship between the first dimension field and the second dimension field.
[0028] Further, according to the method for identifying data table field relationships provided by the present invention, obtaining the description information of the first cell set, the description information of the second cell set, and the relationship information between the first cell set and the second cell set includes:
[0029] Obtaining the index value of the first cell set in the data table to be processed;
[0030] Obtaining the maximum value of the data lengths in each cell of the first cell set;
[0031] Obtaining the minimum value of the data lengths in each cell of the first cell set;
[0032] Obtaining the average value of the data lengths in each cell of the first cell set;
[0033] Obtaining the standard deviation of the data lengths in each cell of the first cell set;
[0034] Obtaining the number of non-repeated data in the first cell set;
[0035] Obtain the number of cells in the first cell set;
[0036] Obtain the index value of the second cell set in the data table to be processed;
[0037] Obtain the maximum value of the lengths of the data in each cell of the second cell set;
[0038] Obtain the minimum value of the lengths of the data in each cell of the second cell set;
[0039] Obtain the average value of the lengths of the data in each cell of the second cell set;
[0040] Obtain the standard deviation of the lengths of the data in each cell of the second cell set;
[0041] Obtain the number of distinct data in the second cell set;
[0042] Obtain the product of the number of distinct data in the first cell set and the number of distinct data in the second cell set;
[0043] Obtain information on whether there is an inclusion relationship between the first cell set and the second cell set.
[0044] Further, according to the data table field relationship recognition method provided by the present invention, constructing a cross - table of the first cell set and the second cell set, and obtaining the descriptive features and correlation features of the cross - table, includes:
[0045] Construct a cross - table of the first cell set and the second cell set;
[0046] Obtain the descriptive features of the cross - table, including: calculating the number of distinct data in the cross - table, calculating the number of non - empty cells in the cross - table;
[0047] Obtain the correlation features of the cross - table, including: calculating the P - value of the chi - square test of the cross - table, calculating the degrees of freedom of the chi - square test of the cross - table, calculating whether the P - value of the chi - square test of the cross - table is less than a preset second threshold, calculating the quotient of the number of cells in the cross - table whose absolute value of the correlation coefficient is greater than or equal to a third threshold and 2; calculating the percentage of the number of cells in the cross - table whose absolute value of the correlation coefficient is greater than or equal to a third threshold in the total number of cells, calculating the average value of the data in each cell of the cross - table, calculating the standard deviation of the average value of the data in each cell of the cross - table.
[0048] Further, according to the method for identifying data table field relationships provided by the present invention, determining the types of each field in the data table to be processed, and determining dimension fields and the cell sets corresponding to the dimension fields according to the types of the fields, includes:
[0049] Determine the types of each field in the data table to be processed;
[0050] For any field in the data table to be processed, when the type of the field meets the first condition, determine the field as a dimension field; wherein, the first condition is a condition for judging whether the types of each field in the data table to be processed belong to dimension fields;
[0051] Obtain the cell set corresponding to the dimension field from the data table to be processed; wherein, the cell set corresponding to the dimension field is the row or column corresponding to the dimension field in the data table to be processed.
[0052] Further, according to the method for identifying data table field relationships provided by the present invention, determining the types of each field in the data table to be processed, includes:
[0053] Obtain the data table to be processed; wherein, the data table includes fields and cells, and the cells include data;
[0054] Determine the type of each cell in the data table according to the data included in the cell;
[0055] Determine the type of the field according to the types of each cell corresponding to the field.
[0056] In a second aspect, the present invention further provides a device for identifying data table field relationships, including:
[0057] A first determination module, configured to determine the types of each field in the data table to be processed, and determine dimension fields and the cell sets corresponding to the dimension fields according to the types of the fields; wherein, the dimension field is a field for describing the meaning represented by the data in the cell;
[0058] An acquisition module, configured to acquire relevant information of the cell sets corresponding to each dimension field, relevant information of the constructed cross table, and information determined according to the types of the dimension fields;
[0059] A determination module, configured to determine the relationships between each dimension field according to the relevant information of the cell sets corresponding to each dimension field, the relevant information of the constructed cross table, and the information determined according to the types of the dimension fields.
[0060] In a third aspect, the present invention further provides an electronic device, including: a processor, a memory, and a bus, wherein,
[0061] The processor and the memory complete communication with each other through the bus;
[0062] The memory stores program instructions executable by the processor, and the processor can execute the steps of the data table field relationship recognition method described in any one of the above by invoking the program instructions.
[0063] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions cause the computer to execute the steps of the data table field relationship recognition method described in any one of the above.
[0064] The present invention provides a method, apparatus, electronic device and storage medium for identifying data table field relationships, which determine the types of each field in a data table to be processed, determine dimension fields and the corresponding cell sets according to the field types, and then obtain the relevant information of the cell sets corresponding to each dimension field, the relevant information of the constructed cross table, and the information determined according to the types of the dimension fields, and determine the relationships between the dimension fields according to the information obtained above. The data table field relationship recognition method provided by the present invention solves the technical problems in the prior art that determining the relationships between dimension fields through a pivot table has a cumbersome operation process and low data table processing efficiency. The present invention improves the accuracy and efficiency of identifying field relationships in a data table, provides a basis for subsequent data table proofreading and field recommendation, improves the efficiency of data table processing, and enhances the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings according to these drawings without creative efforts.
[0066] Figure 1 is a flowchart of the data table field relationship recognition method provided by the present invention;
[0067] Figure 2 is an overall flowchart of the data table field relationship recognition method provided by the present invention;
[0068] Figure 3 is a structural diagram of the data table field relationship recognition apparatus provided by the present invention;
[0069] Figure 4 is a structural diagram of the electronic device provided by the present invention. Specific Embodiments
[0070] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0071] Figure 1 It is a schematic flowchart of the method for identifying the relationship of data table fields provided by an embodiment of the present invention. As Figure 1 shown, the method for identifying the relationship of data table fields provided by the present invention includes the following steps:
[0072] Step 101: Determine the types of each field in the data table to be processed, and determine the dimension fields and the cell sets corresponding to the dimension fields according to the types of the fields; wherein, the dimension field is a field used to describe the meaning represented by the data in the cell.
[0073] In most cases, the data in the same column of the data table has the same attribute information. For example, the cells in the same column are all used to describe the name of the user. Therefore, in this embodiment, the cell set refers to a certain column in the data table. In other embodiments, the cell set may also refer to a certain row in the data table. In this case, the data in the same row of the data table has the same attribute information, such as the cells in the same row are all used to describe the name of the user.
[0074] After confirming the types of each field in the data table to be processed, it is possible to determine which fields belong to the dimension fields and which fields belong to the measure fields according to the types of each field. Among them, the dimension field is a field used to describe what the problem is, and essentially belongs to the "classification field", while the measure field is a field used to describe the quantity, and belongs to the quantization field. For example, the mobile phone number type and the ID card number type belong to the dimension fields, and the dimension fields also include other field types, which can be specifically seen in the following embodiments. The numerical type belongs to the measure field, which can also be specifically seen in the following embodiments.
[0075] Step 102: Obtain the relevant information of the cell sets corresponding to each dimension field, the relevant information of the constructed cross table, and the information determined according to the types of the dimension fields.
[0076] In this embodiment, it is necessary to obtain the relevant information of the cell sets corresponding to each dimension field, including description information and relationship information; it is also necessary to construct a cross-tabulation based on any two obtained cell sets and obtain the relevant information of the cross-tabulation, including descriptive features and correlation features. A cross-tabulation is a commonly used classified summary table, and data can be queried using a cross-tabulation, which is very intuitive. The construction of a cross-tabulation belongs to a relatively mature technology in the prior art. For example, data can be organized into a cross-tabulation using SQL at the database end, and the specific process of its construction will not be introduced in detail here.
[0077] Step 103: Determine the relationships between each dimension field based on the relevant information of the cell sets corresponding to each dimension field, the relevant information of the constructed cross-tabulation, and the information determined according to the type of the dimension field.
[0078] In this embodiment, based on all the information obtained in the previous step, the relationships between each dimension field are determined. Among them, the relationship between dimension fields is a way to describe the degree of mutual influence between the data in the cell sets corresponding to any two dimension fields. The relationship between dimension fields can be an inclusion relationship, a non-inclusion relationship and a recommended combination relationship, or a non-inclusion relationship and a non-recommended combination relationship.
[0079] It should be noted that the way to determine the relationship between each dimension field can be analyzed and determined through a preset classification model, or determined by judging the magnitude relationship between all the above information and a preset threshold. It can be specifically set according to actual needs and will not be specifically limited here.
[0080] According to the present invention, a method for identifying the relationship between data table fields is provided. The type of each field in the data table to be processed is determined, the dimension fields and the cell sets corresponding to the dimension fields are determined according to the field type, and then the relevant information of the cell sets corresponding to each dimension field, the relevant information of the constructed cross-tabulation, and the information determined according to the type of the dimension field are obtained. The relationships between each dimension field are determined based on the information obtained above. The method for identifying the relationship between data table fields provided by the present invention solves the technical problems in the prior art that the operation process of determining the relationship between dimension fields through a pivot table is cumbersome and the data table processing efficiency is low. The present invention improves the accuracy and efficiency of identifying the relationship between fields in a data table, provides a basis for subsequent data table proofreading and field recommendation, improves the efficiency of data table processing, and enhances the user experience.
[0081] Based on any of the above embodiments, in this embodiment, obtaining the relevant information of the cell sets corresponding to each dimension field, the relevant information of the constructed cross-table, and the information determined according to the type of the dimension field includes: obtaining the description information of the cell sets corresponding to each dimension field, the relationship information between any two cell sets, the description features and correlation features of the cross-table constructed by any two cell sets, and the enumerated value type of the cell set corresponding to the dimension field determined according to the types of any two dimension fields;
[0082] Correspondingly, according to the relevant information of the cell sets corresponding to each dimension field, the relevant information of the constructed cross-table, and the information determined according to the type of the dimension field, determining the relationship between each dimension field includes:
[0083] Determining the relationship between each dimension field according to the description information of the cell sets corresponding to each dimension field, the relationship information between any two cell sets, the description features and correlation features of the cross-table constructed by any two cell sets, and the enumerated value type of the cell set corresponding to the dimension field determined according to the types of any two dimension fields.
[0084] In this embodiment, after determining the dimension field and the cell set corresponding to the dimension field, it is also necessary to obtain the description information of the cell set corresponding to the dimension field, the relationship information between the cell sets, the description features and correlation features of the cross-table constructed by any two cell sets, and the enumerated value type of the cell set corresponding to the dimension field determined according to the types of any two dimension fields. Among them, in this embodiment, the cell set refers to a column corresponding to a certain field name, and the description information refers to some basic information used to describe the cells in this column, which may include information such as the average value and maximum value of the data length in the cell, and can be specifically obtained through the descriptive statistics function module. The relationship information between the cell sets refers to the inclusion relationship information that can be determined between any two columns according to the data information in the cells in each column.
[0085] In this embodiment, it is also necessary to construct a cross-table according to the cell sets corresponding to any two dimension fields. After the cross-table is constructed, it is necessary to extract the description features and correlation features of the cross-table. Among them, the description features refer to the feature information describing the situation of the cross-table, and the description features can be extracted through the descriptive statistics function module in the table. The correlation features refer to the feature information used to describe the correlation between two cell sets. In this embodiment, the correlation features of the cross-table can be extracted by using the correlation analysis method in the multivariate statistical analysis method. The specific content included in the correlation features is shown in the following embodiments and will not be introduced in detail here.
[0086] It should be noted that the multivariate statistical analysis method is a method for studying the statistical regularity of the interdependence between multiple variables (or multiple factors) in objective things. One of its important bases is multivariate normal analysis, so it is also called multivariate analysis.
[0087] In this embodiment, it is also necessary to select any two dimension fields, and determine the enumeration value type of the cell set corresponding to the two dimension fields according to the types of the two dimension fields. An enumeration value refers to a way of defining an ordered set by using an identifier that lists all values predefined, and the order of these values is the same as the order of the identifiers in the enumeration type description. Assume that the form of the enumeration value is <identifier 1>=<type 1>, such as <n1>=<Time type>, <n2>= <Date type>, <n3>= <string type>, etc. In this embodiment, the types of the two selected dimension fields are time type and date type respectively, so the enumerated value types of the two cell sets obtained are <N2, N1> = <date type, time type>.
[0088] In this embodiment, based on all the information obtained above, the relationships between the respective dimension fields are determined. Among them, the relationships between the dimension fields are a way to describe the degree of mutual influence between the data in the cell sets corresponding to any two dimension fields. The relationships between the dimension fields can be an inclusion relationship, a non-inclusion relationship and a recommended combination relationship, or a non-inclusion relationship and a non-recommended combination relationship, which are not specifically limited herein.
[0089] According to the present invention, a method for identifying the relationships of data table fields is provided. The types of the respective fields in the data table to be processed are determined, the dimension fields and the cell sets corresponding to the dimension fields are determined according to the field types, and then the description information of the cell sets corresponding to the respective dimension fields, the relationship information between any two cell sets, the description features and correlation features of the cross-table constructed by any two cell sets, and the enumerated value types of the corresponding cell sets determined according to the types of any two dimension fields are obtained. The relationships between the respective dimension fields are determined based on the above information. The method for identifying the relationships of data table fields provided by the present invention solves the technical problems in the prior art that when determining the relationships between dimension fields through a pivot table, the entire operation process is cumbersome and the data table processing efficiency is low. The present invention improves the accuracy and efficiency of identifying the relationships of fields in the data table, provides a basis for subsequent data table proofreading and field recommendation, and at the same time improves the data table processing efficiency and enhances the user experience.
[0090] Based on any of the above embodiments, in this embodiment, obtaining the description information of the cell sets corresponding to the respective dimension fields, the relationship information between any two cell sets, the description features and correlation features of the cross-table constructed by any two cell sets, and the enumerated value types of the corresponding cell sets determined according to the types of any two dimension fields includes:
[0091] Obtaining the description information of the first cell set, the description information of the second cell set, and the relationship information between the first cell set and the second cell set; wherein, the first cell set is the cell set corresponding to the first dimension field, and the second cell set is the cell set corresponding to the second dimension field; the first dimension field is any one of the dimension fields in the data table to be processed, and the second dimension field is any one of the dimension fields in the data table to be processed that is different from the first dimension field;
[0092] Construct a cross - table of the first cell set and the second cell set, and obtain the descriptive features and correlation features of the cross - table;
[0093] Determine the enumerated value types of the first cell set and the second cell set according to the types of the first - dimension field and the second - dimension field.
[0094] In this embodiment, after determining the dimension fields and the corresponding cell sets according to the types of the fields in the data table, any two dimension fields and the corresponding cell sets are selected from multiple dimension fields. The two dimension fields are identified as the first - dimension field and the second - dimension field. Among them, the first cell set is the cell set corresponding to the first - dimension field, which is the column corresponding to the first - dimension field in this embodiment. The second cell set is the cell set corresponding to the second - dimension field, which is the column corresponding to the second - dimension field in this embodiment. They are not the same. As shown in the data table of Table 1 below, according to the definitions of the dimension fields and the measure fields, it can be determined that except for the field with the field name "Inventory Age" which belongs to the measure field, the rest belong to the dimension fields. Suppose the field name of the first - dimension field is "Vehicle Information", and the field name of the second - dimension field is "Body Color". Then, the first cell set is the set composed of each cell corresponding to "Vehicle Information", and the specific data is "Z*****47, Z*****48, Z*****49, Z*****37"; the second cell set is the set composed of each cell corresponding to "Body Color", and the specific data is "Glacier Blue, Pearl White, Elegant Black, Mocha Brown". It should be noted that the second - dimension field is any dimension field different from the first - dimension field in the data table, then the second cell set is any cell set different from the first cell set in the data table.
[0095] Table 1
[0096]
[0097] In this embodiment, it is necessary to obtain the descriptive information of the first cell set, the descriptive information of the second cell set, and the relationship information between the first cell set and the second cell set. Among them, the descriptive information is the information obtained through objective recording, induction, analysis, and reasoning of the table data. The descriptive information can be information such as the minimum length and maximum length of the data in the cell. The specific content included in the descriptive information of the first cell set and the descriptive information of the second cell set will be introduced in detail in the following embodiments.
[0098] In this embodiment, it is also necessary to obtain the relationship information between the first cell set and the second cell set to determine whether there is an inclusion relationship between the two. Specifically, it can be determined according to the data of the first cell set and the second cell set. For example, if the data in the first cell set is related to "brand" and the data in the second cell set is related to "series", since the data of "series" is related data under this "brand", it can be determined that the relationship between the two is an inclusion relationship. It should be noted that the specific determination method can be determined according to actual needs and is not specifically limited here.
[0099] In this embodiment, it is also necessary to construct a cross-table based on the first cell set and the second cell set, and obtain the description features and correlation features of the cross-table. Among them, the description features include information such as the number of cells with non-repeated data and the number of non-empty cells in the cross-table. The specific content included in the correlation features can be seen in the following embodiments and will not be elaborated in detail here.
[0100] In this embodiment, it is also necessary to determine the enumerated value types of the first cell set and the second cell set according to the types of the first dimension field and the second dimension field. Assume that the form of the enumerated value is <identifier1> = <type1>, such as <n1>= <Time type>, <n2>= <Date type>, <n3>= <string type>. If the type of the first - dimension field is string type and the type of the second - dimension field is date type, the enumeration value types of the first cell set and the second cell set can be determined. For example, <N2, N3> = <date type, string type>. According to a preset form, the enumeration value types of the corresponding cells in the two sets are determined, and the obtained enumeration value types have a certain order.
[0101] According to the present invention, a method for identifying the relationship of data - table fields is provided. The first cell set and the second cell set are determined from multiple cell sets, the description information of the first cell set, the description information of the second cell set, and the relationship information between them are obtained, a cross - table is constructed, the description features and correlation features of the constructed cross - table, and the enumeration value types of the first cell set and the second cell set are obtained. The method for identifying the relationship of data - table fields provided by the present invention provides data support for subsequent identification of the relationship between dimension fields by obtaining the relevant information of the cell sets corresponding to the dimension fields, indirectly improving the accuracy and efficiency of identifying the field relationship in the data table, and providing a basis for subsequent data - table proofreading and field recommendation.
[0102] Based on any of the above - mentioned embodiments, in this embodiment, according to the description information of the cell sets corresponding to each dimension field, the relationship information between any two cell sets, the description features and correlation features of the cross - table constructed by any two cell sets, and the enumeration value types of the cell sets corresponding to any two dimension fields determined according to the types of any two dimension fields, the relationship between each dimension field is determined, including:
[0103] Input the description information of the first cell set, the description information of the second cell set, the relationship information between the first cell set and the second cell set, the description features of the cross - table, the correlation features of the cross - table, and the enumeration value types of the first cell set and the second cell set into a pre - trained first classification model to obtain a first probability value indicating whether there is an inclusion relationship between the first dimension field and the second dimension field;
[0104] Input the description information of the first cell set, the description information of the second cell set, the relationship information between the first cell set and the second cell set, the description features of the cross - table, the correlation features of the cross - table, and the enumeration value types of the first cell set and the second cell set into a pre - trained second classification model to obtain a second probability value indicating whether combination is recommended in the case of non - inclusion relationship between the first dimension field and the second dimension field;
[0105] Determine the relationship between the first dimension field and the second dimension field according to the first probability value and the second probability value;
[0106] Among them, the first classification model is trained based on the description information of the first sample cell set, the description information of the second sample cell set, the relationship information between the first sample cell set and the second sample cell set, the description features of the contingency table formed by the first sample cell set and the second sample cell set, the correlation features of the contingency table formed by the first sample cell set and the second sample cell set, the enumerated value types of the first sample cell set and the second sample cell set, and the label information indicating whether there is an inclusion relationship between the first sample cell set and the second sample cell set;
[0107] The second classification model is trained based on the description information of the first sample cell set, the description information of the second sample cell set, the relationship information between the first sample cell set and the second sample cell set, the description features of the contingency table formed by the first sample cell set and the second sample cell set, the correlation features of the contingency table formed by the first sample cell set and the second sample cell set, the enumerated value types of the first sample cell set and the second sample cell set, and the label information indicating whether combination is recommended under the non-inclusion relationship between the first sample cell set and the second sample cell set.
[0108] In this embodiment, in order to confirm the relationship between the first dimension field and the second dimension field, the obtained description information of the first cell set, the description information of the second cell set, the relationship information between the first cell set and the second cell set, the description features of the contingency table, the correlation features of the contingency table, and the enumerated value types of the first cell set and the second cell set are input into the pre-trained first classification model to obtain a first probability value indicating whether there is an inclusion relationship between the first dimension field and the second dimension field, where the first classification model is a classification model for judging the inclusion relationship; at the same time, all the above information also needs to be input into the second classification model to obtain a second probability value, and the second classification model is a classification model for judging whether combination analysis is recommended under the non-inclusion relationship, and then the relationship between the first dimension field and the second dimension field is determined according to the obtained first probability value and the second probability value.
[0109] It should be noted that the relationship between dimension fields can be an inclusion relationship. The inclusion relationship is usually a genus-species relationship, which refers to a subordinate relationship. For example, in the same category, the range of A is smaller, the range of B is larger, and the range of B includes the range of A, indicating that A and B have an inclusion relationship. It can also be a relationship of whether combination is recommended under a non-inclusion relationship. For example, if it is determined that there is no inclusion relationship between two fields based on the relationship information between the first cell set and the second cell set, all the obtained information can be input into the second classification model to determine the relationship of whether combination is recommended. For dimension fields of time type and region type, since the two are in a non-inclusion relationship, it can be determined whether the two should be combined through the second classification model.
[0110] In this embodiment, both the first classification model and the second classification model are obtained by pre-training training samples using the random forest algorithm. Among them, Random forest refers to a classifier that uses multiple decision trees to train and predict training samples. Using the random forest algorithm, the description information of the first sample cell set, the description information of the second sample cell set, the relationship information between the first sample cell set and the second sample cell set, the description features of the contingency table formed by the first sample cell set and the second sample cell set, the correlation features of the contingency table formed by the first sample cell set and the second sample cell set, the enumerated value types of the first sample cell set and the second sample cell set, and the label information indicating whether there is an inclusion relationship between the first sample cell set and the second sample cell set are used to train the classifier to obtain the first classification model. Similarly, the second classification model is obtained by training the classifier using the pre-obtained second training sample information with the random forest algorithm. The specific training method will not be introduced in detail here.
[0111] It should be noted that the label information refers to the information obtained by annotating the attributes of the training samples. For example, when training a classification model for whether there is an inclusion relationship, it is first determined manually whether the first sample cell set and the second sample cell set have an inclusion relationship and marked. The marked information indicating an inclusion relationship and the information indicating no inclusion relationship are both label information. Similarly, the marked information indicating combination is recommended under a non-inclusion relationship and the marked information indicating combination is not recommended under a non-inclusion relationship can be determined as the label information indicating whether there is an inclusion relationship between the first sample cell set and the second sample cell set.
[0112] According to the method for identifying the relationship of data table fields provided by the present invention, all the obtained information is input into the first classification model obtained by preset training to obtain a first probability value. At the same time, all the obtained information is also input into the second classification model to obtain a second probability value. The relationship information between the first dimension field and the second dimension field is determined according to the first probability value and the second probability value, which can accurately identify the relationship between the dimension fields, simplify the process of identifying the relationship of dimension fields, and improve the accuracy and efficiency of identifying the relationship of dimension fields.
[0113] Based on any of the above embodiments, in this embodiment, determining the relationship information between the first dimension field and the second dimension field according to the first probability value and the second probability value includes:
[0114] When the first probability value is greater than the second probability value, the relationship between the first dimension field and the second dimension field is an inclusion relationship;
[0115] When the first probability value is less than the second probability value, and the second probability value is greater than or equal to a preset first threshold, the relationship between the first dimension field and the second dimension field is a non-inclusion and recommended combination relationship;
[0116] When the first probability value is less than the second probability value, and the second probability value is less than the preset first threshold, the relationship between the first dimension field and the second dimension field is a non-inclusion and non-recommended combination relationship.
[0117] In this embodiment, by comparing the magnitude relationship between the obtained first probability value and the second probability value, the relationship between the first dimension field and the second dimension field is determined. Suppose the preset first threshold is 0.5. When the first probability value is greater than the second probability value, it is directly determined that there is an inclusion relationship between the first dimension field and the second dimension field, and there is no need to compare with the preset first threshold; when the first probability value is less than the second probability value, it is necessary to compare the second probability value with the first threshold for confirmation. For example, when the obtained first probability value is 0.4 and the second probability value is 0.6, that is, the second probability value 0.6 is greater than the first threshold 0.5, the relationship between the first dimension field and the second dimension field is determined to be a non-inclusion relationship and a recommended combination relationship; for another example, when the obtained first probability value is 0.4 and the second probability value is 0.45, that is, the second probability value is greater than the first probability value and less than the first threshold 0.5, the relationship between the first dimension field and the second dimension field is determined to be a non-inclusion relationship and a non-recommended combination relationship. It should be noted that the size of the preset first threshold can be set according to actual needs and is not specifically limited here.
[0118] According to the method for identifying the relationship of data table fields provided by the present invention, by comparing the magnitude relationship between the first probability value and the second probability value obtained and a preset first threshold value, the relationship between the first dimension field and the second dimension field can be accurately identified, the speed and accuracy of dimension field relationship identification are improved, a basis is provided for subsequent data analysis using field relationships, and the processing speed of subsequent data tables is improved.
[0119] Based on any of the above embodiments, in this embodiment, obtaining the description information of the first cell set, the description information of the second cell set, and the relationship information between the first cell set and the second cell set includes:
[0120] Obtaining the index value of the first cell set in the data table to be processed;
[0121] Obtaining the maximum value of the data lengths in each cell of the first cell set;
[0122] Obtaining the minimum value of the data lengths in each cell of the first cell set;
[0123] Obtaining the average value of the data lengths in each cell of the first cell set;
[0124] Obtaining the standard deviation of the data lengths in each cell of the first cell set;
[0125] Obtaining the number of non-repeated data in the first cell set;
[0126] Obtaining the number of cells in the first cell set;
[0127] Obtaining the index value of the second cell set in the data table to be processed;
[0128] Obtaining the maximum value of the data lengths in each cell of the second cell set;
[0129] Obtaining the minimum value of the data lengths in each cell of the second cell set;
[0130] Obtaining the average value of the data lengths in each cell of the second cell set;
[0131] Obtaining the standard deviation of the data lengths in each cell of the second cell set;
[0132] Obtaining the number of non-repeated data in the second cell set;
[0133] Obtaining the product of the number of non-repeated data in the first cell set and the number of non-repeated data in the second cell set;
[0134] Obtaining the information on whether there is an inclusion relationship between the first cell set and the second cell set.
[0135] In this embodiment, the content specifically included in the description information of the first cell set, the content included in the description information of the second cell set, and the relationship information between the first cell set and the second cell set are determined. Among them, the relationship information refers to the information on whether the first cell set and the second cell set have an inclusion relationship and the information obtained by combining the two. For example, the product of the number of non-repeated data in the first cell set and the number of non-repeated data in the second cell set, and the information on whether there is an inclusion relationship between the first cell set and the second cell set. The description information of the first cell set includes: the index value of the first cell set in the data table to be processed, the maximum value of the data length in each cell, the minimum value of the data length in each cell, the average value of the data length in each cell, the standard deviation of the data length in each cell, the number of non-repeated data, and the number of cells; the description information of the second cell set includes: the index value of the second cell set in the data table to be processed, the maximum value of the data length in each cell of the second cell set, the minimum value of the data length in each cell of the second cell set, the average value of the data length in each cell of the second cell set, the standard deviation of the data length in each cell of the second cell set, and the number of non-repeated data in the second cell set.
[0136] It should be noted that in a relational database, an index is a separate, physical storage structure for sorting the values of one or more columns in a database table. It is a collection of the values of one or several columns in a certain table and a list of logical pointers corresponding to the data pages that physically identify these values in the table. The role of an index is equivalent to the table of contents of a book, and the required content can be quickly found according to the page numbers in the table of contents. The index value refers to the position code corresponding to each dimension field and can be used to determine the specific position of the dimension field. For example, index values are set for each field in the data table shown in Table 1 above, as shown in Table 2 below. When the index value of the first cell set is 1 and the index value of the second cell set is 2, the first cell set can be determined to be the first column in the data table shown in Table 1, and the second cell set can be determined to be the second column in the data table of Table 1.
[0137] Table 2
[0138] Field index value Field name Field type 1 Vehicle information eng 2 Vehicle status vn 3 Vehicle series eng 4 Vehicle model eng 5 Body color n 6 Interior color n 7 Engine number eng 8 Age of inventory Number
[0139] In this embodiment, as shown in Table 1 above, after determining the position of the first cell set according to the index value 1 of the first cell set, it is also obtained that the maximum length of the data in each cell of the first cell set is 8, the minimum length is 8, the average length is 8, the standard deviation is 0, and the number of non-repeated data is 4; similarly, according to the index value 2 of the second cell set, the position of the second cell set is determined, and it is obtained that the maximum length of each cell in the second cell set is 4, the minimum length is 2, the average length is 3, the length standard deviation is 1, and the number of non-repeated data is 2.
[0140] Based on the information obtained above, it is also obtained that the product of the number of non-repeated data in the first cell set and the number of non-repeated data in the second cell set is 4 * 2 = 8; and according to the type of the first dimension field corresponding to the first cell set and the type of the second dimension field corresponding to the second cell set, it is determined that there is a non-inclusive relationship between the two.
[0141] According to the method for identifying the relationship of data table fields provided by the present invention, the description information of the first cell set, the description information of the second cell set, and the relationship information between the two are determined, providing data support for accurately determining the relationship between the first dimension field and the second dimension field in the future.
[0142] Based on any of the above embodiments, in this embodiment, a cross-table of the first cell set and the second cell set is constructed, and the description features and correlation features of the cross-table are obtained, including:
[0143] Construct a cross-table of the first cell set and the second cell set;
[0144] Obtain the description features of the cross-table, including: calculating the number of non-repeated data in the cross-table, calculating the number of non-empty cells in the cross-table;
[0145] Obtain the correlation features of the cross-table, including: calculating the P value of the chi-square test of the cross-table, calculating the degrees of freedom of the chi-square test of the cross-table, calculating whether the P value of the chi-square test of the cross-table is less than a preset second threshold, calculating the quotient of the number of cells in the cross-table whose absolute value of the correlation coefficient is greater than or equal to a third threshold and 2; calculating the percentage of the number of cells in the cross-table whose absolute value of the correlation coefficient is greater than or equal to a third threshold in the total number of cells, calculating the average value of the data in each cell of the cross-table, calculating the standard deviation of the average value of the data in each cell of the cross-table.
[0146] In this embodiment, a contingency table is constructed based on the confirmed first cell set and second cell set, and the descriptive features and correlation features of the contingency table are obtained. Among them, calculating the descriptive features of the contingency table at least includes: calculating the number of non-repeated data, and calculating the number of non-empty cells in the contingency table; while calculating and obtaining the correlation features at least includes: calculating the P-value of the chi-square test of the contingency table, calculating the degrees of freedom of the chi-square test of the contingency table, calculating whether the P-value of the chi-square test of the contingency table is less than a preset second threshold, calculating the quotient of the number of cells in the contingency table whose absolute value of the correlation coefficient is greater than or equal to a third threshold and 2; calculating the percentage of the number of cells in the contingency table whose absolute value of the correlation coefficient is greater than or equal to the third threshold in the total number of cells, calculating the average value of the data in each cell of the contingency table, and calculating the standard deviation of the average value of the data in each cell of the contingency table. Among them, the second threshold and the third threshold can be set according to experience. For example, the second threshold is set to 0.5 and the third threshold is set to 0.75. It should be noted that the magnitudes of the second threshold and the third threshold can be set according to actual needs and are not specifically limited herein.
[0147] It should be noted that a contingency table is a commonly used table for classification and summary. Essentially, it is the intersection of rows and columns and is used to present the data on the row as column indicators. The data can be organized into a contingency table using SQL at the database end and then presented in the form of an ordinary report; it can also be directly implemented through Crystal Reports. The chi-square test is a very widely used hypothesis testing method. The chi-square test essentially measures the degree of deviation between the actual observed values and the theoretically inferred values of a statistical sample. The degree of deviation between the actual observed values and the theoretically inferred values determines the size of the chi-square value. If the chi-square value is larger, the deviation between the two is greater; conversely, the deviation between the two is smaller; if the two values are exactly equal, the chi-square value is 0, indicating that the theoretical value is completely consistent. In this embodiment, the contingency table can be tested and processed through the chi-square test method in the prior art to obtain each data value, and the specific processing process is not introduced in detail herein.
[0148] In addition, in this embodiment, the correlation analysis method in the multivariate statistical analysis method is used to extract the correlation features of the contingency table, and the specific extraction method is not introduced in detail.
[0149] According to the method for identifying the relationship between data table fields provided by the present invention, the descriptive features and correlation features of the contingency table constructed based on the first cell set and the second cell set are determined, providing data support for accurately determining the relationship between the first dimension field and the second dimension field in the subsequent process.
[0150] Based on any of the above embodiments, in this embodiment, determining the types of each field in the data table to be processed, and determining the dimension fields and the cell sets corresponding to the dimension fields according to the types of the fields includes:
[0151] Determine the types of each field in the data table to be processed;
[0152] For any field in the data table to be processed, when the type of the field meets the first condition, determine the field as a dimension field; wherein, the first condition is a condition for judging whether the types of each field in the data table to be processed belong to dimension fields;
[0153] Obtain the cell set corresponding to the dimension field from the data table to be processed; wherein, the cell set corresponding to the dimension field is the row or column corresponding to the dimension field in the data table to be processed.
[0154] In this embodiment, the first condition refers to various field types belonging to dimension fields, including any one of the following field types: date and time type, date type, time type, string type, where the string type includes: person name, place name, gerund, English, text-based number, other single string type, mixed type. When it is determined that the type of the first field is a string type, that is, the type of the first field meets the preset first condition, determine the first field as a dimension field, and determine the cell set corresponding to the first field according to the position of the first field in the data table; wherein, a field with a numerical type does not belong to a dimension field. When it is determined that the type of the first field is a numerical type, that is, the type of the first field does not meet the first condition, do not determine the first field as a dimension field, and judge whether it belongs to a metric field. If it belongs, classify it as a metric field. It should be noted that the cell set can be the column corresponding to the first field in the data table, or the row corresponding to the first field in the data table, and no specific limitation is made here.
[0155] According to the data table field relationship recognition method provided by the present invention, by determining the types of each field in the data table, then judging whether the type of the field meets the first condition, if it meets, determining the field as a dimension field, and determining the corresponding cell set according to the position of the field in the data table, it provides data support for the subsequent relationship recognition between each dimension field and improves the speed of data table processing.
[0156] Based on any of the above embodiments, in this embodiment, determining the types of each field in the data table to be processed includes:
[0157] Obtain the data table to be processed; wherein, the data table includes fields and cells, and the cells include data;
[0158] Determine the type of each cell in the data table according to the data contained in the cell;
[0159] Determine the type of the field according to the types of the cells corresponding to the field.
[0160] In this embodiment, according to the data contained in the cells of the data table to be processed obtained, the types of each cell are determined, and then according to the types of all the cells corresponding to a certain field, the type of the field is determined. For example, if the types of all the cells corresponding to the first field are all date types, then the date type is determined as the type of the first field. Among them, for the types of all the cells corresponding to the field, when there are multiple types, the type with the most occurrences can be determined as the type of the field, and the specific determination method is not specifically limited herein.
[0161] According to the method for identifying the relationship between data table fields provided by the present invention, by determining the types of each cell in the data table, and then determining the type of the field according to the types of all the cells corresponding to the field, through the recognition and processing of the types of each cell, the type of the field is determined, which provides data support for the subsequent identification of the relationship between fields in each dimension and improves the processing speed of the data table.
[0162] Based on any of the above embodiments, in this embodiment, as Figure 2 shown, obtain the data table to be recognized, determine the types of each field in the data table, determine the dimension fields and the cell sets corresponding to the dimension fields according to the types of each field. In this embodiment, determine the first cell set and the second cell set. Among them, the cell set is a dimension list. Determine the respective description information and the relationship information between the two according to the obtained first cell set and second cell set. Construct a cross-table according to the first cell set and the second cell set, and determine the description characteristics and correlation characteristics of the cross-table, as well as the enumerated value types of the first cell set and the second cell set. Input all the obtained feature information into the first classification model for scoring to obtain the first probability value, and at the same time input it into the second classification model for scoring to obtain the second probability value. Determine the relationship between the two dimension columns according to the first probability value and the second probability value. Among them, the first classification model is a containment relationship classification model, and the second classification model is a classification model for whether combination analysis is recommended under the non-containment relationship.
[0163] It should be noted that the specific content included in the description information of the first cell set, the description information of the second cell set, the relationship information between the first cell set and the second cell set, the description characteristics and correlation characteristics of the cross-table, and the enumerated value types of the first cell set and the second cell set are as described in the above embodiments, and will not be elaborated in detail herein.
[0164] It should be noted that after obtaining the relationships between the dimension fields, corresponding recommended fields can be generated based on the relationships between the dimension fields. For example, in the case where two dimension fields do not have an inclusion relationship, the field recommendation score is determined based on the correlation between the two dimension fields, the recommended fields are determined according to the magnitude of the recommendation score, and a corresponding pivot table is generated based on the obtained recommended fields to analyze the data in the data table.
[0165] In the already obtained pivot table, the relationship between the newly added dimension field and the existing dimension fields can also be determined, and the newly added dimension field can be set at the corresponding position in the pivot table. If the newly added dimension field and the existing dimension fields have an inclusion relationship, the newly added field is automatically positioned in the same row or the same column; if the newly added dimension field and the existing dimension fields have a non-inclusion relationship and combination is not recommended, then the newly added dimension field and the existing dimension fields are set as one row and one column, and cannot be placed in the same row at the same time. According to the method for identifying the relationship between data table fields provided by the present invention, the newly added dimension field can be automatically positioned, improving the efficiency of data table processing.
[0166] Figure 3 A device for identifying the relationship between data table fields provided by the present invention, as Figure 3 shown, the device for identifying the relationship between data table fields provided by the present invention includes:
[0167] A first determination module 301, configured to determine the types of each field in the data table to be processed, and determine the dimension fields and the cell sets corresponding to the dimension fields according to the types of the fields; wherein, the dimension fields are fields used to describe the meanings represented by the data in the cells.
[0168] An acquisition module 302, configured to acquire the relevant information of the cell sets corresponding to each dimension field, the relevant information of the constructed cross table, and the information determined according to the types of the dimension fields.
[0169] A second determination module 303, configured to determine the relationships between each dimension field according to the relevant information of the cell sets corresponding to each dimension field, the relevant information of the constructed cross table, and the information determined according to the types of the dimension fields.
[0170] According to the device for identifying the relationship between data table fields provided by the present invention, the technical problem in the prior art of determining the relationship between dimension fields through a pivot table, with the entire operation process being cumbersome and the data table processing efficiency being low, is solved. The present invention improves the accuracy and efficiency of identifying the field relationships in the data table, provides a basis for subsequent data table proofreading and field recommendation, and at the same time improves the data table processing efficiency and enhances the user experience.
[0171] Based on any of the above embodiments, in this embodiment, the obtaining module 302 is further configured to:
[0172] Obtain the description information of the cell sets corresponding to each dimension field, the relationship information between any two cell sets, the description features and correlation features of the cross - tables constructed by any two cell sets, and determine the enumerated value type of the cell set corresponding to the dimension field according to the types of any two dimension fields;
[0173] The second determination module 303 is further configured to:
[0174] Determine the relationship between each dimension field according to the description information of the cell sets corresponding to each dimension field, the relationship information between any two cell sets, the description features and correlation features of the cross - tables constructed by any two cell sets, and the enumerated value type of the cell set corresponding to the dimension field determined according to the types of any two dimension fields.
[0175] According to the data table field relationship recognition device provided by the present invention, it solves the technical problems in the prior art that determining the relationship between dimension fields through a pivot table has a cumbersome operation process and low data table processing efficiency. The present invention improves the accuracy and efficiency of field relationship recognition in a data table, provides a basis for subsequent data table proofreading and field recommendation, and at the same time improves the efficiency of data table processing and enhances the user experience.
[0176] Based on any of the above embodiments, in this embodiment, the obtaining module 302 is further configured to:
[0177] Obtain the description information of the first cell set, the description information of the second cell set, and the relationship information between the first cell set and the second cell set; wherein, the first cell set is the cell set corresponding to the first dimension field, and the second cell set is the cell set corresponding to the second dimension field; the first dimension field is any dimension field in the data table to be processed, and the second dimension field is any dimension field in the data table to be processed that is different from the first dimension field;
[0178] Construct a cross - table of the first cell set and the second cell set, and obtain the description features and correlation features of the cross - table;
[0179] Determine the enumerated value type of the first cell set and the second cell set according to the type of the first dimension field and the type of the second dimension field.
[0180] According to the data table field relationship recognition device provided by the present invention, by obtaining the correlation between the first cell set and the second cell set, it provides data support for subsequent determination of the relationship between dimension fields, and at the same time improves the efficiency of data table processing and enhances the user experience.
[0181] Based on any of the above embodiments, in this embodiment, the second determination module 303 is further configured to:
[0182] Input the description information of the first cell set, the description information of the second cell set, the relationship information between the first cell set and the second cell set, the description features of the cross-tabulation, the correlation features of the cross-tabulation, and the enumeration value types of the first cell set and the second cell set into a pre-trained first classification model to obtain a first probability value indicating whether there is an inclusion relationship between the first dimensional field and the second dimensional field;
[0183] Input the description information of the first cell set, the description information of the second cell set, the relationship information between the first cell set and the second cell set, the description features of the cross-tabulation, the correlation features of the cross-tabulation, and the enumeration value types of the first cell set and the second cell set into a pre-trained second classification model to obtain a second probability value indicating whether combination is recommended under the non-inclusion relationship between the first dimensional field and the second dimensional field;
[0184] Determine the relationship between the first dimensional field and the second dimensional field according to the first probability value and the second probability value;
[0185] Wherein, the first classification model is trained based on the description information of the first sample cell set, the description information of the second sample cell set, the relationship information between the first sample cell set and the second sample cell set, the description features of the cross-tabulation formed by the first sample cell set and the second sample cell set, the correlation features of the cross-tabulation formed by the first sample cell set and the second sample cell set, the enumeration value types of the first sample cell set and the second sample cell set, and the label information indicating whether there is an inclusion relationship between the first sample cell set and the second sample cell set;
[0186] The second classification model is trained based on the description information of the first sample cell set, the description information of the second sample cell set, the relationship information between the first sample cell set and the second sample cell set, the description features of the cross-tabulation formed by the first sample cell set and the second sample cell set, the correlation features of the cross-tabulation formed by the first sample cell set and the second sample cell set, the enumeration value types of the first sample cell set and the second sample cell set, and the label information indicating whether combination is recommended under the non-inclusion relationship between the first sample cell set and the second sample cell set.
[0187] According to the data table field relationship recognition device provided by the present invention, all the obtained information is input into the first classification model obtained by preset training to obtain a first probability value. At the same time, all the obtained information is also input into the second classification model to obtain a second probability value. The relationship information between the first dimension field and the second dimension field is determined according to the first probability value and the second probability value, which can accurately identify the relationship between the dimension fields, simplify the process of dimension field relationship recognition, and improve the accuracy and efficiency of dimension field relationship recognition.
[0188] Based on any of the above embodiments, in this embodiment, the second determination module 303 is further configured to:
[0189] When the first probability value is greater than the second probability value, the relationship between the first dimension field and the second dimension field is an inclusion relationship;
[0190] When the first probability value is less than the second probability value and the second probability value is greater than or equal to a preset first threshold, the relationship between the first dimension field and the second dimension field is a non-inclusion and recommended combination relationship;
[0191] When the first probability value is less than the second probability value and the second probability value is less than the preset first threshold, the relationship between the first dimension field and the second dimension field is a non-inclusion and non-recommended combination relationship.
[0192] According to the data table field relationship recognition device provided by the present invention, by comparing the size relationship between the obtained first probability value and the second probability value and the preset first threshold, the relationship between the first dimension field and the second dimension field can be accurately identified, the speed and accuracy of dimension field relationship recognition are improved, a basis is provided for subsequent data analysis using field relationships, and the processing speed of subsequent data tables is improved.
[0193] Based on any of the above embodiments, in this embodiment, the acquisition module 302 is further configured to:
[0194] Obtain the index value of the first cell set in the data table to be processed;
[0195] Obtain the maximum value of the data lengths in each cell of the first cell set;
[0196] Obtain the minimum value of the data lengths in each cell of the first cell set;
[0197] Obtain the average value of the data lengths in each cell of the first cell set;
[0198] Obtain the standard deviation of the data lengths in each cell of the first cell set;
[0199] Obtain the number of non-repeated data in the first cell set;
[0200] Obtain the number of cells in the first cell set;
[0201] Obtain the index value of the second cell set in the data table to be processed;
[0202] Obtain the maximum value of the data lengths in each cell of the second cell set;
[0203] Obtain the minimum value of the data lengths in each cell of the second cell set;
[0204] Obtain the average value of the data lengths in each cell of the second cell set;
[0205] Obtain the standard deviation of the data lengths in each cell of the second cell set;
[0206] Obtain the number of distinct data in the second cell set;
[0207] Obtain the product of the number of distinct data in the first cell set and the number of distinct data in the second cell set;
[0208] Obtain information on whether there is an inclusion relationship between the first cell set and the second cell set.
[0209] According to the data table field relationship recognition device provided by the present invention, determine the description information of the first cell set, the description information of the second cell set, and the relationship information between the two, providing data support for accurately determining the relationship between the first dimension field and the second dimension field subsequently.
[0210] Based on any of the above embodiments, in this embodiment, the obtaining module 302 is further configured to:
[0211] Construct a cross-tabulation of the first cell set and the second cell set;
[0212] Obtain the descriptive features of the cross-tabulation, including: calculating the number of distinct data in the cross-tabulation, calculating the number of non-empty cells in the cross-tabulation;
[0213] Obtain the correlation features of the cross-tabulation, including: calculating the P value of the chi-square test of the cross-tabulation, calculating the degrees of freedom of the chi-square test of the cross-tabulation, calculating whether the P value of the chi-square test of the cross-tabulation is less than a preset second threshold, calculating the quotient of the number of cells in the cross-tabulation where the absolute value of the correlation coefficient is greater than or equal to a third threshold and 2; calculating the percentage of the number of cells in the cross-tabulation where the absolute value of the correlation coefficient is greater than or equal to a third threshold in the total number of cells, calculating the average value of the data in each cell of the cross-tabulation, calculating the standard deviation of the average value of the data in each cell of the cross-tabulation.
[0214] According to the data table field relationship recognition device provided by the present invention, the description signs and correlation features of the cross table constructed based on the first cell set and the second cell set are determined, providing data support for accurately determining the relationship between the first dimension field and the second dimension field subsequently.
[0215] Based on any of the above embodiments, in this embodiment, the first determination module 301 is further configured to:
[0216] Determine the types of each field in the data table to be processed;
[0217] For any field in the data table to be processed, when the type of the field meets the first condition, the field is determined as a dimension field; wherein, the first condition is a condition for judging whether the types of each field in the data table to be processed belong to dimension fields.
[0218] Obtain the cell set corresponding to the dimension field from the data table to be processed; wherein, the cell set corresponding to the dimension field is the row or column corresponding to the dimension field in the data table to be processed.
[0219] According to the data table field relationship recognition device provided by the present invention, by determining the types of each field in the data table, then judging whether the type of the field meets the first condition, if it meets, determining the field as a dimension field, and determining the corresponding cell set according to the position of the field in the data table, it provides data support for the subsequent recognition of the relationship between each dimension field, and improves the speed of data table processing.
[0220] Based on any of the above embodiments, in this embodiment, the first determination module 301 is further configured to:
[0221] Obtain the data table to be processed; wherein, the data table includes fields and cells, and the cells include data.
[0222] Determine the type of each cell in the data table according to the data included in the cell;
[0223] Determine the type of the field according to the types of each cell corresponding to the field.
[0224] According to the data table field relationship recognition device provided by the present invention, by determining the types of each cell in the data table, and then determining the type of the field according to the types of all cells corresponding to the field, and determining the type of the field by identifying and processing the types of each cell, it provides data support for the subsequent recognition of the relationship between each dimension field, and improves the speed of data table processing.
[0225] Since the principle of the device described in the embodiment of the present invention is the same as that of the method described in the above embodiment, the more detailed explanation content will not be elaborated here.
[0226] Figure 4 This is a schematic diagram of the physical structure of the electronic device provided in the embodiment of the present invention. As Figure 4 shown, the present invention provides an electronic device, including: a processor 401, a memory 402, and a bus 403;
[0227] Among them, the processor 401 and the memory 402 complete mutual communication through the bus 403;
[0228] The processor 401 is used to call program instructions in the memory 402 to execute the methods provided in the above method embodiments, for example, including: determining the types of each field in the data table to be processed, determining the dimension fields and the cell sets corresponding to the dimension fields according to the types of the fields; where the dimension fields are the fields used to describe the meanings represented by the data in the cells; obtaining the relevant information of the cell sets corresponding to each dimension field, the relevant information of the cross-table constructed, and the information determined according to the types of the dimension fields; determining the relationships between each dimension field according to the relevant information of the cell sets corresponding to each dimension field, the relevant information of the cross-table constructed, and the information determined according to the types of the dimension fields.
[0229] The embodiment of the present invention provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the methods provided in the above method embodiments, for example, including: determining the types of each field in the data table to be processed, determining the dimension fields and the cell sets corresponding to the dimension fields according to the types of the fields; where the dimension fields are the fields used to describe the meanings represented by the data in the cells; obtaining the relevant information of the cell sets corresponding to each dimension field, the relevant information of the cross-table constructed, and the information determined according to the types of the dimension fields; determining the relationships between each dimension field according to the relevant information of the cell sets corresponding to each dimension field, the relevant information of the cross-table constructed, and the information determined according to the types of the dimension fields.
[0230] The present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided by the above-mentioned various methods. The method includes: determining the types of each field in the data table to be processed, and determining the dimension fields and the cell sets corresponding to the dimension fields according to the types of the fields; wherein, the dimension fields are the fields used to describe the meanings represented by the data in the cells; obtaining the relevant information of the cell sets corresponding to each dimension field, the relevant information of the cross table constructed, and the information determined according to the types of the dimension fields; determining the relationships between each dimension field according to the relevant information of the cell sets corresponding to each dimension field, the relevant information of the cross table constructed, and the information determined according to the types of the dimension fields.
[0231] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.
[0232] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for identifying data table field relationships, characterized in that, Including: Determine the types of each field in the data table to be processed, and determine the dimension fields and the cell sets corresponding to the dimension fields according to the types of the fields; wherein, the dimension fields are the fields used to describe the meanings represented by the data in the cells. Obtain the relevant information of the cell sets corresponding to each dimension field, the relevant information of the constructed cross-table, and the information determined according to the types of the dimension fields. Determine the relationships between each dimension field according to the relevant information of the cell sets corresponding to each dimension field, the relevant information of the constructed cross-table, and the information determined according to the types of the dimension fields. The obtaining the relevant information of the cell sets corresponding to each dimension field, the relevant information of the constructed cross-table, and the information determined according to the types of the dimension fields includes: Obtain the description information of the cell sets corresponding to each dimension field, the relationship information between any two cell sets, the description features and correlation features of the cross-table constructed by any two cell sets, and the enumerated value types of the cell sets corresponding to the dimension fields determined according to the types of any two dimension fields. The determining the relationships between each dimension field according to the relevant information of the cell sets corresponding to each dimension field, the relevant information of the constructed cross-table, and the information determined according to the types of the dimension fields includes: Determine the relationships between each dimension field according to the description information of the cell sets corresponding to each dimension field, the relationship information between any two cell sets, the description features and correlation features of the cross-table constructed by any two cell sets, and the enumerated value types of the cell sets corresponding to the dimension fields determined according to the types of any two dimension fields.
2. The method for identifying the relationship of data table fields according to claim 1, wherein The obtaining the description information of the cell sets corresponding to each dimension field, the relationship information between any two cell sets, the description features and correlation features of the cross-table constructed by any two cell sets, and the enumerated value types of the corresponding cell sets determined according to the types of any two dimension fields includes: Obtain the description information of the first cell set, the description information of the second cell set, and the relationship information between the first cell set and the second cell set; wherein, the first cell set is the cell set corresponding to the first dimension field, and the second cell set is the cell set corresponding to the second dimension field; the first dimension field is any one of the dimension fields in the data table to be processed, and the second dimension field is any one of the dimension fields in the data table different from the first dimension field. Construct a cross-table of the first cell set and the second cell set, and obtain the description features and correlation features of the cross-table. Determine the enumerated value types of the first cell set and the second cell set according to the type of the first dimension field and the type of the second dimension field.
3. The method for identifying the relationship of data table fields according to claim 2, wherein Determining the relationship between each dimension field based on the description information of the cell set corresponding to each dimension field, the relationship information between any two cell sets, the description features and correlation features of the cross-table constructed by any two cell sets, and the enumerated value type of the corresponding cell set determined according to the types of any two dimension fields, including: Inputting the description information of the first cell set, the description information of the second cell set, the relationship information between the first cell set and the second cell set, the description features of the cross-table, the correlation features of the cross-table, and the enumerated value type of the first cell set and the second cell set into a pre-trained first classification model to obtain a first probability value indicating whether there is an inclusion relationship between the first dimension field and the second dimension field; Inputting the description information of the first cell set, the description information of the second cell set, the relationship information between the first cell set and the second cell set, the description features of the cross-table, the correlation features of the cross-table, and the enumerated value type of the first cell set and the second cell set into a pre-trained second classification model to obtain a second probability value indicating whether combination is recommended under the non-inclusion relationship between the first dimension field and the second dimension field; Determining the relationship between the first dimension field and the second dimension field according to the first probability value and the second probability value; Wherein, the first classification model is trained based on the description information of the first sample cell set, the description information of the second sample cell set, the relationship information between the first sample cell set and the second sample cell set, the description features of the cross-table formed by the first sample cell set and the second sample cell set, the correlation features of the cross-table formed by the first sample cell set and the second sample cell set, the enumerated value type of the first sample cell set and the second sample cell set, and the label information indicating whether there is an inclusion relationship between the first sample cell set and the second sample cell set; The second classification model is trained based on the description information of the first sample cell set, the description information of the second sample cell set, the relationship information between the first sample cell set and the second sample cell set, the description features of the cross-table formed by the first sample cell set and the second sample cell set, the correlation features of the cross-table formed by the first sample cell set and the second sample cell set, the enumerated value type of the first sample cell set and the second sample cell set, and the label information indicating whether combination is recommended under the non-inclusion relationship between the first sample cell set and the second sample cell set.
4. The method for identifying the relationship of data table fields according to claim 3, wherein The determining the relationship information between the first dimension field and the second dimension field according to the first probability value and the second probability value includes: When the first probability value is greater than the second probability value, there is an inclusion relationship between the first dimension field and the second dimension field; When the first probability value is less than the second probability value, and the second probability value is greater than or equal to a preset first threshold, the relationship between the first dimension field and the second dimension field is non-inclusive and recommended for combination; When the first probability value is less than the second probability value, and the second probability value is less than the preset first threshold, the relationship between the first dimension field and the second dimension field is non-inclusive and not recommended for combination.
5. The method for identifying the relationship of data table fields according to claim 2, wherein The obtaining of the description information of the first cell set, the description information of the second cell set, and the relationship information between the first cell set and the second cell set includes: Obtaining the index value of the first cell set in the data table to be processed; Obtaining the maximum value of the data lengths in each cell of the first cell set; Obtaining the minimum value of the data lengths in each cell of the first cell set; Obtaining the average value of the data lengths in each cell of the first cell set; Obtaining the standard deviation of the data lengths in each cell of the first cell set; Obtaining the number of non-repeated data in the first cell set; Obtaining the number of cells in the first cell set; Obtaining the index value of the second cell set in the data table to be processed; Obtaining the maximum value of the data lengths in each cell of the second cell set; Obtaining the minimum value of the data lengths in each cell of the second cell set; Obtaining the average value of the data lengths in each cell of the second cell set; Obtaining the standard deviation of the data lengths in each cell of the second cell set; Obtaining the number of non-repeated data in the second cell set; Obtaining the product of the number of non-repeated data in the first cell set and the number of non-repeated data in the second cell set; Obtaining the information on whether the first cell set and the second cell set are in an inclusion relationship.
6. The method for identifying the relationship of data table fields according to claim 2, wherein, The constructing of the cross-table of the first cell set and the second cell set, and obtaining the description features and correlation features of the cross-table includes: Constructing the cross-table of the first cell set and the second cell set; Obtaining the description features of the cross-table, including: calculating the number of non-repeated data in the cross-table, calculating the number of non-empty cells in the cross-table; Obtaining the correlation features of the cross-table, including: calculating the P value of the chi-square test of the cross-table, calculating the degrees of freedom of the chi-square test of the cross-table, calculating whether the P value of the chi-square test of the cross-table is less than a preset second threshold, calculating the quotient of the number of cells in the cross-table whose absolute value of the correlation coefficient is greater than or equal to a third threshold and 2; calculating the percentage of the number of cells in the cross-table whose absolute value of the correlation coefficient is greater than or equal to a third threshold in the total number of cells, calculating the average value of the data in each cell of the cross-table, calculating the standard deviation of the average value of the data in each cell of the cross-table.
7. The method for identifying the relationship of data table fields according to claim 1, wherein Determining the types of each field in the data table to be processed, and determining the dimension fields and the set of cells corresponding to the dimension fields according to the types of the fields, includes: Determining the types of each field in the data table to be processed; For any field in the data table to be processed, when the type of the field meets the first condition, determining the field as a dimension field; wherein, the first condition is a condition for judging whether the types of each field in the data table to be processed belong to dimension fields; Obtaining the set of cells corresponding to the dimension field from the data table to be processed; wherein, the set of cells corresponding to the dimension field is the row or column corresponding to the dimension field in the data table to be processed.
8. The method for identifying the relationship between data table fields according to claim 7, wherein Determining the types of each field in the data table to be processed includes: Obtaining the data table to be processed; wherein, the data table includes fields and cells, and the cells include data; Determining the type of each cell in the data table according to the data included in the cell; Determining the type of the field according to the types of each cell corresponding to the field.
9. A data table field relationship recognition device, characterized in that, Includes: A first determination module, configured to determine the types of each field in the data table to be processed, and determine the dimension fields and the set of cells corresponding to the dimension fields according to the types of the fields; wherein, the dimension field is a field for describing the meaning represented by the data in the cell; An obtaining module, configured to obtain the relevant information of the set of cells corresponding to each dimension field, the relevant information of the constructed cross-table, and the information determined according to the type of the dimension field; A second determination module, configured to determine the relationship between each dimension field according to the relevant information of the set of cells corresponding to each dimension field, the relevant information of the constructed cross-table, and the information determined according to the type of the dimension field; The obtaining module is further configured to: Obtain the description information of the set of cells corresponding to each dimension field, the relationship information between any two sets of cells, the description features and correlation features of the cross-table constructed by any two sets of cells, and the enumerated value type of the set of cells corresponding to the dimension field determined according to the types of any two dimension fields; The second determination module is further configured to: Determine the relationship between each dimension field according to the description information of the set of cells corresponding to each dimension field, the relationship information between any two sets of cells, the description features and correlation features of the cross-table constructed by any two sets of cells, and the enumerated value type of the set of cells corresponding to the dimension field determined according to the types of any two dimension fields.
10. An electronic device, characterized in that, Includes: A processor, a memory, and a bus, wherein, The processor and the memory communicate with each other through the bus; The memory stores program instructions executable by the processor, and the processor can execute the steps of the data table field relationship recognition method according to any one of claims 1 to 8 by calling the program instructions.
11. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the steps of the data table field relationship recognition method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Data model construction method, device and apparatus and computer readable storage medium
CN110457288A
Method and device for processing table data, equipment, medium and product
CN113221519A