Data table field map generation method and device, electronic equipment and storage medium
By determining the dimensions and metric field types in the data table and using a pre-trained model to generate field graphs, the problem of inaccurate data display in massive data tables is solved, improving data table processing efficiency and user experience.
Patent Information
- Application Number
- CN202111223623.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-20
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2041-10-20
AI Technical Summary
Existing technologies cannot effectively display data in data tables when dealing with massive amounts of data, resulting in inaccurate and inefficient display of data table fields, which affects subsequent data table verification and field recommendations.
By determining the types of dimension and metric fields in the data table, a pre-trained model is used to obtain recommendation scores for the fields, perform classification, and generate field graphs, thereby improving the accuracy and efficiency of displaying data table fields.
It enables accurate display of data in massive data tables, improves data table processing efficiency and user experience, and lays the foundation for subsequent data table verification and field recommendation.
Smart Images

Figure CN114003666B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a data table field graph generation method and device, electronic equipment and a storage medium. BACKGROUND
[0002] In data standardization work, with new data tables constantly being connected to the database, and with the rapid development of big data technology, the accuracy of data and the quality of field relationship identification of the data table are crucial to the value that data can produce.
[0003] A data table is composed of fields in the table and data of each field corresponding to the cells. Among them, the fields in the data table are roughly divided into two categories: dimension fields and measurement fields. The dimension field refers to the "classification field", which is used to describe what the problem is. The types of dimension fields include date and time types, date types, time types, string types, etc. The measurement field is used to describe the number of fields, such as numerical types.
[0004] In the prior art, in order to better display the data in the data table, a column chart, a pie chart or a line chart is generally used. This conventional data display method is only suitable for small amounts of data, and is not suitable for massive amounts of data. SUMMARY
[0005] Based on the problems in the prior art, the present application provides a data table field graph generation method and device, electronic equipment and a storage medium, which solves the technical problem of not being able to better display the data in the data table in the case of massive amounts of data, improves the accuracy and efficiency of data table field display, provides a basis for subsequent data table proofreading and field recommendation, and has the advantages of improving data table processing efficiency and improving user experience.
[0006] In a first aspect, the present application provides a data table field graph generation method, comprising:
[0007] determining the types of each field in the data table to be processed, and determining the dimension fields and the measurement fields in the data table to be processed according to the types of the fields; wherein the dimension fields are used to describe the meaning represented by the data in the corresponding cells, and the measurement fields are used to describe the quantity represented by the data in the corresponding cells;
[0008] determining the recommended score of the dimension field according to the type of the dimension field, the characteristic information of the dimension field and the characteristic information of the data table to be processed;
[0009] classifying the dimension fields in the data table to be processed according to the type of the dimension field and the recommended score of the dimension field, to obtain the category of the dimension field;
[0010] determine the category of the metric field, and determine a recommendation score of the metric field according to the feature information of the to-be-processed data table and the feature information of the metric field;
[0011] generate a field graph of the to-be-processed data table according to the category of each dimension field in the to-be-processed data table, the recommendation score of each dimension field, the category of each metric field, and the recommendation score of each metric field.
[0012] Further, the data table field graph generation method provided by the present application comprises the following steps:
[0013] obtaining feature information of the to-be-processed data table;
[0014] obtaining first feature information of a first dimension field in the to-be-processed data table according to the data distribution in a cell set corresponding to the first dimension field and the distribution of the first dimension field in the to-be-processed data table; wherein the first dimension field is any dimension field in the to-be-processed data table;
[0015] obtaining second feature information of the first dimension field according to the type of each dimension field in the to-be-processed data table and the relationship between the first dimension field and other dimension fields in the to-be-processed data table;
[0016] inputting the feature information of the to-be-processed data table, the first feature information of the first dimension field, and the second feature information of the first dimension field into a pre-trained dimension field recommendation model to obtain a recommendation score of the first dimension field;
[0017] wherein the dimension field recommendation model is trained based on the feature information of a sample data table, the first feature information of a second dimension field in the sample data table, the second feature information of the second dimension field, and a recommendation score label of the second dimension field; wherein the second dimension field is any one or more dimension fields in the sample data table.
[0018] Further, the feature information of the to-be-processed data table comprises a description feature of the to-be-processed data table and a classification feature of the to-be-processed data table;
[0019] Correspondingly, the step of obtaining the feature information of the to-be-processed data table comprises the following steps:
[0020] obtaining a description feature of the to-be-processed data table;
[0021] obtaining a classification feature of the to-be-processed data table.
[0022] Further, according to the data table field graph generation method provided by the present application, the first characteristic information of the first dimension field includes description characteristic information and position characteristic information of the first dimension field; accordingly, the first characteristic information of the first dimension field in the to-be-processed data table is obtained according to the data distribution in the cell set corresponding to the first dimension field in the to-be-processed data table and the distribution of the first dimension field in the to-be-processed data table, including:
[0023] The description characteristic information of the first dimension field is obtained.
[0024] The position characteristic information of the first dimension field is obtained.
[0025] Further, according to the data table field graph generation method provided by the present application, the second characteristic information of the first dimension field is obtained according to the type of each dimension field in the to-be-processed data table and the relationship between the first dimension field and other dimension fields in the to-be-processed data table, including:
[0026] The number of dimension fields with the same type as the first dimension field in the to-be-processed data table is obtained according to the type of each dimension field in the to-be-processed data table.
[0027] The number of dimension fields in the to-be-processed data table that have a parent-child relationship with the first dimension field is obtained according to the relationship between the first dimension field and other dimension fields in the to-be-processed data table.
[0028] The number of dimension fields in the to-be-processed data table that have a parent-child relationship with the first dimension field is obtained according to the relationship between the first dimension field and other dimension fields in the to-be-processed data table.
[0029] The number of dimension fields in the to-be-processed data table that have the same type as the first dimension field and have a parent-child relationship with the first dimension field is obtained according to the relationship between the first dimension field and other dimension fields in the to-be-processed data table.
[0030] The number of dimension fields in the to-be-processed data table that have the same type as the first dimension field and have a parent-child relationship with the first dimension field is obtained according to the relationship between the first dimension field and other dimension fields in the to-be-processed data table.
[0031] The enumeration value of the field type is obtained.
[0032] The second characteristic information of the first dimension field in the to-be-processed data table is obtained according to the obtained data.
[0033] Further, the data table field graph generation method provided by the present application comprises the following steps: classifying the dimension fields in the data table to be processed according to the type of the dimension field and the recommended score of the dimension field, and obtaining the category of the dimension field, wherein the classification of the dimension field is obtained according to the type of the dimension field and the recommended score of the dimension field.
[0034] In the case that the type of the first dimension field satisfies the type-category mapping relationship in the type-category mapping relationship set, the category corresponding to the first dimension field is determined; wherein the type-category mapping relationship describes the type and the category having a unique corresponding relationship; the first dimension field is any one of the dimension fields in the data table to be processed;
[0035] In the case that the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, the category corresponding to the first dimension field is determined according to the correlation between the first dimension field and the dimension field whose category has been determined in the data table to be processed and the recommended score of the first dimension field.
[0036] Further, the data table field graph generation method provided by the present application comprises the following steps: classifying the dimension fields in the data table to be processed according to the type of the dimension field and the recommended score of the dimension field, and obtaining the category of the dimension field, wherein the classification of the dimension field is obtained according to the type of the dimension field and the recommended score of the dimension field.
[0037] Correspondingly, in the case that the type of the first dimension field satisfies the type-category mapping relationship in the type-category mapping relationship set, the category corresponding to the first dimension field is determined, comprising the following steps:
[0038] According to the type of the first dimension field, the corresponding type-category mapping relationship in the type-category mapping relationship set is determined.
[0039] According to the determined type-category mapping relationship, the category of the first dimension field is determined.
[0040] Further, the data table field graph generation method provided by the present application comprises the following steps: classifying the dimension fields in the data table to be processed according to the type of the dimension field and the recommended score of the dimension field, and obtaining the category of the dimension field, wherein the classification of the dimension field is obtained according to the type of the dimension field and the recommended score of the dimension field.
[0041] It is judged whether the type of the first dimension field is a noun type.
[0042] In a case where the type of the first dimension field is a noun type, a correlation between the first dimension field and each dimension field determined as a person category is calculated, and a correlation between the first dimension field and each dimension field determined as a place category is calculated;
[0043] A category of the first dimension field is determined according to a comparison result of the first correlation and the second correlation, wherein the first correlation is a maximum value of the correlations between the first dimension field and each dimension field determined as a person category, and the second correlation is a maximum value of the correlations between the first dimension field and each dimension field determined as a place category.
[0044] Further, the data table field graph generation method provided by the present application, the category of the first dimension field is determined according to the comparison result of the first correlation and the second correlation, comprising:
[0045] In a case where the first correlation is greater than a first threshold value and the first correlation is greater than the second correlation, the category of the first dimension field is determined as a person category;
[0046] In a case where the second correlation is greater than the first threshold value and the second correlation is greater than the first correlation, the category of the first dimension field is determined as a place category.
[0047] Further, the data table field graph generation method provided by the present application, in a case where the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, the category corresponding to the first dimension field is determined according to the correlation between the first dimension field and the dimension field whose category has been determined in the data table to be processed and the recommendation score of the first dimension field, comprising:
[0048] In a case where the type of the first dimension field is not a noun type, a correlation between the first dimension field and each dimension field determined as a person category is calculated, a correlation between the first dimension field and each dimension field determined as a place category is calculated, and a correlation between the first dimension field and each dimension field determined as a time category is calculated;
[0049] In a case where the first correlation is greater than a second threshold value, the category of the first dimension field is determined as a person category;
[0050] In a case where the second correlation is greater than a third threshold value, the category of the first dimension field is determined as a place category;
[0051] In a case where the third correlation is greater than a fourth threshold value, the category of the first dimension field is determined as a time category.
[0052] in a case where the category of the first dimension field is not identified as the person category, the place category, or the time category, identifying the category of the first dimension field as an event category;
[0053] wherein the first correlation degree is a maximum value of correlation degrees between the first dimension field and each dimension field that has been determined as a person category; the second correlation degree is a maximum value of correlation degrees between the first dimension field and each dimension field that has been determined as a place category; and the third correlation degree is a maximum value of correlation degrees between the first dimension field and each dimension field that has been determined as a time category.
[0054] Further, according to the data table field graph generation method provided by the present application, in a case where the type of the first dimension field does not satisfy a type category mapping relationship in the type category mapping relationship set, the category corresponding to the first dimension field is determined according to a correlation degree between the first dimension field and a dimension field whose category has been determined in the data table to be processed and a recommendation score of the first dimension field, and the method further comprises:
[0055] obtaining a recommendation score of the first dimension field of the event category;
[0056] in a case where the recommendation score is less than or equal to a fifth threshold value, modifying the category of the first dimension field from the event category to an event category that is not recommended.
[0057] Further, according to the data table field graph generation method provided by the present application, the category of the dimension field is determined, and the recommendation score of the dimension field is determined according to the feature information of the data table to be processed and the feature information of the dimension field, and the method comprises:
[0058] determining the category of the first dimension field as a dimension category; wherein the first dimension field is any one of the dimension fields in the data table to be processed;
[0059] obtaining the feature information of the data table to be processed; wherein the feature information of the data table to be processed comprises a description feature of the data table to be processed and a classification feature of the data table to be processed;
[0060] obtaining first feature information of the first dimension field according to a data distribution in a cell set corresponding to the first dimension field in the data table to be processed and a distribution of the first dimension field in the data table to be processed;
[0061] obtaining second feature information of the first dimension field in the data table to be processed according to a data statistical situation in the cell set corresponding to the first dimension field in the data table to be processed;
[0062] input the feature information of the to-be-processed data table, the first feature information of the first metric field, and the second feature information of the first metric field into a pre-trained metric field recommendation model to obtain a recommendation score of the first metric field;
[0063] The metric field recommendation model is trained based on feature information of sample data tables, first feature information of a second metric field in the sample data tables, second feature information of the second metric field, and a recommendation score label of the second metric field, wherein the second metric field is any one or more metric fields in the sample data tables.
[0064] Further, according to the data table field graph generation method provided in the present application, the first feature information of the first metric field in the to-be-processed data table is obtained according to the data distribution in the cell set corresponding to the first metric field in the to-be-processed data table and the distribution of the first metric field in the to-be-processed data table, and at least includes any one of the following:
[0065] An index value of the first metric field is obtained, wherein the index value is used to describe the position of the first metric field in the to-be-processed data table;
[0066] The number of non-repeated data in the cell set corresponding to the first metric field is obtained.
[0067] The number of cells in the cell set corresponding to the first metric field is obtained.
[0068] The number of non-empty cells in the cell set corresponding to the first metric field is obtained.
[0069] In the to-be-processed data table, the index value of the first metric field is compared with the index values of other metric fields to obtain the number of other metric fields with an index value smaller than that of the first metric field and the number of other metric fields with an index value greater than that of the first metric field.
[0070] In the to-be-processed data table, the index value of the first metric field is compared with the index values of dimension fields to obtain the number of dimension fields with an index value smaller than that of the first metric field and the number of dimension fields with an index value greater than that of the first metric field.
[0071] The first feature information of the first metric field in the to-be-processed data table is obtained according to the obtained data.
[0072] Further, according to the data table field graph generation method provided in the present application, the second feature information of the first metric field in the to-be-processed data table is obtained according to the data statistics in the cell set corresponding to the first metric field in the to-be-processed data table, and at least includes any one of the following:
[0073] obtaining an average value of the numbers in the cell set corresponding to the first metric field;
[0074] obtaining a median of the numbers in the cell set corresponding to the first metric field;
[0075] obtaining a standard deviation of the numbers in the cell set corresponding to the first metric field;
[0076] obtaining a minimum value of the numbers in the cell set corresponding to the first metric field;
[0077] obtaining a maximum value of the numbers in the cell set corresponding to the first metric field;
[0078] obtaining a three-quarter quantile value of the numbers in the cell set corresponding to the first metric field;
[0079] obtaining a one-quarter quantile value of the numbers in the cell set corresponding to the first metric field;
[0080] obtaining a difference between the one-quarter quantile value and the three-quarter quantile value of the numbers in the cell set corresponding to the first metric field.
[0081] In a second aspect, the present application further provides a data table field map generation device, comprising:
[0082] a first determining module, configured to determine types of each field in a to-be-processed data table, determine dimension fields and metric fields in the to-be-processed data table according to the types of the fields; wherein the dimension fields are used to describe meanings of data represented in corresponding cells, and the metric fields are used to describe quantities of data represented in corresponding cells;
[0083] a second determining module, configured to determine a recommendation score of a dimension field according to a type of the dimension field, characteristic information of the dimension field, and characteristic information of the to-be-processed data table;
[0084] a classifying module, configured to classify the dimension fields in the to-be-processed data table according to the types of the dimension fields and the recommendation scores of the dimension fields, and obtain classes of the dimension fields;
[0085] a third determining module, configured to determine classes of the metric fields, and determine recommendation scores of the metric fields according to the characteristic information of the to-be-processed data table and characteristic information of the metric fields;
[0086] a generating module, configured to generate a field map of the to-be-processed data table according to the classes of each dimension field in the to-be-processed data table, the recommendation scores of each dimension field, the classes of each metric field, and the recommendation scores of each metric field.
[0087] In a third aspect, the present application also provides an electronic device, comprising: a processor, a memory and a bus, wherein,
[0088] The processor and the memory communicate with each other through the bus;
[0089] The memory stores program instructions executable by the processor, and the processor invoking the program instructions can execute the steps of the data table field graph generation method according to any one of the above.
[0090] In a fourth aspect, the present application also provides a computer readable storage medium, which stores computer instructions, and the computer instructions make the computer execute the steps of the data table field graph generation method according to any one of the above.
[0091] The present application provides a data table field graph generation method, device, electronic device and storage medium, determines the dimension field and the measurement field according to the type of each field; determines the recommended score of the dimension field according to the type of the dimension field, the characteristic information and the characteristic information of the to-be-processed data table, classifies the dimension field according to the type and the recommended score of the dimension field to obtain the category of the dimension field, determines the category of the measurement field, and determines the recommended score of the measurement field according to the characteristic information of the to-be-processed data table and the characteristic information of the measurement field; finally, generates the field graph of the to-be-processed data table according to the category and the recommended score of each dimension field and the category and the recommended score of each measurement field. The field graph generation method provided by the present application can accurately display the data in the data table with massive data, at the same time, provides a basis for subsequent data table proofreading and field recommendation, improves the efficiency of data table processing, and improves the user experience. BRIEF DESCRIPTION OF DRAWINGS
[0092] In order to more clearly illustrate the technical solutions in the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0093] Figure 1 is a flowchart of a data table field graph generation method provided by the present application;
[0094] Figure 2 is an example graph of a generated field graph provided by the present application;
[0095] Figure 3 is a flowchart of a dimension field obtaining recommended score provided by the present application;
[0096] Figure 4 is a flowchart of a process for obtaining a recommended score of a metric field provided by the present application;
[0097] Figure 5 is a flowchart of a whole process of a data table field map generation method provided by the present application;
[0098] Figure 6 is a structural schematic diagram of a data table field map generation device provided by the present application;
[0099] Figure 7 is a structural schematic diagram of an electronic device provided by the present application. DETAILED DESCRIPTION
[0100] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described below in conjunction with the accompanying drawings in the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0101] Figure 1 is a flowchart of a data table field map generation method provided by the present application, which comprises the following steps: Figure 1
[0102] Step 101: Determine the types of each field in the data table to be processed, and determine the dimension field and the metric field in the data table to be processed according to the types of the fields; wherein the dimension field is used to describe the meaning represented by the data in the corresponding cell, and the metric field is used to describe the quantity represented by the data in the corresponding cell.
[0103] In the embodiment, due to the different types of each cell in the to-be-processed data table, the types of each field of the to-be-processed data table are also different, and each field has different analysis values, such as a salary table and a transcript. In the embodiment, the types of each field in the to-be-processed data table need to be determined, and the types of the fields can be a date-time type, a date type, a time type, a string type, a numerical value type, and the like. Among them, according to the types of each field, it can be determined which fields belong to dimension fields and which fields belong to measure fields. The dimension field is used to describe the meaning represented by the data in the corresponding cell, and essentially belongs to a "classification field". For example, the date-time type, the date type, the time type, and the string type all belong to the dimension field, and the string type specifically includes a person name type, a place name type, a gerund type, an English type, a text type, a numerical type, other types, and a mixed type. The measure field is used to describe the quantity represented by the data in the corresponding cell, and belongs to a quantitative field. For example, the field of the numerical value type belongs to the measure field. It should be noted that the types of the dimension field and the measure field can be set according to actual needs, and are not limited here.
[0104] Step 102: determining a recommended score of the dimension field according to the type of the dimension field, the characteristic information of the dimension field, and the characteristic information of the to-be-processed data table.
[0105] In the embodiment, the characteristic information of the dimension field includes the data distribution of each dimension field, the distribution of each dimension field in the to-be-processed data table, and the characteristic information of the number of dimension fields or measure fields adjacent to the left and right of the dimension field. For example, in Table 1 below, according to the definition of the dimension field and the measure field, it can be determined that, except for the field with the field name "age", which belongs to the measure field, the rest all belong to the dimension field. For example, the cell corresponding to the field with the field name "vehicle status" has specific data information "lock car, lock car, invoice delay, invoice delay", according to which the maximum cell data length is 4, the minimum cell data length is 2, the position information is the second column, and the specific characteristic information can be seen from the detailed description of the following embodiment. The characteristic information of the to-be-processed data table includes descriptive characteristic information of the to-be-processed data table and classification characteristic information of the data table, which is used to describe the basic information of the to-be-processed data table, such as the number of fields with the field type of the person name type in the to-be-processed data table, and the specific content can be seen from the following embodiment, which is not described here. It should be noted that, according to the type of the dimension field, the characteristic information of the dimension field, and the characteristic information of the to-be-processed data table, the recommended score of the dimension field can be determined, and the recommended field or the application to the data analysis in the pivot table can be obtained according to the generated recommended score.
[0106] According to the determined type of the dimension field, the characteristic information of the dimension field and the characteristic information of the to-be-processed data table, a recommendation score of the dimension field is determined. The recommendation score refers to a score value obtained by judging whether the dimension field can be used in data analysis. The greater the score value is, the more valuable the corresponding dimension field is for analysis, and the more recommended it is. It should be noted that in this embodiment, the greater the score value is, the more recommended it is for analysis. In other embodiments, it can be other setting conditions, such as the greater the score value is, the less recommended it is, which is not limited here.
[0107] Table 1
[0108]
[0109] Step 103: According to the type of the dimension field and the recommendation score of the dimension field, the dimension field in the to-be-processed data table is classified to obtain the category of the dimension field.
[0110] In this embodiment, the dimension field is first classified according to the type of the dimension field. This embodiment can divide the dimension field into four categories, which are time category, place category, person category and event category. The dimension field of date type, time type and date-time type belongs to the time category, the dimension field of place name type belongs to the place category, the dimension field of person name type, organization and group type, mobile phone number type, telephone number type, text type and ID number type belongs to the person category, and the dimension field of verb type belongs to the event category. It should be noted that for the dimension field that cannot be accurately classified into the above four categories according to the type of the dimension field, the dimension field can also be determined and analyzed according to the relevance of the dimension field to the dimension field in each category. Details are described in the following embodiments, which are not described in detail here.
[0111] In this embodiment, the category of the dimension field can also be further determined according to the type of the dimension field, the recommendation score and the preset judgment condition. For example, according to the type of the dimension field, it is determined that the category of the dimension field is the event category, and the recommendation score of the dimension field is less than or equal to 0.1. The dimension field is determined as the "non-recommended event" category. It should be noted that the preset judgment condition can be other conditions, which are not limited here.
[0112] Step 104: The category of the measurement field is determined, and the recommendation score of the measurement field is determined according to the characteristic information of the to-be-processed data table and the characteristic information of the measurement field.
[0113] In the embodiment, the metric field is generally a field for describing the number of corresponding cells, such as the cells corresponding to the field names "English", "Chinese", "Chemistry" and the like in the transcript, wherein the data in each cell is a numerical value, that is, the numerical value type field belongs to the metric field. The metric field is set as a category, and is determined as the metric field category. It should be noted that in the embodiment, after the category of the metric field is determined, the recommendation score of the metric field is determined according to the feature information of the to-be-processed data table and the feature information of the metric field. The higher the recommendation score, the higher the order of the recommended analysis, and the more valuable the analysis, such as the recommendation score of the metric field A is 0.12, and the recommendation score of the metric field B is 0.06. Since the recommendation score of the metric field A is higher than that of the metric field B, the order of the recommended analysis of the metric field A is higher than that of the metric field B, and the metric field A is preferentially recommended. The feature information of the to-be-processed data table and the feature information of the metric field can be seen from the following embodiments, and will not be described in detail here.
[0114] It should be noted that the obtained feature information of the to-be-processed data table and the feature information of the metric field are summarized and input into the corresponding classification model for scoring to obtain the score value corresponding to each metric field, and the obtained score value is sorted from large to small to obtain a recommendation score list of the metric field. The score value is between 0 and 1, and the larger the recommendation score, the more valuable the corresponding metric field for analysis.
[0115] Step 105: generating a field graph of the to-be-processed data table according to the categories of the dimension fields in the to-be-processed data table, the recommendation scores of the dimension fields, the categories of the metric fields and the recommendation scores of the metric fields.
[0116] In the embodiment, the field graph of the to-be-processed data table is generated according to the five categories of the dimension fields and the metric field categories obtained in the above steps, and the recommendation scores of the dimension fields and the recommendation scores of the metric fields. It should be noted that the field graph contains information as shown in Table 2 below. The field graph contains category information of the fields, names of the fields, recommendation scores and inclusion relationship information, and the display mode of the generated field graph can be in the form of bubbles, such as Figure 2 As shown, each field name is displayed in the form of a bubble, and the specific display mode can be set according to actual needs, which will not be limited here.
[0117] Table 2
[0118]
[0119]
[0120] According to the application, a data table field map generation method is provided, dimension fields and measurement fields are determined according to the types of the respective fields; the recommendation scores of the dimension fields are determined according to the types of the dimension fields, the characteristic information of the dimension fields and the characteristic information of the data table to be processed, the dimension fields are classified according to the types of the dimension fields and the recommendation scores to obtain the categories of the dimension fields, the categories of the measurement fields are determined, and the recommendation scores of the measurement fields are determined according to the characteristic information of the data table to be processed and the characteristic information of the measurement fields; finally, the field map of the data table to be processed is generated according to the categories and the recommendation scores of the respective dimension fields and the categories and the recommendation scores of the respective measurement fields. The field map generation method provided by the application can accurately display the data in the data table of the massive data, and the user can accurately obtain the recommended analysis and the category of each field according to the data information displayed by the generated field map, quickly find the field to be recommended and analyzed, improve the efficiency of data table processing and enhance the user experience.
[0121] Based on any of the above embodiments, in this embodiment, the recommendation score of the dimension field is determined according to the type of the dimension field, the characteristic information of the dimension field and the characteristic information of the data table to be processed, including:
[0122] The characteristic information of the data table to be processed is obtained; the first characteristic information of the first dimension field in the data table to be processed is obtained according to the data distribution in the cell set corresponding to the first dimension field in the data table to be processed and the distribution of the first dimension field in the data table to be processed; the second characteristic information of the first dimension field is obtained according to the types of the respective dimension fields in the data table to be processed and the relationship between the first dimension field and the other dimension fields in the data table to be processed; the characteristic information of the data table to be processed, the first characteristic information of the first dimension field and the second characteristic information of the first dimension field are input into the pre-trained dimension field recommendation model to obtain the recommendation score of the first dimension field; wherein the dimension field recommendation model is trained based on the characteristic information of the sample data table, the first characteristic information of the second dimension field in the sample data table, the second characteristic information of the second dimension field and the recommendation score label of the second dimension field; wherein the second dimension field is any one or more dimension fields in the sample data table.
[0123] In the embodiment, in order to obtain the recommendation score of the first dimension field, the feature information of the obtained to-be-processed data table, the first feature information and the second feature information of the first dimension field are input into the dimension field recommendation model obtained by pre-training, to obtain the recommendation score of the first dimension field. The dimension field recommendation model is obtained by pre-training the training sample by using the random forest algorithm. The random forest refers to a classifier for training and predicting the training sample by using multiple decision trees. The classifier is trained by using the random forest algorithm on the feature information of the pre-obtained sample data table, the first feature information of the second dimension field in the sample data table, the second feature information of the second dimension field, and the recommendation score label of the second dimension field, to obtain the dimension field recommendation model. The specific training manner is not described in detail herein.
[0124] In the embodiment, the first feature information of the first dimension field refers to the descriptive feature information of the first dimension field, and whether the first dimension field has analysis significance can be determined according to the obtained first feature information. For example, in the remark information, the data length is long or short, some are empty, the average value is small, the standard deviation is large, the minimum value is zero, and the maximum value is large. The dimension field of this type often does not have analysis value. In the embodiment, the first dimension field is any dimension field in the to-be-processed data table. For example, in Table 1, if the field name of the first dimension field is “vehicle information”, according to the data of the cell set corresponding to the first dimension field, the distribution of the data is obtained, that is, the maximum value in the cell is 7, the average value is also 7, and so on. Moreover, the first dimension field is located in the first column of the to-be-processed data table, that is, the index value is 1, the number of dimension fields with an index value greater than 1 is 6, and the number of dimension fields is 1. According to the obtained specific situation information, the first feature information of the first dimension field is obtained. The content included in the first feature information is described in the following embodiment, which is not described in detail herein.
[0125] In the embodiment, the second feature information of the first dimension field also needs to be obtained. The second feature information is information indicating the relationship between the first dimension field and other fields in the to-be-processed data table. According to the type of each dimension field in the to-be-processed data table and the relationship between the first dimension field and other dimension fields, the second feature information of the first dimension field is obtained. The relationship between the first dimension field and other dimension fields can be a parent-child relationship. The parent-child relationship means that one father can have multiple sons, and one son cannot have multiple fathers. It should be noted that the specific content included in the second feature information is described in the following embodiment.
[0126] In this embodiment, the feature information of the data table to be processed, the first feature information and the second feature information of the first dimension field are input into a pre-trained dimension field recommendation model to obtain the recommendation score of the first dimension field. For example, if the feature information obtained in Table 1 above, the first feature information and the second feature information of the first dimension field are input into the dimension field recommendation model, the recommendation scores shown in Table 3 below are obtained. If the first dimension field is a field named "Vehicle Status", the obtained recommendation analysis value is 0.95, which is higher than the recommendation analysis values of other fields. The recommendation analysis order is 1, and it is recommended and analyzed first. It should be noted that the value of the recommendation score is between 0 and 1.
[0127] Table 3
[0128]
[0129]
[0130] According to the data table field graph generation method provided by the present invention, the feature information of the data table to be processed, the first feature information and the second feature information of the first dimension field are input into a pre-trained dimension field recommendation model to obtain the recommendation score of the first dimension field. This allows users to quickly find the field to be analyzed from multiple dimension fields, improves the processing speed and accuracy of dimension field recommendation analysis, and reduces the cost of manually viewing data.
[0131] Based on any of the above embodiments, such as Figure 3 As shown, in this embodiment, the feature information of the data table to be processed includes the descriptive features and classification features of the data table to be processed; correspondingly, obtaining the feature information of the data table to be processed includes:
[0132] Obtain descriptive features of the data table to be processed;
[0133] Obtain the classification features of the data table to be processed.
[0134] In this embodiment, descriptive features and classification features of the data table to be processed are obtained. Descriptive features primarily describe the characteristics of the data table and can be extracted using the descriptive statistics module within the table. Classification features describe information about each field type in the data table, specifically including: enumerated values for each field type, the number of fields of the person / name type, the number of fields of the place / name type, the number of fields of the numeric type, the number of fields of the date type, the number of fields of the time type, and the number of fields of the gerund type. It should be noted that enumerated values refer to defining an ordered set by predefined identifiers listing all values; the order of these values is consistent with the order of the identifiers in the enumeration type specification. Assuming the form of the enumerated values is <identifier1> = <type1>, such as...<n1>= <time type>, <n2>= <date type>, <n3>< Enumerated value type >, assuming that the types of the fields in the to-be-processed data table in the embodiment are time type, date type, noun type and numerical value type, and the enumerated values of the to-be-processed data table are < N1, N2, N5, N3 > = < time type, date type, noun type, numerical value type >.
[0135] It should be noted that the description features of the to-be-processed data table include at least any one of the following: the number of non-empty cell sets in the to-be-processed data table; the maximum value of the number of non-empty cells included in each cell set in the to-be-processed data table; the minimum value of the number of non-repeated data included in each cell set in the to-be-processed data table; the maximum value of the number of non-repeated data included in each cell set in the to-be-processed data table; the average value of the number of non-repeated data included in each cell set in the to-be-processed data table; and the standard deviation of the number of non-repeated data included in each cell set in the to-be-processed data table. The classification features of the to-be-processed data table include at least any one of the following: the enumerated values of the types of the fields in the to-be-processed data table; the number of fields of the person name type in the to-be-processed data table; the number of fields of the place name type in the to-be-processed data table; the number of fields of the numerical value type in the to-be-processed data table; the number of fields of the date type in the to-be-processed data table; the number of fields of the time type in the to-be-processed data table; and the number of fields of the gerund type in the to-be-processed data table.
[0136] In the embodiment, as the to-be-processed data table is shown in Table 1, the feature information of the to-be-processed data table is obtained according to the basic information of each field, and is shown in Table 4. It is obtained that the number of non-empty cell sets of the to-be-processed data table is 8, that is, the cell sets corresponding to the field names "whole vehicle information, vehicle state, vehicle series, vehicle type, vehicle body color, interior color, engine number and storage age" are all non-empty cell sets. The to-be-processed data table shown in Table 1 is analyzed in sequence, and the feature information shown in Table 4 is obtained.
[0137] Table 4
[0138]
[0139] According to the data table field graph generation method provided by the embodiment, the description features and the classification features of the to-be-processed data table are obtained, the obtained information is used for subsequent determination and acquisition of the dimension field recommendation score, the speed of the dimension field data recommendation processing is improved, and data support is provided for subsequent generation of the field graph.
[0140] Based on any of the above embodiments, in the present embodiment, the first feature information of the first dimension field includes description feature information and position feature information of the first dimension field; accordingly, the first feature information of the first dimension field in the to-be-processed data table is obtained according to the data distribution in the cell set corresponding to the first dimension field in the to-be-processed data table and the distribution of the first dimension field in the to-be-processed data table, including:
[0141] obtaining the description feature information of the first dimension field;
[0142] obtaining the position feature information of the first dimension field.
[0143] In the present embodiment, the description feature information of the first dimension field includes at least any of the following: an index value of the first dimension field; wherein the index value is used to describe the position of the first dimension field in the to-be-processed data table; the number of non-repeated data in the cell set corresponding to the first dimension field; the number of cells in the cell set corresponding to the first dimension field; the average length of data contained in each cell in the cell set corresponding to the first dimension field; the minimum length of data contained in each cell in the cell set corresponding to the first dimension field; the maximum length of data contained in each cell in the cell set corresponding to the first dimension field; the standard deviation of the length of data contained in each cell in the cell set corresponding to the first dimension field; the average number of occurrences of non-repeated data in the cell set corresponding to the first dimension field; the minimum number of occurrences of non-repeated data in the cell set corresponding to the first dimension field; the maximum number of occurrences of non-repeated data in the cell set corresponding to the first dimension field; the standard deviation of the number of occurrences of non-repeated data in the cell set corresponding to the first dimension field.
[0144] wherein the position feature information of the first dimension field includes at least any of the following: in the to-be-processed data table, comparing the index value of the first dimension field with the index value of the measure field to obtain the number of measure fields with index values less than the index value of the first dimension field and the number of measure fields with index values greater than the index value of the first dimension field; in the to-be-processed data table, comparing the index value of the first dimension field with the index value of other dimension fields to obtain the number of other dimension fields with index values less than the index value of the first dimension field and the number of other dimension fields with index values greater than the index value of the first dimension field.
[0145] In the embodiment, the first characteristic information of the first dimension field is determined according to the distribution of the first dimension field in the data table to be processed and the data distribution in the cell set corresponding to the first dimension field. The first characteristic information includes the index value of the first dimension field, the number of non-repeated data, the number of cells, the length value of the data contained in each cell, and the like. It should be noted that, as shown in Table 2, the number of occurrences of non-repeated data (items) in the first dimension field is obtained. If the average value of the number of occurrences is 1, it indicates that the analysis value of the dimension field is not great. If the number of non-repeated data in the cell set corresponding to the first dimension field and the difference between the maximum and minimum values of the number of occurrences are both great, it indicates that the dimension field has a great analysis value. Figure 3
[0146] It should be noted that, in data analysis, the data analysis conclusion also has a certain relationship with the position of each field. For example, in general, the data on the right side of the data table often needs to be analyzed. In a relational database, the data to be analyzed is determined by the index value. The index is a separate, physical storage structure for sorting the values of one or more columns in a database table. It is a collection of one or more column values in a table and a list of logical pointers pointing to the data pages in the table that physically identify these values. The role of the index is equivalent to the table of contents of a book, and the required content can be quickly found according to the page number in the table of contents. The index value refers to the position code corresponding to each dimension field, which can be used to determine the specific position of the dimension field. For example, the index values of the fields in the data table shown in Table 1 are set as shown in Table 5 below. When the index value of the cell set corresponding to the first dimension field is 2, the first dimension field is determined to be the second column in the data table shown in Table 1 by the index value.
[0147] Table 5
[0148] Field index value Field name Field type 1 Whole vehicle information eng 2 Vehicle status vn 3 Vehicle series eng 4 Vehicle model eng 5 Vehicle body color n 6 Interior color n 7 Engine number eng 8 Stock life Number
[0149] In the embodiment, as shown in Table 1, after the position of the first dimension field is determined according to the index value 2 of the first dimension field, the number of non-repeated data in the cell set corresponding to the first dimension field is 0, the number of cells is 4, the average length of the data contained in each cell is 3, the minimum length of the data contained in each cell is 2, the maximum length of the data contained in each cell is 4, the standard deviation of the length of the data contained in each cell is 1, the average value of the number of occurrences of non-repeated data is 0, the minimum value of the number of occurrences of non-repeated data is 0, the maximum value of the number of occurrences of non-repeated data is 0, and the standard deviation of the number of occurrences of non-repeated data is 0. It should be noted that the average value of the length of the data contained in each cell in the cell set conforms to the normal distribution. The average value is particularly small or particularly large, which does not have an analysis value. The value within a certain range has an analysis value.
[0150] After obtaining the above information, in this embodiment, it is also needed to compare the index value of the first dimension field with the index value of the metric field in the to-be-processed data table, and obtain the number of metric fields with index values less than the index value of the first dimension field and the number of metric fields with index values greater than the index value of the first dimension field. For example, if the index value of the first dimension field is 2, by comparing and analyzing in the to-be-processed data table, it is obtained that the index value of the metric field with the field name "library age" is 8 and greater than the index value of the first dimension field. In the data table shown in Table 1, the number of metric fields with index values greater than the index value of the first dimension field is 1, and the number of metric fields with index values less than the index value of the first dimension field is 0.
[0151] In the to-be-processed data table, it is also needed to compare the index value of the first dimension field with the index values of other dimension fields, and obtain the number of other dimension fields with index values less than the index value of the first dimension field and the number of other dimension fields with index values greater than the index value of the first dimension field. For example, if the index value of the first dimension field is 2, by analyzing the to-be-processed data table shown in Table 1, it can be obtained that the number of dimension fields with index values less than 2 is 1, and the number of dimension fields with index values greater than 2 is 5. Then, according to the above obtained data, the first characteristic information of the first dimension field is determined. It should be noted that the first dimension field is any dimension field in the to-be-processed data table, which is not limited here.
[0152] According to the data table field graph generation method provided by the present application, the distribution of the first dimension field in the to-be-processed data table and the distribution of the corresponding cell data are determined according to the index value of the first dimension field, and then the first characteristic information of the first dimension field is determined, which provides data support for accurately determining the recommended score of the first dimension field subsequently.
[0153] Based on any of the above embodiments, in this embodiment, the second characteristic information of the first dimension field is obtained according to the types of the dimension fields in the to-be-processed data table and the relationship between the first dimension field and other dimension fields in the to-be-processed data table, including:
[0154] According to the type of each dimension field in the to-be-processed data table, the number of dimension fields in the to-be-processed data table having the same type as the first dimension field is obtained; according to the relationship between the first dimension field and other dimension fields in the to-be-processed data table, the number of dimension fields in the to-be-processed data table having a father-son relationship with the first dimension field is obtained; according to the relationship between the first dimension field and other dimension fields in the to-be-processed data table, the number of dimension fields in the to-be-processed data table having the same type as the first dimension field and having a father-son relationship with the first dimension field is obtained; according to the relationship between the first dimension field and other dimension fields in the to-be-processed data table, the number of dimension fields in the to-be-processed data table having the same type as the first dimension field and having a father-son relationship with the first dimension field is obtained; the enumeration value of the field type is obtained; and the second characteristic information of the first dimension field in the to-be-processed data table is obtained according to the obtained data.
[0155] In the embodiment, the second characteristic information of the first dimension field includes the number of dimension fields in the to-be-processed data table having the same type as the first dimension field, the number of dimension fields having a father-son relationship with the first dimension field, the number of dimension fields having the same type as the first dimension field and having a father-son relationship with the first dimension field, the enumeration value of the field type, and the like. In the data table, the father-son relationship refers to a case where one side can contain the other side or multiple sides. If the dimension field A has a father-son relationship with the first dimension field, it means that the dimension field A contains the first dimension field; or if the dimension field A has a father-son relationship with the first dimension field, it means that the first dimension field contains the dimension field A.
[0156] It should be noted that the father-son relationship refers to the concept of dimension layering and the concept of layering of dimensions of the same category. The concept of layering refers to a mapping sequence that maps a bottom layer concept to a higher layer and more general concept. For example, Canada contains Vancouver, and both belong to a father-son relationship. The father relationship refers to the relationship between a dimension field with a higher layer concept and a dimension field with a lower layer concept. For example, a province is the father of a district and a city, that is, the relationship between a province and a district is a father relationship. The son relationship refers to the relationship between a dimension field with a lower layer concept and a dimension field with a higher layer concept. For example, the relationship between a city and a province is a son relationship, that is, the city is the son of the province. If a dimension field has many father relationships, the analysis value of the dimension field may not be great.
[0157] It should be noted that in the embodiment, the enumeration value corresponds to the type of each dimension field. It is assumed that the form of the enumeration value is <identifier 1> = <type 1>, such as <n1>= <time type>, <n2>= <date type>, <n3>= <value type> or the like, if the type of the first dimension field is a value type, then the enumeration value type of the cell set corresponding to the first dimension field is <n3>=<value type>, the type of the first dimension field is determined by the corresponding identifier N3.
[0158] According to the data table field graph generation method provided in the application, the second feature information of the first dimension field is determined according to the type of the first dimension field and the relationship with other dimension fields in the data table to be processed, thereby providing data support for subsequent accurate determination of the recommendation score of the first dimension field.
[0159] Based on any of the above embodiments, in the present embodiment, the dimension fields in the data table to be processed are classified according to the type of the dimension field and the recommendation score of the dimension field, to obtain the category of the dimension field, including:
[0160] In the case where the type of the first dimension field satisfies the type-category mapping relationship in the type-category mapping relationship set, the category corresponding to the first dimension field is determined; wherein the type-category mapping relationship describes the type and the category having a unique corresponding relationship; the first dimension field is any one of the dimension fields in the data table to be processed; in the case where the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, the category corresponding to the first dimension field is determined according to the relevance between the first dimension field and the dimension field of which the category has been determined in the data table to be processed and the recommendation score of the first dimension field.
[0161] In the present embodiment, the type-category mapping relationship set refers to the set information composed of the corresponding types in each category, and the type-category mapping relationship describes the type and the category having a unique corresponding relationship. For example, the person name type corresponds to the person category, the organization type corresponds to the person category, the mobile phone number type corresponds to the person category, the phone number type corresponds to the person category, the text numerical type corresponds to the person category, and the ID number type corresponds to the person category. In the present embodiment, the type-category mapping relationship set is: the time category includes the date type, the time type and the date-time type; the person category includes the place name type, the person name type, the organization type, the mobile phone number type, the phone number type, the text numerical type and the ID number type; the place category includes the place name type; the event category includes the verb type. It should be noted that the determination of each category and the type corresponding to each category can be set according to actual needs, which is not limited herein.
[0162] In the embodiment, in the case that the first dimension field is not successfully determined as the above four categories according to the type of the first dimension field, the category corresponding to the first dimension field can be determined according to the correlation between the first dimension field and the dimension field with a determined category. It should be noted that the correlation between the first dimension field and the dimension field with a determined category can be calculated by using the method of obtaining the correlation score by using the containing relationship classification model, and the category corresponding to the dimension field with a larger correlation score is determined as the category of the first dimension field. For example, the category of the dimension field A is the time category, the category of the dimension field B is the place category, the correlation score between the first dimension field and the dimension field A is 0.9, and the correlation score between the first dimension field and the dimension field B is 0.5. Since the correlation score between the first dimension field and the dimension field A is greater than the correlation score between the first dimension field and the dimension field B, the category of the first dimension field is determined as the category corresponding to the dimension field A, i.e., the time category.
[0163] In the embodiment, the category of the first dimension field can also be determined by the recommendation score of the first dimension field. For example, the category of the first dimension field is determined as the event category according to the type of the first dimension field, and the recommendation analysis number of the first dimension field is less than or equal to 0.1, and then the category of the first dimension field is determined as the "not recommended event" category. It should be noted that the preset judgment condition can be other conditions, which are not limited herein.
[0164] According to the data table field graph generation method provided by the present application, the category of the first dimension field is determined by judging whether the type of the first dimension field satisfies the type in the type category mapping relationship set, which provides data support for subsequent accurate generation of the field graph, and ensures the accuracy and efficiency of data processing.
[0165] Based on any of the above embodiments, in the embodiment, the type category mapping relationship set includes one or more of the following type category mapping relationships: date type corresponds to time category, time type corresponds to time category, date and time type corresponds to time category, place name type corresponds to place category, person name type corresponds to person category, organization and group type corresponds to person category, mobile phone number type corresponds to person category, telephone number type corresponds to person category, text type numerical value type corresponds to person category, ID number type corresponds to person category, and verb type corresponds to event category. Correspondingly, in the case that the type of the first dimension field satisfies the type category mapping relationship in the type category mapping relationship set, the category corresponding to the first dimension field is determined, including:
[0166] According to the type of the first dimension field, the corresponding type category mapping relationship in the type category mapping relationship set is determined, and the category of the first dimension field is determined according to the determined type category mapping relationship.
[0167] In the embodiment, when the first dimension field satisfies the type-category mapping relationship, the category of the first dimension field can be determined according to the corresponding type-category mapping relationship. For example, when the type of the first dimension field is a gerund type, the mapping relationship of the gerund type corresponding to the event category is satisfied, and the event category is determined as the category of the first dimension field. It should be noted that, in the embodiment, the type-category mapping relationship in the type-category mapping relationship set is as described above, and in other embodiments, other type-category mapping relationships can also be included, such as a house number type corresponding to an event type, which is not limited here.
[0168] According to the data table field graph generation method provided in the embodiment, the category of the first dimension field can be quickly identified by using the category determination method, and the accuracy and efficiency of category identification and determination are improved.
[0169] Based on any of the above embodiments, in the embodiment, when the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, the category corresponding to the first dimension field is determined according to the relevance between the first dimension field and the category-determined dimension field in the data table to be processed and the recommendation score of the first dimension field, including:
[0170] determining whether the type of the first dimension field is a noun type; in the case where the type of the first dimension field is a noun type, calculating the relevance between the first dimension field and each dimension field determined as a person category, and calculating the relevance between the first dimension field and each dimension field determined as a place category; determining the category of the first dimension field according to the comparison result of the first relevance and the second relevance; wherein the first relevance is the maximum value of the relevance between the first dimension field and each dimension field determined as a person category, and the second relevance is the maximum value of the relevance between the first dimension field and each dimension field determined as a place category.
[0171] In the embodiment, when the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set and it is determined that the type of the first dimension field is a noun type, the correlation analysis method in the multivariate statistical analysis method is used to extract the correlation features of the first dimension field and each dimension field determined as the person category and each dimension field determined as the place category, and the obtained correlation features are input into the containing relationship analysis model to obtain a first correlation degree between the first dimension field and the dimension field determined as the person category and a second correlation degree between the first dimension field and the dimension field determined as the place category, and then the category of the first dimension field is determined according to the size relationship between the first correlation degree and the second correlation degree. It should be noted that the first correlation degree is the maximum value of the correlation degrees between the first dimension field and each dimension field determined as the person category, and the second correlation degree is the maximum value of the correlation degrees between the first dimension field and each dimension field determined as the place category.
[0172] According to the data table field graph generation method provided in the embodiment, when the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set and it is determined that the type of the first dimension field is a noun type, only the correlation degrees between the first dimension field and each dimension field determined as the place category and between each dimension field determined as the place category are calculated, so that the category of the first dimension field can be quickly determined, the efficiency of category determination is improved, and data support is provided for generation of the data table field graph.
[0173] Based on any of the above embodiments, in the embodiment, the category of the first dimension field is determined according to the comparison result of the first correlation degree and the second correlation degree, including:
[0174] In the case that the numerical value of the first correlation degree is greater than the first threshold value and the first correlation degree is greater than the second correlation degree, the category of the first dimension field is determined as the person category; in the case that the numerical value of the second correlation degree is greater than the first threshold value and the second correlation degree is greater than the first correlation degree, the category of the first dimension field is determined as the place category.
[0175] In the embodiments, when the value of the first correlation degree is greater than the first threshold value and the first correlation degree is greater than the second correlation degree, the category of the first dimension field is determined as the person category, wherein the first threshold value can be set as 0.5. For example, the value of the first correlation degree is 0.7 and the value of the second correlation degree is 0.4. By comparison, it is determined that the value 0.7 of the first correlation degree is greater than the first threshold value 0.5, and the value 0.7 of the first correlation degree is greater than the value 0.4 of the second correlation degree. Since the first correlation degree is the maximum value calculated from the first dimension field and each dimension field determined as the person category, the category of the first dimension field is determined as the person category. Similarly, when the value of the second correlation degree is greater than the first threshold value and the second correlation degree is greater than the value of the first correlation degree, the category of the first dimension field is determined as the place category. It should be noted that the size of the first threshold value can be set according to actual needs, and can also be 0.3, etc., which is not limited here.
[0176] According to the data table field graph generation method provided by the present application, when the first dimension field does not satisfy the type category mapping relationship in the type category mapping relationship set and the type of the first dimension field is determined as the noun type, the category of the first dimension field can be quickly determined by comparing the size relationship between the first correlation degree, the second correlation degree and the first threshold value, the efficiency of category determination is improved, and data support is provided for the generation of the data table field graph.
[0177] Based on any of the above embodiments, in this embodiment, in the case that the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, the category corresponding to the first dimension field is determined according to the relevance between the first dimension field and the category-determined dimension field in the data table to be processed and the recommendation score of the first dimension field, and further comprising: in the case that the type of the first dimension field is not a noun type, calculating the relevance between the first dimension field and each dimension field determined as a person category, calculating the relevance between the first dimension field and each dimension field determined as a place category, and calculating the relevance between the first dimension field and each dimension field determined as a time category; in the case that the first relevance is greater than a second threshold, determining that the category of the first dimension field is a person category; in the case that the second relevance is greater than a third threshold, determining that the category of the first dimension field is a place category; in the case that the third relevance is greater than a fourth threshold, determining that the category of the first dimension field is a time category; in the case that the category of the first dimension field is not identified as a person category, a place category or a time category, identifying the category of the first dimension field as an event category; wherein the first relevance is the maximum value of the relevance between the first dimension field and each dimension field determined as a person category; the second relevance is the maximum value of the relevance between the first dimension field and each dimension field determined as a place category; and the third relevance is the maximum value of the relevance between the first dimension field and each dimension field determined as a time category.
[0178] In the embodiment, when the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, and the category of the first dimension field is not a noun type, the correlation degrees between the first dimension field and each dimension field determined as the person category, each dimension field determined as the place category, and each dimension field determined as the time category are respectively calculated, and the maximum values are respectively taken to obtain a first correlation degree, a second correlation degree, and a third correlation degree. Assuming that the second threshold value is 0.4, the third threshold value is 0.5, and the fourth threshold value is 0.4, the obtained correlation degrees are compared with the preset threshold values in sequence, when the first correlation degree is 0.5 and greater than the second threshold value 0.4, it is determined that the category of the first dimension field is the person category; if the first correlation degree is less than the second threshold value, the second correlation degree is 0.6 and greater than the third threshold value, it is determined that the category of the first dimension field is the place category; if the first correlation degree is less than the second threshold value and the second correlation degree is less than the third threshold value, the third correlation degree 0.55 is greater than the fourth threshold value, it is determined that the category of the first dimension field is the time category. It should be noted that the sizes of the three threshold values can be set according to actual needs, and are not specifically limited here. Meanwhile, the first correlation degree, the second correlation degree, and the third correlation degree are compared with the preset threshold values in sequence, after the first dimension field is determined as the person category, subsequent comparison and confirmation are not performed, and if the category of the first dimension field is not determined according to the first correlation degree, the second correlation degree, and the third correlation degree, the category of the first dimension field is determined as the event category.
[0179] According to the data table field graph generation method provided in the embodiment, when the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, and the type of the first dimension field is not a noun type, the sizes of the obtained first correlation degree, second correlation degree, and third correlation degree and the preset threshold values are compared, the category of the first dimension field can be quickly determined, the efficiency of category determination is improved, and data support is provided for generation of the data table field graph.
[0180] Based on any of the above embodiments, in the embodiment, when the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, the category corresponding to the first dimension field is determined according to the correlation degrees between the first dimension field and the dimension fields with determined categories in the data table to be processed and the recommendation score of the first dimension field, and the method further includes:
[0181] The recommendation score of the first dimension field of the event category is obtained, and in a case where the recommendation score is less than a fifth threshold value, the category of the first dimension field is modified from the event category to the non-recommended event category.
[0182] In the embodiment, when the category of the first dimension field is the event category, a recommendation score of the first dimension field is obtained, and in a case where the recommendation score is less than a third threshold value, the category of the first dimension field is modified from the event category to the non-recommended event category. Assuming that the fifth threshold value is 0.1, if the recommendation score of the first dimension field obtained is 0.2, the category of the first dimension field is not modified, and if the recommendation score of the first dimension field obtained is 0.08, which is less than the fifth threshold value, the category of the first dimension field is modified from the event category to the non-recommended event category. It should be noted that the size of the fifth threshold value can be set according to actual needs, and is not specifically limited herein.
[0183] According to the data table field graph generation method provided in the embodiment, the recommendation score of the first dimension field with the category of the event category is obtained, and in a case where the recommendation score is less than the fifth threshold value, the category of the first dimension field is modified from the event category to the non-recommended event category, so that the category of the first dimension field is accurately determined, the category of the first dimension field is quickly and accurately determined, the efficiency of category determination is improved, and data support is provided for generation of the data table field graph.
[0184] Based on any of the above embodiments, in the embodiment, the category of the metric field is determined, and the recommendation score of the metric field is determined according to the feature information of the to-be-processed data table and the feature information of the metric field, including: determining the category of the first metric field as the metric category; wherein the first metric field is any one of the metric fields in the to-be-processed data table; obtaining the feature information of the to-be-processed data table; wherein the feature information of the to-be-processed data table includes the description feature of the to-be-processed data table and the classification feature of the to-be-processed data table; obtaining the first feature information of the first metric field according to the data distribution in the cell set corresponding to the first metric field in the to-be-processed data table and the distribution of the first metric field in the to-be-processed data table; obtaining the second feature information of the first metric field in the to-be-processed data table according to the data statistics in the cell set corresponding to the first metric field in the to-be-processed data table; inputting the feature information of the to-be-processed data table, the first feature information of the first metric field, and the second feature information of the first metric field into a pre-trained metric field recommendation model to obtain the recommendation score of the first metric field; wherein the metric field recommendation model is trained based on the feature information of the sample data table, the first feature information of the second metric field in the sample data table, the second feature information of the second metric field, and the recommendation score label of the second metric field; wherein the second metric field is any one or more of the metric fields in the sample data table.
[0185] In the embodiment, in order to obtain the recommended score of the first metric field, the obtained feature information of the to-be-processed data table, the first feature information and the second feature information of the first metric field are input into the metric field recommendation model obtained by pre-training, to obtain the recommended score of the first metric field, and the recommended score is set between 0 and 1. The metric field recommendation model is obtained by pre-training the training sample by using the random forest algorithm. The feature information of the sample data table obtained in advance, the first feature information of the second metric field in the sample data table, the second feature information of the second metric field, and the recommended score label of the second metric field are used to train the classifier by using the random forest algorithm, to obtain the metric field recommendation model. The specific training method is not described in detail here.
[0186] It should be noted that the metric field is a field indicating the quantity in the corresponding cell, and the numerical type field belongs to the metric field. There are more metric category fields in common, such as salary slips and score sheets. In the score sheet, the score values in the cell set corresponding to the fields of English, Chinese, and chemistry, and the score values in the cell set corresponding to the fields of total score, average score, and ranking are obtained. According to the recommended score, the recommended score of "total score, average score, and ranking" is generally higher than that of other fields.
[0187] In the embodiment, the feature information of the to-be-processed data table needs to be obtained, including the description features and classification features of the to-be-processed data table. The description features include the number of non-empty cell sets in the to-be-processed data table, the maximum value of the number of non-empty cells included in each cell set, the minimum value of the number of non-repeated data included in each cell set, the maximum value of the number of non-repeated data included in each cell set, the average value of the number of non-repeated data included in each cell set, and the standard deviation of the number of non-repeated data included in each cell set.
[0188] The classification features of the to-be-processed data table include: the enumeration value of each field type in the to-be-processed data table, the number of fields of the person name type in the to-be-processed data table, the number of fields of the place name type in the to-be-processed data table, the number of fields of the numerical type in the to-be-processed data table, the number of fields of the date type in the to-be-processed data table, the number of fields of the time type in the to-be-processed data table, and the number of fields of the gerund type in the to-be-processed data table.
[0189] In the embodiment, the first characteristic information of the first metric field is determined according to the data distribution in the cell set corresponding to the first metric field and the distribution of the first metric field in the to-be-processed data table, and the second characteristic information of the first metric field is obtained according to the data statistics in the cell set corresponding to the first metric field in the to-be-processed data table. It should be noted that the first characteristic information includes the index value of the first metric field and the like, and the second characteristic information includes the average value of the numbers in the cell set corresponding to the first metric field and the like, which will be described in detail in the following embodiments.
[0190] According to the data table field graph generation method provided by the present application, the characteristic information of the to-be-processed data table, the first characteristic information of the first metric field and the second characteristic information of the first metric field are input into the pre-trained metric field recommendation model to obtain the recommendation score of the first metric field. The present application can quickly obtain the recommendation score of the first metric field, which provides data support for the subsequent generation of field graphs.
[0191] Based on any of the above embodiments, as shown in FIG. 1, in the embodiment, the first characteristic information of the first metric field in the to-be-processed data table is obtained according to the data distribution in the cell set corresponding to the first metric field and the distribution of the first metric field in the to-be-processed data table, and at least includes any one of the following: Figure 4
[0192] The index value of the first metric field is obtained, wherein the index value is used to describe the position of the first metric field in the to-be-processed data table; the number of non-repeated data in the cell set corresponding to the first metric field is obtained; the number of cells in the cell set corresponding to the first metric field is obtained; the number of non-empty cells in the cell set corresponding to the first metric field is obtained; in the to-be-processed data table, the index value of the first metric field is compared with the index values of other metric fields to obtain the number of other metric fields whose index value is less than the index value of the first metric field and the number of other metric fields whose index value is greater than the index value of the first metric field; in the to-be-processed data table, the index value of the first metric field is compared with the index values of dimension fields to obtain the number of dimension fields whose index value is less than the index value of the first metric field and the number of dimension fields whose index value is greater than the index value of the first metric field; and the first characteristic information of the first metric field in the to-be-processed data table is obtained according to the obtained data.
[0193] In the embodiment, the first feature information is obtained according to the data distribution in the cell set corresponding to the first metric field in the to-be-processed data table and the distribution of the first metric field in the to-be-processed data table, and the first feature information includes the index value of the first metric field, the number of non-repeated data in the cell set corresponding to the first metric field, the number of cells, and the number of non-empty cells, and the like. It should be noted that the index value of the metric field needs to be obtained in the embodiment to determine the position of the metric field, and the position of the metric field has an important influence on data analysis. For example, in data summation, the data on the right side of the metric field has more analysis value, and if there is no other metric field on the left side of the current metric field and there is other metric field on the right side, it is often not recommended to analyze, because the metric field may be a serial number. As shown in the to-be-processed data table in Table 6 below, assuming that the field with the field name "basic salary" is the first metric field, wherein the index value of the first metric field is 2, the number of non-repeated data is 1, the number of cells is 3, and the number of non-empty cells is 3.
[0194] Table 6
[0195]
[0196]
[0197] It should be noted that the index value of the first metric field also needs to be compared with the index value of other metric fields, the number of other metric fields with an index value less than the index value of the first metric field and the number of other metric fields with an index value greater than the index value of the first metric field are obtained, and the to-be-processed data table shown in Table 6 includes six metric fields. Since the index value of the first metric field is 2, the field with the index value of 1 is a dimension field, therefore, the number of other metric fields with an index value less than the index value of the first metric field is 0, and the number of other metric fields with an index value greater than the index value of the first metric field is 5.
[0198] Similarly, the index value of the first metric field is compared with the index value of the dimension field, the number of dimension fields with an index value less than the index value of the first metric field and the number of dimension fields with an index value greater than the index value of the first metric field are obtained, and it can be seen from the above Table 6 that the number of dimension fields with an index value less than the index value of the first metric field is 1, and the number of dimension fields with an index value greater than the index value of the first metric field is 0. It should be noted that the first metric field is any metric field in the to-be-processed data table.
[0199] According to the data table field map generation method provided in the present application, the first characteristic information of the first measurement field is determined according to the distribution of the first measurement field in the to-be-processed data table and the distribution of data in the cell set corresponding to the first measurement field, thereby providing data support for subsequent generation of a field map according to data.
[0200] Based on any of the above embodiments, in the present embodiment, the second characteristic information of the first measurement field in the to-be-processed data table is obtained according to the data statistics in the cell set corresponding to the first measurement field, and at least includes any of the following: an average value of the numbers in the cell set corresponding to the first measurement field; a median of the numbers in the cell set corresponding to the first measurement field; a standard deviation of the numbers in the cell set corresponding to the first measurement field; a minimum value of the numbers in the cell set corresponding to the first measurement field; a maximum value of the numbers in the cell set corresponding to the first measurement field; a three-quarter quantile value of the numbers in the cell set corresponding to the first measurement field; a one-quarter quantile value of the numbers in the cell set corresponding to the first measurement field; and a difference between the one-quarter quantile value and the three-quarter quantile value of the numbers in the cell set corresponding to the first measurement field.
[0201] In the present embodiment, the second characteristic information of the first measurement field needs to be obtained according to the data statistics in the cell set corresponding to the first measurement field in the to-be-processed data table. The second characteristic information includes: an average value of the numbers in the cell set corresponding to the first measurement field, a median of the numbers, a standard deviation of the numbers, a minimum value of the numbers, a maximum value of the numbers, a three-quarter quantile value of the numbers, a one-quarter quantile value of the numbers, and a difference between the one-quarter quantile value and the three-quarter quantile value of the numbers. The quantile value is one of the characteristic numbers of a random variable, and has many applications in statistics. In general data analysis, 25 quantiles, 50 quantiles, and 75 quantiles can be calculated. The one-quarter quantile value is a value calculated by 25 quantiles, and the three-quarter quantile value is a value calculated by 75 quantiles. The specific calculation method can be according to the calculation method in the prior art, which is not described here.
[0202] For example, if the first measurement field is the field with the field name "age of the library" in Table 1, the average value of the numbers in the cell set corresponding to the first measurement field is 5.75, the median of the numbers is 6, the standard deviation of the numbers is 1.35, the minimum value of the numbers is 4, the maximum value of the numbers is 7, the three-quarter quantile value of the numbers is 7, the one-quarter quantile value of the numbers is 4.75, and the difference between the one-quarter quantile value and the three-quarter quantile value of the numbers is 2.25.
[0203] According to the data table field graph generation method provided in the application, the second feature information of the first measurement field is obtained through the data statistical situation in the cell set corresponding to the first measurement field, thereby providing data support for generating the field graph according to the data.
[0204] Based on any of the above embodiments, in the present embodiment, as shown in Figure 5 the data table to be recognized is obtained, the types of the fields in the data table are determined, the dimension fields and the cell sets corresponding to the dimension fields are determined according to the types of the fields, and the measurement fields and the cell sets corresponding to the measurement fields are determined. In the present embodiment, the time category, the place category, the person category and the event category are determined according to the types of the determined dimension fields, wherein the dimension fields of the time type, the date type and the date-time type belong to the time category, the dimension fields of the person name type, the organization type, the mobile phone number type, the phone number type, the text numerical value type and the ID number type belong to the person category, and the dimension fields of the verb type belong to the event category.
[0205] For the first dimension field of the noun type, which is not determined to belong to any category according to the types of the dimension fields, the correlation degrees of the first dimension field with the dimension fields determined to belong to the person category and the correlation degrees of the first dimension field with the dimension fields determined to belong to the place category are calculated, and the category of the first dimension field is determined by judging the relationship with the preset threshold.
[0206] For the first dimension field of the type other than the noun type, the correlation degrees of the first dimension field with the dimension fields determined to belong to the person category, the correlation degrees of the first dimension field with the dimension fields determined to belong to the place category and the correlation degrees of the first dimension field with the dimension fields determined to belong to the time category are calculated, and the category of the first dimension field is determined by comparing with the preset threshold. The category of the first dimension field can also be modified by the recommendation score of the first dimension field. When the category of the first dimension field is the event category and the recommendation score is less than the threshold, the category of the first dimension field is modified from the event category to the event category not recommended.
[0207] In the present embodiment, the category of the measurement field is also determined according to the recommendation score of the measurement field. According to the obtained information of each category and the information of the cell sets corresponding to each field contained in each category, the field graph of the data table to be processed is generated.
[0208] It should be noted that the specific content of the feature information of the data table to be processed, the first feature information and the second feature information of the first dimension field, the first feature information and the second feature information of the first measurement field, and the above-mentioned related information is as described in the above embodiments, which will not be described in detail here.
[0209] Figure 6 A data table field map generation device is provided in the present application, as shown in Figure 6 The data table field map generation device provided by the present application comprises:
[0210] The first determining module 601 is configured to determine the types of the fields in the to-be-processed data table, determine the dimension fields and the measure fields in the to-be-processed data table according to the types of the fields, wherein the dimension fields are used to describe the meanings of the data in the corresponding cells, and the measure fields are used to describe the quantities of the data in the corresponding cells; the second determining module 602 is configured to determine the recommended scores of the dimension fields according to the types of the dimension fields, the feature information of the dimension fields and the feature information of the to-be-processed data table; the classifying module 603 is configured to classify the dimension fields in the to-be-processed data table according to the types of the dimension fields and the recommended scores of the dimension fields, and obtain the categories of the dimension fields; the third determining module 604 is configured to determine the categories of the measure fields, and determine the recommended scores of the measure fields according to the feature information of the to-be-processed data table and the feature information of the measure fields; and the generating module 605 is configured to generate the field map of the to-be-processed data table according to the categories of the dimension fields, the recommended scores of the dimension fields, the categories of the measure fields and the recommended scores of the measure fields.
[0211] According to the data table field map generation device provided by the present application, the data in the data table of the massive data can be accurately displayed, the user can accurately obtain the recommended analysis conditions and the category conditions of each field according to the data information displayed by the generated field map, the field that needs to be recommended and analyzed can be quickly found, the efficiency of the data table processing is improved, and the user experience is improved.
[0212] Based on any of the above embodiments, in the present embodiment, the second determining module 602 is further configured to:
[0213] obtain feature information of the to-be-processed data table; obtain first feature information of a first dimension field in the to-be-processed data table according to a data distribution in a cell set corresponding to the first dimension field and a distribution of the first dimension field in the to-be-processed data table, wherein the first dimension field is any one dimension field in the to-be-processed data table; obtain second feature information of the first dimension field according to types of each dimension field in the to-be-processed data table and a relationship between the first dimension field and other dimension fields in the to-be-processed data table; input the feature information of the to-be-processed data table, the first feature information of the first dimension field, and the second feature information of the first dimension field into a pre-trained dimension field recommendation model to obtain a recommendation score of the first dimension field; wherein the dimension field recommendation model is trained based on feature information of a sample data table, first feature information of a second dimension field in the sample data table, second feature information of the second dimension field, and a recommendation score label of the second dimension field; and the second dimension field is any one or more dimension fields in the sample data table.
[0214] According to the data table field graph generation apparatus provided in the application, the feature information of the to-be-processed data table, the first feature information and the second feature information of the first dimension field are input into the pre-trained dimension field recommendation model to obtain the recommendation score of the first dimension field, so that the user can quickly find the field to be analyzed from multiple dimension fields, the processing speed and accuracy of the dimension field recommendation analysis are improved, and the cost of manual data checking is reduced.
[0215] According to any one of the above embodiments, in this embodiment, the second determination module 602 is further configured to:
[0216] obtain a description feature of the to-be-processed data table;
[0217] obtain a classification feature of the to-be-processed data table.
[0218] According to the data table field graph generation apparatus provided in the application, the description feature and the classification feature of the to-be-processed data table are obtained, the obtained information is used for subsequent determination and acquisition of the dimension field recommendation score, the processing speed of the dimension field data recommendation is improved, and data support is provided for subsequent generation of the field graph.
[0219] According to any one of the above embodiments, in this embodiment, the second determination module 602 is further configured to:
[0220] obtain description feature information of the first dimension field;
[0221] obtain position feature information of the first dimension field.
[0222] The data table field graph generation device provided in the application determines the distribution of the first dimension field in the data table to be processed and the distribution of the corresponding cell data through the index value of the first dimension field, and then determines the first feature information of the first dimension field, thereby providing data support for accurately determining the recommendation score of the first dimension field.
[0223] According to any of the above embodiments, in this embodiment, the second determination module 602 is further configured to:
[0224] According to the type of each dimension field in the data table to be processed, the number of dimension fields with the same type as the first dimension field in the data table to be processed is obtained; according to the relationship between the first dimension field and other dimension fields in the data table to be processed, the number of dimension fields in the data table to be processed that have a parent-child relationship with the first dimension field is obtained; according to the relationship between the first dimension field and other dimension fields in the data table to be processed, the number of dimension fields in the data table to be processed that have a parent-child relationship with the first dimension field is obtained; according to the relationship between the first dimension field and other dimension fields in the data table to be processed, the number of dimension fields in the data table to be processed that have the same type as the first dimension field and have a parent-child relationship is obtained; according to the relationship between the first dimension field and other dimension fields in the data table to be processed, the number of dimension fields in the data table to be processed that have the same type as the first dimension field and have a parent-child relationship is obtained; the enumeration value of the field type is obtained; and the second feature information of the first dimension field in the data table to be processed is obtained according to the obtained data.
[0225] The data table field graph generation device provided in the application determines the second feature information of the first dimension field through the type of the first dimension field and the relationship between the first dimension field and other dimension fields in the data table to be processed, thereby providing data support for accurately determining the recommendation score of the first dimension field.
[0226] According to any of the above embodiments, in this embodiment, the classification module 604 is further configured to:
[0227] In the case where the type of the first dimension field satisfies the type-class mapping relationship in the type-class mapping relationship set, the category corresponding to the first dimension field is determined; wherein the type-class mapping relationship describes the type and the category that have a unique corresponding relationship; the first dimension field is any one of the dimension fields in the data table to be processed; in the case where the type of the first dimension field does not satisfy the type-class mapping relationship in the type-class mapping relationship set, the category corresponding to the first dimension field is determined according to the relevance between the first dimension field and the dimension fields in the data table to be processed whose categories have been determined and the recommendation score of the first dimension field.
[0228] The data table field graph generation device provided in the present application can determine the category of the first dimension field by judging whether the type of the first dimension field meets the type in the type-category mapping relationship set, thereby providing data support for subsequent accurate generation of the field graph and ensuring the accuracy and efficiency of data processing.
[0229] According to any one of the above embodiments, in the present embodiment, the classification module 604 is further configured to:
[0230] According to the type of the first dimension field, a corresponding type-category mapping relationship is determined in the type-category mapping relationship set; and according to the determined type-category mapping relationship, the category of the first dimension field is determined.
[0231] The data table field graph generation device provided in the present application can rapidly identify the category of the first dimension field by using the above-mentioned category determination method, thereby improving the accuracy and efficiency of category identification and determination.
[0232] According to any one of the above embodiments, in the present embodiment, the classification module 604 is further configured to:
[0233] In the case where the type of the first dimension field is a noun type, the correlation between the first dimension field and each dimension field that has been determined as a person category is calculated, and the correlation between the first dimension field and each dimension field that has been determined as a place category is calculated; according to the comparison result of the first correlation and the second correlation, the category of the first dimension field is determined; wherein the first correlation is the maximum value of the correlation between the first dimension field and each dimension field that has been determined as a person category, and the second correlation is the maximum value of the correlation between the first dimension field and each dimension field that has been determined as a place category.
[0234] The data table field graph generation device provided in the present application can rapidly determine the category of the first dimension field by only calculating the correlation between the first dimension field and each dimension field that has been determined as a place category and the correlation between the first dimension field and each dimension field that has been determined as a place category in the case where the first dimension field does not meet the type-category mapping relationship in the type-category mapping relationship set and the type of the first dimension field is a noun type, thereby improving the efficiency of category determination and providing data support for generation of the data table field graph.
[0235] According to any one of the above embodiments, in the present embodiment, the classification module 604 is further configured to:
[0236] In a case where the value of the first correlation degree is greater than the first threshold value and the first correlation degree is greater than the second correlation degree, the category of the first dimension field is determined as a person category; in a case where the value of the second correlation degree is greater than the first threshold value and the second correlation degree is greater than the first correlation degree, the category of the first dimension field is determined as a place category.
[0237] According to the data table field graph generation device provided in the present application, in a case where the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set and the type of the first dimension field is a noun type, the category of the first dimension field can be quickly determined by comparing the size relationship among the obtained first correlation degree, second correlation degree and first threshold value, the efficiency of category determination is improved, and data support is provided for generation of the data table field graph.
[0238] Based on any of the above embodiments, in the present embodiment, the classification module 604 is further configured to:
[0239] In a case where the type of the first dimension field is not a noun type, the correlation degrees between the first dimension field and each dimension field determined as a person category, the correlation degrees between the first dimension field and each dimension field determined as a place category, and the correlation degrees between the first dimension field and each dimension field determined as a time category are calculated; in a case where the first correlation degree is greater than a second threshold value, the category of the first dimension field is determined as a person category; in a case where the second correlation degree is greater than a third threshold value, the category of the first dimension field is determined as a place category; in a case where the third correlation degree is greater than a fourth threshold value, the category of the first dimension field is determined as a time category; in a case where the category of the first dimension field is not identified as a person category, place category or time category, the category of the first dimension field is identified as an event category; wherein the first correlation degree is the maximum value of the correlation degrees between the first dimension field and each dimension field determined as a person category; the second correlation degree is the maximum value of the correlation degrees between the first dimension field and each dimension field determined as a place category; and the third correlation degree is the maximum value of the correlation degrees between the first dimension field and each dimension field determined as a time category.
[0240] According to the data table field graph generation device provided in the present application, in a case where the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set and the type of the first dimension field is not a noun type, the category of the first dimension field can be quickly determined by comparing the size relationship among the obtained first correlation degree, second correlation degree, third correlation degree and preset threshold value, the efficiency of category determination is improved, and data support is provided for generation of the data table field graph.
[0241] Based on any of the above embodiments, in the present embodiment, the classification module 604 is further configured to:
[0242] obtaining a recommendation score of the first dimension field of the event category; and modifying the category of the first dimension field from the event category to a non-recommended event category if the recommendation score is less than a fifth threshold.
[0243] According to the data table field graph generation device provided in the present application, the category of the first dimension field is accurately determined by obtaining the recommendation score of the first dimension field of the event category and modifying the category of the first dimension field from the event category to a non-recommended event category if the recommendation score is less than a fifth threshold, so that the category of the first dimension field is quickly and accurately determined, the efficiency of category determination is improved, and data support is provided for the generation of the data table field graph.
[0244] Based on any of the above embodiments, in the present embodiment, the third determination module 604 is further configured to:
[0245] determining the category of the first dimension field as a measurement category; obtaining feature information of the to-be-processed data table; wherein the feature information of the to-be-processed data table includes a description feature of the to-be-processed data table and a classification feature of the to-be-processed data table; obtaining first feature information of the first dimension field according to the data distribution in the cell set corresponding to the first dimension field in the to-be-processed data table and the distribution of the first dimension field in the to-be-processed data table; obtaining second feature information of the first dimension field in the to-be-processed data table according to the data statistics in the cell set corresponding to the first dimension field in the to-be-processed data table; inputting the feature information of the to-be-processed data table, the first feature information of the first dimension field and the second feature information of the first dimension field into a pre-trained measurement field recommendation model to obtain a recommendation score of the first dimension field; wherein the measurement field recommendation model is trained based on the feature information of a sample data table, the first feature information of a second dimension field in the sample data table, the second feature information of the second dimension field and a recommendation score label of the second dimension field; and the second dimension field is any one or more measurement fields in the sample data table.
[0246] According to the data table field graph generation device provided in the present application, the feature information of the to-be-processed data table, the first feature information of the first dimension field and the second feature information of the first dimension field are input into the pre-trained measurement field recommendation model to obtain the recommendation score of the first dimension field. The present application can quickly obtain the recommendation score of the first dimension field and provide data support for the subsequent generation of the field graph.
[0247] Based on any of the above embodiments, in the present embodiment, the third determination module 604 is further configured to obtain any of the following information:
[0248] obtain an index value of the first metric field, wherein the index value is used to describe a position of the first metric field in the data table to be processed, obtain a number of non-repeated data in the cell set corresponding to the first metric field, obtain a number of cells in the cell set corresponding to the first metric field, obtain a number of non-empty cells in the cell set corresponding to the first metric field, compare the index value of the first metric field with index values of other metric fields in the data table to be processed, and obtain a number of the other metric fields with an index value smaller than the index value of the first metric field and a number of the other metric fields with an index value larger than the index value of the first metric field, compare the index value of the first metric field with index values of dimension fields in the data table to be processed, and obtain a number of the dimension fields with an index value smaller than the index value of the first metric field and a number of the dimension fields with an index value larger than the index value of the first metric field, and obtain the first characteristic information of the first metric field in the data table to be processed according to the obtained data.
[0249] According to the data table field graph generation device provided in the application, the first characteristic information of the first metric field is determined according to the distribution of the first metric field in the data table to be processed and the distribution of data in the cell set corresponding to the first metric field, thereby providing data support for generating a field graph according to data.
[0250] According to any of the above embodiments, in this embodiment, the third determination module 604 is further configured to obtain any of the following information:
[0251] obtain an average value of the numbers in the cell set corresponding to the first metric field, obtain a median of the numbers in the cell set corresponding to the first metric field, obtain a standard deviation of the numbers in the cell set corresponding to the first metric field, obtain a minimum value of the numbers in the cell set corresponding to the first metric field, obtain a maximum value of the numbers in the cell set corresponding to the first metric field, obtain a three-quarter quantile value of the numbers in the cell set corresponding to the first metric field, obtain a one-quarter quantile value of the numbers in the cell set corresponding to the first metric field, and obtain a difference between the one-quarter quantile value and the three-quarter quantile value of the numbers in the cell set corresponding to the first metric field.
[0252] According to the data table field graph generation device provided in the application, the second characteristic information of the first metric field is obtained according to the data statistics in the cell set corresponding to the first metric field, thereby providing data support for generating a field graph according to data.
[0253] Since the device described in the application embodiment has the same principle as the method described in the above embodiments, more detailed explanation is not repeated here.
[0254] Figure 7 An electronic device entity structure schematic diagram provided in the embodiments of the present application is shown in FIG. 1. Figure 7 As shown in FIG. 1, the present application provides an electronic device, comprising a processor 701, a memory 702 and a bus 703.
[0255] The processor 701 and the memory 702 communicate with each other through the bus 703.
[0256] The processor 701 is configured to invoke program instructions in the memory 702 to execute the methods provided in the above method embodiments, for example, comprising: determining the types of each field in a to-be-processed data table, determining dimension fields and measure fields in the to-be-processed data table according to the types of the fields; wherein the dimension fields are used to describe the meanings represented by the data in the corresponding cells, and the measure fields are used to describe the quantities represented by the data in the corresponding cells; determining a recommendation score of the dimension fields according to the types of the dimension fields, characteristic information of the dimension fields and characteristic information of the to-be-processed data table; classifying the dimension fields in the to-be-processed data table according to the types of the dimension fields and the recommendation scores of the dimension fields, to obtain the categories of the dimension fields; determining the categories of the measure fields, and determining recommendation scores of the measure fields according to the characteristic information of the to-be-processed data table and the characteristic information of the measure fields; and generating a field graph of the to-be-processed data table according to the categories of each dimension field in the to-be-processed data table, the recommendation scores of each dimension field, the categories of each measure field and the recommendation scores of each measure field.
[0257] The present application provides a computer readable storage medium in the embodiments, the computer readable storage medium stores computer instructions, the computer instructions make the computer execute the methods provided in the above method embodiments, for example, comprising: determining the types of each field in a to-be-processed data table, determining dimension fields and measure fields in the to-be-processed data table according to the types of the fields; wherein the dimension fields are used to describe the meanings represented by the data in the corresponding cells, and the measure fields are used to describe the quantities represented by the data in the corresponding cells; determining a recommendation score of the dimension fields according to the types of the dimension fields, characteristic information of the dimension fields and characteristic information of the to-be-processed data table; classifying the dimension fields in the to-be-processed data table according to the types of the dimension fields and the recommendation scores of the dimension fields, to obtain the categories of the dimension fields; determining the categories of the measure fields, and determining recommendation scores of the measure fields according to the characteristic information of the to-be-processed data table and the characteristic information of the measure fields; and generating a field graph of the to-be-processed data table according to the categories of each dimension field in the to-be-processed data table, the recommendation scores of each dimension field, the categories of each measure field and the recommendation scores of each measure field.
[0258] The application further provides a computer program product, which comprises a computer program stored on a computer readable storage medium, and the computer program comprises program instructions, and when the program instructions are executed by a computer, the computer can execute the method provided by each method described above, and the method comprises the following steps: determining the types of each field in a to-be-processed data table, determining a dimension field and a measure field in the to-be-processed data table according to the types of the fields; the dimension field is used to describe the meaning of the data in a corresponding cell, and the measure field is used to describe the quantity of the data in the corresponding cell; determining a recommended score of the dimension field according to the type of the dimension field, characteristic information of the dimension field and characteristic information of the to-be-processed data table; classifying the dimension field in the to-be-processed data table according to the type of the dimension field and the recommended score of the dimension field, to obtain the category of the dimension field; determining the category of the measure field, and determining a recommended score of the measure field according to the characteristic information of the to-be-processed data table and the characteristic information of the measure field; and generating a field graph of the to-be-processed data table according to the category of each dimension field in the to-be-processed data table, the recommended score of each dimension field, the category of each measure field and the recommended score of each measure field.
[0259] Those skilled in the art can understand that all or part of the steps of the above method embodiments can be completed by program instruction related hardware, and the foregoing program can be stored in a computer readable storage medium, and the program executes the steps of the above method embodiments when executed; and the foregoing storage medium includes ROM, RAM, a magnetic disc or an optical disc and various storage medium capable of storing program codes.
[0260] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A data table field mapping generation method, characterized by, The method comprises the following steps: determining the types of each field in a to-be-processed data table, determining dimension fields and metric fields in the to-be-processed data table according to the types of the fields; wherein the dimension fields are used to describe the meaning of the data in the corresponding cells, and the metric fields are used to describe the quantity of the data in the corresponding cells; determining a recommendation score of a dimension field according to the type of the dimension field, characteristic information of the dimension field, and characteristic information of the to-be-processed data table; classifying the dimension fields in the to-be-processed data table according to the type of the dimension field and the recommendation score of the dimension field, to obtain the category of the dimension field; determining the category of a metric field, and determining a recommendation score of the metric field according to the characteristic information of the to-be-processed data table and the characteristic information of the metric field; wherein the recommendation score of the metric field is obtained by inputting the characteristic information of the to-be-processed data table, first characteristic information of a first metric field, and second characteristic information of the first metric field into a pre-trained metric field recommendation model; the first metric field is any one of the metric fields in the to-be-processed data table; the first characteristic information is obtained according to the data distribution in a cell set corresponding to the first metric field in the to-be-processed data table and the distribution of the first metric field in the to-be-processed data table; and the second characteristic information is obtained according to the data statistics in the cell set corresponding to the first metric field in the to-be-processed data table; generating a field graph of the to-be-processed data table according to the category of each dimension field in the to-be-processed data table, the recommendation score of each dimension field, the category of each metric field, and the recommendation score of each metric field.
2. The data table field mapping generation method according to claim 1, characterized by, The method comprises the following steps: obtaining the characteristic information of the to-be-processed data table; obtaining first characteristic information of a first dimension field in the to-be-processed data table according to the data distribution in a cell set corresponding to the first dimension field in the to-be-processed data table and the distribution of the first dimension field in the to-be-processed data table; wherein the first dimension field is any one of the dimension fields in the to-be-processed data table; obtaining second characteristic information of the first dimension field according to the type of each dimension field in the to-be-processed data table and the relationship between the first dimension field and other dimension fields in the to-be-processed data table; inputting the characteristic information of the to-be-processed data table, the first characteristic information of the first dimension field, and the second characteristic information of the first dimension field into a pre-trained dimension field recommendation model to obtain the recommendation score of the first dimension field; wherein the dimension field recommendation model is trained based on the characteristic information of a sample data table, first characteristic information of a second dimension field in the sample data table, second characteristic information of the second dimension field, and a recommendation score label of the second dimension field; wherein the second dimension field is any one or more of the dimension fields in the sample data table.
3. The data table field mapping generation method according to claim 2, characterized by, The feature information of the to-be-processed data table includes a description feature of the to-be-processed data table and a classification feature of the to-be-processed data table. Correspondingly, the obtaining of the feature information of the to-be-processed data table includes: obtaining the description feature of the to-be-processed data table; obtaining the classification feature of the to-be-processed data table.
4. The data table field mapping generation method of claim 2, wherein, The first feature information of the first dimension field includes description feature information and position feature information of the first dimension field. Correspondingly, the obtaining of the first feature information of the first dimension field in the to-be-processed data table according to the data distribution in the cell set corresponding to the first dimension field in the to-be-processed data table and the distribution of the first dimension field in the to-be-processed data table includes: obtaining the description feature information of the first dimension field; obtaining the position feature information of the first dimension field.
5. The data table field mapping generation method of claim 2, wherein, The obtaining of the second feature information of the first dimension field according to the types of the dimension fields in the to-be-processed data table and the relationship between the first dimension field and other dimension fields in the to-be-processed data table includes: obtaining the number of the dimension fields with the same type as the first dimension field in the to-be-processed data table according to the types of the dimension fields in the to-be-processed data table; obtaining the number of the dimension fields in the to-be-processed data table that have a parent-child relationship with the first dimension field according to the relationship between the first dimension field and other dimension fields in the to-be-processed data table; obtaining the number of the dimension fields in the to-be-processed data table that have a parent-child relationship with the first dimension field according to the relationship between the first dimension field and other dimension fields in the to-be-processed data table; obtaining the number of the dimension fields in the to-be-processed data table that have a parent-child relationship with the first dimension field according to the relationship between the first dimension field and other dimension fields in the to-be-processed data table; obtaining the number of the dimension fields in the to-be-processed data table that have a parent-child relationship with the first dimension field according to the relationship between the first dimension field and other dimension fields in the to-be-processed data table; obtaining the number of the dimension fields in the to-be-processed data table that have a parent-child relationship with the first dimension field according to the relationship between the first dimension field and other dimension fields in the to-be-processed data table; obtaining the number of the dimension fields in the to-be-processed data table that have a parent-child relationship with the first dimension field according to the relationship between the first dimension field and other dimension fields in the to-be-processed data table; 6. The data table field mapping generation method of claim 1, wherein, obtaining the enumeration value of the field type; obtaining the second feature information of the first dimension field in the to-be-processed data table according to the obtained data. The classification of the dimension field in the to-be-processed data table according to the type of the dimension field and the recommendation score of the dimension field includes: in a case where the type of the first dimension field satisfies a type-class mapping relationship in the type-class mapping relationship set, determining the class corresponding to the first dimension field; wherein the type-class mapping relationship describes a type and a class that have a unique corresponding relationship; the first dimension field is any one of the dimension fields in the to-be-processed data table; in a case where the type of the first dimension field does not satisfy a type-class mapping relationship in the type-class mapping relationship set, determining the class corresponding to the first dimension field according to the relevance between the first dimension field and the dimension field in the to-be-processed data table whose class has been determined and the recommendation score of the first dimension field.
7. The data table field mapping generation method according to claim 6, characterized by, The type-category mapping relationship set includes one or more of the following type-category mapping relationships: a date type corresponds to a time category, a time type corresponds to a time category, a date-time type corresponds to a time category, a place name type corresponds to a place category, a person name type corresponds to a person category, an organization group type corresponds to a person category, a mobile phone number type corresponds to a person category, a telephone number type corresponds to a person category, a text class numerical value type corresponds to a person category, an ID number type corresponds to a person category, a verb type corresponds to an event category; Correspondingly, the type-category mapping relationship set includes one or more of the following type-category mapping relationships: a date type corresponds to a time category, a time type corresponds to a time category, a date-time type corresponds to a time category, a place name type corresponds to a place category, a person name type corresponds to a person category, an organization group type corresponds to a person category, a mobile phone number type corresponds to a person category, a telephone number type corresponds to a person category, a text class numerical value type corresponds to a person category, an ID number type corresponds to a person category, a verb type corresponds to an event category; According to the type of the first dimension field, a corresponding type-category mapping relationship is determined in the type-category mapping relationship set; According to the determined type-category mapping relationship, the category of the first dimension field is determined.
8. The data table field mapping generation method according to claim 7, characterized by, In the case that the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, the category corresponding to the first dimension field is determined according to the relevance between the first dimension field and the dimension field with a determined category in the to-be-processed data table and the recommendation score of the first dimension field, including: determine whether the type of the first dimension field is a noun type; In the case that the type of the first dimension field is a noun type, the relevance between the first dimension field and each dimension field determined as a person category is calculated, and the relevance between the first dimension field and each dimension field determined as a place category is calculated; According to the comparison result of the first relevance and the second relevance, the category of the first dimension field is determined; wherein the first relevance is the maximum value of the relevance between the first dimension field and each dimension field determined as a person category, and the second relevance is the maximum value of the relevance between the first dimension field and each dimension field determined as a place category.
9. The data table field mapping generation method according to claim 8, characterized by, According to the comparison result of the first relevance and the second relevance, the category of the first dimension field is determined, including: In the case that the value of the first relevance is greater than the first threshold value and the first relevance is greater than the second relevance, the category of the first dimension field is determined as a person category; In the case that the value of the second relevance is greater than the first threshold value and the second relevance is greater than the first relevance, the category of the first dimension field is determined as a place category.
10. The data table field mapping generation method of claim 7, wherein, In the case that the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, the category corresponding to the first dimension field is determined according to the relevance between the first dimension field and the dimension field with a determined category in the to-be-processed data table and the recommendation score of the first dimension field, including: In the case where the type of the first dimension field is not a noun type, the correlation between the first dimension field and each dimension field determined as a person category is calculated, the correlation between the first dimension field and each dimension field determined as a place category is calculated, the correlation between the first dimension field and each dimension field determined as a time category is calculated, and; In the case where the first correlation is greater than a second threshold value, the category of the first dimension field is determined as a person category; In the case where the second correlation is greater than a third threshold value, the category of the first dimension field is determined as a place category; In the case where the third correlation is greater than a fourth threshold value, the category of the first dimension field is determined as a time category; In the case where the category of the first dimension field is not identified as the person category, the place category, or the time category, the category of the first dimension field is identified as an event category; The first correlation is the maximum value of the correlation between the first dimension field and each dimension field determined as a person category; the second correlation is the maximum value of the correlation between the first dimension field and each dimension field determined as a place category; and the third correlation is the maximum value of the correlation between the first dimension field and each dimension field determined as a time category.
11. The data table field mapping generation method according to claim 10, characterized by, In the case where the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, the category corresponding to the first dimension field is determined according to the correlation between the first dimension field and the dimension field whose category has been determined in the to-be-processed data table and the recommendation score of the first dimension field, which further comprises: Obtaining the recommendation score of the first dimension field of the event category; In the case where the recommendation score is less than or equal to a fifth threshold value, the category of the first dimension field is modified from the event category to the non-recommended event category.
12. The data table field mapping generation method of claim 1, wherein, The category of the first dimension field is determined as a measure category, the feature information of the to-be-processed data table is obtained, wherein the feature information of the to-be-processed data table comprises the description feature of the to-be-processed data table and the classification feature of the to-be-processed data table, the first feature information of the first measure field is obtained according to the data distribution in the cell set corresponding to the first measure field in the to-be-processed data table and the distribution of the first measure field in the to-be-processed data table, the second feature information of the first measure field in the to-be-processed data table is obtained according to the data statistics in the cell set corresponding to the first measure field in the to-be-processed data table, and the recommendation score of the first measure field is obtained by inputting the feature information of the to-be-processed data table, the first feature information of the first measure field, and the second feature information of the first measure field into a pre-trained measure field recommendation model. The category of the first dimension field is determined as a measure category, the feature information of the to-be-processed data table is obtained, wherein the feature information of the to-be-processed data table comprises the description feature of the to-be-processed data table and the classification feature of the to-be-processed data table, the first feature information of the first measure field is obtained according to the data distribution in the cell set corresponding to the first measure field in the to-be-processed data table and the distribution of the first measure field in the to-be-processed data table, the second feature information of the first measure field in the to-be-processed data table is obtained according to the data statistics in the cell set corresponding to the first measure field in the to-be-processed data table, and the recommendation score of the first measure field is obtained by inputting the feature information of the to-be-processed data table, the first feature information of the first measure field, and the second feature information of the first measure field into a pre-trained measure field recommendation model. The metric field recommendation model is trained based on feature information of a sample data table, first feature information of a second metric field in the sample data table, second feature information of the second metric field, and a recommendation score label of the second metric field.
13. The data table field mapping generation method of claim 12, wherein, The first feature information of the first metric field in the to-be-processed data table is obtained according to a data distribution in a cell set corresponding to the first metric field in the to-be-processed data table and a distribution of the first metric field in the to-be-processed data table, and at least includes any one of the following: An index value of the first metric field is obtained, where the index value is used to describe a position of the first metric field in the to-be-processed data table. A number of non-repeated data in the cell set corresponding to the first metric field is obtained. A number of cells in the cell set corresponding to the first metric field is obtained. A number of non-empty cells in the cell set corresponding to the first metric field is obtained. In the to-be-processed data table, an index value of the first metric field is compared with index values of other metric fields to obtain a number of other metric fields with an index value smaller than that of the first metric field and a number of other metric fields with an index value greater than that of the first metric field. In the to-be-processed data table, an index value of the first metric field is compared with index values of dimension fields to obtain a number of dimension fields with an index value smaller than that of the first metric field and a number of dimension fields with an index value greater than that of the first metric field. The first feature information of the first metric field in the to-be-processed data table is obtained according to the obtained data.
14. The data table field mapping generation method of claim 12, wherein, The second feature information of the first metric field in the to-be-processed data table is obtained according to a data statistical situation in the cell set corresponding to the first metric field in the to-be-processed data table, and at least includes any one of the following: An average value of numbers in the cell set corresponding to the first metric field is obtained. A median of numbers in the cell set corresponding to the first metric field is obtained. A standard deviation of numbers in the cell set corresponding to the first metric field is obtained. A minimum value of numbers in the cell set corresponding to the first metric field is obtained. A maximum value of numbers in the cell set corresponding to the first metric field is obtained. A three-quarter quantile value of numbers in the cell set corresponding to the first metric field is obtained. A one-quarter quantile value of numbers in the cell set corresponding to the first metric field is obtained. A difference between the one-quarter quantile value and the three-quarter quantile value of the numbers in the cell set corresponding to the first metric field is obtained.
15. A data table field map generation apparatus characterized by comprising: The first determining module is configured to determine types of fields in a to-be-processed data table, determine dimension fields and metric fields in the to-be-processed data table according to the types of the fields, and determine the types of the fields in the to-be-processed data table. The second determining module is configured to determine a recommendation score of the dimension field according to the type of the dimension field, the feature information of the dimension field, and the feature information of the to-be-processed data table; The classifying module is configured to classify the dimension fields in the to-be-processed data table according to the type of the dimension field and the recommendation score of the dimension field, to obtain a category of the dimension field; The third determining module is configured to determine a category of the metric field, and determine a recommendation score of the metric field according to the feature information of the to-be-processed data table and the feature information of the metric field; wherein the recommendation score of the metric field is obtained by a pre-trained metric field recommendation model based on the feature information of the to-be-processed data table, first feature information of a first metric field, and second feature information of the first metric field; the first metric field is any one of the metric fields in the to-be-processed data table; the first feature information is obtained according to a data distribution in a cell set corresponding to the first metric field in the to-be-processed data table and a distribution of the first metric field in the to-be-processed data table; and the second feature information is obtained according to a data statistical condition in the cell set corresponding to the first metric field in the to-be-processed data table; The generating module is configured to generate a field map of the to-be-processed data table according to the category of each dimension field in the to-be-processed data table, the recommendation score of each dimension field, the category of each metric field, and the recommendation score of each metric field.
16. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, The processor executes the program to implement the steps of the data table field map generation method in any one of claims 1 to 14.
17. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the data table field map generation method in any one of claims 1 to 14.
Citation Information
Patent Citations
Chart recommendation method, device and electronic equipment
CN110489449A
Method and device for assisting user in exploring data set and data table
CN110874644A