Pivot table visualization processing method and apparatus, device, and medium
By generating a source data table of the visual analysis results of the PivotTable and determining the recommended chart types and chart conclusions, the problem of high threshold for PivotTable visualization processing is solved, and automated and efficient visual analysis is achieved.
Patent Information
- Application Number
- CN202111223229.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-20
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2041-10-20
AI Technical Summary
In the existing technology, the visualization processing threshold of PivotTable is relatively high, and users need to analyze and organize the tables themselves for visualization, which increases labor costs.
This paper provides a visualization processing method for pivot tables. By obtaining the values and row and column fields of the pivot table, a data cross-tab is generated. The source data table of the visualization analysis results is generated based on the cross-tab, the recommended chart type and chart conclusion are determined, and the visualization analysis results are finally output.
It realizes the automated visual analysis of pivot tables, reduces labor costs and improves analysis efficiency.
Smart Images

Figure CN114064637B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and more particularly, to a method and device for visualizing a pivot table, an electronic device and a readable storage medium. BACKGROUND
[0002] A pivot table is an interactive table that can perform certain calculations, such as sum and count. The calculations performed are related to the arrangement of data in the pivot table. They are called pivot tables because they can be dynamically rearranged to analyze data in different ways, and row numbers, column labels, and page fields can be rearranged. Each time the layout is changed, the pivot table will immediately recalculate the data according to the new layout. In addition, if the original data changes, the pivot table can be updated.
[0003] In the prior art, the visualization processing of the pivot table has a high threshold, and most of the pivot tables cannot directly perform data analysis and visual display. Users can only analyze and organize new tables according to the table data of the pivot table for visual display, which requires high self-business ability of the users and increases the labor cost. SUMMARY
[0004] An object of the present disclosure is to provide a new technical solution for automatically processing a pivot table to obtain a visual analysis result of the pivot table.
[0005] According to a first aspect of the present disclosure, a method for visualizing a pivot table is provided, comprising:
[0006] obtaining a pivot table, the pivot table having one numerical field and at least two row and column fields;
[0007] generating a corresponding data cross table according to the row and column fields;
[0008] generating a source data table of a visual analysis result of the pivot table according to the data cross table;
[0009] determining a recommended chart type corresponding to the source data table according to table data of the source data table;
[0010] generating a chart conclusion of the source data table according to a first quantity value of a series in the source data table;
[0011] outputting the visual analysis result of the pivot table according to the source data table, the recommended chart type and the chart conclusion.
[0012] Optionally, in the case that the first number of series in the source data table is one, the generating the chart conclusion of the source data table according to the first number of series in the source data table comprises:
[0013] Obtaining a data sequence in the source data table;
[0014] In the case that the category values corresponding to the data items of the data sequence represent time and there is a sequential relationship between the category values corresponding to the data items of the data sequence, determining that the data sequence is a first target data sequence, and the first target data sequence is in time sequence;
[0015] Obtaining a first correlation, the first correlation being a correlation between the data items of the first target data sequence and the category values corresponding to the data items of the first target data sequence;
[0016] According to the first correlation, outputting a chart processing result related to the first correlation;
[0017] According to the chart processing result related to the first correlation, obtaining the chart conclusion of the source data table;
[0018] In the case that the category values corresponding to the data items of the data sequence do not represent time or there is no sequential relationship between the category values corresponding to the data items of the data sequence, generating the chart conclusion of the source data table according to the recommended chart type.
[0019] Optionally, in the case that the recommended chart type is a pie chart,
[0020] The generating the chart conclusion of the source data table according to the recommended chart type further comprises:
[0021] In the case that the proportion of the data items corresponding to one category value of the source data table is greater than or equal to a first threshold value, determining that the chart conclusion of the source data table represents the proportion of the data items corresponding to the category value;
[0022] In the case that the sum of the proportions of the data items corresponding to two category values of the source data table is greater than or equal to the first threshold value, determining that the chart conclusion of the source data table represents the two category values corresponding to the data items whose proportion exceeds the first threshold value and the sum of the proportions of the data items corresponding to the two category values;
[0023] In the case that the sum of the proportions of the data items corresponding to at least three category values of the source data table is greater than or equal to the first threshold value, determining that the chart conclusion of the source data table represents the proportion of the data items corresponding to each category value of the source data table.
[0024] Optionally, in the case that the recommended chart type is a column chart,
[0025] The generating the chart conclusion of the source data table according to the recommended chart type further includes:
[0026] In a case where the second quantity value of the category values of the source data table is greater than or equal to the second threshold value, and the proportion of the data items corresponding to one category value is greater than or equal to the third threshold value or the proportion of the data items corresponding to one category value is greater than the average value by a first multiple, it is determined that the chart conclusion of the source data table indicates that the category value ranks first; wherein the average value is an average value of the data items corresponding to all category values.
[0027] In a case where the second quantity value of the category values of the source data table is greater than the second threshold value, and the proportions of the data items corresponding to two category values are greater than or equal to the third threshold value, it is determined that the chart conclusion of the source data table indicates the two category values corresponding to the data items whose proportions exceed the third threshold value, and the sum of the proportions of the data items corresponding to the two category values.
[0028] In a case where the second quantity value of the category values of the source data table is greater than or equal to the second threshold value, and the proportion of the data items corresponding to one category value is less than the average value by a second multiple, it is determined that the chart conclusion of the source data table indicates that the category value is the lowest.
[0029] In other cases, it is determined that the chart conclusion of the source data table indicates the distribution of the data items corresponding to each category value of the source data table.
[0030] Optionally, in a case where the recommended chart type is a line chart,
[0031] The generating the chart conclusion of the source data table according to the recommended chart type further includes:
[0032] It is determined that the chart conclusion of the source data table indicates the distribution of the data items corresponding to each category value of the source data table.
[0033] Optionally, in a case where the first quantity value of the series in the source data table is at least two, the generating the chart conclusion of the source data table according to the first quantity value of the series in the source data table includes:
[0034] The third correlation of the data items of all series in the source data table is calculated.
[0035] The chart conclusion of the source data table is generated according to the third correlation.
[0036] Optionally, in a case where the first quantity value of the series in the source data table is two,
[0037] The generating the chart conclusion of the source data table according to the third correlation includes:
[0038] in a case where the third correlation of the data items of the two series in the source data table is greater than a fourth threshold value, obtaining a series with a greater sum of data items, and determining that the chart conclusion of the source data table indicates that the series is generally greater;
[0039] in a case where the third correlation of the data items of the two series in the source data table is less than or equal to the fourth threshold value, obtaining a difference of the data items of the two series at each category value, and obtaining a category value with a greatest difference, and determining that the chart conclusion of the source data table indicates that the data items of the two series are inconsistent in distribution, where the category value has the greatest difference.
[0040] Optionally, in a case where the first number of series in the source data table is at least three,
[0041] the generating the chart conclusion of the source data table according to the third correlation comprises:
[0042] in a case where the third correlation of the data items of all series in the source data table is greater than a fourth threshold value, obtaining a series with a greater sum of data items, and determining that the chart conclusion of the source data table indicates that the series is generally greater;
[0043] in a case where the third correlation of the data items of all series in the source data table is less than or equal to the fourth threshold value, determining that the chart conclusion of the source data table indicates a distribution of the data items of each series in the source data table.
[0044] Optionally, the at least two row and column fields comprise at least one row field and at least one column field.
[0045] the generating the corresponding data cross table according to the row and column fields comprises:
[0046] traversing the row fields, and combining the traversed row fields, the numerical value field, and any column field to form at least one first field combination;
[0047] obtaining a data cross table corresponding to the first field combination according to the first field combination and the numerical values in the pivot table.
[0048] Optionally, the generating the source data table of the visual analysis result of the pivot table according to the data cross table further comprises:
[0049] traversing the data cross table;
[0050] performing analysis of variance on the traversed data cross table according to a target direction to obtain a corresponding F value, where the target direction comprises a row direction and / or a column direction.
[0051] When the F value is less than a first preset threshold, data in the data cross table is processed according to the target direction to obtain a first data table;
[0052] The first data table is taken as the source data table.
[0053] Optionally, the source data table of the visual analysis result of the pivot table generated according to the data cross table further includes:
[0054] When the F value is less than a first preset threshold, it is judged whether a proportion of blank cells in the data cross table is less than a second preset threshold;
[0055] When the proportion of blank cells in the data cross table is less than the second preset threshold, a correlation coefficient matrix value of the data cross table in the target direction is calculated.
[0056] When any correlation coefficient matrix value is greater than a third preset threshold, the data cross table is taken as the source data table.
[0057] Optionally, when the proportion of blank cells in the data cross table is less than the second preset threshold, the method further includes:
[0058] A Mahalanobis distance value of data in each data cell in the data cross table is calculated; in the case that the target direction is a row direction, one data cell represents a row of the data cross table; in the case that the target direction is a column direction, one data cell represents a column of the data cross table.
[0059] A degree of freedom value of the data cross table in the target direction is determined.
[0060] According to the Mahalanobis distance value and the degree of freedom value, a P value of data in each data cell in the data cross table is determined by using chi-square test.
[0061] According to a percentage comparison result of data in a data cell with a P value less than a fourth preset threshold and data in all data cells in the data cross table, a second data table is generated.
[0062] The second data table is taken as the source data table.
[0063] Optionally, the row and column field only includes a row field and does not include a column field.
[0064] The data cross table corresponding to the row and column field is generated according to the row and column field, including:
[0065] traversing the row fields, combining the traversed row fields and the value fields to form a second field combination;
[0066] obtaining a data cross table corresponding to the second field combination according to the second field combination and the values in the pivot table.
[0067] Optionally, the method further comprises:
[0068] traversing the data cross table, and judging whether a type of the row field corresponding to the traversed data cross table is a date type;
[0069] when the type of the row field corresponding to the traversed data cross table is the date type, reordering rows in the traversed data cross table according to a time sequence to obtain the source data table;
[0070] when the type of the row field corresponding to the traversed data cross table is not the date type, reordering rows in the traversed data cross table according to values of the value fields to obtain the source data table.
[0071] Optionally, when there are N first target time data sequences and there is a time sequence relationship between the N first target data sequences, N is an integer and N≥2, the method further comprises:
[0072] processing a data item of an nth first target data sequence to obtain target data corresponding to the nth first target data sequence, the nth first target data sequence being any one of the N first target data sequences;
[0073] assigning the target data corresponding to the N first target data sequences as data items to category values according to the time sequence relationship between the N first target data sequences to construct a second target data sequence in time sequence;
[0074] obtaining a second correlation between the data items of the second target data sequence and the category values corresponding to the data items of the second target data sequence;
[0075] outputting a chart processing result corresponding to the second correlation according to the second correlation;
[0076] obtaining a chart conclusion of the source data table according to the chart processing result corresponding to the second correlation.
[0077] Optionally, the outputting the chart processing result corresponding to the first correlation according to the first correlation comprises:
[0078] determine a relationship between two data items adjacent in a sequence direction in the first target data sequence, in a case where the first correlation indicates that the data items of the first target data sequence and the corresponding category values thereof are positively correlated;
[0079] generate a chart processing result corresponding to the first correlation according to the relationship between the two data items adjacent in the sequence direction in the first target data sequence.
[0080] Optionally, the generating the chart processing result corresponding to the first correlation according to the relationship between the two data items adjacent in the sequence direction in the first target data sequence comprises:
[0081] calculating a growth rate of a latter data item relative to a former data item for the two data items adjacent in the first target data sequence;
[0082] determining whether the latter data item is negatively growing according to the growth rate of the latter data item relative to the former data item;
[0083] generating the chart processing result corresponding to the first correlation according to a quantity value of the data items negatively growing in the first target data sequence and a position of the data items negatively growing in the first target data sequence.
[0084] Optionally, the outputting the chart processing result corresponding to the first correlation according to the first correlation comprises:
[0085] determine a relationship between two data items adjacent in a sequence direction in the first target data sequence, in a case where the first correlation indicates that the data items of the first target data sequence and the corresponding category values thereof are negatively correlated;
[0086] generate a chart processing result corresponding to the first correlation according to the relationship between the two data items adjacent in the sequence direction in the first target data sequence.
[0087] Optionally, the generating the chart processing result corresponding to the first correlation according to the relationship between the two data items adjacent in the sequence direction in the first target data sequence comprises:
[0088] calculating a growth rate of a latter data item relative to a former data item for the two data items adjacent in the first target data sequence;
[0089] determining whether the latter data item is positively growing according to the growth rate of the latter data item relative to the former data item;
[0090] According to the number of data items that are growing positively in the first target data sequence and the positions of the data items that are growing positively in the first target data sequence, a chart processing result corresponding to the first correlation is generated.
[0091] Optionally, the outputting of the chart processing result corresponding to the first correlation comprises:
[0092] In a case where the first correlation indicates that the data items of the first target data sequence and their corresponding category values are irrelevant, determining an average value of the data items of the first target data sequence and a maximum value and a minimum value among the data items of the first target data sequence;
[0093] determining a first difference value and a second difference value, the first difference value being a difference between the maximum value and the average value, and the second difference value being a difference between the average value and the minimum value;
[0094] In a case where the first difference value is greater than the second difference value, outputting a category value corresponding to the maximum value;
[0095] In a case where the first difference value is less than the second difference value, outputting a category value corresponding to the minimum value.
[0096] Optionally, the method further comprises:
[0097] For two adjacent data items in the first target data sequence, calculating an increase rate of a latter data item relative to a former data item;
[0098] According to the increase rate of the latter data item relative to the former data item, determining whether the latter data item is growing positively or negatively;
[0099] According to the category value of the data items that are growing positively and the category value of the data items that are growing negatively, outputting a corresponding chart processing result;
[0100] Further according to the corresponding chart processing result, obtaining a chart conclusion of the source data table.
[0101] According to a second aspect of the present disclosure, a data visualization processing device is provided, comprising:
[0102] a pivot table obtaining module configured to obtain a data pivot table, the data pivot table having one numerical field and at least two row and column fields;
[0103] a cross table generating module configured to generate a corresponding data cross table according to the row and column fields;
[0104] a source data table generating module configured to generate a source data table of a visual analysis result of the data pivot table according to the data cross table.
[0105] a chart type determining module, configured to determine a recommended chart type corresponding to the source data table according to table data of the source data table;
[0106] a chart conclusion generating module, configured to generate a chart conclusion of the source data table according to a first quantity value of series in the source data table;
[0107] an analysis result outputting module, configured to output a visual analysis result of the pivot table according to the source data table, the recommended chart type and the chart conclusion.
[0108] According to a third aspect of the present disclosure, an electronic device is provided, comprising:
[0109] the apparatus according to the second aspect of the present disclosure; or
[0110] a processor and a memory, the memory being configured to store instructions for controlling the processor to perform the method according to the first aspect of the present disclosure.
[0111] According to a fourth aspect of the present disclosure, a readable storage medium is provided, which stores a computer program, the computer program being configured to implement the method according to the first aspect of the present disclosure when executed by a processor.
[0112] In the embodiments of the present disclosure, for a pivot table having one numerical field and at least two row and column fields, a corresponding data cross table is generated according to the row and column fields, a source data table of a visual analysis result of the pivot table is generated according to the data cross table, a recommended chart type corresponding to the source data table is determined according to table data of the source data table, a chart conclusion of the source data table is generated according to a first quantity value of series in the source data table, and the visual analysis result of the pivot table is output according to the source data table, the recommended chart type and the chart conclusion, so that the pivot table can be automatically and accurately visualized analyzed, and the labor cost of visual analysis of the pivot table can be reduced.
[0113] Other features of the present disclosure, and its advantages, will become apparent in the following detailed description of exemplary embodiments of the present disclosure, with reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS
[0114] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0115] Figure 1 is a block diagram of one example of a hardware configuration of an electronic device that can be used to implement embodiments of the present disclosure.
[0116] Figure 2A flow chart of a method of visualizing a pivot table is shown according to one embodiment of the present disclosure.
[0117] Figure 3 A schematic diagram of a pivot table is shown according to one embodiment of the present disclosure.
[0118] Figure 4 and Figure 5 A schematic diagram of a source data table is shown according to one embodiment of the present disclosure.
[0119] Figure 6 and Figure 7 A schematic diagram of another source data table is shown according to one embodiment of the present disclosure.
[0120] Figure 8 A block diagram of a device for visualizing a pivot table is shown according to one embodiment of the present disclosure.
[0121] Figure 9 A block diagram of an electronic device is shown according to one embodiment of the present disclosure. DETAILED DESCRIPTION
[0122] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of the components and steps set forth in the embodiments, the numerical expressions, and the numerical values are not limiting to the scope of the present disclosure unless otherwise specifically stated.
[0123] The following description of at least one exemplary embodiment is merely exemplary in nature and is in no way intended to limit the scope of the present disclosure, its application, or uses.
[0124] Techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail herein. However, the techniques, methods, and devices should be considered part of the specification, if appropriate.
[0125] In all of the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments can have different values.
[0126] It should be noted that like references and characters herein relate to like items throughout the figures, and once an item is defined in one figure, it need not be discussed further in subsequent figures.
[0127] <Hardware Configuration>
[0128] Figure 1 is a structural schematic diagram of an electronic device that can be used to implement embodiments of the present disclosure.
[0129] The electronic device 1000 can be a smart phone, a portable computer, a desktop computer, a tablet computer, a server, etc., which is not limited herein.
[0130] The electronic device 1000 can include, but is not limited to, a processor 1100, a memory 1200, an interface device 1300, a communication device 1400, a display device 1500, an input device 1600, a speaker 1700, a microphone 1800, etc. The processor 1100 can be a central processing unit CPU, a graphics processing unit GPU, a microprocessor MCU, etc., configured to execute a computer program, which can be written in an instruction set of an architecture such as x86, Arm, RISC, MIPS, SSE, etc. The memory 1200 can include, for example, a ROM (read only memory), a RAM (random access memory), a non-volatile memory such as a hard disk, etc. The interface device 1300 can include, for example, a USB interface, a serial interface, a parallel interface, etc. The communication device 1400 can be configured to perform wired communication using an optical fiber or a cable, or wireless communication, and can include, for example, WiFi communication, Bluetooth communication, 2G / 3G / 4G / 5G communication, etc. The display device 1500 can be, for example, a liquid crystal display screen, a touch display screen, etc. The input device 1600 can include, for example, a touch screen, a keyboard, a body-sensing input, etc. The speaker 1700 is configured to output an audio signal. The microphone 1800 is configured to acquire an audio signal.
[0131] In the embodiments of the present disclosure, the memory 1200 of the electronic device 1000 is configured to store a computer program for controlling the processor 1100 to operate to implement the method according to the embodiments of the present disclosure. The computer program can be designed by a skilled person according to the solutions disclosed in the present disclosure. How the computer program controls the processor to operate is known in the art, and thus is not described in detail herein. The electronic device 1000 can be installed with a smart operating system (such as Windows, Linux, Android, IOS, etc.) and application software.
[0132] Those skilled in the art should understand that, although a plurality of devices of the electronic device 1000 are shown in Figure 1 the embodiments of the present disclosure, the electronic device 1000 of the embodiments of the present disclosure can only involve part of the devices, for example, only the processor 1100 and the memory 1200, etc.
[0133] In the following, various embodiments and examples according to the present disclosure are described with reference to the accompanying drawings.
[0134] In embodiments of the present disclosure, pivot table is an interactive table that can perform certain calculations, such as sum and count. The calculation performed is related to the arrangement of data in the pivot table. The reason why it is called pivot table is that they can dynamically change their layout to analyze data in different ways, and also rearrange row numbers, column labels and page fields. Each time the layout is changed, the pivot table will immediately recalculate the data according to the new arrangement. In addition, if the original data changes, the pivot table can be updated.
[0135] While the ordinary data table is a grid virtual table (a table representing data in memory) that temporarily saves data.
[0136] Data visualization refers to the process of representing data in large data sets in graphical images and discovering unknown information in them using data analysis and development tools.
[0137] Because the data structure of pivot table is more complex than that of ordinary data table, the existing visualization tools cannot directly visualize the pivot table, that is, the visualization processing method of ordinary data table cannot be applied to pivot table.
[0138] <Method embodiment>
[0139] In this embodiment, a visualization processing method of pivot table is provided. The method is implemented by an electronic device. The electronic device can be an electronic product with a processor and a memory. For example, it can be a desktop computer, a notebook computer, a mobile phone, a tablet computer, etc. In an example, the electronic device can be an electronic device 1000 as shown in Figure 1
[0140] Figure 2 A schematic flowchart of the visualization processing method of pivot table of an embodiment is shown. As shown in Figure 2 The visualization processing method of pivot table of the present embodiment includes the following steps S2100-S2600:
[0141] Step S2100, obtaining a pivot table, the pivot table having a numerical field and at least two row and column fields.
[0142] In this embodiment, the at least two row and column fields can include at least one row field and at least one column field, or at least two row fields and no column field.
[0143] In the pivot table as shown in Figure 3 The pivot table can include a numerical field, two row fields and a column field, wherein the numerical field can include sales quantity value, the row field can include office and store, and the column field can include brand.
[0144] Step S2200: Generate a corresponding data cross table based on the row and column fields.
[0145] A crosstab is a table in matrix format that shows the frequency distribution of variables.
[0146] In an embodiment where the at least two row and column fields include at least one row field and at least one column field, generating a corresponding data cross-tab according to the row and column fields may include steps S2211 to S2212 as follows:
[0147] Step S2211: traverse the row fields, and combine the traversed row fields, value fields, and any column fields to form at least one first field combination.
[0148] In this embodiment, a first field combination may include a value field, a row field, and a column field.
[0149] For example, in Figure 3 In the PivotTable shown, two first field combinations can be obtained. The first first field combination can include sales quantity values, offices, and brands, and the second first field combination can include sales quantity values, stores, and brands.
[0150] Step S2212: Obtain a data crosstab corresponding to the first field combination based on the first field combination and the values in the pivot table.
[0151] In such Figure 3 In the example shown, based on the first first field combination and the values in the pivot table, the resulting data crosstab can be as shown in Table 1 below, and based on the second first field combination and the values in the pivot table, the resulting data crosstab can be as shown in Table 2 below.
[0152] Table 1
[0153] Brand Sales Quantity Value Office Xshang TCx Chuangx Kangx Xin Xer Changx Wanax Office NaN 17 11 10 11 NaN 11 Guangx Office NaN NaN 2 NaN 1 NaN NaN Fuxu Office NaN 5 4 7 15 0 11 Yuxu Office NaN 49 28 21 52 7 37 Chongxu Office 1 50 53 23 43 11 49
[0154] Table 2
[0155]
[0156]
[0157]
[0158] NaN in Tables 1 and 2 above indicates that the sales quantity value is empty.
[0159] In the embodiment in which the at least two row-column fields include at least two row fields and no column field, generating the corresponding data cross table according to the row-column field can include the following steps S2221-S2222:
[0160] Step S2221: traversing the row fields, combining the traversed row fields and the value field to form a second field combination.
[0161] In this embodiment, one second field combination can include the value field and one row field.
[0162] Step S2222: obtaining the data cross table corresponding to the second field combination according to the second field combination and the values in the pivot table.
[0163] Step S2300: generating a source data table of the visual analysis result of the pivot table according to the data cross table.
[0164] The source data table of the visual analysis result of the pivot table is a data table used for visual analysis of the pivot table. Since the data structure of the pivot table is more complex than that of the ordinary data table, the visual processing method of the ordinary data table cannot be applied to the pivot table, and therefore, this embodiment can first generate a source data table used for visual analysis of the pivot table, and then perform visual analysis on the source data table to obtain the visual analysis result of the pivot table.
[0165] The visual analysis result can be a result of clearly and effectively conveying the data content of the pivot table by means of graphical means. In one example, the visual analysis result can be a result obtained by performing visual analysis on the pivot table.
[0166] In the embodiment in which the at least two row-column fields include at least one row field and at least one column field, generating the source data table of the visual analysis result of the pivot table according to the data cross table can include the following steps S2311-S2314:
[0167] Step S2311: traversing the data cross table.
[0168] Step S2312: performing variance analysis on the traversed data cross table according to a target direction to obtain a corresponding F value.
[0169] The target direction includes a row direction and / or a column direction.
[0170] In this embodiment, the row direction can be taken as the target direction, and step S2300 of this embodiment can be executed once; and / or, the column direction can be taken as the target direction, and step S2300 of this embodiment can be executed once.
[0171] Analysis of Variance (ANOVA), also known as "variance analysis", is used for significant test of difference between two or more sample means.
[0172] By performing the analysis of variance on the traversed data cross table according to the target direction, the F value can reflect the difference between the means of the traversed data cross table in the target direction. For example, in the case where the target direction is the row direction, the F value can reflect the difference between the means of the values of all rows of the traversed data cross table. For another example, in the case where the target direction is the column direction, the F value can reflect the difference between the means of the values of all columns of the traversed data cross table.
[0173] In step S2313, when the F value is less than the first preset threshold, the values in the traversed data cross table are processed according to the target direction to obtain a first data table.
[0174] In this embodiment, the first preset threshold can be set in advance according to application scenarios or specific requirements. For example, the first preset threshold can be 0.5.
[0175] Processing the values in the data cross table according to the target direction can be processing the values in the traversed data cross table according to the value field and the row and column field corresponding to the target direction, that is, obtaining the first data table.
[0176] In the case where the traversed cross data table is the above table 1, the target direction is the row direction, and the F value of the cross data table in the row direction is less than the first preset threshold, the value field of the traversed cross data table is the sales quantity value, and the row and column field corresponding to the row direction is the office. Therefore, the values of the sales quantity value corresponding to each office in the traversed data cross table can be summed up according to the sales quantity value and the office, and the obtained first data table can be as shown in the following table 3.
[0177] Table 3
[0178] Office Chongxu Office Yuxu Office Wanax Office Fuxu Office Guangx Office Sales Quantity Value 230 194 60 42 3
[0179] In the case where the traversed cross data table is the above table 1, the target direction is the row direction, and the F value of the cross data table in the row direction is less than the first preset threshold, the value field of the traversed cross data table is the sales quantity value, and the row and column field corresponding to the row direction is the office. Therefore, the values of the sales quantity value corresponding to each office in the traversed data cross table can be summed up according to the sales quantity value and the office, and the obtained first data table can be as shown in the following table 3.
[0180] In the case that the traversed cross-data table is the above table 1, the target direction is the column direction, and the F value of the cross-data table in the column direction is less than the first preset threshold value, the numerical field of the traversed cross-data table is the sales quantity value, and the row-column field corresponding to the column direction is the brand. Therefore, the sales quantity value corresponding to each brand in the traversed cross-data table can be summed to obtain a first data table as shown in the following table 4.
[0181] Table 4
[0182] Brand Xin TCx Changx Chuangx Kangx Xer Xshang Sales Quantity Value 122 121 108 98 61 18 1
[0183] In the case that the traversed cross-data table is the above table 2, the target direction is the row direction, and the F value of the cross-data table in the row direction is less than the first preset threshold value, the numerical field of the traversed cross-data table is the sales quantity value, and the row-column field corresponding to the row direction is the store. Therefore, the sales quantity value corresponding to each store in the traversed cross-data table can be summed to obtain a first data table as shown in the following table 5.
[0184] Table 5
[0185]
[0186]
[0187]
[0188] In step S2314, the first data table is taken as a source data table.
[0189] Further, the source data table of the visual analysis result of the pivot table generated according to the cross-data table can further include the following steps S2315-S2317.
[0190] In step S2315, when the F value is less than the first threshold value, it is judged whether the proportion value of the blank cells in the traversed cross-data table is less than a second preset threshold value.
[0191] The blank cells can be cells with the value of NaN in the cross-data table.
[0192] The proportion value of the blank cells can be the ratio of a first quantity value and a second quantity value, wherein the first quantity value is the quantity value of the cells with the value of NaN in the traversed cross-data table, and the second quantity value is the quantity value of the cells with the value in the traversed cross-data table.
[0193] The second preset threshold value can be set according to the application scenario or specific requirements in advance. For example, the second preset threshold value can be 50%.
[0194] Step S2316, when the proportion of the blank cells in the traversed data cross table is less than the second preset threshold, calculating the correlation coefficient matrix value of the traversed data cross table in the target direction.
[0195] In the embodiment, the value matrix can be constructed according to the values in the traversed data cross table.
[0196] In the case where the target direction is the column direction, the correlation coefficient matrix can be composed of the correlation coefficients between the columns of the value matrix. That is, the element in the i-th row and the j-th column of the correlation coefficient matrix is the correlation coefficient between the i-th column and the j-th column of the value matrix. In the case where the target direction is the row direction, the correlation coefficient matrix can be composed of the correlation coefficients between the columns of the transpose matrix of the value matrix. That is, the element in the i-th row and the j-th column of the correlation coefficient matrix is the correlation coefficient between the i-th column and the j-th column of the transpose matrix of the value matrix.
[0197] The correlation coefficient matrix value can be the value of each element in the correlation coefficient matrix.
[0198] Step S2317, when any correlation coefficient matrix value is greater than the third preset threshold, taking the traversed data cross table as the source data table.
[0199] In the embodiment, in the case where the value of each element in the correlation coefficient matrix is greater than the third threshold, the traversed data cross table can be taken as the source data table.
[0200] The third preset threshold can be set according to the application scenario or specific requirements in advance. For example, the third preset threshold can be 0.75.
[0201] Further, after step S2315 is executed, the following steps S2318-S23112 can be executed when the proportion of the blank cells in the traversed data cross table is less than the second preset threshold. In the embodiment, steps S2316-S2317 and steps S2318-S23112 can be executed simultaneously, or only steps S2316-S2317 or steps S2318-S23112 can be executed.
[0202] Step S2318, calculating the Mahalanobis distance value of each data cell in the traversed data cross table and the traversed data cross table.
[0203] In the case where the target direction is the row direction, one data cell represents a row of the traversed data cross table; in the case where the target direction is the column direction, one data cell represents a column of the traversed data cross table.
[0204] The Mahalanobis distance can be defined as the degree of difference between two random variables subject to the same distribution and whose covariance matrix is Σ. It represents the distance between a point and a distribution. It is an effective method for calculating the similarity of two unknown sample sets.
[0205] In the case where the target direction is the row direction, the Mahalanobis distance value of the row vector composed of the values of each row in the traversed data cross table and the value matrix composed of the traversed data cross table can be calculated. In the case where the target direction is the column direction, the Mahalanobis distance value of the column vector composed of the values of each column in the traversed data cross table and the value matrix composed of the traversed data cross table can be calculated.
[0206] Step S2319, determine the degree of freedom value of the traversed data cross table in the target direction.
[0207] In this embodiment, determining the degree of freedom value of the traversed data cross table in the target direction can include: constructing a value matrix according to the values in the traversed data cross table; and calculating the degree of freedom value of the value matrix in the target direction.
[0208] For the value matrix composed of the traversed data cross table, the degree of freedom value can reflect the state of mutual constraint of each data cell in the value matrix.
[0209] Step S23110, according to the Mahalanobis distance value and the degree of freedom value, using chi-square test to determine the P value of the data of each data cell in the traversed data cross table.
[0210] The chi-square test is to determine the deviation between the actual observation value and the theoretical inference value of the sample. The deviation between the actual observation value and the theoretical inference value determines the size of the chi-square value. If the chi-square value is larger, the deviation between the two is larger; on the contrary, the smaller the deviation between the two is; if the two values are completely equal, the chi-square value is 0, indicating that the theoretical value is completely consistent.
[0211] Step S23111, according to the comparison result of the percentage of the data cells with P value less than the fourth preset threshold value and all data cells of the traversed data cross table, generate a second data table.
[0212] In this embodiment, the fourth preset threshold value can be set in advance according to the application scenario or specific requirements. For example, the fourth preset threshold value can be 0.005.
[0213] In the case that the target direction is the row direction, according to the percentage comparison result of the data units with the P value less than the fourth preset threshold value and the data of all data units of the traversed data cross table, the value of the row corresponding to any column in the second data table can be obtained according to the percentage of the value of the row corresponding to the column in the second data table, which is obtained according to the percentage of the value of the row corresponding to the column in the second data table, and the sum of the values of all rows corresponding to the column.
[0214] In the case that the traversed data cross table is Table 2, the target direction is the row direction, and the P values of the rows corresponding to all stores, the row corresponding to the store of Heavy XX Company 1, and the row corresponding to the store of Heavy XX Company 2 are less than the fourth preset threshold value, the value of the corresponding element in the second data table can be obtained according to the ratio between the value of each brand corresponding to the row and the sum of the values of each brand.
[0215] Table 6
[0216] Store TCx Xin Changx Chuangx Kangx Xer Xshang All 23% 23% 20% 19% 12% 3% 0.0% Chongxx Company 1 Store 64% 14% 9% 9% 5% 0.0% 0.0% Chongxx Company 2 Store 8% 17% 21% 17% 8% 25% 4%
[0217] Step S23112, taking the second data table as the source data table.
[0218] In the embodiment in which the at least two row and column fields include at least two row fields and do not include column fields, generating the source data table of the visual analysis result of the pivot table according to the data cross table can include the following steps S2321-S2323:
[0219] Step S2321, traversing the data cross table, and judging whether the type of the row field corresponding to the traversed data cross table is a date type.
[0220] Step S2322, when the type of the row field corresponding to the traversed data cross table is a date type, reordering the rows in the traversed data cross table according to the time sequence to obtain the source data table.
[0221] In the case that the traversed data cross table is shown in Table 7, the source data table obtained by reordering the rows in the traversed data cross table according to the time sequence can be shown in Table 8.
[0222] Table 7
[0223] Date 1st 5th 2nd 4th 3rd Sales Quantity Value 230 194 60 42 3
[0224] Table 8
[0225] Date 1st 2nd 3rd 4th 5th Sales Quantity Value 230 60 3 42 194
[0226] Step S2323, when the type of the row field corresponding to the traversed data cross table is not a date type, reordering the rows in the traversed data cross table according to the values of the numerical fields to obtain the source data table.
[0227] In the case that the traversed data cross table is as shown in Table 9, the rows in the traversed data cross table are reordered according to the values of the value fields, and the obtained source data table can be as shown in Table 10.
[0228] Table 9
[0229] Office Yuxu Office Fuxu Office Wanax Office Guangx Office Chongxu Office Sales Quantity Value 194 42 60 3 230
[0230] Table 10
[0231] Office Chongxu Office Yuxu Office Wanax Office Fuxu Office Guangx Office Sales Quantity Value 230 194 60 42 3
[0232] In this embodiment, by reordering the rows in the data cross table, the user can conveniently view the source data table.
[0233] In step S2400, the recommended chart type corresponding to the source data table is determined according to the table data of the source data table.
[0234] In an embodiment of the present disclosure, the recommended chart type corresponding to the source data table is determined according to the table data of the source data table, which can include steps S2410-S2430 as shown below:
[0235] In step S2410, each series value column and each category column suitable for chart creation are determined from each column of the table data of the source data table according to a predetermined column determination manner.
[0236] In a specific application, the table data is the data in the source data table, and the table data of the source data table is arranged in a vertical direction. It can be understood that in order to identify the type of each column of the table data of the source data table, each column of data needs to be traversed and determined, and therefore the table data of the source data table needs to be arranged in a vertical direction. The vertical direction means that each column represents a type of data.
[0237] After the table data of the source data table is determined, since the series value column is relied on to draw the image area of the chart and the category column is relied on to draw the label area of the chart when creating the chart, in order to recommend the chart information, in this step, each series value column and each category column suitable for chart creation are determined from each column of the table data according to a predetermined column determination manner.
[0238] There are various specific implementation manners for determining each series value column and each category column suitable for chart creation from each column of the table data according to a predetermined column determination manner. In order to make the scheme clear and the layout clear, the specific implementation manner for determining each series value column and each category column suitable for chart creation from each column of the table data according to a predetermined column determination manner will be introduced in the following.
[0239] Step S2420, for each category column, based on the feature data of the category column and the feature data of the target column, determine the recommended result for each chart type when creating a chart with the category column and each series value column, wherein the target column includes one or more of the series value columns.
[0240] After determining each series value column and each category column, since the determination of the recommended result for each chart type needs to input the feature data of the series value column and the feature data of the category column, the feature data of each category column and the feature data of the target column are determined first.
[0241] Among them, the feature data in this step includes one or more of the following data:
[0242] Data type, maximum character length, Chinese / English character length in the cell of the maximum character length, number of cells with non-empty content, number of cells with numerical content greater than the average value of the column, number of cells with numerical content less than half of the average value of the column, whether the entire column data type is a non-numeric type, and, in the case of the entire column data type being a numeric type, whether the sum of the entire column data is a specific value, whether the column composed of the entire column data is an increasing sequence, whether the column composed of the entire column data is a decreasing sequence.
[0243] Among them, the data type, the maximum character length, the Chinese / English character length in the cell of the maximum character length, and the number of cells with non-empty content are determined by the table processing client traversing the cell content of the entire column.
[0244] The determination process of the number of cells with numerical content greater than the average value of the column is: the table processing client traverses the entire column cell content, judges whether the entire column cell content contains non-numeric content; if it contains non-numeric content, the result is 0; if it does not contain non-numeric content, calculate the average value of the entire column cell content, and calculate the number of cells greater than the average value according to each cell content, and the number calculated as the result.
[0245] The determination process of the number of cells with numerical content less than half of the average value of the column is: the table processing client traverses the entire column cell content, judges whether the entire column cell content contains non-numeric content; if it contains non-numeric content, the result is 0; if it does not contain non-numeric content, calculate half of the average value of the entire column cell content, and calculate the number of cells less than half of the average value according to each cell content, and the number calculated as the result.
[0246] In the case that the data type of the whole column is a numerical type, whether the sum of the whole column data is a specific value, whether the column composed of the whole column data is an increasing sequence, and whether the column composed of the whole column data is a decreasing sequence are determined by the table processing client traversing the whole column cell content, and in the case that the whole column cell content is all numerical content, whether the sum of the whole column cell is a specific value, whether the column composed of the whole column cell content is an increasing sequence, and whether the column composed of the whole column cell content is a decreasing sequence are calculated.
[0247] For example, the data type can be text, numerical value, date, time, etc., and the specific value can be 1, 10, 100, 1000, etc., and the specific value can be set according to the actual situation.
[0248] In addition, in the process of determining the feature data of each category column and the feature data of the target column, the target column can be one or more columns in the series value column. For example, the determined feature data can be the feature data of each category column and the feature data of the first column series value, or the feature data of each category column and the feature data of the first column series value, the second column series value column.
[0249] In addition, the specific display form of the determined recommendation result for each chart type exists in multiple forms. For example, the specific display form of the recommendation result can be a percentage representing the degree of recommendation, a decimal representing the degree of recommendation, a recommended / non-recommended result content, a most recommended / relatively recommended / non-recommended result content, etc.
[0250] The specific implementation of the recommendation result for each chart type when creating a chart with the category column and each series value column is described below.
[0251] Step S2430, based on the determined recommendation result, output the recommended chart type corresponding to the source data table.
[0252] The recommended chart type is used to represent the recommendation result for each chart type when creating a chart with the category column and each series value column. The chart type can include a line chart, a pie chart, a column chart, etc.
[0253] In a specific embodiment, the recommended chart type can be displayed in the form of a pop-up window, a table, a prompt box, a functional entrance of a selectable option, etc.
[0254] The above embodiment realizes table data based on a table processing client determined source data table, determines series value columns and category columns in the table data, and outputs recommended chart types for each category column and series value column, so that the user can select a more appropriate type when establishing a chart based on the recommended chart type, and the category column and series value column in the chart, avoiding repeated operations of the user, thereby improving the efficiency of chart creation.
[0255] In step S2500, a chart conclusion of the source data table is generated according to the first number of series in the source data table.
[0256] In the embodiment in which the first number of series in the source data table is one, generating the chart conclusion of the source data table according to the first number of series in the source data table can include steps S2510-S2560 as shown below:
[0257] In step S2510, a data sequence in the source data table is obtained.
[0258] Figure 4 And Figure 5 are the same source data table, wherein Figure 4 shows a graphical form of the source data table, Figure 5 shows a tabular form of the source data table. The source data table includes a plurality of data sequences, wherein, Figure 4 And Figure 5 only one of the data sequences is shown, which is {9700, 109267, 6800, 20864, 17602, 68028, 27400, 5000, 92855, 27400, 15650, 15520, 32450}, a total of 13 data items, and the category values corresponding to the data items represent time, for example, the category value "1" represents "February 8", the category value "2" represents "February 11", and so on.
[0259] In step S2520, in a case where the category values corresponding to the data items of the data sequence represent time and the category values corresponding to the data items of the data sequence have a sequential relationship, the data sequence is determined as a first target data sequence, wherein the first target data sequence is in time order.
[0260] In one example, whether the category values corresponding to the data items of the data sequence represent time includes: obtaining a graph type corresponding to the data sequence. In a case where the graph type corresponding to the data sequence is a line graph, it is determined that the category values corresponding to the data items of the data sequence represent time.
[0261] Referring to Figure 4 And Figure 5As shown in the data sequence, 13 data items correspond to 13 category values respectively "1", "2", …, "13", there is an order relationship, and the category values represent time, so it can be determined that the data sequence is in time sequence, and the data sequence can be determined as the first target sequence.
[0262] In step S2530, the first correlation is obtained, wherein the first correlation is the correlation between the data items of the first target data sequence and the category values corresponding to the data items of the first target data sequence.
[0263] In one example, the correlation coefficient between the data items of the first target data sequence and the category values corresponding thereto can be calculated, and whether the data items of the first target data sequence and the category values corresponding thereto are positively correlated, negatively correlated, or not correlated can be determined according to the correlation coefficient.
[0264] In one example, the correlation coefficient takes a value between -1 and 1, 0 indicates no correlation, and the greater the absolute value of the correlation coefficient indicates the greater the first correlation, a positive value tends to be positively correlated, and a negative value tends to be negatively correlated. In one example, the first correlation coefficient can be the covariance between the data items of the first target data sequence and the category values corresponding thereto. The covariance can reflect the correlation between two variables X and Y, that is, whether the change trends of the two variables are consistent. If variable X becomes larger, and variable Y also becomes larger, it indicates that the two variables are positively correlated, and the covariance is positive. If variable X becomes larger, and variable Y becomes smaller, it indicates that the two variables are negatively correlated, and the covariance is negative. In this example, the data items of the first target data sequence are taken as variable X, and the corresponding category values are taken as variable Y, and the covariance between the two is calculated.
[0265] In one example, if the first correlation coefficient between the data items of the data sequence and the category values corresponding to the data items is greater than a positive threshold, it is determined that the data items of the data sequence and the category values corresponding to the data items are positively correlated. The positive threshold is, for example, 0.7.
[0266] In one example, if the first correlation coefficient between the data items of the data sequence and the category values corresponding to the data items is less than a negative threshold, it is determined that the data items of the data sequence and the category values corresponding to the data items are negatively correlated. The negative threshold is, for example, -0.7.
[0267] In one example, if the first correlation coefficient between the data items of the data sequence and the category values corresponding to the data items is less than or equal to a positive threshold and greater than or equal to a negative threshold, it is determined that the data items of the data sequence and the category values corresponding to the data items are not correlated. The positive threshold is, for example, 0.7, and the negative threshold is, for example, -0.7, that is, if the first correlation coefficient is less than or equal to 0.7 and greater than or equal to -0.7, it is determined that the data items of the data sequence and the category values corresponding to the data items are not correlated.
[0268] At step S2540, according to the first correlation, a chart processing result corresponding to the first correlation is outputted.
[0269] The step S2540 is exemplified as follows:
[0270] In one example, the step S2540 includes steps S2541-S2542.
[0271] At step S2541, in a case where the first correlation indicates that the data items of the first target data sequence and the corresponding category values are positively correlated, a relationship between two data items adjacent in the order direction in the first target data sequence is determined.
[0272] At step S2542, according to the relationship between the two data items adjacent in the order direction in the first target data sequence, a chart processing result corresponding to the first correlation is generated.
[0273] In the step S2541, the relationship between the two data items adjacent in the order direction in the first target data sequence can include: for the two data items adjacent in the first target data sequence, the growth rate of the latter data item relative to the former data item is calculated.
[0274] For the i-th data item in the first target data sequence, Ki=(Si-Si-1) / Si-1, where Si is the i-th data item in the first target data sequence, Si-1 is the (i-1)-th data item in the first target data sequence, and i is an integer and i≥2. Si-1 and Si are adjacent in the order direction, and Ki is the growth rate of the i-th data item relative to the (i-1)-th data item in the first target data sequence.
[0275] In the step S2542, according to the relationship between the two data items adjacent in the order direction in the first target data sequence, the chart processing result corresponding to the first correlation can include: according to the growth rate of the latter data item relative to the former data item, it is determined whether the latter data item is negative growth; according to the number value of the data items with negative growth in the first target data sequence and the position of the data items with negative growth in the first target data sequence, the chart processing result corresponding to the first correlation is generated.
[0276] If Ki is positive, it is determined that the i-th data item Si is positive growth. If Ki is negative, it is determined that the i-th data item Si is negative growth.
[0277] According to the number value of the data items with negative growth in the first target data sequence and the position of the data items with negative growth in the first target data sequence, the chart processing result corresponding to the first correlation can be, for example:
[0278] If the number of data items with negative growth in the first target data sequence is zero, the generated chart processing conclusion is steady growth.
[0279] If the number of data items with negative growth in the first target data sequence is 1 and the only data item with negative growth is the last data item of the first target data sequence, the generated chart processing conclusion includes overall growth but slight slowdown in recent growth.
[0280] If the number of data items with negative growth in the first target data sequence is 1, assuming that the time represented by the category value corresponding to the only data item is T11, the generated chart processing conclusion includes continuous growth at all times except for a decline at time T11.
[0281] If the number of data items with negative growth in the first target data sequence is greater than or equal to 1, the data item with the minimum growth rate is selected, and assuming that the time represented by the category value corresponding to the data item with the minimum growth rate is T12, the generated chart processing conclusion includes the fastest decline at time T12.
[0282] For a data sequence in a chart with time as the axis, the chart processing method in the above examples can be used to mine key information to generate relevant data conclusions, without the need for excessive human intervention, thereby achieving automation of chart processing, saving chart processing time, and bringing convenience to users.
[0283] In one example, step S2540 includes steps S2543-S2544.
[0284] Step S2543, in the case where the first correlation represents a negative correlation between a data item of the first target data sequence and its corresponding category value, determining the relationship between two adjacent data items in the first target data sequence in the sequential direction.
[0285] Step S2544, generating a chart processing result corresponding to the first correlation according to the relationship between the two adjacent data items in the first target data sequence in the sequential direction.
[0286] In step S2543, determining the relationship between the two adjacent data items in the first target data sequence in the sequential direction can include: for the two adjacent data items in the first target data sequence, calculating the growth rate of the latter data item relative to the former data item.
[0287] For the i-th data item in the first target data sequence, Ki=(Si-Si-1) / Si-1, where Si is the i-th data item in the first target data sequence, Si-1 is the (i-1)-th data item in the first target data sequence, i is an integer and i≥2. Si-1 and Si are adjacent in the order direction, and Ki is the growth rate of the i-th data item in the first target data sequence relative to the (i-1)-th data item.
[0288] In step S2544, the chart processing result corresponding to the first correlation is generated according to the relationship between two adjacent data items in the first target data sequence in the order direction, which can include: determining whether the latter data item is increasing according to the growth rate of the latter data item relative to the former data item; and generating the chart processing result corresponding to the first correlation according to the number value of the data items increasing in the first target data sequence and the position of the data items increasing in the first target data sequence.
[0289] If Ki is positive, it is determined that the i-th data item Si is increasing. If Ki is negative, it is determined that the i-th data item Si is decreasing.
[0290] The chart processing result corresponding to the first correlation is generated according to the number value of the data items increasing in the first target data sequence and the position of the data items increasing in the first target data sequence, which can be, for example:
[0291] If the number value of the data items increasing in the first target data sequence is zero, the generated chart processing conclusion includes: steady decline.
[0292] If the number value of the data items increasing in the first target data sequence is 1 and the only data item increasing is the last data item of the first target data sequence, the generated chart processing conclusion includes: overall decline but the decline rate slows down slightly in the recent period.
[0293] If the number value of the data items increasing in the first target data sequence is 1, assuming that the category value corresponding to the only data item represents time T21, the generated chart processing conclusion includes: continuous decline at other times except for the growth at time T21.
[0294] If the number value of the data items increasing in the first target data sequence is greater than or equal to 1, the data item with the largest growth rate is selected, and assuming that the category value corresponding to the data item with the largest growth rate represents time T22, the generated chart processing conclusion includes: fastest growth at time T22.
[0295] For the data sequence with time as the axis in the chart, the chart processing method of the above example can mine key information to generate relevant data conclusions without too much human intervention, realize the automation of chart processing, save chart processing time, and bring convenience to users.
[0296] In one example, step S2540 includes steps S2545-S2547.
[0297] Step S2545, in the case where the first correlation represents that the data item of the first target data sequence and its corresponding category value are irrelevant, determining the average value of the data item of the first target data sequence and the maximum value and the minimum value in the data item of the first target data sequence.
[0298] Step S2546, determining the first difference value and the second difference value, the first difference value being the difference between the maximum value and the average value, and the second difference value being the difference between the average value and the minimum value.
[0299] Step S2547, in the case where the first difference value is greater than the second difference value, outputting the category value corresponding to the maximum value. In the case where the first difference value is less than the second difference value, outputting the category value corresponding to the minimum value.
[0300] For example, in the case where the first difference value is greater than the second difference value, the category value corresponding to the maximum value is outputted, and it is assumed that the category value corresponding to the maximum value is time T31, and the generated chart processing result includes: the maximum value at time T31.
[0301] For example, in the case where the first difference value is less than the second difference value, the category value corresponding to the minimum value is time T32, and the generated chart processing result includes: the minimum value at time T32.
[0302] For the data sequence with time as the axis in the chart, the chart processing method of the above example can mine key information to generate relevant data conclusions without too much human intervention, realize the automation of chart processing, save chart processing time, and bring convenience to users.
[0303] Step S2550, obtaining the chart conclusion of the source data table according to the chart processing result related to the first correlation.
[0304] In one embodiment of the present disclosure, the chart processing result related to the first correlation can be taken as the chart conclusion of the source data table.
[0305] In one embodiment of the present disclosure, the method of the present embodiment can further include steps S3100-S3400.
[0306] Step S3100, for two adjacent data items in the first target data sequence, calculating the growth rate of the latter data item relative to the former data item.
[0307] In step S3100, for the i-th data item in the first target data sequence, Ki=(Si-Si-1) / Si-1, where Si is the i-th data item in the first target data sequence, Si-1 is the (i-1)-th data item in the first target data sequence, i is an integer and i≥2. Si-1 and Si are adjacent in the order direction, and Ki is the growth rate of the i-th data item in the first target data sequence relative to the (i-1)-th data item.
[0308] In step S3200, according to the growth rate of the latter data item relative to the former data item, it is determined whether the latter data item is positive growth or negative growth. In step S3200, according to the relationship between the two data items adjacent in the order direction in the first target data sequence, the chart processing result corresponding to the first correlation is generated, which can include: according to the growth rate of the latter data item relative to the former data item, it is determined whether the latter data item is negative growth; according to the number of negative growth data items in the first target data sequence and the position of the negative growth data item in the first target data sequence, the chart processing result corresponding to the first correlation is generated.
[0309] If Ki is positive, it is determined that the i-th data item Si is positive growth. If Ki is negative, it is determined that the i-th data item Si is negative growth.
[0310] In step S3300, according to the category value of the number of positive growth data items and the category value of the negative growth data items, the corresponding chart processing result is output.
[0311] In step S3300, the positive growth data items are collected according to the same type, and the collected positive growth data items are arranged in time sequence or other sequence, and the continuous and uninterrupted at least one positive growth data item is taken as a collection data, and after all the collection data are arranged, the processing result corresponding to the type value of the positive growth data item is output, and the processing result can describe how many times and time periods of the positive growth in the chart processing. Similarly, the negative growth data items are collected according to the same type, and the collected negative growth data items are arranged in time sequence or other sequence, and the continuous and uninterrupted at least one negative growth data item is taken as a collection data, and after all the collection data are arranged, the processing result corresponding to the type value of the negative growth data item is output, and the processing result can describe how many times and time periods of the negative growth in the chart processing. Further, the processing result of the corresponding chart can be output as the comparison result of the number of times of the positive growth and the number of times of the negative growth, and the length of the time period of the positive growth and the length of the time period of the negative growth, and the length of the time period is the total time of all time periods.
[0312] In step S3400, the chart conclusion of the source data table is obtained according to the corresponding chart processing result.
[0313] In this embodiment, the chart processing result corresponding to the first correlation and the corresponding chart processing result can be taken as the chart conclusion of the source data table.
[0314] In the case where there are N first target time data sequences and there is a time sequence relationship between the N first target data sequences (the N first target data sequences can belong to the same chart or multiple charts), N is an integer and N≥2, the method can include steps S3500-S3900 as shown below.
[0315] In step S3500, the data items of the nth first target data sequence are processed to obtain the target data corresponding to the nth first target data sequence, and the nth first target data sequence is any one of the N first target data sequences.
[0316] The target data corresponding to the first target data sequence is data determined according to the data items in the first target data sequence. For example, the target data corresponding to the first target data sequence can be the average of all data items in the first target data sequence. For example, the target data corresponding to the first target data sequence can be the maximum of all data items in the first target data sequence. For example, the target data corresponding to the first target data sequence can be the minimum of all data items in the first target data sequence. For example, the target data corresponding to the first target data sequence can be the difference between the maximum and the minimum.
[0317] In step S3600, the target data corresponding to the N first target data sequences are assigned as data items to category values according to the time sequence relationship between the N first target data sequences to construct a second target data sequence in time sequence.
[0318] The time sequence relationship between the N first target data sequences can be determined according to the time ranges corresponding to the first target data sequences. For example, the time range corresponding to the first first target data sequence is the first half of the year, and the time range corresponding to the second first target data sequence is the second half of the year, and the time sequence relationship between the two is that the first first target data sequence is in the front and the second first target data sequence is in the back.
[0319] The target data corresponding to the N first target data sequences are assigned as N new data items, and these new data items are arranged into a new sequence according to the time sequence relationship between the N first target data sequences. The data items in the new sequence are assigned category values, so that the category values corresponding to the data items in the new sequence have a sequence relationship, thereby obtaining a second target data sequence in time sequence.
[0320] In step S3700, a second correlation is obtained, which is the correlation between the data items of the second target data sequence and the category values corresponding to the data items of the second target data sequence.
[0321] In step S3800, according to the second correlation, a chart processing result corresponding to the second correlation is output.
[0322] In this example, if there are multiple first target time data sequences and there is a time sequence relationship between these first target data sequences, the target data corresponding to these first target data sequences are assigned as data items to category values to construct a second target data sequence in time sequence. Then through steps S3700-S3800, based on a similar manner as the aforementioned steps S2530-S2540, a chart processing result for the second target data sequence is output.
[0323] In this example, comprehensive analysis can be performed on multiple time-axis data sequences in the chart, the relationship between the multiple time-axis data sequences is mined to generate relevant data conclusions, without too much human intervention, the automation of chart processing is realized, the chart processing time is saved, and convenience is brought to the user.
[0324] In step S3900, the chart conclusion of the source data table is obtained according to the chart processing result corresponding to the second correlation.
[0325] In this embodiment, the chart processing result corresponding to the first correlation and the chart processing result corresponding to the second correlation can be used as the chart conclusion of the source data table.
[0326] In step S2560, when the category value corresponding to the data item of the data sequence does not represent time, or the category value corresponding to the data item of the data sequence does not have a sequential relationship, the chart conclusion of the source data table is generated according to the recommended chart type.
[0327] In the embodiment in which the recommended chart type is a pie chart, generating the chart conclusion of the source data table according to the recommended chart type can include steps S2561-S2563 as shown below:
[0328] In step S2561, when the proportion of the data item corresponding to one category value of the source data table is greater than or equal to a first threshold value, it is determined that the chart conclusion of the source data table represents the proportion of the data item corresponding to the category value.
[0329] In this embodiment, the first threshold value can be set according to the application scenario or specific requirements in advance, for example, the first threshold value can be 50%.
[0330] Figure 6 And Figure 7 are the same source data table, wherein Figure 6 shows the graphical form of the source data table, Figure 7 shows the table form of the source data table.
[0331] Referring to Figure 6 and Figure 7 , the proportion of the data item of the A department is greater than 50%, and it can be determined that the chart conclusion of the source data table is that the count of the A department accounts for 53% of the total.
[0332] In step S2562, when the sum of the proportions of the data items corresponding to two category values of the source data table is greater than or equal to a first threshold value, it is determined that the chart conclusion of the source data table represents the two category values corresponding to the data items whose proportions exceed the first threshold value, and the sum of the proportions of the data items corresponding to the two category values.
[0333] In the embodiment, if the sum of the proportions of the data items corresponding to the category value 1 and the category value 2 is greater than or equal to the first threshold value, the chart conclusion of the source data table can be: more than 50% of the proportions are concentrated in the category value 1 and the category value 2, the proportion of the data items corresponding to the category value 1 is *%, and the proportion of the data items corresponding to the category value 2 is *%. Wherein * represents the specific proportion value of the data items corresponding to the corresponding category value.
[0334] In step S2563, in the case that the sum of the proportions of the data items corresponding to at least three category values of the source data table is greater than or equal to the first threshold value, it is determined that the chart conclusion of the source data table represents the proportion condition of the data items corresponding to each category value of the source data table.
[0335] In the embodiment, in the case that the category values of the source data table include the category value 1 to the category value m (wherein m is an integer greater than 2) and the sum of the proportions of the data items corresponding to at least three category values is greater than or equal to 50%, the chart conclusion of the source data table can be: the proportion of the data items corresponding to the category value 1 is *%, the proportion of the data items corresponding to the category value 2 is *%, and the proportion of the data items corresponding to the category value m is *%. Wherein * represents the specific proportion value of the data items corresponding to the corresponding category value.
[0336] In the embodiment in which the recommended chart type is the column chart, according to the recommended chart type, the chart conclusion of the source data table can be generated, which can include steps S2564-S2567 as shown below:
[0337] In step S2564, in the case that the second number value of the category values of the source data table is greater than or equal to the second threshold value, and the proportion of the data items corresponding to one category value is greater than or equal to the third threshold value, or the proportion of the data items corresponding to one category value is greater than the first multiple of the average value, it is determined that the chart conclusion of the source data table represents that the category value ranks first; wherein the average value is the average value of the data items corresponding to all category values.
[0338] In the embodiment, the second threshold value, the third threshold value and the first multiple are respectively set according to the application scenario or specific requirements in advance, for example, the second threshold value can be 2, the third threshold value can be 50%, and the first multiple can be 1.5.
[0339] In the case that the second number value of the category values of the source data table is greater than or equal to 2, and the sum of the proportions of the data items corresponding to the category value m is greater than or equal to 50%, the chart conclusion of the source data table can be: the category value m ranks first. Or, in the case that the second number value of the category values of the source data table is greater than or equal to 2, and the proportion of the data items corresponding to the category value m is greater than the average value, the chart conclusion of the source data table can be: the category value m ranks first. Wherein the average value is the average value of the data items corresponding to all category values in the source data table.
[0340] In the step S2565, in the case that the second quantity value of the category values of the source data table is greater than the second threshold value, and the data item proportion corresponding to the two category values is greater than or equal to the third threshold value, it is determined that the chart conclusion of the source data table indicates that the two category values corresponding to the data items with the proportion exceeding the third threshold value, and the sum of the data item proportions corresponding to the two category values.
[0341] In the embodiment, if the second quantity value of the category values of the source data table is greater than or equal to 2, and the sum of the data item proportions corresponding to the category value 1 and the category value 2 is greater than or equal to the third threshold value, the chart conclusion of the source data table can be: the proportions exceeding 50% are concentrated in the category value 1 and the category value 2, the proportion of the data items corresponding to the category value 1 is *%, and the proportion of the data items corresponding to the category value 2 is *%. Wherein, * represents the specific proportion value of the data items corresponding to the corresponding category value.
[0342] In the step S2566, in the case that the second quantity value of the category values of the source data table is greater than or equal to the second threshold value, and the data item proportion corresponding to one category value is less than the second multiple of the average value, it is determined that the chart conclusion of the source data table indicates that the category value is the lowest.
[0343] In the embodiment, the second multiple is set according to the application scenario or specific requirements in advance, for example, the second multiple can be 2 / 3.
[0344] In the case that the second quantity value of the category values of the source data table is greater than or equal to 2, and the data item proportion corresponding to the category value m is less than 2 / 3 of the average value, the chart conclusion of the source data table can be: the category value m is the lowest. Wherein, the average value is the average value of the data items corresponding to all category values in the source data table.
[0345] In the step S2567, in other cases, it is determined that the chart conclusion of the source data table indicates the distribution of the data items corresponding to each category value of the source data table.
[0346] The other cases in the embodiment can be cases other than the following four cases: the first case, the case that the second quantity value of the category values of the source data table is greater than or equal to the second threshold value, and the data item proportion corresponding to one category value is greater than or equal to the third threshold value; the second case, the case that the second quantity value of the category values of the source data table is greater than or equal to the second threshold value, and the data item proportion corresponding to one category value is greater than the first multiple of the average value; the third case, the case that the second quantity value of the category values of the source data table is greater than the second threshold value, and the data item proportions corresponding to the two category values are greater than or equal to the third threshold value; and the fourth case, the case that the second quantity value of the category values of the source data table is greater than or equal to the second threshold value, and the data item proportion corresponding to one category value is less than the second multiple of the average value.
[0347] In other cases, the category values of the source data table include category value 1 to category value m, and the chart conclusion of the source data table can be that the proportion of data items corresponding to category value 1 is *%, the proportion of data items corresponding to category value 2 is *%, and the proportion of data items corresponding to category value m is *%. Wherein, * represents the specific proportion value of the data items corresponding to the corresponding category value.
[0348] In the embodiment in which the recommended chart type is a line chart, according to the recommended chart type, generating the chart conclusion of the source data table can include the following step S2568: determining that the chart conclusion of the source data table represents the distribution of data items corresponding to each category value of the source data table.
[0349] In this embodiment, the category values of the source data table include category value 1 to category value m, and the chart conclusion of the source data table can be that the proportion of data items corresponding to category value 1 is *%, the proportion of data items corresponding to category value 2 is *%, and the proportion of data items corresponding to category value m is *%. Wherein, * represents the specific proportion value of the data items corresponding to the corresponding category value.
[0350] In the embodiment in which the first number of series in the source data table is at least two, according to the first number of series in the source data table, generating the chart conclusion of the source data table can include the following steps S2570-S2580:
[0351] Step S2570, calculating the third correlation of data items of all series in the source data table.
[0352] In one example, the third correlation coefficient can be the covariance between all series of data items.
[0353] In another example, the third correlation coefficient can be the Euclidean distance between all series of data items.
[0354] Step S2580, generating the chart conclusion of the source data table according to the third correlation.
[0355] In the embodiment in which the first number of series in the source data table is two, according to the third correlation, generating the chart conclusion of the source data table can include the following steps S2581-S2582:
[0356] Step S2581, in the case where the third correlation of data items of the two series in the source data table is greater than a fourth threshold value, obtaining the series with a larger total sum of data items, and determining that the chart conclusion of the source data table represents that the series is generally larger.
[0357] In this embodiment, the fourth threshold value is set in advance according to application scenarios or specific requirements, for example, the fourth threshold value can be 0.7.
[0358] In the embodiment, the series of the source data table can include series 1 and series 2, if the third correlation of the data items of series 1 and series 2 is greater than 0.7, and the sum of the data items of series 1 is greater than the sum of the data items of series 2, the chart conclusion in the source data table can be that series 1 is generally larger.
[0359] In step S2582, in the case that the third correlation of the data items of the two series in the source data table is less than or equal to the fourth threshold value, the difference of the data items of the two series on each category value is obtained, and the category value with the largest difference is obtained, and it is determined that the chart conclusion of the source data table indicates that the distribution of the data items of the two series is inconsistent, wherein the difference of the category value is the largest.
[0360] In the embodiment, the series of the source data table can include series 1 and series 2, and the category values include category value 1 to category value m, if the third correlation of the data items of series 1 and series 2 is less than or equal to 0.7, the differences of series 1 and series 2 on category value 1 to category value m are X1 to Xn respectively, and the difference of series 1 and series 2 on category value m is the largest, the chart conclusion in the source data table can be that the difference of series 1 and series 2 on category value m is the largest.
[0361] In the embodiment in which the first number of values of the series in the source data table is at least three, generating the chart conclusion of the source data table according to the third correlation can include steps S2583 and S2584 as follows:
[0362] In step S2583, in the case that the third correlation of the data items of all the series in the source data table is greater than the fourth threshold value, the series with the larger sum of the data items is obtained, and it is determined that the chart conclusion of the source data table indicates that the series is generally larger.
[0363] In the embodiment, the series of the source data table can include series 1 to series k (wherein k is an integer greater than 2), if the third correlation of the data items of series 1 to series k is greater than 0.7, and the sum of the data items of series k, the chart conclusion in the source data table can be that series k is generally larger.
[0364] In step S2584, in the case that the third correlation of the data items of all the series in the source data table is less than or equal to the fourth threshold value, it is determined that the chart conclusion of the source data table indicates the distribution of the data items of each series in the source data table.
[0365] In the embodiment, the series of the source data table can include series 1 to series k (wherein k is an integer greater than 2), and the chart conclusion of the source data table can be: the proportion of the sum of data items of series 1 is *%, the proportion of the sum of data items of series 1 is *%, …, the proportion of the sum of data items of series 1 is *%; …, the proportion of the sum of data items of series k is *%, the proportion of the sum of data items of series k is *%, …, the proportion of the sum of data items of series k is *%. Wherein * represents the specific proportion value of the sum of data items of the corresponding series.
[0366] In step S2600, the visual analysis result of the pivot table is output according to the source data table, the recommended chart type and the chart conclusion.
[0367] In the embodiment, the visual analysis result of the pivot table can be a list of the source data table, the chart of the source data table conforming to the recommended chart type and the chart conclusion.
[0368] Further, the visual analysis result of the pivot table can be displayed in the electronic device, and can also be a file representing the visual analysis result generated and stored in the electronic device, and can also be output to other electronic devices for display.
[0369] According to the embodiments of the present disclosure, for the pivot table with one numerical field and at least two row and column fields, the corresponding data cross table is generated according to the row and column fields, the source data table of the visual analysis result of the pivot table is generated according to the data cross table, the recommended chart type corresponding to the source data table is determined according to the table data of the source data table, the chart conclusion of the source data table is generated according to the first number of series in the source data table, and the visual analysis result of the pivot table is output according to the source data table, the recommended chart type and the chart conclusion. In this way, the pivot table can be automatically and accurately visualized, and the labor cost of visual analysis of the pivot table can be reduced.
[0370] <Device Embodiment>
[0371] In the embodiment, a visual processing device 5000 for pivot table is provided, as shown in Figure 8As shown, the system includes a pivot table obtaining module 5100, a cross table generating module 5200, a source data table generating module 5300, a chart type determining module 5400, a chart conclusion generating module 5500, and an analysis result output module 5600. The pivot table obtaining module 5100 is configured to obtain a data pivot table, the data pivot table having one numerical field and at least two row and column fields. The cross table generating module 5200 is configured to generate a corresponding data cross table according to the row and column fields. The source data table generating module 5300 is configured to generate a source data table of a visual analysis result of the data pivot table according to the data cross table. The chart type determining module 5400 is configured to determine a recommended chart type corresponding to the source data table according to table data of the source data table. The chart conclusion generating module 5500 is configured to generate a chart conclusion of the source data table according to a first number of series in the source data table. The analysis result output module 5600 is configured to output the visual analysis result of the data pivot table according to the source data table, the recommended chart type, and the chart conclusion.
[0372] In an embodiment of the present disclosure, when the first number of series in the source data table is one, the chart conclusion generating module 5500 further includes:
[0373] a sequence obtaining unit configured to obtain a data sequence in the source data table;
[0374] a sequence determining unit configured to determine, when category values corresponding to data items of the data sequence represent time and the category values corresponding to the data items of the data sequence have a sequential relationship, that the data series is a first target data sequence, the first target data sequence being in time order;
[0375] a first correlation obtaining unit configured to obtain a first correlation between the data items of the first target data sequence and the category values corresponding to the data items of the first target data sequence;
[0376] a first result obtaining unit configured to output a chart processing result related to the first correlation according to the first correlation;
[0377] a first conclusion generating unit configured to obtain a chart conclusion of the source data table according to the chart processing result related to the first correlation;
[0378] a second conclusion generating unit configured to generate the chart conclusion of the source data table according to the recommended chart type when the category values corresponding to the data items of the data sequence do not represent time or the category values corresponding to the data items of the data sequence do not have a sequential relationship.
[0379] In an embodiment of the present disclosure, when the recommended chart type is a pie chart,
[0380] The second conclusion generating unit further comprises:
[0381] A first sub-unit is configured to determine, in a case where a proportion of data items corresponding to one category value of the source data table is greater than or equal to a first threshold, that the chart conclusion of the source data table represents the proportion of data items corresponding to the category value;
[0382] A second sub-unit is configured to determine, in a case where a sum of proportions of data items corresponding to two category values of the source data table is greater than or equal to the first threshold, that the chart conclusion of the source data table represents the two category values corresponding to the data items whose proportions exceed the first threshold and the sum of the proportions of the data items corresponding to the two category values.
[0383] A third sub-unit is configured to determine, in a case where a sum of proportions of data items corresponding to at least three category values of the source data table is greater than or equal to the first threshold, that the chart conclusion of the source data table represents the proportion of data items corresponding to each category value of the source data table.
[0384] In an embodiment of the present disclosure, in a case where the recommended chart type is a column chart, the second conclusion generating unit further comprises:
[0385] A fourth sub-unit is configured to determine, in a case where a second number of category values of the source data table is greater than or equal to a second threshold, and a proportion of data items corresponding to one category value is greater than or equal to a third threshold or the proportion of data items corresponding to one category value is greater than a first multiple of an average value, that the chart conclusion of the source data table represents that the category value ranks first; wherein the average value is an average value of data items corresponding to all category values.
[0386] A fifth sub-unit is configured to determine, in a case where the second number of category values of the source data table is greater than the second threshold, and proportions of data items corresponding to two category values are greater than or equal to the third threshold, that the chart conclusion of the source data table represents the two category values corresponding to the data items whose proportions exceed the third threshold and a sum of the proportions of the data items corresponding to the two category values.
[0387] A sixth sub-unit is configured to determine, in a case where the second number of category values of the source data table is greater than or equal to the second threshold, and a proportion of data items corresponding to one category value is less than a second multiple of the average value, that the chart conclusion of the source data table represents that the category value is the lowest.
[0388] A seventh sub-unit is configured to determine, in other cases, that the chart conclusion of the source data table represents a distribution of data items corresponding to each category value of the source data table.
[0389] In an embodiment of the present disclosure, in the case where the recommended chart type is a line chart, the second conclusion generation unit further comprises:
[0390] An eighth sub-unit configured to determine that the chart conclusion of the source data table represents distribution of data items corresponding to each category value of the source data table.
[0391] In an embodiment of the present disclosure, in the case where the first number of series in the source data table is at least two, the chart conclusion generation module 5500 further comprises:
[0392] A third correlation calculation unit configured to calculate a third correlation of data items of all series in the source data table.
[0393] A third conclusion generation unit configured to generate a chart conclusion of the source data table according to the third correlation.
[0394] In an embodiment of the present disclosure, in the case where the first number of series in the source data table is two, the third conclusion generation unit comprises:
[0395] A ninth sub-unit configured to, in the case where the third correlation of data items of the two series in the source data table is greater than a fourth threshold value, obtain a series with a greater sum of data items, and determine that the chart conclusion of the source data table represents the series as a whole.
[0396] A tenth sub-unit configured to, in the case where the third correlation of data items of the two series in the source data table is less than or equal to the fourth threshold value, obtain a difference of data items of the two series at each category value, and obtain a category value with a maximum difference, and determine that the chart conclusion of the source data table represents that the distribution of data items of the two series is inconsistent, wherein the category value has the maximum difference.
[0397] In an embodiment of the present disclosure, in the case where the first number of series in the source data table is at least three, the third conclusion generation unit comprises:
[0398] An eleventh sub-unit configured to, in the case where the third correlation of data items of all series in the source data table is greater than a fourth threshold value, obtain a series with a greater sum of data items, and determine that the chart conclusion of the source data table represents the series as a whole.
[0399] A twelfth sub-unit configured to, in the case where the third correlation of data items of all series in the source data table is less than or equal to the fourth threshold value, determine that the chart conclusion of the source data table represents distribution of data items of each series of the source data table.
[0400] In an embodiment of the present disclosure, the at least two row and column fields comprise at least one row field and at least one column field.
[0401] The cross table generation module 5200 can include:
[0402] a first combination unit, configured to traverse the row field, and combine the traversed row field, the value field and any column field to form at least one first field combination;
[0403] a first cross table obtaining unit, configured to obtain a data cross table corresponding to the first field combination according to the first field combination and the value in the pivot table.
[0404] In an embodiment of the present disclosure, the source data table generation module 5300 can further include:
[0405] a first traversal unit, configured to traverse the data cross table;
[0406] an F value obtaining unit, configured to perform analysis of variance on the traversed data cross table according to the target direction to obtain a corresponding F value; wherein the target direction includes a row direction and / or a column direction;
[0407] a first processing unit, configured to, when the F value is less than a first preset threshold, process the value in the traversed data cross table according to the target direction to obtain a first data table;
[0408] a first source table obtaining unit, configured to take the first data table as a source data table.
[0409] In an embodiment of the present disclosure, the source data table generation module 5300 can further include:
[0410] a judging unit, configured to, when the F value is less than the first preset threshold, judge whether a proportion value of blank cells in the traversed data cross table is less than a second preset threshold;
[0411] a first calculation unit, configured to, when the proportion value of blank cells in the traversed data cross table is less than the second preset threshold, calculate a correlation coefficient matrix value of the traversed data cross table in the target direction;
[0412] a second source table obtaining unit, configured to, when any correlation coefficient matrix value is greater than a third preset threshold, take the traversed data cross table as a source data table.
[0413] In an embodiment of the present disclosure, the source data table generation module 5300 can further include:
[0414] The second computing unit is configured to calculate a Mahalanobis distance value of data in each data unit of the traversed data cross table and the traversed data cross table; in a case where the target direction is a row direction, one data unit represents one row of the traversed data cross table; in a case where the target direction is a column direction, one data unit represents one column of the traversed data cross table.
[0415] The third computing unit is configured to determine a degree of freedom value of the traversed data cross table in the target direction.
[0416] The P value determining unit is configured to determine, according to the Mahalanobis distance value and the degree of freedom value, a P value of the data in each data unit of the traversed data cross table by using chi-square test.
[0417] The second processing unit is configured to generate a second data table according to a comparison result of a percentage of data in the data unit with a P value less than a fourth preset threshold and data in all data units of the traversed data cross table.
[0418] The third source table obtaining unit is configured to obtain the second data table as a source data table.
[0419] In an embodiment of the present disclosure, the row-column field only contains a row field and does not contain a column field.
[0420] The cross table generating module 5200 can further include:
[0421] The second combination unit is configured to traverse the row field, and combine the traversed row field and the value field to form a second field combination.
[0422] The second cross table obtaining unit is configured to obtain a data cross table corresponding to the second field combination according to the second field combination and the value in the pivot table.
[0423] In an embodiment of the present disclosure, the source data table generating module 5300 can further include:
[0424] The second traversal unit is configured to traverse the data cross table, and determine whether a type of a row field corresponding to the traversed data cross table is a date type.
[0425] The fourth source table obtaining unit is configured to, when the type of the corresponding row field is the date type, reorder rows in the traversed data cross table according to a time sequence to obtain the source data table.
[0426] The fifth source table obtaining unit is configured to, when the type of the corresponding row field is not the date type, reorder the rows in the traversed data cross table according to the value of the value field to obtain the source data table.
[0427] In an embodiment of the present disclosure, the visualization processing apparatus 5000 further includes a determining module.
[0428] The determining module is configured to acquire a data sequence in the source data table; and determine that the data sequence is a first target data sequence in a case where category values corresponding to data items of the data sequence represent time and the category values corresponding to the data items of the data sequence have a sequential relationship.
[0429] In one example, the visualization processing apparatus 5000 further includes a first processing module, a constructing module, a third acquiring module, and a second output module.
[0430] The first processing module is configured to, in a case where N first target time data sequences exist and there is a time sequential relationship between the N first target data sequences, N being an integer and N≥2, process data items of an nth first target data sequence to obtain target data corresponding to the nth first target data sequence, the nth first target data sequence being any one of the N first target data sequences.
[0431] The constructing module is configured to, according to the time sequential relationship between the N first target data sequences, assign the target data corresponding to the N first target data sequences as data items to category values to construct a second target data sequence in time sequence.
[0432] The third acquiring module is configured to obtain a second correlation, the second correlation being a correlation between a data item of the second target data sequence and a category value corresponding to the data item of the second target data sequence.
[0433] The second output module is configured to output a chart processing result corresponding to the second correlation according to the second correlation.
[0434] The first conclusion generating unit further obtains a chart conclusion of the source data table according to the chart processing result corresponding to the second correlation.
[0435] In one example, the outputting of the chart processing result corresponding to the first correlation according to the first correlation includes: determining a relationship between two data items adjacent in a sequential direction in the first target data sequence in a case where the first correlation represents that the data items of the first target data sequence and the category values corresponding thereto are positively correlated; and generating the chart processing result corresponding to the first correlation according to the relationship between the two data items adjacent in the sequential direction in the first target data sequence.
[0436] In this example, determining the relationship between two adjacent data items in the first target data sequence in the sequential direction includes: for the two adjacent data items in the first target data sequence, calculating the growth rate of the latter data item relative to the former data item. Generating the chart processing result corresponding to the first correlation according to the relationship between the two adjacent data items in the first target data sequence in the sequential direction includes: determining whether the latter data item is a negative growth according to the growth rate of the latter data item relative to the former data item; and generating the chart processing result corresponding to the first correlation according to the number of the data items with negative growth in the first target data sequence and the positions of the data items with negative growth in the first target data sequence.
[0437] In one example, outputting the chart processing result corresponding to the first correlation according to the first correlation includes: in the case that the first correlation represents that the data items of the first target data sequence and the category values corresponding thereto are negatively correlated, determining the relationship between two adjacent data items in the first target data sequence in the sequential direction; and generating the chart processing result corresponding to the first correlation according to the relationship between the two adjacent data items in the first target data sequence in the sequential direction.
[0438] In this example, determining the relationship between two adjacent data items in the first target data sequence in the sequential direction includes: for the two adjacent data items in the first target data sequence, calculating the growth rate of the latter data item relative to the former data item. Generating the chart processing result corresponding to the first correlation according to the relationship between the two adjacent data items in the first target data sequence in the sequential direction includes: determining whether the latter data item is a negative growth according to the growth rate of the latter data item relative to the former data item; and generating the chart processing result corresponding to the first correlation according to the number of the data items with negative growth in the first target data sequence and the positions of the data items with negative growth in the first target data sequence.
[0439] In one example, outputting the chart processing result corresponding to the first correlation according to the first correlation includes: in the case that the first correlation represents that the data items of the first target data sequence and the category values corresponding thereto are not correlated, determining the average value of the data items of the first target data sequence, and the maximum value and the minimum value among the data items of the first target data sequence; determining a first difference value and a second difference value, the first difference value being the difference between the maximum value and the average value, and the second difference value being the difference between the average value and the minimum value; in the case that the first difference value is greater than the second difference value, outputting the category value corresponding to the maximum value; and in the case that the first difference value is less than the second difference value, outputting the category value corresponding to the minimum value.
[0440] In one example, the visualization processing apparatus 5000 further includes a second processing module and a third output module.
[0441] The second processing module is configured to calculate, for two adjacent data items in the first target data sequence, an increase rate of a latter data item relative to a former data item; and determine whether the latter data item is positively increasing or negatively increasing according to the increase rate of the latter data item relative to the former data item.
[0442] The third output module is configured to output a corresponding chart processing result according to the category value of the positively increasing data items and the category value of the negatively increasing data items.
[0443] The first conclusion generating unit is further configured to obtain a chart conclusion of the source data table according to the corresponding chart processing result.
[0444] Those skilled in the art should understand that the data visualization processing device 5000 of the pivot table can be implemented in various ways. For example, the data visualization processing device 5000 of the pivot table can be implemented by configuring a processor with instructions. For example, the instructions can be stored in a ROM, and when the device is started, the instructions are read from the ROM to a programmable device to implement the data visualization processing device 5000 of the pivot table. For example, the data visualization processing device 5000 of the pivot table can be fixed in a special device (such as an ASIC). The data visualization processing device 5000 of the pivot table can be divided into independent units, or they can be combined together. The data visualization processing device 5000 of the pivot table can be implemented by one of the above-mentioned various implementation ways, or can be implemented by a combination of two or more of the above-mentioned various implementation ways.
[0445] In this embodiment, the data visualization processing device 5000 of the pivot table can have various implementation forms. For example, the data visualization processing device 5000 of the pivot table can be any function module running in a software product or application program that provides a data visualization processing service of a pivot table, or a peripheral embedded part, plug-in, patch of the software product or application program, or the software product or application program itself.
[0446] <Embodiment of electronic device>
[0447] The present disclosure also provides an electronic device 6000.
[0448] In one embodiment, the electronic device 6000 can include the aforementioned data visualization processing device 5000 of the pivot table.
[0449] In another embodiment, the electronic device 6000 can further include a processor 6100 and a memory 6200 as shown in Figure 9 The memory 6200 is configured to store executable instructions; and the instructions are configured to control the processor 6100 to perform the aforementioned data visualization processing method of the pivot table.
[0450] In this embodiment, the electronic device 6000 can be any electronic product having a processor 6100 and a memory 6200, such as a mobile phone, a tablet computer, a palm computer, a desktop computer, a notebook computer, a workstation, a game console, a server, and the like.
[0451] <Readable storage medium embodiment>
[0452] In this embodiment, a readable storage medium having a computer program stored thereon is also provided, and the computer program, when executed by a processor, implements the data visualization processing method of the pivot table according to any embodiment of the present disclosure.
[0453] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.
[0454] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a magnetically encoded device such as magnetic strip cards, an optically encoded device such as a compact disc (CD) or DVD, and / or any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0455] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0456] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0457] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0458] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other data storage device. When the computer readable program instructions are loaded into the computer and other programmable data processing apparatus, a series of operational steps are implemented that provide processes such that the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0459] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0460] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of an instruction, and the module, program segment or part of the instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are all equivalent.
[0461] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, their practical applications, or technical improvements in the marketplace, or to enable other persons skilled in the art to understand the embodiments disclosed herein. The scope of the present disclosure is defined by the appended claims.
Claims
1. A method for visualizing a pivot table, characterized in that: include: Obtain a pivot table, wherein the pivot table has a value field and at least two row and column fields; Generate a corresponding data cross table according to the row and column fields; Generate a source data table for visual analysis results of the pivot table based on the data cross table; wherein the source data table is a data table used for visual analysis of the pivot table; Determining a recommended chart type corresponding to the source data table based on the table data of the source data table; generating a chart conclusion of the source data table according to a first quantity value of a series in the source data table; Outputting a visual analysis result of the pivot table according to the source data table, the recommended chart type, and the chart conclusion; When the first quantity value of the series in the source data table is one, generating a chart conclusion of the source data table according to the first quantity value of the series in the source data table includes: Obtain the data sequence in the source data table; When the category values corresponding to the data items of the data sequence represent time and the category values corresponding to the data items of the data sequence have a sequential relationship, determining that the data sequence is a first target data sequence, and the first target data sequence is in time order; Obtaining a first correlation, where the first correlation is a correlation between a data item of the first target data sequence and a category value corresponding to the data item of the first target data sequence; outputting a chart processing result related to the first correlation according to the first correlation; Obtaining a chart conclusion of the source data table based on a chart processing result related to the first correlation; When the category values corresponding to the data items of the data sequence do not represent time or the category values corresponding to the data items of the data sequence do not have a sequential relationship, generating a chart conclusion of the source data table according to the recommended chart type; When there are at least two first quantity values of the series in the source data table, generating a chart conclusion of the source data table according to the first quantity values of the series in the source data table includes: Calculating the third correlation of data items of all series in the source data table; A graphical conclusion of the source data table is generated based on the third correlation.
2. The method according to claim 1, characterized in that In the case where the recommended chart type is a pie chart, Generating a chart conclusion of the source data table according to the recommended chart type further includes: When a proportion of data items corresponding to a category value in the source data table is greater than or equal to a first threshold, determining that a chart conclusion of the source data table represents a proportion of data items corresponding to the category value; If the total proportion of data items corresponding to the two category values of the source data table is greater than or equal to the first threshold, determining that the chart conclusion of the source data table represents the two category values corresponding to the data items whose proportions exceed the first threshold, and the total proportion of data items corresponding to the two category values; When the total proportion of data items corresponding to at least three category values of the source data table is greater than or equal to the first threshold, it is determined that the chart conclusion of the source data table represents the proportion of data items corresponding to each category value of the source data table.
3. The method according to claim 1, characterized in that When the recommended chart type is a bar chart, Generating a chart conclusion of the source data table according to the recommended chart type further includes: In a case where the second quantity value of the category value in the source data table is greater than or equal to the second threshold value, and the proportion of data items corresponding to one category value is greater than or equal to the third threshold value, or in a case where the second quantity value of the category value in the source data table is greater than or equal to the second threshold value, and the proportion of data items corresponding to one category value is greater than a first multiple of the average value, it is determined that the conclusion of the chart of the source data table indicates that the category value ranks first; wherein the average value is the average value of the data items corresponding to all category values; If the second quantity value of the category value of the source data table is greater than the second threshold value, and the proportion of the data items corresponding to the two category values is greater than or equal to the third threshold value, determine that the chart conclusion of the source data table represents the two category values corresponding to the data items whose proportion exceeds the third threshold value, and the total proportion of the data items corresponding to the two category values; When a second quantity value of the category value of the source data table is greater than or equal to a second threshold value, and a proportion of data items corresponding to a category value is less than a second multiple of the average value, determining that the conclusion of the chart of the source data table indicates that the category value is the lowest; In other cases, determining a chart conclusion of the source data table represents a distribution of data items corresponding to each category value of the source data table.
4. The method according to claim 1, wherein When the recommended chart type is a line chart, Generating a chart conclusion of the source data table according to the recommended chart type further includes: Determining a chart conclusion of the source data table represents a distribution of data items corresponding to each category value of the source data table.
5. The method according to claim 1, wherein If the first quantity value of the series in the source data table is two, Generating a chart conclusion of the source data table according to the third correlation includes: When the third correlation between the data items of the two series in the source data table is greater than a fourth threshold, obtaining the series with the larger sum of the data items, and determining that the conclusion of the chart of the source data table indicates that the series is larger overall; When the third correlation between the data items of the two series in the source data table is less than or equal to the fourth threshold, the difference between the data items of the two series in each category value is obtained, and the category value with the largest difference is obtained, and it is determined that the chart conclusion of the source data table indicates that the distribution of the data items of the two series is inconsistent, and the difference in the category value is the largest.
6. The method according to claim 1, characterized in that When the first quantity value of the series in the source data table is at least three, Generating a chart conclusion of the source data table according to the third correlation includes: When the third correlation of the data items of all series in the source data table is greater than a fourth threshold, obtaining a series with a larger sum of data items, and determining that the conclusion of the chart of the source data table indicates that the series is larger overall; When the third correlation of the data items of all series in the source data table is less than or equal to the fourth threshold, it is determined that the chart conclusion of the source data table represents the distribution of the data items of each series in the source data table.
7. The method according to claim 1, characterized in that The at least two row and column fields include at least one row field and at least one column field; Generating a corresponding data cross table according to the row and column fields includes: Traversing the row fields, and combining the traversed row fields, the value fields, and any one of the column fields to form at least one first field combination; A data crosstab corresponding to the first field combination is obtained according to the first field combination and the values in the pivot table.
8. The method according to claim 7, characterized in that Generating a source data table for visual analysis results of the pivot table based on the data cross table also includes: Traversing the data cross table; According to the target direction, performing variance analysis on the traversed data cross table to obtain a corresponding F value; wherein the target direction includes a row direction and / or a column direction; When the F value is less than a first preset threshold, the values in the traversed data cross table are processed according to the target direction to obtain a first data table; The first data table is used as the source data table.
9. The method according to claim 8, characterized in that Generating a source data table for visual analysis results of the pivot table based on the data cross table also includes: When the F value is less than a first preset threshold, determining whether the proportion of blank cells in the traversed data cross table is less than a second preset threshold; When the proportion of blank cells in the traversed data cross table is less than the second preset threshold, calculating the correlation coefficient matrix value of the traversed data cross table in the target direction; When any of the correlation coefficient matrix values is greater than a third preset threshold, the traversed data cross table is used as the source data table.
10. The method according to claim 9, characterized in that When the proportion of blank cells in the traversed data cross table is less than the second preset threshold, the method further includes: Calculating the Mahalanobis distance between the data of each data unit in the traversed data cross table and the traversed data cross table; wherein, when the target direction is the row direction, one data unit represents one row of the traversed data cross table; when the target direction is the column direction, one data unit represents one column of the traversed data cross table; Determining the degree of freedom value of the traversed data cross table in the target direction; Determine the P value of the data of each data unit in the traversed data cross table using a chi-square test according to the Mahalanobis distance value and the degree of freedom value; Generate a second data table based on a comparison result of the data units whose P values are less than a fourth preset threshold and the data percentages of all the data units in the traversed data cross table; The second data table is used as the source data table.
11. The method according to claim 1, wherein The row and column fields only include row fields and do not include column fields; Generating a corresponding data cross table according to the row and column fields includes: Traversing the row fields, and combining the traversed row fields and the value field to form a second field combination; According to the second field combination and the values in the pivot table, a data cross table corresponding to the second field combination is obtained.
12. The method according to claim 11, characterized in that Generating a source data table for visual analysis results of the pivot table based on the data cross table also includes: Traversing the data cross table, and determining whether the type of the row field corresponding to the traversed data cross table is a date type; When the type of the corresponding row field is a date type, reordering the rows in the traversed data cross table according to the time sequence to obtain the source data table; When the type of the corresponding row field is not a date type, the rows in the traversed data cross table are reordered according to the value of the value field to obtain the source data table.
13. The method according to claim 1, wherein In a case where there are N first target time data sequences and there is a time sequence relationship between the N first target data sequences, N is an integer and N≥2, the method further includes: Processing data items of an nth first target data sequence to obtain target data corresponding to the nth first target data sequence, where the nth first target data sequence is any one of the N first target data sequences; According to the time sequence relationship between the N first target data sequences, target data corresponding to the N first target data sequences are assigned category values as data items to construct a second target data sequence in time sequence; obtaining a second correlation, where the second correlation is a correlation between a data item of the second target data sequence and a category value corresponding to the data item of the second target data sequence; outputting a chart processing result corresponding to the second correlation according to the second correlation; Furthermore, according to the chart processing result corresponding to the second correlation, a chart conclusion of the source data table is obtained.
14. The method according to claim 1, wherein Outputting a chart processing result corresponding to the first correlation according to the first correlation includes: determining a relationship between two adjacent data items in a sequential direction in the first target data sequence when the first correlation represents a positive correlation between a data item of the first target data sequence and a corresponding category value; A chart processing result corresponding to the first correlation is generated according to the relationship between two data items adjacent in the order direction in the first target data sequence.
15. The method according to claim 14, characterized in that Generating a chart processing result corresponding to the first correlation according to the relationship between two data items adjacent in a sequential direction in the first target data sequence includes: For two adjacent data items in the first target data sequence, calculating the growth rate of the latter data item relative to the former data item; According to the growth rate of the latter data item relative to the previous data item, determine whether the latter data item is negative growth; A chart processing result corresponding to the first correlation is generated according to the number of negatively growing data items in the first target data sequence and the positions of the negatively growing data items in the first target data sequence.
16. The method according to claim 1, characterized in that Outputting a chart processing result corresponding to the first correlation according to the first correlation includes: determining a relationship between two adjacent data items in a sequential direction in the first target data sequence when the first correlation indicates that a data item of the first target data sequence and a corresponding category value are negatively correlated; A chart processing result corresponding to the first correlation is generated according to the relationship between two data items adjacent in the order direction in the first target data sequence.
17. The method according to claim 16, characterized in that Generating a chart processing result corresponding to the first correlation according to the relationship between two data items adjacent in a sequential direction in the first target data sequence includes: For two adjacent data items in the first target data sequence, calculating the growth rate of the latter data item relative to the former data item; Determine whether the latter data item is positive growth based on the growth rate of the latter data item relative to the former data item; A chart processing result corresponding to the first correlation is generated according to the quantity value of the positively increasing data items in the first target data sequence and the position of the positively increasing data items in the first target data sequence.
18. The method according to claim 1, wherein Outputting a chart processing result corresponding to the first correlation according to the first correlation includes: determining an average value of the data items of the first target data sequence and a maximum value and a minimum value among the data items of the first target data sequence when the first correlation represents that the data items of the first target data sequence and their corresponding category values are uncorrelated; Determine a first difference and a second difference, the first difference being a difference between the maximum value and the average value, and the second difference being a difference between the average value and the minimum value; When the first difference is greater than the second difference, outputting the category value corresponding to the maximum value; When the first difference is smaller than the second difference, the category value corresponding to the minimum value is output.
19. The method according to claim 1, wherein The method further comprises: For two adjacent data items in the first target data sequence, calculating the growth rate of the latter data item relative to the former data item; According to the growth rate of the latter data item relative to the previous data item, determine whether the latter data item has positive growth or negative growth; Output the corresponding chart processing results according to the category values of the positively growing data items and the category values of the negatively growing data items; Furthermore, according to the corresponding chart processing results, a chart conclusion of the source data table is obtained.
20. A visualization processing device for a pivot table, characterized in that: include: A pivot table acquisition module is used to acquire a pivot table, wherein the pivot table has a value field and at least two row and column fields; A cross table generation module, used for generating a corresponding data cross table according to the row and column fields; A source data table generating module, configured to generate a source data table of the visual analysis result of the pivot table according to the data cross table; wherein the source data table is a data table used for visual analysis of the pivot table; A chart type determination module, configured to determine a recommended chart type corresponding to the source data table based on the table data of the source data table; A chart conclusion generating module, configured to generate a chart conclusion of the source data table according to a first quantity value of a series in the source data table; An analysis result output module, configured to output a visual analysis result of the pivot table based on the source data table, the recommended chart type, and the chart conclusion; When the first quantity value of the series in the source data table is one, the chart conclusion generating module includes: A sequence acquisition unit, configured to acquire a data sequence from the source data table; a sequence determining unit, configured to determine that the data sequence is a first target data sequence, where the first target data sequence is ordered by time, if the category values corresponding to the data items of the data sequence represent time and the category values corresponding to the data items of the data sequence have a sequential relationship; a first correlation obtaining unit, configured to obtain a first correlation, where the first correlation is a correlation between a data item of the first target data sequence and a category value corresponding to the data item of the first target data sequence; a first result obtaining unit, configured to output a chart processing result related to the first correlation according to the first correlation; a first conclusion generating unit, configured to obtain a chart conclusion of the source data table according to a chart processing result related to the first correlation; a second conclusion generating unit, configured to generate a chart conclusion for the source data table according to the recommended chart type when the category values corresponding to the data items of the data sequence do not represent time or the category values corresponding to the data items of the data sequence do not have an order relationship; When the first quantity value of the series in the source data table is at least two, the chart conclusion generating module includes: A third correlation calculation unit, configured to calculate a third correlation of all series of data items in the source data table; A third conclusion generating unit is configured to generate a chart conclusion of the source data table according to the third correlation.
21. An electronic device, characterized in that: include: The device according to claim 20; or, A processor and a memory, wherein the memory is used to store instructions, and the instructions are used to control the processor to execute the method according to any one of claims 1 to 19.
22. A readable storage medium, characterized in that A computer program is stored thereon, which implements the method according to any one of claims 1 to 19 when executed by a processor.
Citation Information
Patent Citations
Chart generating method and device, computer equipment and storage medium
CN107688664A
Data processing method, device, computer readable medium and electronic equipment
CN113010582A