Big Data Chart Report Production Method and Device Based on Sampling Calculation

By using sampling calculations during the production of big data chart report, sampling data and obtaining target data in the sampling table, the problem of slow production speed of chart report is solved, and a more efficient production process and resource saving is achieved.

CN114201948BActive Publication Date: 2025-07-08SHANGHAI ZHONGTONGJI NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111475777.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-06
Publication Date
2025-07-08
Estimated Expiration
2041-12-06

AI Technical Summary

Technical Problem

In the prior art, the production speed of big data chart reports is slow, resulting in inefficiency and waste of resources.

Method used

Using a sampling calculation method, the sampling data is extracted from the original data table to the sampling table, and the target data is obtained from the sampling table to make a chart report.

Benefits of technology

Improves the speed and smooth experience of chart reports, reduces resource waste, and does not affect product functionality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114201948B_ABST
    Figure CN114201948B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and device for making a big data chart report based on sampling calculation, belonging to the technical field of big data analysis. The method and device for making a big data chart report based on sampling calculation extract sampling data from the original data table to a sampling table based on a preset sampling rule; receive a query request, and obtain target data in the sampling table; and make a chart report according to the target data, thereby solving the technical problem of slow production speed of chart reports in the prior art. In the whole working process of this application, both the production efficiency and smooth experience of big data chart reports are improved, resource waste is reduced, and product functions are not affected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of big data analysis, and particularly relates to a method and device for making a big data chart report based on sampling calculation. Background Art

[0002] Big data analysis chart reports are one of the most important outputs in the big data industry. The main task of big data analysis is to calculate valuable data from massive data and provide it to decision-makers in the form of chart reports for real-time browsing and viewing, facilitating problem discovery and decision-making assistance. For example, bar charts, pie charts, line charts, two-dimensional tables, crosstabs, etc.

[0003] In related technologies, analysts usually make chart reports through a big data analysis platform. During the process of making chart reports, they will frequently query the underlying massive data, and each operation step requires previewing the effect, resulting in slow speed and wasted time.

[0004] Therefore, how to improve the production speed of chart reports has become a technical problem to be urgently solved in the existing technology. Summary of the Invention

[0005] The present invention provides a method and device for making a big data chart report based on sampling calculation to solve the technical problem of slow production speed of chart reports in the existing technology.

[0006] The technical solution provided by the present invention is as follows:

[0007] On the one hand, a method for making a big data chart report based on sampling calculation includes:

[0008] Based on a preset sampling rule, extract sampling data from the original data table to the sampling table;

[0009] Receive a query request and obtain target data in the sampling table;

[0010] Make a chart report according to the target data.

[0011] Optionally, the extracting sampling data from the original data table to the sampling table based on a preset sampling rule includes:

[0012] Based on a Java timed task, set the scanning interval duration, and scan all original data tables according to the scanning interval duration;

[0013] When there is no corresponding sampling table in any original data table, create a sampling table and add _sample to the suffix of the original data table without a corresponding sampling table as the table name of the sampling table;

[0014] Scan the data in the original data table, perform sampling calculations according to the preset upper limit of the sampling number and the preset sampling rules, and obtain the sampling data;

[0015] Save the sampling data to the sampling table.

[0016] Optionally, the preset sampling rules include:

[0017] Group the data fields, calculate the probability of each group of data fields, and perform sampling according to the probability from high to low;

[0018] Among them, the grouping rules for each type of data field include: the dimension field is grouped by a single value, and the same values are classified into one group; the date field is grouped according to the time precision; the data field is based on the average value, standard deviation, minimum value, maximum value of the statistical field and the preset progressive quantile values; determine the approximate distribution of the data, and perform range grouping according to the principle that the larger the data volume, the finer the grouping.

[0019] Optionally, the receiving the query request and obtaining the target data in the sampling table includes:

[0020] Receive the http query request, and obtain the mode flag in the http query request header;

[0021] Inject the mode flag into the Java thread context;

[0022] Obtain the mode flag in the Java thread context. When the mode flag is the sampling mode, replace all the original data tables with the sampling table;

[0023] Perform data query based on the replaced sampling table set.

[0024] Optionally, the making the chart report according to the target data includes:

[0025] Return the target data to the front end, render the corresponding chart, and obtain the chart report for the user to preview and view.

[0026] In another aspect, a big data chart report making device based on sampling calculation includes: a processor and a memory connected to the processor;

[0027] The memory is used to store a computer program, and the computer program is at least used to execute the big data chart report making method based on sampling calculation described in any one of the above;

[0028] The processor is used to call and execute the computer program in the memory.

[0029] The beneficial effects of the present invention are:

[0030] The method and device for making big data chart reports based on sampling calculation provided by the embodiments of the present invention extract sampling data from the original data table to a sampling table based on a preset sampling rule; receive a query request, and obtain target data in the sampling table; make a chart report according to the target data, thereby solving the technical problem of slow production speed of chart reports in the prior art. The entire working process of this application not only improves the production efficiency and smooth experience of big data chart reports, but also reduces resource waste and does not affect product functions. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0032] Figure 1 It is a schematic flowchart of a method for making a big data chart report based on sampling calculation provided by an embodiment of the present invention;

[0033] Figure 2 It is a schematic flowchart of a sampling query provided by an embodiment of the present invention;

[0034] Figure 3 It is a schematic structural diagram of a device for making a big data chart report based on sampling calculation provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be described in detail below. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other implementation manners obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope protected by the present invention.

[0036] Big data analysis chart reports are one of the most important outputs in the big data industry. The main task of big data analysis is to calculate valuable data from massive data and provide it to decision-makers in the form of chart reports for real-time browsing and viewing, facilitating problem discovery and decision-making assistance. For example, bar charts, pie charts, line charts, two-dimensional tables, crosstabs, etc.

[0037] In related technologies, analysts usually create chart reports through big data analysis platforms. In the process of creating chart reports, they frequently query the underlying massive data. At this time, analysts are most concerned about the structure, display style and display content of the chart report, and do not care about the correctness of the content; in this process, speed and efficiency are the most important, and it is ideal for various operations to produce results in seconds and see the effects quickly; but the reality is: in the process of creating chart reports, every step of the operation, in order to see the effect, must query the underlying massive business data in real time, resulting in slow speed and operation jams; the experience of getting results in seconds has become an unattainable dream.

[0038] At present, the big data analysis platforms used in the industry to produce chart reports are mainly divided into two categories: (1) Professional "big data analysis platforms" provided by third-party companies. Data is stored in third-party companies, and the use of their analysis and "chart report" production capabilities is paid according to data storage volume, query traffic, resource quota, etc.; the query speed is often proportional to the payment and data volume. Considering issues such as data security, capital investment, and cost-effectiveness, this is not a good choice for companies with certain IT technical strength, and is gradually being abandoned by medium and large companies. (2) "Big data analysis platforms" built by enterprises themselves. Enterprises with certain technical strength generally choose to build their own "big data analysis platforms" to analyze and produce "chart reports". As the amount of data accumulates, they want to improve the speed of big data analysis and improve the fluency and efficiency of "chart report" production. They increasingly rely on the underlying stacking of machines and increasing computing resources to solve the problem. This in turn leads to an increase in the company's investment costs and creates financial pressure, which eventually becomes a bottleneck for speed improvement. After all, it is not a long-term solution.

[0039] Therefore, how to improve the speed of making chart reports has become a technical problem that needs to be solved urgently in the existing technology.

[0040] Based on this, an embodiment of the present invention provides a method and device for producing a big data chart report based on sampling calculation.

[0041] Embodiment 1:

[0042] An embodiment of the present invention provides a method for producing a large data chart report based on sampling calculation.

[0043] Figure 1 A flowchart of a method for making a large data chart report based on sampling calculation provided by an embodiment of the present invention is shown in FIG. Figure 1 The method provided by the embodiment of the present invention may include the following steps:

[0044] S11. Based on the preset sampling rules, extract the sampling data from the original data table to the sampling table.

[0045] In some embodiments, it may include: setting a scanning interval duration based on a Java timed task, and scanning all original data tables according to the scanning interval duration; when there is no corresponding sampling table for any original data table, creating a sampling table, and adding _sample to the suffix of the original data table without a corresponding sampling table as the table name of the sampling table; scanning the data in the original data table, performing sampling calculation according to a preset sampling number upper limit and a preset sampling rule to obtain sampling data; and saving the sampling data into the sampling table.

[0046] For example, (1) a Java timed task can be started to scan all original data tables regularly every day; (2) for each original data table, if there is no corresponding sampling table, a sampling table with the same table structure is automatically created, and _sample is added to the suffix of the original data table as the table name of the sampling table; (3) scan the data in the original data table, perform sampling calculation according to the sampling number upper limit and the sampling rule to obtain the sampled data; (4) save the sampling data into the sampling table. Among them, steps (2), (3), and (4) can be executed irregularly according to a pre-specified frequency.

[0047] It should be noted that according to actual business needs and underlying computing capabilities, the sampling number upper limit and the sampling calculation frequency can be configured and adjusted. If resources are abundant, the larger the sampling number upper limit and the sampling calculation frequency, the better, and the closer to the calculation result of the actual data.

[0048] In some embodiments, the preset sampling rule includes: grouping data fields, calculating the probability of each group of data fields, and performing sampling according to the probability from high to low; among them, the grouping rule for each type of data field includes: dimension fields are grouped by a single value, and the same values are grouped into one group; date fields are grouped according to the time precision; data fields are based on the average value, standard deviation, minimum value, maximum value of the statistical field, and preset progressive percentile values; determine the approximate distribution of the data, and perform range grouping according to the principle that the larger the data volume, the finer the grouping.

[0049] For example, the sampling rule may include the following content: the general sampling principle is to group and calculate the probability of field values, and perform sampling according to the probability from high to low. The higher the probability of a group, the more sampling is done; specifically, for each type of data field, the grouping rule is as follows: (1) dimension fields (such as strings): grouped by a single value, and the same values are grouped into one group; (2) date fields: grouped according to the specific time precision (such as days, hours, minutes, etc.); (3) data fields (such as amount, weight): first calculate the average value, standard deviation, minimum value, maximum value of the field, and progressive percentile values from 1 to 5 (such as: 5%, 10%, 15%......95%, the granularity is adjusted according to the data volume); finally, obtain the approximate distribution of the data, and perform range grouping according to the principle that the larger the data volume, the finer the grouping.

[0050] S12. Receive a query request and obtain target data from the sampling table.

[0051] In some embodiments, receiving a query request and obtaining target data from the sampling table includes: receiving an http query request, obtaining a mode flag from the http query request header; implanting the mode flag into the Java thread context; obtaining the mode flag in the Java thread context, and when the mode flag is the sampling mode, replacing all original data tables with the sampling table; performing data query based on the replaced set of sampling tables.

[0052] Figure 2 The following is a schematic flowchart of a sampling query provided by an embodiment of the present invention. Refer to Figure 2 , this application can be based on a Java-based web big data analysis platform for solution description and implementation. The sampling query (i.e., step S12) can specifically include the following: the user can select the actual calculation mode or the sampling mode on the web front-end page; when initiating a query data request, automatically implant the mode flag into the http request header and pass it to the backend; the backend automatically obtains the mode flag from the http request header and places it in the thread context; the intermediate business processing layer does not need to perform any processing, and only needs to intercept the request in the data query layer at the bottom of the system to obtain the query SQL; obtain the mode flag from the Java thread context. If it is the sampling calculation mode, automatically replace all original data tables in the executed SQL with the sampling table (according to the rule, just add the suffix “_sample” to the table name of the “original data table”); send the replaced SQL to the underlying query engine for data query calculation; the query engine queries the sampling table. Since the amount of data is very small, the result can be obtained in seconds and returned quickly, saving the time-consuming of querying and calculating from the massive original data tables, and also saving a large amount of underlying physical resources.

[0053] S13. Generate a chart report based on the target data.

[0054] In some embodiments, generating a chart report based on the target data includes: returning the target data to the front end, rendering the corresponding chart, and obtaining the chart report for the user to preview and view.

[0055] For example, the sampling calculation result can be directly returned to the front end to render the corresponding chart for the producer to quickly preview the effect. After all the reports are made, switch the mode back to the actual calculation mode to view the real data.

[0056] The big data chart report production method and device based on sampling calculation provided by the embodiments of the present invention extract sampling data from the original data table to the sampling table based on a preset sampling rule; receive a query request, obtain target data in the sampling table; and produce a chart report according to the target data. Thereby, the technical problem of slow production speed of chart reports in the prior art is solved. The entire working process of this application not only improves the production efficiency and smooth experience of big data chart reports, but also reduces resource waste and does not affect product functions.

[0057] Embodiment 2:

[0058] Based on one general inventive concept, the embodiments of the present invention also provide a big data chart report production device based on sampling calculation.

[0059] Figure 3 For the structural schematic diagram of a big data chart report production device based on sampling calculation provided by the embodiments of the present invention, please refer to Figure 3 A big data chart report production device based on sampling calculation provided by the embodiments of the present invention includes: a processor 31 and a memory 32 connected to the processor.

[0060] The memory 32 is used to store a computer program, and the computer program is at least used for the big data chart report production method based on sampling calculation described in any one of the above embodiments;

[0061] The processor 31 is used to call and execute the computer program in the memory.

[0062] Based on one general inventive concept, the embodiments of the present invention also provide a storage medium.

[0063] A storage medium stores a computer program, and when the computer program is executed by a processor, each step in the above-mentioned big data chart report production method based on sampling calculation is implemented.

[0064] As mentioned above, only the specific embodiments of the present invention are described, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claimed rights.

[0065] It can be understood that the same or similar parts in the above embodiments can be referred to each other, and the content not detailed in some embodiments can be seen in the same or similar content in other embodiments.

[0066] It should be noted that in the description of the present invention, the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "a plurality of" refers to at least two.

[0067] Any process or method description shown in the flowchart or described in other ways herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in the reverse order according to the involved functions, rather than in the order shown or discussed. This should be understood by those skilled in the technical field of the embodiments of the present invention.

[0068] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0069] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program. The said program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0070] In addition, each functional unit in various embodiments of the present invention can be integrated into one processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0071] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.

[0072] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0073] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for making a big data chart report based on sampling calculation, characterized in that, Including: Extracting sampling data from the original data table to the sampling table based on a preset sampling rule; Setting a scanning interval duration based on a Java timed task and scanning all original data tables according to the scanning interval duration; When there is no corresponding sampling table in any original data table, creating a sampling table and adding _sample to the suffix of the original data table without a corresponding sampling table as the table name of the sampling table; Scanning the data in the original data table, performing sampling calculations according to a preset sampling number upper limit and a preset sampling rule to obtain sampling data. The preset sampling rule includes: grouping data fields, calculating the probability of each group of data fields, and performing sampling according to the probability level. Among them, the grouping rule for each type of data field includes: dimension fields are grouped by a single value, and the same values are grouped into one group; date fields are grouped according to the time precision; data fields are based on the average value, standard deviation, minimum value, maximum value, and preset progressive quantile values of statistical fields; determining the approximate distribution of the data, and performing range grouping according to the principle that the larger the data volume, the finer the grouping; Saving the sampling data to the sampling table; Receiving a query request and obtaining target data from the sampling table; Making a chart report according to the target data.

2. The method according to claim 1, wherein The receiving a query request and obtaining target data from the sampling table includes: Receiving an http query request and obtaining a mode flag from the http query request header; Implanting the mode flag into the Java thread context; Obtaining the mode flag in the Java thread context. When the mode flag is in the sampling mode, replacing all original data tables with sampling tables; Performing data query based on the replaced set of sampling tables.

3. The method according to claim 1, characterized in that The making a chart report according to the target data includes: Returning the target data to the front end, rendering the corresponding chart, and obtaining a chart report for the user to preview and view.

4. A big data chart report production device based on sampling calculation, characterized in that Including: A processor and a memory connected to the processor; The memory is used to store a computer program, and the computer program is at least used to execute the method for making a big data chart report based on sampling calculation according to any one of claims 1 to 3; The processor is used to call and execute the computer program in the memory.

Citation Information

Patent Citations

  • Method, device and system for revealing operation result

    CN101739410A

  • Data display method and system

    CN108197297A