Data report generation method and device and storage medium

By automatically selecting field combinations from the target data table and generating visual charts and text descriptions, the problem of relying on manual operations in the existing technology is solved, and efficient and easy-to-understand data report generation is achieved, which improves data analysis efficiency and report readability.

CN120508591APending Publication Date: 2025-08-19ZHUHAI KINGSOFT OFFICE SOFTWARE +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510393054.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing data analysis methods and data report production methods rely on manual operation, which is time-consuming and labor-intensive, and the produced data reports have flaws in terms of accuracy and completeness.

Method used

By automatically selecting the target field combination from the fields of the target data table, visual charts and text descriptions are generated, and integrated into data reports.

Benefits of technology

It greatly reduces the time for manual participation in data analysis and report production, significantly improves the overall efficiency of data analysis, and presents complex data relationships in an intuitive and easy-to-understand way, enhancing the readability of data reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508591A_ABST
    Figure CN120508591A_ABST
Patent Text Reader

Abstract

The invention relates to a data report generation method and device and a storage medium. The method comprises the steps that a target field combination is selected from fields of a target data table; based on the data item of each field in the target field combination, material information of the target field combination is generated, and the material information comprises a visual chart and text description; and integrating the material information of the target field combination to form a data report of the target data table. Therefore, the data can be automatically processed, the visual and easy-to-understand data report can be generated, the time for manual participation in data analysis and report making is greatly shortened, the overall efficiency of data analysis is remarkably improved, meanwhile, the complex data relation can be presented in a visual and easy-to-understand mode, the readability of the data report is enhanced, and the data analysis efficiency is improved. And the user is helped to quickly understand information and trend behind the data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular to a data report generation method, device, and storage medium. Background Art

[0002] In the field of data processing and analysis, extracting key information from complex tables and generating valuable data reports is a common demand.

[0003] However, existing data extraction and analysis methods rely heavily on manual operations. Specifically, users first need to manually screen and organize the field combinations that need to be analyzed according to their needs. This process is not only time-consuming and labor-intensive, but human factors may also lead to inaccurate or missed field selections. Next, in order to generate data reports, users need to use a variety of data processing methods such as functions and classifications to convert data into intuitive charts. This step also requires a lot of manual operations and places high demands on the user's data processing capabilities. Finally, after generating the charts, users also need to summarize the charts and layout them to generate data reports. This step is time-consuming and requires users to have certain design and typesetting capabilities.

[0004] It can be seen that the existing data analysis methods and data report production methods rely heavily on manual operations, which are not only time-consuming and labor-intensive, but the data reports produced are also prone to defects in accuracy and completeness due to human factors. Summary of the Invention

[0005] The present application provides a data report generation method, device and storage medium to solve the technical problems that existing data analysis methods and data report production methods rely heavily on manual operations, are time-consuming and labor-intensive, and the produced data reports also have defects in accuracy and completeness.

[0006] In a first aspect, the present application provides a method for generating a data report, the method comprising:

[0007] Select the target field combination from the fields of the target data table;

[0008] Generating material information of the target field combination based on the data items of each field in the target field combination, wherein the material information includes a visual chart and a text description;

[0009] The material information of the target field combination is integrated to form a data report of the target data table.

[0010] In a possible implementation, selecting a target field combination from the fields of the target data table includes:

[0011] Classifying the fields related to the same analysis target in the target data table into the same field class to obtain multiple field classes;

[0012] Build a target field combination using the fields in the field class.

[0013] In a possible implementation, the constructing a target field combination using the fields in the field class includes:

[0014] Combine the fields in the field class according to the set field combination type to obtain multiple candidate field combinations;

[0015] Determining a recommendation score for the candidate field combination;

[0016] A target field combination is selected from a plurality of the candidate field combinations according to the recommendation scores of the candidate field combinations.

[0017] In a possible implementation, determining the recommendation score of the candidate field combination includes:

[0018] Determining a recommendation score for each field in the candidate field combination, and determining a first recommendation score for the candidate field combination based on the recommendation score of each field;

[0019] determining a second recommendation score of the candidate field combination according to the dimension fields in the candidate field combination;

[0020] A final recommendation score of the candidate field combination is determined according to the first recommendation score and the second recommendation score.

[0021] In a possible implementation, the target field combination selected from the fields of the target data table includes:

[0022] Obtaining the associated data table of the target data table;

[0023] Obtaining a historically recommended field combination of the associated data table;

[0024] According to the historically recommended field combinations, a target field combination is selected from the fields of the target data table.

[0025] In a possible implementation, selecting a target field combination from the fields of a target data table based on the historically recommended field combination includes:

[0026] Selecting, from the fields of the target data table, a field combination that matches the historically recommended field combination as a candidate field combination;

[0027] Determining a recommendation score for the candidate field combination;

[0028] A target field combination is selected from a plurality of the candidate field combinations according to the recommendation scores of the candidate field combinations.

[0029] In a possible implementation, determining the recommendation score of the candidate field combination includes:

[0030] Determining a recommendation score for each field in the candidate field combination, and determining a first recommendation score for the candidate field combination based on the recommendation score of each field;

[0031] Determining a historical recommendation count of the candidate field combination, and determining a second recommendation score of the candidate field combination based on the historical recommendation count;

[0032] A final recommendation score of the candidate field combination is determined according to the first recommendation score and the second recommendation score.

[0033] In a possible implementation, generating the material information of the target field combination based on the data items of each field in the target field combination includes:

[0034] Determining a combination type of the target field combination according to the number of dimension fields and metric fields in the target field combination;

[0035] A target analysis rule is determined according to the combination type, and data items corresponding to the target field combination are analyzed according to the target analysis rule to obtain material information of the target field combination.

[0036] In a possible implementation, integrating the material information of the target field combination to form a data report of the target data table includes:

[0037] Determining a graph category of the target field combination;

[0038] Based on the atlas category of the target field combination, the material information of the target field combination is sorted according to a preset atlas category order and a set content organization structure to form a data report of the target data table.

[0039] In a possible implementation, determining the graph category of the target field combination includes:

[0040] If the type of the first dimension field satisfies a type-category mapping relationship in a preset type-category mapping relationship set, determining a category corresponding to the first dimension field; wherein the type-category mapping relationship describes a type and a category with a unique corresponding relationship; and the first dimension field is a dimension field in the target field combination;

[0041] If the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, determining the category corresponding to the first dimension field based on the correlation between the first dimension field and the dimension field of the determined category in the target data table and the recommendation score of the first dimension field;

[0042] The category corresponding to the first dimension field is determined as the graph category of the target field combination.

[0043] In a possible implementation, based on the atlas category of the target field combination, the material information of the target field combination is sorted according to a preset atlas category order and a set content organization structure to form a data report of the target data table, including:

[0044] Based on the atlas category of the target field combination, the material information of the target field combination is input into the data report generation model according to a preset atlas category order to obtain a data report of the target data table.

[0045] In a second aspect, the present application provides a data report generating device, the device comprising:

[0046] The field combination module is used to select the target field combination from the fields of the target data table;

[0047] A material generation module, configured to generate material information of the target field combination based on the data items of each field in the target field combination, wherein the material information includes a visual chart and a text description;

[0048] The material integration module is used to integrate the material information of the target field combination to form a data report of the target data table.

[0049] In a possible implementation, the field combination module includes:

[0050] A field classification unit is used to classify the fields related to the same analysis target in the target data table into the same field class to obtain multiple field classes;

[0051] The field combination unit is used to construct a target field combination using the fields in the field class.

[0052] In a possible implementation, the field combination unit includes:

[0053] A combining unit, configured to combine the fields in the field class according to a set field combination type to obtain a plurality of candidate field combinations;

[0054] an evaluation unit, configured to determine a recommendation score for the candidate field combination;

[0055] The screening unit is configured to select a target field combination from a plurality of the candidate field combinations according to the recommendation scores of the candidate field combinations.

[0056] In a possible implementation, the evaluation unit is specifically configured to:

[0057] Determining a recommendation score for each field in the candidate field combination, and determining a first recommendation score for the candidate field combination based on the recommendation score of each field;

[0058] determining a second recommendation score of the candidate field combination according to the dimension fields in the candidate field combination;

[0059] A final recommendation score of the candidate field combination is determined according to the first recommendation score and the second recommendation score.

[0060] In a possible implementation, the field combination module includes:

[0061] An associated table acquisition unit, configured to acquire an associated data table of the target data table;

[0062] A historical recommendation combination acquisition unit, configured to acquire a historical recommendation field combination of the associated data table;

[0063] The combination unit is used to select a target field combination from the fields of the target data table according to the historically recommended field combination.

[0064] In a possible implementation, the combination unit includes:

[0065] a combining subunit, configured to select, from the fields of the target data table, a field combination that matches the historically recommended field combination as a candidate field combination;

[0066] An evaluation subunit, configured to determine a recommendation score for the candidate field combination;

[0067] The screening subunit is configured to select a target field combination from a plurality of the candidate field combinations according to the recommendation scores of the candidate field combinations.

[0068] In one possible implementation, the evaluation subunit is specifically configured to:

[0069] Determining a recommendation score for each field in the candidate field combination, and determining a first recommendation score for the candidate field combination based on the recommendation score of each field;

[0070] Determining a historical recommendation count of the candidate field combination, and determining a second recommendation score of the candidate field combination based on the historical recommendation count;

[0071] A final recommendation score of the candidate field combination is determined according to the first recommendation score and the second recommendation score.

[0072] In a possible implementation, the material generation module includes:

[0073] a combination type determining unit, configured to determine a combination type of the target field combination according to the number of dimension fields and metric fields in the target field combination;

[0074] A generating unit is configured to determine a target analysis rule according to the combination type, analyze the data items corresponding to the target field combination according to the target analysis rule, and generate material information of the target field combination according to the analysis result.

[0075] In a possible implementation, the material integration module includes:

[0076] A graph category determination unit, configured to determine the graph category of the target field combination;

[0077] The integration unit is used to organize the material information of the target field combination according to the preset atlas category order and the set content organization structure based on the atlas category of the target field combination to form a data report of the target data table.

[0078] In a possible implementation, the atlas category determination unit is specifically configured to:

[0079] If the type of the first dimension field satisfies a type-category mapping relationship in a preset type-category mapping relationship set, determining a category corresponding to the first dimension field; wherein the type-category mapping relationship describes a type and a category with a unique corresponding relationship; and the first dimension field is a dimension field in the target field combination;

[0080] If the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, determining the category corresponding to the first dimension field based on the correlation between the first dimension field and the dimension field of the determined category in the target data table and the recommendation score of the first dimension field;

[0081] The category corresponding to the first dimension field is determined as the graph category of the target field combination.

[0082] In a possible implementation manner, the integration unit is specifically used to:

[0083] Based on the atlas category of the target field combination, the material information of the target field combination is input into the data report generation model according to a preset atlas category order to obtain a data report of the target data table.

[0084] In a third aspect, the present application provides an electronic device comprising: a processor and a memory, wherein the processor is configured to execute a data report generation program stored in the memory to implement the data report generation method described in any one of the first aspects.

[0085] In a fourth aspect, the present application provides a storage medium storing one or more programs, which can be executed by one or more processors to implement the data report generation method described in any one of the first aspects.

[0086] The above-mentioned technical solution provided by the embodiment of the present application has the following advantages compared with the existing technology: the method provided by the embodiment of the present application automatically selects the target field combination from the fields of the target data table, and generates visual charts and text descriptions based on the data items of each field in the target field combination, and finally integrates the visual charts and text descriptions of these target field combinations to form a data report. This automated process can quickly process large amounts of data and generate intuitive and easy-to-understand data reports, thereby greatly reducing the time for manual participation in data analysis and report production, and significantly improving the overall efficiency of data analysis. At the same time, it can present complex data relationships in an intuitive and easy-to-understand manner, thereby enhancing the readability of the data report and helping users understand the information and trends behind the data more quickly. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0088] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0089] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0090] Figure 1 A flow chart of an embodiment of a data report generation method provided in an embodiment of the present application;

[0091] Figure 2 is an example of a target data table;

[0092] Figure 3 for Figure 1 Examples of visualization charts corresponding to each target field combination in the data table shown;

[0093] Figure 4 A flowchart of another method for generating a data report according to an embodiment of the present application;

[0094] Figure 5 A flowchart of another method for generating a data report according to an embodiment of the present application;

[0095] Figure 6 A flowchart of another method for generating a data report according to an embodiment of the present application;

[0096] Figure 7 A block diagram of an embodiment of a data report generating device provided in an embodiment of the present application;

[0097] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0098] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0099] The disclosure below provides many different embodiments or examples for implementing different structures of the present application. In order to simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, these are merely examples and are not intended to limit the present application. In addition, the present application may repeat reference numbers and / or letters in different examples. Such repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or settings discussed.

[0100] In order to solve the technical problems in the prior art that data analysis methods and data report production methods rely heavily on manual operations, are time-consuming and labor-intensive, and the produced data reports also have flaws in accuracy and completeness, the present application provides a data report generation method, device and storage medium, which can automatically process data and generate intuitive and easy-to-understand data reports, thereby greatly reducing the time for manual participation in data analysis and report production, significantly improving the overall efficiency of data analysis, and at the same time being able to present complex data relationships in an intuitive and easy-to-understand manner, thereby enhancing the readability of data reports and helping users to more quickly understand the information and trends behind the data.

[0101] Figure 1 This is a flow chart of an embodiment of a data report generation method provided in an embodiment of the present application. Figure 1 As shown, the method includes the following steps:

[0102] Step 101: Select a target field combination from the fields of a target data table.

[0103] The target data table refers to the data table to be analyzed. It can be formatted in a variety of formats, including but not limited to spreadsheet formats such as .xls and .xlsx. A target data table typically consists of two main elements: fields and data items. Fields in a data table can be categorized into metric fields and dimension fields based on the information they carry. Metric fields are quantitative fields used to describe quantities or scales. The data types of data items in metric fields can be numeric (including integers and decimals) or percentages. Metric fields in a data table form an indispensable quantitative foundation for data analysis. Dimension fields play a categorization and description role in a data table; that is, dimension fields are used to categorize and describe non-quantitative fields in a data table. The data types of data items in dimension fields can be text (such as region and name) or date (such as year, month, day, quarter, and week). Depending on the data type of the data items in dimension fields, dimension fields in a data table can be categorized into time dimension fields and non-time dimension fields. Time dimension fields are fields whose data items are of the time type and are used to record time information. Time dimension fields form the basis for identifying trends or patterns in data evolution over time during data analysis. Non-time dimension fields refer to other dimension fields in the data table except the time dimension field.

[0104] Based on this, the target field combination, as the field combination for subsequent in-depth analysis and reporting, is particularly important. This combination is typically required to be able to deeply explore and reveal hidden information or trends in the target data table from a specific analytical dimension or perspective. The target field combination may include one or more fields, including metric fields, dimension fields, or a combination of dimension and metric fields. Different target field combinations represent different analytical dimensions or perspectives.

[0105] For example, see Figure 1 , is an example of a target data table.

[0106] By executing step 101, Figure 1 The following target field combinations are selected from the target data table of the example: {processing date, quantity}, {complaint date, quantity}, {customer evaluation, quantity}, {complaint type, customer evaluation, quantity}, {SKU, customer evaluation, quantity}, {complaint type, quantity}, {complaint reason, quantity}, {product SPU, quantity}.

[0107] As for how to select the target field combination from the fields of the target data table, we will explain it in the following text. Figure 4 and Figure 5 The process shown is explained in detail and will not be described in detail here.

[0108] Step 102: Generate material information of the target field combination based on the data items of each field in the target field combination, wherein the material information includes a visual chart and a text description.

[0109] The purpose of step 102 is to generate a set of material information based on the data items in each field of the target field combination. This material information can present the data in an intuitive and easy-to-understand manner, enhancing the depth and breadth of data interpretation. The material information includes visual charts and text descriptions. These two types of information form a visualization scheme for the analysis results of the target field combination.

[0110] When generating material information for target field combinations, you can select the appropriate chart type based on the data characteristics and analysis requirements (for example, a bar chart is used to compare the quantities of each category, a line chart is good at revealing the trend of data changes over time, and a pie chart intuitively shows the proportion of each part in the whole). At the same time, a concise text description is generated for each chart. This description not only explains the data content and trends shown in the chart in detail, but also further enhances the readability and comprehensibility of the data.

[0111] On this basis, the material information of the target field combination lays the foundation for subsequent data analysis and data report generation, and is an important tool for users to understand data and gain insights into trends.

[0112] For example, see Figure 3 ,for Figure 1 Examples of visualization charts corresponding to each target field combination in the data table shown. Also, see Table 1 below for examples of text descriptions corresponding to each target field combination:

[0113] Table 1

[0114]

[0115]

[0116] As for how to select the target field combination from the fields of the target data table, we will explain it in the following text. Figure 6 The process shown is explained in detail and will not be described in detail here.

[0117] Step 103: Integrate the material information of the target field combination to form a data report of the target data table.

[0118] Step 103 is intended to systematically integrate the material information of the target field combination, thereby generating a complete and clearly structured data report.

[0119] In one embodiment, the source information can be integrated according to a predefined report structure. First, the content corresponding to each component of the report structure is extracted from the source information of the target field combination. Then, the content of each component is sorted and integrated according to the predefined report structure, thereby forming a complete and clearly structured data report.

[0120] For example, the established report structure includes key sections such as the document title, data summary, data brief, detailed analysis, and data recommendations. The data summary section provides a concise and clear overview of the report's core content and key findings. The data brief uses concise language to quickly present key data characteristics and trends. The detailed analysis section uses charts and text descriptions combining multiple target fields to deeply analyze the data's inherent patterns and potential connections. The data recommendations section provides targeted suggestions through text descriptions.

[0121] As for how to integrate the material information of the target field combination to form the data report of the target data table, a detailed explanation will be given below through specific embodiments, which will not be described in detail here.

[0122] The technical solution provided in the embodiment of the present application automatically selects target field combinations from the fields of the target data table, and generates visual charts and text descriptions based on the data items of each field in the target field combination, and finally integrates the visual charts and text descriptions of these target field combinations into a data report. This automated process can quickly process large amounts of data and generate intuitive and easy-to-understand data reports, thereby greatly reducing the time for manual participation in data analysis and report production, significantly improving the overall efficiency of data analysis, and at the same time being able to present complex data relationships in an intuitive and easy-to-understand manner, thereby enhancing the readability of the data report and helping users understand the information and trends behind the data more quickly.

[0123] In one embodiment, selecting a target field combination from the fields of the target data table includes: selecting different target field combinations from the fields of the target data table according to different analysis targets.

[0124] Specifically, see Figure 4 , Figure 4 This is a flow chart of another embodiment of a data report generation method provided in an embodiment of the present application. Figure 4 The process shown in Figure 1 Based on the process shown in the figure, a method for selecting a target field combination from the fields of the target data table is described. Figure 4 As shown, the following steps are included:

[0125] Step 401: Classify the fields related to the same analysis target in the target data table into the same field class to obtain multiple field classes.

[0126] Step 401 identifies fields in the target data table that are related to the same analysis objective and groups them into the same field category, making the data table structure clearer and facilitating subsequent data processing and analysis. Fields related to an analysis objective are fields that directly or indirectly support, measure, or describe a specific analysis objective.

[0127] In one embodiment, a pre-trained large model is used to classify fields in the target data table that are related to the same analysis target into the same field class, thereby obtaining multiple field classes. For example, the following large model prompt word template is designed:

[0128] [Task] You are an expert and need to combine dimensions and metrics to generate new field groups. Requirements: 1) Dimensions and metrics can be combined repeatedly, up to n groups. Provide a list of fields for each group, and provide the purpose and reason for the grouping. 2) The number of fields in each group must be greater than {^minCount^} and less than {^maxCount^}. [Dimension] {^dimTitle^} [Metric] {^dataTitle^} [Output Format] Output results strictly in JSON format. No additional content is allowed. [Bold text is not allowed.] [Output Format Example] [{"titleList":[],"goal":"","reason":""}]

[0129] When applying, input parameters into the above prompt word template according to the target data table and actual analysis requirements. Figure 2 Taking the data table shown as an example, the input parameters include:

[0130]

[0131] For ease of understanding, the output of the above large model is presented in Table 2 below:

[0132] Table 2

[0133]

[0134] Step 402: Build a target field combination using the fields in each field class.

[0135] Since the fields in the same field class are closely related to the same analysis target, using the fields in the same field class to construct target field combinations can significantly improve the pertinence and efficiency of data analysis.

[0136] In one embodiment, the specific implementation of constructing a target field combination using fields in a field class includes: combining the fields in the field class according to a set field combination type to obtain multiple candidate field combinations; determining the recommendation scores of the candidate field combinations; and selecting a target field combination from the multiple candidate field combinations based on the recommendation scores of the candidate field combinations.

[0137] Among them, according to the analysis objectives and business needs, a series of field combination types are preset, and these types may be determined based on factors such as the nature of the data, business logic, or analysis habits. For example, some combination types may focus on time series analysis, and some combination types may focus on cross-analysis between different dimensions. For example, the preset field combination types include: 1 dimension 1 measure (that is, one dimension field and one measure field constitute a field combination), 1 dimension (that is, one dimension field independently constitutes a field combination), 2 dimensions 1 measure (that is, two dimension fields and one dimension field constitute a field combination), and 2 dimensions (that is, two dimension fields constitute a field combination).

[0138] Combining fields in a field class according to the specified field combination types to obtain multiple candidate field combinations means exhaustively trying to combine the fields in the field class according to each specified field combination type. This means that the actual meaning or feasibility of the combinations is not considered during this step; the fields are simply mechanically combined according to different combination types. For this reason, the field combinations obtained in this step are considered candidate field combinations and do not directly participate in subsequent data analysis and report generation.

[0139] Furthermore, for each candidate field combination, a recommendation score is determined for it. The recommendation score refers to the score value obtained by judging whether the candidate field combination can be used in data analysis. The larger the score value, the greater the value of the corresponding candidate field combination for analysis, and the more recommended it is. On this basis, the highest-scoring candidates can be selected from multiple candidate combinations as target field combinations based on the recommendation scores of the candidate field combinations. The target field combination will participate in subsequent data analysis and report generation. It should be noted that the recommendation score in this embodiment means that the larger the score value, the more recommended it is for analysis. In other embodiments, it can be other settings, such as the larger the score value, the less recommended it is. This is not specifically limited here.

[0140] This embodiment of the application implements an efficient and intelligent method for constructing target field combinations through a series of steps, including setting the field combination type, generating candidate field combinations, determining the recommendation score, and selecting the target field combination. This method not only improves data processing efficiency but also ensures that the final selected combination has high practicality and value.

[0141] Figure 4The process shown here first groups the fields in the target data table that are closely related to a specific analysis objective into a single field class, thereby forming multiple field classes. Then, for each field class, the fields within it are used to construct the target field combination. This approach not only ensures the high relevance of the selected fields to the target analysis task but also greatly facilitates subsequent analysis and the compilation of data reports, resulting in more accurate and insightful analysis results.

[0142] In one embodiment, determining a recommendation score for a candidate field combination includes determining recommendation scores for each candidate field combination based on information from different modalities, and then determining a final recommendation score for the candidate field combination based on the recommendation scores determined for each modality. This approach comprehensively considers multiple information sources, improving the accuracy and comprehensiveness of recommendations.

[0143] As an optional implementation, a recommendation score is determined for each field in a candidate field combination, and a first recommendation score for the candidate field combination is determined based on the recommendation scores of each field. Subsequently, a second recommendation score for the candidate field combination is determined based on the dimension fields in the candidate field combination. If the candidate field combination includes one dimension field, the number of occurrences of the dimension field in multiple candidate field combinations is obtained, and the second recommendation score for the candidate field combination is determined based on the number of occurrences. If the candidate field combination includes two dimension fields, a cross-tabulation is constructed based on the two dimension fields, and the second recommendation score for the candidate field combination is determined based on the data fill rate of the cross-tabulation. Finally, a final recommendation score for the candidate field combination is determined based on the first and second recommendation scores. For example, the average of the first and second recommendation scores can be calculated as the final recommendation score for the candidate field combination. Alternatively, the first and second recommendation scores can be weighted and summed to obtain the final recommendation score for the candidate field combination. In addition to this, other implementations that combine the first and second recommendation scores to obtain the final recommendation score for the candidate field combination can also be used in practical applications, and this embodiment of the present application is not limited thereto.

[0144] Determining the recommendation score of the dimension field in the candidate field combination includes: determining the recommendation score of the dimension field according to the type of the dimension field, characteristic information of the dimension field, and characteristic information of the target data table.

[0145] Dimension field types include, but are not limited to, date and time types, date types, time types, and string types. String types include, specifically, person name types, place name types, gerund types, English types, text-type and numeric types, other types, and mixed types.

[0146] The characteristic information of a dimension field includes the data distribution of each dimension field, the distribution of each dimension field in the target data table, and characteristic information about the number of dimension fields or metric fields adjacent to the dimension field. Exemplarily, first characteristic information of a first dimension field in the target data table is obtained based on the data distribution in the cell set corresponding to the first dimension field in the target data table and the distribution of the first dimension field in the target data table; wherein the first dimension field is any dimension field in the target data table; and second characteristic information of the first dimension field is obtained based on the type of each dimension field in the target data table and the relationship between the first dimension field and other dimension fields in the target data table.

[0147] In this embodiment, the first characteristic information of the first dimension field refers to descriptive characteristic information of the first dimension field. The obtained first characteristic information can be used to determine whether the first dimension field has analytical significance. For example, in the remark information, if the data length varies, some are empty, the average value is small, the standard deviation is large, the minimum value is zero, and the maximum value is large, then this type of dimension field is often not valuable for analysis.

[0148] Specifically, the first characteristic information of the first dimension field includes descriptive characteristic information and positional characteristic information of the first dimension field; accordingly, based on the data distribution in the cell set corresponding to the first dimension field in the target data table and the distribution of the first dimension field in the data table to be processed, the first characteristic information of the first dimension field in the target data table is obtained, including: obtaining the descriptive characteristic information of the first dimension field; obtaining the positional characteristic information of the first dimension field. Among them, the descriptive feature information of the first dimension field is obtained, and the descriptive features include at least any one of the following: the index value of the first dimension field; wherein the index value is used to describe the position of the first dimension field in the target data table; the number of non-repeated data in the cell set corresponding to the first dimension field; the number of cells in the cell set corresponding to the first dimension field; the average length of the data contained in each cell in the cell set corresponding to the first dimension field; the minimum length of the data contained in each cell in the cell set corresponding to the first dimension field; the maximum length of the data contained in each cell in the cell set corresponding to the first dimension field; the standard deviation of the length of the data contained in each cell in the cell set corresponding to the first dimension field; the average number of occurrences of non-repeated data in the cell set corresponding to the first dimension field; the minimum number of occurrences of non-repeated data in the cell set corresponding to the first dimension field; the maximum number of occurrences of non-repeated data in the cell set corresponding to the first dimension field; the standard deviation of the number of occurrences of non-repeated data in the cell set corresponding to the first dimension field. The position feature information includes at least any one of the following: in the target data table, comparing the index value of the first dimension field with the index value of the measurement field, obtaining the number of measurement fields whose index values are less than the index value of the first dimension field and the number of measurement fields whose index values are greater than the index value of the first dimension field; in the target data table, comparing the index value of the first dimension field with the index values of other dimension fields, obtaining the number of other dimension fields whose index values are less than the index value of the first dimension field and the number of other dimension fields whose index values are greater than the index value of the first dimension field.

[0149] The second characteristic information of the first dimension field represents the relationship between the first dimension field and other fields in the target data table. This second characteristic information is derived based on the type of each dimension field in the target data table and the relationship between the first dimension field and other dimension fields. The relationship between the first dimension field and other dimension fields can be a parent-child relationship, meaning that a father can have multiple sons, but a son cannot have multiple fathers. Specifically, according to the type of each dimension field in the target data table and the relationship between the first dimension field and other dimension fields in the target data table, the second characteristic information of the first dimension field is obtained, including: according to the type of each dimension field in the target data table, the number of dimension fields with the same type as the first dimension field is obtained; according to the relationship between the first dimension field and other dimension fields in the target data table, the number of dimension fields in the target data table that have a father relationship with the first dimension field is obtained; according to the relationship between the first dimension field and other dimension fields in the target data table, the number of dimension fields in the target data table that have a son relationship with the first dimension field is obtained; according to the relationship between the first dimension field and other dimension fields in the target data table, the number of dimension fields in the target data table that have the same type as the first dimension field and are in a father relationship with the first dimension field is obtained; according to the relationship between the first dimension field and other dimension fields in the target data table, the number of dimension fields in the target data table that have the same type as the first dimension field and are in a son relationship with the first dimension field is obtained; the enumeration value of the field type is obtained; and according to the obtained data, the second characteristic information of the first dimension field in the target data table is obtained. In this embodiment, the second characteristic information of the first dimension field includes the number of dimension fields of the same type as the first dimension field in the target data table, the number of dimension fields that are in a parent relationship with the first dimension field, the number of dimension fields that are in a child relationship with the first dimension field, the number of dimension fields that are of the same type as the first dimension field and are in a parent relationship, the number of dimension fields that are of the same type as the first dimension field and are in a child relationship, and enumerated values of the field type. Among them, the parent-child relationship in the data table refers to the situation where one party can contain another party or multiple parties. If dimension field A is in a parent relationship with the first dimension field, it means that dimension field A contains the first dimension field; or if dimension field A is in a child relationship with the first dimension field, it means that the first dimension field contains dimension field A.

[0150] It's important to note that parent-child relationships refer to the concepts of dimension hierarchies and the concepts of similar dimension hierarchies. Concept hierarchies are a sequence of mappings that map lower-level concepts to higher-level, more general concepts. For example, "Canada" includes Vancouver, and these two concepts are in a parent-child relationship. A parent relationship refers to the relationship between a dimension field with a higher-level concept and a dimension field with a lower-level concept. For example, a province is the parent of both a district and a city; the relationship between a province and a district is a parent relationship. A child relationship refers to the relationship between a dimension field with a lower-level concept and a dimension field with a higher-level concept. For example, the relationship between a city and a province is a child relationship; the city is the son of the province. If a dimension field has many parent relationships, it may not be very valuable for analysis.

[0151] The characteristic information of the target data table includes the descriptive characteristic information of the target data table, the classification characteristic information of the data table, etc., which are used to express the basic information of the target data table. Among them, the descriptive characteristics are mainly used to describe the characteristic information of the target data table situation, and the descriptive characteristics can be extracted through the descriptive statistics function module in the table. The descriptive characteristics include at least any one of the following: the number of non-empty cell sets in the target data table; the maximum number of non-empty cells contained in each cell set in the target data table; the minimum number of non-repeated data contained in each cell set in the target data table. The classification characteristics are information that describes each field type in the target data table, specifically including: enumeration values of each field type, the number of fields of the name type, the number of fields of the place name type, the number of fields of the numerical type, the number of fields of the date type, the number of fields of the time type, and the number of fields of the gerund type. It should be noted that the enumeration value refers to a method of defining an ordered set by pre-defining identifiers that list all values. The order of these values is consistent with the order of the identifiers in the enumeration type description. Assume that the enumeration value is in the form of <identifier1>=<type1>, such as <n1>= <Time type>, <n2>= <Date type>, <n3>= <value type>, etc., assuming that the types of the fields in the target data table in this embodiment are time type, date type, noun type, and value type, the enumeration value of the target data table is<N1、N2、N5、N3> =<time type, date type, noun type, numerical type>.

[0152] The recommendation score of the dimension field is determined based on the determined type of dimension field, the characteristic information of the dimension field, and the characteristic information of the data table to be processed. The recommendation score refers to the score value obtained by judging whether the dimension field can be used in data analysis. The larger the score value, the greater the value of the corresponding dimension field for analysis and the more recommended it is. It should be noted that in this embodiment, the recommendation score means that the larger the score value, the more recommended it is for analysis. In other embodiments, it can be other settings, such as the larger the score value, the less recommended it is. This is not specifically limited here.

[0153] In this embodiment, the recommendation score of the dimension field is determined based on the type of the dimension field, the characteristic information of the dimension field, and the characteristic information of the target data table, including: inputting the characteristic information of the target data table, the first characteristic information of the first dimension field, and the second characteristic information of the first dimension field into a pre-trained dimension field recommendation model to obtain the recommendation score of the first dimension field; wherein the dimension field recommendation model is trained based on the characteristic information of the sample data table, the first characteristic information of the second dimension field in the sample data table, the second characteristic information of the second dimension field, and the recommendation score label of the second dimension field; wherein the second dimension field is any one or more dimension fields in the sample data table.

[0154] The dimension field recommendation model is pre-trained using the random forest algorithm on training samples. A random forest is a classifier that uses multiple decision trees to train and predict training samples. The classifier is trained using the random forest algorithm on the pre-obtained feature information of the sample data table, the first feature information of the second dimension field in the sample data table, the second feature information of the second dimension field, and the recommendation score label of the second dimension field. The specific training method is not described in detail here.

[0155] Determining the recommendation score of the metric field in the candidate field combination includes: determining a category of the metric field, and determining the recommendation score of the metric field based on characteristic information of the target data table and characteristic information of the metric field.

[0156] Among them, the field of numeric type belongs to the measurement field, so the measurement field is set to a category, that is, the numeric type.

[0157] Obtaining characteristic information of the measurement field includes: obtaining first characteristic information of the first measurement field according to data distribution in a cell set corresponding to the first measurement field in the target data table and distribution of the first measurement field in the target data table; and obtaining second characteristic information of the first measurement field in the target data table according to data statistics in a cell set corresponding to the first measurement field in the target data table.

[0158] Obtaining first characteristic information of the first metric field in the target data table based on data distribution in the cell set corresponding to the first metric field in the target data table and distribution of the first metric field in the target data table includes at least any one of the following: obtaining an index value of the first metric field, wherein the index value is used to describe the position of the first metric field in the target data table; obtaining the number of non-repeated data in the cell set corresponding to the first metric field; obtaining the number of cells in the cell set corresponding to the first metric field; obtaining the number of non-empty cells in the cell set corresponding to the first metric field; comparing the index value of the first metric field with the index values of other metric fields in the target data table to obtain the number of other metric fields whose index values are less than the index value of the first metric field and the number of other metric fields whose index values are greater than the index value of the first metric field; comparing the index value of the first metric field with the index values of dimension fields in the target data table to obtain the number of dimension fields whose index values are less than the index value of the first metric field and the number of dimension fields whose index values are greater than the index value of the first metric field; and obtaining the first characteristic information of the first metric field in the target data table based on the obtained data.

[0159] In this embodiment, first characteristic information is obtained based on the data distribution in the cell set corresponding to the first metric field in the target data table and the distribution of the first metric field in the target data table. The first characteristic information includes: the index value of the first metric field, the number of non-repeated data in the cell set corresponding to the first metric field, the number of cells, and the number of non-empty cells. It should be noted that in this embodiment, it is necessary to obtain the index value of the metric field to determine the position of the metric field. The position of the metric field has a significant impact on data analysis. For example, in data summation, the data to the right of the metric field is more valuable for analysis. If there are no other metric fields to the left of the current metric field and there are other metric fields to the right, analysis is often not recommended because the metric field may be a serial number. The first metric field is any metric field in the target data table.

[0160] A recommendation score for the metric field is determined based on the determined metric field category, the metric field's characteristic information, and the characteristic information of the target data table. The recommendation score is a score calculated to determine whether the metric field can be used for data analysis. A higher score indicates that the corresponding metric field is more valuable for analysis and is more recommended. It should be noted that in this embodiment, the recommendation score indicates that a higher score indicates a higher recommendation for analysis. In other embodiments, this can be set differently, such as a higher score indicates a lower recommendation, and this is not specifically limited here.

[0161] In this embodiment, the recommendation score of the metric field is determined based on the category of the metric field, the characteristic information of the metric field, and the characteristic information of the target data table, including: inputting the characteristic information of the target data table, the first characteristic information of the first metric field, and the second characteristic information of the first metric field into a pre-trained metric field recommendation model to obtain the recommendation score of the first metric field; wherein the metric field recommendation model is trained based on the characteristic information of the sample data table, the first characteristic information of the second metric field in the sample data table, the second characteristic information of the second metric field, and the recommendation score label of the second metric field; wherein the second metric field is any one or more metric fields in the sample data table.

[0162] The metric field recommendation model is pre-trained using the random forest algorithm on training samples. This algorithm trains a classifier using the previously acquired feature information from the sample data table, the first feature information from the second metric field in the sample data table, the second feature information from the second metric field, and the recommendation score label for the second metric field. The specific training method is not described in detail here.

[0163] As an optional implementation, determining the first recommendation score of the candidate field combination according to the recommendation score of each field includes: calculating the average of the recommendation scores of each field, and determining the average as the first recommendation score of the candidate field combination.

[0164] The above describes how to determine the first recommendation score of a candidate field combination. The following describes how to determine the second recommendation score of a candidate field combination. This discussion is divided into two cases:

[0165] One case is: when the candidate field combination includes one dimension field, the number of occurrences of the dimension field in multiple candidate field combinations is obtained, and the second recommendation score of the candidate field combination is determined according to the number of occurrences.

[0166] In data analysis scenarios, the number of occurrences of a single dimension field often reflects its importance and information value. Therefore, when a candidate field combination contains only one dimension field (and optionally a metric field), its second recommendation score can be determined by counting the frequency of occurrence of that dimension field across all candidate field combinations. For example, all candidate field combinations can be traversed, the number of occurrences of each dimension field can be counted, and the resulting counts can be normalized to form the second recommendation score. Normalization ensures that the score is within a reasonable range, facilitating subsequent use.

[0167] By considering the number of occurrences of dimension fields, the system can more accurately identify dimension fields that appear frequently and may have higher information value, thereby improving the accuracy of recommendations.

[0168] Another case is: when the candidate field combination includes two dimension fields, a cross table is constructed based on the two dimension fields, and the second recommendation score of the candidate field combination is determined based on the data filling rate of the cross table.

[0169] The cross-tab can show the correlation and data integrity between two dimension fields, while the data fill rate reflects the richness of this related information. Then, for the candidate field combination containing two dimension fields, a corresponding cross-tab is constructed. The rows and columns of the cross-tab represent the different values of the two dimension fields, and the cells are filled with corresponding data values (such as count, sum, etc.). Subsequently, the proportion of non-empty cells in the cross-tab is counted, that is, the data fill rate. Finally, the calculated data fill rate is used as the second recommendation score for the candidate field combination. Similarly, the data fill rate can be normalized to ensure the rationality and comparability of the score.

[0170] By considering the data fill rate of the cross-tab, the system can identify candidate field combinations with more complete data and stronger correlation, thereby improving the data quality of the recommendation results.

[0171] In one embodiment, selecting a target field combination from the fields of the target data table includes: selecting the target field combination from the fields of the target data table according to historically recommended field combinations.

[0172] Specifically, see Figure 5 , Figure 5 This is a flow chart of another embodiment of a data report generation method provided in an embodiment of the present application. Figure 5 The process shown in Figure 1 Based on the process shown in the figure, another implementation method of selecting the target field combination from the fields of the target data table is described. Figure 5 As shown, the following steps are included:

[0173] Step 501: Obtain the associated data table of the target data table.

[0174] In step 501 , the purpose of obtaining the associated data table of the target data table is to provide additional context information for the subsequent selection of the target field combination, which helps to more accurately select the target field combination from the fields of the target data table.

[0175] Exemplarily, the associated data tables of the target data table may include the following two types of data tables:

[0176] 1. Data tables with similar header structures to the target data table. These tables contain the same or similar fields as the target data table. By analyzing the recommendations for these similar fields in the associated data table, we can provide auxiliary reference for the recommendations for the corresponding fields in the target data table.

[0177] Second, data tables that have business logic relationships with the target data table. For example, a related data table might be the result of pre- or post-processing of the target data table, or another data entity that participates in a business process with the target data table. By understanding these business logic relationships, you can more accurately assess the business value of each field in the target data table.

[0178] 3. A data table that records the status or change trends of the target data table over a certain period of time. By analyzing this historical data, you can identify the timeliness, stability, and potential development trends of field combinations, providing valuable reference for selecting current field combinations.

[0179] Different acquisition methods can be used for related data tables of different categories. For example, for the first type of data tables mentioned above, SQL query statements can be used to retrieve other data tables with similar structures to the target data table. This usually involves comparing table structures (such as field names and data types) to identify similarities. For the second type of data tables mentioned above, other data tables that have business logic associations with the target data table can be identified based on the upstream and downstream relationships of the business process. For the third type of data tables mentioned above, historical data tables corresponding to the target data table can be searched in a data warehouse or historical database. This is merely an exemplary explanation, and the embodiments of the present application are not limited to this.

[0180] Step 502: Obtain historical recommended field combinations of the associated data table.

[0181] Step 502 aims to collect field combinations in the associated data table that have been frequently recommended or have performed well in the past, i.e., historically recommended field combinations. These historically recommended field combinations may be based on past user behavior, business rules, insights from the data science team, or predictions from machine learning models, without limitation.

[0182] For example, historical records of the associated data table are retrieved from the database, and historically recommended field combinations of the associated data table are obtained from the historical records.

[0183] By collecting historical recommended field combinations of related data tables, we can help understand which field combinations have proven to be effective in the past, thereby providing valuable reference for the selection of current field combinations.

[0184] Step 503: Select a target field combination from the fields of the target data table based on the historically recommended field combinations.

[0185] In one embodiment, an exemplary implementation of selecting a target field combination from the fields of a target data table based on historically recommended field combinations includes: selecting a field combination that matches the historically recommended field combination from the fields of the target data table as a candidate field combination; determining the recommendation score of the candidate field combination; and selecting a target field combination from multiple candidate field combinations based on the recommendation score of the candidate field combination.

[0186] Among them, selecting a field combination that matches the historically recommended field combination from the fields of the target data table includes: checking whether each field in the historically recommended field combination directly appears in the target data table; if so, selecting the corresponding fields from the target data table to construct a subsequent field combination. It should be noted here that field names may have aliases between the historical data table and the target data table. Therefore, in the above-mentioned inspection process, "directly appearing" does not mean exactly the same, but also includes the case where there is a mapping relationship between them. In addition, the correspondence of the fields in the business logic also needs to be considered. For example, if a field in the historically recommended field combination represents sales, and the corresponding field in the target data table represents cost, then the two fields do not correspond in business logic, and the field combination should be excluded.

[0187] For example, suppose the historically recommended field combination is [Product Category, Sales Amount], and the target data table's field list is [Customer ID, Transaction Date, Product Category, Transaction Amount, Payment Method]. Based on the historically recommended field combination, the target field combination [Product Category, Transaction Amount] can be selected from the target data table's fields.

[0188] Among them, an exemplary implementation of determining the recommendation score of a candidate field combination includes: determining the recommendation score of each field in the candidate field combination, and determining a first recommendation score of the candidate field combination based on the recommendation score of each field; determining the historical number of recommendations for the candidate field combination, and determining a second recommendation score for the candidate field combination based on the historical number of recommendations; and determining a final recommendation score for the candidate field combination based on the first recommendation score and the second recommendation score.

[0189] Here, the specific implementation of determining the recommendation score of each field in the candidate field combination and determining the first recommendation score of the candidate field combination based on the recommendation score of each field can be referred to the above description and will not be repeated here.

[0190] The number of historical recommendations for a candidate field combination refers to how often the candidate field combination has been recommended or used in the past. The number of historical recommendations is an indicator of the effectiveness and popularity of a field combination.

[0191] The second recommendation score is a score assigned to a candidate field combination based on the number of historical recommendations. As an optional method, the historical recommendation counts can be normalized to form the second recommendation score. Normalization ensures that the score is within a reasonable range for subsequent use.

[0192] For example, the average of the first recommendation score and the second recommendation score can be calculated as the final recommendation score of the candidate field combination. Alternatively, the first recommendation score and the second recommendation score are weighted and summed to obtain the final recommendation score of the candidate field combination. Alternatively, the first recommendation score and the second recommendation score are multiplied to obtain the final recommendation score of the candidate field combination. In addition, in actual applications, other implementation methods of fusing the first recommendation score and the second recommendation score to obtain the final recommendation score of the candidate field combination can also be adopted, and the embodiments of the present application are not limited to this.

[0193] Figure 5 The process shown, by obtaining the associated data table of the target data table, obtaining the historical recommended field combinations of the associated data table, and selecting the target field combination from the fields of the target data table based on the historical recommended field combinations, can refer to the historical recommended field combinations of the associated data table to more accurately select those field combinations that are highly relevant to the target data table and can reveal deep business logic, thereby improving the accuracy and effectiveness of data analysis.

[0194] In one embodiment, based on the data items of each field in the target field combination, material information of the target field combination is generated, including: determining the target analysis rules corresponding to the target field combination, analyzing the data items corresponding to the target field combination according to the target analysis rules, and obtaining the material information of the target field combination based on the analysis results.

[0195] Specifically, see Figure 6 , Figure 6 This is a flow chart of another embodiment of a data report generation method provided in an embodiment of the present application. Figure 6 The process shown in Figure 1 Based on the process shown in the figure, this paper describes how to generate the target field combination material information based on the data items of each field in the target field combination. Figure 6 As shown, the following steps are included:

[0196] Step 601: Determine the combination type of the target field combination according to the number of dimension fields and metric fields in the target field combination.

[0197] Step 602: Determine a target analysis rule according to the combination type, and analyze the data items corresponding to the target field combination according to the target analysis rule to obtain material information of the target field combination.

[0198] First of all, the target field combination here can include execution Figure 4 The process and Figure 5 The summary of target field combinations obtained by the process.

[0199] A combination type refers to a field combination category classified according to the number of dimension fields and metric fields in the target field combination. In one embodiment, the following three combination types are set: a target field combination containing only metric fields, a target field combination containing a time dimension field, and a target field combination containing no time dimension field.

[0200] Target analysis rules are a series of guiding principles developed based on the combination type and data analysis purpose, which are used to guide how to analyze the data items corresponding to the target field combination. These rules may include: Data aggregation rules: specify how to perform aggregation calculations on metric fields (such as sum, average, maximum, minimum, etc.), as well as the aggregation level (such as by day, week, month, etc.); Data filtering rules: define which data items should be included in the analysis and which should be excluded. This may involve filtering dimension fields, such as only analyzing data for a specific time period or a specific region; Data sorting rules: specify how to sort the analysis results to highlight the most important information. This may be based on the value of the metric field or the order of the dimension fields; Data visualization rules: determine how to present the analysis results in the form of graphics or tables for easy understanding and interpretation; and so on.

[0201] The following discusses the data analysis process under different combination types:

[0202] The first combination type: The target field combination contains only metric fields. This can be divided into two cases: the target field combination contains only one metric field, and the target field combination contains two or more metric fields.

[0203] If there is only one metric field in the target field combination, first determine the data distribution of the metric field data in the data table to be analyzed. See Table 3 for an illustration of the data distribution and its determination method:

[0204] Table 3

[0205]

[0206]

[0207] Subsequently, the measurement fields are statistically analyzed according to the data distribution method to obtain statistical analysis data. Among them, for the normal distribution, the following statistical analysis data can be obtained: upper quartile, lower quartile, and 95% confidence interval, etc. For the 90 / 100 rule, the following statistical analysis data can be obtained: maximum value, minimum value, and values greater than or equal to 88% of the total, etc. For the 80 / 20 rule, the following statistical analysis data can be obtained: maximum value, minimum value, and values greater than or equal to 77% of the total, etc. In addition, for any of the data distribution methods exemplified in Table 3 above (including no clear distribution), the following statistical analysis data can be obtained: number of cells, number of non-blank cells, sum of non-blank cell values, average of non-blank cell values, and median of non-blank cell values, etc. Generate material information based on the statistical analysis data.

[0208] When there are two or more metric fields in the target field combination, for each metric field, multiple data distribution intervals corresponding to the metric field are determined based on the metric field data in the metric field. In one embodiment, the multiple data distribution intervals corresponding to the metric field are determined in the following manner: first, the maximum value in the metric field is subtracted from the minimum value, and the number of intervals is determined based on the difference. For example, when the difference is greater than or equal to 100, the number of intervals is determined to be 10, and when the difference is less than or equal to 20, the number of intervals is determined to be 5. In other cases, the number of intervals is obtained by dividing the number of metric field data by a set value (for example, 8) and rounding. Then, the interval width is determined based on the number of intervals, where the interval width refers to the difference between the upper limit and the lower limit of the interval. Finally, the range of each interval, that is, the multiple data distribution intervals mentioned above, can be obtained based on the maximum value, the minimum value, and the interval width.

[0209] Subsequently, the number of metric field data falling within each data distribution interval is determined, and a data distribution chart corresponding to the metric field is generated based on the multiple data distribution intervals and numbers. The above data distribution chart can be a bar chart, pie chart, etc., which is not limited in this embodiment of the application.

[0210] The second combination type: the target field combination includes the time dimension field

[0211] A time dimension field is a field whose data type is time and is used to record time information. Time dimension fields form the basis for analyzing trends or patterns in data over time.

[0212] In one embodiment, the metric data items in the metric field are processed according to different time intervals based on the time data items in the time dimension field, and a visual analysis result of the data table is generated based on the processing results of the target metric field in different time intervals.

[0213] The different time intervals include at least two of the following time intervals: day, week, month, quarter, year, and other customized time periods (such as the last week, the last two weeks, the last 30 days, etc.).

[0214] The above processing refers to aggregation processing. Taking aggregation processing as an example, the aggregation processing method can be summation, averaging, maximum / minimum value calculation, or more complex statistical operations. The specific processing method depends on the analysis requirements and business logic and is not limited here.

[0215] The visual analysis results of the data table include, but are not limited to: the aggregation results of the metric fields in different time intervals, the interpretation data generated based on the aggregation results of the metric fields in different time intervals, the interpretation data generated based on the aggregation results of the metric fields in a certain time interval, and so on. The interpretation data is the derived data obtained after in-depth interpretation and analysis of the aggregation results, such as growth rate, month-on-month change, year-on-year change, etc. By interpreting the data, the reasons and trends of the changes in the target metric fields can be revealed in more depth, thereby providing more powerful data support for user decision-making. In an exemplary solution, a trained artificial intelligence model can be used to conduct in-depth interpretation and analysis of the aggregation results to obtain interpretation data, which is not limited here. The presentation form of the visual analysis results includes, but is not limited to: a combination of one or more forms such as data tables, visual charts, and text reports. Among them, the visual charts can be in the form of charts, curve charts, bar charts, pie charts, etc. The specific form can depend on the characteristics of the data and the purpose of the analysis, which is not limited here.

[0216] When analyzing data trends over time, the year-over-year (YoY) growth rate is a common and important metric. It provides an intuitive way to compare data changes between two time periods, revealing growth or decline in data between the two time periods. The year-over-year growth rate is typically expressed as a percentage or ratio and is calculated as follows:

[0217] Month-on-month change rate = (current period value - previous period value) / previous period value * 100%

[0218] Accordingly, in one embodiment, generating a visual analysis result of a data table based on the processing results of a metric field in different time intervals may include performing the following processing for each set time interval: determining, based on the aggregation results of the metric field in that time interval, month-over-month change information of the metric field corresponding to that time interval; and then generating a visual analysis result of the data table corresponding to that time interval based on the processing results of the metric field in that time interval and the month-over-month change information. The month-over-month change information includes the month-over-month change rate of the processing results of the metric field in two consecutive time periods.

[0219] In one embodiment, the text description of the target field combination including the time dimension field includes but does not include: description of data rise and fall trends, growth turning points and the highest value reached, decline turning points and the lowest value reached, etc.

[0220] By processing measurement data according to the time dimension and automatically generating a solution for visual analysis results, users do not need to master and set up various complex date functions, data aggregation formulas or visualization tools, etc., which greatly reduces the threshold and complexity of data analysis, improves the efficiency and accuracy of data analysis, and provides users with a more convenient and intuitive data analysis method.

[0221] The third combination type: The target field combination does not include a time dimension field. This can be divided into two cases for discussion: the target field combination contains only one non-time dimension field, the target field combination contains two or more non-time dimension fields, and the target field combination contains multiple non-time dimension fields and one measure field.

[0222] In the case where there is only one non-time dimension in the target field combination, multiple data items in the non-time dimension field are clustered to obtain at least two clusters, each cluster including at least one data item. The purpose of the above clustering is to cluster the dimension field data that perform similarly or similarly under each metric field in the same dimension field into one cluster. Subsequently, the clustering parameters for each metric field corresponding to each data item in the dimension field are determined. Based on the clustering parameters corresponding to each data item, multiple data items are clustered to obtain at least two clusters. The following processing is performed for each metric field: analogy analysis is performed on at least two clusters based on the metric field data in the metric field to obtain analogy analysis data for the metric field corresponding to each cluster; analogy analysis data corresponding to the metric field corresponding to each cluster is obtained. Finally, material information is generated based on the analogy analysis data corresponding to the dimensional field.

[0223] Analogy analysis can compare similar things under different conditions, and provide an intuitive and simple way to highlight the differences and gaps between similar things under different conditions, thereby assisting users in making comparisons and decisions.

[0224] When there are two or more non-time dimension fields in the target field combination, first, classify the dimension fields in the target field combination, for example, into categories such as personnel, region, date, and others. Then, use corresponding data analysis methods for dimension fields of different categories. For example, for dimension fields of the date category, you can use the above-mentioned related description method to generate a time trend analysis chart; for dimension fields of the personnel, region, and other categories, you can use the above-mentioned related description method to generate analogy analysis data. Then, all data generated during the entire data analysis process, including but not limited to: first-level titles, inferred description statements, statistical analysis data, charts (such as time trend analysis charts, analogy analysis data, data distribution charts), and tables, are used as material information.

[0225] When a target field combination contains multiple non-time dimension fields and one metric field, a cross-dimensional field group and corresponding cross-tab are obtained from the target field combination. A cross-dimensional field group contains a cross-relationship between the two dimension fields. A cross-relationship refers to one of the five relationships between concepts in logic. Simply put, in the relationship between concepts A and B, if some A is B and some A is not B, and some B is A and some B is not A, then the two concepts A and B have a cross-relationship. For example, the concepts of food and plant have a cross-relationship. For a detailed explanation of the cross-relationship, those skilled in the art can refer to the relevant descriptions in the prior art and will not be elaborated here. A cross-tab is a commonly used classification summary table. Using a cross-tab for data querying is very intuitive and clear, and is widely used. When constructing a cross-tab, one dimension field in the cross-dimensional field group is used as a row, and another dimension field as a column. The intersection of rows and columns allows for various summary calculations on the data, such as count, sum, average, maximum, minimum, and so on.

[0226] Afterwards, association analysis is performed on the cross-tab to obtain the material information of the target field combination. Among them, association analysis, also known as association mining, is a simple and practical analysis technique that aims to discover the associations (or correlations) existing in large data sets, thereby describing the regularities and patterns of the simultaneous occurrence of certain attributes in a thing.

[0227] In one embodiment, integrating the material information of the target field combination to form a data report of the target data table includes: determining the integration order between the material information of multiple target field combinations, and integrating the material information of the target field combination to form a data report of the target data table according to the integration order.

[0228] Specifically, an exemplary implementation of integrating the material information of multiple target field combinations to form a data report of a target data table includes: determining the graph category of the target field combination; based on the graph category of the target field combination, organizing the material information of the multiple target field combinations according to a preset graph category order and a set content organization structure to form a data report of the target data table.

[0229] Among them, the atlas category is a classification system for data and information. During use, data and information can be divided into different categories based on factors such as the nature, source, and purpose of the data. These categories include but are not limited to time, place, person, event or thing, others, and numerical values. Among them, time records the specific time point or time period when the event occurs, which is the basis for constructing the event timeline. Place indicates the specific location or area where the event occurs, providing a spatial background for the event. People identify the key individuals or groups involved in the event, which is a key element in understanding the motivation and outcome of the event. Events or things describe specific events that occurred or objects and phenomena that exist, and are the core content of the data report. Others cover other important information that does not belong to the above specific categories and is used to supplement and improve the data report. Numerical values record specific data or statistical results related to the event, providing quantitative evidence and data support.

[0230] The division of graph categories not only helps organize and manage data, but also ensures that data reports are structured and logical.

[0231] For example, the preset graph category order is time, place, person, event or thing, other, and value. The data report that forms the target data table by organizing the material information of multiple target fields according to the preset graph category order and the set content organization structure is set based on the general rules of user cognition. This order helps readers grasp the core content of the data report more quickly and gradually deepen their understanding in a logical order.

[0232] In one embodiment, based on the graph category of the target field combination, the material information of multiple target field combinations is sorted according to a preset graph category order and a set content organization structure to form a data report of the target data table. An exemplary implementation includes: based on the graph category of the target field combination, the material information of multiple target field combinations is input into the data report generation model according to a preset graph category order to obtain a data report of the target data table.

[0233] For example, the data report generation model and the target data table header field list are used to generate the document title and data summary section of the data report. The data report generation model and the target field combination of material information are used to generate the data briefing, detailed analysis, and data suggestion sections of the data report, and a visual chart is configured in the detailed analysis section.

[0234] In one embodiment, an exemplary implementation of determining the graph category of a target field combination includes: when the type of the first dimension field satisfies a type-category mapping relationship in a preset type-category mapping relationship set, determining the category corresponding to the first dimension field; wherein the type-category mapping relationship describes a type and category with a unique correspondence; the first dimension field is a dimension field in the target field combination or a dimension field with the highest recommendation score among multiple dimension fields; when the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, determining the category corresponding to the first dimension field based on the correlation between the first dimension field and the dimension field whose category has been determined in the target data table and the recommendation score of the first dimension field; and determining the category corresponding to the first dimension field as the graph category of the target field combination.

[0235] In this embodiment, the type category mapping relationship set includes one or more of the following type category mapping relationships: date type corresponds to time category, time type corresponds to time category, date time type corresponds to time category, place name type corresponds to place category, person name type corresponds to person category, organization group type corresponds to person category, mobile phone number type corresponds to person category, telephone number type corresponds to person category, text type numerical type corresponds to person category, ID card number type corresponds to person category, verb type corresponds to event category; accordingly, when the type of the first dimension field satisfies the type category mapping relationship in the type category mapping relationship set, determining the category corresponding to the first dimension field includes: determining the corresponding type category mapping relationship in the type category mapping relationship set according to the type of the first dimension field; and determining the category of the first dimension field according to the determined type category mapping relationship.

[0236] In this embodiment, when the first dimension field satisfies the type-category mapping relationship, the category of the first dimension field can be determined based on the corresponding type-category mapping relationship. For example, when the type of the first dimension field is a gerund type, the mapping relationship of the gerund type corresponding to the event category is satisfied, and the event category is determined as the category of the first dimension field. It should be noted that in this embodiment, the type-category mapping relationship in the type-category mapping relationship set is as described above. In other embodiments, other type-category mapping relationships may also be included, such as room number type corresponding to event type, etc., which are not specifically limited here.

[0237] In this embodiment, when the first dimension field is not successfully determined as one of the four categories mentioned above based on the type of the first dimension field, the category corresponding to the first dimension field can be determined based on the correlation between the first dimension field and the dimension field whose category has been determined. It should be noted that the correlation between the first dimension field and the dimension field whose category has been determined can be calculated using a method of obtaining a correlation score using a classification model with an inclusion relationship, and the category corresponding to the dimension field with a larger correlation score is determined as the category of the first dimension field. If the category of dimension field A is the time category and the category of dimension field B is the location category, the correlation score of the first dimension field with dimension field A is calculated to be 0.9, and the correlation score with dimension field B is 0.5. Since the correlation score with dimension field A is greater than the correlation score with dimension field B, the category of the first dimension field is determined to be the category corresponding to dimension field A, that is, the time category.

[0238] Figure 7 This is a block diagram of an embodiment of a data report generating device provided in an embodiment of the present application. Figure 7 As shown, the device includes:

[0239] The field combination module 71 is used to select a target field combination from the fields of the target data table;

[0240] A material generation module 72 is configured to generate material information of the target field combination based on the data items of each field in the target field combination, wherein the material information includes a visual chart and a text description;

[0241] The material integration module 73 is used to integrate the material information of the target field combination to form a data report of the target data table.

[0242] In a possible implementation, the field combination module 71 includes:

[0243] A field classification unit is used to classify the fields related to the same analysis target in the target data table into the same field class to obtain multiple field classes;

[0244] The field combination unit is used to construct a target field combination using the fields in the field class.

[0245] In a possible implementation, the field combination unit includes:

[0246] A combining unit, configured to combine the fields in the field class according to a set field combination type to obtain a plurality of candidate field combinations;

[0247] an evaluation unit, configured to determine a recommendation score for the candidate field combination;

[0248] The screening unit is configured to select a target field combination from a plurality of the candidate field combinations according to the recommendation scores of the candidate field combinations.

[0249] In a possible implementation, the evaluation unit is specifically configured to:

[0250] Determining a recommendation score for each field in the candidate field combination, and determining a first recommendation score for the candidate field combination based on the recommendation score of each field;

[0251] determining a second recommendation score of the candidate field combination according to the dimension fields in the candidate field combination;

[0252] A final recommendation score of the candidate field combination is determined according to the first recommendation score and the second recommendation score.

[0253] In a possible implementation, the field combination module 71 includes:

[0254] An associated table acquisition unit, configured to acquire an associated data table of the target data table;

[0255] A historical recommendation combination acquisition unit, configured to acquire a historical recommendation field combination of the associated data table;

[0256] The combination unit is used to select a target field combination from the fields of the target data table according to the historically recommended field combination.

[0257] In a possible implementation, the combination unit includes:

[0258] a combining subunit, configured to select, from the fields of the target data table, a field combination that matches the historically recommended field combination as a candidate field combination;

[0259] An evaluation subunit, configured to determine a recommendation score for the candidate field combination;

[0260] The screening subunit is configured to select a target field combination from a plurality of the candidate field combinations according to the recommendation scores of the candidate field combinations.

[0261] In one possible implementation, the evaluation subunit is specifically configured to:

[0262] Determining a recommendation score for each field in the candidate field combination, and determining a first recommendation score for the candidate field combination based on the recommendation score of each field;

[0263] Determining a historical recommendation count of the candidate field combination, and determining a second recommendation score of the candidate field combination based on the historical recommendation count;

[0264] A final recommendation score of the candidate field combination is determined according to the first recommendation score and the second recommendation score.

[0265] In a possible implementation, the material generation module 72 includes:

[0266] a combination type determining unit, configured to determine a combination type of the target field combination according to the number of dimension fields and metric fields in the target field combination;

[0267] A generating unit is configured to determine a target analysis rule according to the combination type, analyze the data items corresponding to the target field combination according to the target analysis rule, and generate material information of the target field combination according to the analysis result.

[0268] In one possible implementation, the material integration module 73 includes:

[0269] A graph category determination unit, configured to determine the graph category of the target field combination;

[0270] The integration unit is used to organize the material information of the target field combination according to the preset atlas category order and the set content organization structure based on the atlas category of the target field combination to form a data report of the target data table.

[0271] In a possible implementation, the atlas category determination unit is specifically configured to:

[0272] If the type of the first dimension field satisfies a type-category mapping relationship in a preset type-category mapping relationship set, determining a category corresponding to the first dimension field; wherein the type-category mapping relationship describes a type and a category with a unique corresponding relationship; and the first dimension field is a dimension field in the target field combination;

[0273] If the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, determining the category corresponding to the first dimension field based on the correlation between the first dimension field and the dimension field of the determined category in the target data table and the recommendation score of the first dimension field;

[0274] The category corresponding to the first dimension field is determined as the graph category of the target field combination.

[0275] In a possible implementation manner, the integration unit is specifically used to:

[0276] Based on the atlas category of the target field combination, the material information of the target field combination is input into the data report generation model according to a preset atlas category order to obtain a data report of the target data table.

[0277] like Figure 8 As shown, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.

[0278] Memory 113, for storing computer programs;

[0279] In one embodiment of the present application, the processor 111 is configured to execute a program stored in the memory 113 to implement the data report generation method provided in any of the aforementioned method embodiments, including:

[0280] Select the target field combination from the fields of the target data table;

[0281] Generating material information of the target field combination based on the data items of each field in the target field combination, wherein the material information includes a visual chart and a text description;

[0282] The material information of the target field combination is integrated to form a data report of the target data table.

[0283] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method for generating a data report provided in any of the aforementioned method embodiments are implemented.

[0284] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0285] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiment.

[0286] It should be understood that the terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "comprise", "include", "contain" and "have" are inclusive and therefore specify the presence of stated features, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the specific order described or illustrated, unless the order of execution is clearly indicated. It should also be understood that additional or alternative steps may be used.

[0287] The foregoing is merely a list of specific embodiments of the present application, intended to enable those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the broadest scope consistent with the principles and novel features of the present application.

Claims

1. A data report generation method, characterized in that: The method comprises: Select the target field combination from the fields of the target data table; Generating material information of the target field combination based on the data items of each field in the target field combination, wherein the material information includes a visual chart and a text description; The material information of the target field combination is integrated to form a data report of the target data table.

2. The method according to claim 1, characterized in that The step of selecting a target field combination from the fields of the target data table includes: Classifying the fields related to the same analysis target in the target data table into the same field class to obtain multiple field classes; Build a target field combination using the fields in the field class.

3. The method according to claim 2, characterized in that The step of constructing a target field combination by using the fields in the field class includes: Combine the fields in the field class according to the set field combination type to obtain multiple candidate field combinations; Determining a recommendation score for the candidate field combination; A target field combination is selected from a plurality of the candidate field combinations according to the recommendation scores of the candidate field combinations.

4. The method according to claim 3, characterized in that Determining the recommendation score of the candidate field combination includes: Determining a recommendation score for each field in the candidate field combination, and determining a first recommendation score for the candidate field combination based on the recommendation score of each field; determining a second recommendation score of the candidate field combination according to the dimension fields in the candidate field combination; A final recommendation score of the candidate field combination is determined according to the first recommendation score and the second recommendation score.

5. The method according to claim 1, wherein The target field combination selected from the fields of the target data table includes: Obtaining the associated data table of the target data table; Obtaining a historically recommended field combination of the associated data table; According to the historically recommended field combinations, a target field combination is selected from the fields of the target data table.

6. The method according to claim 5, characterized in that The step of selecting a target field combination from the fields of the target data table according to the historically recommended field combination includes: Selecting, from the fields of the target data table, a field combination that matches the historically recommended field combination as a candidate field combination; Determining a recommendation score for the candidate field combination; A target field combination is selected from a plurality of the candidate field combinations according to the recommendation scores of the candidate field combinations.

7. The method according to claim 6, characterized in that Determining the recommendation score of the candidate field combination includes: Determining a recommendation score for each field in the candidate field combination, and determining a first recommendation score for the candidate field combination based on the recommendation score of each field; Determining a historical recommendation count of the candidate field combination, and determining a second recommendation score of the candidate field combination based on the historical recommendation count; A final recommendation score of the candidate field combination is determined according to the first recommendation score and the second recommendation score.

8. The method according to claim 1, characterized in that The generating of the material information of the target field combination based on the data items of each field in the target field combination includes: Determining a combination type of the target field combination according to the number of dimension fields and metric fields in the target field combination; A target analysis rule is determined according to the combination type, and data items corresponding to the target field combination are analyzed according to the target analysis rule, and material information of the target field combination is obtained according to the analysis result.

9. The method according to claim 1, characterized in that The step of integrating the material information of the target field combination to form a data report of the target data table includes: Determining a graph category of the target field combination; Based on the atlas category of the target field combination, the material information of the target field combination is sorted according to a preset atlas category order and a set content organization structure to form a data report of the target data table.

10. The method according to claim 9, characterized in that Determining the graph category of the target field combination includes: If the type of the first dimension field satisfies a type-category mapping relationship in a preset type-category mapping relationship set, determining a category corresponding to the first dimension field; wherein the type-category mapping relationship describes a type and a category with a unique corresponding relationship; and the first dimension field is a dimension field in the target field combination; If the type of the first dimension field does not satisfy the type-category mapping relationship in the type-category mapping relationship set, determining the category corresponding to the first dimension field based on the correlation between the first dimension field and the dimension field of the determined category in the target data table and the recommendation score of the first dimension field; The category corresponding to the first dimension field is determined as the graph category of the target field combination.

11. The method according to claim 9, characterized in that The graph category based on the target field combination organizes the material information of the target field combination according to a preset graph category order and a set content organization structure to form a data report of the target data table, including: Based on the atlas category of the target field combination, the material information of the target field combination is input into the data report generation model according to a preset atlas category order to obtain a data report of the target data table.

12. A data report generating device, characterized in that: The device comprises: The field combination module is used to select the target field combination from the fields of the target data table; A material generation module, configured to generate material information of the target field combination based on the data items of each field in the target field combination, wherein the material information includes a visual chart and a text description; The material integration module is used to integrate the material information of the target field combination to form a data report of the target data table.

13. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the data report generating method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Data report generation method and device, electronic equipment and storage medium

    CN115345138A

  • Data report generation method and device, electronic equipment and storage medium

    CN115345139A

  • Visual generation method and system for electric power cost analysis report

    CN119718519A