Data processing method, device and equipment and computer readable storage medium
By enumeration of data topics, matching fact types and visualization of business data sets, the problems of incomplete data extraction and poor chart readability in the prior art are solved, and more efficient and easy-to-read data visualization is achieved.
Patent Information
- Application Number
- CN202510125016.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-23
AI Technical Summary
Existing business data visualization technology relies on manual creation of charts, resulting in incomplete and inaccurate data extraction and poor chart readability.
By obtaining a business data set including multiple dimensions and metrics, data topic enumeration is performed, the full number of fact types are matched, data facts are extracted, and data facts are visualized according to the adapted chart type.
Improves the completeness and accuracy of data fact extraction and enhances the readability of charts after data visualization.
Smart Images

Figure CN120030082A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet technology, and in particular to a data processing method, apparatus, device, and computer-readable storage medium. Background Art
[0002] Business data visualization can significantly improve data comprehensibility and analysis efficiency by converting complex data into intuitive charts.
[0003] Existing business data visualization technologies usually rely on manual creation of charts, but manual creation of charts may result in incomplete and inaccurate data extraction or data analysis, so the readability of charts after data visualization is poor. Summary of the invention
[0004] The embodiments of the present application provide a data processing method, apparatus, device, and computer-readable storage medium, which can not only improve the completeness and accuracy of data fact extraction, but also enhance the readability of charts after data fact visualization.
[0005] On the one hand, an embodiment of the present application provides a data processing method, including:
[0006] Obtain a business data set including data corresponding to d dimensions and data corresponding to m metrics; d and m are both positive integers;
[0007] Perform data subject enumeration processing on the business data set to obtain a data subject set; the dimensions included in each data subject in the data subject set belong to d dimensions; the metrics included in each data subject in the data subject set belong to m metrics;
[0008] Each data subject in the data subject set is matched with the full amount of fact types to obtain a first data subject and a first fact type that are successfully matched; the first data subject belongs to the data subject set; the first fact type belongs to the full amount of fact types;
[0009] Extracting facts from the first data subject according to the first fact type to obtain data facts corresponding to the business data set;
[0010] The data fact is visualized according to a chart type adapted to the first fact type to obtain a visualization chart for displaying the data fact.
[0011] On the one hand, an embodiment of the present application provides a data processing device, the device comprising:
[0012] An acquisition module, used to acquire a business data set including data corresponding to d dimensions and data corresponding to m metrics; d and m are both positive integers;
[0013] A processing module is used to perform data subject enumeration processing on the business data set to obtain a data subject set; the dimensions included in each data subject in the data subject set belong to d dimensions; the metrics included in each data subject in the data subject set belong to m metrics;
[0014] The processing module is further used to match each data subject in the data subject set with the full amount of fact types to obtain a first data subject and a first fact type that are successfully matched; the first data subject belongs to the data subject set; and the first fact type belongs to the full amount of fact types;
[0015] An extraction module, configured to extract facts from a first data subject according to a first fact type, to obtain data facts corresponding to a business data set;
[0016] The processing module is further used to perform visualization processing on the data fact according to a chart type adapted to the first fact type, so as to obtain a visualization chart for displaying the data fact.
[0017] In a possible implementation, the processing module performs data subject enumeration processing on the business data set to obtain a data subject set for performing the following operations:
[0018] Perform subspace enumeration processing on the business data set to obtain a subspace set; the subspace set includes the first subspace; the first subspace includes data on b dimensions in the business data set; b is a positive integer, and b is less than or equal to d; the b dimensions include dimension P q , q is a positive integer, and q is less than or equal to b;
[0019] If the first subspace includes dimension P q The full amount of data on the q Determine the partition dimension of the first subspace;
[0020] Performing data topic enumeration processing on the first subspace, the segmentation dimension, and the m metrics to obtain data topics corresponding to the first subspace; the data topics corresponding to the first subspace include the first subspace and the segmentation dimension;
[0021] The data topics corresponding to the first subspace and the data topics corresponding to the second subspace are combined into a data topic set; the second subspace refers to a subspace in the subspace set except the first subspace.
[0022] In a possible implementation, the processing module matches each data subject in the data subject set with the full set of fact types to obtain a first data subject and a first fact type that are successfully matched, and performs the following operations:
[0023] Perform influence analysis on each data topic in the data topic set to obtain the influence value corresponding to each data topic in the data topic set;
[0024] Compare the influence value corresponding to each data topic in the data topic set with the influence threshold;
[0025] The data topics in the data topic set whose influence values are greater than or equal to the influence threshold are matched with all fact types to obtain the first data topic and the first fact type that are successfully matched.
[0026] In a possible implementation, the data subject in the data subject set includes a second data subject; the second data subject is any data subject in the data subject set; the second data subject includes a first metric; the first metric belongs to m metrics;
[0027] The processing module performs influence analysis on each data topic in the data topic set to obtain an influence value corresponding to each data topic in the data topic set, which is used to perform the following operations:
[0028] According to the segmentation dimension in the second data subject, the subspace in the second data subject is segmented to obtain multiple brother subspaces corresponding to the second data subject; the data of the multiple brother subspaces on the segmentation dimension are different from each other, and the data of the multiple brother subspaces on the dimensions other than the segmentation dimension are the same;
[0029] In the business data set, full data of the first metric is obtained, and influence analysis is performed on multiple brother subspaces according to the full data of the first metric to obtain influence values corresponding to the multiple brother subspaces respectively;
[0030] The influence values corresponding to the multiple brother subspaces are averaged to obtain the influence value corresponding to the second data topic.
[0031] In a possible implementation, the multiple brother subspaces include brother subspace A c , c is a positive integer, and c is less than or equal to the total number of multiple sibling subspaces;
[0032] The processing module performs influence analysis on the multiple brother subspaces according to the full amount of data of the first metric, and obtains influence values corresponding to the multiple brother subspaces, which are used to perform the following operations:
[0033] Sum all the data of the first metric to obtain total data corresponding to the first metric;
[0034] In the full data of the first metric, obtain the sibling subspace A c Data on the first measure;
[0035] Brother subspace A c The data on the first metric and the total data are processed by ratio to obtain the sibling subspace A c The corresponding influence value.
[0036] In a possible implementation, the number of the first metrics is at least two; the at least two first metrics include the second metric, and the second metric is any one of the at least two first metrics; the total data includes the total data corresponding to the second metric;
[0037] Processing module for sibling subspace A c The data on the first metric and the total data are processed by ratio to obtain the sibling subspace A c The corresponding influence values are used to perform the following operations:
[0038] Brother subspace A c The data on the second metric and the total data corresponding to the second metric are processed by ratio to obtain the sibling subspace A c The influence value on the second metric;
[0039] Get the sibling subspace A c The influence values corresponding to at least two first metrics are weighted averaged to obtain the sibling subspace A. c The corresponding influence value.
[0040] In a possible implementation, the extraction module extracts facts from the first data subject according to the first fact type to obtain data facts corresponding to the business data set, and performs the following operations:
[0041] Performing importance analysis on the first data topic to obtain an importance score corresponding to the first data topic;
[0042] If the importance score corresponding to the first data subject is equal to or greater than the importance score threshold, facts are extracted from the first data subject according to the pattern verification strategy corresponding to the first fact type to obtain data facts corresponding to the business data set.
[0043] In a possible implementation, the number of the first data topics is at least two; the at least two first data topics include data topic G h and data subject G i , h and i are both positive integers, and h and i are both less than or equal to the total number of at least two first data topics, and h is different from i;
[0044] The extraction module performs importance analysis on the first data topic to obtain an importance score corresponding to the first data topic, which is used to perform the following operations:
[0045] If the data subject G h The influence value is greater than or equal to the data subject G i The influence value of the first importance assessment task is determined to have a higher execution priority than the second importance assessment task; the first importance assessment task is used to complete the data subject G h The second importance assessment task is used to complete the data theme G i Analysis of the importance of
[0046] Add the first importance evaluation task and the second importance evaluation task to the task queue in descending order of execution priority;
[0047] Determine the idle number corresponding to the idle threads in the parallel execution thread pool; if the idle number is equal to or greater than the parallel execution thread threshold, obtain the third importance evaluation task from the task queue in descending order of execution priority through the idle thread; the number of the third importance evaluation task is equal to or less than the idle number; the third importance evaluation task is used to complete the data topic G in at least two first data topics j The importance analysis of ; j is a positive integer, and j is less than or equal to the total number of at least two first data topics;
[0048] Through the idle thread, the third importance evaluation task is executed in parallel to obtain the data topic G j The corresponding importance score; wherein the importance score corresponding to a data topic is obtained by performing a third importance assessment task in an idle trip;
[0049] The first fact type includes the data subject G j The second fact type that matches successfully;
[0050] If the importance score corresponding to the first data topic is equal to or greater than the importance score threshold, the extraction module extracts facts from the first data topic according to the pattern verification strategy corresponding to the first fact type, and obtains data facts corresponding to the business data set, which are used to perform the following operations:
[0051] If the data subject G j If the corresponding importance score is equal to or greater than the importance score threshold, then according to the pattern verification strategy corresponding to the second fact type, the data subject G j Perform fact extraction to obtain data facts corresponding to the business data set.
[0052] In a possible implementation, the extraction module performs importance analysis on the first data topic to obtain an importance score corresponding to the first data topic, which is used to perform the following operations:
[0053] According to the segmentation dimension in the first data subject, the subspace in the first data subject is segmented to obtain multiple brother subspaces corresponding to the first data subject; the data of the multiple brother subspaces on the segmentation dimension are different from each other, and the data of the multiple brother subspaces on the dimensions other than the segmentation dimension are the same;
[0054] In the business data set, obtain the full amount of data of the metric z in the first data topic; the metric z belongs to m metrics;
[0055] According to the full data of the metric z and multiple sibling subspaces, an importance analysis is performed on the first data topic to obtain an importance score corresponding to the first data topic.
[0056] In a possible implementation, the extraction module performs importance analysis on the first data topic according to the full data of the metric z and the multiple sibling subspaces to obtain an importance score corresponding to the first data topic, which is used to perform the following operations:
[0057] In the full amount of data of metric z, data corresponding to multiple sibling subspaces on metric z are obtained, and the obtained multiple data are determined as a result set;
[0058] Perform a significance test on the result set to obtain a significance value corresponding to the first data topic;
[0059] The significance value corresponding to the first data topic and the influence value corresponding to the first data topic are multiplied to obtain the importance score corresponding to the first data topic.
[0060] In a possible implementation, the extraction module performs a significance test on the result set to obtain a significance value corresponding to the first data topic, which is used to perform the following operations:
[0061] Perform a point significance test on the result set to obtain the point significance value, and perform a slope significance test on the result set to obtain the slope significance value;
[0062] If the point significance value is greater than or equal to the slope significance value, the point significance value is determined as the significance value corresponding to the first data theme;
[0063] If the point significance value is less than the slope significance value, the slope significance value is determined as the significance value corresponding to the first data topic.
[0064] In a possible implementation, after the processing module obtains the visualization chart, the processing module is further configured to perform the following operations:
[0065] If the total number of visualization charts is at least two, the at least two visualization charts are sorted to obtain a chart optimization sequence of the visualization charts;
[0066] Perform animation mapping processing on each visualization chart in the chart optimization sequence to obtain a chart animation;
[0067] Perform video creation processing on the chart animation to obtain a visual video for presenting data facts.
[0068] In a possible implementation, the processing module sorts at least two visualization charts to obtain a chart optimization sequence of the visualization charts, which is used to perform the following operations:
[0069] According to the total number of at least two visualization charts, determine the sorting number X; X is a positive integer greater than 1;
[0070] Sort at least two visual charts X times to obtain X chart sequences; each of the X chart sequences includes at least two visual charts, and the chart sequences corresponding to the X chart sequences are different; the X chart sequences include chart sequence K n , n is a positive integer, and n is less than or equal to X;
[0071] For the chart sequence K n Perform cost analysis and obtain chart sequence K n The corresponding sequence cost;
[0072] Determine the sequence costs corresponding to the X chart sequences respectively, and obtain the minimum sequence cost from the X sequence costs;
[0073] The chart sequence corresponding to the minimum sequence cost among the X chart sequences is determined as the chart optimization sequence.
[0074] In a possible implementation, the processing module processes the chart sequence K n Perform cost analysis and obtain chart sequence K n The corresponding sequence costs are used to perform the following operations:
[0075] For the chart sequence K n Perform chart conversion cost analysis to obtain chart sequence K n The corresponding chart conversion cost;
[0076] For the chart sequence K n Perform graph filtering cost analysis to obtain graph sequence K n The corresponding chart filtering cost;
[0077] The chart conversion cost and the chart filtering cost are summed to obtain the chart sequence K n The corresponding sequence cost.
[0078] In a possible implementation, the processing module processes the chart sequence K n Perform chart conversion cost analysis to obtain chart sequence Kn The corresponding chart conversion costs for performing the following operations:
[0079] Get chart sequence K n Every two adjacent visualization charts in n Each two adjacent visualization charts in the method include an adjacent first visualization chart and a second visualization chart;
[0080] If there is a conversion relationship between the first visualization chart and the second visualization chart, the conversion cost corresponding to the conversion relationship is determined as the chart conversion cost of the first visualization chart and the second visualization chart;
[0081] If there is no conversion relationship between the first visualization chart and the second visualization chart, determining a chart path for connecting the first visualization chart and the second visualization chart, wherein every two adjacent visualization charts in the chart path have a conversion relationship;
[0082] Determining chart conversion costs of the first visualization chart and the second visualization chart according to conversion costs corresponding to conversion relationships in the chart path;
[0083] For the chart sequence K n The chart conversion costs corresponding to each two adjacent visualization charts in the summation process are processed to obtain the chart sequence K n The corresponding chart conversion cost.
[0084] In a possible implementation, the processing module processes the chart sequence K n Perform graph filtering cost analysis to obtain graph sequence K n The corresponding chart filtering costs are used to perform the following operations:
[0085] Get chart sequence K n Every two adjacent visualization charts in n Each two adjacent visualization charts in the method include an adjacent first visualization chart and a second visualization chart;
[0086] If the dimension displayed by the first visualization chart is the same as the dimension displayed by the second visualization chart, determining the first value as a chart distance between the first visualization chart and the second visualization chart;
[0087] If the dimension displayed by the first visualization chart is different from the dimension displayed by the second visualization chart, the second value is determined as the chart distance between the first visualization chart and the second visualization chart; the first value is greater than the second value;
[0088] Determine the chart sequence K n The chart distance corresponding to each two adjacent visualization charts in ;
[0089] According to the chart sequence K n The chart distances corresponding to each two adjacent visualization charts in the image are used to determine the chart sequence K. n The corresponding chart filtering cost.
[0090] On one hand, the present application provides a computer device, including: a processor, a memory, and a network interface;
[0091] The above-mentioned processor is connected to the above-mentioned memory and the above-mentioned network interface, wherein the above-mentioned network interface is used to provide a data communication function, the above-mentioned memory is used to store a computer program, and the above-mentioned processor is used to call the above-mentioned computer program so that the computer device executes the method in the embodiment of the present application.
[0092] On one hand, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program is suitable for being loaded by a processor and executing the method in the embodiment of the present application.
[0093] On the one hand, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium; a processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method in the embodiment of the present application.
[0094] The embodiment of the present application can extract all data subjects of the business data set by enumerating data subjects of the business data set, thereby improving the analysis completeness of the business data set; by matching the data subjects in the data subject set and the full fact types, the first fact type matching the first data subject can be accurately and completely determined; by extracting facts from the first data subject through the first fact type, the accuracy of the data facts can be improved; by visualizing the data facts through a chart type adapted to the first fact type, the readability of the data visualization can be enhanced; from the above, it can be seen that the embodiment of the present application can not only improve the completeness and accuracy of data fact extraction, but also enhance the readability of the chart after data fact visualization. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0096] Figure 1It is a schematic diagram of a system architecture provided by an embodiment of the present application;
[0097] Figure 2 This is a flow diagram of a data processing method provided in an embodiment of the present application. Figure 1 ;
[0098] Figure 3 This is a data processing scenario provided by an embodiment of the present application. Figure 1 ;
[0099] Figure 4 This is a data processing scenario provided by an embodiment of the present application. Figure 2 ;
[0100] Figure 5a This is a data distribution diagram provided in an embodiment of the present application;
[0101] Figure 5b is a schematic diagram of a Gaussian distribution provided in an embodiment of the present application;
[0102] Figure 6a is a schematic diagram of a straight line fitted by linear regression provided in an embodiment of the present application;
[0103] Figure 6b It is a schematic diagram of a Logistic distribution provided in an embodiment of the present application;
[0104] Figure 7 This is a flow diagram of a data processing method provided in an embodiment of the present application. Figure 2 ;
[0105] Figure 8 This is a flow diagram of a data processing method provided in an embodiment of the present application. Figure 3 ;
[0106] Fig. 9 This is a schematic diagram of an automated working pipeline provided in an embodiment of the present application;
[0107] Fig.10 is a schematic diagram of a data fact extraction algorithm framework provided in an embodiment of the present application;
[0108] Fig.11 This is a data processing scenario provided by an embodiment of the present application. Figure 3 ;
[0109] Fig.12 is a schematic diagram of a video synthesis scene provided by an embodiment of the present application;
[0110] Fig.13 It is a schematic diagram of a framework of a method for visualizing video presentation based on business data provided in an embodiment of the present application;
[0111] Fig.14 is a structural schematic diagram of a data processing device provided in an embodiment of the present application;
[0112] Fig.15 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0113] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0114] See also Figure 1 , Figure 1 Schematic diagram of a system architecture provided by an embodiment of the present application. Figure 1 As shown, the system may include a business server 100 and a terminal device cluster. The terminal device cluster may include: terminal device 200a, terminal device 200b, terminal device 200c, ..., terminal device 200n. It can be understood that the above system may include one or more terminal devices, and this application does not limit the number of terminal devices.
[0115] There may be communication connections between the terminal device clusters, for example, there is a communication connection between the terminal device 200a and the terminal device 200b, and there is a communication connection between the terminal device 200a and the terminal device 200c. At the same time, any terminal device in the terminal device cluster may have a communication connection with the business server 100, for example, there is a communication connection between the terminal device 200a and the business server 100, wherein the above communication connection does not limit the connection method, and may be directly or indirectly connected by wired communication, directly or indirectly connected by wireless communication, or by other methods, and this application does not limit it here.
[0116] It should be understood that Figure 1 Each terminal device in the terminal device cluster shown in FIG. 1 may be installed with an application client. When the application client runs in each terminal device, it may be respectively connected to the above-mentioned Figure 1 The business server 100 shown performs data interaction, that is, the above-mentioned communication connection. Among them, the application client can be a video application, social application, digital resource application, payment application, financial application, game application, shopping application, novel application, browser and other application clients, and this application does not limit this.
[0117] The application client may be an independent client or an embedded sub-client integrated in a client (e.g., a video client), which is not limited here. Taking the financial management application as an example, the business server 100 may be a collection of multiple servers including a background server corresponding to the financial management application, a data processing server, etc. Therefore, each terminal device may transmit data with the business server 100 through the application client corresponding to the financial management application. For example, each terminal device may upload financial management data to the business server 100 through the application client of the financial management application, and then the business server 100 may perform fact processing on the financial management data to obtain the data facts corresponding to the financial management data, and then may perform visualization processing on the data facts to obtain a visualization chart for displaying the data facts.
[0118] It is understandable that in the specific implementation of this application, related data such as user information (such as business data) is involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant local laws, regulations and standards.
[0119] To facilitate subsequent understanding and description, the embodiments of the present application can be Figure 1 An example of a terminal device is selected from the terminal device cluster shown, for example, terminal device 200a is used for description. When a data visualization instruction for a business data set is received in the application client, terminal device 200a can generate a data visualization request corresponding to the data visualization instruction in the application client. The business data set includes data corresponding to d dimensions and data corresponding to m metrics, that is, each row of business data in the business data set includes sub-data corresponding to d dimensions and sub-data corresponding to m metrics, and d and m are both positive integers.
[0120] The embodiments of the present application do not limit the content of the business data set and can be determined according to the actual application scenario. For example, if the application client is a financial application, the business data set can be a financial data set; if the application client is a school system, the business data set can be a student performance data set; if the application client is an advertising application, the business data set can be an advertising conversion data set.
[0121] Through the application client, the terminal device 200a can send a data visualization request to the business server 100. The business server 100 obtains the business data set according to the data visualization request. The embodiment of the present application does not limit the way in which the business server 100 obtains the business data set, which can be determined according to the actual application scenario. One feasible way to obtain is that the business server 100 stores the business data set; another feasible way to obtain is that the data visualization request includes the business data set, so the business server 100 can obtain the business data set from the data visualization request; another feasible way to obtain is that the data visualization request includes the storage path of the business data set, so the business server 100 can obtain the storage path from the data visualization request, access the storage path, and obtain the business data set.
[0122] The business server 100 performs data subject enumeration processing on the business data set, that is, generates all data subjects of the business data set to obtain a data subject set; the dimensions included in each data subject in the data subject set belong to d dimensions; the metrics included in each data subject in the data subject set belong to m metrics. A data subject includes one or more dimensions and one or more metrics.
[0123] The business server 100 matches each data subject in the data subject set with the full fact type to obtain the first data subject and the first fact type that are successfully matched; the first data subject belongs to the data subject set; the first fact type belongs to the full fact type. Among them, the full fact type refers to all fact types, and the fact type (facttype) is the data fact type, which can also be called the perspective (sight). Matching a data subject with a fact type is to verify whether the data subject meets the requirements of observing data distribution from this perspective.
[0124] According to the first fact type, the business server 100 extracts facts from the first data subject to obtain data facts corresponding to the business data set. Data facts refer to verified, objective, specific information or statistical results extracted from the business data set, which are the basis for building more complex analysis and decision-making.
[0125] Further, according to the chart type adapted to the first fact type, the business server 100 visualizes the data fact to obtain a visualized chart for displaying the data fact. The chart type refers to the type of chart, such as a bar chart, a line chart, a pie chart, a scatter chart, etc.
[0126] The business server 100 may return the visualization chart to the terminal device 200a. After receiving the visualization chart, the terminal device 200a may display the visualization chart for presenting data facts on its corresponding screen.
[0127] Optionally, the business server 100 returns the data facts corresponding to the business data set and the first fact type to the terminal device 200a. At this time, the terminal device 200a can visualize the data facts according to the chart type adapted to the first fact type to obtain a visualized chart.
[0128] Optionally, the business server 100 returns the successfully matched first data subject and the first fact type to the terminal device 200a. At this time, the terminal device 200a can extract facts from the first data subject according to the first fact type to obtain data facts corresponding to the business data set; and visualize the data facts according to a chart type adapted to the first fact type to obtain a visualized chart.
[0129] Optionally, the business server 100 can return the data subject set and the full set of fact types to the terminal device 200a. The data processing process after the terminal device 200a obtains the data subject set and the full set of fact types is the same as the data processing process of the business server 100, so please refer to the above description and will not be repeated here.
[0130] Optionally, if the terminal device 200a stores a business data set or can call a business data set, and the terminal device 200a has computing capabilities, then when the above data visualization instruction is received, the above data processing process can be executed locally. Therefore, please refer to the above description for the specific implementation process, which will not be repeated here.
[0131] In summary, the data processing method provided in this application can be Figure 1 Executed by any terminal device in the terminal device cluster in Figure 1 The business server in the Figure 1 The terminal devices and business servers in the terminal device cluster cooperate to execute, and the devices used to execute the data processing method in this application can be collectively referred to as computer devices.
[0132] Among them, the business server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud databases, cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.
[0133] Figure 1 The terminal devices shown include but are not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. Among them, the terminal devices and the service servers can be directly or indirectly connected in a wired or wireless manner, and the embodiments of the present application are not limited here.
[0134] For further information, see Figure 2 , Figure 2 This is a flow diagram of a data processing method provided in an embodiment of the present application. Figure 1 The implementation process of the data processing method can be carried out in the service server, in the terminal device, or interactively in the terminal device and the service server, which is not limited here. The terminal device can be the above-mentioned Figure 1 For any terminal device in the terminal device cluster of the corresponding embodiment, the service server may be the above-mentioned Figure 1 The corresponding business server 100 of the embodiment.
[0135] For the sake of ease of description and understanding, the present application embodiment takes the method executed by a business server as an example for explanation, that is, the business server executes the data processing method as a computer device. Figure 2 As shown, the data processing method may at least include the following steps S101 to S105.
[0136] Step S101, obtaining a business data set including data corresponding to d dimensions and data corresponding to m metrics; d and m are both positive integers.
[0137] Specifically, in data analysis and data visualization, dimension and measure are two important concepts. Dimension is a categorical attribute of data, which is used to group, classify or describe data. Dimension is usually a discrete, non-numeric field, such as city, product category, but can also be a numeric field, such as time dimensions such as year and month. Measures are numerical attributes of data, which are used to quantify and measure data. Measures are usually continuous, numerically operable fields, which are used to calculate, summarize or analyze data, such as total sales, total profit, conversion rate, etc.
[0138] The data corresponding to the dimension described in the embodiment of the present application refers to the dimension value of the dimension. For example, if the dimension is product, the dimension value is Game 1 and Game 2. The data corresponding to the metric refers to the metric value corresponding to the metric. For example, if the metric is recharge amount, the metric value is 300 (i.e., the recharge amount is 300).
[0139] For ease of understanding and description, please refer to Figure 3 , Figure 3 This is a data processing scenario provided by an embodiment of the present application. Figure 1 . Figure 3 Let d be 3, let m be 3, the three dimensions are transaction date, resource code for indicating digital resources and trading market for resource transactions, and the three metrics are first order price (also called opening price), highest price and lowest price.
[0140] like Figure 3 As shown, the business data set 20a includes 6 groups of records (i.e., 6 rows of business data). The first group of records is the first date, the first unit price of resource X...X in market 1 is 150, the highest price is 155, and the lowest price is 145; the second group of records is the first date, the first unit price of resource Y...Y in market 1 is 2800, the highest price is 2850, and the lowest price is 2770; the third group of records is the first date, the first unit price of resource Z...Z in market 1 is 300, the highest price is 2850, and the lowest price is 2770. The first price of resource X...X in market 1 is 305, the lowest price is 295; the fourth group of records is the second date, the first price of resource X...X in market 1 is 152, the highest price is 157, and the lowest price is 149; the fifth group of records is the second date, the first price of resource Y...Y in market 1 is 2810, the highest price is 2860, and the lowest price is 2790; the sixth group of records is the second date, the first price of resource Z...Z in market 1 is 303, the highest price is 308, and the lowest price is 300. The first date and the second date are different.
[0141] dom(D i ) can represent the domain of the i-th column dimension, which refers to the value range of the i-th column dimension, where the value range of i is [1, d]. Figure 3 The example is described in the following way. The value range of the first column dimension (i.e. dom(D 1 )) is [first date, second date], the value range of the second column dimension (i.e. dom(D 2 )) is [X...X, Y...Y, Z...Z], and the value range of the third column dimension (i.e. dom(D 3 )) is [Market 1].
[0142] Step S102, enumerate the data topics of the business data set to obtain a data topic set; the dimensions included in each data topic in the data topic set belong to d dimensions; the metrics included in each data topic in the data topic set belong to m metrics.
[0143] Specifically, the business data set is enumerated by subspace enumeration to obtain a subspace set; the subspace set includes a first subspace; the first subspace includes data on b dimensions in the business data set; b is a positive integer, and b is less than or equal to d; the b dimensions include dimension P q , q is a positive integer, and q is less than or equal to b; if the first subspace includes dimension P q The full amount of data on the qDetermine the partition dimension of the first subspace; perform data topic enumeration processing on the first subspace, the partition dimension and m metrics to obtain data topics corresponding to the first subspace; the data topics corresponding to the first subspace include the first subspace and the partition dimension; combine the data topics corresponding to the first subspace and the data topics corresponding to the second subspace into a data topic set; the second subspace refers to the subspace in the subspace set except the first subspace.
[0144] In a business data set, a subspace refers to a space composed of some or all dimensions in the business data set, and it can also be understood that any dimension of a non-empty subset can constitute a subspace of the business data set. The subspace described in the embodiment of the present application represents a filter on a dimension, which can be expressed by the following formula (1):
[0145] Sub={S[1],...,S[b]}, b∈[1,d] (1)
[0146] Wherein, Sub in formula (1) represents the subspace subspace, S[1] represents the value of the first dimension of the subspace subspace (which is not equivalent to the first dimension of the business data set, because the first dimension of the subspace may be a non-first dimension of the business data set), S[1] is any value or all values in the domain of the first dimension, S[b] represents the value of the bth dimension of the subspace subspace, and S[b] is any value or all values in the domain of the bth dimension.
[0147] The subspace set includes all subspaces of the business data set, and any subspace has at least one dimension with all values (the symbol "*" is used below and in the accompanying drawings to represent all values). Figure 3 The computer device performs subspace enumeration processing on the business data set 20a to obtain a subspace set. The subspace enumeration processing refers to generating all subspaces of the business data set 20a and combining them into a subspace set. Figure 3 Two subspaces in the subspace set are illustrated. The first subspace (i.e. Figure 3 Sub1) in includes two dimensions. The first dimension is the transaction date, which can take all values, including Figure 3 The first date and the second date of the example. The second dimension is the resource code, which can take all values, including Figure 3 Example X...X, Y...Y, Z...Z.
[0148] The second subspace (i.e. Figure 3 Sub2 in ) includes two dimensions, the first dimension is the transaction date, the value is the first date, and the second dimension is the resource code, the value is all values. Obviously, Figure 3The two subspaces of the example include the same dimensions, but with different dimension values, so Figure 3 Sub1 and Sub2 in represent different subspaces.
[0149] The first subspace described above can be any subspace in the subspace set. Figure 3 The subspace Sub1 in the example is described as the first subspace. The subspace Sub1 includes two dimensions. The dimension P q It can be understood as the transaction date or resource code in the subspace Sub1. Since the subspace Sub1 includes the full amount of data of the transaction date, that is, all values, the transaction date can be used as the breakdown dimension of the subspace Sub1; the computer equipment has subspace Sub1, transaction date and 3 ( Figure 3 Perform data topic enumeration processing on m example 3) metrics to obtain data topics corresponding to the subspace Sub1, wherein the data topic enumeration processing refers to generating all data topics corresponding to the subspace Sub1.
[0150] The data topic represents the data scope that the data fact focuses on and is defined as a triple, as shown in the following formula (2):
[0151] subspace = {subspace, partition dimension, one or more metrics} (2)
[0152] In the scenario where the transaction date is used as the segmentation dimension of subspace Sub1, the data topics corresponding to subspace Sub1 include Figure 3 The example subject 1 includes a subspace Sub1, with transaction date as the segmentation dimension, and first unit price as the measure.
[0153] Similarly, since subspace Sub1 includes all data of resource codes, that is, all values, resource codes can be used as the breakdown dimension of subspace Sub1; the computer device performs data subject enumeration processing on subspace Sub1, resource codes, and three metrics to obtain the data subject corresponding to subspace Sub1. In the scenario where resource codes are used as the breakdown dimension of subspace Sub1, the data subjects corresponding to subspace Sub1 include Figure 3 The subject 2 of the example includes a subspace Sub1, with resource code as the segmentation dimension, and first unit price as the metric.
[0154] Obviously, Figure 3 The two data topics (i.e., Topic 1 and Topic 2) in the example include the same subspace and metric, but different partitioning dimensions. Figure 3 The topics 1 and 2 in the figure represent different data topics.
[0155] It is important to understand that Figure 3In the example of d=3, m=3, the number of subspaces generated by the computer device for the business data set 20a is much larger than Figure 3 The sum of the number of subspaces Sub1 and Subspaces Sub2 of the example (i.e. 2); Figure 3 It mainly emphasizes that although the dimensions included are the same, different dimension values represent different subspaces.
[0156] When the transaction date is used as the segmentation dimension of the subspace Sub1, the number of data topics generated by the computer device for the subspace Sub1 is equal to 7, specifically 3+3+1, where the first number 3 can be understood as the number of data topics generated when the measurement number is 1, the second number 3 can be understood as the number of data topics generated when the measurement number is 2, and the third number 1 can be understood as the number of data topics generated when the measurement number is 3; Figure 3 The main emphasis is that although the subspaces and metrics involved are the same, different data themes are represented when the partitioning dimensions are different.
[0157] Step S103, match each data subject in the data subject set with the full set of fact types to obtain a first data subject and a first fact type that are successfully matched; the first data subject belongs to the data subject set; the first fact type belongs to the full set of fact types.
[0158] Specifically, an influence analysis is performed on each data topic in the data topic set to obtain an influence value corresponding to each data topic in the data topic set; the influence value corresponding to each data topic in the data topic set is compared with an influence threshold; the data topics in the data topic set whose influence values are greater than or equal to the influence threshold are matched with all fact types to obtain the first data topic and the first fact type that are successfully matched.
[0159] Among them, the data topics in the data topic set include the second data topic; the second data topic is any data topic in the data topic set; the second data topic includes the first metric; the first metric belongs to m metrics; an influence analysis is performed on each data topic in the data topic set, and the specific process of obtaining the influence value corresponding to each data topic in the data topic set may include: according to the segmentation dimension in the second data topic, the subspace in the second data topic is segmented to obtain multiple brother subspaces corresponding to the second data topic; the data of the multiple brother subspaces on the segmentation dimension are different from each other, and the data of the multiple brother subspaces on the dimensions other than the segmentation dimension are the same; in the business data set, the full data of the first metric is obtained, and according to the full data of the first metric, the influence analysis is performed on the multiple brother subspaces respectively to obtain the influence values corresponding to the multiple brother subspaces respectively; the influence values corresponding to the multiple brother subspaces are averaged to obtain the influence value corresponding to the second data topic.
[0160] Among them, multiple brother subspaces include brother subspace A c , c is a positive integer, and c is less than or equal to the total number of the multiple brother subspaces; according to the full amount data of the first metric, the influence analysis is performed on the multiple brother subspaces respectively, and the specific process of obtaining the influence values corresponding to the multiple brother subspaces may include: summing the full amount data of the first metric to obtain the total data corresponding to the first metric; in the full amount data of the first metric, obtaining the brother subspace A c Data on the first metric; for the sibling subspace A c The data on the first metric and the total data are processed by ratio to obtain the sibling subspace A c The corresponding influence value.
[0161] In this step, the processing process of each data topic by the computer device is the same, so the embodiment of the present application is described using one data topic as an example, and the processing process of the remaining data topics can be found in the description below.
[0162] Different data fact types correspond to different data fact analysis perspectives. For example, the insight type "Outstanding #1" corresponds to the perspective of finding "leading values that are significantly higher than other values." Specifying the data fact type is necessary to further quantify the interestingness of the data fact. For example, given a data set, the inspection strategy is different for different types such as trends or outliers.
[0163] It is understandable that by enumerating the subspace of the business data set, the computer device can obtain a larger number of subspaces, and by enumerating the data topics of the metrics in the subspace and the business data set, a larger number of data topics can be obtained. However, some data topics may not have analytical value because some data topics are not important to the business data set. Therefore, the embodiment of the present application matches the data topics with analytical value with the full amount of fact types, which can improve the analysis efficiency and focus generation of the business data set.
[0164] From a business perspective, the influence value reflects the importance of the insight topic to the entire data set. The value range is the interval [0, 1]. The larger the influence value of a data topic, the greater the influence (importance) of the data topic on the business data set. The smaller the influence value of a data topic, the smaller the influence (importance) of the data topic on the business data set.
[0165] The process of calculating the influence value of each data topic by a computer device is the same, so the present application embodiment takes one data topic as an example for description, and the influence value calculation process of the remaining data topics can be referred to the description below. Figure 3 , the second data subject can be exemplified as Figure 3 In Topic 1, the first measure included in the second data topic is the first unit price included in Topic 1, and the segmentation dimension included in the second data topic is the transaction date included in Topic 1.
[0166] According to the segmentation dimension of the transaction date in Topic 1, the computer device performs segmentation processing on the subspace in Topic 1, that is, the subspace Sub1 = [transaction date: *, resource code: *] is segmented to obtain two brother subspaces corresponding to Topic 1, such as Figure 3 The example sibling subspace Sub1.1 = [transaction date: first date, resource code: *], and the sibling subspace Sub1.2 = [transaction date: second date, resource code: *]. Obviously, Figure 3 The values of the two brother subspaces in the example on the split dimension (i.e., transaction date) are data in the domain of this dimension (the split dimension is also a dimension), so the data of the two brother subspaces on the split dimension are different from each other, and the data of the two brother subspaces on dimensions other than the split dimension are the same.
[0167] Please see again Figure 3 In the business data set 20a, the computer device obtains the full amount of data of the first unit price (i.e., the column data of the first unit price), which can be expressed as [150, 2800, 300, 152, 2810, 303]. Topic 1 corresponds to 2 sibling subspaces, such as Figure 3 The example of brother subspace Sub1.1 and brother subspace Sub1.2, the above brother subspace Ac Can be Figure 3 The example's sibling subspace Sub1.1 or sibling subspace Sub1.2. Please also refer to Figure 3 as well as Figure 4 , Figure 4 This is a data processing scenario provided by an embodiment of the present application. Figure 2 .like Figure 4 As shown, the computer device processes the first metric ( Figure 3 The example is the sum of all the data of the first unit price to obtain the total data corresponding to the first metric, such as Figure 4 The total first unit price shown = 150 + 2800 + 300 + 152 + 2810 + 303 = 6515.
[0168] In the full data of the first unit price [150, 2800, 300, 152, 2810, 303], the computer device obtains the data of the first unit price of the brother subspace Sub1.1, including 150, 2800 and 300. Since the number of data is 3, 150, 2800 and 300 are summed to obtain the total data of the first unit price of the brother subspace Sub1.1, which is 3250. For further information, please refer to Figure 4 , the computer equipment performs ratio processing on the total data 3250 and the total data of the first unit price (i.e., the total first unit price 6515), and obtains the influence value of the sibling subspace Sub1.1 in topic 1, that is, Figure 4 0.499 of them.
[0169] Similarly, in the full data of the first unit price [150, 2800, 300, 152, 2810, 303], the computer device obtains the data of the first unit price of the brother subspace Sub1.2, including 152, 2810 and 303. Since the number of data is 3, 152, 2810 and 303 are summed to obtain the total data of the first unit price of the brother subspace Sub1.2 3265. For further information, please refer to Figure 4 , the computer equipment performs ratio processing on the total data 3265 and the total data of the first unit price (i.e., the total first unit price 6515), and obtains the influence value of the sibling subspace Sub1.2 in topic 1, that is, Figure 4 0.501 of .
[0170] Furthermore, computer equipment Figure 3 as well as Figure 4 The influence values corresponding to the two sibling subspaces of the example are averaged, such as Figure 4 The influence value corresponding to topic 1 is obtained by taking (0.499+0.501) / 2, that is, Figure 4 0.5 of .
[0171] If a subspace has multiple dimensions that take all values, then the subspace has different partitioning dimensions, such as Figure 3 The example subspace Sub1 has two segmentation dimensions, namely transaction date and resource code. It is understandable that by segmenting the same subspace with different segmentation dimensions, different sibling subspace groups may be obtained (a sibling subspace group can be understood as the above-mentioned multiple sibling subspaces). Figure 3 Taking the subspace Sub1 of the example as an example, when the transaction date is used as the segmentation dimension, because the transaction date includes two values, namely the first date and the second date, the subspace Sub1 corresponds to two brother subspaces, namely Figure 3 When the resource code is used as the segmentation dimension, the resource code includes three values, namely Figure 3 In the example, X...X, Y...Y and Z...Z, the subspace Sub2 corresponds to three sibling subspaces, as shown below:
[0172] Brother subspace Sub2.1: [Transaction date: *, Resource code: X...X]
[0173] Brother subspace Sub1.2: [Transaction date: *, Resource code: Y...Y]
[0174] Brother subspace Sub1.3: [Transaction date: *, Resource code: Z...Z]
[0175] The process of the computer device determining the influence value of Topic 2 is the same as the process of determining the influence value of Topic 1, so it is not repeated here. Please refer to the description of Topic 1 above.
[0176] Please see again Figure 4 , the computer device compares the influence value 0.5 corresponding to topic 1 with the influence threshold. The influence threshold is an adjustable parameter, and the specific value can be set according to the actual application scenario. If the influence value 0.5 is less than the influence threshold, for example, the influence threshold is 0.7, the computer device deletes topic 1 from the data topic set and does not perform subsequent operations on topic 1; if the influence value 0.5 is equal to or greater than the influence threshold, for example, the influence threshold is 0.5, the computer device matches topic 1 with the full amount of fact types. Figure 4 The example full fact types include value, trend, ..., and categorization.
[0177] If there is a fact type that successfully matches Topic 1 in the full set of fact types, Topic 1 will be used as the first data topic, and the fact type that successfully matches Topic 1 will be used as the first fact type. A data topic can successfully match one or more fact types. If there is no fact type that successfully matches Topic 1 in the full set of fact types, Topic 1 will be deleted from the data topic set, and Topic 1 will not be processed later.
[0178] Fact Type refers to the way of classifying data facts, which is used to describe the nature, purpose and role of data facts in data analysis. Data fact types help analysts better understand the characteristics of data, so as to choose appropriate methods for analysis and visualization. Full fact types refer to all fact types. By referring to the classification of insights in analysis tasks, fact classification and automatic insight research, this application embodiment designs 10 data fact types. These 10 fact types cover most of the visualization charts of data stories. Please refer to Table 1. Table 1 shows the 10 fact types provided in this application embodiment and the pattern verification strategies corresponding to the 10 fact types.
[0179] Table 1
[0180]
[0181] The following explains the application scenarios of the 10 data fact types in Table 1:
[0182] Value: used to answer simple statistical questions about data attributes. For example, count: "What is the total number of digital resources in XXX", average: "What is the average rate of return of digital resources in XXX"
[0183] Proportion: used to analyze the proportion or proportional relationship of data. For example, "What is the proportion of digital resources with positive annual growth rate to the total number of digital resources?", "What is the proportion of digital resources invested in the technology sector to all digital resources?"
[0184] Difference: used to analyze the differences between different data. For example, "What is the difference in the total number of investors between digital resource A and digital resource B?", "What is the difference in the rate of return of digital resources in two different regions?"
[0185] Distribution: used to analyze the distribution characteristics of data. For example, "How is the distribution of investors in digital resource A in various regions?", "How is the distribution of digital resource yields?"
[0186] Trend: used to analyze the changes in data over time. For example, "What is the growth trend of digital resource A from January XXX to December XXX?", "What is the average rate of return of digital resources in the past five years?"
[0187] Rank: used to analyze the relative position or ranking of data. For example, "What is the revenue ranking of digital resource A among similar digital resources?", "Which digital resources ranked in the top 10 in terms of revenue in XXX year?"
[0188] Association: used to analyze the correlation or association between data. For example, "Is there a relationship between the rate of return of digital resource A and the increase of XX sector?", "Is there a correlation between the number of investors in digital resource A and the scale of digital resources?"
[0189] Extreme: used to analyze the maximum, minimum, or extreme values in the data. For example, "What are the top three digital resources with the highest returns in XXX?", "What are the top three digital resources with the highest losses in XXX?"
[0190] Categorization: used to filter or categorize data. For example, "What are the digital resources that invest in the technology sector?", "What are the digital resources with a return rate of more than 10%?"
[0191] Outlier: used to detect anomalies in data. For example, "Are there any anomalies in the increase or decrease of digital resources?", "Which digital resources have a significant deviation from other digital resources in terms of their yields?"
[0192] Step S104: extract facts from the first data subject according to the first fact type to obtain data facts corresponding to the business data set.
[0193] Specifically, an importance analysis is performed on the first data subject to obtain an importance score corresponding to the first data subject; if the importance score corresponding to the first data subject is equal to or greater than the importance score threshold, facts are extracted from the first data subject according to the pattern verification strategy corresponding to the first fact type to obtain data facts corresponding to the business data set.
[0194] Among them, the specific process of performing importance analysis on the first data topic and obtaining the importance score corresponding to the first data topic may include: according to the segmentation dimension in the first data topic, segmenting the subspace in the first data topic to obtain multiple brother subspaces corresponding to the first data topic; the data of the multiple brother subspaces on the segmentation dimension are different from each other, and the data of the multiple brother subspaces on dimensions other than the segmentation dimension are the same; in the business data set, obtaining the full data of metric z in the first data topic; metric z belongs to m metrics; based on the full data of metric z and the multiple brother subspaces, performing importance analysis on the first data topic to obtain the importance score corresponding to the first data topic.
[0195] Among them, the specific process of performing importance analysis on the first data topic based on the full data of metric z and multiple brother subspaces to obtain the importance score corresponding to the first data topic may include: in the full data of metric z, obtaining the data corresponding to multiple brother subspaces on metric z respectively, and determining the obtained multiple data as a result set; performing a significance test on the result set to obtain the significance value corresponding to the first data topic; multiplying the significance value corresponding to the first data topic and the influence value corresponding to the first data topic to obtain the importance score corresponding to the first data topic.
[0196] As can be seen from step S103, the first data subject includes the data subject that successfully matches the fact type in the data subject set, and the number of the first data subject can be one or more. When there are multiple first data subjects, the processing process of each first data subject by the computer device is the same, so the following description is based on one first data subject.
[0197] The first data topic may not have analytical value, because there are some data topics that are not important to the business data set. Therefore, the embodiment of the present application extracts facts from data topics that have analytical value, which can improve the analysis efficiency and focus capture of the business data set. The embodiment of the present application assigns an appropriate score to the first data topic to quantify its "interestingness". Intuitively, the interestingness of a data fact is judged by two factors, namely the influence (Imp) value and the significance (SigT) value. These two indicators can measure the value of the data topic. Among them, the influence value is to calculate the influence of the data topic. The specific calculation process can be found in the description of step S103 above, which will not be repeated here. The significance value reflects the importance of a data topic in terms of statistical characteristics. The higher the score, the less common and unexpected the data fact corresponding to the data topic, and the more it can arouse the interest of users. This is the data fact that deserves special attention.
[0198] The specific implementation process of the computer device obtaining the full data of measurement z in the first data topic in the business data set is the same as the specific implementation process of the computer device obtaining the full data of the first unit price in topic 1 in the business data set 20a in step S103, so it will not be repeated here.
[0199] Similarly, the specific implementation process of the computer device obtaining the data corresponding to the multiple brother subspaces on the metric z from the full data of the metric z is the same as the specific implementation process of the computer device obtaining the data corresponding to the first unit price of the two brother subspaces from the full data of the first unit price in step S103, so it will not be repeated here. If there are multiple metrics z, the computer device generates a result set corresponding to each metric z, so multiple result sets are generated. It can be understood that the computer device processes each result set in the same way, so it will not be repeated here. The computer device performs mean processing on the multiple significance values corresponding to the data theme, and determines the mean as the final significance value of the data theme.
[0200] A data in the result set refers to the data of a sibling subspace on a measure in the data topic (if there are multiple data, it is the sum of multiple data, for example Figure 4 The total data of the brother subspace Sub1.1 on the first unit price is 3250). The embodiment of the present application does not limit the method of performing significance test on the result set, and can be set according to the actual application. The point significance value of the data theme is positively correlated with the significance of analyzing the data theme.
[0201] A method of performing a significance test on a result set is as follows: a computer device performs a point significance test on the result set to obtain a point significance value, and determines the point significance value as the significance value corresponding to the first data theme. Assume that the result set S = {s 1 ,s 2 ,…,s r}, where r is equal to the number of sibling subspaces corresponding to a data topic, for example Figure 4 The subject 1 of the example includes 2 sibling subspaces, then r=2.
[0202] The result data of the result set usually follows a power law distribution. Therefore, the null hypothesis (H0) of the result set is set as follows in this embodiment of the application:
[0203] H0:S follows a power-law distribution with Gaussian noises
[0204] x follows a power-law distribution of Gaussian noise
[0205] When H0 is true, the significance reveals that the maximum value in the result set S is significantly different from the rest of the values. The significance calculation process of the result set is mainly divided into the following key steps:
[0206] 1. Data sorting: Sort the data set S in descending order to identify the maximum value. For example, for a data set S = {1, 2, 3, 4, 5, 16}, the sorted S = {16, 5, 4, 3, 2, 1}.
[0207] 2. Maximum value identification: Identify the maximum value s from the sorted data max , in the example of step 1, s max =16.
[0208] 3. Power law distribution fitting: After removing the maximum value, the remaining data (S\{s max}) to fit a power-law distribution. This usually involves using statistical methods to estimate the parameters of the power-law distribution, and the fitting process may use least squares, maximum likelihood estimation, or other statistical fitting techniques. If S\{s max} is normal, please refer to Figure 5a , Figure 5a This is a data distribution diagram provided in an embodiment of the present application.
[0209] 4. Predict the maximum value: Use the fitted power law distribution parameters to predict the maximum estimate s' max This may involve calculating the quantiles of a power law distribution, or using the inverse CDF to find the value corresponding to s max s' max .
[0210] 5. Calculate the prediction error ε: After fitting the power law distribution, calculate the difference between the predicted value and the actual observed value of each data point, that is, the prediction error ε i =s' i -s i , the residual (i.e., prediction error) reflects the deviation between the prediction and the actual data. max is the maximum value s' predicted by max Subtract the actual maximum value s max to determine, that is, ε max =s' max -s max .
[0211] 6. Gaussian distribution assumption: The prediction error ε is assumed to approximately follow the Gaussian distribution N(μ, δ). μ represents the mean value of the prediction error, and δ represents the standard deviation of the prediction error. Please also refer to Figure 5b , Figure 5b It is a schematic diagram of a Gaussian distribution provided in an embodiment of the present application.
[0212] 7. p-value calculation: Under Gaussian distribution N(μ, δ), the calculated prediction error is greater than the observed ε max The probability of max |N(μ,δ)).
[0213] 8. Point significance calculation: Point significance value = 1-p, which means that under the condition that H0 is true, the observed s max With the predicted value s max The rarity of the difference between .
[0214] Another way to perform a significance test on the result set is as follows: the computer device performs a slope significance test on the result set to obtain a slope significance value, and determines the slope significance value as the significance value corresponding to the first data topic. Assume that the result set S = {s 1 ,s 2 ,…,s r}, where r is equal to the number of sibling subspaces corresponding to a data topic.
[0215] Usually the data trend of the result set is neither rising nor falling. Therefore, the null hypothesis H0 is set as:
[0216] H0:S forms a shape with slope≈0
[0217] S forms a shape with a slope
[0218] In data analysis, analysts are more likely to be attracted by obvious upward / downward trends, and the slope of a trend that changes dramatically is very different from the slope of a gentle trend of 0. Therefore, the p-value should measure how much the slope differs from 0. The computer equipment fits the result set S to a straight line through linear regression analysis, please refer to Figure 6a , Figure 6a It is a schematic diagram of a straight line fitted by linear regression provided in an embodiment of the present application.
[0219] The computer equipment calculates the slope and the fitting degree r of the fitting line 2 In the present embodiment, the Logistic distribution L(μ,λ) (a continuous probability distribution) is used to simulate the distribution of the slope of the straight line, where the parameters μ and λ are constant parameters. Figure 6b , Figure 6bis a schematic diagram of a Logistic distribution provided in an embodiment of the present application. Figure 6a The p-value is the probability that the slope value is equal to or greater than the slope of the observed upward trend, so the p-value is calculated using the following formula (3).
[0220] p=Pr(s>|slope||L(μ,λ)) (3)
[0221] Wherein, s in formula (3) represents the slope, and |slope| represents the absolute value of the fitting slope.
[0222] After obtaining the P value, define the slope significance as: r 2 (1-p), where the goodness of fit value r 2 are used as weights.
[0223] Another way to perform a significance test on the result set is as follows: perform a point significance test on the result set to obtain a point significance value, and perform a slope significance test on the result set to obtain a slope significance value; if the point significance value is greater than or equal to the slope significance value, then the point significance value is determined as the significance value corresponding to the first data theme; if the point significance value is less than the slope significance value, then the slope significance value is determined as the significance value corresponding to the first data theme.
[0224] After determining the significance value corresponding to the first data topic and the influence value corresponding to the first data topic, the computer device can calculate the importance score of the first data topic by the following formula (4).
[0225] Score((S,D i ), T) = Imp(SG(S,D i ))*SigT(Φ) (4)
[0226] Among them, (S,D i ) in which S represents subspace, D i represents the segmentation dimension of the subspace, SG(S,D i ) represents a set of sibling groups, which includes the above-mentioned multiple sibling subspaces, T represents the fact type type, which can also be called the insight angle, Imp(SG(S,D i )) represents the influence value of the data topic, and SigT(Φ) represents the significance value of the data set Φ, that is, the significance value corresponding to the data topic.
[0227] The importance score threshold is an adjustable parameter, and the specific value can be set according to the actual application scenario.
[0228] Step S105 , visualizing the data fact according to a chart type adapted to the first fact type, to obtain a visualized chart for displaying the data fact.
[0229] Specifically, if the first data subject has multiple successfully matched fact types, then the number of first fact types corresponding to the first data subject is multiple, and the process of fact extraction of the first data subject by the computer device is the same according to each first fact type. Therefore, if there are multiple first fact types, the computer device obtains the pattern verification strategy corresponding to each first fact type, and extracts facts from the first data subject according to each pattern verification strategy to obtain data facts corresponding to the business data set. Therefore, if there are multiple first fact types, the computer device can generate multiple data facts corresponding to the first data subject, where one data fact is obtained according to one fact type and the first data subject.
[0230] The data fact can be modeled as a five-tuple, as shown in the following formula (5):
[0231] fact={subspace(s), breakdown, measure(s), type, score} (5)
[0232] Please also see Figure 7 , Figure 7 This is a flow chart of a data processing method provided in an embodiment of the present application. Figure 2 .like Figure 7 As shown, the data processing method may include steps S201 to S211.
[0233] Step S201: input a business data set.
[0234] Step S202, enumerate data subjects.
[0235] First, list the subspaces, determine the segmentation dimensions of the subspaces, and combine the subspaces, segmentation dimensions, and one or more metrics into a data topic.
[0236] Step S203, calculating the influence value of the data topic.
[0237] Step S204, determining whether the influence value of the data topic is equal to or greater than the influence threshold.
[0238] If the influence value of the data topic is equal to or greater than the influence threshold, execute step S205; if the influence value of the data topic is less than the influence threshold, re-execute step S202. It can be understood that the data topics listed in the re-execution of step S202 are different from the data topics listed in the historical execution of step S202, that is, a new data topic will be listed after each execution of step S202.
[0239] The influence value of a data topic can reflect the importance and influence of the subspace of the data topic on the measurement of the data topic. The embodiment of the present application filters out data topics with low influence and low importance, which can not only reduce the subsequent calculation amount, but also highlight data facts with great influence and importance.
[0240] Step S205, enumerate each fact type.
[0241] Please refer to Table 1 again. The embodiment of the present application enumerates 10 fact types, namely Value, Proportion, Difference, Distribution, Trend, Rank, Association, Extreme, Categorization, and Outlier.
[0242] Step S206, determining whether the data subject matches the trend type successfully.
[0243] If the data subject matches the trend type successfully, step S207 is executed. If the data subject does not match the trend type, that is, the matching fails, step S205 is re-executed. It can be understood that the fact types listed in step S205 are different from the currently listed fact types (i.e., trend types). For example, after re-execution of step S205, associated types, also known as related types, are listed.
[0244] The correlation model studies the degree of correlation between two or more variables and uses correlation functions, such as the Pearson correlation coefficient analysis function, to express the correlation between phenomena. Therefore, the data subject has two measurement numbers.
[0245] It is understandable that, for the same data subject, if each fact type fails to match, after step S206 is executed, step S202 is re-executed, that is, a new data subject is re-listed.
[0246] Step S207, calculating the significance value of the data topic.
[0247] Using the result set of data topics, the computer device calculates a significance value for the data topics.
[0248] Step S208, calculating the importance score of the data topic.
[0249] The computer device determines the product of the influence value corresponding to the data topic and the significance value of the data topic as the importance score of the data topic.
[0250] Step S209, determining whether the importance score of the data topic is greater than or equal to the importance score threshold.
[0251] If the importance score of the data topic is greater than or equal to the importance score threshold, step S210 is executed; if the importance score of the data topic is less than the importance score threshold, step S202 is re-executed because the importance of the data topic is low.
[0252] Different from the order of executing the steps described above, another feasible implementation is to first execute steps S207-S209. If the importance score of the data topic is greater than or equal to the importance score threshold, then execute steps S205-S206; if the importance score of the data topic is less than the importance score threshold, then re-execute step S202.
[0253] Step S210, extracting data facts using a pattern verification strategy corresponding to the trend type.
[0254] Please refer to Table 1 again. The pattern verification strategy corresponding to the trend type can be the Mann-Kendall trend verification method. The generated derivative word Derived Value is an upward trend or a downward trend in a certain dimension. For example, a data fact is that from January XXX to December XXX, the growth trend of digital resource A increased.
[0255] Step S211, using the line chart type corresponding to the trend type, visualize the data facts to obtain a line chart.
[0256] The embodiment of the present application designs 10 data fact types, and each type can pre-set its adapted chart type. For example, for the Value type, a horizontal bar chart can be used to visualize the data fact; for the Trend type, a line chart can be used to display the data fact; and for the Association type, a scatter plot visualization can be used to convey associated information. The embodiment of the present application allows data analysts to pre-configure the chart type adapted to the data fact.
[0257] As can be seen from the above, the embodiments of the present application can not only improve the completeness and accuracy of data fact extraction, but also enhance the readability of the chart after data fact visualization. In addition, by determining the influence value and importance score, data topics with low analysis value can be screened out, thereby improving the analysis efficiency of the data.
[0258] See also Figure 8 , Figure 8This is a flow diagram of a data processing method provided in an embodiment of the present application. Figure 3 The method may be performed by a business server (eg, Figure 1 The service server 100 shown in the figure may also be executed by a terminal device (for example, the above Figure 1 The terminal device 200a) shown in the figure can also be executed by the service server and the terminal device interactively. For ease of understanding, the embodiment of the present application takes the method executed by the service server as an example for explanation, that is, the service server executes the data processing method as a computer device. Figure 8 As shown, the process of the data processing method includes steps S301 to S308.
[0259] Step S301, obtaining a business data set including data corresponding to d dimensions and data corresponding to m metrics; d and m are both positive integers.
[0260] Specifically, multidimensional data is conceptually organized in a table format, consisting of multiple rows of records, each of which is represented by a set of attributes. Dimensions are categorical or ordinal columns, and measures can perform certain aggregation operations, such as summing SUM, calculating average AVG, etc. A multidimensional data set can be represented as R(D, M), where D = <D 1 , …, D d > is the set of dimensions, and M is the set of metrics. dom(D i ) means D i The domain of , that is, the value range of the i-th column dimension.
[0261] In order to realize the automation of the business data visualization video system, the embodiment of the present application innovatively proposes an automated work pipeline. Please refer to Fig. 9 , Fig. 9 Schematic diagram of an automated work pipeline provided in an embodiment of the present application. Fig. 9 As shown in the figure, the automated work pipeline mainly consists of two parts. The first part is to make a story: the process from data to animation can be regarded as making a story; the second part is to create a video: the process from animation to video can be regarded as creating a video.
[0262] The story making part consists of the first stage and the second stage. In the first stage, the automatic insight module automatically extracts data facts. The system analyzes the data characteristics of the original data set, and uses statistical algorithms to analyze the metadata attributes of the original data to provide a basis for automatic insight. The automatic insight function of the system uses multiple pattern extractors to extract high-quality data story films within a limited time. The generated data facts are visualized as a series of titled visual charts in the story editor, which makes it convenient for data analysts to intuitively understand the meaning of each data story film, and are passed into the sequence optimizer as input. In summary, the first stage realizes automatic data analysis (Auto Analyze) and editing facts (Edit Facts), that is, editing the facts obtained from the analysis. For the specific implementation process, please refer to the description of steps S302-S305 below.
[0263] In the second stage, the sequence optimization module automatically arranges and organizes the data facts. In this stage, there are two ways to perform the logical sorting of the data story film: manual and automatic. The latter method realizes automatic sorting of facts (AutoSequence) and collection of sorted facts (Gather order facts). For the specific implementation process, please refer to the description of step S306 below.
[0264] The video creation part consists of the third and fourth stages. In the third stage, the chart animation is generated. The system automatically maps the data story sequence to chart animations (Map facts to animations) and automatically edits the appropriate duration for the data fact animations (Edit Animations duration). For the specific implementation process, please refer to the description of step S307 below.
[0265] The fourth stage is to perform data video synthesis (Compose). The user can upload multimedia content that enriches the audio video, and can control the duration of the final video, and can download the generated video. For the specific implementation process, please refer to the description of step S308 below.
[0266] Step S302, enumerate the data topics of the business data set to obtain a data topic set; the dimensions included in each data topic in the data topic set belong to d dimensions; the metrics included in each data topic in the data topic set belong to m metrics.
[0267] Specifically, when performing data analysis, the change of the cube unit is usually analyzed by changing a single dimension, and the single dimension here is the segmentation dimension. Therefore, the embodiment of the present application defines a brother group (including multiple brother subspaces) to cover subspaces that differ only in the segmentation dimension. Given a subspace Sub and a dimension D i, a sibling group is defined as a group of dom(D i ), also referred to as D i For the brother group SG (S, D i ) is the breakdown dimension. i ) can be expressed by the following formula (6):
[0268]
[0269] Where S' in formula (6) represents the sibling subspace, for example Figure 3 Example of sibling subspace sub1.1 and sibling subspace sub1.2.
[0270] In order to extract data facts from business data sets efficiently and quickly, this application embodiment proposes a set of automated data fact extraction algorithm frameworks. Please refer to Fig.10 , Fig.10 Schematic diagram of a data fact extraction algorithm framework provided by an embodiment of the present application. Fig.10 As shown, the data fact extraction algorithm can be divided into three stages: stage 1, topic search and task generation; stage 2, data fact pattern calculation; stage 3, data fact visualization. The first two stages can be executed simultaneously in parallel within the time budget. When the time exceeds the time budget, the data facts with importance scores equal to or greater than the importance score threshold are output in stage 3 for visualization. For details, please refer to the description of steps S303 to S305 below.
[0271] Step S303, match each data subject in the data subject set with the full set of fact types to obtain a first data subject and a first fact type that are successfully matched; the first data subject belongs to the data subject set; the first fact type belongs to the full set of fact types.
[0272] Specifically, the number of the first metrics is at least two; the at least two first metrics include the second metric, and the second metric is any one of the at least two first metrics; the total data includes the total data corresponding to the second metric; for the brother subspace A c The data on the first metric and the total data are processed by ratio to obtain the sibling subspace A c The corresponding influence values include: c The data on the second metric and the total data corresponding to the second metric are processed by ratio to obtain the sibling subspace A c Influence value on the second metric; Get the sibling subspace A c The influence values corresponding to at least two first metrics are weighted averaged to obtain the sibling subspace A.c The corresponding influence value.
[0273] The first metric mentioned above refers to the metric in the data subject. Figure 2 Step S103 of the corresponding embodiment mainly describes that the data subject includes a metric. In actual application, a data subject may include multiple metrics. When a data subject includes multiple metrics, the process of the computer device calculating the influence value of the data subject is similar to that described in step S102 above, except that the influence value of the subspace on each metric is first counted, and the multiple influence values are weighted averaged to obtain the influence value corresponding to the subspace. Among them, the weighted value corresponding to each metric can be set according to the actual application scenario.
[0274] Please see again Fig.10 ,In the first stage, i.e., the stage of topic search and task generation, ,the topic searcher obtains the business data set and tries to list all the subspaces of the business data set. Each ,subspace is calculated by the impact calculator to calculate the impact score of the corresponding ,data topic and filter out the data topics with low ,impact scores.
[0275] Each data subject with a high influence score is tested for data pattern matching through a pattern matcher. The Checker verifies in advance whether the data subject meets the requirements for observing data distribution from a certain perspective, which refers to the data fact type. If the subject does not meet the requirements, it will not be processed further and will be deleted or discarded directly.
[0276] Data fact types cover common problems and requirements in data analysis and can be used to understand and interpret data from different perspectives. They can be applied to the following scenarios:
[0277] Descriptive analysis: Understand the basic characteristics of data through types such as Value, Proportion, and Distribution.
[0278] Comparative analysis: Compare the differences between different data through Difference, Rank and other types.
[0279] Trend analysis: Use the Trend type to analyze how data changes over time.
[0280] Association analysis: Explore the relationship between data through the Association type.
[0281] Anomaly detection: Identify anomalies in data through outlier types.
[0282] Categorization and filtering: Use the Categorization type to filter data that meets specific conditions.
[0283] The above data fact types are the basis of data analysis and data visualization, and can help analysts understand and utilize data more comprehensively.
[0284] The type matcher Checker verifies the automatic insight perspective that each data range is suitable for in advance. The data subject and the data fact type that needs to be verified are passed to the type matcher to obtain a returned Boolean value. If the Boolean value is true, it means that the type is compliant. If it is false, it means that the type is not compliant. Then the extraction task of the pattern in this data subject range can be pruned to optimize the performance of the system. For example, Checker (subject, Trend) = false, Checker (subject, outlier) = true, which indicates that from the perspective of the outlier, there is a valid pattern in the data subject range, so a calculation task is generated to extract the outlier data fact type for this data range. Because the subsequent data aggregation and pattern extraction calculation amount is much larger than the previous topic influence calculation, this can save the system's computing resources and improve the efficiency of extracting data facts.
[0285] Step S304: extract facts from the first data subject according to the first fact type to obtain data facts corresponding to the business data set.
[0286] Specifically, the number of the first data subjects is at least two; the at least two first data subjects include data subject G h and data subject G i , h and i are both positive integers, and h and i are both less than or equal to the total number of at least two first data topics, and h is different from i; performing importance analysis on the first data topic to obtain the importance score corresponding to the first data topic may include: if the data topic G h The influence value is greater than or equal to the data subject G i The influence value of the first importance assessment task is determined to have a higher execution priority than the second importance assessment task; the first importance assessment task is used to complete the data subject G h The second importance assessment task is used to complete the data theme G iimportance analysis; adding the first importance evaluation task and the second importance evaluation task to the task queue in order of execution priority from high to low; determining the idle number corresponding to the idle threads in the parallel execution thread pool, and if the idle number is equal to or greater than the parallel execution thread threshold, obtaining the third importance evaluation task from the task queue in order of execution priority from high to low through the idle thread; the number of the third importance evaluation task is equal to or less than the idle number; the third importance evaluation task is used to complete the data topic G in at least two first data topics j The importance analysis of j is a positive integer, and j is less than or equal to the total number of at least two first data topics; through the idle thread, the third importance evaluation task is executed in parallel to obtain the data topic G j The corresponding importance score; wherein the importance score corresponding to a data topic is obtained by performing a third importance assessment task in an idle trip.
[0287] The first fact type includes the data subject G j The second fact type that matches successfully; if the importance score corresponding to the first data subject is equal to or greater than the importance score threshold, then according to the pattern verification strategy corresponding to the first fact type, the first data subject is extracted for facts, and the specific process of obtaining the data facts corresponding to the business data set may include: if the data subject G j If the corresponding importance score is equal to or greater than the importance score threshold, then according to the pattern verification strategy corresponding to the second fact type, the data subject G j Perform fact extraction to obtain data facts corresponding to the business data set.
[0288] Please see again Fig.10 ,After passing the pattern matcher, the computer device combines the subspace with the metric, generates the importance evaluation tasks, and sorts them according to the ,influence values, with the larger ones in the front and the smaller ones in the back. ,The generated tasks are stored in a task queue and are executed in the
[0289] In the second stage, i.e. the data fact pattern calculation stage, the tasks in the task queue are calculated in parallel by a group of dedicated worker threads. The worker threads obtain high-priority importance assessment tasks from the task queue. When executing an importance assessment task through a thread, the significance value of the data topic corresponding to the task is first calculated through the significance calculator, and the importance score of the data topic is obtained by multiplying the significance value of the data topic and the influence value of the data topic.
[0290] The data topics whose importance scores are greater than or equal to the importance score threshold are passed as parameters to the fact extractor for pattern extraction to obtain data facts.
[0291] The fact extractor uses different pattern verification algorithms (see Table 1) to extract data facts. Pattern verification algorithms mainly use mathematical statistics, pattern recognition and other technologies to verify whether the data subject composed of subspaces, dimensions and metrics meets certain rules. For example, for trend types, the Mann-Kendall trend test is used to extract data facts.
[0292] The pattern verification strategy for each data type, the embodiment of the present application presets some commonly used analysis algorithms in exploratory analysis to perform pattern verification (refer to Table 1), but the pattern verification strategy is extensible, and data analysts can define some personalized algorithms according to business characteristics to meet the needs of pattern recognition.
[0293] Step S305 , visualizing the data fact according to a chart type adapted to the first fact type, to obtain a visualized chart for displaying the data fact.
[0294] Specifically, the computer device stores data facts as input data for the visual stage of the data facts, and realizes the visualization of the data facts to the data analyst. Fig.10 This example shows a bar chart showing sales in different months.
[0295] The computer device may pre-set a chart type for each fact type. When visualizing a data fact, a visualization chart of the chart type set for the fact type is generated based on which fact type it matches.
[0296] Narrative text is indispensable in narrative visualization videos, which can promote user data cognition, convey insights, and allow users to quickly capture data information. Therefore, in addition to presenting data facts with visual charts, data facts can also rely on brief descriptions to enrich the content, so that users can quickly understand data facts through text descriptions. The embodiment of the present application adopts a grammatical template-based method to generate a corresponding description for each data fact type. The computer device obtains the grammatical template corresponding to the fact type, fills in the grammatical template according to the fact type and the data fact, and obtains the description text corresponding to the visualization chart.
[0297] Step S306: if the total number of visualization charts is at least two, then sorting is performed on the at least two visualization charts to obtain a chart optimization sequence of the visualization charts.
[0298] Specifically, according to the total number of at least two visual charts, the sorting number X is determined; X is a positive integer greater than 1; at least two visual charts are sorted X times to obtain X chart sequences; the X chart sequences each include at least two visual charts, and the chart sequences corresponding to the X chart sequences are different; the X chart sequences include chart sequence K n , n is a positive integer, and n is less than or equal to X; for the chart sequence K n Perform cost analysis and obtain chart sequence K n corresponding sequence costs; determining the sequence costs corresponding to the X chart sequences respectively, and obtaining the minimum sequence cost from the X sequence costs; determining the chart sequence corresponding to the minimum sequence cost among the X chart sequences as the chart optimization sequence.
[0299] Among them, for the chart sequence K n Perform cost analysis and obtain chart sequence K n The specific process of the corresponding sequence cost may include: n Perform chart conversion cost analysis to obtain chart sequence K n The corresponding chart conversion cost; for the chart sequence K n Perform graph filtering cost analysis to obtain graph sequence K n The corresponding chart filtering cost; sum the chart conversion cost and the chart filtering cost to obtain the chart sequence K n The corresponding sequence cost.
[0300] Among them, for the chart sequence K n Perform chart conversion cost analysis to obtain chart sequence K n The specific process of the corresponding chart conversion cost may include: obtaining the chart sequence K n Every two adjacent visualization charts in n Each two adjacent visualization charts in the graph include an adjacent first visualization chart and a second visualization chart; if there is a conversion relationship between the first visualization chart and the second visualization chart, the conversion cost corresponding to the conversion relationship is determined as the chart conversion cost of the first visualization chart and the second visualization chart; if there is no conversion relationship between the first visualization chart and the second visualization chart, a chart path for connecting the first visualization chart and the second visualization chart is determined; each two adjacent visualization charts in the chart path have a conversion relationship; according to the conversion cost corresponding to the conversion relationship in the chart path, the chart conversion cost of the first visualization chart and the second visualization chart is determined; for the chart sequence K n The chart conversion costs corresponding to each two adjacent visualization charts in the summation process are processed to obtain the chart sequence K n The corresponding chart conversion cost.
[0301] Among them, for the chart sequence K n Perform graph filtering cost analysis to obtain graph sequence K n The specific process of corresponding chart filtering cost may include: obtaining chart sequence K n Every two adjacent visualization charts in n Each two adjacent visualization charts in the graph include an adjacent first visualization chart and a second visualization chart; if the dimension displayed by the first visualization chart is the same as the dimension displayed by the second visualization chart, the first value is determined as the chart distance between the first visualization chart and the second visualization chart; if the dimension displayed by the first visualization chart is different from the dimension displayed by the second visualization chart, the second value is determined as the chart distance between the first visualization chart and the second visualization chart; the first value is greater than the second value; determine the chart sequence K n The chart distance corresponding to each two adjacent visualization charts in the graph sequence K n The chart distances corresponding to each two adjacent visualization charts in the image are used to determine the chart sequence K. n The corresponding chart filtering cost.
[0302] The computer device obtains at least two visualization charts, determines a directly connected visualization chart having a conversion relationship among the at least two visualization charts, and determines a conversion cost of the directly connected visualization chart according to the conversion relationship of the directly connected visualization charts. Fig.11 , Fig.11 This is a data processing scenario provided by an embodiment of the present application. Figure 3 . Fig.11 The example has 4 visualization charts, namely Fig.11 Figure a, Figure b, Figure c, and Figure d in the figure.
[0303] By comparing the data facts expressed by the four charts, the computer device determines that there is a mark type conversion relationship between Chart A and Chart B, that is, by performing a mark type conversion on Chart A (Chart B), Chart B (Chart A) can be obtained, so Chart A and Chart B are a directly connected visualization chart. Among them, the conversion cost of the mark type conversion is 1, so the chart conversion cost between Chart A and Chart B is 1. Mark type conversion involves changing the basic graphic elements used to display data in the chart, such as switching from a bar chart to a line chart or a scatter chart. Mark type conversions usually have little impact on user perception because they mainly change the visual representation of data without changing the encoding method of the data.
[0304] Please see again Fig.11, by comparing the data facts expressed by the four charts, the computer device determines that there is a data conversion relationship between chart a and chart c, that is, by performing data conversion on chart a (chart c), chart c (chart a) can be obtained, so chart a and chart c are a directly connected visualization chart. Among them, the conversion cost of data conversion is 2, so the chart conversion cost between chart a and chart c is 2. Similarly, the computer device determines that there is a data conversion relationship between chart b and chart c, that is, by performing data conversion on chart b (chart c), chart c (chart b) can be obtained, so chart b and chart c are a directly connected visualization chart. Among them, the conversion cost of data conversion is 2, so the chart conversion cost between chart b and chart c is 2. Data transformation includes operations on data, such as filtering, aggregation, sorting, etc., which can change the collection of data or the way it is organized. Data transformation may affect users' understanding of data because they change the context of the data or the data points emphasized.
[0305] Please see again Fig.11 By comparing the data facts expressed by the four charts, the computer device determines that there is a visual coding conversion relationship between Chart d and Chart c, that is, Chart d (Chart c) is converted into a visual coding, and Chart c (Chart d), so Chart d and Chart c are a directly connected visual chart. Among them, the conversion cost of the visual coding conversion is 3, so the chart conversion cost between Chart d and Chart c is 3. Visual coding conversion involves changing the mapping relationship between data fields and visual channels (such as color, size, position, etc.). This conversion has the greatest impact on users because it changes the way data is presented and may significantly affect users' interpretation and analysis of data. For example, changing a numerical variable from mapping to the height of a bar chart to mapping to the depth of color will change users' understanding of the relative size and distribution of data.
[0306] In summary, tag type conversion is the lowest-cost editing operation, followed by data transformation, and the highest-cost editing operation is visual encoding operations, because visual encoding operations essentially change the data fields being visualized.
[0307] Further, the computer device determines an indirect visualization graph without a conversion relationship among at least two visualization graphs, determines a connection path of the indirect visualization graph based on the direct visualization graph, and determines a conversion cost of the indirect visualization graph based on the connection path of the indirect visualization graph and the conversion cost of the direct visualization graph.
[0308] Please see again Fig.11In chart b, chart c and chart d, chart b and chart c have direct conversion relationship with chart a, and there is no conversion relationship between chart d and chart a. Therefore, chart a and chart d are an indirect visualization chart. Through data conversion between chart a and chart c, and visualization coding conversion between chart c and chart d, chart a can be converted to chart d. Therefore, the chart conversion cost between chart a and chart d is 2+3=5.
[0309] In chart a, chart c and chart d, chart a and chart c have direct conversion relationships with chart b, and there is no conversion relationship between chart d and chart b. Therefore, chart b and chart d are an indirect visualization chart. Chart b can be converted to chart d through data conversion between chart b and chart c, and visualization coding conversion between chart c and chart d. Therefore, the chart conversion cost between chart b and chart d is 2+3=5.
[0310] Further, the computer device constructs at least two nodes for representing at least two visual graphs; wherein one node is used to represent one visual graph; and according to the conversion relationship between the at least two visual graphs, the edge weight of the directed edge between each two nodes in the at least two nodes is determined. As described above, the embodiment of the present application can determine the graph conversion cost between directly connected visual graphs according to the conversion relationship between directly connected visual graphs, and can also determine the graph path for connecting indirectly connected visual graphs, for example Fig.11 The graph paths of the example graph a and graph b are: graph a - graph c - graph b, so the graph conversion cost between the indirect visualization graphs can also be determined.
[0311] According to at least two nodes and the edge weights of the directed edges between each two nodes, the computer device can construct a directed graph model, such as Fig.11 The directed graph model 80a in the example may be a GraphScape model, which is a directed graph model for visualization design space, and is mainly used to support automatic reasoning about visualization similarity and ranking. The model uses Vega-Lite language (a high-level declarative language for describing the specifications of interactive visualization diagrams) to model a single diagram and uses a directed graph to represent the transformation relationship between diagrams.
[0312] According to the total number of at least two visualization charts, the computer device determines the sorting number X, wherein X=u!, that is, X is equal to the factorial of u, and u represents the total number of at least two visualization charts. Fig.11In the embodiment of the present application, four charts are taken as an example, and 4*3*2*1=24 sorting results can be obtained by sorting the four charts, that is, 24 chart sequences. For example, the chart sequence is abcd, which means that chart a is converted to chart b, chart b is converted to chart c, and chart c is converted to chart d.
[0313] The process of cost analysis for each chart sequence by the computer device is the same, so this step describes a chart sequence. The sequence cost of a chart sequence consists of two parts: one is the chart conversion cost of the chart sequence, which reflects the difficulty of editing and conversion of the chart sequence, and the other is the filtering conversion cost of the chart sequence, which reflects the degree of duplication of the dimensions of the chart sequence.
[0314] The chart conversion cost T of a chart sequence can be determined by the following formula (7):
[0315]
[0316] Where r in formula (7) represents the number of charts in the chart sequence, for example Fig.11 The number of examples with at least two visualization charts is 4; (v i , v i+1 ) represents two adjacent visualization charts, T(v i , v i+1 ) represents the chart transition cost between two adjacent visualization charts.
[0317] Please see again Fig.11 , according to formula (7), the chart conversion cost of the chart sequence sorted by abcd is as follows:
[0318] T(node a, node b)+T(node b, node c)+T(node c, node d)=1+2+3=6
[0319] Please see again Fig.11 , according to formula (7), the chart conversion cost of the chart sequence sorted by adcb is as follows:
[0320] T(node a, node d)+T(node d, node c)+T(node c, node b)=5+3+2=11 The filtering conversion cost F of the graph sequence can be determined by the following formula (8) and formula (9):
[0321]
[0322] Wherein, |V| in formulas (8)-(9) represents the number of charts in the chart sequence, for example Fig.11 The number of examples with at least two visualization charts is 4; (v i-1 , v i) represents two adjacent visualization charts, D i-1 Represents a node (representing a visual graph) v i-1 The dimension, D i Represents a node (representing a visual graph) v i The dimension, L i-1 Represents a node (representing a visual graph) v i-1 The number in the chart sequence, L i Represents a node (representing a visual graph) v i In the sequence of numbers in the chart, the first value is equal to 0 and the second value is equal to -1; d(v i-1 , v i ) represents the distance between two adjacent visualization charts; α is an adjustable parameter, such as 0.1, which is used to avoid d(v i-1 , v i ) is equal to 0; the |V|-1 in the denominator is the distance between adjacent nodes.
[0323] Combining formula (7) and formula (8), the sequence cost of the chart sequence can be determined by the following formula (10):
[0324]
[0325] Step S307 , performing animation mapping processing on each visual chart in the chart optimization sequence to obtain a chart animation.
[0326] Specifically, according to relevant research on chart animation, chart animation can improve the readability and comprehensibility of data graphics, because dynamic transformation can make data changes more intuitive and easy to understand, and can help people better discover patterns and trends in the data. Chart animation can effectively improve the attractiveness of data visualization, because dynamic transformation can bring dynamic and vivid elements to data graphics, making them more attractive.
[0327] The automatic generation of animation in the embodiment of the present application includes four design principles: (1) Staging: avoid changing too much content at one time, and try to change only one type of information in the chart at a time. For example, the animation conversion between a scatter plot and a bar chart can be completed in three stages. (2) Compatibility. Visual charts may be difficult for readers to recognize due to the animation effects. It is necessary to ensure the compatibility of the animation effects and charts. If inappropriate animation effects are used on charts, it will only have a counterproductive effect. (3) Necessity: It is necessary to avoid unnecessary animation effects and only change the places that need to be changed. Those redundant animation effects will not only distract people, but may also confuse the audience. (4) Meaningfulness: Ensure that the charts during the animation process are meaningful, that is, it is necessary to ensure that the visual encoding of the charts is effective. Based on animation research, the embodiment of the present application uses staged animation to gradually present the chart animation transition.
[0328] In order to better highlight the details of data changes, further analysis is needed on the details of the animation. A better way to highlight the details of the animation is to split the transition between two time slices (i.e., visual charts) into multiple steps, which is the concept of staged animation.
[0329] In order to achieve a smoother animation transition process as much as possible, the embodiment of the present application encapsulates an animation automatic generator AnimationManager on the basis of the visualization Vizzu library, mainly to increase the support for animation duration configuration, and to support the implementation of staged animations for complex visualizations. It is mainly based on the sequence optimization module SequenceManager, the output sequence result JSON object (JavaScript Object Notation, JavaScript object representation), extracts the corresponding Transitions atomic operation, and guides AnimationManager to build and generate transition conversions. AnimationSpliter (animation filter) is required to traverse all transition operation groups between each group of visualizations, then identify the transition category, extract the visual channel channel and mapping rules, and then call the AnimationGenrator class to generate an animation driver, and allocate a suitable duration to separate the start time of each different transition operation, otherwise transform in parallel, and finally trigger AnimationExector (a component for processing animation-related tasks) to execute the animation.
[0330] Step S308, performing video creation processing on the chart animation to obtain a visual video for displaying data facts.
[0331] Specifically, the business data story has been formed into a chart animation through the aforementioned steps and is presented in the view. Before the video synthesis module exports the video, the duration of the animation clip is automatically configured to ensure that the optimal video length is obtained. There are three issues to consider in the duration configuration: First, because too long a presentation time will make the audience lose interest, the data video length should be moderate. Second, the animation should not be too fast. Each animation transition should last about one second so that people can perceive the visual changes. Third, people usually want to pay more attention to important facts when watching videos, so important (reflected in high importance scores) data facts should have longer time.
[0332] The recording of the animation is implemented using the MediaStream recorder API, which is an application programming interface (API) for recording audio and video streams in web applications. However, the video recorded by the MediaStream Recorder API is in the WEBM format (an open source media container format). In addition to the automatic generation of the video, the data analysis platform or data analysis system also supports users to upload audio and switch the video format. Therefore, more customized development such as video transcoding is required. The embodiment of this application decides to use FFmpeg.wasm to implement video transcoding and audio synthesis functions.
[0333] FFmpeg.wasm is a project that transplants the functions of FFmpeg (a very powerful open source multimedia processing tool, widely used for recording, converting, streaming and editing audio and video) to WebAssembly (a low-level, platform-independent bytecode format that runs in the sandbox environment of the browser. It allows developers to compile code written in languages such as C, C++, Rust into WebAssembly bytecode, so that it can run in the browser at a speed close to that of native code) and JavaScript (a dynamic scripting language that runs in the browser, mainly used to implement the interactive functions of web pages), allowing audio and video processing directly in the browser environment without installing additional plug-ins or software. It is based on WebAssembly (Wasm) technology, compiling the core functions of FFmpeg into WebAssembly bytecode, thereby achieving efficient audio and video processing in the browser.
[0334] Please also see Fig.12 , Fig.12 FIG. 1 is a schematic diagram of a video synthesis scene provided by an embodiment of the present application. Fig.12As shown, through the application interface 80d of the browser (e.g., MediaRecorder API), the computer device captures a media stream, which may include a video clip 80b, or a video clip 80b (i.e., a chart animation) and an audio 80c. The computer device divides the chart animation into multiple data chunks, which are usually Blob objects representing binary data.
[0335] Each chunk (data block) generated by the MediaRecorder API is collected into an array allChunks, which contains all the data blocks generated during recording. All the collected data blocks allChunks are passed to the open source multimedia processing tool FFmpeg.wasm. FFmpeg.wasm is a WebAssembly module that provides FFmpeg's audio and video processing functions and can be run directly in the browser to process these data blocks. After FFmpeg.wasm processes all the data blocks, it generates a new Blob object that contains the entire processed audio and video file. Fig.12 As shown, this Blob object can be downloaded directly in the browser or used for other purposes. Users can download the processed audio and video files to their local computer.
[0336] This application embodiment proposes a visual video presentation solution based on business data. Through an automated business data visualization video generation system, it realizes automatic extraction of data facts, sequence optimization, animation mapping and arrangement, and video synthesis. It innovatively proposes an automated workflow to improve the efficiency and attractiveness of business data visualization. Please also refer to Fig.13 , Fig.13 Schematic diagram of a visual video presentation method based on business data provided by an embodiment of the present application. Fig.13 As shown, the automatic insight module, with business data as input, can automatically mine interesting data patterns and data features from the business data, rank, recommend and visualize the most important insights (perspectives), and simplify the tedious and repetitive exploratory analysis process.
[0337] Extracting interesting trends and data patterns from data is the first step in creating data stories. The main task of automatic insight in the embodiment of the present application is to use machines to use basic statistical indicators (such as mean, median, standard deviation, etc.) to describe the characteristics of business data. For example, correlation analysis, time series analysis, etc. are used to automatically explore data and discover different types of interesting business data facts.
[0338] Automatic insights have the following advantages: (1) Efficiency: Automatic insights can quickly analyze large amounts of data, greatly improving analysis efficiency. (2) Accuracy: Automatic insights use technologies such as machine learning and artificial intelligence to analyze data more accurately and avoid errors in human analysis. (3) Real-time: Automatic insights can analyze data in real time, discover data changes and trends in a timely manner, and make decisions more timely and accurate. It can help people better understand data and make decisions faster and more accurately. Therefore, this system implements an automatic insight generator that can automatically extract meaningful data facts, so that ordinary users can also quickly grasp the characteristics of the data set and create data stories more quickly.
[0339] The automatic insight module mainly includes data fact definition and modeling, which can be seen above. Figure 2 Step S102 in the corresponding embodiment also includes data fact significance calculation analysis, which can be the scenario above. Figure 2 In the corresponding step S104 of the embodiment, the importance of data facts can be further calculated by calculating the significance of facts.
[0340] The automatic insight module outputs visualization charts to the sequence optimization module. The order in which data visualization is presented to the target audience may affect understanding and memory. The optimal sequence can reduce cognitive load and promote better understanding and recall of the valuable information of the visual data story, while a less than ideal story sequence hinders user understanding and recall. Therefore, it is necessary to introduce an effective sequence optimizer in data story video generation. In this application, the Graphscape model can be used to recommend the most optimized visualization sequence through conversion cost calculation and filtering cost calculation.
[0341] The Graphscape model represents visualizations as graph nodes and transitions between nodes as directed edges. The model proposes a cost estimate of the difficulty of interpreting a conversion to a target visualization given a source visualization, and the conversion cost of a visualization directed graph can be derived by sorting editing operations such as conversions between tag types, data conversions, and encodings between visualization pairs.
[0342] In this application, the data video generation system implements the recommendation of the best visualization sequence through the Graphscape model, and also provides useful guidance and inspiration for creators to design more effective data story sequences from a psychological perspective, making the generated data stories easier for the audience to understand and remember. Therefore, it is proposed to use the Graphscape model to recommend the most optimized visualization sequence to users, and consider the consistency of both perception and content. By finding the sequence with the lowest cost, a smooth sequence that can be used for visual synthesis is obtained.
[0343] The sequence optimization module outputs the chart optimization sequence to the animation mapping module. The animation mapping module outputs one or more chart animations. A chart animation can be understood as a video clip. The dimensions included in a chart animation are the same, so the sequence cost corresponding to a chart animation is low.
[0344] The video synthesis module synthesizes the chart animation output by the animation mapping module, that is, creates a video, obtains a complete initial video, and outputs the initial video to the audio synthesis module. The audio synthesis module synthesizes the initial video and audio file and outputs the video.
[0345] This application can also be combined with the following technologies to achieve richer applications:
[0346] 1. Interactive and immersive data visualization technology:
[0347] Using AR (Augmented Reality) / VR (Virtual Reality) technology, data is placed in three-dimensional space, providing a more intuitive way to understand data through immersive interaction.
[0348] Implementing classic visualization charts in VR allows users to directly experience the true scale of the data.
[0349] 2. Cross-integration of information visualization and pedagogy:
[0350] Drawing the statistical analysis methods in R language into comics makes learning R language more interesting. R language is a programming language and software environment for statistical computing and graphical representation.
[0351] Design comic tutorials to help people better learn complex charts, such as parallel coordinates, box-and-whisker plots, etc.
[0352] 3. Application of Artificial Intelligence (AI):
[0353] Use AI technology to efficiently process large amounts of business data and generate accurate and personalized visualization results.
[0354] From the above, we can see that efficient visualization of business data through automated work pipelines not only improves the accuracy of data understanding and decision-making, enhances the attractiveness and readability of data visualization, but also lowers the technical threshold, allowing ordinary users to quickly grasp the characteristics of data sets and create data stories. At the same time, it optimizes the data display sequence, reduces cognitive load, promotes better understanding and memory of information, and ultimately improves the efficiency and quality of business data video production.
[0355] For further information, see Fig.14 , Fig.14 1 is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. The above-mentioned data processing device 1 can be used to execute the corresponding steps in the method provided in an embodiment of the present application. Fig.14 As shown, the data processing device 1 may include:
[0356] An acquisition module 11 is used to acquire a business data set including data corresponding to d dimensions and data corresponding to m metrics; d and m are both positive integers;
[0357] The processing module 12 is used to perform data subject enumeration processing on the business data set to obtain a data subject set; the dimensions included in each data subject in the data subject set belong to d dimensions; the metrics included in each data subject in the data subject set belong to m metrics;
[0358] The processing module 12 is further used to match each data subject in the data subject set with the full amount of fact types to obtain a first data subject and a first fact type that are successfully matched; the first data subject belongs to the data subject set; and the first fact type belongs to the full amount of fact types;
[0359] An extraction module 13, configured to extract facts from the first data subject according to the first fact type, and obtain data facts corresponding to the business data set;
[0360] The processing module 12 is further used to perform visualization processing on the data fact according to a chart type adapted to the first fact type, so as to obtain a visualization chart for displaying the data fact.
[0361] The specific description of the data processing device 1 can be found in the above-mentioned data processing method, which will not be repeated here.
[0362] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0363] It can be seen from the above that the embodiments of the present application can not only improve the completeness and accuracy of data fact extraction, but also enhance the readability of charts after data fact visualization.
[0364] For further information, see Fig.15 , Fig.15is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device may be Figure 1 The terminal device or service server shown in FIG. Fig.15 As shown, the computer device 1000 may include: at least one processor 1001, such as a CPU, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002. The communication bus 1002 is used to realize connection and communication between these components.
[0365] In some embodiments, the user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and the network interface 1004 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 1005 may also be at least one storage device located away from the aforementioned processor 1001.
[0366] like Fig.15 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a device control application.
[0367] exist Fig.15 In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to implement the method steps in each embodiment of the present application.
[0368] It should be understood that the computer device 1000 described in the embodiment of the present application can execute the description of the data processing method or device in the above embodiments, which will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated.
[0369] The present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the description of the data processing method or device in the above embodiments is implemented, which will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated.
[0370] The computer-readable storage medium may be the data processing device provided in any of the aforementioned embodiments or the internal storage unit of the computer device, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., provided on the computer device.
[0371] Furthermore, the computer-readable storage medium may also include both an internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0372] The present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device can execute the description of the data processing method or device in the above embodiments, which will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated.
[0373] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, device, product, or equipment that includes a series of steps or units is not limited to the listed steps or modules, but optionally includes steps or modules that are not listed, or optionally includes other step units inherent to these processes, methods, devices, products, or equipment.
[0374] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
Claims
1. A data processing method, characterized in that: include: Obtain a business data set including data corresponding to d dimensions and data corresponding to m metrics; d and m are both positive integers; Performing data topic enumeration processing on the business data set to obtain a data topic set; the dimensions included in each data topic in the data topic set belong to the d dimensions; the metrics included in each data topic in the data topic set belong to the m metrics; Matching each data subject in the data subject set with the full set of fact types to obtain a first data subject and a first fact type that are successfully matched; the first data subject belongs to the data subject set; and the first fact type belongs to the full set of fact types; According to the first fact type, extract facts from the first data subject to obtain data facts corresponding to the business data set; The data fact is visualized according to a chart type adapted to the first fact type to obtain a visualization chart for displaying the data fact.
2. The method according to claim 1, characterized in that The data subject enumeration process is performed on the business data set to obtain a data subject set, including: Subspace enumeration is performed on the business data set to obtain a subspace set; the subspace set includes a first subspace; the first subspace includes data on b dimensions in the business data set; b is a positive integer, and b is less than or equal to d; the b dimensions include dimension P q , q is a positive integer, and q is less than or equal to b; If the first subspace includes the dimension P q The full amount of data on the q Determining a segmentation dimension of the first subspace; Performing data topic enumeration processing on the first subspace, the segmentation dimension, and the m metrics to obtain data topics corresponding to the first subspace; the data topics corresponding to the first subspace include the first subspace and the segmentation dimension; The data topics corresponding to the first subspace and the data topics corresponding to the second subspace are combined into a data topic set; the second subspace refers to a subspace in the subspace set except the first subspace.
3. The method according to claim 1, characterized in that The matching process of each data subject in the data subject set with the full set of fact types to obtain a first data subject and a first fact type that are successfully matched includes: Performing influence analysis on each data topic in the data topic set to obtain an influence value corresponding to each data topic in the data topic set; Comparing the influence value corresponding to each data topic in the data topic set with the influence threshold; The data topics in the data topic set whose influence values are greater than or equal to the influence threshold are matched with all fact types to obtain a first data topic and a first fact type that are successfully matched.
4. The method according to claim 3, characterized in that The data topics in the set of data topics include a second data topic; the second data topic includes a first metric; The performing influence analysis on each data subject in the data subject set to obtain an influence value corresponding to each data subject in the data subject set includes: According to the segmentation dimension in the second data subject, the subspace in the second data subject is segmented to obtain a plurality of brother subspaces corresponding to the second data subject; the data of the plurality of brother subspaces on the segmentation dimension are different from each other, and the data of the plurality of brother subspaces on dimensions other than the segmentation dimension are the same; In the business data set, full data of the first metric is obtained, and influence analysis is performed on the multiple brother subspaces respectively according to the full data of the first metric to obtain influence values corresponding to the multiple brother subspaces respectively; The influence values corresponding to the multiple brother subspaces are averaged to obtain the influence value corresponding to the second data topic.
5. The method according to claim 4, characterized in that The multiple brother subspaces include brother subspace A c , c is a positive integer, and c is less than or equal to the total number of the plurality of brother subspaces; The performing influence analysis on the multiple brother subspaces respectively according to the full amount of data of the first metric to obtain influence values corresponding to the multiple brother subspaces respectively includes: Summing all data of the first metric to obtain total data corresponding to the first metric; In the full amount of data of the first metric, obtain the brother subspace A c data on the first metric; For the sibling subspace A c The data on the first metric and the total data are processed by ratio to obtain the brother subspace A c The corresponding influence value.
6. The method according to claim 5, characterized in that The number of the first metrics is at least two; the at least two first metrics include the second metric; the total data includes the total data corresponding to the second metric; The pair of sibling subspaces A c The data on the first metric and the total data are processed by ratio to obtain the brother subspace A c The corresponding influence values include: For the sibling subspace A c The data on the second metric and the total data corresponding to the second metric are processed by ratio to obtain the brother subspace A c The influence value on the second metric; Get the sibling subspace A c The influence values corresponding to the at least two first metrics are weighted averaged to obtain the sibling subspace A. c The corresponding influence value.
7. The method according to claim 1, characterized in that The extracting facts from the first data subject according to the first fact type to obtain data facts corresponding to the business data set includes: Performing importance analysis on the first data topic to obtain an importance score corresponding to the first data topic; If the importance score corresponding to the first data topic is equal to or greater than the importance score threshold, facts are extracted from the first data topic according to the pattern verification strategy corresponding to the first fact type to obtain data facts corresponding to the business data set.
8. The method according to claim 7, characterized in that The number of the first data topics is at least two; the at least two first data topics include data topic G h and data subject G i , h and i are both positive integers, and h and i are both less than or equal to the total number of the at least two first data topics, and h is different from i; The performing importance analysis on the first data subject to obtain an importance score corresponding to the first data subject includes: If the data subject G h The influence value is greater than or equal to the data subject G i The influence value of the first importance assessment task is determined to have a higher execution priority than the second importance assessment task; the first importance assessment task is used to complete the data subject G h The second importance assessment task is used to complete the data subject G i Importance analysis; Adding the first importance assessment task and the second importance assessment task to a task queue in descending order of execution priority; A third importance assessment task is obtained from the task queue in descending order of execution priority; the third importance assessment task is used to complete the data topic G of the at least two first data topics. j The importance analysis of ; j is a positive integer, and j is less than or equal to the total number of the at least two first data topics; The third importance assessment task is performed in parallel to obtain the data topic G j The corresponding importance score.
9. The method according to claim 7, characterized in that: The performing importance analysis on the first data subject to obtain an importance score corresponding to the first data subject includes: According to the segmentation dimension in the first data subject, segment the subspace in the first data subject to obtain a plurality of brother subspaces corresponding to the first data subject; In the business data set, obtain the full amount of data of the metric z in the first data subject; the metric z belongs to the m metrics; According to the full data of the metric z and the multiple sibling subspaces, an importance analysis is performed on the first data topic to obtain an importance score corresponding to the first data topic.
10. The method according to claim 9, characterized in that The performing importance analysis on the first data topic according to the full amount of data of the metric z and the multiple sibling subspaces to obtain the importance score corresponding to the first data topic includes: From the full amount of data of the metric z, data corresponding to the multiple sibling subspaces on the metric z are obtained, and the obtained multiple data are determined as a result set; Performing a significance test on the result set to obtain a significance value corresponding to the first data topic; The significance value corresponding to the first data topic and the influence value corresponding to the first data topic are multiplied to obtain the importance score corresponding to the first data topic.
11. The method according to claim 10, characterized in that The performing a significance test on the result set to obtain a significance value corresponding to the first data topic includes: Performing a point significance test on the result set to obtain a point significance value, and performing a slope significance test on the result set to obtain a slope significance value; If the point significance value is greater than or equal to the slope significance value, the point significance value is determined as the significance value corresponding to the first data theme; If the point significance value is less than the slope significance value, the slope significance value is determined as the significance value corresponding to the first data theme.
12. The method according to claim 1, characterized in that After obtaining the visualization chart, the method further includes: If the total number of the visualization charts is at least two, sorting the at least two visualization charts to obtain a chart optimization sequence of the visualization charts; Performing animation mapping processing on each visualization chart in the chart optimization sequence to obtain a chart animation; The chart animation is processed into a video to obtain a visual video for displaying the data facts.
13. The method according to claim 12, characterized in that The step of sorting at least two visual charts to obtain a chart optimization sequence of the visual charts includes: According to the total number of at least two visualization charts, determine the sorting number X; X is a positive integer greater than 1; Sorting the at least two visual charts X times to obtain X chart sequences; the X chart sequences all include the at least two visual charts, and the chart sequences corresponding to the X chart sequences are all different; the X chart sequences include chart sequence K n , n is a positive integer, and n is less than or equal to X; For the graph sequence K n Perform cost analysis and obtain the chart sequence K n The corresponding sequence cost; Determine the sequence costs corresponding to the X chart sequences respectively, and obtain the minimum sequence cost from the X sequence costs; The chart sequence corresponding to the minimum sequence cost among the X chart sequences is determined as a chart optimization sequence.
14. The method according to claim 13, characterized in that The graph sequence K n Perform cost analysis and obtain the chart sequence K n The corresponding sequence costs include: For the graph sequence K n Perform chart conversion cost analysis to obtain the chart sequence K n The corresponding chart conversion cost; For the graph sequence K n Perform graph filtering cost analysis to obtain the graph sequence K n The corresponding chart filtering cost; The chart conversion cost and the chart filtering cost are summed to obtain the chart sequence K n The corresponding sequence cost.
15. The method according to claim 14, characterized in that The graph sequence K n Perform chart conversion cost analysis to obtain the chart sequence K n The corresponding chart conversion costs include: Get the chart sequence K n Every two adjacent visualization charts in n Each two adjacent visualization charts in the method include an adjacent first visualization chart and a second visualization chart; If a conversion relationship exists between the first visualization chart and the second visualization chart, a conversion cost corresponding to the conversion relationship is determined as a chart conversion cost of the first visualization chart and the second visualization chart; If there is no conversion relationship between the first visualization chart and the second visualization chart, determining a chart path for connecting the first visualization chart and the second visualization chart, wherein every two adjacent visualization charts in the chart path have a conversion relationship; Determining chart conversion costs of the first visual chart and the second visual chart according to conversion costs corresponding to conversion relationships in the chart path; For the graph sequence K n The chart conversion costs corresponding to each two adjacent visualization charts in the summation process are processed to obtain the chart sequence K n Corresponding chart conversion costs.
16. The method according to claim 14, characterized in that The graph sequence K n Perform graph filtering cost analysis to obtain the graph sequence K n The corresponding chart filtering costs include: Get the chart sequence K n Every two adjacent visualization charts in n Each two adjacent visualization charts in the method include an adjacent first visualization chart and a second visualization chart; If the dimension displayed by the first visualization chart is the same as the dimension displayed by the second visualization chart, determining the first value as a chart distance between the first visualization chart and the second visualization chart; If the dimension displayed by the first visualization chart is different from the dimension displayed by the second visualization chart, the second value is determined as the chart distance between the first visualization chart and the second visualization chart; the first value is greater than the second value; Determine the chart sequence K n The chart distance corresponding to each two adjacent visualization charts in ; According to the chart sequence K n The chart sequence K is determined by the chart distance corresponding to each two adjacent visualization charts in n The corresponding chart filtering cost.
17. A data processing device, characterized in that: include: An acquisition module, used to acquire a business data set including data corresponding to d dimensions and data corresponding to m metrics; d and m are both positive integers; a processing module, configured to perform data topic enumeration processing on the business data set to obtain a data topic set; the dimensions included in each data topic in the data topic set all belong to the d dimensions; and the metrics included in each data topic in the data topic set all belong to the m metrics; The processing module is further used to match each data subject in the data subject set with the full amount of fact types to obtain a first data subject and a first fact type that are successfully matched; the first data subject belongs to the data subject set; the first fact type belongs to the full amount of fact types; an extraction module, configured to extract facts from the first data subject according to the first fact type, to obtain data facts corresponding to the business data set; The processing module is further used to perform visualization processing on the data fact according to a chart type adapted to the first fact type, so as to obtain a visualization chart for displaying the data fact.
18. A computer device, characterized in that: include: Processor, memory, and network interface; The processor is connected to the memory and the network interface, wherein the network interface is used to provide a data communication function, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method described in any one of claims 1-16.
19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 16.
20. A computer program product, characterized in that The computer program product comprises a computer program, which is stored in a computer-readable storage medium. The computer program is suitable for being read and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 16.