A data analysis method
By acquiring pattern recognition and metadata analysis from heterogeneous data sources, a mapping table between mind map nodes and data fields is generated, solving the problem of the disconnect between mind mapping tools and computation in existing technologies. This achieves intelligent and reliable data analysis, and improves data modeling efficiency and interpretability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FUTURE TV CO LTD
- Filing Date
- 2026-03-20
- Publication Date
- 2026-06-16
AI Technical Summary
In existing data analysis technologies, mind mapping tools are disconnected from the computational process, relationship discovery lacks a multi-evidence fusion mechanism, layout updates are not stable enough, analysis results have weak interpretability, and report generation relies on manual operation, resulting in low levels of intelligence, difficulty in cross-tool collaboration, and difficulty in ensuring the credibility and security of data analysis work.
By acquiring multiple heterogeneous data sources selected by the user, performing pattern recognition and metadata analysis, a mapping table between mind map nodes and data fields is generated. Combining the data analysis results and the preset mind map structure, the optimal data fields corresponding to each node are automatically matched and bound to generate a complete mapping table containing the business entity context, realizing intelligent mapping from raw data to knowledge structure.
It improves the efficiency and accuracy of data modeling, reduces human intervention, enhances the intelligence and credibility of data analysis, improves the interpretability and interactivity of the system, and avoids analytical bias.
Smart Images

Figure CN122220784A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and more specifically, to a data analysis method. Background Technology
[0002] In a data-driven business environment, enterprises are facing increasingly complex analytical needs. Business decisions require the integration of multi-source, heterogeneous data from multiple dimensions, including market, operations, and risk. This not only demands the rapid identification of key influencing factors and the uncovering of potential development patterns, but also necessitates support for dynamic adjustments to decisions and cross-departmental collaborative analysis, thereby improving the scientific rigor and timeliness of decision-making.
[0003] Currently, data analysis technologies primarily utilize visualization tools that generate charts by dragging and dropping tables, mind mapping software that organizes knowledge using nodes and boundaries, data mining platforms that arrange analysis steps using flowcharts, and graph tools that conduct community and path analysis based on graph structures for data processing.
[0004] However, existing technologies suffer from a disconnect between mind mapping tools and computational processes. Relationship discovery lacks multi-evidence fusion mechanisms, layout updates are unstable, and analysis results have weak interpretability. Report generation relies on manual operation, and data governance capabilities are fragmented and disjointed. This results in low levels of intelligence in existing systems, difficulties in cross-tool collaboration, and challenges in ensuring the credibility and security of data analysis. Summary of the Invention
[0005] The purpose of this application is to provide a data analysis method to address the shortcomings of the prior art, thereby improving the intelligence and reliability of data processing.
[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows: In a first aspect, one embodiment of this application provides a data analysis method, the method comprising: Retrieve the target data source selected by the user from multiple heterogeneous data sources; Pattern recognition is performed based on the metadata of the target data source to obtain pattern description information in the target data source. The pattern description information is used to indicate the structure of each data object in the target data source and the relationship between the data objects. Based on the pattern description information, obtain multiple data fields from the target data source; Based on the pattern description information and the information of the multiple data fields, obtain the data profiling results of the target data source; Based on the data analysis results and the preset mind map structure, the data fields corresponding to each mind map node in the preset mind map structure are determined from the multiple data fields; Based on each mind map node, its corresponding data field, and the business entity of the corresponding data field, a mapping table is generated for each mind map node and its data field, wherein the business entity of the corresponding data field is used to indicate the data source of the corresponding data field.
[0007] Optionally, the method further includes: Based on the data values of each data field in the target data source, obtain the field profile of each data field; Based on the field profiles of each data field, a quality assessment is performed on each data field to obtain a quality label for each data field.
[0008] Optionally, the method further includes: Displays each mind map node, its corresponding data field, the matching score of each mind map node and its corresponding data field, and the quality label of the corresponding data field.
[0009] Optionally, the method further includes: Based on the data values of each mind map node in the mapping table and the corresponding data field, calculate the correlation parameter of every two nodes among all unconnected nodes in the preset mind map structure; Based on the business entities of the data fields corresponding to each mind map node in the mapping table, calculate the potential association parameters between every two nodes; Based on the correlation parameters between every two nodes and the potential association parameters between every two nodes, candidate relationships of the preset mind map structure and the confidence level of each candidate relationship are obtained. A candidate relation set is generated based on the candidate relations and the confidence level of each candidate relation.
[0010] Optionally, the step of calculating the correlation parameter between every two nodes among all unconnected nodes in the preset mind map structure based on the data values of each mind map node in the mapping table and the corresponding data field includes: If the data field values of the first and second nodes among all the unconnected nodes are numerical or discrete, then a statistical algorithm is used to obtain a correlation score. If the data values of the corresponding data fields of the third and fourth nodes among all unconnected nodes are time-series data, a causal structure algorithm is used to obtain a causal score, wherein the correlation parameter includes the correlation score and the causal score.
[0011] Optionally, the method further includes: Based on the candidate set of relationships, an incremental relationship suggestion list for the preset mind map structure is generated. The incremental relationship suggestion list includes: the target variable relationship to be added, the two mind map nodes corresponding to the target edge relationship, and the direction of the target edge relationship.
[0012] Optionally, the method further includes: Based on the candidate relation set, the candidate relations are sequentially added to the preset mind map structure to obtain the updated mind map structure. The updated mind map structure is displayed, and the candidate relationships are distinguished within the updated mind map structure.
[0013] Optionally, the method further includes: After the user confirms the updated mind map structure, the operator template corresponding to each mind map node in the updated mind map structure is matched according to the updated mind map structure. The execution order of each operator template is determined based on the relationship between the nodes in the updated mind map structure. Based on the execution order of the various operator templates, a directed acyclic graph of the updated mind map structure is generated. Each node in the directed acyclic graph corresponds to an operator template, and the edges between nodes are used to represent the data dependencies between operator templates.
[0014] Optionally, the method further includes: The operator templates are executed sequentially according to the directed acyclic graph to obtain the intermediate and final results of each brain graph node in the updated brain graph structure. The intermediate results and the final result of each brain map node in the updated brain map structure are displayed.
[0015] Optionally, the method further includes: The intermediate and final results of each brain map node in the updated brain map structure are determined by using a preset feature algorithm to determine the feature contribution of each brain map node in the updated brain map structure. Based on the feature contributions of each brain map node in the updated brain map structure, a data feature evidence chain structure is generated, wherein the data feature evidence chain structure is used to indicate the path of each brain map node in the updated brain map structure. Based on the intermediate results and final results of each mind map node in the updated mind map structure, as well as the data feature evidence chain structure, a data analysis report is generated using a preset report template.
[0016] Secondly, another embodiment of this application provides a data analysis apparatus, the apparatus comprising: The first acquisition module is used to acquire the target data source selected by the user from multiple heterogeneous data sources; The first determining module is used to perform pattern recognition based on the metadata of the target data source to obtain pattern description information in the target data source. The pattern description information is used to indicate the structure of each data object in the target data source and the relationship between the data objects. The second acquisition module is used to acquire multiple data fields from the target data source based on the pattern description information; The third acquisition module is used to acquire the data analysis results of the target data source based on the pattern description information and the information of the multiple data fields; The second determining module is used to determine the data field corresponding to each mind map node in the preset mind map structure from the multiple data fields based on the data analysis results and the preset mind map structure. The generation module is used to generate a mapping table between each mind map node and the corresponding data field based on each mind map node, the corresponding data field of each mind map node, and the business entity of the corresponding data field, wherein the business entity of the corresponding data field is used to indicate the data source of the corresponding data field.
[0017] Thirdly, another embodiment of this application provides a computer device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the steps of any of the data analysis methods described in the first aspect above.
[0018] Fourthly, another embodiment of this application provides a storage medium storing a computer program, which, when executed by a processor, performs the steps of any of the data analysis methods described in the first aspect above.
[0019] The beneficial effects of this application are: This application provides a data analysis method that obtains a target data source selected by the user from multiple heterogeneous data sources; performs pattern recognition based on the metadata of the target data source to obtain pattern description information in the target data source, which indicates the structure of each data object in the target data source and the relationships between data objects; obtains multiple data fields in the target data source based on the pattern description information; obtains data profiling results of the target data source based on the pattern description information and the information of multiple data fields; determines the data fields corresponding to each mind map node in the preset mind map structure from multiple data fields based on the data profiling results and the preset mind map structure; and generates a mapping table of each mind map node and data field based on each mind map node, the corresponding data field of each mind map node, and the business entity of the corresponding data field. This application achieves intelligent mapping from raw data to knowledge structure by integrating the metadata of multiple heterogeneous data sources. It can accurately extract the structure and relationships of data objects based on pattern recognition, and also performs semantic understanding and quality assessment of fields based on data profiling results. Then, it automatically matches and binds the optimal data fields corresponding to each node according to the preset mind map structure, and finally generates a complete mapping table containing the context of business entities. The whole process reduces manual intervention and improves the efficiency and accuracy of data modeling. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A flowchart illustrating a data analysis method provided in an embodiment of this application; Figure 2 A flowchart illustrating the process of determining field quality labels in a data analysis method provided in this application embodiment; Figure 3 A flowchart illustrating the process of determining a candidate set of relations in a data analysis method provided in this application embodiment; Figure 4 A schematic diagram illustrating the process of determining the correlation parameters between every two nodes in data analysis, provided as an embodiment of this application; Figure 5 A flowchart illustrating the display of candidate relationships in a data analysis method provided in this application embodiment; Figure 6 A schematic diagram of the process for generating a directed acyclic graph in a data analysis method provided in an embodiment of this application; Figure 7A flowchart illustrating the display of results in a data analysis method provided in an embodiment of this application; Figure 8 This application provides a schematic diagram of the process for generating a data analysis report in a data analysis method. Figure 9 This is a schematic diagram of the structure of a data analysis device provided in an embodiment of this application; Figure 10 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0023] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0024] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0025] To clearly describe the method provided in the embodiments of this application, the data analysis method will be described below in conjunction with several accompanying drawings. Figure 1 This is a flowchart illustrating a data analysis method provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes: Step 101: Obtain the target data source selected by the user from multiple heterogeneous data sources.
[0026] Heterogeneous data sources include relational database tables, semi-structured data, graph databases, and file-based data sources, etc., and this application embodiment does not limit them. The target data source is the data source that the user determines needs to be processed.
[0027] Optionally, a target data source can be selected from multiple heterogeneous data sources based on the user's chosen target data source. Users can select this source through a selection interface or by entering the corresponding target data source information.
[0028] Step 102: Perform pattern recognition based on the metadata of the target data source to obtain pattern description information from the target data source. Schema description information is used to indicate the structure of each data object in the target data source and the relationships between them. Schema description information refers to the collection used in the database to define and organize data structures. It describes the structure of various tables, fields, constraints, indexes, and other objects in the database, as well as the relationships between them, enabling data to be stored, queried, and managed effectively.
[0029] Optionally, pattern recognition is performed based on the field names, types, and constraints of the target data source's metadata to obtain pattern description information from the target data source. The field name represents the name of a column in the data table, used to identify the type of information stored in that column. The type refers to the types and formats of data that the field can store. Constraints are the rules restricting the field. Pattern description information includes primary keys, foreign keys, time fields, enumeration fields, etc., but this embodiment does not impose any limitations on these.
[0030] Step 103: Based on the schema description information, obtain multiple data fields from the target data source.
[0031] Optionally, the schema description information can be traversed to obtain multiple data fields from the target data source.
[0032] Step 104: Based on the pattern description information and information from multiple data fields, obtain the data profiling results of the target data source.
[0033] The data analysis results include: pattern description information for each data field, field profile for each data field, and data quality for each data field.
[0034] Optionally, data profiling results from the target data source can be obtained based on the pattern description information and sample data from multiple data fields.
[0035] Step 105: Based on the data analysis results and the preset mind map structure, determine the data fields corresponding to each mind map node in the preset mind map structure from multiple data fields.
[0036] The preset mind map structure specifies the displayed themes, sub-themes, and their hierarchical relationships. Mind map nodes are multiple nodes within the preset mind map structure, determined by the user's definition and specifically based on the data analysis scenario. For example, in a shopping scenario, mind map nodes could be "customer," "order," "product," etc.
[0037] Optionally, based on the data analysis results and the preset mind map structure, a preset business dictionary / thesaurus is used to match multiple data fields with each mind map node in the preset mind map structure to determine the data fields corresponding to each mind map node in the preset mind map structure. For example, fields containing strings such as "cust", "customer", and "user" are matched to the "customer" node type.
[0038] Step 106: Generate a mapping table for each mind map node and its corresponding data field based on each mind map node, the corresponding data field, and the business entity of the corresponding data field.
[0039] The business entity corresponding to the data field indicates the data source of that data field. The mapping table is a ternary mapping table, which represents the correspondence between mind map nodes, their corresponding data fields, and the business entities associated with those data fields. A mind map node can correspond to one or more data fields.
[0040] Optionally, the corresponding data fields of each mind map node are determined based on each mind map node, and the business entities of the corresponding data fields are determined based on the corresponding data fields, thereby generating a ternary mapping table of each mind map node and data field, and displaying the ternary mapping table and the preset mind map structure.
[0041] In this embodiment, a target data source selected by the user from multiple heterogeneous data sources is obtained; pattern recognition is performed based on the metadata of the target data source to obtain pattern description information in the target data source, which indicates the structure of each data object in the target data source and the relationships between data objects; multiple data fields in the target data source are obtained based on the pattern description information; data profiling results of the target data source are obtained based on the pattern description information and the information of multiple data fields; based on the data profiling results and a preset mind map structure, the data fields corresponding to each mind map node in the preset mind map structure are determined from the multiple data fields; a mapping table of each mind map node and data field is generated based on each mind map node, the corresponding data field of each mind map node, and the business entity of the corresponding data field. This application integrates the metadata of multiple heterogeneous data sources to achieve intelligent mapping from raw data to knowledge structure. It can accurately extract the structure and relationships of data objects based on pattern recognition, and also performs semantic understanding and quality assessment of fields based on data profiling results. Then, it automatically matches and binds the optimal data fields corresponding to each node according to the preset mind map structure, and finally generates a complete mapping table containing the context of business entities. The whole process reduces manual intervention and improves the efficiency and accuracy of data modeling.
[0042] Based on the above embodiments, this application also provides a process for determining field quality labels in a data analysis method. Figure 2 This is a flowchart illustrating the process of determining field quality labels in a data analysis method provided in an embodiment of this application, as shown below. Figure 2 As shown, based on steps 101-105 above, the method further includes: Step 201: Obtain the field profile of each data field based on the data values of each data field in the target data source.
[0043] The data values of each data field are the specific data content contained in the actual storage of the data field. Field profiling is used to collect structured information and characterize the features of a field.
[0044] Optionally, by creating a profile of each data field based on indicators such as the distribution of values, missing rate, proportion of outliers, and number of unique values in each data field of the target data source, a field profile can be obtained for each data field.
[0045] Step 202: Based on the field profile of each data field, perform a quality assessment on each data field to obtain the quality label of each data field.
[0046] The quality labels include: high quality, missing or suspected identification fields, etc.
[0047] Optionally, based on the field profile of each data field, a preset rule template is used to perform a quality assessment on each data field, obtaining a quality label for each data field. The field profile and quality label for each data field are then output. For example, when the data fields are ID card number, mobile phone number, and timestamp, the corresponding preset rule template can be a regular expression rule.
[0048] In this embodiment of the application, by constructing field profiles based on the actual data values of each data field in the target data source, and performing rule judgments based on the field profiles, the quality assessment and labeling output of each data field can be automated, which can comprehensively and objectively reflect the data quality status, reduce the cost of manual review, and improve the efficiency of data credibility identification.
[0049] Based on the above embodiments, this application also provides a process for displaying quality labels in a data analysis method. In addition to steps 201-202 above, the method further includes: Displays each mind map node, its corresponding data field, the matching score between each mind map node and its corresponding data field, and the quality label of the corresponding data field.
[0050] Optionally, the system displays each mind map node, its corresponding data field, the matching score of each mind map node and its corresponding data field, and the quality label of the corresponding data field. Based on the matching score and a matching score threshold, the system filters the corresponding data fields of the mind map nodes, removing data fields whose matching scores are less than the threshold. If there is ambiguity in the matching scores of corresponding data fields, the system selects fields based on their quality labels, removing data fields with quality disclosure in the quality labels.
[0051] In this embodiment, by visually presenting each mind map node, its corresponding data field, the matching score between the two, and the quality label of the field, the transparency and interpretability of the data mapping process are achieved. This allows users to intuitively judge the accuracy and reliability of the binding of each node, improves the credibility and user-friendliness of the system, and avoids analysis bias caused by incorrect binding or poor data.
[0052] Based on the above embodiments, this application also provides a process for determining a candidate set of relations in a data analysis method. Figure 3 This is a flowchart illustrating the process of determining a candidate set of relations in a data analysis method provided in an embodiment of this application, as shown below. Figure 3 As shown, the method also includes: Step 301: Based on the data values of each mind map node and its corresponding data field in the mapping table, calculate the correlation parameters of every two nodes among all unconnected nodes in the preset mind map structure.
[0053] In this context, any two nodes in the current mind map structure that are not connected by an edge are considered to be related. The correlation parameter is a numerical indicator representing the degree of association between the data fields corresponding to the two unconnected nodes.
[0054] Optionally, every two nodes among all unconnected nodes in the preset mind map structure are identified, and the corresponding data field for each pair of nodes is determined based on each mind map node in the mapping table. The data value of the corresponding data field is then determined from the target data source based on this data field. Finally, the correlation parameter for each pair of nodes is determined based on the corresponding data field and its value.
[0055] Optionally, based on the data values of each mind map node and its corresponding data field in the mapping table, and the data values of all fields in the target data source, the correlation parameters between the fields in the target data source that are not in the mapping table and the fields in the mapping table are determined. Based on the correlation parameters, new nodes are added to the preset mind map structure. The new nodes are nodes corresponding to one or more fields in the target data source that are not in the mapping table.
[0056] Step 302: Calculate the potential association parameters between every two nodes based on the business entities of the data fields corresponding to each mind map node in the mapping table.
[0057] Among them, the potential association parameter indicates that although two nodes are not directly connected, their fields come from the same or related business entities and thus have the possibility of indirect connection.
[0058] Optionally, based on the business entities in the data fields corresponding to each mind map node in the mapping table, the potential association parameters between every two nodes are calculated using a graph algorithm. The graph algorithm can be co-occurrence analysis, community detection, shortest path, similarity calculation, etc., and this embodiment does not impose any limitations on it.
[0059] Step 303: Based on the correlation parameters between every two nodes and the potential association parameters between every two nodes, obtain the candidate relationships of the preset mind map structure and the confidence level of each candidate relationship.
[0060] The confidence score is a weighted value obtained by combining the relevance parameter and the potential association parameter, and its range is between 0 and 1. The confidence score is a quantitative indicator used to describe the reliability and probability of a genuine association between two unconnected nodes.
[0061] Optionally, the correlation parameters of every two nodes and the potential association parameters between every two nodes are weighted to obtain the confidence level of every two nodes based on the weighting result. Based on the confidence level of every two nodes, every two nodes with a confidence level higher than a preset confidence threshold are determined as candidate relationships of the preset mind map structure and the confidence level of each candidate relationship.
[0062] Step 304: Generate a candidate relation set based on the candidate relations and the confidence level of each candidate relation.
[0063] The candidate relation set is sorted according to confidence level.
[0064] Optionally, each candidate relation is sorted according to its confidence level to generate a candidate relation set, which includes every two nodes, the order between the two nodes, the edge between the two nodes, and the confidence level of the two nodes.
[0065] In this embodiment, the discovery of potential relationships between unconnected nodes in a preset mind map structure enhances the automated construction and dynamic expansion capabilities of knowledge graphs or mind maps, avoids misjudgments caused by relying solely on data or structure, and improves the accuracy and interpretability of recommendation results.
[0066] Based on the above embodiments, this application also provides a process for determining the correlation parameters between every two nodes in data analysis. Figure 4 This application provides a schematic diagram of a data analysis process for determining the correlation parameters between every two nodes, as illustrated in the embodiments of this application. Figure 4 As shown, in step 301 above, based on the data values of each mind map node and its corresponding data field in the mapping table, the correlation parameters between every two nodes among all unconnected nodes in the preset mind map structure are calculated, including: Step 401: If the data field values of the first and second nodes among all unconnected nodes are numerical or discrete data, then a statistical algorithm is used to obtain the correlation score.
[0067] Numerical data refers to continuous or discrete numerical fields. Discrete data has a finite number of values and is not continuous. Statistical algorithms are mathematical methods used to measure whether a linear or non-linear relationship exists between two nodes, and the output is a correlation score. Statistical algorithms can be correlation coefficients, mutual information, etc., and this application does not limit them.
[0068] Optionally, if the data fields corresponding to the first and second nodes among all unconnected nodes are numerical, then the correlation score is obtained using the Pearson correlation coefficient or the Spearman coefficient. The larger the absolute value of the correlation score, the stronger the correlation.
[0069] Optionally, if the data values of the corresponding data fields of the first and second nodes among all unconnected nodes are discrete data, then a correlation score is obtained through a chi-square test or mutual information. The closer the correlation score is to 1, the stronger the correlation.
[0070] Optionally, if the data fields corresponding to the first and second nodes among all unconnected nodes are numerical and discrete data, then a correlation score is obtained through analysis of variance or a rank-sum test. The closer the correlation score is to 1, the more significant the group difference.
[0071] Step 402: If the data values of the third and fourth nodes among all unconnected nodes are time-series data, use the causal structure algorithm to obtain the causal score.
[0072] The correlation parameters include correlation score and causality score. Time-series data means there is a causal relationship between the third and fourth nodes; for example, the third node occurs before the fourth. Causal structure algorithms can include Granger causality tests, transition entropy, dynamic time warping, and causal inference, etc., and this application does not limit these methods. The causality score is used to quantify the reliability of the causal association between nodes in time-series data.
[0073] Optionally, if the data values of the corresponding data fields of the third and fourth nodes among all unconnected nodes are time-series data, a causal structure algorithm is used to determine the causal relationship between the third and fourth nodes and obtain a causal score.
[0074] In this embodiment, differentiated analysis algorithms are employed for different types of data features, improving the scientific rigor and accuracy of node relationship identification. For numerical or discrete data, statistical algorithms are used to calculate correlation scores, effectively capturing the collaborative change trends between variables. For time-series data, a causal structure algorithm is introduced to calculate causal scores, identifying potential driving relationships between variables. By incorporating both correlation and causal scores into the correlation parameter system, the system's ability to understand complex business logic is enhanced, avoiding the misjudgment of incidental correlations as valid connections.
[0075] Based on the above embodiments, this application also provides a process for generating an incremental relationship suggestion list in a data analysis method. In addition to steps 301-304 above, the method further includes: Based on the candidate set of relationships, generate an incremental list of relationship suggestions with a pre-defined mind map structure.
[0076] The incremental relationship suggestion list includes: the suggested target variable relationship, the two mind map nodes corresponding to the target edge relationship, and the direction of the target edge relationship.
[0077] Optionally, an incremental relationship suggestion list with a preset mind map structure is generated based on the confidence levels of the relationship candidates in the relationship candidate set. Specifically, by setting a confidence threshold, relationship candidates with confidence levels greater than a certain threshold are added to the incremental relationship suggestion list, and the relationship candidates are sorted according to their confidence levels to obtain the incremental relationship suggestion list. The incremental relationship suggestion list includes: the target variable relationship to be added, the two mind map nodes corresponding to the target edge relationship, and the direction of the target edge relationship.
[0078] In the embodiments of this application, potential associations with high confidence can be recommended, which improves the accuracy and interpretability of relationship discovery, helps users identify key driving factors and influence paths, and avoids bias caused by relying solely on experience or intuition.
[0079] Based on the above embodiments, this application also provides a process for displaying candidate relationships in a data analysis method. Figure 5 This is a flowchart illustrating the display of candidate relationships in a data analysis method provided in an embodiment of this application, as shown below. Figure 5 As shown, based on steps 301-304 above, the method further includes: Step 501: Based on the candidate relation set, add candidate relations sequentially to the preset mind map structure to obtain the updated mind map structure.
[0080] Optionally, candidate relationships are added sequentially according to their confidence level, and an incremental force-guided layout or multi-level layout algorithm is adopted. Priority is given to maintaining the relative positions of existing nodes, with only minor adjustments made in local areas. Furthermore, the node displacement in each iteration is limited to achieve local stability of the layout and avoid frequent jumps in the overall graph.
[0081] Step 502: Display the updated mind map structure and differentiate candidate relationships within the updated mind map structure.
[0082] Among these features, candidate relationships can be displayed using different colors or line styles.
[0083] Optionally, the updated mind map structure can be displayed, and candidate relationships can be differentiated by using different colors or line styles. Options such as "Accept," "Reject," and "Pending" can be added next to the candidate relationship, allowing the user's selection to determine the action taken on that relationship. Labels or notes can also be added to candidate relationships to indicate that they are newly added content.
[0084] In this embodiment, candidate relationships are incrementally added sequentially to a preset mind map structure based on the candidate relationship set, and the newly added relationships are visually distinguished. This enables a gradual evolution and enhanced visualization of the mind map structure without disrupting the user's existing cognitive layout. This approach maintains the local stability of the graph, avoids user disorientation caused by frequent and significant structural changes, and improves the comprehensibility of the human-computer collaborative analysis process.
[0085] Based on the above embodiments, this application also provides a process for generating directed acyclic graphs in a data analysis method. Figure 6 This is a flowchart illustrating the generation of a directed acyclic graph in a data analysis method provided in this application embodiment. Based on steps 501-502 above, the method further includes: Step 601: After the user confirms the updated mind map structure, match the operator template corresponding to each mind map node in the updated mind map structure.
[0086] The operator template is an operator template obtained by matching through a preset operator template library. The operator can be a data filtering operator, a summary operator, a correlation operator, a graph query operator, a statistical analysis operator, or a machine learning model, etc. This application embodiment does not limit this.
[0087] Optionally, after the user confirms the updated mind map structure, that is, the current mind map structure is the updated mind map structure, the operator template corresponding to each mind map node in the updated mind map structure is matched based on the preset operator template library.
[0088] Step 602: Determine the execution order of each operator template based on the relationship between the nodes in the updated mind map structure.
[0089] The relationships between nodes in the updated mind map structure are the paths and hierarchical relationships between them.
[0090] Optionally, the updated mind map structure is parsed into a directed graph model with topological relationships based on the relationships between the nodes in the updated mind map structure. Each mind map node represents a business entity, and each edge represents the association between entities. Based on the path and hierarchical structure between nodes in the updated mind map structure, a data dependency graph between operators is constructed. The output generated by the upstream node serves as the input of the downstream node. The path length and nesting level determine the execution priority, thereby obtaining the execution order of the operator template.
[0091] Step 603: Generate a directed acyclic graph of the updated mind map structure according to the execution order of each operator template.
[0092] In a directed acyclic graph (DAG), each node corresponds to an operator template, and the edges between nodes represent the data dependencies between operator templates. A DAG is a graph with direction but no cycles.
[0093] Optionally, all the determined operator templates are treated as nodes, and directed edges are added according to data dependencies to form a directed acyclic graph. Each operator template extracts configuration parameters, such as field source, time range, filtering conditions, and algorithm type, from the mapping table and the updated mind map structure, and embeds them into the node metadata in the directed acyclic graph to form an executable operator instance.
[0094] Optionally, for repeated subprocesses in a directed acyclic graph, a cache and increment-based computation strategy can be implemented to reuse historical computation results when the input remains unchanged.
[0095] In this embodiment, the updated mind map structure confirmed by the user is automatically mapped into a directed acyclic graph with a clear execution order. This transforms the data analysis system from a visual semantic model into an executable analysis process, improving its intelligence and human-machine collaboration efficiency. It avoids the errors and inefficiencies of manually orchestrating complex processes and enhances the system's responsiveness.
[0096] Based on the above embodiments, this application also provides a flowchart for displaying results in a data analysis method. Figure 7 This is a flowchart illustrating the display of results in a data analysis method provided in an embodiment of this application, as shown below. Figure 7 As shown, based on steps 601-603 above, the method further includes: Step 701: Execute each operator template sequentially according to the directed acyclic graph to obtain the intermediate and final results of each brain graph node in the updated brain graph structure.
[0097] The intermediate results represent the phased data output by each operator during the execution of the directed acyclic graph (DAG). The final result represents the output generated by the last one or more target operators in the DAG.
[0098] Optionally, the scheduling engine executes each operator template sequentially according to the directed acyclic graph to obtain intermediate and final results for each brain graph node in the updated brain graph structure.
[0099] Step 702: Display the intermediate and final results of each node in the updated mind map structure.
[0100] Optionally, after each operator template is executed, its output is marked as an intermediate result, and lineage information is recorded. Key values from the operator template execution process, corresponding intermediate results, and the final result are then backfilled into the corresponding mind map node and its edges, and displayed on the mind map node using numerical values, colors, icons, etc. Key values include statistical indicators, labels, confidence levels, and time windows. Lineage information includes the source operator, processing logic, and timestamp.
[0101] Optionally, when the parameters of a node or edge are directly modified in the real-world interface, the system updates the corresponding operator parameters in reverse and triggers incremental recalculation of the affected subgraph, ensuring a two-way binding closed loop of editing upon execution and execution upon backfilling. The parameters of the node or edge can be thresholds, time windows, or filtering conditions; this embodiment does not impose any limitations on these.
[0102] Optionally, after displaying the intermediate and final results of each brain map node in the updated brain map structure, an execution log of the directed acyclic graph is generated based on the calculation process.
[0103] Based on the above embodiments, this application also provides a process for generating a data analysis report in a data analysis method. Figure 8 This is a schematic diagram of the process for generating a data analysis report in a data analysis method provided in an embodiment of this application, as shown below. Figure 8 As shown, based on steps 701-702 above, the method further includes: Step 801: Based on the intermediate and final results of each brain map node in the updated brain map structure, a preset feature algorithm is used to determine the feature contribution of each brain map node in the updated brain map structure.
[0104] Feature contribution refers to the degree of influence of each input variable on the output result in a model or statistical analysis. Preset feature algorithms are used to calculate the interpretability of feature contributions. These can be SHAP (SHapley Additive ex Planations), based on game theory, which fairly allocates the contribution value of each feature; Permutation Importance, which observes the degree of performance degradation of a model after shuffling a feature; LIME, Tree Interpreter, etc.
[0105] Optionally, based on intermediate and final results, a preset feature algorithm is used to determine the feature contribution of each mind map node, identifying the target analysis task type. If the final result comes from a machine learning model, SHAP or PermutationImportance is enabled; if it is statistical inference, standardized coefficients are directly extracted as the contribution; if it is output from a rule engine, the contribution is restored according to weighted rules. The original features used by the model are traced back to their corresponding mind map nodes and fields, and the preset feature algorithm is used to output the contribution value and confidence interval of each feature. When multiple fields belong to the same mind map node, their feature contributions are weighted and merged to obtain the overall influence score of the node, and a feature contribution score table for each mind map node is output.
[0106] Step 802: Generate a data feature evidence chain structure based on the feature contributions of each brain map node in the updated brain map structure.
[0107] The data feature evidence chain structure is used to indicate the path of each node in the updated mind map structure. It traces back to the original data and key features through the final conclusion. Using the mind map as a skeleton, it records the dependencies between conclusion nodes, key paths, decision nodes, core features, and original fields, forming a verifiable evidence path.
[0108] Optionally, the conclusion node containing the final result is selected as the starting point of the evidence chain. A lineage and feature contribution ranking are performed using a directed acyclic graph (DAG) to identify the upstream node with the greatest impact on the conclusion. The path is organized into a directed tree or chain structure, while preserving the links of the original data samples. The path of the evidence chain is highlighted in the mind map, and supporting data screenshots, distribution comparison charts, and control group information can be viewed at any stage. A structured data feature evidence chain is output.
[0109] Step 803: Based on the intermediate and final results of each mind map node in the updated mind map structure and the data feature evidence chain structure, generate a data analysis report using a preset report template.
[0110] The preset report template is a predefined document structure framework, including chapter titles, chart placeholders, indicator description templates, condition judgment rules, etc., and supports automatic content filling to generate reports.
[0111] Optionally, a corresponding preset report template is obtained, and the intermediate and final results of each mind map node in the updated mind map structure, as well as the data feature evidence chain structure, are filled into the preset report template using natural language understanding to generate a data analysis report. The data report can be in Portable Document Format (PDF), PowerPoint Presentation, or Hypertext Markup Language (HTML), etc., and this application embodiment does not limit this.
[0112] Optionally, role-based access control can be applied throughout the entire process of data access, analysis execution, and result display, restricting the datasets and functions accessible to different roles. Specifically, anonymization strategies can be implemented for sensitive fields, and the anonymization rules can be recorded in the lineage information. The anonymization strategy can include masking, de-identification, etc.
[0113] Optionally, the data lineage from the data source to the analysis results, as well as the audit logs of all critical operations, can be determined to facilitate traceability and compliance review.
[0114] Optionally, in cross-departmental or cross-domain collaborative scenarios, federated learning or privacy-preserving computation schemes can be enabled as needed to complete joint modeling without directly exposing the original data.
[0115] In this embodiment of the application, by combining the intermediate results and the final results in the analysis process with a preset feature algorithm to automatically calculate the feature contribution of each mind map node, and generating a traceable data feature evidence chain structure based on the contribution, a closed loop from conclusion to critical path is realized, which lowers the threshold for non-professional users to understand complex analysis results, enhances the credibility of decision-making and compliance audit capabilities, and reduces the time cost of manually writing reports.
[0116] Based on the same inventive concept, this application also provides a data analysis device corresponding to the data analysis method. Since the principle of the device in this application is similar to the data analysis method described above in this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0117] Figure 9 This is a schematic diagram of the structure of a data analysis device provided in an embodiment of this application, as shown below. Figure 9 As shown, the device includes: The first acquisition module 901 is used to acquire the target data source selected by the user from multiple heterogeneous data sources; The first determining module 902 is used to perform pattern recognition based on the metadata of the target data source to obtain pattern description information in the target data source. The pattern description information is used to indicate the structure of each data object in the target data source and the relationship between the data objects. The second acquisition module 903 is used to acquire multiple data fields from the target data source based on the pattern description information; The third acquisition module 904 is used to acquire the data analysis results of the target data source based on the pattern description information and information from multiple data fields; The second determining module 905 is used to determine the data fields corresponding to each mind map node in the preset mind map structure from multiple data fields based on the data analysis results and the preset mind map structure. The generation module 906 is used to generate a mapping table between each mind map node and the corresponding data field based on each mind map node, the corresponding data field of each mind map node, and the business entity of the corresponding data field. The business entity of the corresponding data field is used to indicate the data source of the corresponding data field.
[0118] In one possible implementation, the generation module 906 is further configured to: obtain a field profile of each data field based on the data values of each data field in the target data source; Based on the field profile of each data field, a quality assessment is performed on each data field to obtain a quality label for each data field.
[0119] In one possible implementation, the generation module 906 is further configured to: display each mind map node, its corresponding data field, the matching score of each mind map node and its corresponding data field, and the quality label of the corresponding data field.
[0120] In one possible implementation, the generation module 906 is further configured to: calculate the correlation parameter of every two nodes among all unconnected nodes in the preset mind map structure based on the data values of each mind map node and the corresponding data field in the mapping table. Based on the business entities of the data fields corresponding to each mind map node in the mapping table, calculate the potential association parameters between every two nodes; Based on the correlation parameters between every two nodes and the potential association parameters between every two nodes, candidate relationships of the preset mind map structure and the confidence level of each candidate relationship are obtained. A candidate relation set is generated based on the candidate relations and the confidence level of each candidate relation.
[0121] In one possible implementation, the generation module 906 is further configured to: if the data values of the corresponding data fields of the first and second nodes among all unconnected nodes are numerical or discrete data, then use a statistical algorithm to obtain a correlation score; If the data fields of the third and fourth nodes among all unconnected nodes are time-series data, the causal structure algorithm is used to obtain the causal score, where the correlation parameter includes the correlation score and the causal score.
[0122] In one possible implementation, the generation module 906 is further configured to: generate an incremental relation suggestion list with a preset mind map structure based on the relation candidate set, the incremental relation suggestion list including: the target variable relation to be added, the two mind map nodes corresponding to the target edge relation, and the direction of the target edge relation.
[0123] In one possible implementation, the generation module 906 is further configured to: add candidate relations sequentially to the preset mind map structure according to the candidate relation set, so as to obtain the updated mind map structure; Display the updated mind map structure and differentiate candidate relationships within the updated mind map structure.
[0124] In one possible implementation, the generation module 906 is further configured to: after the user confirms the updated mind map structure, match the operator template corresponding to each mind map node in the updated mind map structure according to the updated mind map structure. Based on the relationships between the nodes in the updated mind map structure, determine the execution order of each operator template; Based on the execution order of each operator template, a directed acyclic graph (DAG) of the updated mind map structure is generated. Each node in the DAG corresponds to an operator template, and the edges between nodes are used to represent the data dependencies between operator templates.
[0125] In one possible implementation, the generation module 906 is further configured to: execute each operator template sequentially according to the directed acyclic graph to obtain intermediate results and final results of each brain graph node in the updated brain graph structure; This displays the intermediate and final results of each node in the updated mind map structure.
[0126] In one possible implementation, the generation module 906 is further configured to: determine the feature contribution of each brain map node in the updated brain map structure by using a preset feature algorithm based on the intermediate and final results of each brain map node in the updated brain map structure. Based on the feature contributions of each brain map node in the updated brain map structure, a data feature evidence chain structure is generated, which is used to indicate the path of each brain map node in the updated brain map structure. Based on the intermediate and final results of each mind map node in the updated mind map structure, as well as the data feature evidence chain structure, a data analysis report is generated using a preset report template.
[0127] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.
[0128] This application also provides a computer device. Figure 10 This application provides a schematic diagram of the structure of a computer device, as shown in the embodiment of the present application. Figure 10 As shown, the computer device includes a processor 1001 and a memory 1002, and optionally, a bus 1003. The memory 1002 stores machine-readable instructions executable by the processor 1001. When the computer device is running, the processor 1001 and the memory 1002 communicate via the bus 1003. When the machine-readable instructions are executed by the processor 1001, the steps of the aforementioned data analysis method are performed.
[0129] This application also provides a computer-readable storage medium storing a computer program, which, when run by a processor, executes the steps of the above-described data analysis method.
[0130] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.
[0131] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0132] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A data analysis method, characterized in that, The method includes: Retrieve the target data source selected by the user from multiple heterogeneous data sources; Pattern recognition is performed based on the metadata of the target data source to obtain pattern description information in the target data source. The pattern description information is used to indicate the structure of each data object in the target data source and the relationship between the data objects. Based on the pattern description information, obtain multiple data fields from the target data source; Based on the pattern description information and the information of the multiple data fields, obtain the data profiling results of the target data source; Based on the data analysis results and the preset mind map structure, the data fields corresponding to each mind map node in the preset mind map structure are determined from the multiple data fields; Based on each mind map node, its corresponding data field, and the business entity of the corresponding data field, a mapping table is generated for each mind map node and its data field, wherein the business entity of the corresponding data field is used to indicate the data source of the corresponding data field.
2. The method according to claim 1, characterized in that, The method further includes: Based on the data values of each data field in the target data source, obtain the field profile of each data field; Based on the field profiles of each data field, a quality assessment is performed on each data field to obtain a quality label for each data field.
3. The method according to claim 2, characterized in that, The method further includes: Displays each mind map node, its corresponding data field, the matching score of each mind map node and its corresponding data field, and the quality label of the corresponding data field.
4. The method according to claim 1, characterized in that, The method further includes: Based on the data values of each mind map node in the mapping table and the corresponding data fields, calculate the correlation parameter of every two nodes among all unconnected nodes in the preset mind map structure; Based on the business entities of the data fields corresponding to each mind map node in the mapping table, calculate the potential association parameters between every two nodes; Based on the correlation parameters between every two nodes and the potential association parameters between every two nodes, candidate relationships of the preset mind map structure and the confidence level of each candidate relationship are obtained. A candidate relation set is generated based on the candidate relations and the confidence level of each candidate relation.
5. The method according to claim 4, characterized in that, The step of calculating the correlation parameter between every two nodes among all unconnected nodes in the preset mind map structure based on the data values of each mind map node in the mapping table and the corresponding data field includes: If the data field values of the first and second nodes among all the unconnected nodes are numerical or discrete, then a statistical algorithm is used to obtain a correlation score. If the data values of the corresponding data fields of the third and fourth nodes among all unconnected nodes are time-series data, a causal structure algorithm is used to obtain a causal score, wherein the correlation parameter includes the correlation score and the causal score.
6. The method according to claim 4, characterized in that, The method further includes: Based on the candidate set of relationships, an incremental relationship suggestion list for the preset mind map structure is generated. The incremental relationship suggestion list includes: the target variable relationship to be added, the two mind map nodes corresponding to the target edge relationship, and the direction of the target edge relationship.
7. The method according to claim 4, characterized in that, The method further includes: Based on the candidate relation set, the candidate relations are sequentially added to the preset mind map structure to obtain the updated mind map structure. The updated mind map structure is displayed, and the candidate relationships are distinguished within the updated mind map structure.
8. The method according to claim 7, characterized in that, The method further includes: After the user confirms the updated mind map structure, the operator template corresponding to each mind map node in the updated mind map structure is matched according to the updated mind map structure. The execution order of each operator template is determined based on the relationship between the nodes in the updated mind map structure. Based on the execution order of the various operator templates, a directed acyclic graph of the updated mind map structure is generated. Each node in the directed acyclic graph corresponds to an operator template, and the edges between nodes are used to represent the data dependencies between operator templates.
9. The method according to claim 8, characterized in that, The method further includes: The operator templates are executed sequentially according to the directed acyclic graph to obtain the intermediate and final results of each brain graph node in the updated brain graph structure. The intermediate results and the final result of each brain map node in the updated brain map structure are displayed.
10. The method according to claim 9, characterized in that, The method further includes: Based on the intermediate results and the final results of each brain map node in the updated brain map structure, a preset feature algorithm is used to determine the feature contribution of each brain map node in the updated brain map structure. Based on the feature contributions of each brain map node in the updated brain map structure, a data feature evidence chain structure is generated, wherein the data feature evidence chain structure is used to indicate the path of each brain map node in the updated brain map structure. Based on the intermediate results and final results of each mind map node in the updated mind map structure, as well as the data feature evidence chain structure, a data analysis report is generated using a preset report template.