Implementation method of multi-dimensional data visual analysis platform based on semantic driving
By utilizing topological data analysis theory and topologically-aware dynamic knowledge graphs, the problem of insufficient understanding of implicit structures in multidimensional data visualization systems is solved, achieving high-quality data visualization and adaptive optimization, thereby improving analysis efficiency and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-03-31
AI Technical Summary
Existing data visualization systems lack a deep understanding of the implicit structures and relationships in multidimensional data scenarios, cannot effectively preserve data structure characteristics, and lack adaptive optimization capabilities, resulting in inaccurate analysis results and poor user experience.
Employing topological data analysis theory, semantic parsing is performed through a topology-aware dynamic knowledge graph. Combined with a topology-preserving visualization mapping mechanism, end-to-end automatic conversion from natural language queries to high-quality data visualization is achieved, including semantic parsing, multi-dimensional data fusion processing, and visualization mapping.
It improves semantic understanding capabilities, preserves data structure features, generates more insightful and intuitive visualizations, lowers the barrier to entry, and achieves adaptive optimization and improved analysis efficiency.
Smart Images

Figure CN121764945A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis and visualization technology, specifically to a method for implementing a semantic-driven multidimensional data visualization analysis platform, and more particularly to a multidimensional data intelligent analysis method that combines natural language processing, topological data analysis, and visualization technology. Background Technology
[0002] With the advent of the big data era, multidimensional data analysis and visualization have become key means of gaining data insights. Traditional data visualization methods typically require users to possess professional data analysis knowledge and visualization design skills, using complex query languages or interface operations for data processing and visualization configuration. This approach is not only costly to learn but also inefficient, making it difficult to meet the demand for rapid data insights.
[0003] In existing technologies, some systems attempt to simplify the data analysis process through natural language interfaces, but these systems primarily focus on the accuracy of semantic parsing, neglecting the structural features in the semantic space. Traditional natural language-driven visualization systems typically employ rule-based or simple machine learning methods, which fail to deeply understand complex semantic relationships, especially the implicit structures and associations in multidimensional data scenarios.
[0004] Furthermore, existing data visualization systems often employ predefined templates and mapping rules in the data processing and visualization mapping stages, lacking in-depth analysis and preservation of the inherent structural characteristics of the data. This results in visualizations that, while conforming to formal requirements, may lose key structural information from the data, affecting the accuracy and intuitiveness of the analysis results.
[0005] Another problem with traditional technologies is the lack of an effective feedback optimization mechanism. The system cannot automatically adjust and optimize its internal model based on user interaction, making it difficult to achieve continuous performance improvement and user experience enhancement.
[0006] Therefore, there is an urgent need for a multi-dimensional data visualization analysis platform that can deeply understand semantic relationships, preserve data structure features, support intelligent visualization mapping, and has adaptive optimization capabilities. Summary of the Invention
[0007] The purpose of this invention is to provide a semantically driven multidimensional data visualization analysis platform implementation method. By introducing topological data analysis theory to enhance semantic understanding capabilities, a data processing and visualization mapping mechanism that preserves topological structure is established, enabling end-to-end automatic conversion from natural language queries to high-quality data visualization.
[0008] This invention proposes a method for implementing a semantically driven multidimensional data visual analysis platform, comprising:
[0009] Obtain the natural language query input by the user;
[0010] Based on the natural language query, semantic parsing is performed through a topology-aware dynamic knowledge graph to obtain semantic parsing results. The topology-aware dynamic knowledge graph represents the topological relationship of the semantic structure by calculating the persistent coherence features of the semantic space.
[0011] Based on the semantic parsing results, multidimensional data fusion processing is performed to obtain the processed data view;
[0012] Based on the processed data view, a visualization scheme is generated through a topology-preserving visualization mapping mechanism, wherein the topology-preserving visualization mapping mechanism ensures that the visualization result retains the topological structure features of the original data; and
[0013] The visualization results are rendered according to the visualization scheme, and the topology-aware dynamic knowledge graph and the topology-preserving visualization mapping mechanism are optimized in response to user interaction operations.
[0014] As a preferred method, semantic parsing using topology-aware dynamic knowledge graphs includes:
[0015] The natural language query is mapped to a multi-level semantic vector space to form a semantic point cloud;
[0016] A simplified complex is constructed on the semantic point cloud, and persistent homology groups of different dimensions are calculated to obtain topological features such as connectivity, loops, and holes.
[0017] Based on the aforementioned topological features, nodes and relationships in the knowledge graph are updated using a topology-aware graph neural network to obtain a semantic representation that integrates the topological structure; and
[0018] Based on the semantic representation, the query intent and data requirements are determined, and the semantic parsing result is generated.
[0019] Preferably, the multi-level semantic vector space includes word-level, phrase-level, sentence-level, and concept-level semantic representations. The natural language query is encoded from different levels of abstraction through a hierarchical Transformer architecture, with each level corresponding to a different granularity of semantic understanding.
[0020] Preferably, the topology-aware graph neural network includes:
[0021] A topology-enhanced message passing mechanism integrates topological features into the node representation update process;
[0022] A multi-scale topological feature integration layer dynamically adjusts the importance of topological features at different scales based on the query context; and
[0023] The update strategy for maintaining topological consistency ensures that key topological structures are preserved during the knowledge graph update process.
[0024] As a preferred option, performing multidimensional data fusion processing includes:
[0025] Based on the semantic parsing results, identify the data sources related to the query and evaluate their topological matching degree;
[0026] Perform topology-preserving data fusion on the identified data sources to ensure that key topology structures are maintained during the fusion process;
[0027] Based on the fused data, topology-feature-driven data transformation and aggregation operations are performed, including topology-persistent dimensionality reduction, topology-preserving data normalization, and topology-aware anomaly detection; and
[0028] Generate a data view with topological annotations as input for the visualization mapping.
[0029] As a preferred approach, the visualization scheme generated through a topology-preserving visualization mapping mechanism includes:
[0030] Establish a mapping relationship between topological features and visual channels, where topological features of different dimensions are mapped to different types of visual representations;
[0031] Perform topological characteristic analysis on templates in the visualization template library, and select the optimal matching template based on topological similarity;
[0032] Based on the selected template and topological features, generate the visual channel parameter configuration; and
[0033] Perform topology consistency verification, evaluate the topology fidelity of the visualization scheme, and improve the visual representation of the topology structure through iterative optimization.
[0034] Preferably, the mapping of the different dimensions of topological features to different types of visual representations includes:
[0035] 0. Durable features are primarily mapped to the color, shape, and transparency visual channels;
[0036] 1. Durable features are primarily mapped to spatial location, connectivity, and boundary representations of the visual channel; and
[0037] 2. Durable features are mainly mapped to hierarchical structures, grouped regions, and background processing visual channels.
[0038] 8. The method according to claim 6, wherein performing topology consistency verification includes:
[0039] Calculate the topological features of the visualization results and compare them with the topological features of the original data view;
[0040] Topological fidelity is assessed from dimensions such as connectivity preservation, ring structure preservation, hole preservation, persistent correlation, and overall structural similarity; and
[0041] When the topology fidelity is lower than the preset threshold, adjust the visual channel parameters and regenerate the visualization scheme until a satisfactory topology fidelity is achieved or the maximum number of iterations is reached.
[0042] As a preferred option, system performance optimization is also included, including:
[0043] Implement a multi-level caching mechanism to cache semantic parsing results, topological features, and visualization templates;
[0044] A parallel computing framework is used to accelerate topology feature extraction and analysis;
[0045] Progressive rendering is implemented based on topological importance, prioritizing the display of elements with higher topological importance; and
[0046] It employs a layered detail technique to dynamically adjust the rendering details of the topology based on the interaction context.
[0047] Preferably, the mechanism for optimizing the topology-aware dynamic knowledge graph and the topology-preserving visualization mapping in response to user interaction includes:
[0048] Collect user interaction data with visualization results and extract the implicit topological preferences within it;
[0049] Update the importance weights of topological features in the knowledge graph based on the topological preferences;
[0050] Adjusting topological feature extraction parameters and visualization mapping strategies; and
[0051] For frequent query patterns, relevant topological features and visualization configurations are pre-calculated and cached to improve system response speed.
[0052] The present invention has the following beneficial effects:
[0053] 1. Enhanced semantic understanding: Through topology-aware dynamic knowledge graphs, the system can capture complex structural relationships in the semantic space, improving the accuracy of understanding user query intent. In particular, for complex queries containing implicit structures, the accuracy rate can be improved by up to 35%.
[0054] 2. Preserve data structure features: Data processing methods based on topology features ensure that key topology structures are preserved during data fusion, transformation and aggregation, making data analysis results more accurate and reliable.
[0055] 3. Improve visualization quality: The topology-preserving visualization mapping mechanism can map the inherent structural features of data to visual representations, generating more insightful and intuitive visualization results, making it easier for users to identify patterns, trends and anomalies in the data.
[0056] 4. Achieve adaptive optimization: By collecting user interaction data, the system automatically adjusts and optimizes its internal model, forming a closed-loop feedback mechanism to continuously improve system performance and user experience.
[0057] 5. Lowering the barrier to entry: Users only need to express their analysis needs through natural language, without needing to master professional data analysis and visualization knowledge, to obtain high-quality data visualization results, which greatly lowers the barrier to entry for data analysis.
[0058] 6. Improved analytical efficiency: End-to-end automated processing significantly reduces the time from data to insights, and the time required for users to understand visualizations is reduced by an average of 30%, resulting in a significant improvement in analytical efficiency. Attached Figure Description
[0059] Figure 1 This is a system architecture diagram of a semantically driven multidimensional data visual analysis platform in an embodiment of the present invention;
[0060] Figure 2 This is a flowchart illustrating the construction process of a multi-level semantic vector space in an embodiment of the present invention.
[0061] Figure 3 This is a flowchart of the multi-dimensional data fusion processing in an embodiment of the present invention;
[0062] Figure 4 This is a schematic diagram of the topology-preserving visualization mapping mechanism in an embodiment of the present invention;
[0063] Figure 5 This is a flowchart of the topology consistency verification and optimization process in an embodiment of the present invention; Detailed Implementation
[0064] Please refer to the attached document. Figure 1-5 The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0065] Reference Figure 1This invention provides a method for implementing a semantically driven multidimensional data visualization analysis platform. The method includes: acquiring a natural language query input by a user; performing semantic parsing based on the natural language query using a topology-aware dynamic knowledge graph to obtain a semantic parsing result; performing multidimensional data fusion processing based on the semantic parsing result to obtain a processed data view; generating a visualization scheme based on the processed data view using a topology-preserving visualization mapping mechanism; and rendering the visualization result according to the visualization scheme, and optimizing the topology-aware dynamic knowledge graph and the topology-preserving visualization mapping mechanism in response to user interaction operations.
[0066] In this invention, a key innovation is the topology-aware dynamic knowledge graph, which represents the topological relationships of semantic structures by computing persistent cohomology features of the semantic space. Persistent cohomology theory is an important method in topological data analysis, capable of capturing structural features in data, such as connectivity, loops, and holes. These topological features are of great significance for understanding complex relationships in the semantic space.
[0067] Similarly, the topology-preserving visualization mapping mechanism is another important innovation, ensuring that the visualization results retain the topological structural features of the original data. This structure-preserving mapping method can more accurately reflect the inherent relationships in the data, improving the insight and interpretability of the visualization results.
[0068] As a specific application scenario, consider financial analysts using this system to analyze market trends. When an analyst inputs a natural language query such as "analyze the correlation patterns between stock prices and trading volumes across various industries over the past year," the system uses a topology-aware dynamic knowledge graph to understand the query intent, identifying key concepts such as industry, stock price, trading volume, and correlation patterns, as well as their relationships. Then, it performs multi-dimensional data fusion, integrating stock price and trading volume data from different industries while preserving the topological characteristics of the data. Finally, the system generates visualizations that clearly show the correlation patterns between stock prices and trading volumes across various industries, including cyclical patterns, clustering structures, and outliers, enabling analysts to intuitively discover market patterns and investment opportunities.
[0069] Reference Figure 2 The semantic parsing process of the topology-aware dynamic knowledge graph of the present invention includes: mapping the natural language query to a multi-level semantic vector space to form a semantic point cloud; constructing a simplified complex on the semantic point cloud and calculating persistent homology groups of different dimensions to obtain topological features such as connectivity, loops, and holes; updating the nodes and relationships in the knowledge graph through a topology-aware graph neural network based on the topological features to obtain a semantic representation of the fused topological structure; and determining the query intent and data requirements according to the semantic representation to generate the semantic parsing result.
[0070] In a preferred embodiment of the present invention, the multi-level semantic vector space includes word-level, phrase-level, sentence-level, and concept-level semantic representations. A hierarchical Transformer architecture encodes the natural language query at different levels of abstraction, with each level corresponding to a different granularity of semantic understanding. Specifically, word-level representation captures the basic semantics of words, phrase-level representation captures the combined semantics of word groups, sentence-level representation captures the overall semantics of complete sentences, and concept-level representation captures abstract domain concepts and knowledge.
[0071] The process of constructing a multi-level semantic vector space can be represented as follows:
[0072] ,
[0073] in: It is a multi-level semantic vector space, which is a set containing four different levels of vector space; Let be a word-level semantic vector space, representing semantic representations at the word level, with dimensions of . ; Let be a phrase-level semantic vector space, representing phrase-level semantic representations, with dimension . ; Let be a sentence-level semantic vector space, representing sentence-level semantic representations, with dimension . ; Let be a concept-level semantic vector space, representing the semantic representation at the domain concept level, with dimensions of . .
[0074] After mapping natural language queries to a multi-level semantic vector space, the resulting semantic point cloud can be represented as:
[0075] ,
[0076] in: A semantic point cloud is a set of points; For the first A coordinate vector of semantic points represents a point in the semantic space; This represents the total number of semantic points, which is related to the number of key concepts in the query. The dimension of the semantic space is usually chosen as 128 or 256, to balance expressive power and computational efficiency. express 3D real space.
[0077] In constructing a simplified complex on a semantic point cloud, this invention employs the Vietoris-Rips complex, a commonly used complex constructed based on point-to-point distance. Specifically, for a given threshold parameter... If two points and The distance between them is less than If each pair of points is connected by an edge, then there is an edge between them. If every pair of three points is connected by an edge, then they form a triangle, and so on.
[0078] To compute persistent homology groups of different dimensions, this invention uses standard methods from persistent homology theory. First, with the parameters... The increase in [something] tracks the appearance and disappearance of topological features such as connected components, loops, and holes. The persistence of these features (i.e., their existence) [is also considered]. The range reflects their importance. Persistence diagrams are used to visualize the birth and death times of these features, while persistence barcodes use the length of the bars to represent the persistence of the features.
[0079] In the context of financial data analysis, a long-term homology characteristic of 0 (connected components) may correspond to different asset classes or market segments, a long-term homology characteristic of 1 (loops) may reflect market cycles or feedback loops, while a long-term homology characteristic of 2 (holes) may represent loopholes in the market structure or investment opportunities that have not been fully explored.
[0080] The topology-aware graph neural network of the present invention includes: a topology-enhanced message passing mechanism that integrates topological features into the node representation update process; a multi-scale topological feature integration layer that dynamically adjusts the importance of topological features at different scales according to the query context; and a topology consistency-preserving update strategy that ensures that key topological structures are maintained during the knowledge graph update process.
[0081] Topology-enhanced message passing is a core innovation of graph neural networks (GNNs), overcoming the limitation of traditional GNNs that only consider local connections. In traditional GNNs, node representation updates typically take the following form:
[0082] ,
[0083] in: For nodes In the The layer's representation vector has dimensions of ; For nodes The set of neighbors, containing nodes All directly connected nodes; This is the normalization constant, typically taking the value of [value missing]. ,in Represents a node The number of neighbors; For the first The learnable weight matrix of the layer has dimensions of ; It is a non-linear activation function, such as ReLU or tanh; Indicates a node All neighbors Perform a summation operation.
[0084] The topology-enhanced message passing mechanism of this invention incorporates the influence of topological features, and the update formula is modified as follows:
[0085] ,
[0086] in: For nodes In the The updated representation vector of the layer has dimensions of ; The edge weights are based on topological importance, representing the node's edge weight. For nodes The degree of influence, with a value range of [0,1]; This is the normalization constant, defined as above; For the first The learnable weight matrix of the layer has dimensions of ; For nodes In the The layer's representation vector has dimensions of ; For nodes The topological feature weight represents the importance of the topological feature to the node representation update, and its value ranges from [0,1]. For nodes In the The topological feature vector of the layer has a dimension of ; It is a non-linear activation function; Indicates a node All neighbors Perform a summation operation. In this way, the node representation update process takes into account both local connectivity information and global topology information.
[0087] In the multi-scale topological feature integration layer, the system dynamically adjusts the importance of topological features at different scales based on the query context. Specifically, for the scale parameter... Calculate the corresponding topological feature vector , , ..., Then, their weighted sum is calculated using an attention mechanism:
[0088] ,
[0089] in: For nodes The comprehensive topological feature vector, with dimension . ; For scale Importance weights represent the importance of topological features at this scale, satisfying... and ; For nodes In scale The topological feature vector below has a dimension of ; Indicates all Summation is performed on each scale. In practice, Typically, three to five values are chosen to balance expressive power and computational efficiency.
[0090] The topology consistency-preserving update strategy ensures that key topological structures are maintained during knowledge graph updates. When updating the knowledge graph, the system first calculates the topology signature of the current knowledge graph, then evaluates the topology consistency score of each candidate update scheme, and selects the scheme with the highest score for execution. The topology consistency score can be represented as:
[0091] ,
[0092] in: The topology consistency score ranges from [0,1], with a larger value indicating better preservation of the topology. For dimension The weights represent the importance of the topological features in that dimension, satisfying the following conditions: and ; Dimensions before update A persistent graph represents the set of birth and death times of topological features; For the updated dimensions Persistent graph; This is a similarity metric between persistent graphs, such as the normalized inverse of the Wasserstein distance; This indicates a summation operation across three dimensions: 0, 1, and 2. In practical applications, , and It can be configured according to the characteristics of a specific field. For example, in a financial data analysis scenario, it can be configured as follows: , , This is to emphasize the importance of loop structures (such as market cycles).
[0093] Reference Figure 3 The multidimensional data fusion processing procedure of the present invention includes: identifying data sources related to the query based on the semantic parsing results and evaluating their topological matching degree; performing topology-preserving data fusion on the identified data sources to ensure that key topological structures are maintained during the fusion process; performing topology feature-driven data transformation and aggregation operations based on the fused data, including topology-persistent dimensional reduction, topology-preserving data normalization, and topology-aware anomaly detection; and generating a data view with topological annotations as input for visualization mapping.
[0094] Evaluating the topology matching degree of the data source is the first step in multidimensional data fusion processing. It ensures that the selected data source is topologically compatible with the query semantics. Topology matching degree can be expressed as:
[0095] ,
[0096] in: For data source With query The topological matching degree, with a value range of A higher value indicates a higher degree of matching. For dimension The weights represent the importance of the topological features in that dimension, satisfying the following conditions: and ; For data source Dimensions Persistent graph; For query Dimensions Persistent graph; A similarity metric function between persistent graphs; This indicates a summation operation across 0, 1, and 2 dimensions. A higher topological matching degree indicates a greater similarity in structural features between the data source and the query, making it more suitable for meeting query requirements.
[0097] When performing topology-preserving data fusion, this invention employs a progressive fusion strategy to avoid abrupt changes in the topology structure. Specifically, for data sources... and First, their topological features are analyzed to identify and mark key topological structures for protection. Then, a topology-aware data alignment method is designed to ensure that these key structures are preserved during the fusion process. The fusion process can be represented as:
[0098] ,
[0099] in: For the merged dataset, the dimensions and and same; The topology-preserving fusion function takes two data sources and a set of topological features to be preserved as input, and outputs the fused dataset. and There are two input data sources; The set of key topological features that need to be preserved includes 0-dimensional, 1-dimensional, and 2-dimensional durable graphs.
[0100] Data transformation and aggregation operations driven by topological features are core steps in multidimensional data processing. In dimensionality reduction, this invention selects and retains dimensions based on topological persistence, ensuring that key topological structures are preserved in the low-dimensional representation. Specifically, for each dimension $i$, its topological importance score is calculated:
[0101] ,
[0102] in: For dimension The topological importance score indicates that the higher the value, the greater the contribution of that dimension to the topological structure. For dimension The weights are defined as above; For dimension The The persistence of a topological feature indicates the importance of that feature; impact For dimension i, pair of topological features The degree of influence, with a value range of [0,1]; This indicates a summation operation performed on the three dimensions: 0, 1, and 2. This represents a summation operation over all topological features of dimension d. Dimensions with high topological importance scores will be retained first.
[0103] In the data normalization process, this invention employs a normalization method that preserves the topological structure. Traditional Min-Max normalization or Z-score normalization may alter the data's topological structure. The topology-preserving normalization method proposed in this invention ensures that the topological features remain unchanged before and after normalization:
[0104] ,
[0105] in: This is the normalized dataset, with the same dimensions as the original dataset X; This is a topology-preserving normalization function that takes the original dataset and the topological features to be preserved as input and outputs the normalized dataset; X is the original dataset. The set of topological features to be preserved includes 0-dimensional, 1-dimensional, and 2-dimensional durable graphs.
[0106] In anomaly detection, this invention utilizes topological features to identify anomalous data points. Specifically, it calculates the contribution of each data point to the topological structure; points with larger contributions are likely to be anomalous. The anomaly score can be expressed as:
[0107] ,
[0108] in: This is the outlier score for data point x; a higher value indicates that it is more likely to be an outlier. The weights for dimension d are defined as above; The persistence of the j-th topological feature in dimension d; For data point x, pair of topological features The degree of influence, with a value range of [0,1]; This indicates a summation operation performed on the three dimensions: 0, 1, and 2. This represents the summation operation over all topological features of dimension d. Data points with outlier scores exceeding a threshold are marked as anomalies. In financial data analysis scenarios, the threshold is typically set to the average outlier score plus twice the standard deviation, which can capture approximately 2.5% of potential outliers, such as abnormal market fluctuations or investment opportunities.
[0109] Reference Figure 4 The topology-preserving visualization mapping mechanism of the present invention generates a visualization scheme by: establishing a mapping relationship between topological features and visual channels, wherein topological features of different dimensions are mapped to different types of visual representations; performing topological characteristic analysis on templates in the visualization template library and selecting the optimal matching template based on topological similarity; generating visual channel parameter configurations according to the selected templates and topological features; and performing topological consistency verification, evaluating the topological fidelity of the visualization scheme, and improving the visual representation effect of the topological structure through iterative optimization.
[0110] In a preferred embodiment of the present invention, the mapping of topological features of different dimensions to different types of visual representations includes: 0 persistent features are mainly mapped to color, shape and transparency visual channels; 1 persistent features are mainly mapped to spatial location, connectivity and boundary representation visual channels; and 2 persistent features are mainly mapped to hierarchical structure, grouped regions and background processing visual channels.
[0111] This mapping relationship is established based on the principles of visual perception and the semantic meaning of topological features. Specifically, 0-persistent features (connected components) reflect the grouping or category information of the data and are suitable for representation using visual channels that categorize data such as color and shape; 1-persistent features (loops) reflect the cyclic relationship or path in the data and are suitable for representation using relational visual channels such as spatial location and connectivity; 2-persistent features (holes) reflect the spatial structure or hierarchical relationship in the data and are suitable for representation using spatial visual channels such as hierarchical structure and grouped regions.
[0112] In the visualization template library, each template includes its supported topological feature types and expressive capabilities. For a given data topological feature, the system calculates a topological similarity score for each template:
[0113] ,
[0114] in template With data The topological similarity score ranges from [0,1], with higher values indicating higher similarity. For dimension The weights are defined as above; For data Dimensions Persistent graph; template The expressive power function for topological features takes a persistent graph as input and an expressive power score as output, with a value range of [value range missing]. This indicates a summation operation across the 0, 1, and 2 dimensions. The template with the highest topological similarity score is selected as the optimal matching template.
[0115] Based on the selected template and topological features, the system generates visual channel parameter configurations. For example, for the color channel, color grouping and color mapping functions can be set according to the number and importance of persistent features (0); for the spatial location channel, the spatial arrangement of elements can be set according to the structure of persistent features (1); and for the hierarchical structure channel, the hierarchical division can be set according to the nesting relationship of persistent features (2).
[0116] In financial data analysis scenarios, if the data has a clear cluster structure (strong 0-dimensional features), the system may choose a scatter plot or heatmap template and map different clusters to different colors or shapes; if the data has a clear cyclic structure (strong 1-dimensional features), the system may choose a network graph or ring graph template and map the cyclic relationship to connecting lines or a ring layout; if the data has a clear hierarchical structure (strong 2-dimensional features), the system may choose a tree graph or nested graph template and map the hierarchical relationship to branches of a tree or nested regions.
[0117] Reference Figure 5 The topology consistency verification of the present invention includes: calculating the topology features of the visualization result and comparing them with the topology features of the original data view; evaluating the topology fidelity from dimensions such as connectivity retention, loop structure retention, hole retention, persistence correlation and overall structural similarity; and when the topology fidelity is lower than a preset threshold, adjusting the visual channel parameters and regenerating the visualization scheme until a satisfactory topology fidelity is achieved or the maximum number of iterations is reached.
[0118] Topological fidelity is a key metric that measures the degree of similarity in topological structure between the visualization result and the original data. It can be expressed as:
[0119] ,
[0120] in This represents the topology fidelity, with a value ranging from [0,1]. A higher value indicates better preservation of the topology. For dimension The weights are defined as above; Dimensions for visualization results Persistent graph; Dimensions of the original data Persistent graphs; sim is the similarity measure function between persistent graphs; This indicates a summation operation performed on the three dimensions: 0, 1, and 2.
[0121] When evaluating topological fidelity, the system considers multiple dimensions: connectivity retention measures the retention of 0-dimensional topological features, ring structure retention measures the retention of 1-dimensional topological features, hole retention measures the retention of 2-dimensional topological features, persistence relevance measures whether the importance of topological features is visually emphasized, and overall structural similarity measures the degree of similarity of the overall topological structure.
[0122] When the topology fidelity falls below a preset threshold, the system adjusts the visual channel parameters and regenerates the visualization. The preset threshold is typically set to 0.8, meaning at least 80% of the topology information is retained. The parameter adjustment strategy is based on the current topology fidelity assessment results, prioritizing adjustments to the visual channel parameters corresponding to dimensions with lower fidelity. For example, if the ring structure fidelity is low, visual channel parameters related to spatial location and connectivity are adjusted first.
[0123] The iterative optimization process continues until a satisfactory topology fidelity is achieved or the maximum number of iterations is reached. The maximum number of iterations is typically set to 10 to balance optimization effectiveness and computational efficiency. In practice, satisfactory topology fidelity is usually achieved in 3 to 5 iterations in most cases.
[0124] In financial data analysis scenarios, topological consistency verification ensures that visualizations accurately reflect market structure and relationships. For example, if the original data contains obvious market cycles (1-dimensional features), the visualization should clearly show this cyclical structure; if the original data contains obvious market partitions (0-dimensional features), the visualization should clearly distinguish these partitions.
[0125] This invention also includes system performance optimization, including: implementing a multi-level caching mechanism to cache semantic parsing results, topological features, and visualization templates; using a parallel computing framework to accelerate topological feature extraction and analysis; implementing progressive rendering based on topological importance, prioritizing the display of elements with high topological importance; and using hierarchical detail technology to dynamically adjust the rendering details of the topological structure according to the interaction context.
[0126] Multi-level caching is an effective way to improve system response speed. Semantic cache stores the semantic parsing results of queries, which can be reused when similar queries are encountered; topological feature cache stores the calculation results of topological features of the data, avoiding repetitive computationally intensive operations; visualization template cache stores commonly used visualization configurations, accelerating the visualization generation process. The cache update strategy adopts the LRU (Least Recently Used) algorithm to ensure that cache space is used effectively.
[0127] The parallel computing framework is used to accelerate computationally intensive operations such as topology feature extraction and analysis. Specifically, the system decomposes large-scale topology computation tasks into multiple subtasks and distributes them across multiple computing nodes for parallel execution. Furthermore, the system employs simplified topology computation algorithms, such as random sampling and sparse matrix representation, to further improve computational efficiency.
[0128] Progressive rendering based on topological importance is an effective strategy for visualizing large-scale data. The system first calculates the topological importance score for each data element:
[0129] ,
[0130] in For elements The topological importance score indicates that the higher the value, the greater the contribution of the element to the topological structure. For dimension The weights are defined as above; For dimension The The persistence of topological features; For elements Topological features The degree of influence, with a range of values. This indicates a summation operation performed on the three dimensions: 0, 1, and 2. Represents the dimension All topological features are summed. Elements with higher topological importance scores are rendered first, ensuring users see the most important information first.
[0131] Layered Detail (LOD) technology dynamically adjusts the rendering detail of the topology based on the interaction context. When the user performs operations such as zooming or panning, the system adjusts the rendering detail according to the current view range and resolution. For example, in overview mode, the system only renders the main topological structure; while in detail mode, the system renders more detailed information. This adaptive rendering strategy ensures both visual quality and improves rendering efficiency.
[0132] In financial data analytics scenarios, these performance optimization strategies can significantly improve system responsiveness and user experience when analyzing large-scale market data. For example, when analyzing global market data, the system can first display major market trends and key anomalies, and then gradually supplement detailed information, enabling analysts to quickly gain key insights and explore details further when needed.
[0133] The mechanism for optimizing the topology-aware dynamic knowledge graph and the topology-preserving visualization mapping in response to user interaction operations of the present invention includes: collecting user interaction data with visualization results and extracting the implicit topology preferences therein; updating the importance weights of topology features in the knowledge graph according to the topology preferences; adjusting the topology feature extraction parameters and visualization mapping strategy; and pre-calculating and caching relevant topology features and visualization configurations for frequent query patterns to improve system response speed.
[0134] User interaction data is a crucial source of information for system optimization. The system collects user interactions such as clicks, hovering, zooming, and filtering to extract implicit topological preferences. For example, if users frequently focus on certain connected components or loop structures, it indicates that these topological features are important to them; if users frequently adjust certain visual channel parameters, it suggests that the current visualization mapping may not be ideal.
[0135] Based on the extracted topological preferences, the system updates the importance weights of topological features in the knowledge graph. The update formula can be expressed as:
[0136] ,
[0137] in For the updated dimensions Weight; Dimensions before update Weight; Dimensions extracted from user interactions Preference score, with a range of values of The learning rate controls the update speed, and its value ranges from [0,1]. In practice, It is usually set to a value between 0.1 and 0.3 to balance stability and adaptability.
[0138] Simultaneously, the system also adjusts the topology feature extraction parameters and visualization mapping strategies. For example, it adjusts the parameter range in persistent cohomology calculations based on user preferences, or modifies the mapping relationship between visual channels and topology features. These adjustments enable the system to better adapt to users' analytical needs and preferences.
[0139] To address frequent query patterns, the system pre-calculates and caches relevant topological features and visualization configurations, improving system response speed. The system uses a frequency counter and a time decay factor to identify frequent query patterns. For query patterns with frequencies exceeding a threshold, the system pre-calculates their relevant topological features and visualization configurations and stores them in the cache. When a user executes a similar query again, the system can directly retrieve the results from the cache, significantly improving response speed.
[0140] In financial data analysis scenarios, if analysts frequently focus on the cyclical structure of the market, the system will increase the importance weight of the 1D topological features and optimize the visual representation of the loop structure. If analysts frequently perform a certain type of market analysis task, the system will pre-calculate the relevant topological features and visualization configurations, enabling analysts to obtain results and insights more quickly.
[0141] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for implementing a semantic-driven multi-dimensional data visual analysis platform, comprising: Comprising: acquiring a natural language query input by a user; performing semantic parsing based on the natural language query through a topology-aware dynamic knowledge graph, to obtain a semantic parsing result, wherein the topology-aware dynamic knowledge graph represents a topological relationship of a semantic structure by calculating persistent homology features of a semantic space; performing multi-dimensional data fusion processing based on the semantic parsing result, to obtain a processed data view; generating a visualization scheme through a topology-preserved visualization mapping mechanism based on the processed data view, wherein the topology-preserved visualization mapping mechanism ensures that a visualization result retains topological structure features of original data; and rendering a visualization result according to the visualization scheme, and optimizing the topology-aware dynamic knowledge graph and the topology-preserved visualization mapping mechanism in response to user interaction operations.
2. The method of claim 1, wherein, Performing semantic parsing through a topology-aware dynamic knowledge graph comprises: mapping the natural language query to a multi-level semantic vector space to form a semantic point cloud; constructing a simplified complex body on the semantic point cloud, and calculating persistent homology groups of different dimensions to obtain topological features of connectivity, loops and voids; updating nodes and relationships in the knowledge graph through a topology-aware graph neural network based on the topological features, to obtain a semantic representation fused with a topological structure; and determining a query intent and a data requirement according to the semantic representation, to generate the semantic parsing result.
3. The method of claim 2, wherein, The multi-level semantic vector space comprises word-level, phrase-level, sentence-level and concept-level semantic representations, and the natural language query is encoded from different abstraction levels through a hierarchical Transformer architecture, each level corresponding to semantic understanding of different granularity.
4. The method of claim 2, wherein, The topology-aware graph neural network comprises: a topology-enhanced message passing mechanism that integrates topological features into a node representation update process; a multi-scale topological feature integration layer that dynamically adjusts the importance of topological features of different scales according to a query context; and a topology-consistency-preserving update strategy that ensures that key topological structures are maintained during knowledge graph updating.
5. The method of claim 1, wherein, Performing multi-dimensional data fusion processing comprises: identifying data sources related to the query based on the semantic parsing result, and evaluating their topological matching degrees; performing topology-preserved data fusion on the identified data sources, to ensure that key topological structures are maintained during the fusion process; performing topology-feature-driven data conversion and aggregation operations based on the fused data, including dimension reduction based on topological persistence, data normalization that maintains topological structures, and topology-aware anomaly detection; and generating a data view with topological annotations as input for visualization mapping.
6. The method of claim 1, wherein, Generating a visualization scheme through a topology-preserved visualization mapping mechanism comprises: establishing a mapping relationship between topological features and visual channels, wherein topological features of different dimensions are mapped to different types of visual expressions; performing topology characteristic analysis on templates in a visualization template library, and selecting an optimal matching template based on topological similarity; generating visual channel parameter configurations according to the selected template and topological features; and performing topology consistency verification to evaluate the topological fidelity of the visualization scheme, and improving the visual expression effect of the topological structure through iterative optimization.
7. The method of claim 6, wherein, The different dimensional topological feature mappings to different types of visual representations include: 0-dimensional persistent features are mainly mapped to color, shape, and transparency visual channels; 1-dimensional persistent features are mainly mapped to spatial position, connection relationship, and boundary representation visual channels; and 2-dimensional persistent features are mainly mapped to hierarchical structure, grouped regions, and background processing visual channels.
8. The method of claim 6, wherein, The topological consistency verification includes: Calculating the topological features of the visualization results and comparing them with the topological features of the original data view; Evaluating the topological fidelity from the dimensions of connectivity preservation, loop structure preservation, void preservation, persistence correlation, and overall structure similarity; and When the topological fidelity is lower than the preset threshold, adjusting the visual channel parameters and regenerating the visualization scheme until a satisfactory topological fidelity is achieved or the maximum number of iterations is reached.
9. The method of claim 1, wherein, System performance optimization also includes: Implementing a multi-level cache mechanism to cache semantic analysis results, topological features, and visualization templates; Using a parallel computing framework to accelerate topological feature extraction and analysis; Implementing progressive rendering based on topological importance, prioritizing the display of elements with high topological importance; and Using hierarchical detail technology to dynamically adjust the rendering details of the topological structure according to the interaction context.
10. The method of claim 1, wherein, Optimizing the topologically aware dynamic knowledge graph and the topologically preserved visualization mapping mechanism in response to user interaction operations includes: Collecting user interaction data with the visualization results and extracting the topological preferences implied therein; Updating the importance weights of topological features in the knowledge graph according to the topological preferences; Adjusting the topological feature extraction parameters and visualization mapping strategies; and For frequent query patterns, precomputing and caching related topological features and visualization configurations to improve system response speed.