Artificial intelligence data analysis method and system based on machine learning

By using a causal graph model with Riemannian manifold embedding and Ricci curvature perturbation, the problem of causal relationship modeling for multi-source heterogeneous data is solved, achieving high-precision and interpretable analysis of causal relationships, and is suitable for causal analysis and visualization of complex systems.

CN120822418AInactive Publication Date: 2025-10-21XINGHAN LINK TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510961201.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-13
Publication Date
2025-10-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing causal analysis methods suffer from low modeling accuracy and reliability in nested causal relationship modeling of multi-source heterogeneous data due to large differences in data structure, inconsistent distribution, high-dimensional redundancy, heterogeneous features, and nonlinear relationships. Furthermore, they lack structural stability assessment mechanisms, making it difficult to characterize the dynamic linkages between path levels and multi-level causal tensions. Counterfactual simulations also lack interpretability.

Method used

A causal graph model is constructed using a method based on Riemannian manifold embedding, Gromov–Hausdorff measure, and Ricci curvature perturbation. The distance between sample points is calculated through manifold mapping and Gromov–Hausdorff measure. The structure of the causal graph is dynamically adjusted by combining node causal potential energy and Ricci curvature perturbation mechanism, thereby realizing multi-scale causal path-dependent feature extraction and counterfactual inference.

Benefits of technology

It enhances the transparency and credibility of causal analysis, possesses strong causal interpretability, high structural stability, and high accuracy of inference results. It can perform in-depth causal analysis and visualization in complex systems, thereby improving the accuracy and interpretability of causal relationship modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822418A_ABST
    Figure CN120822418A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence data analysis method and system based on machine learning, and the method comprises the following steps: 1, collecting and standardizing original data, and carrying out the embedding processing, and obtaining a data feature vector; 2, mapping the data feature vector into a Riemannian manifold space; 3, calculating the distance between every two sample points by using Gromov-Hausdorff measurement, and establishing an initial causal latent map; 4, constructing a causal potential energy tensor field; 5, obtaining an evolved causal latent map by adopting a Ricci curvature disturbance mechanism; 6, calculating a geodesic distance in the evolved causal latent map as an information propagation path; 7, constructing a causal relationship model; and 8, executing anti-fact reasoning based on the causal relationship model, simulating the causal influence of input variable disturbance on an output result, and performing visual output. According to the method, Riemannian manifold embedding and curvature evolution methods are fused, and intelligent data analysis of causal structure modeling is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis and mining technology, and in particular to an artificial intelligence data analysis method and system based on machine learning. Background Art

[0002] With the growing demand for artificial intelligence and complex system modeling, nested causal relationship modeling for multi-source heterogeneous data and intelligent reasoning technologies driven by high-dimensional data have attracted widespread attention. Existing causal analysis methods mainly rely on traditional graph structure modeling based on the Euclidean space assumption or shallow semantic feature learning for causal inference. However, the following problems are common in practical applications: The collected multi-source data have large structural differences and inconsistent distribution, and there are high-dimensional redundancy, feature heterogeneity and nonlinear relationships, which makes it difficult for the embedded feature representation to maintain the original local geometric structure, affecting the accuracy of modeling and the credibility of causal analysis; the correlation structure between multi-source data shows complex nested and cross-scale dependency characteristics at different scales. Existing causal path modeling technologies mostly use single-scale static graph structures, which are difficult to characterize the dynamic linkage between path levels and multi-level causal tensions; the causal graph construction process lacks a structural stability assessment mechanism and is easily affected by local noise disturbances and edge weight anomalies, resulting in instability of the graph topology and distortion of connection relationships, reducing the reliability of causal reasoning; for counterfactual intervention simulations of input variables, traditional methods mostly use regularized disturbances, ignoring the structural dependencies in the causal path and the potential energy constraints between nodes, resulting in a lack of interpretability of the simulation results and an inability to truly reflect the actual causal effects of input changes on output results.

[0003] Therefore, how to provide an artificial intelligence data analysis method and system based on machine learning is an urgent problem that those skilled in the art need to solve. Summary of the Invention

[0004] One object of the present invention is to propose an artificial intelligence data analysis method and system based on machine learning. The present invention integrates manifold embedding, Gromov–Hausdorff measure and Ricci curvature perturbation methods to construct a causal graph model with continuous geometric structure and dynamic evolution capabilities, realizes causal relationship modeling and counterfactual reasoning of multi-source heterogeneous data, and has the advantages of strong causal interpretability, high structural stability and high accuracy of reasoning results. It is particularly suitable for deep causal analysis and visual expression between variables in complex systems, and improves the transparency and credibility of artificial intelligence data analysis.

[0005] According to an embodiment of the present invention, an artificial intelligence data analysis method based on machine learning includes the following steps: Step 1: Collect and standardize raw data from multiple heterogeneous data sources, and use an encoder to embed the raw data to obtain a data feature vector; Step 2: Map the data feature vector as a sample point to the Riemannian manifold space; Step 3: In the Riemannian manifold space, the distance between each two sample points is calculated using the Gromov–Hausdorff measure, and an initial causal latent graph is established based on the distance. Step 4: In the initial causal potential graph, node causal potential is generated by combining node attributes to construct a causal potential tensor field; Step 5: Based on the causal potential energy tensor field, the initial causal latent graph is topologically evolved using the Ricci curvature perturbation mechanism to obtain an evolved causal latent graph; Step 6: Calculate the geodesic distance in the evolved causal latent graph as the information propagation path; Step 7: Perform nested reasoning on the nodes on the information propagation path, extract multi-scale causal path dependency features, and form a causal relationship model; Step 8: Based on the causal relationship model, perform counterfactual reasoning, simulate the causal impact of input variable disturbance on output results, and perform visual output.

[0006] Optionally, the multi-source heterogeneous data sources include structured data, unstructured text data, image data and time series data; the collection and standardization steps include filling in missing values ​​and normalizing different types of data respectively to generate original data with a unified structure.

[0007] Optionally, the encoder is used to embed the original data to obtain a data feature vector, specifically: the collected and standardized original data is input into the encoder, the encoder includes an input layer, several hidden layers and an embedding output layer, which is used to extract high-dimensional semantic features in the original data; in the encoder, the input original data is transformed layer by layer through a nonlinear activation function to compress the feature dimensions of the original data, and the embedding output layer outputs the data feature vector.

[0008] Optionally, the step three: in the Riemannian manifold space, using the Gromov–Hausdorff measure to calculate the distance between each two sample points, and establishing an initial causal latent graph based on the distance, specifically: Based on the set of sample points mapped to the Riemannian manifold space, constructing a first distance matrix, wherein the first distance matrix records the geodesic distance of each pair of sample points in the Riemannian manifold space; Based on the data eigenvectors in the original feature space, a second distance matrix is ​​constructed, where the second distance matrix records the Euclidean distances of corresponding sample point pairs in the original feature space; For each pair of sample points, the distance values ​​in the first distance matrix and the second distance matrix are compared and calculated using the Gromov–Hausdorff measure. The Gromov–Hausdorff measure is used to quantify the maximum deviation value of the distance structure between two point sets in all metric space correspondences. The Gromov–Hausdorff measurement process includes: finding an optimal corresponding mapping scheme within a set range of sample points so that the maximum distance difference between the mapped first distance matrix element value and the corresponding second distance matrix element value is minimized to form a structural consistency measurement value; Writing the structural consistency measure value into a structural measure matrix to represent the degree of similarity in geometric structure between any two sample points; According to the numerical distribution of the structural measure matrix and the set distance threshold, sample point pairs whose structural consistency measure values ​​are less than the distance threshold are screened, and graph connection edges are established between the sample point pairs to generate an initial causal latent graph.

[0009] Optionally, the step of finding an optimal corresponding mapping solution is as follows: Set an initial point pair mapping scheme to pair each sample point in the first distance matrix with a sample point in the second distance matrix; Based on the initial point pair mapping scheme, calculating the distance differences of all corresponding point pairs in the two matrices to obtain the total distance difference under the current mapping scheme; During the iteration process, new candidate mapping schemes are generated by transforming the point-pair relationships in the current mapping scheme, including exchange, replacement, and local adjustment; For each candidate mapping solution, recalculate the total distance difference and compare it with the total distance difference in the previous iteration. The candidate mapping solution with the smallest total distance difference is retained as the current optimal mapping solution. The mapping scheme iteration process is terminated when any of the following convergence conditions is met: The total distance difference is less than the preset error threshold; The total distance difference change in several consecutive iterations is less than the preset gradient threshold; Reaching the preset maximum number of iterations; Finally, the mapping scheme that meets the convergence conditions is selected as the optimal corresponding mapping scheme.

[0010] Optionally, the fourth step is to generate node causal potential energy in the initial causal potential graph by combining node attributes and constructing a causal potential energy tensor field, specifically: The node attributes include the node's coordinate position in the Riemannian manifold space, the local point density value, and the average geodesic distance to the adjacent nodes. The node attributes are combined and transformed by a nonlinear activation function to output the node causal potential energy. The causal potential of all nodes is organized into a set of scalar data. Based on the topological structure of the initial causal latent graph, the causal potential gradient between nodes is constructed according to the geodesic adjacency relationship between nodes to form a causal potential tensor field, in which each tensor element represents the causal potential difference and directionality between adjacent node pairs.

[0011] Optionally, the Ricci curvature perturbation mechanism is specifically: In the initial causal potential graph, the adjacent nodes connected by each edge are connected by combining the information of the causal potential tensor field. With node To evaluate the stability of the causal relationship between them, the Ricci curvature value of the edge is calculated: ; in, Represents adjacent node pairs The Ricci curvature value, Representation node With node Geodesic distance in Riemannian manifold space, Represented by nodes With node The quality value of the adjacent distribution is normalized and distributed according to the causal potential of each adjacent node, so that the nodes with high causal potential account for a larger proportion in the distribution. It represents the 1-Wasserstein distance calculated according to the definition of adjacency distribution, which is used to characterize the degree of tension between two nodes in the causal neighborhood structure; The calculated Ricci curvature value is compared with the set negative curvature threshold. When the Ricci curvature value of the edge is less than the set negative curvature threshold, it is determined to be in a causal structural tension state and a topological perturbation operation is performed. The topology perturbation operation includes at least one of the following methods: deleting edges; reducing the connection weights of edges; inserting relay nodes to reconstruct the causal path structure; After the topology perturbation operation is completed, the graph structure connection relationship is updated to form an evolved causal latent graph.

[0012] Optionally, the seventh step is to perform nested reasoning on the nodes on the information propagation path, extract multi-scale causal path dependency features, and form a causal relationship model, specifically: Each information propagation path is divided into multiple scale regions according to the path length and local connection density, including: Local scale area: The path length is less than the preset length threshold, and the node hop count is less than the preset hop count threshold, and only the range of the adjacent direct propagation nodes is covered; Mesoscale region: The path length is greater than or equal to the preset length threshold or the node hop count is greater than or equal to the preset hop count threshold, and the propagation range is limited to a single connected subgraph; Global scale region: the path length is greater than or equal to the preset length threshold and the node hop count is greater than or equal to the preset hop count threshold, and the path spans multiple connected subgraphs; For each propagation path segment within each scale region, a nested causal inference module is constructed. The nested inference module is implemented using a multi-layer causal representation network structure, where each layer is responsible for modeling the conditional causal relationship of the output nodes of the previous layer. In the nested causal inference module, the causal state of each node in the path is used as context input and updated layer by layer through gated recursive units to capture multi-scale causal path dependency features; The multi-scale causal path dependency features output by each scale region are subjected to feature splicing and attention aggregation to form a causal relationship model, which is used to characterize the causal path linkage effect in different scale regions.

[0013] Optionally, the eighth step is to perform counterfactual reasoning based on the causal relationship model, simulate the causal impact of the input variable disturbance on the output result, and perform visual output, specifically: Apply perturbations to one variable in the input variable set, including value replacement, distribution sampling shift, and logical state flipping, to simulate non-realistic value scenarios of the variable under counterfactual conditions; The perturbed variable vector is input as the counterfactual input vector into the constructed causal relationship model, and the response result of the output variable is obtained by combining the structural dependency in the causal path; Repeat multiple rounds of perturbations, independently perturb each input variable, record the change value of the output variable under each set of perturbations, form a disturbance-response mapping relationship, and visualize the disturbance-response mapping relationship.

[0014] An artificial intelligence data analysis method based on machine learning according to an embodiment of the present invention includes the following modules: Data preprocessing module, used to collect and standardize raw data from multiple heterogeneous data sources; A feature embedding module is used to embed the original data using an encoder to obtain a data feature vector; A manifold mapping module is used to map the data feature vector to a Riemannian manifold space to construct a set of sample points with a local geometric structure; The causal graph construction module is used to calculate the structural differences between sample points based on the Gromov–Hausdorff measure and establish the initial causal latent graph; A potential energy calculation module is used to generate node causal potential energy in the initial causal potential graph by combining node attributes and constructing a causal potential energy tensor field; A graph evolution module, configured to perform topological evolution on the initial causal latent graph based on the causal potential energy tensor field using a Ricci curvature perturbation mechanism to obtain an evolved causal latent graph; The causal modeling module is used to extract the information propagation path in the evolved causal latent graph and perform multi-scale nested reasoning to form a causal relationship model; The reasoning and visualization module is used to perform counterfactual reasoning based on the causal relationship model, simulate the causal impact of input variable disturbances on output results, and perform visual output.

[0015] The beneficial effects of the present invention are: The present invention integrates non-Euclidean embedding and Ricci curvature-driven causal graph topological evolution to address the problems of inconsistent feature space structure, insufficient causal path representation, and poor interpretability of counterfactual reasoning in multi-source heterogeneous data. Riemannian manifold mapping and the Gromov–Hausdorff measure are used to construct a causal latent graph that maintains local geometric consistency. Combined with structural tension modeling based on node potential fields, the connection structure of edges in the graph is dynamically adjusted through the Ricci curvature perturbation mechanism to achieve adaptive optimization of the causal graph in structurally unstable areas. In the causal modeling link, a multi-scale information propagation path partitioning strategy is introduced, and a nested causal network is constructed in combination with a gated recursive reasoning unit. The conditional dependencies in the local to global path are extracted layer by layer, effectively improving the modeling capability of cross-scale causal dependencies. In the counterfactual reasoning process, a perturbation simulation framework based on causal structural constraints is constructed to perform multiple types of perturbations such as distribution shift and state flipping on the input variables. The response changes of the results are simulated while maintaining causal connectivity. The sensitivity map of the output variables is analyzed through multiple rounds of perturbation-response analysis to achieve a quantitative expression of the causal contribution of different input changes to the output results in a complex variable system. Ultimately, an integrated artificial intelligence causal analysis system is formed with a stable causal structure, clear path hierarchy, and semantically interpretable reasoning results. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is an overall flow chart of an artificial intelligence data analysis method based on machine learning proposed by the present invention; Figure 2This is a structural diagram of an artificial intelligence data analysis system based on machine learning proposed by the present invention. DETAILED DESCRIPTION

[0017] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0018] refer to Figure 1 , an artificial intelligence data analysis method based on machine learning, including: Step 1: Collect and standardize raw data from multiple heterogeneous data sources, and use an encoder to embed the raw data to obtain a data feature vector; Step 2: Map the data feature vector as a sample point to the Riemannian manifold space; Step 3: In the Riemannian manifold space, the distance between each two sample points is calculated using the Gromov–Hausdorff measure, and an initial causal latent graph is established based on the distance. Step 4: In the initial causal potential graph, node causal potential is generated by combining node attributes to construct a causal potential tensor field; Step 5: Based on the causal potential energy tensor field, the initial causal latent graph is topologically evolved using the Ricci curvature perturbation mechanism to obtain an evolved causal latent graph; Step 6: Calculate the geodesic distance in the evolved causal latent graph as the information propagation path; Step 7: Perform nested reasoning on the nodes on the information propagation path, extract multi-scale causal path dependency features, and form a causal relationship model; Step 8: Based on the causal relationship model, perform counterfactual reasoning, simulate the causal impact of input variable disturbance on output results, and perform visual output.

[0019] This invention effectively improves the ability to extract causal relationships from multi-source heterogeneous data. In particular, the introduction of Riemannian manifold space and the Ricci curvature perturbation mechanism makes the graph structure no longer static, but rather capable of dynamically adapting to changes in data distribution, enhancing modeling flexibility and structural robustness. Furthermore, the combination of geodesic extraction of information propagation paths and multi-scale nested reasoning makes causal paths not only hierarchical but also context-dependent, providing a stable foundation for counterfactual simulations. This improves the interpretability of data analysis results, the accuracy of counterfactual simulations, and the generalization performance of inference models, demonstrating practical value in complex causal scenarios.

[0020] In this embodiment, the multi-source heterogeneous data sources include structured data, unstructured text data, image data and time series data; the collection and standardization steps include filling missing values ​​and normalizing different types of data to generate original data with a unified structure.

[0021] In this embodiment, the encoder is used to embed the original data to obtain a data feature vector. Specifically, the collected and standardized original data is input into the encoder, and the encoder includes an input layer, several hidden layers and an embedding output layer for extracting high-dimensional semantic features in the original data; in the encoder, the input original data is transformed layer by layer through a nonlinear activation function to compress the feature dimensions of the original data, and the embedding output layer outputs the data feature vector.

[0022] The proposed embedding processing scheme, based on a neural network encoder structure, effectively compresses raw high-dimensional data and extracts deep semantic features. By constructing a network architecture consisting of an input layer, hidden layers, and an embedded output layer, and using nonlinear activation functions for multi-layer abstraction, the original data is significantly reduced in dimensionality while retaining key information. This improves the efficiency and quality of manifold mapping and graph construction. It not only has good versatility for data of different modalities, but also supports the compact representation of high-dimensional sparse information, significantly reducing computational complexity. The embedded feature vectors can be used as high-quality causal analysis input, improving model stability and interpretability.

[0023] In this embodiment, the step three is: in the Riemannian manifold space, the distance between each two sample points is calculated using the Gromov–Hausdorff measure, and an initial causal latent graph is established based on the distance, specifically: Based on the set of sample points mapped to the Riemannian manifold space, constructing a first distance matrix, wherein the first distance matrix records the geodesic distance of each pair of sample points in the Riemannian manifold space; Based on the data eigenvectors in the original feature space, a second distance matrix is ​​constructed, where the second distance matrix records the Euclidean distances of corresponding sample point pairs in the original feature space; For each pair of sample points, the distance values ​​in the first distance matrix and the second distance matrix are compared and calculated using the Gromov–Hausdorff measure. The Gromov–Hausdorff measure is used to quantify the maximum deviation value of the distance structure between two point sets in all metric space correspondences. The Gromov–Hausdorff measurement process includes: finding an optimal corresponding mapping scheme within a set range of sample points so that the maximum distance difference between the mapped first distance matrix element value and the corresponding second distance matrix element value is minimized to form a structural consistency measurement value; Writing the structural consistency measure value into a structural measure matrix to represent the degree of similarity in geometric structure between any two sample points; According to the numerical distribution of the structural measure matrix and the set distance threshold, sample point pairs whose structural consistency measure values ​​are less than the distance threshold are screened, and graph connection edges are established between the sample point pairs to generate an initial causal latent graph.

[0024] Constructing a causal latent graph using the Gromov–Hausdorff measure effectively addresses the challenge of measuring sample structural consistency in manifold space. Using a dual-distance matrix comparison method, combining distance information between the original Euclidean space and the embedded Riemannian space, the method measures the structural similarity between pairs of sample points. This method then removes spurious connections while retaining node-edge connections with true geometric relationships. This avoids the high-dimensional distortions and connectivity errors common in traditional graph construction methods, enhancing the true representation of the graph structure. By setting a distance threshold to filter connected edges, the method ensures a more sparse and reasonable graph topology, effectively improving the accuracy of causal relationship mining and providing a stable foundation for subsequent causal potential modeling and graph evolution.

[0025] In this embodiment, the search for an optimal corresponding mapping solution is specifically as follows: Set an initial point pair mapping scheme to pair each sample point in the first distance matrix with a sample point in the second distance matrix; Based on the initial point pair mapping scheme, calculating the distance differences of all corresponding point pairs in the two matrices to obtain the total distance difference under the current mapping scheme; During the iteration process, new candidate mapping schemes are generated by transforming the point-pair relationships in the current mapping scheme, including exchange, replacement, and local adjustment; For each candidate mapping solution, recalculate the total distance difference and compare it with the total distance difference in the previous iteration. The candidate mapping solution with the smallest total distance difference is retained as the current optimal mapping solution. The mapping scheme iteration process is terminated when any of the following convergence conditions is met: The total distance difference is less than the preset error threshold; The total distance difference change in several consecutive iterations is less than the preset gradient threshold; Reaching the preset maximum number of iterations; Finally, the mapping scheme that meets the convergence conditions is selected as the optimal corresponding mapping scheme.

[0026] An iterative optimization-based matching mechanism is employed to determine the optimal corresponding mapping scheme, significantly improving the accuracy of finding structurally consistent correspondences between different distance matrices. Through initial mapping settings, iterative optimization of distance errors, and dynamic scheme screening, this process effectively converges to a minimum global error state, ensuring good stability and convergence of the measurement process. The set convergence conditions (error threshold, gradient threshold, and maximum number of iterations) provide a reliable basis for termination of the entire process, avoiding local optimality or over-computation. The overall mechanism is characterized by strong generalization, high efficiency, and good adjustability, making it suitable for matching measurement tasks for large-scale samples and supporting the stable construction of initial causal graphs.

[0027] In this embodiment, the fourth step is to generate node causal potential energy in the initial causal potential graph by combining node attributes and constructing a causal potential energy tensor field, specifically: The node attributes include the node's coordinate position in the Riemannian manifold space, the local point density value, and the average geodesic distance to the adjacent nodes. The node attributes are combined and transformed through a nonlinear activation function to output the node causal potential energy: ; in, Representation node The causal potential of Representation node The coordinate position in the Riemannian manifold space, Representation node The local point density value, Representation node The mean geodesic distance to adjacent nodes, and represents the weight matrix, and represents the bias term, represents a nonlinear activation function; The causal potential of all nodes is organized into a set of scalar data. Based on the topological structure of the initial causal latent graph, the causal potential gradient between nodes is constructed according to the geodesic adjacency relationship between nodes to form a causal potential tensor field, in which each tensor element represents the causal potential difference and directionality between adjacent node pairs.

[0028] This step introduces the concept of node causal potential energy. By combining the node's coordinates in Riemann space, local point density, and average geodesic distance to adjacent nodes, a multi-dimensional attribute combination is formed, and the node causal potential energy is generated using nonlinear activation function mapping. By constructing a causal potential energy tensor field, the directional gradient information of the causal potential energy is captured on the graph topology structure, providing quantitative support for topological adjustment. Compared with the traditional unweighted modeling method based on the graph structure itself, the present invention can give nodes stronger physical interpretability and dynamic adjustment capabilities, while improving the ability to express structural heterogeneity in the graph, laying a foundation for clear physical meaning and numerical stability for the curvature-based evolution mechanism.

[0029] In this embodiment, the Ricci curvature perturbation mechanism is specifically as follows: In the initial causal potential graph, the adjacent nodes connected by each edge are connected by combining the information of the causal potential tensor field. With node To evaluate the stability of the causal relationship between them, the Ricci curvature value of the edge is calculated: ; in, Represents adjacent node pairs The Ricci curvature value, Representation node With node Geodesic distance in Riemannian manifold space, Represented by nodes With node The quality value of the adjacent distribution is normalized and distributed according to the causal potential of each adjacent node, so that the nodes with high causal potential account for a larger proportion in the distribution. It represents the 1-Wasserstein distance calculated according to the definition of adjacency distribution, which is used to characterize the degree of tension between two nodes in the causal neighborhood structure; The calculated Ricci curvature value is compared with the set negative curvature threshold. When the Ricci curvature value of the edge is less than the set negative curvature threshold, it is determined to be in a causal structural tension state and a topological perturbation operation is performed. The topology perturbation operation includes at least one of the following methods: deleting edges; reducing the connection weights of edges; inserting relay nodes to reconstruct the causal path structure; After the topology perturbation operation is completed, the graph structure connection relationship is updated to form an evolved causal latent graph.

[0030] This step introduces a Ricci curvature perturbation mechanism during graph evolution. By evaluating the stability of each edge under causal tension, it achieves adaptive topological adjustment of the causal graph. By constructing an adjacency probability distribution based on the normalized distribution of causal potential energy and calculating the 1-Wasserstein distance, the Ricci curvature value, which reflects the stability of edge connections, is derived. When an edge is in a negative curvature state, indicating structural tension or conflict in the causal path, operations such as edge weight reduction, connection disconnection, or the insertion of relay nodes are automatically performed to alleviate structural conflicts and improve information transmission stability. The Ricci curvature perturbation mechanism is dynamic and locally sensitive, automatically reconstructing the causal topology based on the evolution of the graph structure, effectively improving the adaptability and robustness of the causal model to abnormal patterns or mutated data structures.

[0031] In this embodiment, the seventh step is to perform nested reasoning on the nodes on the information propagation path, extract multi-scale causal path dependency features, and form a causal relationship model, specifically: Each information propagation path is divided into multiple scale regions according to the path length and local connection density, including: Local scale area: The path length is less than the preset length threshold, and the node hop count is less than the preset hop count threshold, and only the range of the adjacent direct propagation nodes is covered; Mesoscale region: The path length is greater than or equal to the preset length threshold or the node hop count is greater than or equal to the preset hop count threshold, and the propagation range is limited to a single connected subgraph; Global scale region: the path length is greater than or equal to the preset length threshold and the node hop count is greater than or equal to the preset hop count threshold, and the path spans multiple connected subgraphs; For each propagation path segment within each scale region, a nested causal reasoning module is constructed. The nested reasoning module is implemented using a multi-layer causal representation network structure, where each layer is responsible for modeling the conditional causal relationship of the output nodes of the previous layer. In the nested causal inference module, the causal state of each node in the path is used as context input and updated layer by layer through gated recursive units to capture multi-scale causal path dependency features; The multi-scale causal path dependency features output by each scale region are subjected to feature splicing and attention aggregation to form a causal relationship model, which is used to characterize the causal path linkage effect in different scale regions.

[0032] This step builds a multi-scale nested inference structure to fully exploit dependencies at different levels in the causal path, improving the causal model's ability to express complex chain structures. Based on path length and local connectivity density, the model divides the inference into three regions: local, mesoscale, and global. Each region is independently modeled using a nested causal network. State changes in the causal path are dynamically captured using gated recursive units. In the final stage, feature splicing and attention aggregation are performed to output a unified causal representation. This addresses the single-scale and weakly coupled path problems of traditional causal models, improves the ability to extract key causal links in multi-level causal chains, and offers significant advantages for real-world data with complex causal structures.

[0033] In this embodiment, the eighth step is to perform counterfactual reasoning based on the causal relationship model, simulate the causal impact of the input variable disturbance on the output result, and perform visual output, specifically: Apply perturbations to one variable in the input variable set, including value replacement, distribution sampling shift, and logical state flipping, to simulate non-realistic value scenarios of the variable under counterfactual conditions; The perturbed variable vector is input as the counterfactual input vector into the constructed causal relationship model, and the response result of the output variable is obtained by combining the structural dependency in the causal path; Repeat multiple rounds of perturbations, independently perturb each input variable, record the change value of the output variable under each set of perturbations, form a disturbance-response mapping relationship, and visualize the disturbance-response mapping relationship.

[0034] This step introduces a counterfactual reasoning mechanism that can simulate the causal impact of different input variable disturbances on the output results, effectively improving the model's interpretability and decision-making guidance capabilities. By performing non-real disturbances on variables through numerical replacement, distribution sampling, or logical reversal, and combining existing causal models for response reasoning, a disturbance-response mapping is formed. This not only identifies the sensitivity of key dependent variables to changes in results, but also supports output prediction under counterfactual conditions, enabling quantitative analysis of "what would happen if a certain variable were different." Through the visual output of results, it provides an intuitive basis for intervention control, causal discovery, and strategy optimization of complex systems. It is an important analytical tool for high-demand scenarios such as intelligent diagnosis, financial analysis, and policy simulation.

[0035] refer to Figure 2 , an artificial intelligence data analysis system based on machine learning, including the following modules: Data preprocessing module, used to collect and standardize raw data from multiple heterogeneous data sources; A feature embedding module is used to embed the original data using an encoder to obtain a data feature vector; A manifold mapping module is used to map the data feature vector to a Riemannian manifold space to construct a set of sample points with a local geometric structure; The causal graph construction module is used to calculate the structural differences between sample points based on the Gromov–Hausdorff measure and establish the initial causal latent graph; A potential energy calculation module is used to generate node causal potential energy in the initial causal potential graph by combining node attributes and constructing a causal potential energy tensor field; A graph evolution module, configured to perform topological evolution on the initial causal latent graph based on the causal potential energy tensor field using a Ricci curvature perturbation mechanism to obtain an evolved causal latent graph; The causal modeling module is used to extract the information propagation path in the evolved causal latent graph and perform multi-scale nested reasoning to form a causal relationship model; The reasoning and visualization module is used to perform counterfactual reasoning based on the causal relationship model, simulate the causal impact of input variable disturbances on output results, and perform visual output.

[0036] This system forms a complete closed-loop analysis system from raw data to causal decision-making, achieving high coupling and data flow between modules, and supporting automated modeling of the entire process from multi-source data to causal graphs. The system has a clear structure, good modular scalability, and easy engineering deployment, facilitating its application in practical fields such as healthcare, finance, and industry. Through systematic design, it effectively reduces the cost of multi-stage manual parameter adjustment while improving the generation efficiency and stability of causal analysis models, providing solid technical support for building explainable artificial intelligence systems.

[0037] Example 1: To verify the feasibility of this invention, we applied it to equipment operation monitoring and fault diagnosis in a large-scale intelligent manufacturing workshop. Specifically, we addressed the causal analysis modeling problem of the high-dimensional, multi-source, and heterogeneous operational data generated by multiple automated CNC machine tools during long-term operation. The goal was to accurately identify the potential causal paths leading to equipment failures, considering the interplay of multiple factors, and to conduct counterfactual simulations and predictions, thereby enabling early warning and optimizing maintenance decisions.

[0038] In practical applications, traditional equipment condition monitoring relies on fixed threshold judgments or linear regression predictions, which fail to fully explore the potential causal structure between sensors and explain how complex environmental changes affect equipment failure rates. This is especially true when faced with issues such as abnormal spindle temperature, servo current fluctuations, and machining error fluctuations. Conventional models are weak in fusing multi-source data and lack the ability to interpret causal relationships. This results in a high rate of missed alerts and an even higher rate of false alarms in early warning systems.

[0039] In this embodiment, the system accesses five types of core equipment data sources, including: temperature sensors (spindle / cooling / ambient temperature), vibration accelerometers, servo current and voltage records, tool wear images (RGB), and machining error indicators.

[0040] These data come from seven intelligent processing equipment that are running continuously. The monitoring period is 60 consecutive days. The sampling frequency of each device is 1Hz. The total data volume is approximately 2.5TB, covering structured time series data, image data and a small amount of text labels (operator records).

[0041] During data preprocessing, the system uses multi-channel interpolation, resampling, and missing value filling to align the five signal types to a 1Hz frequency and standardize them to zero mean and unit variance. During feature embedding, a fusion encoder network is used to automatically extract feature representations for each data type, outputting a unified 128-dimensional feature vector space.

[0042] The manifold mapping module then maps this high-dimensional feature space onto a continuously differentiable Riemannian manifold, preserving local geometric adjacency. During the construction phase, the Gromov–Hausdorff measure is applied to minimize the structural difference between the original space and the manifold space, generating an initial causal graph that is both structurally conformal and geometrically consistent. The system then constructs a tensor field based on the node causal potentials, employing the Ricci curvature perturbation mechanism to determine the tension state of causal relationships within the graph and dynamically evolve the topology, thereby improving the stability of causal expression.

[0043] Table 1 Comparison of key indicators of equipment failure prediction between traditional methods and the present invention Table 1 compares the key performance indicators of four typical equipment condition monitoring and fault prediction technologies: a traditional threshold-based alarm method, a machine learning method based on feature engineering, a prediction method based on a deep attention mechanism, and the proposed method that integrates the Gromov–Hausdorff measure and the Ricci curvature perturbation mechanism. In terms of accuracy, the traditional method achieved only 61.3%, while the proposed method reached 89.4%, an improvement of nearly 28 percentage points. The false negative rate dropped significantly from 27.5% in the traditional method to 7.6%, demonstrating that the system has a stronger ability to capture key causal paths and significantly reduces the risk of missed fault detection. Furthermore, the average lead time for warning increased from 3.7 minutes to 12.4 minutes, meaning that the system can issue reliable warnings within a longer time window before a fault occurs, ensuring sufficient response time for equipment maintenance. Furthermore, the model stability score shows that the causal latent graph constructed in this invention exhibits high robustness in the face of complex perturbations and data drift, due to its curvature-driven structural self-adjustment capabilities.

[0044] After completing causal modeling, the system also conducted counterfactual reasoning tests on multiple key variables. For example, if the rise rate of "spindle temperature" slowed down within 10 minutes, or the instantaneous change of "servo current" weakened, would the probability of failure be reduced? Table 2 Counterfactual simulation results after variable disturbance Table 2 above demonstrates the effectiveness of the causal model in its "if...then..." reasoning capabilities by simulating counterfactual perturbations of multiple key input variables. The simulations selected the spindle temperature rise rate, current fluctuation frequency, ambient temperature difference amplitude, and edge pixel values ​​of the tool wear image as input variables. While keeping other conditions constant, each variable was individually perturbed and the resulting output failure probability was analyzed. For example, when the spindle temperature rise rate was reduced from the original 1.6°C / min to 0.9°C / min, the model's output failure probability decreased from 67.2% to 39.5%, a 27.7% reduction. This indicates that this variable significantly impacts failure risk under the current operating conditions. This hypothetical intervention capability is due to the causal model's accurate modeling of multi-scale path dependency characteristics and the expressiveness of the nested causal reasoning module, directly reflecting the technical contributions of claims 7 and 8. Furthermore, the system supports continuous variable perturbations and graphically outputs the perturbation-response mapping, enabling users to intuitively identify high-risk variables and their intervention space, making it highly interpretable and operational in industrial applications.

[0045] This embodiment verifies the comprehensive performance advantages of the technology in multi-source heterogeneous data processing, causal relationship modeling, and predictive reasoning by applying the method of the present invention to the scenario of complex industrial equipment operation status monitoring. By embedding the encoder to extract deep semantic features, the high-dimensional features are mapped to the Riemannian manifold space, and then the structural consistency-driven causal latent graph is constructed in combination with the Gromov–Hausdorff measure. The dynamic evolution of the graph structure is achieved based on the Ricci curvature perturbation mechanism, which effectively enhances the system's ability to identify potential causal paths. Actual simulation data show that in terms of key indicators such as prediction accuracy, warning lead time, variable sensitivity, and model stability, the present invention is superior to existing mainstream methods, fully demonstrating the practicality and advancement of the fusion analysis of causal modeling and geometric structure. The system not only improves the prediction performance, but also has good interpretability and operational flexibility, providing strong support for intelligent decision-making in complex industrial systems.

[0046] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. An artificial intelligence data analysis method based on machine learning, characterized in that: include: Step 1: Collect and standardize raw data from multiple heterogeneous data sources, and use an encoder to embed the raw data to obtain a data feature vector; Step 2: Map the data feature vector as a sample point to the Riemannian manifold space; Step 3: In the Riemannian manifold space, the distance between each two sample points is calculated using the Gromov–Hausdorff measure, and an initial causal latent graph is established based on the distance. Step 4: In the initial causal potential graph, node causal potential is generated by combining node attributes to construct a causal potential tensor field; Step 5: Based on the causal potential energy tensor field, the initial causal latent graph is topologically evolved using the Ricci curvature perturbation mechanism to obtain an evolved causal latent graph; Step 6: Calculate the geodesic distance in the evolved causal latent graph as the information propagation path; Step 7: Perform nested reasoning on the nodes on the information propagation path, extract multi-scale causal path dependency features, and form a causal relationship model; Step 8: Based on the causal relationship model, perform counterfactual reasoning, simulate the causal impact of input variable disturbance on output results, and perform visual output.

2. The artificial intelligence data analysis method based on machine learning according to claim 1, characterized in that: The multi-source heterogeneous data sources include structured data, unstructured text data, image data and time series data; the collection and standardization steps include filling missing values ​​and normalizing different types of data to generate original data with a unified structure.

3. The artificial intelligence data analysis method based on machine learning according to claim 1, characterized in that: The encoder is used to embed the original data to obtain a data feature vector, specifically: the collected and standardized original data is input into the encoder, the encoder includes an input layer, several hidden layers and an embedding output layer, and is used to extract high-dimensional semantic features in the original data; in the encoder, the input original data is transformed layer by layer through a nonlinear activation function to compress the feature dimensions of the original data, and the embedding output layer outputs the data feature vector.

4. The artificial intelligence data analysis method based on machine learning according to claim 1, characterized in that: Step 3: In the Riemannian manifold space, the distance between each two sample points is calculated using the Gromov–Hausdorff measure, and an initial causal latent graph is established based on the distance, specifically: Based on the set of sample points mapped to the Riemannian manifold space, constructing a first distance matrix, wherein the first distance matrix records the geodesic distance of each pair of sample points in the Riemannian manifold space; Based on the data eigenvectors in the original feature space, a second distance matrix is ​​constructed, where the second distance matrix records the Euclidean distances of corresponding sample point pairs in the original feature space; For each pair of sample points, the distance values ​​in the first distance matrix and the second distance matrix are compared and calculated using the Gromov–Hausdorff measure. The Gromov–Hausdorff measure is used to quantify the maximum deviation value of the distance structure between two point sets in all metric space correspondences. The Gromov–Hausdorff measurement process includes: finding an optimal corresponding mapping scheme within a set range of sample points so that the maximum distance difference between the mapped first distance matrix element value and the corresponding second distance matrix element value is minimized to form a structural consistency measurement value; Writing the structural consistency measurement value into the structural measurement matrix to represent the similarity between any two sample points in geometric structure; According to the numerical distribution of the structural measure matrix and the set distance threshold, sample point pairs whose structural consistency measure values ​​are less than the distance threshold are screened, and graph connection edges are established between the sample point pairs to generate an initial causal latent graph.

5. The artificial intelligence data analysis method based on machine learning according to claim 4, characterized in that: The specific steps of finding an optimal corresponding mapping solution are as follows: Set an initial point pair mapping scheme to pair each sample point in the first distance matrix with a sample point in the second distance matrix; Based on the initial point pair mapping scheme, calculating the distance differences of all corresponding point pairs in the two matrices to obtain the total distance difference under the current mapping scheme; In the iterative process, new candidate mapping schemes are generated by transforming the point-pair relationships in the current mapping scheme, including exchange, replacement, and local adjustment; For each candidate mapping solution, recalculate the total distance difference and compare it with the total distance difference in the previous iteration. The candidate mapping solution with the smallest total distance difference is retained as the current optimal mapping solution. The mapping scheme iteration process is terminated when any of the following convergence conditions is met: The total distance difference is less than the preset error threshold; The total distance difference change in several consecutive iterations is less than the preset gradient threshold; Reaching the preset maximum number of iterations; Finally, the mapping scheme that meets the convergence conditions is selected as the optimal corresponding mapping scheme.

6. The artificial intelligence data analysis method based on machine learning according to claim 1, characterized in that: The fourth step is to generate node causal potential energy in the initial causal potential graph by combining node attributes and constructing a causal potential energy tensor field, specifically: The node attributes include the node's coordinate position in the Riemannian manifold space, the local point density value, and the average geodesic distance to the adjacent nodes. The node attributes are combined and transformed by a nonlinear activation function to output the node causal potential energy. The causal potential of all nodes is organized into a set of scalar data. Based on the topological structure of the initial causal latent graph, the causal potential gradient between nodes is constructed according to the geodesic adjacency relationship between nodes to form a causal potential tensor field, in which each tensor element represents the causal potential difference and directionality between adjacent node pairs.

7. The artificial intelligence data analysis method based on machine learning according to claim 1, characterized in that: The Ricci curvature perturbation mechanism is specifically: In the initial causal potential graph, the adjacent nodes connected by each edge are connected by combining the information of the causal potential tensor field. With node To evaluate the stability of the causal relationship between them, the Ricci curvature value of the edge is calculated: ; in, Represents adjacent node pairs The Ricci curvature value, Representation node With node Geodesic distance in Riemannian manifold space, Represented by nodes With node The quality value of the adjacent distribution is normalized and distributed according to the causal potential of each adjacent node. represents the 1-Wasserstein distance calculated according to the definition of the neighbor distribution; The calculated Ricci curvature value is compared with the set negative curvature threshold. When the Ricci curvature value of the edge is less than the set negative curvature threshold, it is determined to be in a causal structural tension state and a topological perturbation operation is performed. The topology perturbation operation includes at least one of the following methods: deleting edges; reducing the connection weights of edges; inserting relay nodes to reconstruct the causal path structure; After the topology perturbation operation is completed, the graph structure connection relationship is updated to form an evolved causal latent graph.

8. The artificial intelligence data analysis method based on machine learning according to claim 1, characterized in that: Step 7: Performing nested reasoning on the nodes on the information propagation path, extracting multi-scale causal path dependency features, and forming a causal relationship model, specifically: Each information propagation path is divided into multiple scale regions according to the path length and local connection density, including: Local scale area: The path length is less than the preset length threshold, and the node hop count is less than the preset hop count threshold, and only the range of the adjacent direct propagation nodes is covered; Mesoscale region: The path length is greater than or equal to the preset length threshold or the node hop count is greater than or equal to the preset hop count threshold, and the propagation range is limited to a single connected subgraph; Global scale region: the path length is greater than or equal to the preset length threshold and the node hop count is greater than or equal to the preset hop count threshold, and the path spans multiple connected subgraphs; For each propagation path segment within each scale region, a nested causal inference module is constructed. The nested inference module is implemented using a multi-layer causal representation network structure, where each layer is responsible for modeling the conditional causal relationship of the output nodes of the previous layer. In the nested causal inference module, the causal state of each node in the path is used as context input and updated layer by layer through gated recursive units to capture multi-scale causal path dependency features; The multi-scale causal path dependency features output by each scale region are subjected to feature splicing and attention aggregation to form a causal relationship model, which is used to characterize the causal path linkage effect in different scale regions.

9. The artificial intelligence data analysis method based on machine learning according to claim 1, characterized in that: Step 8: Based on the causal relationship model, counterfactual reasoning is performed to simulate the causal impact of the input variable disturbance on the output result, and visual output is performed, specifically: Apply perturbations to one variable in the input variable set, including value replacement, distribution sampling shift, and logical state flipping, to simulate non-realistic value scenarios of the variable under counterfactual conditions; The perturbed variable vector is input as the counterfactual input vector into the constructed causal relationship model, and the response result of the output variable is obtained by combining the structural dependency in the causal path; Repeat multiple rounds of perturbations, independently perturb each input variable, record the change value of the output variable under each set of perturbations, form a disturbance-response mapping relationship, and visualize the disturbance-response mapping relationship.

10. The artificial intelligence data analysis system based on machine learning according to claim 1, executing the artificial intelligence data analysis method based on machine learning according to any one of claims 1 to 9, characterized in that: Includes the following modules: Data preprocessing module, used to collect and standardize raw data from multiple heterogeneous data sources; A feature embedding module is used to embed the original data using an encoder to obtain a data feature vector; A manifold mapping module is used to map the data feature vector to a Riemannian manifold space to construct a set of sample points with a local geometric structure; The causal graph construction module is used to calculate the structural differences between sample points based on the Gromov–Hausdorff measure and establish the initial causal latent graph; A potential energy calculation module is used to generate node causal potential energy in the initial causal potential graph by combining node attributes and constructing a causal potential energy tensor field; A graph evolution module, configured to perform topological evolution on the initial causal latent graph based on the causal potential energy tensor field using a Ricci curvature perturbation mechanism to obtain an evolved causal latent graph; The causal modeling module is used to extract the information propagation path in the evolved causal latent graph and perform multi-scale nested reasoning to form a causal relationship model; The reasoning and visualization module is used to perform counterfactual reasoning based on the causal relationship model, simulate the causal impact of input variable disturbances on output results, and perform visual output.

Citation Information

Cited By

  • Intelligent sensing wireless network dynamic coverage optimization system

    CN121174189A

  • Vein thrombosis risk assessment method based on large language model

    CN121281734A

  • Satellite image water resource dynamic monitoring method and system based on deep learning

    CN121582793A

  • Multi-scale coal seam water spatial variation analysis and main control factor identification system oriented to group mine combined mining

    CN121659838A

  • Adaptive control parameter optimization method for multi-field industrial control system

    CN121721964A