Enterprise big data mining method and system based on artificial intelligence
Through multimodal hypergraph neural network and dynamic game adversarial interpretation framework, the integration and analysis problems of multi-source heterogeneous data are solved, and an efficient and accurate enterprise-level intelligent decision-making map is generated, which improves the effectiveness and accuracy of data mining.
Patent Information
- Application Number
- CN202510741713.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-05
AI Technical Summary
Traditional data mining methods are difficult to effectively integrate and analyze multi-source heterogeneous data, resulting in insufficient accuracy and effectiveness of mining results, especially poor performance when dealing with the complexity and dynamic changes of cross-modal features.
Multimodal hypergraph neural network is used for dynamic modal alignment, and a dynamic hypergraph construction algorithm driven by self-attention is fused across modal features, combined with orthogonal adversarial manifold learning and dynamic game adversarial interpretation framework, a multimodal joint embedding tensor with space-consistent multimodal joint embedding tensor is generated, and causal correlation mining is carried out, and an enterprise-level intelligent decision-making map is finally output.
It improves the integration efficiency of multimodal data and the accuracy of intelligent analysis, generates enterprise-level decision support with counterfactual robustness, and enhances the implicit discovery of business logic and the interpretability of decisions.
Smart Images

Figure CN120296158A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data mining, and particularly relates to an enterprise big data mining method and system based on artificial intelligence. Background Art
[0002] In the current era of information explosion and digital transformation, enterprises are facing unprecedented data challenges. During their operation, enterprises generate and accumulate a large amount of multi-source heterogeneous data, including structured data and unstructured data, such as sales records, customer feedback, market research, social media interactions, sensor data, etc. These data contain rich business value and can provide important bases for enterprise decision-making. However, how to effectively extract useful information and insights from these complex data has become an important issue in enterprise operation and strategic formulation. Traditional data mining methods mostly rely on linear models and single-type data analysis. However, in the face of the increasingly complex business environment, the limitations of these methods are gradually revealed. Especially when dealing with multi-source heterogeneous data, traditional methods often struggle to adapt to the dynamic changes of the data and the complexity of cross-modal features, resulting in a significant reduction in the accuracy and effectiveness of the mining results. Summary of the Invention
[0003] The purpose of the present invention is to provide an enterprise big data mining method and system based on artificial intelligence to solve the deficiencies in the prior art, which can efficiently integrate multi-modal data and perform intelligent analysis, and improve the accuracy and effectiveness of the mining results.
[0004] An embodiment of the present application provides an enterprise big data mining method based on artificial intelligence, and the method includes: According to the enterprise multi-source heterogeneous data, perform dynamic modal alignment using a multi-modal hypergraph neural network, fuse cross-modal features through a self-attention-driven dynamic hypergraph construction algorithm, and generate a spatio-temporally consistent multi-modal joint embedding tensor. Among them, the dynamic hypergraph construction algorithm introduces an inter-modal causal inference mechanism and eliminates the time drift noise of cross-domain data through tensor decomposition; Input the multi-modal joint embedding tensor into an orthogonal adversarial manifold learning module, construct a manifold projection space for high-dimensional sparse data based on a transfer learning framework, perform adversarial correction on the feature distribution through an orthogonal-constrained generative adversarial network, and generate a low-dimensional compact and class-separable enhanced semantic embedding vector. Among them, the orthogonal constraint forces the weight matrices of the generator and the discriminator to be orthogonal through Frobenius norm regularization, suppressing mode collapse; Perform spatio-temporal causal association mining on the semantic embedding vectors, extract multi-hop association rules using a deep probabilistic inference model optimized by meta-paths, dynamically adjust the causal threshold through reinforcement learning, and output a spatio-temporal causal meta-path graph containing implicit business logic. Among them, the deep probabilistic inference model jointly optimizes the confidence and causal strength of meta-paths through a path reward mechanism and Monte Carlo tree search; Input the spatio-temporal causal meta-path graph into a dynamic game adversarial interpretation framework, generate an optimal decision-making strategy based on an attention-driven policy optimization mechanism, balance rule interpretability and business value gain through an adversarial perturbation compensation algorithm constrained by KL divergence, and finally output an enterprise-level intelligent decision-making graph with counterfactual robustness. Among them, the framework jointly optimizes the stability of the interpretation boundary and decision confidence through an implicit policy gradient algorithm, and introduces a manifold projection technique for adversarial samples to enhance decision robustness.
[0005] Optionally, based on the enterprise's multi-source heterogeneous data, use a multi-modal hypergraph neural network for dynamic modal alignment, fuse cross-modal features through a self-attention-driven dynamic hypergraph construction algorithm, and generate a spatio-temporally consistent multi-modal joint embedding tensor. Among them, the dynamic hypergraph construction algorithm introduces an inter-modal causal inference mechanism to eliminate the time drift noise of cross-domain data through tensor decomposition, including: Based on the enterprise's multi-source heterogeneous data including structured transaction records, unstructured customer feedback texts, and Internet of Things time series data, use a modal alignment tensor decomposition algorithm to perform timestamp alignment and semantic disambiguation on the multi-source heterogeneous data, and generate an initial tensor with cross-modal temporal alignment; Input the initial tensor into a self-attention-driven dynamic hypergraph construction module, calculate the cross-modal feature correlation weights through an inter-modal causal inference mechanism, and filter the time drift noise using a causal mask matrix to output a dynamic hypergraph adjacency matrix; Perform spatio-temporal joint embedding learning on the dynamic hypergraph adjacency matrix, use a multi-head graph attention network to aggregate the spatio-temporal dependence relationships of cross-modal nodes, and generate a graph embedding vector with multi-modal feature fusion; Input the graph embedding vector into a tensor rank constraint compression module, eliminate redundant features through non-negative matrix factorization and low-rank projection, and finally output a spatio-temporally consistent multi-modal joint embedding tensor.
[0006] Optionally, input the multi-modal joint embedding tensor into an orthogonal adversarial manifold learning module, construct a manifold projection space for high-dimensional sparse data based on a transfer learning framework, perform adversarial correction on the feature distribution through an orthogonal constraint generative adversarial network, and generate a low-dimensional compact semantic embedding vector with enhanced class separability. Among them, the orthogonal constraint forces the weight matrices of the generator and discriminator to be orthogonal through Frobenius norm regularization to suppress mode collapse, including: Construct a manifold projection space based on transfer learning according to the multi-modal joint embedding tensor, and use the adversarial domain adaptation algorithm to align the feature distributions of the source domain and the target domain to generate an initial manifold projection vector; Input the initial manifold projection vector into an orthogonal constraint generative adversarial network, and force the weight matrices of the generator and the discriminator to be orthogonalized through Frobenius norm regularization, and output the adversarial feature distribution after suppressing mode collapse; Dynamically correct the adversarial feature distribution, use the Wasserstein distance to measure the difference in feature separability, and optimize the discriminator decision boundary through the gradient penalty mechanism to generate an intermediate feature vector with enhanced class separability; Input the intermediate feature vector into a manifold compactification module, and use a spectral clustering-driven dimensionality reduction algorithm to compress high-dimensional sparse features, and finally output a low-dimensional compact semantic embedding vector.
[0007] Optionally, perform spatio-temporal causal association mining on the semantic embedding vector, use a meta-path optimized deep probabilistic inference model to extract multi-hop association rules, dynamically adjust the causal threshold through reinforcement learning, and output a spatio-temporal causal meta-path graph containing implicit business logic, where the deep probabilistic inference model jointly optimizes the confidence and causal strength of the meta-path through a path reward mechanism and Monte Carlo tree search, including: Construct a meta-path optimized deep probabilistic inference model according to the semantic embedding vector, extract initial multi-hop association rules through a path walking algorithm, and generate a candidate meta-path set; Perform reinforcement learning on the candidate meta-path set, dynamically adjust the causal threshold based on the temporal difference error, screen out meta-paths with a confidence level higher than a preset value through a path reward mechanism, and output an optimized causal rule set; Perform Monte Carlo tree search on the causal rule set, evaluate the causal strength by simulating counterfactual scenarios, and correct the path weights in combination with the Bayesian posterior probability to generate a meta-path probability graph of spatio-temporal causal association; Input the meta-path probability graph into a graph pruning module, and use the information entropy threshold to eliminate low-significance edges, and finally output a spatio-temporal causal meta-path graph containing implicit business logic.
[0008] Optionally, input the spatio-temporal causal meta-path graph into a dynamic game adversarial interpretation framework, generate an optimal decision-making strategy based on an attention-driven strategy optimization mechanism, balance the rule interpretability and business value gain through a KL-divergence-constrained adversarial perturbation compensation algorithm, and finally output an enterprise-level intelligent decision-making graph with counterfactual robustness, where the framework jointly optimizes the stability of the interpretation boundary and the decision confidence through an implicit policy gradient algorithm, and introduces a manifold projection technology of adversarial samples to enhance decision robustness, including: Construct an attention-driven policy optimization model based on the spatio-temporal causal meta-path graph, dynamically allocate the priority weights of rule explanations through the multi-head self-attention mechanism, and generate an initial decision policy vector; Input the initial decision policy vector into the adversarial perturbation compensation module constrained by KL divergence, optimize the geometric distribution of the explanation boundary through the implicit policy gradient algorithm, and output an anti-perturbation enhanced intermediate decision policy; Perform dynamic game equilibrium solving on the intermediate decision policy, adopt the Nash equilibrium iterative algorithm to balance rule interpretability and business value gain, and generate a decision policy graph verified for counterfactual robustness; Input the decision policy graph into the manifold projection adversarial enhancement module, compress abnormal samples through the discriminator feature space of the generative adversarial network, and finally output an enterprise-level intelligent decision graph.
[0009] Another embodiment of the present application provides an enterprise big data mining system based on artificial intelligence, and the system includes: A fusion module for dynamically aligning modalities according to enterprise multi-source heterogeneous data using a multi-modal hypergraph neural network, fusing cross-modal features through a self-attention-driven dynamic hypergraph construction algorithm, and generating a spatio-temporally consistent multi-modal joint embedding tensor. Among them, the dynamic hypergraph construction algorithm introduces an inter-modal causal reasoning mechanism and eliminates the time drift noise of cross-domain data through tensor decomposition; A correction module for inputting the multi-modal joint embedding tensor into an orthogonal adversarial manifold learning module, constructing a manifold projection space for high-dimensional sparse data based on the transfer learning framework, and performing adversarial correction on the feature distribution through an orthogonal-constrained generative adversarial network to generate a low-dimensional compact and class-separable enhanced semantic embedding vector. Among them, the orthogonal constraint forces the weight matrices of the generator and the discriminator to be orthogonal through Frobenius norm regularization, suppressing mode collapse; An adjustment module for mining spatio-temporal causal associations of the semantic embedding vector, extracting multi-hop association rules using a meta-path optimized deep probabilistic inference model, dynamically adjusting the causal threshold through reinforcement learning, and outputting a spatio-temporal causal meta-path graph containing implicit business logic. Among them, the deep probabilistic inference model jointly optimizes the confidence and causal strength of the meta-path through a path reward mechanism and Monte Carlo tree search; An output module for inputting the spatio-temporal causal meta-path graph into a dynamic game adversarial interpretation framework, generating an optimal decision policy based on an attention-driven policy optimization mechanism, balancing rule interpretability and business value gain through a KL divergence-constrained adversarial perturbation compensation algorithm, and finally outputting an enterprise-level intelligent decision graph with counterfactual robustness. Among them, the framework jointly optimizes the stability of the explanation boundary and the decision confidence through an implicit policy gradient algorithm, and introduces the manifold projection technology of adversarial samples to enhance decision robustness.
[0010] Another embodiment of the present application provides a storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the method described in any one of the above when running.
[0011] Another embodiment of the present application provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method described in any one of the above.
[0012] Compared with the prior art, an enterprise big data mining method based on artificial intelligence provided by the present invention, according to enterprise multi-source heterogeneous data, uses a multi-modal hypergraph neural network for dynamic modal alignment to generate a spatio-temporally consistent multi-modal joint embedding tensor; inputs the multi-modal joint embedding tensor into an orthogonal adversarial manifold learning module to generate a low-dimensional compact and class-separability-enhanced semantic embedding vector; mines spatio-temporal causal associations of the semantic embedding vector to output a spatio-temporal causal meta-path graph containing implicit business logic; inputs the spatio-temporal causal meta-path graph into a dynamic game adversarial interpretation framework, and finally outputs an enterprise-level intelligent decision-making graph with counterfactual robustness, so as to be able to efficiently integrate multi-modal data and perform intelligent analysis, improving the accuracy and effectiveness of mining results. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a hardware structure block diagram of a computer terminal for an enterprise big data mining method based on artificial intelligence provided by an embodiment of the present invention; Figure 2 It is a flowchart of an enterprise big data mining method based on artificial intelligence provided by an embodiment of the present invention; Figure 3 It is a structural diagram of an enterprise big data mining system based on artificial intelligence provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention, and cannot be construed as a limitation to the present invention.
[0015] An embodiment of the present invention first provides an enterprise big data mining method based on artificial intelligence. This method can be applied to an electronic device, such as a computer terminal, specifically, an ordinary computer, etc.
[0016] The following takes running on a computer terminal as an example to describe it in detail. Figure 1 It is a hardware structure block diagram of a computer terminal for an enterprise big data mining method based on artificial intelligence provided by an embodiment of the present invention. As Figure 1As shown in the figure, the computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory.
[0017] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, which when executed, can cause the processor to execute any one of the enterprise big data mining methods based on artificial intelligence.
[0018] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0019] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, it can cause the processor to execute any one of the enterprise big data mining methods based on artificial intelligence.
[0020] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 1 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0021] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0022] See Figure 2 , the embodiments of the present invention provide an enterprise big data mining method based on artificial intelligence, which may include the following steps: S201, according to the enterprise multi-source heterogeneous data, use a multi-modal hypergraph neural network for dynamic modal alignment, fuse cross-modal features through a self-attention-driven dynamic hypergraph construction algorithm, and generate a spatio-temporally consistent multi-modal joint embedding tensor. Among them, the dynamic hypergraph construction algorithm introduces a causal inference mechanism between modalities and eliminates the time drift noise of cross-domain data through tensor decomposition; This method integrates data from different sources within an enterprise (such as structured transaction records, unstructured text, time-series sensor data, etc.) through a hypergraph neural network, uses the self-attention mechanism to dynamically calculate the correlation weights between modalities, and combines causal reasoning to eliminate the noise interference caused by time asynchrony. Finally, a spatio-temporally aligned multi-modal joint feature representation is generated, solving the problem of difficult unified modeling of multi-source heterogeneous data in enterprise big data, ensuring that information in different modalities is consistent in time and semantics, and providing high-quality feature input for subsequent in-depth mining.
[0023] Specifically, for the multi-source heterogeneous data of an enterprise, including structured transaction records, unstructured customer feedback text, and Internet of Things time-series data, the modal alignment tensor decomposition algorithm can be used to perform timestamp alignment and semantic disambiguation on the multi-source heterogeneous data, generating an initial tensor with cross-modal temporal alignment. This method unifies the timestamps of different data sources through tensor decomposition technology and eliminates semantic ambiguities (such as different definitions of "order volume" in financial and operational systems), generating a preliminarily aligned multi-modal data tensor, solving the problems of time asynchrony and semantic inconsistency of multi-source data, and laying a foundation for subsequent feature fusion.
[0024] The modal alignment of enterprise multi-source heterogeneous data is the basis for constructing a unified semantic space. First, for structured transaction records (such as daily sales turnover, inventory changes), fields such as timestamps, transaction amounts, and product categories need to be extracted, and the clock deviation of different business systems is aligned through the Dynamic Time Warping (DTW) algorithm. For example, when there is a minute-level difference between the transaction timestamp in the sales system and the shipping time in the logistics system, DTW matches the time axes of the two through non-linear stretching to ensure the temporal consistency of the same order event in different modal data.
[0025] For unstructured customer feedback text (such as customer service conversation records, social media comments), a pre-trained language model (BERT) is used for semantic embedding. First, key information such as product names and sentiment polarities is extracted through named entity recognition (NER), and then ambiguity is eliminated by combining context-aware word vector mapping. For example, the word "apple" needs to be semantically judged through context in feedback from the electronics and food fields and mapped to different embedding subspaces respectively. At the same time, time expressions in the text (such as "last Wednesday", "three months ago") are standardized and converted into unified timestamps to align with the time axis of structured data.
[0026] The processing of Internet of Things time-series data (such as the temperature and humidity of production line sensors) needs to solve the problem of sampling frequency differences. Multiscale wavelet transform is used to downsample high-frequency sensor data (such as vibration signals sampled 1000 times per second) to the same minute-level granularity as transaction records, while retaining the characteristics of key frequency bands. For example, the low-frequency component (reflecting daily periodic changes) and high-frequency component (reflecting instantaneous equipment anomalies) of temperature data are reconstructed separately to ensure the time correlation with business events.
[0027] After the above preprocessing, the three types of data are integrated into a three-dimensional tensor (time × feature × modality) through the modal alignment tensor decomposition algorithm. Specifically, each time slice contains structured numerical features (such as transaction amounts), text embedding vectors (768-dimensional BERT outputs), and sensor statistics (mean, variance). Tensor decomposition uses the Tucker decomposition model, and the core tensor dimensions are set to (time × 50 × 3). It is iteratively optimized by the alternating least squares (ALS) method to eliminate the time drift noise of cross-modal data. The finally generated initial tensor has the characteristics of cross-modal time series alignment. For example, a certain customer complaint event can be accurately associated with abnormal sensor readings and sales decline records in the same period.
[0028] The initial tensor is input into the self-attention-driven dynamic hypergraph construction module. The cross-modal feature correlation weights are calculated through the inter-modal causal reasoning mechanism, and the time drift noise is filtered using the causal mask matrix to output the dynamic hypergraph adjacency matrix; This method uses the self-attention mechanism to dynamically learn the causal relationships between different modalities (such as the correlation strength between "customer complaint text" and "sales decline"), and eliminates the noise interference caused by time misalignment through the causal mask, enhancing the rationality of cross-modal feature fusion and avoiding false associations caused by time drift.
[0029] The goal of dynamic hypergraph construction is to capture the non-linear associations between cross-modal features. First, the initial tensor is sliced along the time axis, and the feature vectors at each time step are input into the multi-head self-attention mechanism (Multi-Head Self-Attention). The number of attention heads is set to 8, and each head independently calculates the correlation weights between different modalities. For example, one attention head may focus on the cross-modal association between "customer negative emotion - sudden increase in production line temperature", while another head captures the causal relationship between "promotion activity - inventory fluctuation".
[0030] In order to enhance the robustness of causal reasoning, the inter-modal causal mask matrix is introduced. Based on the principle of time priority, the matrix prohibits the causal influence of subsequent events on previous events. For example, a customer complaint at time t cannot affect the sensor reading at time t-1, so the mask weight of the corresponding position is set to 0. The mask matrix is generated using the time-lagged Granger Causality Test, and the causal paths with a significance higher than p=0.05 are screened by the F statistic, and the remaining paths are suppressed.
[0031] The adjacency matrix update of the dynamic hypergraph is implemented through the gated recurrent unit (GRU). The attention weight of each time step is input into the GRU together with the historical hypergraph state. The forget gate in the update formula controls the retention ratio of historical information (for example, setting it to 0.7 means that 70% of the historical association weights are inherited). For example, when it is detected that the causal strength of "declining customer satisfaction → declining sales" continues to increase for three consecutive time steps, the GRU will dynamically increase the weight of the path while attenuating occasional noise paths (such as false associations caused by a single sensor false alarm).
[0032] The final output dynamic hypergraph adjacency matrix contains the weight information of nodes (features) and hyperedges (cross-modal associations). For example, a hyperedge may connect three nodes: "Social media negative sentiment score ≥ 0.8", "Production line temperature standard deviation > 5°C", and "Return rate increased by 20% on the same day". Its weight of 0.92 indicates a strong causal relationship. This matrix serves as the basis for subsequent graph embedding learning.
[0033] Performing spatiotemporal joint embedding learning on the dynamic hypergraph adjacency matrix, aggregating the spatiotemporal dependencies of cross-modal nodes using a multi-head graph attention network, and generating a graph embedding vector of multimodal feature fusion; This method captures the temporal and spatial dependencies of nodes in different modalities through a graph attention network (such as the spatiotemporal correlation between "equipment failure in a certain area" and "customer loss in the area"), generates a fused graph embedding representation, realizes deep interactive modeling of multimodal data, and extracts more business-significant spatiotemporal features.
[0034] The core of spatiotemporal joint embedding learning is to capture the spatial association and temporal evolution patterns between features at the same time. First, the dynamic hypergraph adjacency matrix is input into the spatiotemporal graph convolutional network (ST-GCN), which contains two branches: spatial convolution and temporal convolution. The spatial convolution layer uses Chebyshev polynomial approximation (order K=3) to aggregate neighborhood features on the hypergraph structure; the temporal convolution layer uses dilated causal convolution with a dilation factor of 2 to capture long-term dependencies. For example, for the abnormal temperature pattern of a certain production equipment, spatial convolution identifies the nodes associated with the upstream and downstream processes, while temporal convolution traces the periodic fluctuations in the past 24 hours.
[0035] To enhance cross-modal feature fusion, a heterogeneous multi-head graph attention mechanism (HMGAT) is designed. Each attention head focuses on feature interactions of specific modal combinations: Head 1: Structured data → Text modality (e.g., the emotional correlation between transaction volume and customer service text); Head 2: Text modality → Sensor data (e.g., the spectral correlation between complaint keywords and device vibration); Head 3: Sensor data → Structured data (e.g., the correlation between temperature peaks and inventory consumption rate).
[0036] The attention scores of each head are calculated with cosine similarity weighting and activated by LeakyReLU (negative slope 0.2). For example, in Head 1, the cosine similarity between the daily sales volume vector of a product and the corresponding customer service text embedding vector is 0.85, which becomes the attention weight of this edge after normalization.
[0037] Finally, the outputs of each head are merged through a gated fusion mechanism (Gate Fusion), and the gating weights are dynamically learned by a fully connected layer. For example, when the importance of the text modality in recent data significantly increases (e.g., a large number of customer complaints), the gating mechanism will automatically increase the fusion weights of Head 1 and Head 2. The generated graph embedding vector has a dimension of 256 and contains a compressed representation of cross-modal spatio-temporal dependencies. For example, an embedding vector may encode a composite pattern such as "accumulation of customer negative emotions in the past week → decline in production line efficiency → inventory backlog".
[0038] Input the graph embedding vector into a tensor rank constraint compression module, and eliminate redundant features through non-negative matrix factorization and low-rank projection, and finally output a spatio-temporally consistent multi-modal joint embedding tensor.
[0039] This method removes redundant information in features (such as highly correlated sales metrics) through low-rank decomposition technology, retains the most discriminative core features, improves the simplicity and generalization ability of feature representation, and reduces the complexity of subsequent calculations.
[0040] The goal of tensor rank constraint compression is to remove noise and redundant information in the embedding vector. First, the graph embedding vector is reorganized into a three-dimensional tensor (time × node × feature) according to time steps and input into a non-negative matrix factorization (NMF) module. The dimension of the basis matrix of NMF is set to (time × 50), and the coefficient matrix is (50 × node × feature), which is iteratively optimized through a multiplicative update rule, and the non-negative constraint is enforced to enhance physical interpretability. For example, after decomposition, 50 basis patterns may be obtained, and basis pattern 3 corresponds to "customer sentiment - positive sales feedback loop during promotion period".
[0041] To further reduce the dimension, the Low-Rank Projection (LRP) technique is adopted. The first 30 principal components of the tensor are extracted through Singular Value Decomposition (SVD) (retaining 95% of the variance), and the dimension of the projection matrix is (256×30). For example, the feature representing "the increase in defective product rate caused by equipment aging" in a certain high-dimensional embedding vector is compressed into 3 principal components in the low-rank space, corresponding to "vibration frequency shift", "temperature baseline drift", and "extended production cycle" respectively.
[0042] The dimension of the finally output multi-modal joint embedding tensor is (time×node×30), with spatio-temporal consistency. For example, at the time slice t = 2023-10-05, the embedding vector of the node "East China Warehouse" may reflect the chain effect of "logistics delay → sharp increase in customer complaints → emergency inventory transfer", and smoothly transition with the embedding vectors of the previous and subsequent time steps. This tensor serves as a unified input for downstream tasks (such as causal reasoning and decision optimization), supporting the entire process of enterprise-level intelligent decision-making.
[0043] S202, input the multi-modal joint embedding tensor into the orthogonal adversarial manifold learning module, construct a manifold projection space for high-dimensional sparse data based on the transfer learning framework, and perform adversarial correction on the feature distribution through an orthogonal-constrained generative adversarial network to generate a low-dimensional, compact, and semantically enhanced embedding vector with enhanced class separability. Among them, the orthogonal constraint forces the weight matrices of the generator and discriminator to be orthogonal through Frobenius norm regularization, suppressing mode collapse; This method maps high-dimensional sparse data to a low-dimensional manifold space through a generative adversarial network (GAN), and introduces orthogonal constraints to prevent the model from falling into mode collapse, making the generated features both compact and effectively distinguishable among different classes, improving the robustness and interpretability of feature representation, avoiding overfitting problems caused by data sparsity, and enhancing the classification and clustering effects in different business scenarios.
[0044] Specifically, a manifold projection space based on transfer learning can be constructed according to the multi-modal joint embedding tensor, and an adversarial domain adaptation algorithm can be used to align the feature distributions of the source domain and the target domain to generate an initial manifold projection vector; This method maps data in different business scenarios (such as different branches or time periods) to a unified feature space through adversarial learning techniques, eliminates the bias caused by distribution differences, enhances the generalization ability of the model in cross-scenario applications, and avoids performance degradation caused by changes in data distribution.
[0045] The multi-modal joint embedding tensor (size example: Batch Size×Time Steps×Features = 128×30×512) first constructs a manifold projection space through a transfer learning framework. The core goal of this step is to address the feature distribution differences between the source domain (such as historical sales data) and the target domain (such as real-time market data). The specific implementation is divided into three stages: Inter-domain feature alignment: Using the Adversarial Domain Adaptation (ADA) algorithm, input the source domain and target domain data into a feature extractor with shared weights (such as a variant of ResNet-18) to generate initial feature vectors (dimension 256). The discriminator network (consisting of 3 fully connected layers with a hidden layer dimension of 128) drives the alignment of feature distributions through a binary classification task (distinguishing the source domain from the target domain). For example, in a retail scenario, the source domain may be the sales logs of last year's "Double Eleven", and the target domain is the real-time transaction flow of this year's "618" promotion. Through ADA, the model can identify common consumption patterns across time periods.
[0046] Manifold projection optimization: Use the Maximum Mean Discrepancy (MMD) as the inter-domain difference metric, and combine the Radial Basis Function (RBF kernel bandwidth set to 0.5) to calculate the distribution distance. In each training iteration, the optimizer (Adam, learning rate 0.001) simultaneously minimizes the classification loss (cross-entropy) and the MMD loss (weight coefficient 0.3), forcing the projected features to overlap in the manifold space. For example, when dealing with the cross-modal alignment of customer review texts and sales figures, MMD can effectively eliminate the non-linear distribution shift between the text sentiment polarity (such as positive / negative) and the numerical sales volume.
[0047] Dynamic temperature scaling: To address the problem of class imbalance (such as the data volume of popular products being much larger than that of long-tail products), introduce a temperature scaling parameter (initial value 1.0), and dynamically adjust the density of the feature space according to the class frequency. The temperature value for low-frequency classes is reduced to 0.7, making their feature vectors more compactly distributed in the manifold space. This parameter is optimized through grid search using the F1-score of the validation set (step size 0.1).
[0048] The finally generated initial manifold projection vector (dimension 128) has cross-domain invariance. For example, it can map the "user purchasing power" features of different quarters to the same clustering region in the manifold space, laying the foundation for subsequent adversarial training.
[0049] Input the initial manifold projection vector into an orthogonal constraint generative adversarial network, and orthogonalize the weight matrices of the generator and discriminator through Frobenius norm regularization to output an adversarial feature distribution with mode collapse suppression; This method prevents the generative adversarial network from falling into mode collapse (such as generating a single feature) through orthogonal constraints, ensuring feature diversity, enhancing the richness and distinctiveness of feature representations, and avoiding model degradation.
[0050] The structural design of the Orthogonal GAN (O-GAN) focuses on solving the mode collapse problem of traditional GANs. The specific implementation details are as follows: Network architecture configuration: Generator: A 4-layer transposed convolutional network is adopted (number of channels 512→256→128→64), and each layer is followed by spectral normalization (SN) and LeakyReLU activation (negative slope 0.2). The input noise vector has a dimension of 100, and the output feature map size is aligned with the manifold projection vector.
[0051] Discriminator: A 5-layer convolutional network (number of channels 64→128→256→512→1), and gradient penalty (GP coefficient 10) is used to stabilize the training process.
[0052] Orthogonal constraint implementation: Frobenius norm regularization is introduced in the fully connected layer of the discriminator (for example, the last layer with dimension 512→1). Specifically, for the weight matrix W (size 512×1), its orthogonality loss is calculated: Loss_orth = ||W^T W - I||_F^2. Here, I is the identity matrix, and the regularization coefficient is set to 0.01. This constraint forces the weight vectors to be close to the orthogonal basis, thereby enhancing the feature decoupling ability of the discriminator. For example, when dealing with cross-modal features, orthogonality can ensure that abstract concepts such as "price sensitivity" and "brand loyalty" are independently modeled in the discriminator's decision space.
[0053] Adversarial training strategy: The Wasserstein GAN with Gradient Penalty (WGAN-GP) framework is adopted, and the training ratio of the generator and the discriminator is set to 1:5. 128 samples are input in each batch, and through alternating optimization, the adversarial samples generated by the generator (such as simulating user churn features) gradually approach the real data distribution. For example, when generating customer churn warning signals, O-GAN can generate both high-risk and low-risk samples of churn, avoiding the single output caused by mode collapse in traditional GANs.
[0054] After training is completed, the intermediate layer features of the discriminator (taking the output of the 3rd convolutional layer with a dimension of 256) are extracted as the adversarial feature distribution after mode collapse suppression. This distribution has better class separability. For example, it can expand the distance between the feature clusters of "high-value customers" and "ordinary customers" by 2.3 times (verified by t-SNE visualization).
[0055] Dynamically correct the adversarial feature distribution, use the Wasserstein distance to measure the difference in feature separability, and optimize the discriminator decision boundary through the gradient penalty mechanism to generate an intermediate feature vector with enhanced class separability; This method evaluates the quality of the feature distribution through the Wasserstein distance, optimizes the discriminator to generate more easily classifiable features, strengthens the separability of features, and improves the accuracy of downstream tasks (such as classification or clustering).
[0056] The dynamic correction phase aims to further sharpen the class boundaries and eliminate the interference of abnormal features. The specific process includes: Wasserstein distance optimization: Calculate the Wasserstein-1 distance (Earth Mover's Distance) between the real data distribution Pr and the generated distribution Pg as a quantitative indicator of feature separability. Through the Lipschitz continuity constraint (gradient penalty coefficient λ = 10) of the Critic network (with the same structure as the discriminator), ensure the stability of distance estimation. For example, in the customer segmentation scenario, this distance can effectively reflect the separation degree between the "enterprise customer" and "personal customer" groups.
[0057] Decision boundary sharpening: Hard sample mining: Select the 10% samples closest to the decision boundary from the adversarial feature distribution (sorted by the absolute value of the Critic output value) and enhance their features. For example, for the "potential churn customers" samples near the boundary, generate synthetic samples through MixUp data augmentation (mixing ratio α = 0.2) to increase the data density in the decision boundary area.
[0058] Gradient-guided correction: During backpropagation, apply Jacobian Regularization (coefficient 0.1) to the feature vector to constrain the smoothness of the feature space in the local area. This can prevent the decision boundary from being overly tortuous. For example, it can avoid misclassifying multiple purchase behaviors of the same user into different categories.
[0059] Dynamic Temperature Scheduling: Introduce the Cosine Annealing strategy to adjust the learning rate (initial 0.0002, 50 epochs), combined with the early stopping mechanism (terminate training when the validation set accuracy has not improved for 5 consecutive epochs) to prevent overfitting. At the same time, the Feature Projection Temperature linearly decays from 1.0 to 0.5, gradually enhancing the clustering compactness of the feature space.
[0060] The intermediate feature vectors (dimension 256) after dynamic correction exhibit better class separability. For example, in the product recommendation scenario, the distance between the feature cluster centers of different categories (such as 3C digital vs. beauty and skincare) increases by 40%, while the within-class variance decreases by 25%.
[0061] Input the intermediate feature vectors into the manifold compactification module, and use a spectral clustering-driven dimensionality reduction algorithm to compress the high-dimensional sparse features, and finally output low-dimensional compact semantic embedding vectors.
[0062] This method compresses high-dimensional features into a low-dimensional space through spectral clustering technology, while retaining the core semantic information, improving computational efficiency and enhancing the interpretability of features, facilitating business personnel to understand and use.
[0063] The goal of the manifold compactification module is to compress the 256-dimensional intermediate features into a more interpretable low-dimensional space (such as 32 dimensions), while maintaining the integrity of semantic information. The specific implementation technologies include: Spectral Clustering Pre-segmentation: Construct a feature similarity matrix: Use an adaptive Gaussian kernel (bandwidth σ is determined by the median of the k-nearest neighbor distances, k = 15) to calculate the similarity between samples. For example, in the customer segmentation scenario, the similarity matrix can capture the implicit associations between "high consumption, low frequency" and "low consumption, high frequency" groups.
[0064] Laplacian Matrix Construction: Symmetrically normalize the similarity matrix to generate the graph Laplacian matrix L = D - ¹ / ²(D - W) D - ¹ / ², where D is the degree matrix and W is the similarity matrix.
[0065] Feature Decomposition: Select the eigenvectors corresponding to the first 10 smallest eigenvalues of L to form the initial low-dimensional embedding (dimension 10). This step projects the high-dimensional features into the spectral space for subsequent non-linear dimensionality reduction.
[0066] Multi-layer Autoencoder Optimization: Encoder Structure: A 4-layer fully connected network (256 → 128 → 64 → 32), with BatchNorm and Swish activation functions after each layer.
[0067] Decoder Structure: A symmetric 4-layer decoding network (32 → 64 → 128 → 256).
[0068] Loss Function: Combining reconstruction loss (MSE) and spectral clustering loss (KL divergence, weight 0.5), forcing the embedded vectors after dimensionality reduction to retain the original information and conform to the structural characteristics of the spectral space. For example, in the supply chain prediction task, this design can ensure that key semantic concepts such as "transportation delay risk" are linearly separable in the low-dimensional space.
[0069] Manifold Smoothing Post-Processing: Locally Linear Embedding (LLE) Fine-Tuning: Reconstruct the local neighborhood of the 32-dimensional vectors after dimensionality reduction (number of neighbors k = 20) to eliminate distortions that may be introduced by the autoencoder.
[0070] Density Peak Clustering: Calculate the local density (radius ε = 0.5) and relative distance of each sample to identify the center points of clusters. For example, in financial fraud detection, this method can highlight the central positions of abnormal transaction clusters.
[0071] The finally output 32-dimensional semantic embedded vectors have the following characteristics: In the e-commerce scenario, the cosine similarity of the embedded vectors of the same user in different sessions exceeds 0.85, while the similarity between different users is less than 0.3, proving that its compactness and separability meet the requirements of downstream tasks.
[0072] S203, perform spatio-temporal causal association mining on the semantic embedded vectors, adopt a deep probability inference model optimized by meta-paths to extract multi-hop association rules, dynamically adjust the causal threshold through reinforcement learning, and output a spatio-temporal causal meta-path graph containing implicit business logic, where the deep probability inference model jointly optimizes the confidence and causal strength of meta-paths through a path reward mechanism and Monte Carlo tree search; This method uses meta-path analysis technology to mine multi-hop association rules in enterprise data (such as "Customer A → Product B → Market Trend C"), and dynamically optimizes the confidence of causal relationships through reinforcement learning. Finally, it generates a causal graph reflecting business logic, revealing implicit business rules that are difficult to discover by traditional analysis methods, assisting enterprise decision-makers to understand complex business relationships, and optimizing strategic planning and resource allocation.
[0073] Specifically, according to the semantic embedded vectors, a deep probability inference model optimized by meta-paths can be constructed, and initial multi-hop association rules can be extracted through a path walking algorithm to generate a candidate meta-path set; This method utilizes meta-path analysis technology to simulate multi-hop path walks (such as "customer → product → market trend") based on semantic embedding vectors, mining potential association rules from enterprise data to form a preliminary set of candidate meta-paths. It can discover long-chain business logics that are difficult to capture by traditional analysis methods, providing rich candidate association relationships for subsequent causal reasoning and avoiding the omission of important implicit business rules.
[0074] Semantic embedding vectors are low-dimensional compact feature representations generated by previous steps, containing the fused semantic information of enterprise multi-source data (such as transaction records, customer texts, sensor time series). The core goal of the meta-path optimized deep probabilistic inference model is to mine cross-entity and cross-modal multi-hop association rules from these vectors.
[0075] Meta-path walk algorithm design: Adopt a random walk strategy with restart (Restart Random Walk with Meta-Path Constraints) to perform path exploration on the constructed enterprise knowledge graph. Each walk step is limited to 3 hops (for example, "customer A → purchase → product B → belongs to → category C → associated with → promotion activity D"), and introduce a transfer probability based on attention weights to control the direction. For example, when walking to a product node, the model will dynamically adjust the probability of the next step towards the "category" or "supplier" node according to the customer purchase frequency (numerical feature in the embedding vector) and the sentiment polarity of the product description text (semantic feature in the embedding vector), and the weight ratio can be set to 7:3.
[0076] Deep probabilistic inference model construction: Use the variational autoencoder (VAE) framework to encode the sequences obtained from path walks into a latent probability distribution. The encoder uses a bidirectional LSTM (Long Short-Term Memory Network) to capture path context, with the hidden layer dimension set to 256, and the decoder generates candidate meta-paths through Monte Carlo sampling. For example, for the path "user → click → advertisement → associated with → marketing channel", the model will calculate the conditional probability of this path appearing in historical data (such as 0.85), and perform weighted scoring in combination with features such as path length and node type diversity.
[0077] Candidate meta-path generation: Through threshold filtering (such as retaining paths with a confidence level ≥ 0.7) and diversity control (at most retaining the top 3 paths of the same type), a set of candidate meta-paths is finally formed. For example, in a retail scenario, the following typical meta-paths may be generated: Path 1: customer → purchase → commodity → viewed → similar commodities → associated with → discount activity (confidence 0.82); Path 2: supplier → supply → commodity → complained about → customer service work order → associated with → logistics delay (confidence 0.75).
[0078] Each meta - path is attached with multi - dimensional attribute labels, including path length, node type combination, time decay coefficient, etc.
[0079] Perform reinforcement learning on the set of candidate meta - paths, dynamically adjust the causal threshold based on temporal - difference error, and screen out meta - paths with confidence higher than the preset value through a path reward mechanism, and output an optimized causal rule set; This method uses reinforcement learning to dynamically evaluate the causal strength of each meta - path, screens out high - confidence association rules (such as "customer complaints → product defects → sales decline") through a reward mechanism, eliminates low - correlation paths, optimizes the causal rule set, improves the reliability of causal relationships, ensures that the finally output meta - path graph only contains high - confidence business logics, reduces noise interference, and improves the credibility of decision - making support.
[0080] The reinforcement learning framework aims to dynamically optimize the causal threshold of meta - paths (i.e., the minimum confidence to judge whether a causal relationship holds) to balance rule coverage and accuracy. The specific process is as follows: Definition of state space and action space: State: A three - dimensional vector composed of the current causal threshold (initial value 0.7), the average confidence of candidate meta - paths (such as 0.74), and the path diversity index (calculated by Shannon entropy, range 0 - 1).
[0081] Action: The step size for adjusting the causal threshold, set to ±0.05 (for example, increasing from 0.7 to 0.75 or decreasing to 0.65).
[0082] Design of reward function: The reward value is calculated by weighted summation of three parts: Reward for rule accuracy: The proportion of meta - paths triggered during the verification period (such as 7 days) that lead to the expected results in actual business (for example, if the actual purchase conversion rate increases by 15% after path 1 is triggered, the reward is +0.3).
[0083] Penalty for coverage: If the number of valid paths is less than 20% of the total number due to too high a threshold, the penalty is - 0.2.
[0084] Reward for diversity: For every 0.1 increase in the entropy value of path type distribution, the reward is +0.1.
[0085] Temporal - difference (TD) learning and policy optimization: The Q - learning algorithm is used to update the action - value function, with the learning rate set to 0.01 and the discount factor γ = 0.9. For example, when the threshold is adjusted from 0.7 to 0.75, the rule accuracy increases but the coverage decreases. The system will update the Q - table according to the TD error (the difference between the actual reward and the predicted Q - value) and gradually converge to the optimal strategy.
[0086] Path reward mechanism screening: Perform secondary filtering on the meta - paths passing through the dynamic threshold, using the Path Importance Score (PIS): PIS = 0.6×confidence + 0.3×causal strength + 0.1×business weight Among them, the causal strength is calculated through the Granger causality test (with the lag order set to 3), and the business weight is labeled by domain experts (for example, the supply - chain path weight is set to 0.8, and the marketing path is 0.6). Finally, the paths with PIS≥0.65 are retained to form the optimized causal rule set.
[0087] Perform Monte Carlo tree search on the causal rule set, evaluate the causal strength by simulating counterfactual scenarios, and combine Bayesian posterior probability to correct the path weight to generate a meta - path probability map of spatio - temporal causal associations; This method simulates causal relationships in different business scenarios (such as "if the product is not defective, will the sales volume recover") through Monte Carlo tree search, combines Bayesian probability to dynamically correct the path weight, generates a probability map reflecting the true causal strength, enhances the robustness of causal reasoning, avoids overfitting or spurious associations, enables the map to adapt to changes in different business environments, and improves the adaptability of decision - making.
[0088] Monte Carlo tree search (MCTS) is used to evaluate the causal effect of meta - paths in counterfactual reasoning, which is specifically divided into four stages: Selection: Starting from the root node (the initial causal rule set), select child nodes through the UCB1 (Upper Confidence Bound) formula. For example, the UCB value of a certain path is calculated as: ; Among them, Q is the cumulative reward of this path, N is the number of visits, T is the total number of visits, and the exploration coefficient c is set to 1.414. Preferentially select the path with a high UCB value to enter the next layer.
[0089] Expansion: When encountering an incompletely explored node, expand new child nodes. For example, for the path "Customer → Complaint → Work Order → Association → Logistics Delay", generate a counterfactual scenario: Assume that the logistics response time is shortened by 20%, and simulate the change in the customer complaint rate. Random perturbations need to be injected during expansion (such as ±5% fluctuations in logistics parameters).
[0090] Simulation: Run the counterfactual scenario 1000 times in a virtual environment and record the causal strength index. For example, if the simulation shows that the complaint rate drops by 12% after the logistics is accelerated, the causal strength is recorded as 0.12 (normalized to the interval [0,1]).
[0091] Backpropagation: Update the Q-values and visit counts of the path nodes in the reverse direction based on the simulation results. At the same time, use the Bayesian posterior probability to correct the path weights: Prior distribution: Assume that the path weights follow a Beta distribution (α = 2, β = 2 represents a neutral prior).
[0092] Posterior update: According to the number of successful times (such as 800 valid times) and the number of failed times (200 invalid times) obtained from the simulation, update to Beta(802, 202).
[0093] Finally, the path weight takes the posterior mean (802 / (802 + 202) ≈ 0.799).
[0094] Through multiple rounds of iteration (such as 10000 simulations), generate a meta-path probability map. The weight of each edge in the map represents the probability strength of the causal relationship. For example: Edge A (Supply Chain Delay → Production Interruption): Probability 0.92; Edge B (Promotion Activity → Short-Term Sales Volume): Probability 0.85 but accompanied by a confidence interval [0.78, 0.89].
[0095] Input the meta-path probability map into the map pruning module, use the information entropy threshold to eliminate low-significance edges, and finally output a spatio-temporal causal meta-path map containing implicit business logic.
[0096] This method is based on information entropy analysis, automatically prunes low-significance edges (such as weakly associated or paths without actual business significance), retains the key causal chains, forms a concise and high-value meta-path map, improves the interpretability and practicality of the map, enables decision-makers to focus on the most influential business logic, avoids information overload, and improves decision-making efficiency.
[0097] The goal of map pruning is to remove noisy edges and redundant connections, improve the interpretability and computational efficiency of the map. The specific process is as follows: Information entropy calculation: Calculate the information entropy for the out-edge set of each node to measure the degree of chaos in the connections. For example, a certain supplier node has 3 out-edges: Edge 1: Delivery delay → Production problem (weight 0.9); Edge 2: Delivery delay → Inventory backlog (weight 0.7); Edge 3: Delivery delay → Customer complaint (weight 0.4).
[0098] The probability distribution is normalized to [0.45, 0.35, 0.20], and the information entropy is: H = -(0.45ln0.45 + 0.35ln0.35 + 0.20ln0.20) ≈ 1.03 Set an entropy threshold (such as 0.7), and nodes with values higher than this need to be pruned.
[0099] Significance test: Use Fisher's exact test to evaluate the statistical significance of the edges. For example, a certain edge appears 200 times in the historical data, and 180 of them lead to the expected results. The p-value is calculated as 1.2×10 -6 , which is lower than the significance level (α = 0.01), so it is retained.
[0100] Spectral clustering for redundancy removal: Perform spectral clustering on the retained edges. Set the number of clusters K = 5 and the feature vector dimension to 50. Only retain the edge with the highest weight among those with a similarity higher than 0.8 within the same cluster. For example, within a cluster, there are three similar paths: "Customer churn → Revenue decline", "Customer churn → Market share reduction", "Customer churn → Stock price drop". Only retain the first one (weight 0.92).
[0101] Strengthening spatio-temporal constraints: For edges with time decay characteristics (such as the short-term effect of promotional activities), impose time window constraints. For example, if the causal effect of a certain path decays to less than 30% of the initial value after 30 days, add a timeliness label to it and display it as a dashed line in the graph.
[0102] The finally output spatio-temporal causal meta-path graph can be directly used for decision support. For example, in the supply chain optimization scenario, the graph shows that the path "Raw material price increase → Supplier switching delay → Production plan adjustment" has a high causal strength (0.89). The system will recommend establishing an alternative supplier library in advance and trigger the automatic adjustment of procurement strategies.
[0103] S204. Input the spatio-temporal causal meta-path graph into the dynamic game adversarial interpretation framework, generate the optimal decision-making strategy based on the attention-driven strategy optimization mechanism, balance the rule interpretability and business value gain through the adversarial perturbation compensation algorithm constrained by KL divergence, and finally output an enterprise-level intelligent decision-making graph with counterfactual robustness. Among them, the framework jointly optimizes the stability of the interpretation boundary and the decision confidence through the implicit policy gradient algorithm, and introduces the manifold projection technology of adversarial samples to enhance the decision robustness.
[0104] This method combines game theory and adversarial learning techniques to maximize business value while ensuring decision interpretability, verifies the robustness of decisions through counterfactual analysis, and finally generates an intelligent decision-making graph that can guide actual business, providing interpretable and high-value decision support, avoiding the uncontrollable risks of black-box AI models, and enhancing the model's resistance to abnormal data and adversarial attacks.
[0105] Specifically, an attention-driven strategy optimization model can be constructed based on the spatio-temporal causal meta-path graph, and the priority weights of rule interpretation are dynamically allocated through the multi-head self-attention mechanism to generate the initial decision-making strategy vector. This method uses the multi-head self-attention mechanism to analyze the key paths in the causal graph (such as "supply chain delay → cost increase → profit decrease"), dynamically allocates the priorities of different rules, forms a preliminary decision-making strategy, ensures that the decision-making strategy focuses on the most influential causal relationships, avoids treating all rules equally, and improves the pertinence and effectiveness of the strategy.
[0106] The spatio-temporal causal meta-path graph stores implicit business logic in the form of a graph structure. The nodes represent business entities (such as customers, products, supply chain nodes), and the edges represent causal associations (such as "the increase in customer complaint rate → the increase in product return rate"). To construct an attention-driven strategy optimization model, the nodes and meta-paths in the graph are first encoded into high-dimensional embedding vectors. The embedding dimension of each node is set to 512 dimensions, and the features within its 3-hop neighborhood are aggregated through the GraphSAGE algorithm. For example, the embedding of a certain customer node may fuse its purchase records, the sentiment analysis results of complaint texts, and the quality indicators of associated products.
[0107] The multi-head self-attention mechanism (Multi-Head Self-Attention, MHSA) is used at this stage to dynamically evaluate the interpretation priorities of different meta-paths. The model sets 8 independent attention heads, and each head focuses on different-dimensional causal association patterns. For example, head 1 may focus on the lag effect in the time dimension (such as the impact of a promotion activity on quarterly sales being delayed by 2 months), and head 2 focuses on the regional diffusion effect in the space dimension (such as the impact intensity of a supply chain interruption in a certain region on the national distribution network). The calculation process of each head is as follows: Query-Key-Value Generation: The node embedding vectors are linearly transformed to generate Query, Key, and Value matrices, all with dimensions of 64×8 (i.e., 8 total heads and 64 single-head dimensions).
[0108] Attention Weight Calculation: After taking the dot product of the Query and Key and scaling (scaling factor is √64), the weight distribution is obtained through Softmax normalization. For example, the attention weight of a certain meta-path "raw material price increase → production cost increase → pricing strategy adjustment" may be as high as 0.8, while the weight of "employee satisfaction → production efficiency" is only 0.2, reflecting that the former is more important in the current business scenario.
[0109] Feature Fusion: The outputs of each head are concatenated and then fused through a fully connected layer to generate a comprehensive attention weight matrix.
[0110] Finally, the initial decision strategy vector is composed of the weighted sum of each node embedding and its corresponding attention weight. For example, for the "pricing strategy adjustment" node, its strategy vector may include features such as cost sensitivity coefficient (0.75) and market competition intensity (0.63), and the weights are determined by its hubness in multiple meta-paths. The dimension of this vector is 1024, covering the key driving factors of business decisions.
[0111] Input the initial decision strategy vector into the adversarial perturbation compensation module with KL divergence constraint, optimize the geometric distribution of the explanation boundary through the implicit policy gradient algorithm, and output the anti-perturbation enhanced intermediate decision strategy; This method optimizes the stability of the decision strategy through adversarial learning techniques, enabling it to maintain a reasonable explanation boundary even in the face of data noise or adversarial attacks, avoiding policy failure due to minor perturbations, enhancing the robustness of the decision strategy, ensuring reliable execution in complex or adversarial business environments, and reducing decision risks.
[0112] The goal of the adversarial perturbation compensation module is to enhance the robustness of the decision strategy to input noise while maintaining the interpretability of the rules. The specific implementation is divided into three stages: Adversarial Perturbation Generation: The fast gradient sign method (FGSM) is used to add perturbations to the initial strategy vector. The perturbation amplitude ε is set to 0.1 (corresponding to 10% of the maximum eigenvalue change), and the direction is determined by the gradient of the policy loss function. For example, if the gradient of the "cost sensitivity coefficient" dimension of a certain policy vector is positive, the perturbation will deliberately increase this coefficient to test the stability of the decision.
[0113] KL Divergence Constraint: Restrict the deviation between the perturbed policy distribution and the original distribution through the Kullback-Leibler Divergence (KL Divergence), with the threshold set at 0.05 (i.e., the distribution difference does not exceed 5%). During specific calculations, consider the policy vector as a Gaussian distribution and calculate the KL divergence of its mean μ and covariance Σ. For example, if the original policy vector has μ = [0.5, 0.3,...] and the perturbed μ' = [0.52, 0.29,...], if the KL value exceeds the threshold, reduce the perturbation amplitude.
[0114] Implicit Policy Gradient Optimization: Use the PPO (Proximal Policy Optimization) algorithm to update the policy parameters. In each iteration, sample 100 sets of policies from the perturbed policy distribution and evaluate their commercial value gain (such as the percentage increase in profit margin) and interpretability score (such as the matching degree between the decision rule and the meta-path graph). Adjust the gradient update step size through importance sampling, with the learning rate set at 0.001 to ensure the stability of policy updates.
[0115] The finally generated intermediate decision-making policy has stronger anti-interference ability. For example, in the test, when 5% noise is injected into the input data (simulating data acquisition errors), the commercial value fluctuation of the perturbed policy is reduced from ±15% to ±3%, and the interpretation weight of the key decision rule (such as "prioritize the supply of high-profit products") remains stable.
[0116] Solve the dynamic game equilibrium for the intermediate decision-making policy, and use the Nash equilibrium iteration algorithm to balance rule interpretability and commercial value gain to generate a decision-making policy graph for counterfactual robustness verification; This method combines game theory ideas to find the optimal balance between interpretability (such as clear decision rules) and commercial value (such as profit maximization), and verifies the robustness of the policy through counterfactual analysis, providing a decision-making solution that is both easy to understand and can bring actual commercial benefits, avoiding sacrificing business effects due to excessive pursuit of interpretability, or vice versa.
[0117] The dynamic game model models rule interpretability (scored by business experts) and commercial value gain (quantified by financial indicators) as the interests of two game parties. The process of solving the Nash equilibrium is as follows: Payment Matrix Construction: The strategy space of the interpretability party is to adjust the matching threshold between the decision rule and the meta-path graph (0.6 - 0.9, step size 0.05).
[0118] The strategy space of the commercial value party is to adjust the resource allocation weight (such as the market budget ratio from 20% to 50%).
[0119] The payment value is calculated through Monte Carlo simulation. For example, when the matching threshold is 0.8 and the market budget accounts for 40%, the interpretability score is 85 points and the profit margin is increased by 12%.
[0120] Iterative solution: The initial strategy pair is set as (matching threshold 0.7, market budget 35%).
[0121] In each round of iteration, both sides select the optimal response according to the opponent's strategy: The interpretability side fixes the market budget and searches for the matching threshold that maximizes its own payment (such as increasing from 0.7 to 0.75).
[0122] The business value side fixes the matching threshold and adjusts the budget allocation (such as increasing from 35% to 38%).
[0123] The convergence condition is set as the strategy change being less than 1% for 10 consecutive rounds or reaching the maximum number of iterations of 1000 times.
[0124] Counterfactual verification: Conduct counterfactual tests on the equilibrium strategy. For example, assume that "a certain supplier suddenly runs out of stock". The system traces the causal chain based on the meta-path graph (such as "the supplier rating drops → the probability of delivery delay increases"), and automatically generates emergency strategies (such as enabling a backup supplier and adjusting the production plan). The verification metrics include the supply chain recovery time after the strategy is executed (target < 48 hours) and the interpretability consistency score (required > 80 points).
[0125] In the final output decision strategy graph, each node is associated with the optimal strategy parameters and the counterfactual test results. For example, the strategy parameters of a certain node are a matching threshold of 0.78 and a market budget of 42%. The average recovery time in 5 counterfactual scenarios is 36 hours and the interpretability score is 82 points.
[0126] Input the decision strategy graph into the manifold projection adversarial enhancement module, compress abnormal samples through the discriminator feature space of the generative adversarial network, and finally output the enterprise-level intelligent decision graph.
[0127] This method uses the discriminator of the generative adversarial network (GAN) to identify and filter abnormal samples (such as extreme market conditions or data anomalies), ensuring that the final decision graph has high reliability in real business scenarios, further improving the practicality of the decision graph, enabling it to adapt to the uncertainties in the real business environment, and providing stable and implementable intelligent decision support for enterprises.
[0128] The manifold projection adversarial enhancement module identifies and corrects abnormal samples in the strategy through the generative adversarial network (GAN). The GAN structure is as follows: Generator: A 4-layer fully connected network with a 128-dimensional noise vector as input and a 1024-dimensional synthetic policy vector as output. The activation function is LeakyReLU (negative slope 0.2).
[0129] Discriminator: A 3-layer convolutional network (kernel size 3×3, stride 2), with the output being the probability of policy authenticity (0~1), and the Sigmoid activation function is used in the last layer.
[0130] Two key mechanisms are introduced in the training process: Anomaly sample detection: The intermediate layer features of the discriminator (output of the second layer) are used to construct the manifold space of the policy vector. The Mahalanobis Distance between the input policy and the manifold center is calculated, and the threshold is set to 3σ (σ is the standard deviation of the distances in the training set). For example, a certain policy may have a Mahalanobis Distance as high as 5.2 due to containing contradictory rules (such as "increasing advertising investment and cutting the market budget simultaneously"), and is determined to be abnormal.
[0131] Manifold projection correction: Gradient descent optimization is performed on the abnormal policy to project it onto the normal manifold space. The optimization objective is to minimize the L2 distance (weight 0.7) between the projected policy and the original policy and the discriminator feature distance (weight 0.3), with the upper limit of the number of iterations being 50 times. For example, after projection, a certain abnormal policy has its contradictory rules corrected to "increase advertising investment in the eastern region and at the same time reduce the budget of inefficient markets in the western region".
[0132] In the finally output enterprise-level intelligent decision-making graph, each node is attached with the anti-interference policy version and the manifold space coordinates. For example, the original policy vector of a certain node has coordinates [0.3, 0.6] in the manifold space. After adversarial enhancement, its anomaly score drops from 0.89 to 0.12, and the commercial value gain is stably in the range of 8% - 10%. All policies are stored on the blockchain to ensure that the decision-making process is traceable and auditable.
[0133] It can be seen that according to the enterprise's multi-source heterogeneous data, a multi-modal hypergraph neural network is used for dynamic modal alignment to generate a spatio-temporally consistent multi-modal joint embedding tensor; the multi-modal joint embedding tensor is input into the orthogonal adversarial manifold learning module to generate a low-dimensional compact and class-separability-enhanced semantic embedding vector; spatio-temporal causal association mining is performed on the semantic embedding vector to output a spatio-temporal causal meta-path graph containing implicit business logic; the spatio-temporal causal meta-path graph is input into the dynamic game adversarial interpretation framework, and finally an enterprise-level intelligent decision-making graph with counterfactual robustness is output, so as to be able to efficiently integrate multi-modal data and perform intelligent analysis, improving the accuracy and effectiveness of the mining results.
[0134] Another embodiment of the present invention provides an enterprise big data mining system based on artificial intelligence, see Figure 3, the system may include: A fusion module 301, configured to perform dynamic modality alignment using a multimodal hypergraph neural network based on enterprise multi-source heterogeneous data, fuse cross-modal features through a self-attention-driven dynamic hypergraph construction algorithm, and generate a spatio-temporally consistent multimodal joint embedding tensor. Wherein, the dynamic hypergraph construction algorithm introduces an inter-modal causal reasoning mechanism, and eliminates the temporal drift noise of cross-domain data through tensor decomposition; A correction module 302, configured to input the multimodal joint embedding tensor into an orthogonal adversarial manifold learning module, construct a manifold projection space for high-dimensional sparse data based on a transfer learning framework, perform adversarial correction on the feature distribution through an orthogonal constraint generative adversarial network, and generate a low-dimensional compact and class-separable semantic embedding vector. Wherein, the orthogonal constraint forces the weight matrices of the generator and the discriminator to be orthogonal through Frobenius norm regularization, suppressing mode collapse; An adjustment module 303, configured to perform spatio-temporal causal association mining on the semantic embedding vector, extract multi-hop association rules using a depth probability inference model optimized by a meta-path, dynamically adjust the causal threshold through reinforcement learning, and output a spatio-temporal causal meta-path graph containing implicit business logic. Wherein, the depth probability inference model jointly optimizes the confidence and causal strength of the meta-path through a path reward mechanism and Monte Carlo tree search; An output module 304, configured to input the spatio-temporal causal meta-path graph into a dynamic game adversarial interpretation framework, generate an optimal decision-making strategy based on an attention-driven policy optimization mechanism, balance rule interpretability and business value gain through an adversarial perturbation compensation algorithm constrained by KL divergence, and finally output an enterprise-level intelligent decision-making graph with counterfactual robustness. Wherein, the framework jointly optimizes the stability of the interpretation boundary and the decision confidence through an implicit policy gradient algorithm, and introduces a manifold projection technology for adversarial samples to enhance decision robustness.
[0135] It can be seen that, based on enterprise multi-source heterogeneous data, dynamic modality alignment is performed using a multimodal hypergraph neural network to generate a spatio-temporally consistent multimodal joint embedding tensor; the multimodal joint embedding tensor is input into an orthogonal adversarial manifold learning module to generate a low-dimensional compact and class-separable semantic embedding vector; spatio-temporal causal association mining is performed on the semantic embedding vector to output a spatio-temporal causal meta-path graph containing implicit business logic; the spatio-temporal causal meta-path graph is input into a dynamic game adversarial interpretation framework, and finally an enterprise-level intelligent decision-making graph with counterfactual robustness is output, thereby enabling efficient integration of multimodal data and intelligent analysis, and improving the accuracy and effectiveness of mining results.
[0136] An embodiment of the present invention also provides a storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0137] Specifically, in this embodiment, the above storage medium may be configured to store a computer program for executing the following steps: S201, according to the enterprise multi-source heterogeneous data, perform dynamic modality alignment using a multi-modal hypergraph neural network, fuse cross-modal features through a self-attention-driven dynamic hypergraph construction algorithm, and generate a spatio-temporally consistent multi-modal joint embedding tensor. Wherein, the dynamic hypergraph construction algorithm introduces an inter-modal causal reasoning mechanism and eliminates the time drift noise of cross-domain data through tensor decomposition; S202, input the multi-modal joint embedding tensor into an orthogonal adversarial manifold learning module, construct a manifold projection space for high-dimensional sparse data based on a transfer learning framework, perform adversarial correction on the feature distribution through an orthogonal constraint generative adversarial network, and generate a low-dimensional compact and class-separable semantic embedding vector. Wherein, the orthogonal constraint forces the weight matrices of the generator and the discriminator to be orthogonal through Frobenius norm regularization, suppressing mode collapse; S203, perform spatio-temporal causal association mining on the semantic embedding vector, extract multi-hop association rules using a meta-path optimized deep probabilistic inference model, dynamically adjust the causal threshold through reinforcement learning, and output a spatio-temporal causal meta-path graph containing implicit business logic. Wherein, the deep probabilistic inference model jointly optimizes the confidence and causal strength of the meta-path through a path reward mechanism and Monte Carlo tree search; S204, input the spatio-temporal causal meta-path graph into a dynamic game adversarial interpretation framework, generate an optimal decision-making strategy based on an attention-driven strategy optimization mechanism, balance the rule interpretability and business value gain through a KL-divergence-constrained adversarial perturbation compensation algorithm, and finally output an enterprise-level intelligent decision-making graph with counterfactual robustness. Wherein, the framework jointly optimizes the stability of the interpretation boundary and the decision confidence through an implicit policy gradient algorithm, and introduces a manifold projection technique for adversarial samples to enhance decision robustness.
[0138] It can be seen that, based on the enterprise's multi-source heterogeneous data, a multi-modal hypergraph neural network is used for dynamic modal alignment to generate a spatio-temporally consistent multi-modal joint embedding tensor; the multi-modal joint embedding tensor is input into an orthogonal adversarial manifold learning module to generate a low-dimensional compact semantic embedding vector with enhanced class separability; spatio-temporal causal association mining is performed on the semantic embedding vector to output a spatio-temporal causal meta-path graph containing implicit business logic; the spatio-temporal causal meta-path graph is input into a dynamic game adversarial interpretation framework, and finally an enterprise-level intelligent decision-making graph with counterfactual robustness is output, so as to efficiently integrate multi-modal data and perform intelligent analysis, improving the accuracy and effectiveness of mining results.
[0139] An embodiment of the present invention also provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0140] Specifically, the above electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0141] Specifically, in this embodiment, the above processor may be configured to execute the following steps through a computer program: S201, based on the enterprise's multi-source heterogeneous data, use a multi-modal hypergraph neural network for dynamic modal alignment, fuse cross-modal features through a self-attention-driven dynamic hypergraph construction algorithm, and generate a spatio-temporally consistent multi-modal joint embedding tensor. Among them, the dynamic hypergraph construction algorithm introduces an inter-modal causal reasoning mechanism to eliminate the time drift noise of cross-domain data through tensor decomposition; S202, input the multi-modal joint embedding tensor into an orthogonal adversarial manifold learning module, construct a manifold projection space for high-dimensional sparse data based on a transfer learning framework, perform adversarial correction on the feature distribution through an orthogonal-constrained generative adversarial network, and generate a low-dimensional compact semantic embedding vector with enhanced class separability. Among them, the orthogonal constraint forces the weight matrices of the generator and the discriminator to be orthogonal through Frobenius norm regularization, suppressing mode collapse; S203, perform spatio-temporal causal association mining on the semantic embedding vector, use a meta-path optimized deep probabilistic inference model to extract multi-hop association rules, dynamically adjust the causal threshold through reinforcement learning, and output a spatio-temporal causal meta-path graph containing implicit business logic. Among them, the deep probabilistic inference model jointly optimizes the confidence and causal strength of the meta-path through a path reward mechanism and Monte Carlo tree search; S204. Input the spatio-temporal causal meta-path graph into the dynamic game adversarial interpretation framework, generate the optimal decision-making strategy based on the attention-driven strategy optimization mechanism, balance the rule interpretability and business value gain through the adversarial perturbation compensation algorithm constrained by KL divergence, and finally output the enterprise-level intelligent decision-making graph with counterfactual robustness. Among them, the framework jointly optimizes the stability of the interpretation boundary and the decision confidence through the implicit policy gradient algorithm, and introduces the manifold projection technology of adversarial samples to enhance the decision robustness.
[0142] It can be seen that according to the enterprise's multi-source heterogeneous data, the multi-modal hypergraph neural network is used for dynamic modal alignment to generate a spatio-temporally consistent multi-modal joint embedding tensor; input the multi-modal joint embedding tensor into the orthogonal adversarial manifold learning module to generate a low-dimensional compact and class-separability-enhanced semantic embedding vector; conduct spatio-temporal causal correlation mining on the semantic embedding vector to output a spatio-temporal causal meta-path graph containing implicit business logic; input the spatio-temporal causal meta-path graph into the dynamic game adversarial interpretation framework, and finally output the enterprise-level intelligent decision-making graph with counterfactual robustness, so as to efficiently integrate multi-modal data and conduct intelligent analysis, and improve the accuracy and effectiveness of the mining results.
[0143] The above has detailed the structure, features and effects of the present invention according to the embodiments shown in the drawings. The above is only the preferred embodiment of the present invention, but the present invention is not limited to the scope shown in the drawings. Any changes made according to the concept of the present invention, or modified into equivalent embodiments with equivalent changes, still within the spirit covered by the description and the drawings, shall fall within the protection scope of the present invention.
Claims
1. An enterprise big data mining method based on artificial intelligence, characterized in that, The method includes: Based on the enterprise's multi-source heterogeneous data, a multi-modal hypergraph neural network is used for dynamic modal alignment. By means of a self-attention-driven dynamic hypergraph construction algorithm, cross-modal features are fused to generate a spatio-temporally consistent multi-modal joint embedding tensor. Among them, the dynamic hypergraph construction algorithm introduces an inter-modal causal reasoning mechanism, and eliminates the time drift noise of cross-domain data through tensor decomposition; The multi-modal joint embedding tensor is input into an orthogonal adversarial manifold learning module. Based on a transfer learning framework, a manifold projection space for high-dimensional sparse data is constructed. Through an orthogonal-constrained generative adversarial network, the feature distribution is adversarially corrected to generate a low-dimensional compact semantic embedding vector with enhanced class separability. Among them, the orthogonal constraint forces the weight matrices of the generator and the discriminator to be orthogonal through Frobenius norm regularization, suppressing mode collapse; Spatio-temporal causal association mining is performed on the semantic embedding vector. A depth probability inference model optimized by meta-path is used to extract multi-hop association rules, and the causal threshold is dynamically adjusted through reinforcement learning, and a spatio-temporal causal meta-path graph containing implicit business logic is output. Among them, the depth probability inference model jointly optimizes the confidence and causal strength of the meta-path through a path reward mechanism and Monte Carlo tree search; The spatio-temporal causal meta-path graph is input into a dynamic game adversarial interpretation framework. Based on an attention-driven policy optimization mechanism, an optimal decision-making strategy is generated. Through an adversarial perturbation compensation algorithm constrained by KL divergence, the rule interpretability and business value gain are balanced, and finally an enterprise-level intelligent decision-making graph with counterfactual robustness is output. Among them, the framework jointly optimizes the stability of the interpretation boundary and the decision confidence through an implicit policy gradient algorithm, and introduces a manifold projection technique for adversarial samples to enhance decision robustness.
2. The method according to claim 1, wherein The step of based on the enterprise's multi-source heterogeneous data, using a multi-modal hypergraph neural network for dynamic modal alignment, fusing cross-modal features by means of a self-attention-driven dynamic hypergraph construction algorithm, and generating a spatio-temporally consistent multi-modal joint embedding tensor, where the dynamic hypergraph construction algorithm introduces an inter-modal causal reasoning mechanism and eliminates the time drift noise of cross-domain data through tensor decomposition, includes: Based on the enterprise's multi-source heterogeneous data including structured transaction records, unstructured customer feedback texts, and IoT time-series data, a modal alignment tensor decomposition algorithm is used to perform timestamp alignment and semantic disambiguation on the multi-source heterogeneous data, generating an initial tensor with cross-modal temporal alignment; The initial tensor is input into a self-attention-driven dynamic hypergraph construction module. Through an inter-modal causal reasoning mechanism, the cross-modal feature association weights are calculated, and the causal mask matrix is used to filter the time drift noise, outputting a dynamic hypergraph adjacency matrix; Spatio-temporal joint embedding learning is performed on the dynamic hypergraph adjacency matrix. A multi-head graph attention network is used to aggregate the spatio-temporal dependence relationships of cross-modal nodes, generating a graph embedding vector with multi-modal feature fusion; The graph embedding vector is input into a tensor rank constraint compression module. Through non-negative matrix factorization and low-rank projection, redundant features are eliminated, and finally a spatio-temporally consistent multi-modal joint embedding tensor is output.
3. The method according to claim 2, wherein Input the multi-modal joint embedding tensor into the orthogonal adversarial manifold learning module, construct a manifold projection space for high-dimensional sparse data based on the transfer learning framework, and perform adversarial correction on the feature distribution through a generative adversarial network with orthogonal constraints to generate a low-dimensional compact semantic embedding vector with enhanced class separability. Among them, the orthogonal constraint forces the weight matrices of the generator and discriminator to be orthogonal through Frobenius norm regularization, suppressing mode collapse, including: Construct a manifold projection space based on transfer learning according to the multi-modal joint embedding tensor, and use an adversarial domain adaptation algorithm to align the feature distributions of the source domain and the target domain to generate an initial manifold projection vector; Input the initial manifold projection vector into the orthogonal constraint generative adversarial network, and force the weight matrices of the generator and discriminator to be orthogonal through Frobenius norm regularization, and output the adversarial feature distribution after suppressing mode collapse; Perform dynamic correction on the adversarial feature distribution, use the Wasserstein distance to measure the difference in feature separability, and optimize the discriminator decision boundary through the gradient penalty mechanism to generate an intermediate feature vector with enhanced class separability; Input the intermediate feature vector into the manifold compactification module, and use a spectral clustering-driven dimensionality reduction algorithm to compress the high-dimensional sparse features, and finally output a low-dimensional compact semantic embedding vector.
4. The method according to claim 3, wherein Perform spatio-temporal causal association mining on the semantic embedding vector, use a deep probabilistic inference model optimized by meta-paths to extract multi-hop association rules, and dynamically adjust the causal threshold through reinforcement learning to output a spatio-temporal causal meta-path graph containing implicit business logic. Among them, the deep probabilistic inference model jointly optimizes the confidence and causal strength of the meta-path through a path reward mechanism and Monte Carlo tree search, including: Construct a deep probabilistic inference model optimized by meta-paths according to the semantic embedding vector, and extract initial multi-hop association rules through a path walking algorithm to generate a candidate meta-path set; Perform reinforcement learning on the candidate meta-path set, dynamically adjust the causal threshold based on the temporal difference error, and screen out meta-paths with confidence higher than the preset value through the path reward mechanism to output an optimized causal rule set; Perform Monte Carlo tree search on the causal rule set, evaluate the causal strength by simulating counterfactual scenarios, and correct the path weights in combination with the Bayesian posterior probability to generate a meta-path probability graph of spatio-temporal causal association; Input the meta-path probability graph into the graph pruning module, and use the information entropy threshold to eliminate low-significance edges, and finally output a spatio-temporal causal meta-path graph containing implicit business logic.
5. The method according to claim 4, characterized in that Input the spatio-temporal causal meta-path graph into the dynamic game adversarial interpretation framework, generate an optimal decision-making strategy based on the attention-driven policy optimization mechanism, and balance the rule interpretability and business value gain through the adversarial perturbation compensation algorithm constrained by the KL divergence. Finally, output an enterprise-level intelligent decision-making graph with counterfactual robustness. Among them, the framework jointly optimizes the stability of the interpretation boundary and the decision confidence through the implicit policy gradient algorithm, and introduces the manifold projection technology of adversarial samples to enhance the decision-making robustness, including: Construct an attention-driven policy optimization model based on the spatio-temporal causal meta-path graph, dynamically allocate priority weights for rule explanations through the multi-head self-attention mechanism, and generate an initial decision policy vector; Input the initial decision policy vector into the adversarial perturbation compensation module constrained by KL divergence, optimize the geometric distribution of the explanation boundary through the implicit policy gradient algorithm, and output an anti-perturbation enhanced intermediate decision policy; Solve the dynamic game equilibrium for the intermediate decision policy, adopt the Nash equilibrium iterative algorithm to balance rule interpretability and business value gain, and generate a counterfactual robustness verification decision policy graph; Input the decision policy graph into the manifold projection adversarial enhancement module, compress abnormal samples through the discriminator feature space of the generative adversarial network, and finally output an enterprise-level intelligent decision graph.
6. An enterprise big data mining system based on artificial intelligence, characterized in that, The system includes: A fusion module for dynamically aligning modalities using a multi-modal hypergraph neural network based on enterprise multi-source heterogeneous data, fusing cross-modal features through a self-attention-driven dynamic hypergraph construction algorithm, and generating a spatio-temporally consistent multi-modal joint embedding tensor. Among them, the dynamic hypergraph construction algorithm introduces an inter-modal causal reasoning mechanism to eliminate temporal drift noise in cross-domain data through tensor decomposition; A calibration module for inputting the multi-modal joint embedding tensor into an orthogonal adversarial manifold learning module, constructing a manifold projection space for high-dimensional sparse data based on the transfer learning framework, and performing adversarial calibration on the feature distribution through an orthogonal constraint generative adversarial network to generate a low-dimensional compact and class-separable enhanced semantic embedding vector. Among them, the orthogonal constraint forces the weight matrices of the generator and discriminator to be orthogonal through Frobenius norm regularization to suppress mode collapse; An adjustment module for mining spatio-temporal causal associations of the semantic embedding vector, extracting multi-hop association rules using a meta-path optimized deep probabilistic inference model, dynamically adjusting the causal threshold through reinforcement learning, and outputting a spatio-temporal causal meta-path graph containing implicit business logic. Among them, the deep probabilistic inference model jointly optimizes the confidence and causal strength of the meta-path through a path reward mechanism and Monte Carlo tree search; An output module for inputting the spatio-temporal causal meta-path graph into a dynamic game adversarial interpretation framework, generating an optimal decision policy based on an attention-driven policy optimization mechanism, balancing rule interpretability and business value gain through a KL divergence-constrained adversarial perturbation compensation algorithm, and finally outputting an enterprise-level intelligent decision graph with counterfactual robustness. Among them, the framework jointly optimizes the stability of the explanation boundary and the decision confidence through an implicit policy gradient algorithm, and introduces a manifold projection technique for adversarial samples to enhance decision robustness.
7. The system according to claim 6, characterized in that The fusion module is specifically used for: Based on the enterprise's multi-source heterogeneous data including structured transaction records, unstructured customer feedback texts, and IoT time-series data, use the modality alignment tensor decomposition algorithm to perform timestamp alignment and semantic disambiguation on the multi-source heterogeneous data, and generate an initial tensor with cross-modal temporal alignment; Input the initial tensor into the self-attention-driven dynamic hypergraph construction module, calculate the cross-modal feature correlation weights through the inter-modal causal reasoning mechanism, filter the time drift noise using the causal mask matrix, and output the dynamic hypergraph adjacency matrix; Perform spatio-temporal joint embedding learning on the dynamic hypergraph adjacency matrix, adopt a multi-head graph attention network to aggregate the spatio-temporal dependencies of cross-modal nodes, and generate a graph embedding vector for multi-modal feature fusion; Input the graph embedding vector into the tensor rank constraint compression module, eliminate redundant features through non-negative matrix factorization and low-rank projection, and finally output a spatio-temporally consistent multi-modal joint embedding tensor.
8. The system according to claim 7, wherein The correction module is specifically used for: Construct a manifold projection space based on transfer learning according to the multi-modal joint embedding tensor, adopt an adversarial domain adaptation algorithm to align the feature distributions of the source domain and the target domain, and generate an initial manifold projection vector; Input the initial manifold projection vector into the orthogonal constraint generative adversarial network, enforce the orthogonality of the weight matrices of the generator and the discriminator through Frobenius norm regularization, and output the adversarial feature distribution after suppressing mode collapse; Perform dynamic correction on the adversarial feature distribution, measure the difference in feature separability using the Wasserstein distance, and optimize the discriminator decision boundary through the gradient penalty mechanism to generate an intermediate feature vector with enhanced class separability; Input the intermediate feature vector into the manifold compactification module, adopt a spectral clustering-driven dimensionality reduction algorithm to compress high-dimensional sparse features, and finally output a low-dimensional compact semantic embedding vector.
9. A storage medium, characterized in that, A computer program is stored in the storage medium, wherein the computer program is set to execute the method according to any one of claims 1-5 when running.
10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is set to run the computer program to execute the method according to any one of claims 1-5.
Citation Information
Patent Citations
Urban resource prediction method and system based on multi-modal city knowledge graph
CN115796331A
Customer portrait key data mining method and system based on space-time big data
CN118797542A
Video analysis method based on deep learning platform and related equipment
CN118918522A
Dynamic data pipeline construction method based on artificial intelligence and multi-modal data processing
CN119830200A
Bidding and tendering data intelligent analysis method and system based on AI technology and storage medium
CN119850315A
Cited By
Multi-modal industrial data exception collaborative detection method and system
CN120524396A
Enterprise intelligent decision-making method and system driven by causal atlas
CN120542981A
Multi-dimensional data-driven subscription service user loss risk and value combined prediction method
CN120611840A
Industrial digital real-time risk control and decision support system based on block chain
CN120611985A
Blockchain-based industrial digitization real-time risk control and decision support system
CN120611985B