Decision model construction method based on big data environment
Through the joint model of causal discovery and neural symbols, combined with the knowledge graph, the problem of insufficient explanatory nature of decision-making models in the big data environment is solved, and decision-making support with high accuracy and interpretability is achieved, which is suitable for dynamic optimization of complex economic systems.
Patent Information
- Application Number
- CN202510410039.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
AI Technical Summary
The existing data-driven decision-making methods lack explanatory and causal research in the big data environment, resulting in the decision-making process being distrusted and accepted.
Using a combination of causal discovery, neural network computing and symbolic logic, a decision-making model based on the big data environment is constructed, and data-driven and knowledge-driven collaborative learning is realized through multi-source heterogeneous data fusion, knowledge graph generation, neural symbol joint model training and multi-dimensional verification.
Build a decision-making model with high prediction accuracy and good interpretation, which can be dynamically optimized under real-time data flow and expert feedback, providing transparent, impartial and open decision-making support.
Smart Images

Figure CN120258453A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of economic data decision-making, and more specifically relates to a method for constructing a decision-making model based on a big data environment. Background Art
[0002] In modern economic activities, the complexity and uncertainty of economic decision-making problems are very common. The decision-making problems in complex economic systems are characterized by multi-source heterogeneous data, complex interactions between variables, and uncertainty factors to be borne, etc. Therefore, it is particularly important to scientifically and effectively process these decision-making problems.
[0003] Traditional decision-making models are generally based on the assumption of rational people, that is, it is assumed that decision-makers are completely rational and can choose the action that maximizes their own benefits from all available actions. However, in reality, due to factors such as incomplete information and limited thinking ability, the decision-making behavior of economic agents usually deviates from the expectation of complete rationality.
[0004] In this context, data-driven decision-making has become an important decision-making strategy. Data-driven decision-making uses big data and artificial intelligence technologies to extract decision-making knowledge from historical data and provide decision-making support. However, existing data-driven decision-making methods are mainly based on statistical learning models, and there is insufficient research on the interpretability and causal relationship of the decision-making process, which brings difficulties to the trust and acceptance of decision-makers.
[0005] To address this problem, the present invention proposes a method for constructing a decision-making model based on a big data environment. This method comprehensively uses causal discovery, neural network operations, and symbolic logic, and with the help of a knowledge graph, organically combines data-driven and knowledge-driven to achieve structured and highly interpretable decision-making support in a complex economic decision-making environment. Summary of the Invention
[0006] The technical problem to be solved by the present invention is how to effectively integrate multi-source heterogeneous data in a big data environment, and through causal discovery and deep learning, construct a decision-making model that not only has high prediction accuracy but also has good interpretability. At the same time, solve the dynamic optimization problem of such models under real-time data streams and expert feedback, and construct a reliable, efficient and continuously optimizable decision-making support system.
[0007] To achieve the above object, the present invention is implemented by adopting the following technical solutions:
[0008] The method includes:
[0009] Obtain multi-source heterogeneous data, including sensor time-series data, expert rule bases, and device causal prior knowledge, and construct a standardized training data set with timestamps;
[0010] Extract the explicit causal relationships between variables from historical data through a causal discovery algorithm, and combine the causal priors in the equipment manual to generate an interpretable knowledge graph;
[0011] Construct a neuro-symbolic joint model, which includes a neural network module for feature extraction and a rule reasoning module based on symbolic logic, and design a knowledge fusion layer to dynamically coordinate the outputs of the two types of modules;
[0012] Adopt a joint training strategy with constrained optimization, and synchronously update the neural network parameters and rule confidence weights during the backpropagation process to achieve collaborative learning of data-driven and knowledge-driven;
[0013] Conduct multi-dimensional logical verification on the trained model, including rule conflict detection, counterfactual reasoning testing, and decision path traceability verification;
[0014] Based on real-time data streams and expert feedback, adopt an incremental conflict detection and rule correction mechanism to dynamically optimize the model and build a closed-loop iterative decision support system.
[0015] In one solution, the causal discovery algorithm adopts a constrained causal discovery framework, specifically including:
[0016] Generate an initial causal skeleton through the PC algorithm, and use the known causal relationships in the equipment manual as constraint conditions to screen candidate causal edges;
[0017] Adopt a Bayesian scoring function based on conditional independence testing to quantify the strength of causal relationships, and perform probability modeling on multi-hop causal paths between sensor data and equipment component nodes;
[0018] Construct a knowledge graph with weights and directions, where the node types include sensor variables, physical components, and fault types, and the edge weights reflect the statistical significance of causal relationships and the verification status of domain experts.
[0019] In one solution, the structure of the neuro-symbolic joint model includes:
[0020] The neural network module adopts a spatio-temporal graph convolutional network, and its graph structure is constructed based on the generated knowledge graph. The input layer receives sensor time-series data and extracts spatio-temporal correlation features;
[0021] The symbolic reasoning module encodes expert rules into differentiable logic programs, and each rule includes sensor threshold judgment of preconditions, time persistence constraints, and fault type mapping of conclusions;
[0022] The knowledge fusion layer includes an attention gating mechanism, which dynamically calculates the weighted fusion coefficient of the neural network prediction result and the rule reasoning result according to the feature distribution of the current input data.
[0023] In one solution, the joint training strategy includes the following optimization processes:
[0024] Design a multi-task loss function, including the neural network prediction error, the matching degree between the rule inference result and the annotation data, and the embedding consistency loss of the causal relationship in the knowledge graph;
[0025] Perform differentiable relaxation processing on the rule condition thresholds in the symbolic reasoning module, allowing the update of threshold parameters and rule application priorities during backpropagation;
[0026] Execute rule simplification operations after each iterative training, delete redundant rules with confidence weights lower than the preset threshold, and merge logically equivalent or highly relevant rule entries.
[0027] In one solution, the multi-dimensional logical verification includes:
[0028] In the static verification stage, use a model detection tool to perform deadlock analysis and decision path reachability verification on the expert rule base, and identify conflicting rule combinations;
[0029] In the dynamic verification stage, construct adversarial test cases through a counterfactual sample generator to evaluate the decision-making robustness of the model in scenarios of data distribution shift;
[0030] For the rule defects found during the verification process, automatically generate rule correction suggestions with context constraints and submit them to the expert for review.
[0031] In one solution, deploy a multi-modal processing pipeline to perform layout analysis and semantic parsing on PDF process documents and handwritten inspection records; first, extract text and table areas through OCR technology, and then use a domain-adapted BERT model trained by injecting device coding manual corpus for entity relationship extraction, converting unstructured data into a structured JSON format containing device parameters and operation instructions; among them, the reconstruction accuracy of the table area is not less than 98%, and the F1 value of entity relationship extraction is not less than 0.92.
[0032] In one solution, based on the constrained causal discovery algorithm GES, construct a causal graph model under spatio-temporal constraints and verify the causal strength through a structural equation model; for the multi-variable time series relationship in the device operation data, use a method combining time-delay mutual information and Granger causality test to extract the causal chain between device state parameters and calculate the causal effect size; each rule in the generated decision rule base is associated with a confidence weight, and the confidence threshold is not less than 0.85.
[0033] Advantages of the present invention:
[0034] Through causal discovery and neuro-symbolic joint models, a decision-making model can be efficiently constructed in a big data environment, greatly reducing the complexity and computational burden of model establishment.
[0035] Adopting a joint training strategy with constrained optimization, it can optimize both the neural network model and the symbolic logic model simultaneously, updating the parameters synchronously during the backpropagation process, and improving the accuracy of the model.
[0036] By generating a knowledge graph to explain the roles of the neural network and symbolic logic in the model, it provides logical verification of the model results, enhances the interpretability of the model, and is conducive to human understanding and acceptance.
[0037] Conducting multi-dimensional logical verification on the trained model, including rule conflict detection, counterfactual reasoning tests, and decision path traceability verification, makes the decision-making process more transparent, fair, and open.
[0038] Through an incremental conflict detection and rule correction mechanism, based on real-time data streams and expert feedback, the model is dynamically optimized, enabling the decision-making model to continuously adapt to and learn new situations, making the decision-making more accurate and real-time.
[0039] Therefore, the method for constructing a decision-making model based on a big data environment designed by this invention can not only efficiently process big data, but also establish a decision-making model with high accuracy, strong interpretability, and traceability, having great application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a flowchart of the method of this invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] To facilitate the understanding of this invention, the following will describe this invention more comprehensively with reference to the relevant drawings. The typical embodiments of this invention are shown in the drawings. However, this invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of this invention more thorough and comprehensive.
[0042] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as understood by those skilled in the technical field to which this invention belongs. The terms used in the description of this invention in this specification are only for the purpose of describing specific embodiments and are not intended to limit this invention. To facilitate the understanding of this invention, the following will describe this invention more comprehensively with reference to the relevant drawings. The typical embodiments of this invention are shown in the drawings. However, this invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of this invention more thorough and comprehensive.
[0043] As Figure 1 shown, a method for constructing a decision-making model based on a big data environment includes the following steps:
[0044] Step 1: Use a multi-source heterogeneous data fusion module to clean and align data from various types of report data, operation data, and production data to achieve spatio-temporal unification processing and generate standardized data.
[0045] When implementing the multi-source heterogeneous data fusion module, first, a distributed data lake needs to be constructed as the underlying storage architecture, and the characteristics of Delta Lake are used to achieve ACID transaction support and data version control. For raw data from different sources such as sensor networks, business systems, and scanned documents, a parallel access channel is established through the unified stream and batch processing engine of Flink. Among them, the time-series data stream is ingested in real-time through the Kafka message queue, while the historical batch data is loaded in slices through a distributed file system (such as HDFS). In the data cleaning stage, a dynamic anomaly detection algorithm based on a sliding window is designed for structured operation data, and an improved 3σ rule combined with an isolation forest model is used to identify outliers. At the same time, a multi-mode filling strategy is implemented for missing data, and time series prediction (Prophet model), spatial interpolation (Kriging algorithm), or associated field reasoning is used respectively according to the business meaning of the fields for completion. For unstructured report data (such as PDF process documents, handwritten inspection records), a multi-modal processing pipeline is deployed. First, a layout analysis module based on PaddleOCR is used to extract text and table areas, and then a BERT model trained for domain adaptation (injecting professional corpora such as device coding manuals) is used for entity relationship extraction, and finally, it is converted into a structured JSON format. In the spatio-temporal alignment link, a unified spatio-temporal coordinate system needs to be constructed. In the spatial dimension, the Gaussian-Kruger projection is used to convert GPS / Beidou positioning data with different precisions, and in the time dimension, the device clock deviation is calibrated through the dynamic time warping algorithm (DTW). For the data stream with asynchronous sampling, time series resampling based on Lagrange interpolation is implemented. After the basic processing is completed, various data fields are mapped to a unified semantic model using predefined ontology mapping rules, and data type constraints are enforced through Apache Avro. Finally, a standardized Parquet file containing spatio-temporal stamps, data source fingerprints, and quality assessment labels is generated and written into a partition storage structure optimized by Z-Order. At the same time, a data lineage graph is generated for subsequent traceability analysis. The entire process needs to embed data quality monitoring points to calculate integrity, consistency, and accuracy indicators in real-time. When field conflicts or confidence levels lower than the threshold are detected, an artificial review process is triggered to ensure the reliability of the output data.
[0046] Step 2: Based on the causal reinforcement learning framework and knowledge graph embedding technology, extract implicit causal features from the standardized data containing various relevant indicators, and construct a symbolic rule constraint set.
[0047] First, construct a dynamic causal graph model based on the standardized data, and adopt a two-layer optimization framework that combines the structural equation model (SEM) and reinforcement learning. At the bottom layer, through the causal discovery engine of the DoWhy library, conduct causal structure learning on the spatio-temporal correlation indicators in the standardized data. For the high-dimensional feature space, adopt an improved GES algorithm, introduce spatio-temporal constraint terms to optimize the search process, and its objective function is defined as:
[0048] min G∈DAG ∑(X i -f(PA(X i ))) 2 +λ1||W⊙A||1+λ2Φ(t,s)
[0049] where PA(X i ) represents the set of parent nodes of node X i , W is the weight of the adjacency matrix, and Φ(t,s) is the spatio-temporal consistency penalty term. In the knowledge graph embedding link, elements such as device entities and process parameters are modeled as a heterogeneous graph structure, and the RotatE model is used to learn entity relationships in the complex space. The embedding loss function is defined as:
[0050]
[0051] where h, r, t ∈ C d are complex vectors, represents the Hadamard product, and γ is the margin hyperparameter. The causal reinforcement learning framework fuses the causal graph and knowledge graph information through the Q-learning strategy, and designs the reward function R = α·I(X→Y)+β·Sim(e x ,e y ), where the mutual information I(X→Y) calculates the causal effect P(Y|do(X)) through do-calculus, and Sim(e x ,e y ) is the cosine similarity of the knowledge graph embedding vectors. The generation of the symbolic rule constraint set adopts probabilistic soft logic (PSL), and transforms the causal edge X→Y into a weighted logical rule:
[0052]
[0053] The weight w is jointly determined by the causal strength and the knowledge graph relationship score, and the rule confidence is solved through maximum likelihood estimation where It is a rule instantiation indication function. The finally constructed constraint set contains temporal logic expressions
[0054] ◇[0,Δ](X≥θ→◇[0,δ]Y≤μ)
[0055] and spatial propagation rules
[0056]
[0057] and other composite forms. Probabilistic collaborative reasoning between rules is realized through a Markov logic network to form an interpretable symbolic decision boundary.
[0058] Step 3: Design a neuro-symbolic joint reasoning system. By implementing stable model solving through dynamic logic programming and combining with the non-linear prediction channel of the neural network, cross-modal fusion of symbolic rules and data-driven is achieved.
[0059] In the implementation of Step 3, the neuro-symbolic joint reasoning system adopts a dual-channel architecture design. Its core consists of a dynamic logic programming solver and a graph attention causal network (GACN). The system first encodes the symbolic rule constraint set generated in Step 2 into a temporal logic program and uses the Clingo engine for ground derivation under the stable model semantics. Each atomic proposition p(t, s, v) carries a confidence weight generated by the LSTM-QNN network, defining the truth degree of the proposition
[0060]
[0061] where is the temporal feature encoding vector, coming from the knowledge graph embedding projection. The dynamic logic program is updated in real time through a time-sliding window mechanism. Its program rules are formalized as:
[0062]
[0063] where the energy function
[0064]
[0065] fuses the symbolic constraint term and the neural network prediction deviation, and β is the temperature coefficient. The neural channel adopts a multi-layer causal graph attention network, and its message passing formula is:
[0066]
[0067] The attention coefficient α_{ij} is modulated by the activation degree of the symbolic rule, and the calculation formula is
[0068] α ij =softmax(σ(a T [Ws ψ(r ij )||W n k ij ]))
[0069] , where ψ(r ij ) is the rule triggering function. When there is a symbolic constraint X i →X j Output rule confidence p {ij} , k {ij} is the original attention key value of the neural network. Cross-modal fusion is achieved through the tensor product interaction layer, and the joint reasoning result is defined as:
[0070]
[0071] Where SPN is the Dirichlet distribution output of the symbol probability network, GAP is the global average pooling feature of the neural channel, represents the Kronecker product expansion, ò is the learnable modal balance coefficient, and SM(Π) is the satisfaction score returned by the stable model solver. The system is jointly optimized by the alternating direction multiplier method (ADMM), and the objective function includes the symbolic consistency loss Ls = KL(q(Π)||p(Π|R)) and the prediction loss Among them, satr(·) represents the satisfaction mapping function of rule r in the neural output space. The final reasoning process forms a closed-loop optimization: the symbolic rules generate constraint energy surfaces through lattice grammar parsing to guide the gradient update direction of the neural network; at the same time, the new causal patterns recognized by the neural network are converted into incremental logic programs through the Jena rule engine, continuously improving the symbolic knowledge base and forming a two-way enhanced cognitive reasoning system.
[0072] Step 4: Verify the logical compliance of the model, including static rule conflict detection, dynamic coverage testing, and counterfactual consistency testing, and use formal methods to ensure that the decision output complies with the predetermined rules.
[0073] In the logic compliance verification of step 4, a static rule conflict detection framework based on linear temporal logic (LTL) is first established, and the symbolic rule constraint set generated in step 2 is converted into a state transition model through Büchi automaton. The verification space is defined as a five-tuple M = (S, Σ, δ, s0, F), where the state s∈S corresponds to the truth value combination of the rule atomic proposition, and the migration condition δ is jointly determined by the knowledge graph embedding similarity threshold and causal strength. The rule set is formally verified using the NuSMV model detector, and the core verification condition is expressed as the CTL formula: The deadlocks and contradictions in the reachable state space are calculated by the symbolic model checking algorithm. Dynamic coverage testing uses a hybrid verification method to construct a coverage measurement function Among them, the trigger function
[0074] trig(r|x) = σ(α·sim(ex,er)+β·P(Y|do(X)))
[0075] It is determined by the joint activation probability of the neural network prediction output and the symbolic rule. Coverage trajectory analysis is performed on the standardized dataset through Monte Carlo sampling. When r ∈ R such that Cov(r) < θc, the rule reparameterization mechanism based on attention weights is triggered to adjust the rule weights
[0076]
[0077] The counterfactual consistency test constructs a causal intervention verification pipeline and designs a counterfactual generator for the key decision variable X
[0078] CFG(X) = GN(ò)+GS(Δ), where the neural network generator GN generates counterfactual samples within the data distribution through adversarial training, and the symbolic generator GS performs do-calculus intervention based on the causal graph in step 1 to generate logically compliant counterfactuals. The verification metric is defined as the consistency score
[0079] Consist = E{x~CFG(X)}[(argmax f N (x) ≡ f S (x))]
[0080] where f N is the neural network prediction channel, f S is the symbolic rule inference channel. When Consist < η, the rule distillation module is automatically activated, and bidirectional knowledge alignment is achieved by minimizing the KL divergence
[0081]
[0082] Finally, the verification system integrates the results of the three stages to generate a formal proof certificate, and its compliance boundary is determined by the differential logic equation where ψ(R,x) is the activation mapping of the rule in the input space x, and εL is the preset acceptable risk threshold. The Z3 solver is used to verify the validity of this inequality under all possible input patterns to ensure that the decision-making system meets the safety constraints of the predefined rules
[0083] Step 5: Based on real-time data streams and expert feedback, an incremental conflict detection and rule correction mechanism is adopted to dynamically optimize the model and construct a closed-loop iterative decision support system
[0084] A dynamic evolution architecture based on an event bus is constructed. A distributed stream processing engine is used to capture multi-source heterogeneous data streams in real time, and sensor signals, business logs, and expert feedback are jointly encoded into timestamped knowledge graph incremental events. When a new data packet arrives, the lightweight edge inference module first performs local causal inference, uses the causal discovery algorithm in step 1 to online detect potential new association patterns, and proposes candidate constraint conditions through a differentially private protected rule hypothesis generator. These temporary rules and the existing rule base perform incremental conflict detection at the in-memory computing layer, and use an improved Retract-Update-Propagate (RUP) algorithm to dynamically maintain consistency: if it is found that the new rule is logically contradictory to the existing system, the system automatically activates the multi-armed bandit strategy, and under the premise of keeping the core constraints unchanged, dynamically adjusts the rule confidence weights according to real-time performance indicators, and at the same time pushes the conflict graph and correction suggestions to domain experts through the human-machine collaboration interface. The rule updates approved by experts are injected into the neuro-symbolic joint inference system through the hot deployment mechanism, triggering incremental retraining of the dual-channel model - on the symbolic reasoning side, the logic program slicing technology is used to only re-solve the stable model for the sub-graph related to the affected rules; on the neural network side, based on the elastic weight consolidation algorithm, when fine-tuning the graph attention network parameters, a regularization constraint of causal importance is imposed to prevent new knowledge from covering the original cognition. The enhanced sandbox environment is synchronously started in the closed-loop verification link, and the output of the updated model is compared with the historical decision-making trajectory for counterfactual analysis, and a credibility assessment report is generated through the compliance verification pipeline in step 4. Scenarios with abnormal fluctuations exceeding the threshold are automatically rolled back to the previous stable version. The final formed dynamic knowledge base is traced and managed through a blockchain-assisted version tree, and each rule iteration is associated with data fingerprints, expert signatures, and verification summaries, forming an interpretable, auditable, and evolvable decision support ecosystem, while continuously absorbing fresh knowledge and maintaining the overall logical completeness and operational stability of the system.
[0085] Example:
[0086] Taking the electricity market as an example, in order to effectively manage and dispatch electricity resources, market participants need to make decisions by predicting factors such as electricity prices and electricity demand. The data involved includes public information such as historical electricity prices, electricity demand data, climate data, holidays, as well as feedback data from major power companies and end-to-end users.
[0087] First, by collecting these multi-source heterogeneous data, including historical transaction data in the electricity market (electricity prices, trading volumes, etc. with timestamps), meteorological data (temperature, rainfall, wind power, etc.), public holiday data, expert manuals (stipulating the operating rules of the electricity market and the working principles of electrical equipment), and equipment data (operating status of power plant equipment, such as generator load, power, efficiency, etc.), a training data set for the power system is constructed.
[0088] Then, an explicit causal relationship between variables such as electricity demand, electricity supply, climate conditions, and public holidays is extracted from historical data through a causal discovery algorithm. Combining the operating rules of the electricity market and the working principles of electrical equipment, a knowledge graph in the economic field is generated.
[0089] Next, a neuro-symbolic joint model is constructed. The neural network module is used to extract features from electricity market transaction data, meteorological data, etc., and the symbolic logic model is used to perform reasoning based on the operating rules of the electricity market. Then, the output of the neural network and the symbolic model is dynamically coordinated through a knowledge fusion layer.
[0090] Then, a joint training strategy with constrained optimization is adopted. During the backpropagation process, the neural network parameters and rule confidence are updated simultaneously to achieve collaborative learning driven by data and knowledge.
[0091] Then, logical verification is performed on the trained model, including rule conflict detection, counterfactual reasoning tests, etc., to ensure the traceability of the decision-making path.
[0092] Finally, based on real-time electricity market data and expert feedback, the model is dynamically optimized through an incremental conflict detection and rule correction mechanism to adjust the decisions of the electricity market in real time.
[0093] Through the above steps, efficient, accurate, and transparent decision support for the electricity market can be achieved.
[0094] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.
[0095] It should be understood that the detailed description of the technical solutions of the present invention with the aid of the preferred embodiments is illustrative rather than restrictive. Those of ordinary skill in the art can modify the technical solutions recorded in each embodiment based on reading the specification of the present invention, or perform equivalent substitution on some of the technical features; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A method for constructing a decision-making model based on a big data environment, characterized in that: The described method includes: Obtain multi-source heterogeneous data, including sensor time-series data, expert rule bases, and device causal prior knowledge, and construct a standardized training dataset with timestamps; Extract explicit causal relationships between variables from historical data through a causal discovery algorithm, and generate an interpretable knowledge graph in combination with the causal prior of the device manual; Construct a neuro-symbolic joint model, which includes a neural network module for feature extraction and a rule inference module based on symbolic logic, and design a knowledge fusion layer to dynamically coordinate the outputs of the two types of modules; Adopt a joint training strategy with constrained optimization, and synchronously update the neural network parameters and rule confidence weights during the backpropagation process to achieve collaborative learning of data-driven and knowledge-driven; Conduct multi-dimensional logical verification on the trained model, including rule conflict detection, counterfactual reasoning tests, and decision path traceability verification; Based on real-time data streams and expert feedback, adopt an incremental conflict detection and rule correction mechanism to dynamically optimize the model and construct a closed-loop iterative decision support system.
2. The method for constructing a decision-making model based on a big data environment according to claim 1, wherein The described causal discovery algorithm adopts a constrained causal discovery framework, specifically including: Generate an initial causal skeleton through the PC algorithm, and use the known causal relationships in the device manual as constraint conditions to screen candidate causal edges; Adopt a Bayesian scoring function based on conditional independence testing to quantify the strength of causal relationships, and perform probability modeling on multi-hop causal paths between sensor data and device component nodes; Construct a knowledge graph with weights and directions, where the node types include sensor variables, physical components, and fault types, and the edge weights reflect the statistical significance of causal relationships and the verification status of domain experts.
3. A method for constructing a decision-making model based on a big data environment according to claim 1, characterized in that: The structure of the described neuro-symbolic joint model includes: The neural network module adopts a spatio-temporal graph convolutional network, whose graph structure is constructed based on the generated knowledge graph, and the input layer receives sensor time-series data and extracts spatio-temporal correlation features; The symbolic reasoning module encodes expert rules into differentiable logic programs, and each rule includes a sensor threshold judgment of the precondition, a time persistence constraint, and a fault type mapping of the conclusion; The knowledge fusion layer includes an attention gating mechanism, which dynamically calculates the weighted fusion coefficient of the neural network prediction result and the rule inference result according to the feature distribution of the current input data.
4. A method for constructing a decision-making model based on a big data environment according to claim 1, characterized in that: The described joint training strategy includes the following optimization processes: Design a multi-task loss function, including the neural network prediction error, the matching degree between the rule inference result and the labeled data, and the embedding consistency loss of causal relationships in the knowledge graph; Perform differentiable relaxation processing on the rule condition thresholds in the symbolic reasoning module, allowing the threshold parameters and rule application priorities to be updated during backpropagation; Execute rule simplification operations after each iterative training, delete redundant rules with confidence weights lower than the preset threshold, and merge logically equivalent or highly relevant rule entries.
5. A method for constructing a decision-making model based on a big data environment according to claim 1, characterized in that: The described multi-dimensional logical verification includes: In the static verification stage, use a model checking tool to perform deadlock analysis and decision path reachability verification on the expert rule base, and identify conflicting rule combinations; In the dynamic verification stage, construct adversarial test cases through a counterfactual sample generator to evaluate the decision robustness of the model in scenarios of data distribution shift; For the rule defects found in the verification process, automatically generate rule correction suggestions with context constraints and submit them to experts for review.
6. The method for constructing a decision-making model based on a big data environment according to claim 1, wherein: Deploy a multi-modal processing pipeline to perform layout analysis and semantic parsing on PDF process documents and handwritten inspection records; first, use OCR technology to extract text and table areas, and then use a domain-adapted BERT model trained by injecting equipment coding manual corpus to perform entity relationship extraction, converting unstructured data into a structured JSON format containing equipment parameters and operation instructions; among them, the reconstruction accuracy of the table area is not less than 98%, and the F1 value of entity relationship extraction is not less than 0.
92.
7. A method for constructing a decision-making model based on a big data environment according to claim 2, characterized in that: Based on the constraint-based causal discovery algorithm GES, construct a causal graph model under spatio-temporal constraint conditions and verify the causal strength through a structural equation model. For the multi-variable time series relationship in the equipment operation data, adopt a method combining time-delay mutual information and Granger causality test to extract the causal chain between equipment state parameters and calculate the causal effect size; each rule in the generated decision rule library is associated with a confidence weight, and the confidence threshold is not less than 0.85.
Citation Information
Cited By
Scheduling strategy selection large model training method based on reinforcement learning
CN120525020A
Power grid intelligent decision-making method, model construction method, equipment and medium
CN120542551A
Activity evaluation processing method and device, storage medium and electronic equipment
CN120598440A
Chip packaging test production line performance control method based on reinforcement learning
CN120631674A
Intelligent decision support system based on big data and artificial intelligence
CN120851666A