Data processing method and device based on non-causal prior, equipment and storage medium

By extracting non-causal prior knowledge from LLMs in high-dimensional noisy environments and suppressing spurious edges, the accuracy and stability issues of causal discovery algorithms are resolved, and more efficient causal relationship recognition is achieved.

CN121706960APending Publication Date: 2026-03-20NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511888409.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing causal discovery algorithms suffer from decreased learning accuracy in high-dimensional noisy scenarios, issues such as false dependencies and confusion of true causal paths, and large language models exhibit instability and illusion errors in causal directional reasoning.

Method used

We employ a non-causal prior-based data processing method, extracting non-causal prior knowledge from large language models (LLMs) through a structured causal hint framework (SCPF), and penalizing spurious edges in the causal graph adjacency matrix. Combined with an interactive inference engine and a penalty function, we suppress the data processing from moving towards erroneous edges.

Benefits of technology

It significantly improves the accuracy and robustness of causal discovery algorithms, effectively reduces the number of erroneous edges in high-dimensional scenarios, and enhances the accuracy and stability of causal graph construction without the need for expensive manual annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121706960A_ABST
    Figure CN121706960A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of knowledge processing, and discloses a data processing method and device based on non-causal priori, equipment and a storage medium, and the method comprises the steps: screening variable pairs with a non-causal relationship, and obtaining non-causal priori knowledge; on the basis of non-causal priori knowledge, data processing is inhibited from developing to an error edge, and the error edge is composed of variable pairs with a non-causal relationship. The device, the equipment and the storage medium all correspond to the method. According to the method, the performance of various causal discovery algorithms is improved, and the robust performance is shown in real application; non-causal edges are inhibited in the optimization process, so that the number of error edges caused by data noise is effectively reduced, especially in a high-dimensional scene; the reasoning ability of LLMs is combined with a data-driven causal discovery algorithm, and reliable constraints can be generated without expensive manual annotation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of knowledge processing technology, specifically a data processing method, apparatus, device, and storage medium based on non-causal priors. Background Technology

[0002] As data dimensionality increases, the adverse effects of data noise intensify significantly, leading to a substantial decrease in the learning accuracy of traditional causal discovery algorithms. This decline is primarily manifested in the introduction of spurious dependencies and the obfuscation of genuine causal paths. While Large Language Models (LLMs) may mitigate these issues through tacit knowledge extraction, existing methods mainly focus on causal directionality, neglecting the crucial role of non-causal priors.

[0003] Therefore, there is an urgent need for a new technical solution for data processing based on non-causal priors. Summary of the Invention

[0004] The purpose of this application is to provide a data processing method, apparatus, device and storage medium based on non-causal priors, so as to solve the technical problem of low data processing quality in high-dimensional noise scenarios in the prior art.

[0005] To achieve the above objectives, this application proposes a data processing method based on non-causal priors, the method comprising: By filtering out variable pairs with non-causal relationships, non-causal prior knowledge can be obtained; Based on the aforementioned non-causal prior knowledge, data processing is prevented from developing into erroneous edges, which are composed of variable pairs with non-causal relationships.

[0006] Preferably, the screening of variables with non-causal relationships is performed based on a preset structured causal hint framework, which performs causal relationship analysis based on an interactive reasoning engine to obtain the non-causal prior knowledge.

[0007] Preferably, the suppression of data processing from developing erroneous edges is based on penalizing false edges in the adjacency matrix of the causal graph, wherein the false edges consist of variable pairs that are processed as causal in the data processing but are actually non-causal.

[0008] Preferably, the structured causal suggestion framework includes: The instruction optimization module is used to define task boundaries and research objectives based on preset guidance strategies, and to limit the application scope of knowledge generation in data processing. The causal context building module is used to introduce the relationship between two variables in a variable pair; The metadata representation module is used to provide a structured description of variables, including variable name, variable description, and possible value states; The interactive reasoning engine is used to simulate the human reasoning process to derive the non-causal prior knowledge; A standardized output interface is used to generate a standard format, including variable names and response results, for output based on the aforementioned non-causal prior knowledge.

[0009] Preferably, the guidance strategy is subject to specific domain limitations.

[0010] Preferably, the penalty for the false edge includes restricting the data processing objective to an acyclicity constraint and a preset penalty function, which is used to limit the strength of the causal relationship between the two variables in a variable pair.

[0011] Preferably, the strength of the causal relationship between the two variables in the limiting variable pair is such that the strength of the limiting causal relationship tends to be non-causal.

[0012] To achieve the above objectives, this application also proposes a data processing apparatus based on non-causal priors, which applies the data processing method based on non-causal priors as described above. The apparatus includes: The filtering module is used to filter variable pairs with non-causal relationships to obtain non-causal prior knowledge; The suppression module is used to suppress data processing from developing into erroneous edges based on the non-causal prior knowledge, wherein the erroneous edges consist of pairs of variables with non-causal relationships.

[0013] To achieve the above objectives, this application also proposes an apparatus comprising at least one processor, at least one memory, and a data bus; The processor and the memory communicate with each other via the data bus; The memory stores program instructions that can be executed by the processor, which invokes the program instructions to execute the non-causal prior-based data processing method described above.

[0014] To achieve the above objectives, this application also proposes a storage medium on which a computer program is stored, which, when executed by a processor, implements the non-causal prior-based data processing method described above.

[0015] Beneficial effects: The data processing method, apparatus, device and storage medium based on non-causal priors in this application improve the performance of various causal discovery algorithms and demonstrate robust performance in real-world applications; by suppressing non-causal edges during the optimization process, the number of erroneous edges caused by data noise is effectively reduced, especially in high-dimensional scenarios; by combining the inference capabilities of LLMs with data-driven causal discovery algorithms, reliable constraints can be generated without expensive manual annotation. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart illustrating a non-causal prior-based data processing method provided in an embodiment of this application; Figure 2 A structural block diagram of the structured causal suggestion framework provided in the embodiments of this application; Figure 3 This is a schematic diagram illustrating the specific process of the LACD framework provided in the embodiments of this application; Figure 4 This is a template for an instruction optimization module provided in an embodiment of this application; Figure 5 This is a template for a causal context construction module provided in the embodiments of this application; Figure 6 This is a template for the metadata representation module provided in the embodiments of this application; Figure 7 This is a structural block diagram of a data processing device based on non-causal priors provided in an embodiment of this application.

[0018] The implementation, functional features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] The technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0020] In this document, the term "comprising" is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0021] Causal discovery is a method for identifying causal relationships between variables, with the primary goal of inferring the underlying causal graph structure from observational data. The mainstream methods in this field can be categorized into the following three paradigms: (1) Constraint-based causal discovery, such as the PC algorithm (Peter-Clark), which constructs a causal skeleton through conditional independence tests; (2) Scoring-based causal discovery, such as the GES algorithm (Greedy Equivalence Search), which uses scoring criteria such as Bayesian information criteria to optimize graph structure scores; (3) Causal discovery based on functional causal model.

[0022] These methods represent causal relationships by using the direction of edges in a Directed Acyclic Graph (DAG), which has significant application value in fields such as medical decision support systems and economic policy simulation.

[0023] However, traditional causal discovery methods typically face four main challenges when dealing with high-dimensional data: (1) Combinatorial explosion problem, that is, the search space of potential causal relationships grows exponentially with the increase of dimensionality; (2) Computational scalability constraints, because the complexity of high-dimensional conditional independence tests reaches ,in, It is the number of variables. It is the size of the condition set; (3) Structural equivalence class problem: The problem of unidentifiable causal direction caused by Markov equivalence class becomes more prominent as the data dimension increases; (4) Data quality degradation, data noise and data sparsity are common in high-dimensional scenarios and can affect the estimation of causal relationships.

[0024] To address these challenges, existing improvements, such as continuous optimization methods like the NOTEARS algorithm, learn directed acyclic graphs (DAGs) through numerical optimization and leverage deep learning to enhance causal identification capabilities. Other improvements, such as dimensionality reduction techniques and prior knowledge integration, have also demonstrated performance enhancements. However, their ability to identify causal relationships remains limited by the noise in high-dimensional observation data.

[0025] Incorporating prior knowledge is a promising approach to addressing the challenges of high-dimensional causal discovery. The integration of domain knowledge can effectively constrain the hypothesis space for causal structure learning while mitigating the risk of overfitting in data-driven causal discovery methods. Traditional methods typically employ hard constraints, such as pre-specified causal edges or time variable ordering, but their reliance on manual annotation limits scalability. Recent advances in language model-based causal discovery methods have enabled automated knowledge acquisition through semantic reasoning. However, inference instability and illusion errors still plague these methods, severely limiting their applicability to practical applications.

[0026] To facilitate the explanation of the non-causal prior-based data processing method in this embodiment, we will first describe the research progress of causal discovery algorithms and large language models in causal discovery.

[0027] We will specifically elaborate on the causal discovery method based on continuous optimization. Causal discovery based on continuous optimization aims to discover causal relationships from observational data. The learning function causal model (FCM) is as follows: in, This represents the number of independent and identically distributed samples. Represents a random vector The dimensions of the FCM. The FCM includes: A causal directed acyclic graph ,in, Represents a set of variables, directed edges Encoding causal relationships; A set of structural equations Map the parent variable to each node: in, express In a directed acyclic graph The parent variable set in It is a learnable function (linear or nonlinear). Represents mutually independent noise variables. Based on Markov properties, variables joint probability density It can be decomposed into: Based on equation (2), a causal graph was established. The equivalence relationship between the data generation process and the data generation process.

[0028] With the introduction of the functional causal model, the combinatorial graph search problem can be transformed into a continuous optimization problem. The aforementioned NOTEARS algorithm successfully transforms the traditional combinatorial optimization problem into a continuous optimization problem. Let... express The weighted adjacency matrix, where, Quantification from arrive The causal strength. Its goal is to minimize the observed data. Least squares loss between the data generated by the functional causal model and the data: To ensure the acyclicity of the DAG, the NOTEARS algorithm introduces a smooth acyclicity constraint. This transforms the loss function into an equality-constrained problem that can be optimized using the augmented Lagrange method: in, and These are used to control the intensity of penalties for constraint violations and to implement sparse regularization, respectively.

[0029] However, continuous optimization methods often violate the identifiability assumption, and their global optimization methods often produce incorrect local causal structures, generating DAGs with reverse or false causal edges.

[0030] Among the current advancements in the application of large-scale language models (LLMs) for causal discovery, LLMs have become a key tool due to their ability to extract implicit causal knowledge from unstructured text data. Preliminary studies have shown that LLMs can effectively infer pairwise causal relationships through semantic analysis, thus complementing traditional statistical methods. Subsequent methodological advances, such as CORR2CAUSE (Correlation to Causation), formalize causal inference by transforming causal graphs into natural language representations for LLM-based reasoning. Further developments, including COAT (Causal Representation On Assistance Tactics) and CMA (Causal Mediation Analysis), have further advanced the field by combining LLM's world knowledge with causal graph learning algorithms to address challenges in latent variable identification and metastructure modeling. Hybrid approaches, such as multi-agent systems and causal agent frameworks, improve robustness through modular reasoning and iterative optimization, achieving accuracy exceeding 80% in complex tasks. Recent innovations, such as Statistical Causal Prompt (SCP) and Retrieval-Augmented Generation (RAG), illustrate how the semantic abstraction capabilities of LLM can synergize with constraint-based methods, such as PC algorithms, to address the problems of data sparsity and sampling bias.

[0031] However, LLM's tendency to produce model illusions can lead to incorrect causal edges in its output. These models may generate spurious associations due to text co-occurrence bias (e.g., incorrectly linking "ice cream sales" with "drowning rate") or misjudge the direction of causality (e.g., being unsure whether education leads to income or income leads to education). Although corrective strategies like the pre-existing Meek rules can mitigate local errors, LLM remains sensitive to indirect causal relationships and spurious paths caused by latent variables. Furthermore, the causal assumptions generated by LLM often lack a verifiable semantic basis, making them difficult to correct using traditional statistical methods. Simultaneously, the causal graphs output by large language models are prone to loops, violating the acyclicity assumption. Therefore, utilizing non-causal prior knowledge from the output of large language models is another important research direction.

[0032] refer to Figure 1 , Figure 1 This is a flowchart of the non-causal prior-based data processing method in this embodiment.

[0033] To overcome the above limitations, such as Figure 1As shown, this embodiment discloses a data processing method based on non-causal priors, the method comprising: By filtering out variable pairs with non-causal relationships, non-causal prior knowledge can be obtained; Based on the aforementioned non-causal prior knowledge, data processing is prevented from developing into erroneous edges, which are composed of variable pairs with non-causal relationships.

[0034] This embodiment of the data processing method based on non-causal priors actually discloses an innovative framework called LLM-Augmented Causal Discovery (LACD). This framework shifts the focus from identifying causal relationships to identifying non-causal relationships in high-dimensional noisy scenes. The non-causal prior-based data processing of this framework formalizes the non-causal priors of language model inference into regularization constraints within a continuous optimization framework. Specifically, LACD introduces two core components: (1) The SCPF mechanism (Structured Causal Prompt Framework) systematically extracts non-causal variable pairs from large language models through structured prompts and semantic reasoning; (2) An optimization model based on constraint relaxation can dynamically penalize false edges in the causal adjacency matrix.

[0035] Specifically, regarding the SCPF mechanism, the filtering of variables with non-causal relationships is based on a preset structured causal hint framework. This structured causal hint framework performs causal relationship analysis based on an interactive reasoning engine to obtain the non-causal prior knowledge.

[0036] Reference Figure 2 , Figure 2 This is a structural block diagram of the structured cause-effect hint framework in this embodiment.

[0037] like Figure 2 As shown, in a preferred embodiment of this invention, the structured causal suggestion framework includes: The instruction optimization module is used to define task boundaries and research objectives based on a preset guidance strategy, and to limit the application scope of knowledge generation in data processing; wherein the guidance strategy is restricted to a specific domain. The causal context building module is used to introduce the relationship between two variables in a variable pair; The metadata representation module is used to provide a structured description of variables, including variable name, variable description, and possible value states; The interactive reasoning engine is used to simulate the human reasoning process to derive the non-causal prior knowledge; A standardized output interface is used to generate a standard format, including variable names and response results, for output based on the aforementioned non-causal prior knowledge.

[0038] Specifically, for the constraint-relaxation-based optimization model, the suppression of data processing from developing into erroneous edges is based on penalizing false edges in the adjacency matrix of the causal graph. The false edges consist of variable pairs that are processed as having a causal relationship in the data processing but are actually not causal.

[0039] In a preferred embodiment of this invention, the penalty applied to the false edge includes restricting the data processing objective to an acyclicity constraint and a preset penalty function. The penalty function is used to limit the strength of the causal relationship between the two variables in a variable pair. The limitation of the strength of the causal relationship between the two variables in a variable pair is such that the strength of the causal relationship tends to be non-causal.

[0040] Based on the above, we disclose a novel learning framework called LLM Augmented Causal Discovery (LACD) for processing data based on non-causal priors. This framework effectively integrates non-causal prior knowledge from LLM. Architecturally, LACD includes: (1) Non-causal prior knowledge is obtained from LLM using the Structured Causal Hint Framework (SCPF); (2) A constrained optimization framework that integrates non-causal prior knowledge during the causal graph learning process.

[0041] The application of this embodiment based on non-causal priors for data processing will now be described.

[0042] Reference Figure 3 , Figure 3 This is a schematic diagram illustrating the specific process of the LACD framework in this embodiment.

[0043] like Figure 3 As shown, we use a five-node causal discovery task as an example. For each variable pair, the SCPF framework generates a prompt template through the instruction optimization module, causal context construction module, and metadata representation module. This prompt template is then input into the interactive inference engine to build an agent to perform LLM inference. Finally, the LLM output is converted into a JSON file for the algorithm to read. After obtaining the LLM inference results, we extract information from variable pairs that respond "no causal relationship," such as... Figure 3 In V2-V5 and V4-V5, the constraints are transformed into non-causal prior constraints and learned together with causal discovery algorithms based on observation data to obtain more accurate results.

[0044] This embodiment discloses a Structured Causal Hint Framework (SCPF) for causal reasoning, aiming to improve the accuracy and interpretability of LLM in causal inference tasks through a systematic knowledge-guided mechanism. The framework consists of five core modules that together construct a complete causal reasoning logic chain.

[0045] Reference Figure 4 , Figure 4 This is the instruction optimization module template for this embodiment.

[0046] Instruction Optimization Module: This module employs domain-specific guidance strategies to clarify task boundaries and research objectives, and to limit the application scope of model knowledge generation. For example, in a medical diagnosis scenario, the construction of structured instruction templates would be as follows: Figure 4 As shown, the part within square brackets can be replaced according to different domain backgrounds. This design effectively limits knowledge generation to the target domain by restricting the scope of the professional field, thereby reducing cross-domain knowledge interference in the LLM response.

[0047] Reference Figure 5 , Figure 5 This is the template for the causal context building module in this embodiment.

[0048] Causal Context Building Module: This module introduces the possible relationships between two variables, providing richer contextual knowledge for LLM to improve recognition accuracy. Hint template as follows: Figure 5 As shown.

[0049] Reference Figure 6 , Figure 6 This is the template for the metadata representation module in this embodiment.

[0050] Metadata Representation Module: This module provides structured descriptions of research variables, including variable names, descriptions, and possible value states. Through detailed explanations of all variables, LLM enables deep semantic-based reasoning. Hint templates are as follows: Figure 6 As shown.

[0051] Interactive Inference Engine: This embodiment builds an iterative inference mechanism based on the existing ReAct (Reasoning and Action) framework to simulate the human reasoning process. The engine first receives the user's question and background information as input, then uses LLM to generate an initial inference path, followed by executing specific actions to collect additional information or verify the validity of the inference path, and integrating feedback information to update the inference path. This iterative process continues until a satisfactory solution is found or a preset number of iterations is reached, ultimately generating an inference result. This inference result includes at least the non-causal prior knowledge required by this embodiment.

[0052] Standardized output interface: Design a standardized output interface to generate JSON format output containing variable names and response results.

[0053] Based on the above applications, we have achieved the mining of non-causal prior knowledge. Next, we will explain the non-causal prior-guided optimization for causal discovery.

[0054] The mathematical formalization of the aforementioned continuously optimized causal discovery algorithm can be summarized as follows: in, Indicates through model parameters The learned directed acyclic graph (DAG). represent The weighted adjacency matrix, Represents observation data, It is the objective function to be optimized. This represents the acyclicity constraint. Here... Represented as: in, Represents a single sample The loss value, For the sample size, Represents the current cause-effect graph The regularization term, It is a hyperparameter that controls the strength of regularization.

[0055] Unlike existing orientation methods that utilize large language models to identify causal relationships between variables, we address the problem of causal structure learning in high-dimensional noisy environments by using large language models to determine non-causal relationships between variable pairs. The core innovation of this method lies in integrating the non-causal priors of the large language model, which enables the causal discovery algorithm to reliably identify erroneous directed edges and significantly improve the accuracy of causal graph construction.

[0056] Causal discovery algorithms based on continuous optimization typically represent the learning results as an adjacency matrix. In each round of training of the causal discovery algorithm, the adjacency matrix... Used to store the learning results of the current round. . The value is located at Within the interval, this value represents the variable. With variables The strength of the causal relationship between them.

[0057] Based on the SCPF framework, the set of non-causal variable pairs inferred from LLM can be obtained, denoted as: This set This represents variable pairs that are determined to be non-causal through LLM-based semantic analysis. For each variable pair... We use a penalty function Forced This means that the strength of the causal relationship is restricted to tend towards non-causal relationships, and L1 regularization penalty term is used here.

[0058] Based on the above description, the optimization problem of integrating non-causal prior knowledge of LLM is formulated as follows: To solve this constrained optimization problem, we employ the augmented Lagrangian method to reconstruct the objective function as follows: in, This is the penalty factor, with a default value of 1. The penalty mechanism is activated when a continuous optimization algorithm attempts to construct an edge connecting two variables in an iteration, but the large language model, based on its in-depth knowledge and semantic logic analysis, determines that there is no causal relationship between the two variables. This mechanism effectively suppresses the further development of this erroneous edge in subsequent iterations.

[0059] Based on the above, this embodiment constructs a complete LLM-enhanced causal discovery model to realize a data processing method based on non-causal priors. Furthermore, this embodiment provides partial experimental data to demonstrate that the designed non-causal priors can effectively guide the causal discovery algorithm to learn accurate causal structures. As a model-independent framework, our method can significantly improve the performance of continuously optimized baseline causal discovery methods in high-dimensional settings, particularly in terms of causal graph accuracy.

[0060] A. Experimental Setup Datasets: This suitable embodiment uses synthetic datasets and real-world datasets to verify the effectiveness of non-causal prior knowledge as constraints and the effectiveness of LLM prior knowledge, respectively.

[0061] 1. Synthetic Dataset: Randomly generated using the Erdos-Renyi graphical model, used to evaluate the effectiveness of non-causal prior constraints. The dataset exhibits good performance in terms of the number of variables (d={30,40,50}) and sparsity (ERk, where k represents the expected number of edges). The difference lies in d). The nonlinear relationships are modeled using a Gaussian process with a radial basis function kernel (bandwidth = 1), for each structure equation. All contain additive Gaussian noise .

[0062] 2. Real-world datasets: Two well-known benchmark datasets are used: Alarm (d=37 variables) and Hailfinder (d=56 variables).

[0063] Evaluation metrics: This embodiment uses four evaluation metrics to measure the performance of the causal discovery algorithm: FDR (false detection rate), TPR (true positive rate), FPR (false positive rate), and SHD (structural Hamming distance).

[0064] Baseline Methods: We selected three causal discovery algorithms as baseline methods for comparison. Notears-MLP employs a multilayer perceptron for continuous optimization to model nonlinear causal relationships. Gran-DAG utilizes generative adversarial networks to capture complex dependency structures, but its adversarial training mechanism may have scalability limitations in high-dimensional settings. DAG-GNN integrates graph neural networks with variational autoencoders for high-dimensional causal discovery, but is sensitive to noise or sparse graph topologies. These methods provide a comprehensive baseline comparison for evaluating our proposed framework.

[0065] B. Experimental Results of Synthetic Dataset To verify the effectiveness of the proposed non-causal prior constraints, we generated 10 synthetic datasets (each containing n=2000 samples) for each configuration based on the ER graph model. The real causal graphs provided information on non-existent edges for constraint derivation. The algorithm performance was evaluated and compared under unconstrained and constrained conditions. Under the constraints, all edges not present in the real graph were transformed into non-causal priors. For robust statistical evaluation, the Notears-MLP algorithm was used as the baseline model. The results for all ten datasets under each experimental configuration are summarized, and their means and standard deviations are reported, as shown in Table 1.

[0066] Table 1 Experimental results show that in high-dimensional datasets (i.e., scenarios with a large number of nodes), "Notears-MLP + Non-causal" exhibits significant performance improvements over "Notears-MLP" under different numbers of nodes and different sparsity settings. For example, in the ER2 model with d=50 nodes, both FDR and FPR drop to 0, TPR increases by 7.5%, and SHD decreases by 7%.

[0067] These results confirm the crucial role of prior knowledge in causal graph learning, especially for complex graph topologies and high-dimensional data scenarios. The proposed prior constraint formulation effectively enhances the performance of causal structure learning algorithms based on continuous optimization.

[0068] C. Experimental results using real-world datasets To evaluate the algorithm's ability to identify complex causal relationships in high-dimensional scenarios, we conducted comparative experiments on two real-world datasets representing different domains (Alarm and Hailfinder).

[0069] Alarm is a classic dataset widely used in the field of medical diagnosis, containing 37 variables and 46 directed edges, simulating causal relationships in the intensive care unit. Hailfinder comes from the field of weather forecasting, containing 56 variables and 66 directed edges, and aims to simulate complex causal relationships in hail forecasting systems. The selected datasets cover medium-dimensional (20≤d≤50) and high-dimensional (d>50) scenarios, enabling the algorithm performance to be evaluated from three key dimensions: (1) the number of variables, which determines the data dimension; (2) the number of directed edges, which reflects the sparsity of the causal graph; and (3) the generalization ability across different domains.

[0070] For the choice of LLM, we adopted the Deepseek-Reasoner model for variable causal reasoning, which supports context lengths of up to 128K and multi-turn dialogue reasoning.

[0071] The experimental results are shown in Tables 2 and 3.

[0072] Table 2 Table 3 Experimental results show that after introducing prior knowledge from LLM, Notears-MLP and Gran-DAG exhibit significant improvements in key metrics (including FDR, TPR, FPR, and SHD) across both datasets. This confirms the positive role of LLM prior knowledge in enhancing the performance of these two algorithms. However, for DAG-GNN, LACD performs slightly worse on the mid-dimensional Alarm dataset but significantly better on the higher-dimensional Hailfinder dataset. This indicates that DAG-GNN maintains strong performance on mid-dimensional data without prior knowledge, but its performance degrades in high-dimensional datasets due to noise interference, and introducing prior knowledge can significantly improve its performance in such cases.

[0073] Furthermore, after introducing non-causal priors, the TPR of Notears-MLP decreased from 0.2576 to 0.1667. This phenomenon may stem from the conservatism of the prior constraints: while effectively suppressing the generation of non-causal edges, non-causal priors also partially limit the model's ability to explore potential causal relationships. Nevertheless, the introduction of non-causal priors significantly reduced FDR and FPR, and decreased SHD by 34%, indicating that it has a positive effect on improving the overall performance of the model.

[0074] This finding validates the effectiveness of the LACD framework on high-dimensional datasets. Overall, LLM prior knowledge can provide effective guidance in real-world scenarios with high-dimensional data, thereby improving the performance of causal discovery algorithms, although the improvement may vary depending on the chosen baseline algorithm.

[0075] In summary, this embodiment's data-mathematical method based on non-causal priors, by integrating non-causal priors derived from LLMs within the LACD framework, addresses the key challenge of spurious dependencies caused by noise in high-dimensional causal structure learning. Experimental results on synthetic and real datasets demonstrate that LACD significantly outperforms state-of-the-art baseline methods in most cases, effectively reducing erroneous edges learned by causal discovery algorithms in high-dimensional scenarios and improving model performance. The innovation of this framework lies in its systematic transformation of LLM knowledge into non-causal constraints through the Structured Causal Hints Framework (SCPF), thereby suppressing erroneous edges during continuous optimization.

[0076] Based on the above, the data mathematical method based on non-causal priors in this embodiment achieves at least the following technical effects: 1. General Algorithm Enhancement Framework: The LACD framework improves the performance of various causal discovery algorithms (such as increasing the accuracy of DAG-GNN by 26.7% on the Hailfinder dataset) and demonstrates robust performance in real-world applications.

[0077] 2. Robustness to noisy data: By suppressing non-causal edges during the optimization process, LACD effectively reduces the number of erroneous edges caused by data noise, especially in high-dimensional scenarios (d≥50).

[0078] 3. Automated knowledge fusion: The SCPF framework combines the reasoning capabilities of LLMs with data-driven causal discovery algorithms, generating reliable constraints without the need for expensive manual annotation.

[0079] It should be noted that the non-causal prior-based data mathematical method in this embodiment is a general data processing method embedded in various data processing scenarios. It optimizes the processing of raw data, primarily by standardizing non-causal relationships and applying these standardized relationships to constrain data processing processes, such as causal discovery. This advantage enables various fields to obtain higher-quality basic data and data references during data processing, thereby improving data processing quality and user experience.

[0080] Reference Figure 7 , Figure 7 This is a structural block diagram of the data processing device based on non-causal priors in this embodiment.

[0081] like Figure 7As shown, this embodiment also discloses a data processing apparatus based on non-causal priors, which applies the data processing method based on non-causal priors as described above. The apparatus includes: The filtering module is used to filter variable pairs with non-causal relationships to obtain non-causal prior knowledge; The suppression module is used to suppress data processing from developing into erroneous edges based on the non-causal prior knowledge, wherein the erroneous edges consist of pairs of variables with non-causal relationships.

[0082] This embodiment also discloses a device, including at least one processor, at least one memory, and a data bus; The processor and the memory communicate with each other via the data bus; The memory stores program instructions that can be executed by the processor, which invokes the program instructions to execute the non-causal prior-based data processing method described above.

[0083] This embodiment also discloses a storage medium on which a computer program is stored, which, when executed by a processor, implements the data processing method based on non-causal prior as described above.

[0084] It should be noted that the data processing apparatus, device, and storage medium based on non-causal priors in this embodiment correspond to the aforementioned data processing method based on non-causal priors. Therefore, any content not specifically described in the data processing apparatus, device, and storage medium based on non-causal priors in this embodiment, including but not limited to functional definitions, working principles, and technical effects, can be referred to the description in the aforementioned data processing method based on non-causal priors, and will not be repeated here.

[0085] In the embodiments provided in this application, it should be understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, code, or any suitable combination thereof. For hardware implementation, the processor may be implemented in one or more of the following: application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, other electronic units designed to implement the functions described herein, or combinations thereof. For software implementation, some or all of the processes of the embodiments may be performed by a computer program instructing the associated hardware. During implementation, the program may be stored in a computer-readable storage medium or transmitted as one or more instructions or code on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein communication media include any medium that facilitates the transmission of a computer program from one place to another. Storage media may be any available medium accessible to a computer. Computer-readable storage media may include, but are not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code having the form of instructions or data structures and accessible to a computer.

[0086] Finally, it should be noted that the above description is only a preferred embodiment of this application and is not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A data processing method based on non-causal priors, characterized in that, The method includes: By filtering out variable pairs with non-causal relationships, non-causal prior knowledge can be obtained; Based on the aforementioned non-causal prior knowledge, data processing is prevented from developing into erroneous edges, which are composed of variable pairs with non-causal relationships.

2. The data processing method based on non-causal priors according to claim 1, characterized in that, The process of filtering variables with non-causal relationships is based on a preset structured causal hint framework. This structured causal hint framework performs causal relationship analysis based on an interactive reasoning engine to obtain the non-causal prior knowledge.

3. The data processing method based on non-causal priors according to claim 1, characterized in that, The suppression of data processing from developing erroneous edges is based on penalizing false edges in the adjacency matrix of the causal graph. These false edges consist of variable pairs that are processed as causal in the data processing but are actually non-causal.

4. The data processing method based on non-causal priors according to claim 2, characterized in that, The structured causal suggestion framework includes: The instruction optimization module is used to define task boundaries and research objectives based on preset guidance strategies, and to limit the application scope of knowledge generation in data processing. The causal context building module is used to introduce the relationship between two variables in a variable pair; The metadata representation module is used to provide a structured description of variables, including variable name, variable description, and possible value states; The interactive reasoning engine is used to simulate the human reasoning process to derive the non-causal prior knowledge; A standardized output interface is used to generate a standard format, including variable names and response results, for output based on the aforementioned non-causal prior knowledge.

5. The data processing method based on non-causal priors according to claim 4, characterized in that, The guidance strategy is subject to specific domain limitations.

6. The data processing method based on non-causal priors according to claim 3, characterized in that, The penalty imposed on the false edges includes restricting the data processing objective to acyclicity constraints and a preset penalty function, which limits the strength of the causal relationship between the two variables in a variable pair.

7. The data processing method based on non-causal priors according to claim 5, characterized in that, The strength of the causal relationship between the two variables in the restricted variable pair is such that the strength of the causal relationship tends to be non-causal.

8. A data processing apparatus based on non-causal priors, employing the data processing method based on non-causal priors as described in any one of claims 1 to 7, characterized in that, The device includes: The filtering module is used to filter variable pairs with non-causal relationships to obtain non-causal prior knowledge; The suppression module is used to suppress data processing from developing into erroneous edges based on the non-causal prior knowledge, wherein the erroneous edges consist of pairs of variables with non-causal relationships.

9. A device, characterized in that, Includes at least one processor, at least one memory, and a data bus; The processor and the memory communicate with each other via the data bus; The memory stores program instructions that can be executed by the processor, which invokes the program instructions to execute the data processing method based on non-causal priors as described in any one of claims 1 to 7.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the data processing method based on non-causal priors as described in any one of claims 1 to 7.