Causal analysis system and method for controlling variable experiment data

Through structured mapping coding and self-supervised feature distillation mechanisms, the problem of difficulty in distinguishing causal signals and noise is solved, and the robustness and reliability of causal analysis in complex experimental scenarios is improved, reducing the risk of misjudgment.

CN120234537AInactive Publication Date: 2025-07-01GUANGDONG OCEAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510485010.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When processing complex experimental data, it is difficult for the prior art to distinguish causal signals from random noise, resulting in false correlation misjudgment, and the setting of causal judgment thresholds lacks adaptability and insufficient robustness. It is difficult to accurately identify the true causal relationship in the case of high noise, small samples or complex interactions of variables.

Method used

The joint characterization of dependent variables and fruit variables is constructed through structured mapping encoding, and a self-supervised feature distillation mechanism based on deep learning is used to achieve accurate separation of noise and causal signals, and an adaptive decision-making system that links causal effect estimation and dynamic thresholds is established.

Benefits of technology

It improves the robustness of causal analysis, can accurately suppress noise while retaining nonlinear causal characteristics, reduce the risk of misjudgment in high-dimensional variable interactions or small sample data, and provides reliable causal inference support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234537A_ABST
    Figure CN120234537A_ABST
Patent Text Reader

Abstract

The invention discloses a causal analysis system and method for control variable experimental data, and the system achieves the precise stripping of noise and causal signals through the construction of the joint representation of dependent variables and fruit variables through the structural mapping coding, and the adoption of a self-supervision feature distillation mechanism based on deep learning. Meanwhile, a causal effect estimation and dynamic threshold linkage adaptive decision system is established. Through the mode, the causal analysis robustness in a complex experiment scene is improved, so that the system can realize accurate noise suppression on the premise of keeping nonlinear causal characteristics, and can adaptively adjust a judgment standard through coupling optimization of data noise reduction and effect estimation, so that the accuracy of noise suppression is improved. The misjudgment risk in high-dimensional variable interaction or small sample data is effectively reduced, and causal inference support with interpretability and reliability is provided for a control variable experiment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent analysis, and more specifically, to a causal analysis system and method for controlling variable experimental data. Background Art

[0002] In the field of experimental science and data analysis, the control variable method is widely used to infer causal relationships, but existing technologies have significant limitations when dealing with complex experimental data. Traditional methods usually rely on statistical hypothesis testing or regression analysis, which are highly sensitive to data noise. Especially in scenarios with multiple variables and nonlinear associations, noise interference can easily lead to false positive correlations. Existing data denoising techniques often use general filtering algorithms (such as wavelet transforms or low-pass filtering), but such methods have difficulty distinguishing the essential differences between causal signals and random noise in experimental data, and may over-smooth key causal features.

[0003] In addition, the existing technologies rely on empirical values ​​or fixed standards to set the threshold of causal judgment, and have not established an adaptive mechanism that is dynamically linked to the data noise reduction process, resulting in insufficient robustness in complex experimental scenarios. These defects jointly restrict the accuracy and interpretability of the control variable experimental data analysis, especially in the case of high noise, small samples or complex variable interactions, the existing system is difficult to reliably identify the true causal relationship.

[0004] Therefore, an optimized causal analysis system for controlled variable experimental data is desired. Summary of the invention

[0005] In order to solve the above-mentioned technical problems, the present application is proposed. The embodiment of the present application provides a causal analysis system and method for controlled variable experimental data, which constructs a joint representation of dependent variables and effect variables through structured mapping coding, and adopts a self-supervised feature distillation mechanism based on deep learning to achieve accurate separation of noise and causal signals, and at the same time establishes an adaptive decision-making system that links causal effect estimation with dynamic thresholds. In this way, the robustness of causal analysis in complex experimental scenarios is improved, so that the system can not only achieve accurate noise suppression while retaining nonlinear causal characteristics, but also adaptively adjust the judgment criteria through the coupling optimization of data denoising and effect estimation, effectively reducing the risk of misjudgment in high-dimensional variable interactions or small sample data, and providing causal inference support for controlled variable experiments that is both explanatory and reliable.

[0006] According to one aspect of the present application, a causal analysis system for control variable experimental data is provided, which includes: a data acquisition module, used to acquire control variable experimental data, the control variable experimental data including a set of {dependent variable, effect variable}; a noise reduction module, used to perform noise identification and noise removal on the control variable experimental data to obtain the noise-reduced control variable experimental data; a causal effect estimation module, used to input the noise-reduced control variable experimental data into the causal effect estimation module to obtain a causal effect value; and a causal analysis module, used to determine whether there is a causal relationship between the dependent variable and the effect variable based on a comparison between the causal effect value and a preset threshold.

[0007] According to another aspect of the present application, a causal analysis method for control variable experimental data is provided, which includes: obtaining control variable experimental data, the control variable experimental data including a set of {dependent variable, effect variable}; performing noise identification and noise removal on the control variable experimental data to obtain noise-reduced control variable experimental data; inputting the noise-reduced control variable experimental data into a causal effect estimation module to obtain a causal effect value; and determining whether there is a causal relationship between the dependent variable and the effect variable based on a comparison between the causal effect value and a preset threshold.

[0008] Compared with the prior art, the present application provides a causal analysis system and method for controlled variable experimental data, which constructs a joint representation of dependent variables and effect variables through structured mapping coding, uses a self-supervised feature distillation mechanism based on deep learning to achieve accurate separation of noise and causal signals, and establishes an adaptive decision-making system that links causal effect estimation with dynamic thresholds. In this way, the robustness of causal analysis in complex experimental scenarios is improved, so that the system can not only achieve accurate noise suppression while retaining nonlinear causal characteristics, but also adaptively adjust the judgment criteria through the coupling optimization of data denoising and effect estimation, effectively reducing the risk of misjudgment in high-dimensional variable interactions or small sample data, and providing causal inference support for controlled variable experiments that is both explanatory and reliable. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0010] Figure 1 4 is a block diagram of a causal analysis system for controlling variable experimental data according to an embodiment of the present application.

[0011] Figure 2Schematic diagram of data flow of a causal analysis system for controlling variable experimental data according to an embodiment of the present application.

[0012] Figure 3 It is a block diagram of a noise reduction module in a causal analysis system for controlling variable experimental data according to an embodiment of the present application.

[0013] Figure 4 4 is a block diagram of a noise distillation unit in a causal analysis system for controlling variable experimental data according to an embodiment of the present application.

[0014] Figure 5 The flowchart is a causal analysis method for controlling variable experimental data according to an embodiment of the present application. DETAILED DESCRIPTION

[0015] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.

[0016] As shown in this application and claims, unless the context clearly indicates an exception, the words "a", "an", "an" and / or "the" do not refer to the singular and may also include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0017] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are only illustrative, and different aspects of the system and method can use different modules.

[0018] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed accurately in order. On the contrary, various steps may be processed in reverse order or simultaneously as required. Meanwhile, other operations may also be added to these processes, or a certain step or several steps of operations may be removed from these processes.

[0019] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.

[0020] In the technical solution of the present application, a causal analysis system for controlling variable experimental data is proposed. Figure 1 A block diagram of a causal analysis system for controlling variable experimental data according to an embodiment of the present application. Figure 2 FIG. 1 is a data flow diagram of a causal analysis system for controlling variable experimental data according to an embodiment of the present application. Figure 1 and Figure 2 As shown, according to an embodiment of the present application, a causal analysis system 300 for controlling variable experimental data includes: a data acquisition module 310, used to acquire controlled variable experimental data, the controlled variable experimental data including a set of {dependent variable, effect variable}; a noise reduction module 320, used to perform noise identification and noise removal on the controlled variable experimental data to obtain the controlled variable experimental data after noise reduction; a causal effect estimation module 330, used to input the controlled variable experimental data after noise reduction into the causal effect estimation module to obtain a causal effect value; and a causal analysis module 340, used to determine whether there is a causal relationship between the dependent variable and the effect variable based on a comparison between the causal effect value and a preset threshold.

[0021] In particular, the data acquisition module 310 is used to acquire control variable experimental data, and the control variable experimental data includes a set of {dependent variables, effect variables}. Wherein, the dependent variable refers to a variable that is assumed to change due to other factors; in contrast, the effect variable is a result variable that is directly or indirectly caused by the change of the dependent variable. In one example, in a study on the effect of a certain drug on the cure rate of a disease, the use of the drug can be regarded as a dependent variable, and the cure rate of the patient is the effect variable. Therefore, acquiring a set of {dependent variables, effect variables} actually refers to systematically collecting data on all possible dependent variables and effect variables related to a specific causal relationship hypothesis. In particular, the noise reduction module 320 is used to perform noise identification and noise removal on the control variable experimental data to obtain the noise-reduced control variable experimental data. It should be understood that the acquired set of {dependent variables, effect variables} often inevitably contains various forms of noise. These noises may come from a variety of channels, such as measurement errors, environmental interference, inaccuracies in data records, etc. In particular, when dealing with complex, multivariate interactions or nonlinearly associated scenarios, the presence of noise will greatly affect the accuracy of causal relationship inference. Although traditional methods such as wavelet transform or low-pass filtering can remove noise to a certain extent, it is difficult to distinguish the essential difference between causal signals and random noise, which may cause key causal features to be over-smoothed or mistakenly deleted. Therefore, a more intelligent and precise method is needed to achieve effective separation of noise and causal signals. In the technical solution of the present application, by adopting deep learning technology combined with a self-supervised learning mechanism, it is possible to extract prior information from the global causal structure and construct a semantic separation space for noise and causal features, thereby accurately removing noise and retaining the true causal signal. In a specific example of the present application, Figure 3As shown, the noise reduction module 320 includes: a noise distillation unit 321, which is used to perform noise distillation on the control variable experimental data based on the cause-effect variable feature query prompt constraint based on the structured mapping joint encoding feature of the control variable experimental data and the global topological prior information between variables to obtain a noise distillation result; a control variable experimental data reconstruction unit 322, which is used to reconstruct the control variable experimental data based on the noise distillation result to obtain the denoised control variable experimental data.

[0022] Specifically, the noise distillation unit 321 is used to perform noise distillation on the control variable experimental data based on the cause-effect variable feature query prompt constraint based on the structured mapping joint encoding feature of the control variable experimental data and the global topology prior information between the variables to obtain the noise distillation result. In a specific example of the present application, Figure 4 As shown, the noise distillation unit 321 includes: a structured mapping encoding subunit 3221, which is used to perform structured mapping encoding on each {dependent variable, effect variable} in the set of {dependent variable, effect variable} to obtain a set of dependent variable-effect variable structured mapping joint encoding vectors; a noise screening subunit 3222, which is used to input the set of dependent variable-effect variable structured mapping joint encoding vectors into a noise identification and noise removal network based on a feature distillation mechanism to obtain a screened set of dependent variable-effect variable structured mapping joint encoding vectors as a noise distillation result.

[0023] More specifically, the structured mapping encoding subunit 3221 is used to perform structured mapping encoding on each {dependent variable, result variable} in the set of {dependent variable, result variable} to obtain a set of dependent variable-result variable structured mapping joint encoding vectors. It should be understood that traditional data analysis methods often rely on simple statistical tests or regression models to explore the relationship between variables, and this method is powerless when facing complex, nonlinear associations or multivariate interaction scenarios. In particular, in the presence of high dimensions and small sample sizes, simple methods are susceptible to noise interference, resulting in the emergence of pseudo-correlations. In the technical solution of the present application, through structured mapping encoding, the original dependent variables and result variables can be converted into a form that can be understood and processed by a computer, so as to effectively capture the nonlinear associations and complex interactions between variables, so that the potential causal relationship between cause-result variables can be more clearly presented in a multidimensional space. This conversion not only helps to improve the model's ability to understand complex data patterns, but also provides a suitable input format for subsequent deep learning algorithms. In one example, in medical research, the effect of a certain drug (dependent variable) on the cure rate of a specific disease (result variable) may not be linear, but the result of the combined effect of multiple factors such as the patient's age, gender, health status, etc. In this case, a simple linear model is difficult to accurately describe the mutual influence between these variables, and through structured mapping coding, a more comprehensive and detailed representation can be created to help researchers better understand the causal relationship between variables. In addition, structured mapping coding also helps to improve the interpretability of data. Compared with the original data, the encoded vector form is more easily used to explain the causal path and its intensity between variables, which is very valuable for scientific research. In a specific example of the present application, the specific steps of structured mapping coding for each {dependent variable, result variable} in the set of {dependent variable, result variable} are as follows: based on the dependent variable embedding matrix, the dependent variable is structured mapped to obtain a dependent variable embedding coding vector; based on the result variable embedding matrix, the result variable is structured mapped to obtain a result variable embedding coding vector; the dependent variable embedding coding vector and the result variable embedding coding vector are cascaded to obtain a dependent variable-result variable structured mapping joint coding vector.

[0024] More specifically, the noise screening subunit 3222 is used to input the set of dependent variable-result variable structured mapping joint encoding vectors into the noise identification and noise removal network based on the feature distillation mechanism to obtain the screening set of dependent variable-result variable structured mapping joint encoding vectors as the noise distillation result. It should be understood that the traditional denoising method is insufficient in distinguishing causal signals from noise, especially when there are complex nonlinear interactions between variables, and the denoising strategy based solely on statistical distribution or signal smoothing will destroy the structural characteristics of causal associations. Therefore, in the technical solution of the present application, the set of dependent variable-result variable structured mapping joint encoding vectors is input into the noise identification and noise removal network based on the feature distillation mechanism to obtain the screening set of dependent variable-result variable structured mapping joint encoding vectors. That is, by using self-supervised learning to extract prior information from the global causal structure, construct a semantic separation space of noise and causal features, guide the network to focus on the key dimensions of the causal signal through the query prompt mechanism, and capture the local dynamic consistency of the causal association with the help of the neighborhood weight coupling mechanism to solve the defect of homogenization of noise and causal features in traditional methods. In particular, in this process, the global prior constraints of self-supervised learning avoid the problem of traditional methods mistakenly deleting key causal features due to the lack of causal semantic guidance, ensuring the balance between noise removal and causal structure preservation; in addition, the dependency relationship between the local features of the dependent variable and the effect variable and the global context is dynamically established based on the query prompt mechanism, and the expression strength of the causal signal is simultaneously enhanced during the feature distillation process, such as strengthening the local consistency of the causal link through the neighborhood empathy gain mechanism, while filtering out redundant associations caused by random noise or collinearity. Through the precise screening of feature distillation, the denoised control variable experimental data not only eliminates irrelevant noise interference, but also retains the core pattern of causal mapping between cause and effect variables, providing a high-fidelity, low-bias input data basis for subsequent causal effect estimation.

[0025] Specifically, first, the set of dependent variable-effect variable structured mapping joint coding vectors is input into the cause-effect variable prior information self-supervised learning network to obtain the dependent variable-effect variable structured mapping prior information aggregate coding vector. It should be understood that the traditional denoising method has essential defects in capturing the global structural pattern between cause-effect variables. It relies on artificial preset rules or shallow statistical features, and cannot automatically mine the deep semantics of causal associations from unlabeled complex experimental data. Therefore, in the technical solution of the present application, the set of dependent variable-effect variable structured mapping joint coding vectors is input into the cause-effect variable prior information self-supervised learning network to obtain the dependent variable-effect variable structured mapping prior information aggregate coding vector. Through the self-supervised learning network, it is possible to construct a priori knowledge framework of the global causal topology between variables represented by the dependent variable-effect variable structured mapping, so as to dig out the potential distribution law of causal variables and the global context association, such as the implicit relationship between the temporal dependence and the covariate interaction path between variables. In this process, the self-supervised learning network performs unsupervised training on the dependent variable-effect variable structured mapping joint encoding vector through predefined tasks (such as mask prediction or context reconstruction), so that the network learns the topological constraints of the causal conduction path under unsupervised conditions, thereby generating a dependent variable-effect variable structured mapping prior information aggregation encoding vector that integrates global context and implicit relationship features. In a specific example of the present application, the set of dependent variable-effect variable structured mapping joint encoding vectors is input into the cause-effect variable prior information self-supervised learning network with the following self-supervised learning formula to obtain the dependent variable-effect variable structured mapping prior information aggregation encoding vector; wherein, the self-supervised learning formula is: ;in, is the set of dependent variable-result variable structured mapping joint encoding vectors, and They are the first, second, and third vectors in the set of dependent variable-result variable structured mapping joint encoding vectors. and The dependent variable-result variable structured mapping joint encoding vector, For Conduct self-supervision of the prior information of the dependent variable and the outcome variable, and Take The maximum and minimum values ​​in is the adjustment parameter, that is, to prevent the denominator from being zero, It is the first in the set of dependent variable-result variable structured mapping characteristic dynamic range coefficients. The dynamic range coefficient of the structured mapping feature of the dependent variable-result variable, It is the first in the set of dynamic range weight coefficients of the dependent variable-result variable structured mapping feature. A dependent variable - outcome variable structured mapping feature dynamic range weight coefficient, is the number of vectors in the middle vector, and is a dependent variable - outcome variable structured mapping prior information aggregation encoding vector.

[0026] Next, the k-th dependent variable - outcome variable structured mapping joint encoding vector is extracted from the set of dependent variable - outcome variable structured mapping joint encoding vectors as the dependent variable - outcome variable structured mapping joint encoding feature vector to be distilled. It should be understood that traditional noise recognition methods usually process the dataset as a whole and cannot perform refined analysis on the local feature differences of different dependent - outcome variable pairs. Especially when the variable interaction is complex or the noise distribution is uneven, the global noise reduction strategy is prone to misjudgment of key causal features or local noise residue. Therefore, in the technical solution of this application, the k-th dependent variable - outcome variable structured mapping joint encoding vector is extracted from the set of dependent variable - outcome variable structured mapping joint encoding vectors as the dependent variable - outcome variable structured mapping joint encoding feature vector to be distilled. By separately extracting and making distillation decisions for each dependent variable - outcome variable structured mapping joint encoding vector one by one, the fine - grained noise recognition can be achieved, enabling the system to dynamically evaluate the causal association strength of different variable pairs by combining the global prior knowledge of the dependent - outcome variables. This refined processing mode provides a local - global collaborative optimization data basis for subsequent adaptive threshold determination, thus significantly improving the robustness of causal inference in complex experimental scenarios. In a specific example of this application, the k-th dependent variable - outcome variable structured mapping joint encoding vector is extracted from the set of dependent variable - outcome variable structured mapping joint encoding vectors as the dependent variable - outcome variable structured mapping joint encoding feature vector to be distilled according to the following extraction formula, where the extraction formula is: ; where is the -th dependent variable - outcome variable structured mapping joint encoding vector in the set of dependent variable - outcome variable structured mapping joint encoding vectors, and is taking the -th dependent variable - outcome variable structured mapping joint encoding vector as the dependent variable - outcome variable structured mapping joint encoding feature vector to be distilled.

[0027] Further, based on the dependent variable - outcome variable structured mapping prior information aggregation encoding vector, a cause - effect variable feature query hint constraint is imposed on the joint encoding feature vector of the dependent variable - outcome variable structured mapping to be distilled to determine whether to perform distillation on the joint encoding feature vector of the dependent variable - outcome variable structured mapping to be distilled. In an embodiment of the present application, first, the joint encoding feature vector of the dependent variable - outcome variable structured mapping to be distilled and the dependent variable - outcome variable structured mapping prior information aggregation encoding vector are input into the cause - effect variable feature distillation query hint network to obtain a cause - effect variable structured mapping feature distillation query hint implicit encoding matrix. Since traditional noise recognition methods often analyze the structured encoding of individual variable pairs in isolation when dealing with local causal association features, ignoring their semantic consistency with the global causal topology, it leads to misjudgment of implicit noise (such as pseudo - correlation signals caused by confounding factors). In the technical solution of the present application, the joint encoding feature vector of the dependent variable - outcome variable structured mapping to be distilled and the dependent variable - outcome variable structured mapping prior information aggregation encoding vector are input into the cause - effect variable feature distillation query hint network to obtain a cause - effect variable structured mapping feature distillation query hint implicit encoding matrix. That is, through the cause - effect variable feature distillation query hint network, a dynamic interaction mechanism between the local features of the dependent variable - outcome variable to be distilled and the global causal prior of the dependent variable - outcome variable is established, and the causal conduction mode (such as intervention effect strength or time - lag correlation) of the local cause - effect variable pair is implicitly associated with the global structure (such as causal graph path dependence or population - level statistical laws across cause - effect variable pairs) using query hints. Specifically, taking the joint encoding feature vector of the dependent variable - outcome variable structured mapping to be distilled as the query signal, it performs cross - scale interaction with the dependent variable - outcome variable structured mapping prior information aggregation encoding vector carrying global semantics to construct a multi - dimensional mapping relationship (such as the spatio - temporal continuity of the causal conduction path or the synergistic effect of cause - effect variable interaction) between the local and global cause - effect variables in the feature space, and then generates a cause - effect variable structured mapping feature distillation query hint implicit encoding matrix to quantify the compatibility between the local cause - effect variable features and the global causal context. This implicit modeling of global - local feature linkage enables the network to identify noise based on the integrity of causal semantics rather than simply statistical metrics, effectively solving the problem that traditional methods may misdelete weak causal signals or retain implicit noise due to the separation of local and global associations, providing a highly discriminative decision basis for subsequent adaptive distillation, and ultimately improving the signal - to - noise ratio and interpretability of the denoised data in causal effect estimation. In a specific example of the present application, the joint encoding feature vector of the dependent variable - outcome variable structured mapping to be distilled and the dependent variable - outcome variable structured mapping prior information aggregation encoding vector are input into the cause - effect variable feature distillation query hint network according to the following query hint formula to obtain a cause - effect variable structured mapping feature distillation query hint implicit encoding matrix; where the query hint formula is: ; where For and to perform feature distillation query hint processing, is matrix multiplication, is the transposed vector of is the length of is the dependent variable - outcome variable structured mapping feature distillation query hint implicit coding matrix corresponding to

[0028] Next, based on the dependent variable - outcome variable structured mapping feature distillation query prompt implicit coding matrix, determine the dependent variable - outcome variable structured mapping feature distillation intensity descriptor. Preferably, in the technical solution of this application, first, perform local - global feature fusion optimization based on neighborhood empathy on the dependent variable - outcome variable structured mapping feature distillation query prompt implicit coding matrix to obtain an optimized dependent variable - outcome variable structured mapping feature distillation query prompt implicit coding matrix; further, perform matrix trace metric on the optimized dependent variable - outcome variable structured mapping feature distillation query prompt implicit coding matrix to obtain the dependent variable - outcome variable structured mapping feature distillation intensity descriptor. In particular, considering that the dependent variable - outcome variable structured mapping feature distillation query prompt implicit coding matrix itself serves as a high - dimensional semantic representation carrier, although it contains the interaction relationship between causal local features and global causal priors, its complex internal structure is difficult to directly apply to dynamic decision - making. Traditional noise recognition methods rely on fixed rules or single statistics for feature screening and are difficult to adapt to the dynamic coupling characteristics of noise and causal signals in local - global associations in complex experimental scenarios. Especially when there are spatio - temporal heterogeneities or implicit confoundings in the cause - effect variable interaction path, static evaluation indicators are prone to overfitting or underfitting. In the technical solution of this application, by constructing the dependent variable - outcome variable structured mapping feature distillation intensity descriptor, it is possible to quantitatively fuse the local neighborhood consistency (such as the temporal stability of causal conduction) and global semantic compatibility (such as the causal graph topological constraints across variable pairs) contained in the high - dimensional dependent variable - outcome variable structured mapping feature distillation query prompt implicit coding matrix. For example, use neighborhood autocorrelation weights to capture the static structural coherence of the causal path, and combine cosine - type cross - weights to depict the causal flow consistency of dynamic interactions between variables, thereby generating a distillation decision basis that can comprehensively reflect the noise interference intensity and causal feature integrity. Specifically, in this process, model the implicit association between local features and global context at the distribution sensitivity level through an empathy gain coefficient matrix, so as to transform the complex neighborhood dependence relationship in the dependent variable - outcome variable structured mapping feature distillation query prompt implicit coding matrix into an interpretable numerical index, enabling the network to dynamically adjust the distillation strategy according to the optimization potential of local causal features (such as the semantic deviation caused by noise residue) and the repair requirements of the global causal chain (such as the interpretability loss caused by path breakage). In a specific example, it is possible to reconstruct features with enhanced key nodes of causal conduction for high - noise interference, while selectively retaining features with low noise but weak semantic consistency. In this way, the high - dimensional dependent variable - outcome variable structured mapping feature distillation query prompt implicit coding matrix can be mapped to a low - dimensional and operable scalar index, providing a quantitative basis for subsequent adaptive distillation decisions of causal features. In addition, through local empathy gain interpretation analysis, effectively balance the contradiction between the global causal framework and local feature specificity, laying a high - fidelity data foundation for subsequent causal effect estimation.

[0029] More specifically, in this example, based on the dependent variable - outcome variable structured mapping feature distillation query prompt implicit coding matrix, the following optimization is used to determine the dependent variable - outcome variable structured mapping feature distillation intensity descriptor: where the optimization formula is: ; where and are respectively and the corresponding dependent variable - outcome variable structured mapping feature distillation query prompt implicit coding matrices, is the concatenation operation, is to calculate the Frobenius norm, is the corresponding dependent variable - outcome variable structured mapping neighborhood autocorrelation weight coefficient, is the corresponding dependent variable - outcome variable structured mapping neighborhood cosine - type cross - weight coefficient, is the element - wise multiplication by position, is the corresponding dependent variable - outcome variable structured mapping empathy gain coefficient matrix, is the optimized dependent variable - outcome variable structured mapping feature distillation query prompt implicit coding matrix, is the trace metric value of the matrix, is the corresponding dependent variable - outcome variable structured mapping feature distillation intensity descriptor.

[0030] Subsequently, based on the dependent variable - outcome variable structured mapping feature distillation intensity descriptor, it is determined whether to distill the dependent variable - outcome variable structured mapping joint - encoded feature vector to be distilled. It should be understood that traditional noise reduction methods adopt a unified processing strategy and cannot make adaptive decisions according to the dynamic differences in the causal association strength and noise interference degree of individual causal variable pairs. Especially in scenarios where the causal variable interaction is complex or the sample distribution is uneven, the global fixed threshold is likely to cause key weak causal signals to be deleted by mistake or high - noise segments to remain. In the technical solution of this application, by dynamically analyzing the dependent variable - outcome variable structured mapping feature distillation intensity descriptor, an accurate balance between noise suppression and causal feature retention is achieved. For example, in variable pairs with high noise but clear causal conduction paths (such as significant intervention effects but disturbed by measurement errors), the degree of damage to the causal topology caused by noise is quantified based on the descriptor, and distillation is initiated to strengthen the global semantic consistency of local features; while in variable pairs with low noise but weak causal associations (such as marginal conduction paths or small - sample sparse associations), it is determined based on the descriptor to skip distillation to avoid over - suppressing real signals. That is, the self - supervised query - hint mechanism maps the dependent variable - outcome variable structured mapping feature distillation intensity descriptor to a dynamic decision threshold (such as a confidence interval adaptively adjusted according to the experimental scenario), enabling the system to flexibly select the distillation intensity according to the current noise - causal mixed state of the causal variable pair. In this way, the system can suppress random noise while maximizing the integrity and interpretability of the causal conduction path, significantly reducing the over - fitting or under - fitting risks caused by static processing in traditional methods. At the same time, by skipping redundant distillation operations, the computing resource allocation is optimized, and the causal inference efficiency and conclusion reliability of the system in high - noise, small - sample, or variable - interaction - complex scenarios are improved. In a specific example of this application, based on the dependent variable - outcome variable structured mapping feature distillation intensity descriptor, the following screening formula is used to determine whether to distill the dependent variable - outcome variable structured mapping joint - encoded feature vector to be distilled; where the screening formula is: ; where is used to determine whether to distill the dependent variable - outcome variable structured mapping joint - encoded feature vector to be distilled, is a preset threshold.

[0031] Specifically, the control variable experimental data reconstruction unit 322 is configured to reconstruct the control variable experimental data based on the noise distillation result to obtain the denoised control variable experimental data. It should be understood that although the encoded vectors screened by the feature distillation network have removed noise and retained the causal association features, their representation form is vectors in a high-dimensional abstract space and cannot be directly used for causal effect estimation or downstream analysis based on the original data structure. When dealing with such problems, traditional methods often directly model in the encoded space, but ignore the interpretability requirements of the data source domain (such as experimental observations or variable physical meanings), resulting in the disconnection between the causal inference results and the actual scenario. By mapping to the data source domain, the abstract encoded vectors can be restored to denoised data that conforms to the original experimental data distribution and the semantics of cause-effect variables, so as to ensure that the structural features of causal variables (such as the non-linear dependence or temporal association between variables) are not distorted after mapping. In the technical solution of this application, an inverse encoder is used to map each cause-effect variable structured mapping joint encoded vector in the screening set of the cause-effect variable structured mapping joint encoded vectors to the data source domain to obtain the denoised control variable experimental data. Specifically, in this process, through the parametric transformation or generative model (such as an adversarial training network) of the inverse encoder, the screening set of the cause-effect variable structured mapping joint encoded vectors is projected back to the original data dimension. In this way, the data reconstructed by source domain mapping not only eliminates noise interference but also retains the key statistical attributes of the causal path, enabling the subsequent causal effect estimation module to model based on interpretable observations.

[0032] In particular, the causal effect estimation module 330 is used to input the denoised control variable experimental data into the causal effect estimation module to obtain the causal effect value. It should be understood that when the traditional causal effect estimation method directly relies on the original data modeling, it is susceptible to noise interference or variable coupling, resulting in the estimation result being biased towards pseudo-correlation or ignoring weak but real causal signals. The denoised control variable experimental data retains the causal topological characteristics between the cause-effect variables through structured mapping encoding, and removes random noise and redundant associations through self-supervised distillation, so that the causal effect estimation module can focus on the essential association pattern between the cause-effect variables. By inputting the denoised control variable experimental data into the causal effect estimation module, the quantitative representation of the causal effect can be extracted from the high-fidelity data, solving the estimation bias problem caused by insufficient data quality. Here, the causal effect estimation module converts the structured mapping relationship between the cause-effect variables into a computable causal effect indicator by parsing the intrinsic patterns of these encoding vectors. Specifically, this module quantifies the independent influence of each dependent variable on the outcome variable by constructing a causal graph model or a potential outcome framework, for example, by calculating the conditional probability change or marginal effect value of the dependent variable on the outcome variable under the condition of controlling other variables. In this process, the causal effect estimation module adaptively adjusts parameters through dynamic modeling (such as conditional probability inference based on causal graphs or effect decomposition of potential outcome frameworks) combined with the distribution characteristics of the experimental data of the control variables after noise reduction, such as introducing nonlinear kernel functions to capture complex interaction effects, or using attention mechanisms to weight key causal paths. It is worth mentioning that by inputting noise-reduced data, the interference of noise on the effect value calculation is significantly reduced, avoiding the estimation distortion caused by collinearity or noise accumulation in traditional methods in high-dimensional variable scenarios. The model can parse the implicit causal dependencies (such as nonlinear or high-order interactions) between variables from structured coding, and output more explanatory effect values ​​and confidence intervals, thereby providing a robust and interpretable quantitative basis for causal relationship determination.

[0033] In particular, the causal analysis module 340 is used to determine whether there is a causal relationship between the dependent variable and the result variable based on the comparison between the causal effect value and the preset threshold. Considering that multiple factors may affect the occurrence of a certain result at the same time, and there may be interactions between these factors. In this case, it is difficult to accurately judge which variables have a true causal relationship based on intuition or simple statistical methods. By comparing the causal effect value with the preset threshold, pseudo-correlations caused by random fluctuations or noise interference can be effectively filtered out, thereby focusing on relationships that are truly causally significant. In a specific example of the present application, in response to the causal effect value being greater than or equal to the preset threshold, it is determined that there is a causal relationship between the dependent variable and the result variable; in response to the causal effect value being less than the preset threshold, it is determined that there is no causal relationship between the dependent variable and the result variable.

[0034] As described above, the causal analysis system 300 for controlling variable experiment data according to an embodiment of the present application can be implemented in various wireless terminals, such as a server having a causal analysis algorithm for controlling variable experiment data, etc. In a possible implementation manner, the causal analysis system 300 for controlling variable experiment data according to an embodiment of the present application can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the causal analysis system 300 for controlling variable experiment data can be a software module in the operating system of the wireless terminal, or can be an application program developed for the wireless terminal; of course, the causal analysis system 300 for controlling variable experiment data can also be one of the numerous hardware modules of the wireless terminal.

[0035] Alternatively, in another example, the causal analysis system 300 for controlling variable experiment data and the wireless terminal can also be separate devices, and the causal analysis system 300 for controlling variable experiment data can be connected to the wireless terminal through a wired and / or wireless network, and transmit interaction information according to a predefined data format.

[0036] Furthermore, a causal analysis method for controlling variable experiment data is also provided.

[0037] Figure 5 FIG. is a flowchart of a causal analysis method for controlling variable experiment data according to an embodiment of the present application. As Figure 5 shown, the causal analysis method for controlling variable experiment data according to an embodiment of the present application includes the steps of: S1, obtaining control variable experiment data, where the control variable experiment data includes a set of {dependent variable, outcome variable}; S2, performing noise identification and noise removal on the control variable experiment data to obtain denoised control variable experiment data; S3, inputting the denoised control variable experiment data into a causal effect estimation module to obtain a causal effect value; S4, based on the comparison between the causal effect value and a preset threshold, determining whether there is a causal relationship between the dependent variable and the outcome variable.

[0038] In summary, the causal analysis method for controlling variable experiment data according to an embodiment of the present application is elucidated. It constructs a joint representation of the dependent variable and the outcome variable through structured mapping encoding, uses a self-supervised feature distillation mechanism based on deep learning to accurately separate noise from causal signals, and simultaneously establishes an adaptive decision-making system that links causal effect estimation with a dynamic threshold. In this way, the robustness of causal analysis in complex experimental scenarios is improved, enabling the system to not only achieve accurate noise suppression while retaining non-linear causal characteristics, but also adaptively adjust the judgment criteria through the coupling optimization of data denoising and effect estimation, effectively reducing the misjudgment risk in high-dimensional variable interactions or small-sample data, and providing causal inference support with both interpretability and reliability for control variable experiments.

[0039] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A causal analysis system for controlled variable experimental data, characterized in that: include: A data acquisition module is used to acquire control variable experimental data, where the control variable experimental data includes a set of {dependent variables, effect variables}; A noise reduction module is used to identify and remove noise from the control variable experimental data to obtain the noise-reduced control variable experimental data; A causal effect estimation module, used for inputting the denoised control variable experimental data into the causal effect estimation module to obtain a causal effect value; The causal analysis module is used to determine whether there is a causal relationship between the dependent variable and the effect variable based on the comparison between the causal effect value and the preset threshold; wherein the denoising module includes: a noise distillation unit, which is used to perform noise distillation on the control variable experimental data based on the cause-effect variable feature query prompt constraint based on the structured mapping joint encoding feature of the control variable experimental data and the global topological prior information between the variables to obtain the noise distillation result; a control variable experimental data reconstruction unit is used to reconstruct the control variable experimental data based on the noise distillation result to obtain the denoised control variable experimental data.

2. The causal analysis system for controlled variable experimental data according to claim 1, characterized in that: The noise distillation unit includes: a structured mapping encoding subunit, which is used to perform structured mapping encoding on each {dependent variable, effect variable} in the set of {dependent variable, effect variable} to obtain a set of dependent variable-effect variable structured mapping joint encoding vectors; a noise screening subunit, which is used to input the set of dependent variable-effect variable structured mapping joint encoding vectors into a noise identification and noise removal network based on a feature distillation mechanism to obtain a screened set of dependent variable-effect variable structured mapping joint encoding vectors as a noise distillation result.

3. The causal analysis system for controlled variable experimental data according to claim 2, characterized in that: The structured mapping coding subunit is used to: perform structured mapping on the dependent variable based on the dependent variable embedding matrix to obtain a dependent variable embedding coding vector; perform structured mapping on the result variable based on the result variable embedding matrix to obtain a result variable embedding coding vector; and cascade the dependent variable embedding coding vector and the result variable embedding coding vector to obtain a dependent variable-result variable structured mapping joint coding vector.

4. The causal analysis system for controlled variable experimental data according to claim 2, characterized in that: The noise screening subunit includes: a secondary subunit for extracting prior information between variables, which is used to input a set of dependent variable-effect variable structured mapping joint coding vectors into a cause-effect variable prior information self-supervised learning network to obtain a dependent variable-effect variable structured mapping prior information aggregate coding vector; a secondary subunit for extracting variables to be distilled, which is used to extract the kth dependent variable-effect variable structured mapping joint coding vector from the set of dependent variable-effect variable structured mapping joint coding vectors as the dependent variable-effect variable structured mapping joint coding feature vector to be distilled; and a secondary subunit for feature distillation, which is used to perform a cause-effect variable feature query prompt constraint on the dependent variable-effect variable structured mapping joint coding feature vector to be distilled based on the dependent variable-effect variable structured mapping prior information aggregate coding vector.

5. The causal analysis system for controlled variable experimental data according to claim 4, characterized in that: The feature distillation secondary sub-unit includes: a feature distillation query prompt tertiary sub-unit, which is used to input the dependent variable-result variable structured mapping joint encoding feature vector to be distilled and the dependent variable-result variable structured mapping prior information aggregation encoding vector into the cause-result variable feature distillation query prompt network to obtain the dependent variable-result variable structured mapping feature distillation query prompt implicit encoding matrix; a feature distillation strength calculation tertiary sub-unit, which is used to determine the dependent variable-result variable structured mapping feature distillation strength descriptor based on the dependent variable-result variable structured mapping feature distillation query prompt implicit encoding matrix; a feature distillation judgment tertiary sub-unit, which is used to determine whether to distill the dependent variable-result variable structured mapping joint encoding feature vector to be distilled based on the dependent variable-result variable structured mapping feature distillation strength descriptor.

6. The causal analysis system for controlled variable experimental data according to claim 5, characterized in that: The feature distillation strength calculation three-level sub-unit is used to: perform local-global feature fusion optimization based on neighborhood empathy on the dependent variable-result variable structured mapping feature distillation query prompt implicit coding matrix to obtain the optimized dependent variable-result variable structured mapping feature distillation query prompt implicit coding matrix; perform matrix-based trace measurement on the optimized dependent variable-result variable structured mapping feature distillation query prompt implicit coding matrix to obtain the dependent variable-result variable structured mapping feature distillation strength descriptor.

7. The causal analysis system for controlled variable experimental data according to claim 1, characterized in that: The control variable experimental data reconstruction unit is used to: use an inverse encoder to map each dependent variable-result variable structured mapping joint encoding vector in the screening set of dependent variable-result variable structured mapping joint encoding vectors to a data source domain to obtain denoised control variable experimental data.

8. The causal analysis system for controlled variable experimental data according to claim 1, characterized in that: In response to the causal effect value being greater than or equal to a preset threshold, it is determined that there is a causal relationship between the dependent variable and the result variable; in response to the causal effect value being less than the preset threshold, it is determined that there is no causal relationship between the dependent variable and the result variable.

9. A causal analysis method for controlled variable experimental data, characterized in that: include: Obtaining control variable experimental data, the control variable experimental data includes a set of {dependent variables, result variables}; Noise identification and noise removal are performed on the control variable experimental data to obtain the noise-reduced control variable experimental data; the noise-reduced control variable experimental data is input into the causal effect estimation module to obtain the causal effect value; based on the comparison between the causal effect value and the preset threshold, it is determined whether there is a causal relationship between the dependent variable and the result variable.