Dynamic diagnosis and early warning system for regional property education adaptation degree based on causal inference model
Through systematic analysis using causal inference models, the problem of dynamic diagnosis and early warning of regional industry-education fit has been solved, achieving efficient causal diagnosis and intelligent early warning, and supporting the optimization of regional industry-education resources and policy adjustments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-19
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies are insufficient for dynamically and accurately diagnosing the fit between industry and education at the regional level, and the early warning mechanisms lack sensitivity and accuracy, failing to effectively support the optimization of educational resources and policy adjustments.
The system employs a causal inference model, including data fusion and feature engineering, adaptive confounding variable selection, counterfactual reasoning diagnosis, and dynamic early warning decision-making. Through causal-oriented analysis methods combined with a multi-level early warning mechanism, it achieves dynamic diagnosis and intelligent early warning of industry-education fit.
It has improved the causality of industry-education fit diagnosis and the sensitivity of early warning, supported differentiated and forward-looking resource allocation and policy adjustments, and enhanced the scientific and intelligent level of regional industry-education integration governance.
Smart Images

Figure CN121808642A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of educational economics and data science, and in particular to a dynamic diagnosis and early warning system for regional industry-education fit based on a causal inference model. Background Technology
[0002] With the deepening implementation of my country's economic restructuring and upgrading and innovation-driven development strategy, the deep integration of industry and higher education has become a key path to enhance regional competitiveness and promote high-quality employment. Regional industry-education fit, that is, the degree of matching between the talent cultivation output of higher education in a specific region and the dynamic needs of local industrial development, is a core indicator for measuring the effectiveness of industry-education integration. Accurately diagnosing the industry-education fit status and providing early warnings of potential imbalances has significant decision-making support value for education authorities to optimize resource allocation, universities to adjust their professional settings and training models, and enterprises to plan their talent strategies. Currently, the assessment of industry-education fit largely relies on macro-statistical data comparison, questionnaire surveys, or expert experience judgment. While these methods can provide certain descriptive insights, they generally have the following limitations: First, they are mostly static or post-hoc analyses, relying on yearbook-like aggregated data, resulting in poor timeliness and difficulty in achieving dynamic monitoring and forward-looking early warning; second, the analytical methods mainly focus on correlation descriptions, such as calculating the correlation coefficient between the scale of graduates of a certain major and the scale of a certain industry, but correlation does not equal causation. Because of the presence of numerous confounding variables (such as regional economic level, technological development cycle, policy environment, and students' personal choices), drawing conclusions simply based on correlations is prone to misjudgment and cannot answer causal questions such as "whether adjusting the enrollment plan for a certain major can really improve the fit." Finally, industry-education data are usually characterized by high dimensionality (many variables), sparsity (some indicator data are missing), and small sample size (based on cities or districts, with limited sample size). Traditional regression or machine learning models are prone to overfitting or falling into the "curse of dimensionality," making it difficult to robustly identify the real key driving factors.
[0003] Within the existing technological framework, some studies attempt to introduce more complex models. For example, some solutions utilize big data technology to build industry-education integration information platforms, enabling the aggregation and display of supply and demand data; others employ virtual simulation technology to construct teaching systems, optimizing teaching task scheduling by analyzing student behavior sequences. However, the former essentially remains within the realm of data visualization and descriptive statistics, lacking in-depth causal analysis and diagnostic capabilities; the latter focuses on micro-level behavioral optimization in the teaching process, and its technological logic (such as node adjustment based on frequency and failure rate) is suitable for task process improvement, but cannot solve the problem of causal identification in the context of multiple intertwined factors at the macro-regional level. Furthermore, existing early warning mechanisms are mostly based on simple thresholds (such as employment rates below a certain fixed value), failing to comprehensively consider trend changes, volatility, and regional heterogeneity, resulting in insufficient sensitivity and accuracy in early warning. These problems hinder the assessment of regional industry-education fit from "phenomenon description" to "causal diagnosis" and "precise intervention," affecting the scientific nature of industry-education integration policy formulation and the effectiveness of resource allocation.
[0004] Therefore, there is an urgent need in this field to develop a system and method capable of processing high-dimensional, mixed observational data, overcoming the limitations of small samples, dynamically and accurately diagnosing regional industry-education integration from a causal inference perspective, and providing intelligent and differentiated early warnings based on the diagnostic results. This requires breaking through the traditional statistical analysis framework and innovatively combining cutting-edge causal inference models with the specific problems of regional industry-education integration to design a complete analysis and decision support system with a solid mathematical and statistical foundation. Summary of the Invention
[0005] The purpose of this invention is to provide a dynamic diagnosis and early warning system for regional industry-education fit based on a causal inference model, so as to solve the problems existing in the prior art.
[0006] To achieve the above objectives, the present invention provides the following solution: This invention provides a dynamic diagnosis and early warning system for regional industry-education fit based on a causal inference model, comprising: The data fusion and feature engineering module is used to acquire and integrate regional industrial data, educational data, and socio-economic data, standardize and construct features from the data, and output standardized feature matrices, fitness target variables, and time series labels. The adaptive hybrid variable selection module connects the data fusion and feature engineering modules. It receives the feature matrix and target variables, selects core variables with potential causal relationships to fitness from high-dimensional features based on the causal-oriented adaptive LASSO algorithm, and outputs the selected feature subset and the corresponding causal prior weights. The counterfactual reasoning diagnostic engine module, connected to the adaptive confounding variable selection module, receives the selected feature subset and target variable. Based on the enhanced causal forest and dual robust estimation framework, it estimates the average treatment effect and heterogeneous treatment effect under different industry-education policies or conditions, and outputs the fit diagnosis results, including causal effect value, confidence interval and individual treatment effect. The dynamic early warning decision module is connected to the counterfactual reasoning diagnosis engine module. It is used to receive the fit diagnosis results and historical time series data, calculate the fit trend and volatility through a sliding window, and generate multi-level early warning signals in combination with the dynamic threshold algorithm. The visualization and interpretation module connects the counterfactual reasoning diagnostic engine module and the dynamic early warning decision module. It is used to visualize diagnostic results, causal paths and early warning signals, and provide interpretive analysis of model decisions.
[0007] Preferably, in the adaptive heterogeneous variable selection module, the variable selection method includes: Construct a prior cause-effect graph based on domain knowledge or using PC algorithms to determine the initial causal structure among variables and form a set of prior causal variables. Perform ordinary least squares regression on the standardized feature matrix and the target variable to obtain preliminary coefficient estimates; Calculate the adaptive weights; Solve the adaptive LASSO optimization problem with causal constraints, with the objective function as follows: ; in, For the fitness target variable vector, For the standardized feature matrix, For the feature coefficient vector, For the total number of features, , The regularization parameter is determined through cross-validation. For adaptive weights, , The coefficients are those estimated by ordinary least squares. For weight sensitivity parameters, For indicator functions, For the a priori causal set; The selected variables are subjected to significance testing based on the post-selection inference method, and finally a stable feature subset is output.
[0008] Preferably, in the counterfactual reasoning diagnostic engine module, the improved method of the enhanced causal forest model for small sample data includes: During the node splitting process of each causal tree, a Bayesian regularization term is introduced. The splitting criterion, based on the original objective of maximizing the heterogeneity of treatment effects, adds a penalty for insufficient node sample size. The modified splitting criterion function is as follows: ; in, As a measure of the heterogeneity of treatment effects within nodes, The number of samples in the current node. , To adjust hyperparameters; By utilizing observational data from similar regions for transfer learning, and adapting the features of the source domain data to the distribution of processing variables, the training effect of the target domain model is enhanced.
[0009] Preferably, in the counterfactual reasoning diagnostic engine module, the estimator of the average treatment effect calculated using the dual robust estimation framework is defined as: ; in, For the total sample size, This is a binary variable, representing whether or not to accept specific industry-education policy interventions. For the observed fitness results, For the first The covariate vector of each sample, The propensity score is estimated using a logistic regression or machine learning model. The conditional mean results are estimated using an enhanced causal forest. ∈{0,1}; The estimator is modified as follows for small samples: ; in, For effective sample size, For the number of covariates, For the deviation correction term based on the bootstrap method, This is the variance-stabilizing term.
[0010] Preferably, the specific steps for generating the early warning signal in the dynamic early warning decision module include: Set the sliding time window length to In the window =[ , Trends and fluctuations in the internally calculated fitness index: ; ; in, For at a certain point in time The fit index value, For window Mean of intrinsic fitness; Calculate the dynamic early warning threshold, which is composed of a dynamic baseline and an adaptive boundary: ; ; in, For at a certain point in time The dynamic baseline value, As a smoothing factor, This represents the actual fitness observation value at the previous time point. For at a certain point in time The lower warning threshold, For risk parameters, This represents the number of samples within the window. Based on the comparison of current fit, trends and thresholds, multi-level warnings are triggered.
[0011] Preferably, in the data fusion and feature engineering module, the constructed adaptation target variable is a multi-dimensional comprehensive index, and its calculation formula is as follows: ; Among them, weight , , , The degree of job matching is determined by entropy weight method or expert scoring method. It is calculated based on the professional matching rate and salary matching rate. The degree of skill matching is calculated based on the skill coverage rate and skill gap rate. The degree of time and space matching is calculated based on the matching rate of graduates' geographical distribution and job search time delay. The degree of development matching is calculated based on the job promotion space index and the continuing education opportunity index.
[0012] Preferably, the specific method for providing causal path explanation in the visualization and explanation module is as follows: Based on the prior causal graph output by the adaptive confounding variable selection module and the treatment effect estimated by the counterfactual reasoning diagnostic engine module, the Shapley value decomposition method is used to quantify the contribution of each feature variable to the final fit diagnosis result and the individual treatment effect estimation. Complex causal effects are visualized using directed acyclic graphs and waterfall diagrams. Nodes in the diagrams represent variables, edges represent causal relationships and the magnitude of effects, and key causal paths are highlighted.
[0013] Preferably, the system further includes: The model validation and sensitivity analysis module, connected to the counterfactual reasoning diagnostic engine module, is used to perform placebo tests, estimate the "spurious treatment effect" multiple times by randomizing the variables, compare it with the true effect, and calculate the significance value; it performs Rosenbaum boundary sensitivity analysis to quantify the strength of unmeasured confounding factors in the observed data that would be required to overturn the existing causal conclusions; and it employs bootstrap resampling technology to generate confidence intervals for causal effect estimation through repeated sampling, quantifying the uncertainty of the model estimation.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements all or part of the functions of the regional industry-education adaptation dynamic diagnosis and early warning system based on the causal inference model.
[0015] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements all or part of the functions of the regional industry-education adaptation dynamic diagnosis and early warning system based on a causal inference model.
[0016] The present invention achieves the following beneficial technical effects compared to the prior art: This invention provides a dynamic diagnosis and early warning system for regional industry-education integration based on a causal inference model. Addressing the core challenges of complex and limited sample sizes in industry-education data, it employs a causal-oriented adaptive LASSO variable selection mechanism and a dual robust estimator with small-sample correction, effectively improving the model's stability and interpretability in high-dimensional sparse data environments. The system explicitly defines integration diagnosis as a counterfactual causal inference problem, achieving a leap from descriptive statistics to causal diagnosis. A multi-level early warning mechanism integrating trend analysis and dynamic thresholds is designed to sensitively identify abnormal declines in integration and potential risks, supporting differentiated and proactive interventions by management departments. The entire system not only provides powerful data analysis capabilities but also intuitively presents complex causal conclusions through visualization and interpretation modules, significantly enhancing the scientific, precise, and intelligent level of regional industry-education integration governance. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1The framework diagram of the regional industry-education adaptation dynamic diagnosis and early warning system based on the causal inference model provided by the present invention is shown. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] The purpose of this invention is to provide a dynamic diagnosis and early warning system for regional industry-education fit based on a causal inference model, so as to solve the problems existing in the prior art.
[0021] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] Example 1: like Figure 1 The diagram shown illustrates the framework of the regional industry-education adaptation dynamic diagnosis and early warning system based on a causal inference model provided by this invention. This system achieves a closed loop from data integration to intelligent early warning through a data-driven causal inference process. The specific implementation of the system will be described in detail below, in conjunction with this framework.
[0023] The system's implementation begins with the aggregation and structured processing of multi-source data. The data fusion and feature engineering module is responsible for extracting multi-dimensional raw data from regional statistical bureaus, education departments, human resources and social security platforms, and industrial economic databases. This data includes, but is not limited to: the number of job openings in various industries and specialties, job skill requirement tags, median graduate salaries, the number of graduates from various universities and majors, curriculum systems and credit data, student internship and practical participation rates, and macroeconomic indicators such as regional GDP and R&D investment ratios. This module first performs time alignment and spatial matching on the data, followed by missing value imputation and outlier cleaning. Based on this, the module constructs composite features according to business logic, such as calculating "skill coverage matching degree" as a sub-indicator of skill matching degree. Finally, all features are standardized using Z-scores to form a standardized feature matrix.
[0024] Simultaneously, this module calculates the core fitness target variable. In this embodiment, the fitness target variable is a multi-dimensional comprehensive index, and its calculation formula is as follows: ; Among them, weight , , , The degree of job matching is determined using the entropy weight method or expert scoring method. It is calculated based on the professional relevance rate and salary matching rate; the skill matching rate is calculated based on the skill coverage rate and skill gap rate; the spatiotemporal matching rate is calculated based on the graduate geographical distribution matching rate and job search time delay; and the career development matching rate is calculated based on the job promotion space index and continuing education opportunity index. For example, for data from a city in the Yangtze River Delta in 2023, its job matching degree can be calculated by weighting the employment rate and salary matching rate of graduates majoring in "Computer Science and Technology" in the information industry that year. , , , The values are objectively determined using the entropy weight method, for example, by assigning values of 0.35, 0.30, 0.20, and 0.15 respectively, to satisfy the condition that their sum equals 1.
[0025] Further, the preprocessed high-dimensional data is input into the adaptive confounding variable selection module. The core task of this module is to overcome the "curse of dimensionality" caused by high dimensionality and small sample size, and to screen out core variables with potential causal relationships, laying the foundation for subsequent causal inference. The implementation process first constructs a priori causal structure graph based on domain expert knowledge or using the PC (Peter-Clark) causal discovery algorithm. For example, experts might consider "government investment in industry-education integration special funds" as a potential cause affecting "fitness," while "regional per capita GDP" could be both a cause and a confounding factor. These prior-identified variables form a set. Next, the module performs ordinary least squares regression on the high-dimensional data to obtain preliminary coefficient estimates for each feature. Adaptive weights are then calculated using these estimates. , The coefficients are those estimated by ordinary least squares. For the weight sensitivity parameter, in this implementation, we take... =2, which reduces the pressure on large-coefficient features during the penalty, thereby reducing estimation bias. Subsequently, the module solves the adaptive LASSO optimization problem with causal constraints, whose objective function is: ; in, For the fitness target variable vector, For the standardized feature matrix, For the feature coefficient vector, For the total number of features, , The regularization parameter is determined through cross-validation. For adaptive weights, , The coefficients are those estimated by ordinary least squares. For weight sensitivity parameters, For indicator functions, Given a prior causal set. In this function, the first term is the fitting loss; the second term is the standard adaptive LASSO penalty, used to induce sparsity, with parameters... The third point, and the innovation of this invention, is to control the overall sparsity; it imposes an additional penalty on variables that are not in the prior causal set. This allows for the prioritization of theoretically supported causal variables in variable selection, enhancing the model's interpretability and stability. Regularization parameters and Ten-fold cross-validation is used to select variables from a pre-defined grid to minimize the mean squared error of the model on the validation set. Finally, a significance test is performed on the selected variable subset using a post-selection inference method, removing variables with p-values greater than 0.1, and outputting the final stable feature subset with significantly reduced dimensionality and its corresponding causal prior information.
[0026] Furthermore, after obtaining the refined feature set, the counterfactual reasoning diagnostic engine module begins to perform the core causal effect estimation. This module is based on a hybrid framework of enhanced causal forest and dual robust estimation. First, an enhanced causal forest model is constructed. Traditional causal forests are prone to overfitting when splitting tree nodes when dealing with small sample data. To address this, this invention introduces a Bayesian regularization term into the node splitting criterion for each tree. Specifically, during splitting, instead of solely pursuing the maximization of heterogeneity of the processed effects, a modified gain function is used: ; in, As a measure of the heterogeneity of treatment effects within nodes, The number of samples in the current node. , To adjust the hyperparameters, they were set to 0.05 and 0.1 respectively in this implementation. This correction term penalizes splitting when the sample size is small or the number of features is large, effectively avoiding the generation of overly fragmented leaf nodes and improving the model's robustness under small sample conditions. Furthermore, to address the issue of insufficient samples in specific regions, the module introduces a transfer learning mechanism. For example, data from the Pearl River Delta region is used as the source domain for pre-training the model, and then its parameters are transferred to the Yangtze River Delta region for fine-tuning. This adapts the feature distribution and improves the training effect of the target domain model.
[0027] Furthermore, using a trained augmented causal forest, the conditional expectation of the potential outcome for each sample under both treatment and non-treatment conditions can be estimated, i.e. and Simultaneously, propensity scores are estimated using an independent classification model (such as logistic regression). Subsequently, a dual robust estimation method was used to calculate the average treatment effect. The estimator was defined as: ; in, For the total sample size, This is a binary variable, representing whether or not to accept specific industry-education policy interventions. For the observed fitness results, For the first The covariate vector of each sample, The propensity score is estimated using a logistic regression or machine learning model. The conditional mean results are estimated using an enhanced causal forest. ∈{0,1}. The advantage of this estimator is that as long as either the bias towards the scoring model or the outcome model is correctly specified, the estimation result will be consistent, hence the name "double robustness". For small sample scenarios, this invention further modifies this estimator: ; in, For effective sample size, For the number of covariates, For the deviation correction term based on the bootstrap method, This is the variance-stabilizing term; and These are the bias correction term and variance stabilization term, calculated through 500 bootstrapping resampling operations. This correction reduces estimation errors in high-dimensional cases with small samples. Finally, this module outputs the average treatment effect estimate, confidence interval, and heterogeneous treatment effect for each region under different policy interventions, forming a complete fit diagnosis report. For example, the report might state: "After controlling for other variables, implementing the 'Smart Manufacturing Industry Academy' policy can improve the industry-education fit in this region by an average of 0.15 units, with a more significant improvement effect on districts and counties with a better industrial base." Furthermore, the dynamic early warning decision module receives the adaptation time series output by the diagnostic engine and performs real-time monitoring and early warning. This module employs a sliding window mechanism, setting the window length k=12. At each time point t, the window size is calculated. =[ , Fitting trends and fluctuations within: ; ; in, For at a certain point in time The fit index value, For window The mean of the intrinsic fitness. The warning threshold is not a fixed value, but is composed of a dynamic baseline and an adaptive boundary. The dynamic baseline is calculated using exponential smoothing. ; in, For at a certain point in time The dynamic baseline value, This is a smoothing factor with a value of 0.3. This represents the actual fitness observation value at the previous time point. For at a certain point in time The lower warning threshold is calculated as follows: ; in, For risk parameters, This represents the number of samples within the window. The warning logic is: if the current fit At < And the trend If |At-| is less than 0, a red alert is triggered, indicating that the fit has fallen below the dynamic safety boundary and is still deteriorating, requiring immediate activation of the intervention plan; if |At-| is less than 0, a red alert is triggered, indicating that the fit has fallen below the dynamic safety boundary and is still deteriorating, requiring immediate activation of the intervention plan; |>1.5⋅ And | If |>θ (the trend threshold θ is set to 0.02 / month), a yellow alert is triggered, indicating a significant risk of deviation or abnormal fluctuation, requiring closer monitoring and analysis; otherwise, the system status is green and normal.
[0028] Furthermore, the system's visualization and interpretation module transforms the complex diagnostic and early warning results into an intuitive visual interface. Based on a selected subset of features and estimated treatment effects, this module uses the Shapley value decomposition method to quantify the specific contribution of each feature variable (such as "software industry job growth rate" and "proportion of university AI courses") to the final fit score and individual treatment effect estimation, and displays it in the form of waterfall charts. Simultaneously, the module integrates prior knowledge with data-driven causal structures, drawing a directed acyclic graph (DAG). Nodes in the graph represent variables, and the thickness of the edges represents the magnitude of the causal effect. Key causal paths such as "government investment → training base construction → skills matching → comprehensive fit" are highlighted with colors, enabling decision-makers not only to know "what" but also to understand "why."
[0029] Furthermore, to ensure the reliability of the diagnostic conclusions, the system also includes a model validation and sensitivity analysis module. This module performs three key analyses: First, a placebo test is conducted, which involves randomly shuffling the assignment of the treatment variable and repeating the causal inference process 1000 times to obtain the distribution of the "spurious treatment effect." If the estimated true treatment effect is located at the extremes of this distribution, it indicates that the result is not accidental. Second, a Rosenbaum boundary sensitivity analysis is performed. This analysis shows that an unobserved confounding factor with a strength of 2.5 is required to change the currently significant causal conclusion to insignificant one, which indirectly proves the robustness of the conclusion under moderate unobserved confounding. Third, a bootstrap method is used to generate confidence intervals for the causal effect estimate through repeated sampling with replacement, quantifying the uncertainty of the model estimate.
[0030] Through the specific implementation methods described above, this invention constructs a complete technical closed loop from data collection and causal diagnosis to dynamic early warning. The modules are closely integrated, replacing traditional subjective experience-based judgments with a rigorous mathematical framework for causal inference. This effectively solves problems such as difficulty in causal identification, delayed early warnings, and insufficient decision-making basis in regional industry-education integration assessment, providing a scientific and intelligent tool for the precise governance of regional industry-education integration.
[0031] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0032] It should be noted that the components mentioned in the above embodiments are all general standard parts or components known to those skilled in the art. Their structures and principles can be learned by those skilled in the art through technical manuals or conventional experimental methods.
[0033] This invention has used specific examples to illustrate its principles and implementation methods. The above descriptions of the embodiments are only for the purpose of helping to understand the method and core ideas of this invention. Furthermore, those skilled in the art will recognize that, based on the ideas of this invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this invention.
Claims
1. A dynamic diagnosis and early warning system for regional industry-education fit based on a causal inference model, characterized in that, include: The data fusion and feature engineering module is used to acquire and integrate regional industrial data, educational data, and socio-economic data, standardize and construct features from the data, and output standardized feature matrices, fitness target variables, and time series labels. The adaptive hybrid variable selection module connects the data fusion and feature engineering modules. It receives the feature matrix and target variables, selects core variables with potential causal relationships to fitness from high-dimensional features based on the causal-oriented adaptive LASSO algorithm, and outputs the selected feature subset and the corresponding causal prior weights. The counterfactual reasoning diagnostic engine module, connected to the adaptive confounding variable selection module, receives the selected feature subset and target variable. Based on the enhanced causal forest and dual robust estimation framework, it estimates the average treatment effect and heterogeneous treatment effect under different industry-education policies or conditions, and outputs the fit diagnosis results, including causal effect value, confidence interval and individual treatment effect. The dynamic early warning decision module is connected to the counterfactual reasoning diagnosis engine module. It is used to receive the fit diagnosis results and historical time series data, calculate the fit trend and volatility through a sliding window, and generate multi-level early warning signals in combination with the dynamic threshold algorithm. The visualization and interpretation module connects the counterfactual reasoning diagnostic engine module and the dynamic early warning decision module. It is used to visualize diagnostic results, causal paths and early warning signals, and provide interpretive analysis of model decisions.
2. The regional industry-education adaptation dynamic diagnosis and early warning system based on a causal inference model according to claim 1, characterized in that, The variable selection method in the adaptive heterogeneous variable selection module includes: Construct a prior cause-effect graph based on domain knowledge or using PC algorithms to determine the initial causal structure among variables and form a set of prior causal variables. Perform ordinary least squares regression on the standardized feature matrix and the target variable to obtain preliminary coefficient estimates; Calculate the adaptive weights; Solve the adaptive LASSO optimization problem with causal constraints, with the objective function as follows: ; in, For the fitness target variable vector, For the standardized feature matrix, For the feature coefficient vector, For the total number of features, , The regularization parameter is determined through cross-validation. For adaptive weights, , The coefficients are those estimated by ordinary least squares. For weight sensitivity parameters, For indicator functions, For the a priori causal set; The selected variables are subjected to significance testing based on the post-selection inference method, and finally a stable feature subset is output.
3. The regional industry-education adaptation dynamic diagnosis and early warning system based on a causal inference model according to claim 1, characterized in that, The counterfactual reasoning diagnostic engine module includes the following improvements to the enhanced causal forest model for small sample data: During the node splitting process of each causal tree, a Bayesian regularization term is introduced. The splitting criterion, based on the original objective of maximizing the heterogeneity of treatment effects, adds a penalty for insufficient node sample size. The modified splitting criterion function is as follows: ; in, As a measure of the heterogeneity of treatment effects within nodes, The number of samples in the current node. , To adjust hyperparameters; By utilizing observational data from similar regions for transfer learning, and adapting the features of the source domain data to the distribution of processing variables, the training effect of the target domain model is enhanced.
4. The regional industry-education adaptation dynamic diagnosis and early warning system based on a causal inference model according to claim 3, characterized in that, In the counterfactual reasoning diagnostic engine module, the estimator of the average treatment effect calculated using the dual robust estimation framework is defined as follows: ; in, For the total sample size, This is a binary variable, representing whether or not to accept specific industry-education policy interventions. For the observed fitness results, For the first The covariate vector of each sample, The propensity score is estimated using a logistic regression or machine learning model. The conditional mean results are estimated using an enhanced causal forest. ∈{0,1}; The estimator is modified as follows for small samples: ; in, For effective sample size, For the number of covariates, For the deviation correction term based on the bootstrap method, This is the variance-stabilizing term.
5. The regional industry-education fit dynamic diagnosis and early warning system based on a causal inference model according to claim 1, characterized in that, The specific steps for generating early warning signals in the dynamic early warning decision module include: Set the sliding time window length to In the window =[ , Trends and fluctuations in the internally calculated fitness index: ; ; in, For at a certain point in time The fit index value, For window Mean of intrinsic fitness; Calculate the dynamic early warning threshold, which is composed of a dynamic baseline and an adaptive boundary: ; ; in, For at a certain point in time The dynamic baseline value, As a smoothing factor, This represents the actual fitness observation value at the previous time point. For at a certain point in time The lower warning threshold, For risk parameters, This represents the number of samples within the window. Based on the comparison of current fit, trends and thresholds, multi-level warnings are triggered.
6. The regional industry-education fit dynamic diagnosis and early warning system based on a causal inference model according to claim 1, characterized in that, In the data fusion and feature engineering module, the constructed fit target variable is a multi-dimensional comprehensive index, and its calculation formula is as follows: ; Among them, weight , , , The degree of job matching is determined by entropy weight method or expert scoring method. It is calculated based on the professional matching rate and salary matching rate. The degree of skill matching is calculated based on the skill coverage rate and skill gap rate. The degree of time and space matching is calculated based on the matching rate of graduates' geographical distribution and job search time delay. The degree of development matching is calculated based on the job promotion space index and the continuing education opportunity index.
7. The regional industry-education fit dynamic diagnosis and early warning system based on a causal inference model according to claim 1, characterized in that, The specific method for providing causal path explanations in the visualization and explanation module is as follows: Based on the prior causal graph output by the adaptive confounding variable selection module and the treatment effect estimated by the counterfactual reasoning diagnostic engine module, the Shapley value decomposition method is used to quantify the contribution of each feature variable to the final fit diagnosis result and the individual treatment effect estimation. Complex causal effects are visualized using directed acyclic graphs and waterfall diagrams. Nodes in the diagrams represent variables, edges represent causal relationships and the magnitude of effects, and key causal paths are highlighted.
8. The regional industry-education adaptation dynamic diagnosis and early warning system based on a causal inference model according to claim 1, characterized in that, The system also includes: The model validation and sensitivity analysis module, connected to the counterfactual reasoning diagnostic engine module, is used to perform placebo tests, estimate the "spurious treatment effect" multiple times by randomizing the variables, compare it with the true effect, and calculate the significance value; perform Rosenbaum boundary sensitivity analysis to quantify the strength of unmeasured confounding factors in the observed data that would be required to overturn the existing causal conclusions; and employ bootstrap resampling technology to generate confidence intervals for causal effect estimation through repeated sampling, quantifying the uncertainty of the model estimation.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements all or part of the functions of the regional industry-education adaptation dynamic diagnosis and early warning system based on the causal inference model as described in any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements all or part of the functions of the regional industry-education adaptation dynamic diagnosis and early warning system based on the causal inference model as described in any one of claims 1-8.