Social security loss reason tracing method based on causal inference model
By applying a method based on a causal inference model in the analysis of social insurance loss, using the causal directional scoring function and self-supervised learning method, a causal relationship network is built and causal effects is quantified, and the problems of insufficient causal inference ability and complexity of data processing in the existing technology are solved, and a high-accurate tracing of social insurance loss causes and policy formulation support is achieved.
Patent Information
- Application Number
- CN202510370601.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing social security loss analysis technology relies too much on correlation analysis, lacks causal inference ability, and is difficult to accurately trace the causes of loss. Traditional machine learning methods require a large number of manual feature engineering, it is difficult to process complex and heterogeneous social security data, rely on supervised learning, and it is difficult to build an effective model in the absence of labeled data. The existing causal inference methods are insufficiently used in social security loss analysis, and lacks efficient and automated causal relationship identification methods.
The cause traceability method of social insurance loss based on the causal inference model is adopted, and the causal directional scoring function combines social insurance payment behavior and time series constraints to ensure the rationality and accuracy of causal inference. The self-supervised learning method is used to automatically learn and extract the potential causal characteristics in social insurance data, build a causal relationship network and quantify the causal effects of each factor.
It improves the accuracy and interpretability of social security loss analysis, can identify factors that directly and indirectly affect social security loss, quantify the causal effects of various factors, provide scientific policy formulation basis, and reduce the social security loss rate.
Smart Images

Figure CN120235633A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of social security, and in particular, to a method for tracing the causes of social security loss based on a causal inference model. Background Art
[0003] Currently, the research on social security loss mainly relies on statistical analysis and traditional data mining methods. For example, some studies use regression analysis and decision tree machine learning methods to extract features from historical data and construct a social security loss prediction model. Although it can identify the key factors of social security loss to a certain extent, there are still great limitations.
[0004] First of all, the existing methods usually find influencing factors based on correlation analysis, and fail to fully consider the causal relationship between variables. For example, some variables are only the surface features of social security loss, rather than the real causes, resulting in poor interpretability of the prediction results and being difficult to effectively guide policy-making. Secondly, traditional machine learning methods often rely on a large amount of manual feature engineering when dealing with social security data, and social security data is usually heterogeneous, multi-source, and unstructured. Manually constructing features is not only time-consuming and laborious, but also omits key information. In addition, most existing studies rely on supervised learning methods, which require a large amount of labeled data for training. However, in social security management, it is difficult to obtain high-quality labeled data, which affects the generalization ability of the model.
[0005] In recent years, causal inference methods have gradually become an important tool for solving complex social and economic problems. Compared with traditional statistical analysis, causal inference can describe the causal relationship between variables, rather than just staying at the correlation level. However, the current causal inference methods applied to social security loss analysis are still relatively limited, mainly relying on expert experience to construct causal networks, lacking an automated and efficient data-driven modeling method. In addition, the existing causal inference methods are usually based on supervised learning, and for such complex social problems as social security loss, the research on unsupervised causal inference is still in the exploratory stage, and it is difficult to effectively quantify the causal effects of various influencing factors.
[0006] In summary, the existing social security loss analysis technologies mainly have the following problems: over-reliance on correlation analysis, lack of causal inference ability, and difficulty in accurately tracing the causes of loss; traditional machine learning methods require a large amount of manual feature engineering and are difficult to process complex and heterogeneous social security data; relying on supervised learning, it is difficult to construct an effective model in the absence of labeled data; the application of existing causal inference methods in social security loss analysis is insufficient, lacking an efficient and automated causal relationship identification method. Therefore, there is an urgent need for a new technical means to trace the real causes of social security loss in a more scientific and automated way and quantify the causal effects of various factors, so as to provide a reliable basis for optimizing social security policies. Summary of the Invention
[0007] An object of the present invention is to propose a method for tracing the reasons for social security loss based on a causal inference model. The present invention uses a causal directionality scoring function, combined with social security payment behavior and time series constraints, to ensure the rationality and accuracy of causal inference. It can not only identify the factors directly affecting social security loss, but also reveal potential indirect influence paths, improving the credibility of causal inference.
[0008] A method for tracing the reasons for social security loss based on a causal inference model according to an embodiment of the present invention includes the following steps:
[0009] S1. Collect basic information of insured persons, payment records, social security receipt situations, and employer information from multiple heterogeneous data sources to form a preliminary social security data set;
[0010] S2. Preprocess the preliminary social security data set to form a preprocessed social security data set;
[0011] S3. Perform feature construction, variable screening, normalization, and dimensionality reduction processing on the preprocessed social security data set to generate a feature set for model input;
[0012] S4. Use self-supervised learning methods to perform contrast learning and autoencoder training on the feature set, automatically learn and extract potential causal features in social security data, and form self-supervised learning feature representations;
[0013] S5. Based on the self-supervised learning feature representations, use unsupervised causal inference algorithms to construct a causal relationship network, determine the causal relationships between social security loss influencing factors, quantify the effects of each causal relationship, calculate the causal effects of each influencing factor on social security loss prediction, and output the causal relationship network of social security loss influencing factors and the analysis results of their causal effects.
[0014] Optionally, the S1 includes the following steps:
[0015] S11. Collect basic information of insured persons, payment records, social security receipt situations, and employer information from multiple heterogeneous social security data sources, and define a preliminary social security data set;
[0016] S12. Perform social security data identification and classification on the preliminary social security data set, associate social security data according to the insured person ID, and construct a standardized social security data set P of insured persons data :
[0017]
[0018] where p j represents the social security data record of the j-th insured person, and t jis a timestamp, M is the total number of unique insured persons, represents the basic information of the insured person, represents the payment record of the insured person, represents the information of the employer associated with the insured person;
[0019] S13. Classify the social security data on the social security receipt situation and define the social security receipt data set C data :
[0020]
[0021] where c k is the k-th social security receipt data, L is the total number of social security receipt records, t k is a timestamp, p k represents the insured person receiving social security, b k is the type of social security received, including pension, unemployment insurance and medical reimbursement, s k represents the receipt status, including normal receipt, interrupted receipt and non-receipt;
[0022] S14. Based on the social security data set P of the insured person data , the social security receipt data set C data and the social security data of the employer, construct the preliminary social security data set D final :
[0023] D final = {(p j , c k , u m ) | p j ∈P data , c k ∈C data , u m ∈U data};
[0024] where U data represents the social security data set of all employers, and u m represents the employer information associated with the insured person p j ;
[0025] Optionally, the S2 includes the following steps:
[0026] S21. Clean the data of the preliminary social security data set D final to remove noise data and incorrect records, and generate a cleaned social security data set;
[0027] S22. Remove duplicate data records from the cleaned social security data set to form a deduplicated social security data set;
[0028] S23. Adopt a missing value filling strategy for the missing values in the deduplicated social security data set to generate a social security data set with filled missing values.
[0029] S24. Apply an outlier detection algorithm to the social security data set with filled missing values, set an outlier detection threshold θ, and remove abnormal data records to obtain a social security data set after outlier processing.
[0030] S25. Perform format standardization processing on the social security data set after outlier processing, including data type conversion and time format unification, to form a preprocessed social security data set D pre 。
[0031] Optionally, the said S3 includes the following steps:
[0032] S31. Perform feature construction on the preprocessed social security data set D pre Based on the time series features, payment behavior patterns, and social security receipt trends of the social security data, construct a social security feature set F raw :
[0033]
[0034] where N is the total number of data samples, f i represents the i-th social security feature data, represents the payment record feature, represents the employer-related features, represents the social security receipt status feature;
[0035] S32. Perform variable screening on the social security feature set F raw Adopt feature correlation analysis and information gain calculation to remove redundant or invalid features and generate a screened social security feature set F selected ;
[0036] S33. Perform normalization processing on the screened social security feature set F selected Standardize the continuous variables and convert them to the [0, 1] interval to obtain a normalized social security feature set F norm ;
[0037] S34. Perform dimensionality reduction processing on the normalized social security feature set F norm Select the first d principal component features to generate a feature set F final :
[0038]
[0039] where, represents the k-th normalized feature, represents the m-th feature after dimensionality reduction, where w n is the corresponding principal component weight, and d is the number of selected principal components.
[0040] Optionally, S4 includes the following steps:
[0041] S41. Calculate the risk weights for the feature set F final . Based on the indicators reflecting payment anomalies, changes in receiving status, and employer stability in the social security data, calculate the risk score r i of each data sample x i using the risk assessment function R(·), and determine the risk weight factor λ i = f(r i ), where r i represents the risk quantification score of the sample x i , and f(·) is the risk mapping function;
[0042] S42. Use the risk-weighted self-supervised contrastive learning method to generate its enhanced version for each data sample x i from the feature set F final , construct a risk-weighted positive sample pair set and a negative sample set N , and calculate the risk-weighted contrastive loss function r :
[0043]
[0044] where g(·) is the data augmentation function, E(·) is the social security causal feature encoder for extracting potential causal features, sim(·,·) represents the similarity function between vectors, τ is the temperature hyperparameter, |P r | represents the number of positive sample pairs, and λ x is the risk weight factor of the sample x;
[0045] S43. Construct a causal consistency autoencoder to reconstruct the feature set F final . Set the social security causal feature encoder E(·) and the decoder D(·), and introduce a causal consistency regularization term to construct the causal consistency loss function L causal :
[0046]
[0047] where B is the selected sample pair set, r i and r j are the risk scores of the samples x i and x j respectively, and γ ijFor the causal importance weight of the sample pair (x i , x j ), where δ is the preset margin parameter for causal consistency determination;
[0048] S44. Simultaneously perform reconstruction training on the causal consistency autoencoder, and calculate the reconstruction loss function L recon :
[0049]
[0050] where |F final | represents the number of samples;
[0051] S45. Jointly optimize the risk-weighted contrast loss the causal consistency loss L causal and the reconstruction loss L recon by weighted summation to form the jointly optimized overall self-supervised loss function L total :
[0052]
[0053] where α, β, and γ are weight hyperparameters;
[0054] S46. Automatically learn and extract the potential causal features in the social security data through the jointly optimized overall self-supervised loss function to form the self-supervised learning feature representation set F self :
[0055]
[0056] where ξ is the weight of the gradient regularization term, used to regulate the sensitivity of the risk score R(x i ) with respect to the change of the causal feature E(x i ), so that the causal feature can stably reflect the social security loss risk, R(x i ) represents the risk score of the sample x i , represents the gradient of the risk score R(x i ) with respect to the causal feature extracted by the encoder E(·), reflecting the sensitivity of the social security loss risk to the change of the feature representation, and |F final | is the total number of samples in the model input feature set F final .
[0057] Optionally, the S5 includes the following steps:
[0058] S51. Extract the candidate causal variables V of social security loss based on the self-supervised learning feature representation set F self to form the social security loss causal variable set V causal :
[0059] V causal = {v i | v i ∈ F self , ρ(v i , Y) > τ ρ}};
[0060] Among them, ρ(v i , Y) represents the Pearson correlation coefficient between the variable v i and the social security loss result Y, and τ ρ is the set correlation threshold;
[0061] S52. Combine the prior information of the social security business logic to construct the social security loss causal relationship network G causal , whose structure consists of the variable set V causal and the edge set E causal as follows:
[0062] G causal = (V causal , E causal );
[0063] Among them, the causal edge set E causal is constructed by the time constraint of the social security data and the social security payment behavior constraint, and the causal directionality scoring function S causal (v i , v j ) is used to calculate the causal relationship:
[0064]
[0065] Among them, P(v j | v i ) represents the probability that v i occurs when v j occurs, λ prior is the prior constraint weight, and E prior is the causal edge set predefined based on the social security business rules;
[0066] Finally, the optimized social security loss causal relationship network is obtained
[0067] S53. Perform causal effect calculation on the optimized social security loss causal relationship network to calculate the causal effect C i (v eff ) of the causal influence variable v i on the social security loss Y:
[0068]
[0069] Among them, do(·) represents the intervention operation, which represents the expected value of social security loss when the variable v i is under control;
[0070] S54. According to the calculation result of the causal effect, calculate the causal attribution score C score (v i ) to quantify the importance of each influencing factor in social security loss:
[0071] C score (v i ) = |C eff (v i )| × P(v i ) × ω(v i );
[0072] Among them, P(v i ) represents the prior probability of the variable v i , and ω(v i ) is the attribution weight based on social security business rules;
[0073] S55. Combine the causal attribution scores of the influencing factors of social security loss with the key causal paths in the causal network to construct a causal interpretable analysis matrix C matrix for visualizing the main influencing factors and causal paths of social security loss:
[0074]
[0075] Among them, C eff (v i → v j ) represents the indirect impact of the variable v i on the social security loss Y through v j , and finally output the causal relationship network of the influencing factors of social security loss and the analysis result of its causal effect.
[0076] The beneficial effects of the present invention are as follows:
[0077] (1) The present invention adopts a self-supervised contrast learning method. Through a risk-weighted contrast loss function, it can automatically learn and extract potential causal features in social security data. In an unsupervised environment, it can automatically capture key causal variables of social security loss in a data self-driven manner. By using the contrast learning method and the sample pair enhancement strategy, the model can adaptively learn important variables related to causality during the optimization process, thereby improving the accuracy of social security loss analysis.
[0078] (2) The present invention adopts a causal directionality scoring function, combined with social security payment behaviors and time series constraints, to ensure the rationality and accuracy of causal inference. It can not only identify the factors directly affecting social security loss, but also reveal potential indirect influence paths, thereby improving the credibility of causal inference.
[0079] (3) Through a causal effect calculation method, the present invention quantifies the causal effects of various influencing factors on social security loss, and uses an intervention operation to evaluate the causal attribution scores of key variables, enabling the precise determination of the role sizes of various variables in social security loss. In addition, combined with the business logic of social security, the causal effect calculation results are adjusted for attribution weights, making the evaluation results of the influencing factors more reasonable in terms of business. Description of the Drawings
[0080] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0081] Figure 1 It is a flowchart of a method for tracing the causes of social security loss based on a causal inference model proposed by the present invention. Detailed Embodiment
[0082] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0083] Refer to Figure 1 , a method for tracing the causes of social security loss based on a causal inference model, includes the following steps:
[0084] S1. Collect the basic information of insured persons, payment records, social security receipt situations, and employer information from multiple heterogeneous data sources to form a preliminary social security data set;
[0085] S2. Preprocess the preliminary social security data set to form a preprocessed social security data set;
[0086] S3. Perform feature construction, variable screening, normalization, and dimensionality reduction on the preprocessed social security data set to generate a feature set for model input;
[0087] S4. Use a self-supervised learning method to perform contrast learning and autoencoder training on the feature set, automatically learn and extract the potential causal features in the social security data, and form a self-supervised learning feature representation;
[0088] S5. Based on the self-supervised learning feature representation, an unsupervised causal inference algorithm is used to construct a causal relationship network, determine the causal relationships among the influencing factors of social security loss, quantify the effects of each causal relationship, calculate the causal effects of each influencing factor on the prediction of social security loss, and output the causal relationship network of the influencing factors of social security loss and the analysis results of its causal effects.
[0089] In this embodiment, S1 includes the following steps:
[0090] S11. Collect the basic information of insured persons, payment records, social security receipt situations, and employer information from multiple heterogeneous social security data sources, and define a preliminary social security data set;
[0091] S12. Perform social security data identification and classification on the preliminary social security data set, associate the social security data according to the insured person ID, and construct a standardized social security data set P of insured persons data :
[0092]
[0093] where p j represents the social security data record of the jth insured person, t j is the timestamp, M is the total number of unique insured persons, represents the basic information of the insured person, represents the payment record of the insured person, represents the employer information associated with the insured person;
[0094] S13. Classify the social security receipt situation social security data, and define a social security receipt social security data set C data :
[0095]
[0096] where c k is the kth social security receipt data, L is the total number of social security receipt records, t k is the timestamp, p k represents the insured person receiving social security, b k is the type of social security received, including pension, unemployment insurance, and medical reimbursement, s k represents the receipt status, including normal receipt, interrupted receipt, and non-receipt;
[0097] S14. Based on the social security data set P of insured persons data , the social security receipt social security data set C data and the employer social security data, construct a preliminary social security data set D final :
[0098] D final={(p j , c k , u m ) | p j ∈P data , c k ∈C data , u m ∈U data};
[0099] Among them, U data represents the social insurance data set of all employers, and u m represents the employer information associated with the insured person p j .
[0100] In this embodiment, S2 includes the following steps:
[0101] S21. Clean the preliminary social insurance data set D final , remove noise data and error records, and generate a cleaned social insurance data set;
[0102] S22. Remove duplicate data records from the cleaned social insurance data set to form a deduplicated social insurance data set;
[0103] S23. Adopt a missing value filling strategy for the missing values in the deduplicated social insurance data set to generate a social insurance data set with missing values filled;
[0104] S24. Apply an outlier detection algorithm to the social insurance data set with missing values filled, set an outlier detection threshold θ, and remove outlier data records to obtain a social insurance data set after outlier processing;
[0105] S25. Perform format standardization processing on the social insurance data set after outlier processing, including data type conversion and time format unification, to form a preprocessed social insurance data set D pre .
[0106] In this embodiment, S3 includes the following steps:
[0107] S31. Construct features for the preprocessed social insurance data set D pre , and construct a social insurance feature set F raw based on the time series features, payment behavior patterns, and social insurance receiving trends of the social insurance data:
[0108]
[0109] Among them, N is the total amount of data samples, f i represents the i-th social insurance feature data, represents the payment record feature, represents the employer-related feature, Represent the characteristics of the social security receiving status;
[0110] S32. For the social security feature set F raw Perform variable screening, adopt feature correlation analysis and information gain calculation, remove redundant or invalid features, and generate the screened social security feature set F selected ;
[0111] S33. For the screened social security feature set F selected Perform normalization processing, standardize continuous variables, and transform them into the [0, 1] interval to obtain the normalized social security feature set F norm ;
[0112] S34. For the normalized social security feature set F norm Perform dimensionality reduction processing, select the first d principal component features, and generate the feature set F final :
[0113]
[0114] Among them, represents the k-th feature after normalization, represents the m-th feature after dimensionality reduction, w n is the corresponding principal component weight, and d is the number of selected principal components.
[0115] In this embodiment, S4 includes the following steps:
[0116] S41. Calculate the risk weight for the feature set F final According to the indicators reflecting payment anomalies, changes in receiving status, and the stability of employers in the social security data, calculate the risk score r i of each data sample x i through the risk assessment function R(·), and determine the risk weight factor λ i = f(r i ), where r i represents the risk quantification score of the sample x i , and f(·) is the risk mapping function;
[0117] S42. Adopt the risk-weighted self-supervised contrast learning method to generate its enhanced version for each data sample x i from the feature set F final Construct a risk-weighted positive sample pair set and a negative sample set N , and calculate the risk-weighted contrast loss function r
[0118]
[0119] Among them, g(·) is a data augmentation function, E(·) is a social security causal feature encoder for extracting potential causal features, sim(·,·) represents a similarity function between vectors, τ is a temperature hyperparameter, |P r | represents the number of positive sample pairs, and λ x is the risk weight factor of sample x;
[0120] S43. Construct a causal consistency autoencoder to reconstruct the feature set F final , set the social security causal feature encoder E(·) and the decoder D(·), and introduce a causal consistency regularization term to construct a causal consistency loss function L causal :
[0121]
[0122] Among them, B is the selected set of sample pairs, r i and r j are the risk scores of samples x i and x j respectively, γ ij is the causal importance weight of the sample pair (x i , x j ), and δ is a preset margin parameter for causal consistency determination;
[0123] S44. Simultaneously perform reconstruction training on the causal consistency autoencoder, and calculate the reconstruction loss function L recon :
[0124]
[0125] Among them, |F final | represents the number of samples;
[0126] S45. Combine and optimize the risk-weighted contrast loss and the causal consistency loss L causal with the reconstruction loss L recon by weighted summation to form a jointly optimized overall self-supervised loss function L total :
[0127]
[0128] Among them, α, β, and γ are weight hyperparameters;
[0129] S46. Automatically learn and extract potential causal features in social security data through the jointly optimized overall self-supervised loss function to form a self-supervised learning feature representation set F self :
[0130]
[0131] Among them, ξ is the weight of the gradient regularization term, which is used to regulate the sensitivity of the risk score R(x i ) with respect to the change of the causal feature E(x i ), so that the causal feature can stably reflect the social security loss risk. R(x i ) represents the risk score of the sample x i . represents the gradient of the risk score R(x i ) with respect to the causal feature extracted by the encoder E(·), reflecting the sensitivity of the social security loss risk to the change of the feature representation. |F final | is the total number of samples in the model input feature set F final .
[0132] In this embodiment, S5 includes the following steps:
[0133] S51. Based on the self-supervised learning feature representation set F self , extract the candidate causal variables V of social security loss to form the social security loss causal variable set V causal :
[0134] V causal = {v i | v i ∈ F self , ρ(v i , Y) > τ ρ};
[0135] Among them, ρ(v i , Y) represents the Pearson correlation coefficient between the variable v i and the social security loss result Y, and τ ρ is the set correlation threshold;
[0136] S52. Combine the prior information of the social security business logic to construct the social security loss causal relationship network G causal , whose structure is composed of the variable set V causal and the edge set E causal , and is defined as follows:
[0137] G causal = (V causal , E causal );
[0138] Among them, the causal edge set E causal is constructed by the time constraint of the social security data and the social security payment behavior constraint, and the causal directionality scoring function S causal (v i , v j ) is used to calculate the causal relationship:
[0139]
[0140] Among them, P(v j | v i ) represents the probability of v i occurring under the condition that v j occurs. λ prior is the prior constraint weight, and E prior is a predefined causal edge set based on social security business rules;
[0141] Finally, the optimized causal relationship network of social security loss is obtained
[0142] S53. Calculate the causal effect on the optimized causal relationship network of social security loss to calculate the causal effect C i of the causal influence variable v eff on the social security loss Y: i ) :
[0143]
[0144] Among them, do(·) represents the intervention operation, represents the expected value of social security loss when the variable v i is controlled;
[0145] S54. According to the calculation results of the causal effect, calculate the causal attribution score C score (v i ) to quantify the importance of each influencing factor in social security loss:
[0146] C score (v i ) = |C eff (v i )| × P(v i ) × ω(v i ) ;
[0147] Among them, P(v i ) represents the prior probability of the variable v i , and ω(v i ) is the attribution weight based on social security business rules;
[0148] S55. Combine the causal attribution scores of the influencing factors of social security loss with the key causal paths in the causal network to construct a causal interpretable analysis matrix C matrix for visualizing the main influencing factors and causal paths of social security loss:
[0149]
[0150] Among them, Ceff (v i →v j ) represents the variable v i Through v j The indirect impact on the loss of social security Y is analyzed, and finally the causal relationship network of the influencing factors of social security loss and the analysis results of its causal effects are output.
[0151] Example 1:
[0152] From April 2023 to October 2023, the social security management center of a coastal province found in its daily monitoring that the social security loss rate of employees in the manufacturing industry in this region had been rising continuously. Especially in labor-intensive industries such as textiles and machining, the phenomenon of employees' social security payment interruption was prominent. The social security management department noticed that in some enterprises, the average payment cycle of employees had been significantly shortened, and some employees had not paid for several consecutive months. Historical data showed that such phenomena were often related to the business conditions of enterprises, industry fluctuations, and changes in labor contracts.
[0153] To deeply analyze the real reasons for social security loss, the social security management center of this province decided to use the causal inference method proposed in this invention to analyze the data of nearly 100,000 insured persons across the province, with the goal of identifying the core factors affecting social security loss and formulating corresponding policy intervention measures to reduce the social security loss rate.
[0154] On May 1, 2023, the social security management center retrieved the data of 100,000 insured persons in this region during the period from 2021 to 2023, covering the following information:
[0155] Basic information: name, age, gender, place of household registration; payment records: social security payment amount, payment frequency, and interruption times in the past 3 years; social security receipt situation: whether to receive unemployment benefits, pensions, and medical reimbursement records; employer information: enterprise scale, industry category, salary payment records, and layoff situation.
[0156] The system first preprocessed the data and eliminated obviously abnormal records, such as duplicate data and incorrect data. Subsequently, using the self-supervised learning method proposed in this invention, feature construction was performed on the cleaned data, and the following key variables were mainly extracted:
[0157] The payment stability of employees (such as whether there is a payment interruption of more than 3 consecutive months); the enterprise layoff rate in the past 12 months; the change in employees' salaries (whether there is a salary decrease of more than 10%); the individual characteristics of the age and working years of insured persons.
[0158] During the data analysis process, the system noticed that the social insurance loss rate of employees in a certain electronics processing factory (enterprise number E3291) was abnormally high. From May to July 2023, a total of 312 employees in this enterprise lost their social insurance, accounting for 42.3% of the total number of employees. This loss rate far exceeded the industry average level (the industry average loss rate was 18.5%). To further trace the reasons for the loss, the system conducted an in-depth analysis of the enterprise's social insurance payment pattern, salary change trend, and employee mobility situation.
[0159] Through the causal inference model, the system generated the causal network of social insurance loss for this enterprise and identified the following key causal relationships:
[0160] 1. A salary decrease of more than 12% → The probability of social insurance loss in the next 6 months increases by 36.7%;
[0161] 2. The enterprise's layoff rate exceeds 15% → The probability of social insurance loss within the next 3 months increases by 24.5%;
[0162] 3. Failure to pay social insurance for 3 consecutive months → The probability of complete loss in the next year is as high as 81.2%.
[0163] After further analyzing the social insurance data of enterprise E3291, the system found that:
[0164] In May 2023, the average salary of this enterprise decreased by 13.5% (from 4,200 yuan per month to 3,630 yuan)
[0165] During the same period, the enterprise's layoff rate reached 18.9%; 89 employees who lost their social insurance had a salary decrease in the 3 months before the loss; more than 70% of the lost employees experienced an interruption in social insurance payment in the past 6 months.
[0166] This data indicates that the deterioration of the enterprise's operating conditions led to a salary reduction, which in turn caused employees to lose their social insurance.
[0167] To verify the advantages of the method of the present invention, we compared the effects of the traditional correlation analysis method and the causal inference method of the present invention in analyzing social insurance loss.
[0168] The experimental results show that although the traditional method can identify the correlation between social insurance loss and factors such as salary and enterprise layoffs, it cannot determine the causal relationship between them, resulting in the lack of interpretability of the analysis results. In contrast, the method of the present invention can automatically construct a causal relationship network, accurately identify the root factors affecting social insurance loss, and make the analysis results more reliable.
[0169] In August 2023, based on the analysis results of the present invention, the social insurance management center implemented targeted policy intervention measures for high-risk enterprises and high-loss-risk employee groups:
[0170] 1. Provide social insurance premium reduction policies for enterprises with a layoff rate exceeding 10%, reducing the burden on enterprises;
[0171] 2. Provide employee stability subsidies for enterprises with a salary reduction exceeding 10%, encouraging enterprises to improve salary stability;
[0172] 3. Send social insurance warning notices to individuals who have not paid social insurance for 3 consecutive months and provide policies to encourage payment.
[0173] Three months after the implementation of the policy (from August 2023 to October 2023), the social insurance management center monitored the social insurance loss situation in this area again and evaluated the effect of policy intervention.
[0174]
[0175] After the intervention, the social insurance loss rate in the whole region decreased by 7.5%. Among them, the loss rate of enterprise E3291 decreased by 17.6%, and the loss rate of high-risk groups decreased by 13.2%, effectively verifying the accuracy and feasibility of the method of the present invention.
[0176] This embodiment demonstrates the practical application of the present invention in social insurance management, proving that the method can effectively track the real reasons for social insurance losses and provide scientific policy intervention suggestions. Compared with traditional methods, the method of the present invention not only improves the accuracy rate of social insurance loss analysis, but also can reveal potential causal relationships, making policy formulation more accurate and efficient. The application of this technology provides a new direction for future social insurance management, helps to reduce the social insurance loss rate, improve the social insurance coverage rate, and ensure the stable operation of the social security system.
[0177] The present invention adopts a self-supervised contrast learning method. Through a risk-weighted contrast loss function, it can automatically learn and extract potential causal features in social insurance data, be able to automatically capture key causal variables of social insurance losses in an unsupervised environment through a data self-driven method, and use the contrast learning method. Through the sample pair enhancement strategy, the model adaptively learns important variables related to causality during the optimization process, thereby improving the accuracy of social insurance loss analysis.
[0178] The present invention adopts a causal directionality scoring function, combined with social insurance payment behavior and time series constraints, to ensure the rationality and accuracy of causal inference. It can not only identify factors directly affecting social insurance losses, but also reveal potential indirect influence paths, improving the credibility of causal inference.
[0179] Through the causal effect calculation method, the present invention quantifies the causal effects of various influencing factors on social security loss, and uses intervention operations to evaluate the causal attribution scores of key variables, enabling the precise determination of the role of each variable in social security loss. In addition, in combination with the business logic of social security, the causal effect calculation results are adjusted for attribution weights, making the evaluation results of the influencing factors more reasonable in terms of business.
[0180] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.
Claims
1. A method for tracing the causes of social security loss based on a causal inference model, characterized in that: The steps include: S1. Collect basic information of insured persons, payment records, social security receipt status and employer information from multiple heterogeneous data sources to form a preliminary social security data set; S2. preprocessing the preliminary social security data set to form a preprocessed social security data set; S3, performing feature construction, variable screening, normalization and dimensionality reduction processing on the pre-processed social security data set to generate a feature set for model input; S4. Use self-supervised learning methods to conduct comparative learning and autoencoder training on feature sets, automatically learn and extract potential causal features in social security data, and form self-supervised learning feature representation; S5. Based on the self-supervised learning feature representation, an unsupervised causal inference algorithm is used to construct a causal network, determine the causal relationship between the factors affecting social security loss, quantify the effect of each causal relationship, calculate the causal effect of each influencing factor on the prediction of social security loss, and output the causal network of factors affecting social security loss and its causal effect analysis results.
2. According to claim 1, a method for tracing the cause of social security loss based on a causal inference model is characterized in that: The S1 comprises the following steps: S11. Collect basic information of insured persons, payment records, social security receipt status and employer information from multiple heterogeneous social security data sources to define a preliminary social security data set; S12. Identify and classify the preliminary social security data set, associate the social security data according to the insured person ID, and build a standardized social security data set P for insured persons. data : Among them, p j represents the social security data record of the jth insured person, t j is the timestamp, M is the total number of unique insured persons, Indicates the basic information of the insured person. Indicates the payment record of the insured person. Indicates the employer information associated with the insured person; S13. Classify the social security data on social security receipt and define the social security receipt data set C data : Among them, c k is the kth social security collection data, L is the total number of social security collection records, t k is the timestamp, p k Indicates the insured persons who receive social insurance, b k The type of social security received, including pensions, unemployment insurance and medical reimbursement, k Indicates the collection status, including normal collection, interrupted collection and uncollected; S14. Based on the social security data set P of insured persons data 、Social security collection social security data set C data and employer social security data to build a preliminary social security data set D final : D final ={(p j ,c k ,u m )∣p j ∈P data ,c k ∈C data ,u m ∈U data }; Among them, U data Represents the social security data set of all employers, u m Representatives and insured persons j Related employer information.
3. According to claim 1, a method for tracing the cause of social security loss based on a causal inference model is characterized in that: The S2 comprises the following steps: S21. Preliminary social security data set D final Perform data cleaning, remove noise data and erroneous records, and generate a cleaned social security data set; S22, performing deduplication processing on the cleaned social security data set, removing duplicate data records, and forming a deduplicated social security data set; S23, using a missing value filling strategy for the missing values in the social security data set after deduplication, to generate a social security data set after missing value filling; S24, applying an outlier detection algorithm to the social security data set after missing value filling, setting an outlier detection threshold θ, removing abnormal data records, and obtaining a social security data set after outlier processing; S25. Perform format standardization on the social security data set after outlier processing, including data type conversion and time format unification, to form a pre-processed social security data set D pre .
4. According to claim 1, a method for tracing the cause of social security loss based on a causal inference model is characterized in that: The S3 comprises the following steps: S31. Preprocessing social security data set D pre Perform feature construction, and construct the social security feature set F based on the time series characteristics of social security data, payment behavior patterns, and social security collection trends. raw : Among them, N is the total number of data samples, f i represents the social security characteristic data of the ith item, Indicates the payment record characteristics. Represents the relevant characteristics of the employer, Represents the characteristics of social security receipt status; S32, for the social security feature set F raw Perform variable screening, use feature correlation analysis and information gain calculation to remove redundant or invalid features, and generate the screened social security feature set F selected ; S33, the selected social security feature set F selected Perform normalization processing, standardize continuous variables, convert them to the [0,1] interval, and obtain the normalized social security feature set F norm ; S34, the normalized social security feature set F norm Perform dimensionality reduction, select the first d principal component features, and generate the feature set F final : in, represents the normalized kth feature, represents the mth feature after dimensionality reduction, w n is the corresponding principal component weight, and d is the number of principal components selected.
5. According to claim 1, a method for tracing the cause of social security loss based on a causal inference model is characterized in that: The S4 comprises the following steps: S41, feature set F final Calculate the risk weight, and use the risk assessment function R(·) to calculate the risk weight of each data sample x based on the social security data reflecting payment anomalies, changes in payment status, and employer stability indicators. i The risk score of i , and determine the risk weight factor λ based on the risk score i =f(r i ), where r i Represents sample x i The risk quantification score is f(·), and f(·) is the risk mapping function; S42, use risk-weighted self-supervised contrastive learning method for each data sample x i From the feature set F final Generate an enhanced version Construct a risk-weighted positive sample pair set And the negative sample set N r , and calculate the risk-weighted contrast loss function Among them, g(·) is the data enhancement function, E(·) is the social security causal feature encoder, which is used to extract potential causal features, sim(·,·) represents the similarity function between vectors, τ is the temperature hyperparameter, |P r | represents the number of positive sample pairs, λ x is the risk weight factor of sample x; S43, construct a causal consistency autoencoder, and final Reconstruct, set the social security causal feature encoder E(·) and decoder D(·), introduce the causal consistency regularization term, and construct the causal consistency loss function L causal : Among them, B is the selected sample pair set, r i With r j The samples x are i With x j The risk score of ij For the sample pair (x i ,x j ) is the causal importance weight, δ is the preset margin parameter for causal consistency judgment; S44. Perform reconstruction training on the causal consistency autoencoder at the same time and calculate the reconstruction loss function L recon : Among them, |F final | indicates the number of samples; S45. Compare risk weights to losses Causal consistency loss L causal and the reconstruction loss L recon The weighted summation method is used for joint optimization to form the joint optimization overall self-supervisory loss function L total : Among them, α, β and γ are weight hyperparameters; S46. By jointly optimizing the overall self-supervised loss function, we automatically learn and extract the potential causal features in the social security data to form a self-supervised learning feature representation set F self : Among them, ξ is the weight of the gradient regularization term, which is used to regulate the risk score R(x i ) about the causal feature E(x) i ) changes, so that the causal characteristics can stably reflect the risk of social security loss, R(x i ) represents the sample x i The risk score, represents the risk score R(x i ) about the gradient of the causal features extracted by the encoder E(·), reflecting the sensitivity of social security loss risk to changes in feature representation, |F final | Input feature set F to the model final The total number of samples in .
6. The method for tracing the cause of social security loss based on a causal inference model according to claim 1 is characterized in that: The S5 comprises the following steps: S51. Feature representation set F based on self-supervised learning self Extract candidate causal variables V of social security loss and form a set of causal variables V of social security loss causal : V causal ={v i ∣v i ∈F self ,ρ(v i ,Y)>τ ρ }; Among them, ρ(v i ,Y) represents the variable v i The Pearson correlation coefficient between the social security loss result Y, τ ρ is the set relevance threshold; S52. Combine the prior information of social security business logic to construct the social security loss causal relationship network G causal , whose structure consists of a variable set V causal and edge set E causal Composition, defined as follows: G causal =(V causal ,E causal ); Among them, the causal edge set E causal The causal directional scoring function S is used to construct the time constraint of social security data and the social security payment behavior constraint. causal (v i ,v j )Calculate causal relationships: Among them, P(v j ∣v i ) indicates that in v i In case of occurrence j The probability of occurrence, λ prior is the prior constraint weight, E prior It is a causal edge set predefined based on social security business rules; Finally, the social security loss causal relationship network is optimized S53. Optimized causal relationship network of social security loss Perform causal effect calculation and calculate the causal influence variable v i The causal effect C on the loss of social security Y eff (v i ): Among them, do(·) represents the intervention operation, Represents the variable v i the expected value of social security losses when controlled; S54. Calculate the causal attribution score C based on the causal effect calculation results. score (v i ) to quantify the importance of each factor in social security loss: C score (in i )=|C eff (in i )|×P(v i )×ω(v i ); Among them, P(v i ) represents the variable v i The prior probability, ω(v i ) is the attribution weight based on social security business rules; S55. Combine the causal attribution scores of social security loss influencing factors with the key causal paths in the causal network to construct a causal interpretable analysis matrix C matrix , used to visualize the main influencing factors and causal paths of social security loss: Among them, C eff (v i →v j ) represents the variable v i By v j The indirect impact on social security loss Y is finally output as the causal relationship network of factors affecting social security loss and its causal effect analysis results.