Rural finance-oriented LSTM-auto-encoder hybrid risk early warning method
Through the LSTM-autocoder hybrid risk warning method, combined with agricultural production cycle and rural policy cycle information, a long and short-term memory network under a multi-window time scale was built, and asynchronous memory and timing reconstruction were carried out, which solved the lag and misjudgment problems of rural financial risk warning, and achieved more efficient risk identification and response.
Patent Information
- Application Number
- CN202510781675.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-02
AI Technical Summary
The existing rural financial risk warning methods cannot effectively deal with the complex financial risk forms driven by agricultural production laws and policy changes, and there are problems of lagging warnings or frequent misjudgment.
The LSTM-autocoder hybrid risk warning method is adopted to obtain the historical time series of rural financial risk events, perform migration feature extraction of feature spatial distribution, build a long and short-term memory network subnet under the multi-window time scale, perform asynchronous memory state encoding, and use the autoencoder for timing reconstruction, and combine the error-sensitive threshold generation strategy of causal traceability to determine the critical threshold and category of risk events.
It improves the accuracy and interpretation of rural financial risk warnings, enhances the ability to learn complex risk models, improves the accuracy and adaptability of risk warnings, and can form a traceable and interfereable risk response path.
Smart Images

Figure CN120579822A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of risk prediction technology, and more specifically, to an LSTM-autoencoder hybrid risk early warning method for rural finance. Background Art
[0002] Existing early warning methods for rural financial risks generally suffer from insufficient response to cyclical factors, making them ineffective in addressing the complex financial risk landscape driven by agricultural production patterns and policy changes. In rural financial scenarios, risk events exhibit temporal dependencies and regional variations. Traditional risk identification methods based on static indicators or single models cannot accurately capture the dynamic drift of risk across different cycles and spaces, leading to delayed early warnings and frequent misjudgments. Summary of the Invention
[0003] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides an LSTM-autoencoder hybrid risk warning method for rural finance to solve the problems raised in the above-mentioned background technology.
[0004] To achieve the above object, the present invention provides the following technical solutions:
[0005] An LSTM-autoencoder hybrid risk early warning method for rural finance includes the following steps:
[0006] S1: Based on the information of agricultural production cycle and rural policy cycle, the historical time series of rural financial risk events is obtained, and the migration features of the feature space distribution are extracted to form the agricultural policy drift feature set;
[0007] S2: Perform multi-window time-scale layered processing on the agricultural policy drift feature set, build several independent LSTM sub-networks, perform asynchronous memory state encoding on the multi-scale features, and generate a multi-scale latent space encoding set;
[0008] S3: Weighted aggregation of the multi-scale latent space encoding set is input into the autoencoder to reconstruct the time series of risk events and obtain the reconstruction error sequence;
[0009] S4: Using real events of rural financial historical risks, we determine the critical threshold of reconstruction error through an error-sensitive threshold generation strategy based on causal tracing;
[0010] S5: The reconstructed error sequence is judged through the critical threshold, the risk event warning result is output, and the risk category and corresponding influencing factors are determined according to the risk assessment indicator system.
[0011] In a preferred embodiment, S1 is specifically:
[0012] Determine the cycle synchronization time window based on agricultural production cycle information and rural policy cycle information;
[0013] Extract agricultural production cycle characteristics and rural policy cycle characteristics from the historical time series of rural financial risk events based on the cycle synchronization time window;
[0014] The migration characteristics of agricultural production cycle characteristics and rural policy cycle characteristics are evaluated in spatial domain respectively, and the drift characteristic set of agricultural production cycle and rural policy cycle drift characteristic set are obtained;
[0015] The agricultural production cycle drift feature set and the rural policy cycle drift feature set are fused in feature space to form the agricultural policy drift feature set.
[0016] In a preferred embodiment, the cycle synchronization time window is determined based on the agricultural production cycle information and the rural policy cycle information, specifically:
[0017] Divide the time series nodes corresponding to the agricultural production cycle according to the agricultural production cycle information to form an agricultural production cycle node set;
[0018] According to the rural policy cycle information, the time series nodes corresponding to the rural policy cycle are divided to form a rural policy cycle node set;
[0019] Conduct time series cross-comparison analysis on the agricultural production cycle node set and the rural policy cycle node set, and calculate the time synchronization index;
[0020] According to the time synchronization index, several time series nodes with the highest synchronization between the agricultural production cycle node set and the rural policy cycle node set are selected to form the cycle synchronization time window.
[0021] In a preferred embodiment, S2 is specifically:
[0022] Based on the agricultural policy drift feature set, it is divided into the first time scale feature window, the second time scale feature window and the third time scale feature window;
[0023] Based on the first time scale feature window, the second time scale feature window and the third time scale feature window, independent long short-term memory network sub-networks are established respectively;
[0024] The long short-term memory network sub-network is used to asynchronously encode the memory states of the feature sequences of the first time scale feature window, the second time scale feature window, and the third time scale feature window respectively;
[0025] Output the first scale latent space coding set, the second scale latent space coding set and the third scale latent space coding set respectively;
[0026] The time scale feature fusion processing is performed on the first scale latent space code set, the second scale latent space code set and the third scale latent space code set to obtain a multi-scale latent space code set.
[0027] In a preferred embodiment, S3 is specifically:
[0028] Perform feature weight analysis on the multi-scale latent space coding set to obtain the first scale feature weight parameter, the second scale feature weight parameter and the third scale feature weight parameter;
[0029] The multi-scale latent space coding set is weightedly aggregated using the first scale feature weight parameter, the second scale feature weight parameter, and the third scale feature weight parameter to form a weighted fusion feature coding;
[0030] Using the weighted fusion feature code as the input of the autoencoder, a time series reconstruction network model based on the combined structure of long short-term memory network and autoencoder is established;
[0031] The time series reconstruction network model is used to reconstruct the feature coding after weighted fusion to obtain the time series reconstruction output result;
[0032] The data difference between the time series reconstruction output result and the feature coding after weighted fusion is calculated at each time point to obtain the reconstruction error sequence.
[0033] In a preferred embodiment, S4 is specifically:
[0034] According to the real risk events marked by the historical time series of rural financial risk events, the causal correlation analysis is performed on the reconstruction error sequence to obtain the causal correlation feature set of the reconstruction error;
[0035] Based on the causal correlation feature set of the reconstructed error, the error sensitivity weight parameter is constructed to determine the sensitivity of different risk events to error changes;
[0036] Dynamically weight the reconstructed error sequence according to the error sensitivity weight parameter and establish an error dynamic threshold adjustment strategy;
[0037] According to the error dynamic threshold adjustment strategy, the error critical threshold applicable to the current agricultural production cycle information and rural policy cycle information is determined.
[0038] In a preferred embodiment, S5 is specifically:
[0039] The error critical threshold is used to perform risk discrimination processing on the reconstructed error sequence at each time point to determine whether a risk event occurs at each time point;
[0040] Establish a risk assessment indicator system; the indicator system includes economic factor indicators, social environment factor indicators, agricultural production cycle indicators and rural policy cycle indicators;
[0041] According to the risk event time points obtained by risk identification, the risk event data at the corresponding time points are extracted;
[0042] Correlate the risk event data with the risk assessment indicator system to determine the risk category of each risk event;
[0043] Map risk categories and risk assessment indicator systems to determine the influencing factors corresponding to each risk category.
[0044] The technical effects and advantages of the LSTM-autoencoder hybrid risk warning method for rural finance in the present invention are as follows:
[0045] By integrating information on the agricultural production cycle and the rural policy cycle, we can fully extract the structural drift characteristics in the evolution of rural financial risks and enhance the model's sensitivity to cyclical changes; by constructing a long-short-term memory network sub-network under multiple window time scales, we can achieve asynchronous memory and encoding of features at different time scales, effectively capturing short-term fluctuations and long-term trends; the autoencoder reconstructs the multi-scale features in time series, which can enhance the model's learning ability for complex risk patterns and enhance its ability to characterize nonlinear risks; by introducing an error-sensitive threshold generation mechanism based on causal tracing, the accuracy of risk warnings and the adaptability of threshold judgments are improved; combining the risk assessment indicator system to realize risk category identification and influencing factor analysis, it helps to form a traceable and interventionable risk response path; effectively improving the accuracy, interpretability and practical applicability of rural financial risk warnings. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a schematic diagram of an LSTM-autoencoder hybrid risk warning method for rural finance in the present invention. DETAILED DESCRIPTION
[0047] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0048] Example
[0049] Figure 1 The present invention provides an LSTM-autoencoder hybrid risk warning method for rural finance, which includes the following steps:
[0050] S1: Based on the information of agricultural production cycle and rural policy cycle, the historical time series of rural financial risk events is obtained, and the migration features of the feature space distribution are extracted to form the agricultural policy drift feature set;
[0051] S2: Perform multi-window time-scale layered processing on the agricultural policy drift feature set, build several independent LSTM sub-networks, perform asynchronous memory state encoding on the multi-scale features, and generate a multi-scale latent space encoding set;
[0052] S3: Weighted aggregation of the multi-scale latent space encoding set is input into the autoencoder to reconstruct the time series of risk events and obtain the reconstruction error sequence;
[0053] S4: Using real events of rural financial historical risks, we determine the critical threshold of reconstruction error through an error-sensitive threshold generation strategy based on causal tracing;
[0054] S5: The reconstructed error sequence is judged through the critical threshold, the risk event warning result is output, and the risk category and corresponding influencing factors are determined according to the risk assessment indicator system.
[0055] S1: Based on the information of agricultural production cycle and rural policy cycle, we obtain the historical time series of rural financial risk events, extract the migration features of feature space distribution, and form the agricultural policy drift feature set, including:
[0056] Determine the cycle synchronization time window based on agricultural production cycle information and rural policy cycle information;
[0057] The method for determining the cycle synchronization time window is to obtain data on agricultural production cycles and rural policy cycles. Specifically, agricultural production cycle data includes time nodes such as the planting start date, crop growth period, and harvest period of local major crops; rural policy cycle data includes specific policy time nodes such as the implementation date of agricultural subsidy policies, policy adjustment dates, loan policy effective date, and fund disbursement cycle.
[0058] Based on the acquired agricultural production cycle information and the planting patterns of specific crops, such as wheat or corn, the annual agricultural production cycle is divided into distinct time series nodes, forming a set of agricultural production cycle nodes. For example, taking wheat planting as an example, from the sowing period in early October each year to the harvest period in late June of the following year, agricultural production time nodes such as sowing nodes, growing season nodes, irrigation nodes, fertilization nodes, and harvest nodes are divided into a set of agricultural production cycle nodes.
[0059] Based on the acquired rural policy cycle information and specific policies, such as the agricultural subsidy disbursement cycle or the adjustment date of rural credit loan policies, the annual policy implementation cycle is divided into time series nodes to form a rural policy cycle node set. For example, the time nodes for agricultural fund disbursement, the start and end nodes for loan policy implementation, and the subsidy policy adjustment nodes are divided separately to form a rural policy cycle node set.
[0060] A cross-comparison analysis is conducted between the determined set of agricultural production cycle nodes and the set of rural policy cycle nodes. The specific method is to analyze each agricultural production cycle node against each node in the set of rural policy cycle nodes, calculate the temporal matching degree of each corresponding set of nodes, and calculate the synchronization index value by determining whether there is temporal overlap or proximity between them. The calculation formula is: The synchronization index value is equal to the sum of the temporal overlap or proximity between agricultural production cycle nodes and rural policy cycle nodes divided by the total number of matching nodes.
[0061] Based on the calculated synchronization index, several combinations of agricultural production cycle nodes and rural policy cycle nodes with the highest index values are selected from all node combinations. The time nodes included in these combinations together constitute the cycle synchronization time window. Specifically, a high index value indicates that the agricultural production cycle and the rural policy cycle have significant synchronization characteristics at the corresponding time nodes, which is suitable for risk early warning analysis. For example, when the sowing period and the loan disbursement date are highly overlapped, this time node can become an important part of the cycle synchronization time window.
[0062] Extract agricultural production cycle characteristics and rural policy cycle characteristics from the historical time series of rural financial risk events based on the cycle synchronization time window;
[0063] Based on the determined cycle synchronization time window, agricultural production cycle characteristics and rural policy cycle characteristics are extracted from the historical time series of rural financial risk events. Specifically, agricultural production cycle characteristics are extracted by taking as features indicators such as the number of loan applications, number of overdue repayments, and loan purposes corresponding to agricultural production activities within the cycle synchronization time window; rural policy cycle characteristics are extracted and aggregated by taking as features indicators such as loan policy interest rates, subsidy fund size, and policy implementation area within the cycle synchronization time window. This ensures that the extracted features can truly reflect the impact of agricultural production and policy implementation on rural financial risk events.
[0064] The migration characteristics of agricultural production cycle characteristics and rural policy cycle characteristics are evaluated in spatial domain respectively, and the drift characteristic set of agricultural production cycle and rural policy cycle drift characteristic set are obtained;
[0065] Migration feature analysis involves mapping the agricultural production cycle characteristics and rural policy cycle characteristics obtained in different regions and over different time periods into a unified spatial feature domain. The mean and standard deviation of the numerical distribution of these characteristics within the unified spatial feature domain are calculated, and the standardized mean difference between each characteristic across different regions is used as a difference metric. A characteristic is considered a drift feature when the standardized mean difference is greater than the median of all characteristics plus two standard deviations. For example, using the delinquency rate of wheat loan applications in different regions as an example, the degree of variation in the delinquency rate of wheat loans across regions within the unified spatial domain is calculated, and the range of the drift feature is determined based on the data from regions with the largest degree of variation.
[0066] The agricultural production cycle drift feature set and the rural policy cycle drift feature set are fused in feature space to form the agricultural policy drift feature set.
[0067] The feature space fusion processing is implemented, specifically: the feature dimensions of the agricultural production cycle drift feature set and the rural policy cycle drift feature set are uniformly adjusted through feature dimension alignment, so that the two types of drift feature sets have consistent feature dimensions; the feature weighted fusion method is used to fuse the various features of the agricultural production cycle drift feature set and the rural policy cycle drift feature set according to the weight parameters; the fusion calculation method is: each feature is multiplied by the corresponding weight parameter and then added one by one to obtain the fused agricultural policy drift feature set. The fused agricultural policy drift feature set uniformly expresses the joint effect of the agricultural production cycle characteristics and the rural policy cycle characteristics on risk drift.
[0068] S2: Perform multi-window time-scale layered processing on the agricultural policy drift feature set, build several independent LSTM sub-networks, perform asynchronous memory state encoding on the multi-scale features, and generate a multi-scale latent space encoding set, including:
[0069] Based on the agricultural policy drift feature set, it is divided into the first time scale feature window, the second time scale feature window and the third time scale feature window;
[0070] The characteristic window division method is as follows: based on the change rate of different risk characteristics, the agricultural policy drift feature set is divided into three different time-scale characteristic windows: the first time-scale characteristic window, the second time-scale characteristic window, and the third time-scale characteristic window. The division method is as follows: the first time-scale characteristic window covers characteristics that change rapidly and have a short duration, such as the frequency of short-term loan delinquencies or the short-term loan fluctuations caused by sudden adjustments to agricultural subsidy funds; the second time-scale characteristic window covers characteristics with medium-term changes, such as seasonal agricultural production loan applications and changes in policy loan quotas; and the third time-scale characteristic window covers characteristics with long durations and slow changes, such as the long-term trend characteristics of rural credit rating levels and regional economic development indices.
[0071] Based on the first time scale feature window, the second time scale feature window and the third time scale feature window, independent long short-term memory network sub-networks are established respectively;
[0072] Based on the first, second, and third time-scale feature windows, corresponding long-short-term memory (LSTM) sub-network models are constructed. The LSTM has a memory unit structure specifically designed for time series data, effectively capturing long-term dependencies and short-term abnormal fluctuations. The model structure parameters are configured as follows: for the LSTM sub-network established for the first time-scale feature window, a structure with a smaller number of hidden layer nodes is adopted to ensure a rapid response to short-term feature changes; for the LSTM sub-network established for the second time-scale feature window, the number of hidden layer nodes is appropriately increased; and for the LSTM sub-network established for the third time-scale feature window, the number of hidden layer nodes and the number of model layers are further increased to enhance the in-depth capture of long-term trend changes.
[0073] The long short-term memory network sub-network is used to asynchronously encode the memory states of the feature sequences of the first time scale feature window, the second time scale feature window, and the third time scale feature window respectively;
[0074] The asynchronous encoding process is implemented as follows: Agricultural policy drift features in the first, second, and third time-scale feature windows are used as input feature data, and forward computation is performed using independently constructed long-short-term memory (LSTM) subnetworks. Each subnetwork is equipped with three gating mechanisms: a forget gate, an input gate, and an output gate. The forget gate controls the degree of retention of the memory state at the previous moment, the input gate controls the degree of fusion of feature information from the current input, and the output gate controls the degree of output of the currently encoded features. Through this gating mechanism, each LSTM subnetwork obtains the latent space encoding output results for its corresponding time-scale features, forming the first-scale latent space encoding set, the second-scale latent space encoding set, and the third-scale latent space encoding set.
[0075] For example, the first-scale latent space coding set is formed as follows: after the data of the first time-scale feature window is input into the long-short-term memory network sub-network, it is first controlled by the forget gate to delete irrelevant historical memories, and then controlled by the input gate to integrate the new short-term feature information into the existing memory state. Finally, the output gate controls the encoding output, thereby forming the first-scale latent space coding set for short-term feature information; similarly, the second-scale latent space coding set and the third-scale latent space coding set are also generated after being processed by the corresponding long-short-term memory network sub-networks.
[0076] Output the first scale latent space coding set, the second scale latent space coding set and the third scale latent space coding set respectively;
[0077] Performing time scale feature fusion processing on the first scale latent space code set, the second scale latent space code set and the third scale latent space code set to obtain a multi-scale latent space code set;
[0078] The feature fusion processing method is as follows: the first-scale latent space code set, the second-scale latent space code set, and the third-scale latent space code set are adjusted to a unified dimension to ensure the consistency of the dimensions of the three latent space code sets; the three latent space code sets of different scales are fused using the time-scale feature fusion method. The fusion method is as follows: the variance value of the features in each latent space code set is calculated to determine the degree of feature variation; the ratio of the feature variance value to the total variance value of all latent space features is then used as the feature fusion weight coefficient of the scale latent space code set; finally, each feature value of the first-scale latent space code set, the second-scale latent space code set, and the third-scale latent space code set are multiplied by the corresponding feature fusion weight coefficient and added together to obtain the unified fusion multi-scale latent space code set.
[0079] S3: Weighted aggregation of the multi-scale latent space encoding set is input into the autoencoder to reconstruct the time series of risk events and obtain the reconstruction error sequence, including:
[0080] Perform feature weight analysis on the multi-scale latent space coding set to obtain the first scale feature weight parameter, the second scale feature weight parameter and the third scale feature weight parameter;
[0081] The feature weight analysis process involves separately analyzing the first, second, and third scales of the multi-scale latent space code set. Specifically, the contribution value of each feature in each scale latent space code set is calculated. This contribution value is calculated by comparing the degree to which each feature improves risk prediction accuracy. Specifically, the increase in prediction error between the risk prediction result and the actual historical risk event is calculated after removing each feature from the multi-scale latent space code set. A greater increase in prediction error indicates a higher feature contribution value, thereby determining the feature's importance within the corresponding scale code set.
[0082] Based on the contribution values of each feature obtained from the analysis, the first-scale feature weight parameter, the second-scale feature weight parameter, and the third-scale feature weight parameter are determined. The feature weight coefficients are determined by calculating the sum of all feature contribution values in the latent space encoding set at each scale; dividing each feature contribution value by the sum of all feature contribution values in the encoding set at its corresponding scale to obtain the relative importance of the feature in the encoding set at that scale; and taking the average of the relative importance of all features in the latent space encoding set at each scale as the feature weight coefficient for the corresponding scale. For example, when the average relative importance of all features in the latent space encoding set at the first scale is higher, a higher first-scale feature weight parameter is obtained. Similarly, the second-scale and third-scale feature weight parameters are obtained.
[0083] The multi-scale latent space coding set is weightedly aggregated using the first scale feature weight parameter, the second scale feature weight parameter, and the third scale feature weight parameter to form a weighted fusion feature coding;
[0084] Each feature in the first-scale latent space coding set, the second-scale latent space coding set, and the third-scale latent space coding set is multiplied by the corresponding scale feature weight coefficient; the weighted first-scale latent space coding set, the second-scale latent space coding set, and the third-scale latent space coding set are added together at the corresponding positions of each feature to obtain a unified weighted fusion feature coding; the unified weighted fusion feature coding includes the comprehensive impact of the first-scale, second-scale, and third-scale features on risk prediction, ensuring higher accuracy of reconstruction analysis.
[0085] Using the weighted fusion feature code as the input of the autoencoder, a time series reconstruction network model based on the combined structure of long short-term memory network and autoencoder is established;
[0086] The temporal reconstruction network model consists of a long short-term memory network (LSTM) as the encoder portion of the autoencoder and a standard multi-layer feedforward neural network as the decoder portion. The parameters are configured as follows: the LSTM network in the encoder portion has multiple hidden layers and is equipped with forget gates, input gates, and output gates to effectively capture the temporal dependencies in the weighted fusion feature encoding. The feedforward neural network in the decoder portion gradually reconstructs the encoded features through multiple hidden layers, each using a linear activation function to achieve more stable reconstruction output.
[0087] The time series reconstruction network model is used to reconstruct the feature coding after weighted fusion to obtain the time series reconstruction output result;
[0088] The feature coding after weighted fusion is first input into the encoder part of the long short-term memory network. After processing the memory units at multiple moments, the hidden feature representation after time series coding is obtained. The hidden feature representation is input into the decoder part of the autoencoder, and linear reconstruction is performed layer by layer through a multi-layer feedforward neural network to obtain the time series reconstruction output result. The output result is the output feature sequence after the complete reconstruction process of the input feature coding implemented by the time series reconstruction network model.
[0089] The data difference between the time series reconstruction output result and the feature code after weighted fusion is calculated at each time point to obtain the reconstruction error sequence;
[0090] The calculation method is as follows: at the same time point, the numerical difference between the time series reconstruction output and the corresponding weighted fusion feature code is calculated feature by feature. The absolute value of the difference across all features at each time point is taken and summed to obtain the reconstruction error value at each time point. The reconstruction error values of all time points constitute a complete reconstruction error sequence. The resulting reconstruction error sequence serves as an important basis for risk event determination and reflects the accuracy of risk event prediction.
[0091] S4: Using real historical rural financial risk events, we determine the critical threshold of reconstruction error through an error-sensitive threshold generation strategy based on causal tracing, including:
[0092] According to the real risk events marked by the historical time series of rural financial risk events, the causal correlation analysis is performed on the reconstruction error sequence to obtain the causal correlation feature set of the reconstruction error;
[0093] Based on the real risk events clearly marked in the historical time series of rural financial risk events, the causal correlation analysis of the reconstructed error sequence is processed to determine the causal relationship between the reconstructed error sequence and the real risk events. Specifically, the real risk events marked in the historical time series of rural financial risk events are matched and analyzed one by one with the corresponding time nodes of the reconstructed error sequence according to the time nodes; the causal relationship between the occurrence time of each risk event and the corresponding reconstruction error change is calculated and analyzed one by one through the Granger causality test method in the causal inference method; the calculation method is to divide the reconstructed error sequence into two sequences of thirty consecutive time points before the occurrence of the real risk event and thirty time points after the occurrence; the two reconstructed error sequences are linearly fitted by the least squares method, and the trend line slope of each sequence is calculated; the two-sample t-test method in the statistical hypothesis test is used to calculate whether the difference in trend slope before and after the event is statistically significant; the significance level is set to five percent. If the probability value corresponding to the calculated test statistic is less than the significance level threshold, it is determined that there is a statistically significant difference between the risk event and the reconstructed error sequence, thereby confirming the existence of a causal relationship; the features that are tested to have significant trend differences are selected as the causal relationship feature set of the reconstruction error.
[0094] Based on the causal correlation feature set of the reconstructed error, the error sensitivity weight parameter is constructed to determine the sensitivity of different risk events to error changes;
[0095] Based on the causal correlation feature set of the reconstructed error, an error sensitivity weight parameter is constructed to quantify the sensitivity of different risk events to error changes. The error sensitivity weight parameter is constructed by calculating the quantitative index of risk identification sensitivity of each specific feature in the causal correlation feature set of the reconstructed error when different risk events occur. The quantitative sensitivity index is calculated by comparing the average increase in the feature value in a fixed time period before the risk event when each real risk event occurs with the average increase in the feature value during the period before the risk event occurs. The degree of significance difference is obtained through comparison. The higher the degree of significance difference, the more sensitive the feature is to the occurrence of the risk event. The numerical value of the degree of significance difference is used to obtain the error sensitivity weight parameter corresponding to each feature in a standardized manner.
[0096] Dynamically weight the reconstructed error sequence according to the error sensitivity weight parameter and establish an error dynamic threshold adjustment strategy;
[0097] Based on the error sensitivity weight parameter, a dynamic weighted analysis is performed on the reconstructed error sequence to establish a dynamic error threshold adjustment strategy. Specifically, the error value corresponding to each time point in the reconstructed error sequence is weighted by the error sensitivity weight parameter corresponding to the time point to obtain the dynamic weighted error value at each time point. Based on the temporal variation trend of the dynamic weighted error value, an error threshold adjustment strategy that is dynamically adjusted over time is established. Specifically, within a period of time, the variation trend of the dynamic weighted error value is statistically analyzed, and the mean and standard deviation of the dynamic weighted error value in each time window are calculated in real time using a moving window method. The mean value within the window plus a certain multiple of the standard deviation is used as the dynamically adjusted threshold, thus forming an error threshold adjustment strategy that can respond to changes in risk status in real time.
[0098] According to the error dynamic threshold adjustment strategy, determine the error critical threshold applicable to the current agricultural production cycle information and rural policy cycle information;
[0099] Based on the error dynamic threshold adjustment strategy, the critical error threshold applicable to the current agricultural production cycle information and rural policy cycle information is determined. The critical error threshold determination method is as follows: using agricultural production cycle information and rural policy cycle information to determine the cycle synchronization time window, based on the cycle synchronization time window, the dynamic threshold corresponding to each node is calculated for the specific time period of each key node in the agricultural production cycle and the key node in the rural policy cycle. The dynamic thresholds of each node are aggregated and the maximum dynamic threshold is used to determine the critical error threshold applicable to the current time period of the comprehensive agricultural production cycle information and rural policy cycle information.
[0100] S5: The reconstruction error sequence is judged by the critical threshold, and the risk event warning result is output. The risk category and corresponding influencing factors are determined according to the risk assessment indicator system, including:
[0101] The error critical threshold is used to perform risk discrimination processing on the reconstructed error sequence at each time point to determine whether a risk event occurs at each time point;
[0102] The error value in the reconstructed error sequence is compared with the critical error threshold at each time point. If the error value in the reconstructed error sequence exceeds the critical error threshold at a certain time point, it is determined that a risk event exists at that time point. Otherwise, if it does not exceed the critical error threshold, it is determined that no risk event exists at that time point.
[0103] Establish a risk assessment indicator system;
[0104] The indicator system includes indicators for economic factors, social and environmental factors, agricultural production cycles, and rural policy cycles. The indicator system was constructed as follows: Economic indicators include rural loan balances, loan delinquency rates, per capita income levels, and farmers' debt repayment capacity; social and environmental indicators include rural population mobility rates, farmers' credit ratings, and village collective economic activity; agricultural production cycle indicators include crop planting area, crop yield change rates, and agricultural disaster frequency; and rural policy cycle indicators include loan interest rate policies, changes in subsidy funds, and policy coverage changes.
[0105] According to the risk event time points obtained by risk identification, the risk event data at the corresponding time points are extracted;
[0106] Based on the time of the risk event, data on the corresponding risk event is extracted from the historical time series of rural financial risk events. Specifically, if a risk event is determined to have occurred at a certain point in time, detailed data for that point in time and several points immediately before and after it is extracted from the historical time series of rural financial risk events. This data includes changes in loan amounts, details of loan delinquencies, policy implementation, agricultural production data, and indicators of the social environment, ensuring the completeness and accuracy of risk event data.
[0107] Correlate the risk event data with the risk assessment indicator system to determine the risk category of each risk event;
[0108] A feature correlation analysis was performed on the extracted risk event data and the risk assessment indicator system to determine the risk category to which each risk event belongs. This feature correlation analysis method involves using a cluster analysis algorithm, specifically the K-means cluster analysis method; standardizing the risk event data and the risk assessment indicator system data; using the K-means cluster analysis algorithm to divide all risk event data into several categories, with the number of categories determined using the elbow rule; and using Euclidean distance as the distance measure between feature data during the cluster calculation process to determine the category of each risk event through multiple iterations.
[0109] Map risk categories to risk assessment indicator systems to determine the influencing factors for each risk category;
[0110] A factor mapping analysis is conducted between each risk category and the indicators in the risk assessment indicator system to determine the influencing factors corresponding to each risk category. Specifically: the factor mapping analysis is conducted using a correlation coefficient analysis method, specifically selecting the Pearson correlation coefficient as the calculation method; for each risk category, the Pearson correlation coefficient is calculated between the risk event data and each indicator in the economic factor indicator, social environment factor indicator, agricultural production cycle indicator, and rural policy cycle indicator; if the absolute value of the Pearson correlation coefficient between an indicator and a specific risk category exceeds a specific threshold, the indicator is determined to be the main factor affecting the risk category; the threshold setting method is determined through cross-validation of expert experience and historical data to ensure the reliability and accuracy of the identified factors. For example, under the loan risk category, if the per capita income level indicator calculates a high negative correlation coefficient, the per capita income level is determined to be a significant influencing factor of the loan risk category; and under the policy risk category, if the loan interest rate policy indicator calculates a high positive correlation coefficient, the loan interest rate policy is determined to be a significant influencing factor of the policy risk category.
[0111] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field according to actual conditions.
[0112] The above embodiments can be implemented in whole or in part via software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in the embodiments of this application are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0113] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0114] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0115] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0116] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, and may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.
[0117] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0118] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0119] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0120] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A LSTM-autoencoder hybrid risk warning method for rural finance, characterized by: The steps include: S1: Based on the information of agricultural production cycle and rural policy cycle, the historical time series of rural financial risk events is obtained, and the migration features of the feature space distribution are extracted to form the agricultural policy drift feature set; S2: Perform multi-window time-scale layered processing on the agricultural policy drift feature set, build several independent LSTM sub-networks, perform asynchronous memory state encoding on the multi-scale features, and generate a multi-scale latent space encoding set; S3: Weighted aggregation of the multi-scale latent space encoding set is input into the autoencoder to reconstruct the time series of risk events and obtain the reconstruction error sequence; S4: Using real events of rural financial historical risks, we determine the critical threshold of reconstruction error through an error-sensitive threshold generation strategy based on causal tracing; S5: The reconstructed error sequence is judged through the critical threshold, the risk event warning result is output, and the risk category and corresponding influencing factors are determined according to the risk assessment indicator system.
2. The LSTM-autoencoder hybrid risk warning method for rural finance according to claim 1 is characterized in that: S1, specifically: Determine the cycle synchronization time window based on agricultural production cycle information and rural policy cycle information; Extract agricultural production cycle characteristics and rural policy cycle characteristics from the historical time series of rural financial risk events based on the cycle synchronization time window; The migration characteristics of agricultural production cycle characteristics and rural policy cycle characteristics are evaluated in spatial domain respectively, and the drift characteristic set of agricultural production cycle and rural policy cycle drift characteristic set are obtained; The agricultural production cycle drift feature set and the rural policy cycle drift feature set are fused in feature space to form the agricultural policy drift feature set.
3. The LSTM-autoencoder hybrid risk warning method for rural finance according to claim 2, characterized in that: The cycle synchronization time window is determined based on the agricultural production cycle information and rural policy cycle information, specifically: Divide the time series nodes corresponding to the agricultural production cycle according to the agricultural production cycle information to form an agricultural production cycle node set; According to the rural policy cycle information, the time series nodes corresponding to the rural policy cycle are divided to form a rural policy cycle node set; Conduct time series cross-comparison analysis on the agricultural production cycle node set and the rural policy cycle node set, and calculate the time synchronization index; According to the time synchronization index, several time series nodes with the highest synchronization between the agricultural production cycle node set and the rural policy cycle node set are selected to form the cycle synchronization time window.
4. The LSTM-autoencoder hybrid risk warning method for rural finance according to claim 3 is characterized in that: S2, specifically: Based on the agricultural policy drift feature set, it is divided into the first time scale feature window, the second time scale feature window and the third time scale feature window; Based on the first time scale feature window, the second time scale feature window and the third time scale feature window, independent long short-term memory network sub-networks are established respectively; The long short-term memory network sub-network is used to asynchronously encode the memory states of the feature sequences of the first time scale feature window, the second time scale feature window, and the third time scale feature window respectively; Output the first scale latent space coding set, the second scale latent space coding set and the third scale latent space coding set respectively; The time scale feature fusion processing is performed on the first scale latent space code set, the second scale latent space code set and the third scale latent space code set to obtain a multi-scale latent space code set.
5. The LSTM-autoencoder hybrid risk warning method for rural finance according to claim 4 is characterized in that: S3, specifically: Perform feature weight processing on the multi-scale latent space coding set respectively to obtain the first scale feature weight parameter, the second scale feature weight parameter and the third scale feature weight parameter; The multi-scale latent space coding set is weightedly aggregated using the first scale feature weight parameter, the second scale feature weight parameter, and the third scale feature weight parameter to form a weighted fusion feature coding; Using the weighted fusion feature code as the input of the autoencoder, a time series reconstruction network model based on the combined structure of long short-term memory network and autoencoder is established; The time series reconstruction network model is used to reconstruct the feature coding after weighted fusion to obtain the time series reconstruction output result; The data difference between the time series reconstruction output result and the feature coding after weighted fusion is calculated at each time point to obtain the reconstruction error sequence.
6. The LSTM-autoencoder hybrid risk warning method for rural finance according to claim 5, characterized in that: S4, specifically: According to the real risk events marked by the historical time series of rural financial risk events, the causal correlation of the reconstruction error sequence is processed to obtain the causal correlation feature set of the reconstruction error; Based on the causal correlation feature set of the reconstructed error, the error sensitivity weight parameter is constructed to determine the sensitivity of different risk events to error changes; Dynamically weight the reconstructed error sequence according to the error sensitivity weight parameter and establish an error dynamic threshold adjustment strategy; According to the error dynamic threshold adjustment strategy, the error critical threshold applicable to the current agricultural production cycle information and rural policy cycle information is determined.
7. The LSTM-autoencoder hybrid risk warning method for rural finance according to claim 6, characterized in that: S5, specifically: The error critical threshold is used to perform risk discrimination processing on the reconstructed error sequence at each time point to determine whether a risk event occurs at each time point; Establish a risk assessment indicator system; the indicator system includes economic factor indicators, social environment factor indicators, agricultural production cycle indicators and rural policy cycle indicators; According to the risk event time points obtained by risk identification, the risk event data at the corresponding time points are extracted; Correlate the risk event data with the risk assessment indicator system to determine the risk category of each risk event; Map risk categories and risk assessment indicator systems to determine the influencing factors corresponding to each risk category.
Citation Information
Cited By
A network congestion prediction method and an electronic device
CN122395075A