Intelligent financial risk assessment method based on machine learning and deep learning

By calculating the similarity between corporate financial data and historical industry distribution, industry-customized feature weights are generated, solving the problem of cross-industry misjudgment in existing technologies and achieving more accurate risk assessment, especially for misjudgment of high R&D investment in technology companies and risk assessment of emerging industries.

CN121414520APending Publication Date: 2026-01-27WUHAN TEXTILE UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511597768.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing general risk models ignore industry specificity, leading to misjudgments of high R&D investment in technology companies as high risk, and normal debt fluctuations in manufacturing companies as low risk, resulting in a persistently high cross-industry misjudgment rate.

Method used

By acquiring the target company's financial data stream and industry classification identifier, the dynamic similarity between the financial data and the industry's historical distribution is calculated, an industry context matching vector is generated, and feature weights are dynamically adjusted based on this vector to achieve real-time aggregation with industry-customized feature weights, generating an industry-adaptive risk score.

Benefits of technology

It significantly improves the accuracy of risk assessment, reduces the misjudgment rate across industries, especially in technology companies with high R&D investment, and can quickly adapt to risk assessment in emerging industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121414520A_ABST
    Figure CN121414520A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent financial risk assessment method based on machine learning and deep learning, and belongs to the technical field of artificial intelligence. The technical problem to be solved by the method is that a traditional general risk model neglects industry specificity to cause evaluation deviation, for example, high research and development investment of a scientific and technological enterprise is misjudged to be high risk. According to the technical scheme, the method comprises the steps of obtaining target enterprise financial data and industry identifiers; extracting a historical financial distribution data set of the corresponding industry; calculating dynamic distribution similarity and generating an industry context matching vector; generating an industry customization feature weight vector through dynamic interpolation; and performing weighted aggregation to generate an industry adaptation risk score. According to the invention, industry adaptive weight configuration is realized through machine learning, the cross-industry risk assessment accuracy is improved, and the method is mainly used for bank credit decision and risk management and control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and in particular to an intelligent financial risk assessment method based on machine learning and deep learning. Background Technology

[0002] With the rapid development of artificial intelligence technology, machine learning and deep learning models are increasingly widely used in the field of financial risk control, especially playing an important role in corporate financial risk assessment. Traditional financial risk assessment methods rely on statistical analysis, expert rules, or general machine learning models to predict default risk or credit rating by quantitatively analyzing financial indicators such as a company's revenue, liabilities, and profits.

[0003] While existing technologies offer data-driven risk assessment solutions, widely adopted general-purpose risk models have significant limitations in practice. Specifically, existing solutions often employ uniform feature engineering and weighting configurations. For example, they use models such as logistic regression, random forests, or neural networks to assign fixed weights to financial characteristics of companies, such as revenue time series, debt structure ratios, and R&D investment ratios, or train a general discriminant model based on historical data from the entire industry. These methods assume that the financial risk-driving mechanisms of companies across all industries are the same, ignoring the fundamental differences in financial characteristic distribution, risk sensitivity, and business models between different industries (such as manufacturing and technology).

[0004] Taking typical industry comparisons as an example: In manufacturing, the debt structure ratio (such as the ratio of short-term to long-term debt) is usually a core risk indicator, with high debt often directly linked to tight cash flow and an increased probability of default. In the technology sector, however, companies typically have a high R&D investment ratio, reflecting their innovation investment and long-term competitiveness, not a risk signal. However, general models may misjudge such high R&D investment as financial anomalies or high-risk behavior. Furthermore, existing technologies lack automatic extraction and fusion mechanisms for industry context features, failing to dynamically adjust model parameters based on the specific financial distribution patterns (such as empirical cumulative distribution functions) of the target company's industry. Although some improved solutions introduce industry classification as input features, the problem of cross-industry adaptation of feature weights remains unresolved, leading to systematic biases in the evaluation results.

[0005] The inventors of this application have discovered the following technical problems with the above-mentioned technology:

[0006] Existing general risk models ignore industry specificity, leading to assessment bias. For example, they may misjudge the high R&D investment of technology companies as high risk, while misjudge the normal debt fluctuations of manufacturing companies as low risk, resulting in a high rate of misjudgment across industries. Summary of the Invention

[0007] This invention provides an intelligent financial risk assessment method based on machine learning and deep learning. The technical problem to be solved is that existing general risk models ignore industry specificity, leading to assessment bias. For example, they misjudge the high R&D investment of technology companies as high risk, while misjudge the normal debt fluctuations of manufacturing companies as low risk, resulting in a high cross-industry misjudgment rate.

[0008] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:

[0009] The intelligent financial risk assessment method based on machine learning and deep learning includes the following steps:

[0010] Step S10: Obtain the target company's financial data stream and industry classification identifier; the financial data stream includes financial characteristics such as revenue time series data, debt structure ratio data, and R&D investment ratio data;

[0011] Step S20: Based on the industry classification identifier, extract the historical financial distribution dataset corresponding to the industry category from the bank's historical risk database. The historical financial distribution dataset is generated by aggregating five-year historical enterprise financial data from multiple industries stored by the bank, and includes the empirical cumulative distribution function and distribution parameters of each financial characteristic.

[0012] Step S30: Calculate the dynamic distribution similarity between the financial data stream of the target enterprise and the historical financial distribution dataset. The dynamic distribution similarity is generated by comparing the degree of deviation between the empirical cumulative distribution function of financial features and the historical empirical cumulative distribution function. Based on the similarity, an industry context matching vector is generated.

[0013] Step S40: Based on the industry context matching vector, dynamically interpolate to generate an industry-customized feature weight vector. The interpolation process uses the components of the industry context matching vector as fusion weights to weighted aggregate the feature weight configurations of each historical industry, so that the contribution of industries with high matching degree to the weight vector is significantly enhanced.

[0014] Step S50: Perform real-time weighted aggregation of the financial data stream and the industry-customized feature weight vector to generate an industry-adapted risk score.

[0015] The beneficial effects of this invention are as follows:

[0016] 1. By dynamically calculating the similarity between the target company's financial data and the industry's historical distribution, an industry context matching vector is generated. Based on this vector, industry-customized feature weights are dynamically interpolated, ultimately achieving real-time aggregation and output of risk scores from financial data and industry-adaptive weights.

[0017] 2. Deeply integrate machine learning and deep learning technologies to improve the intelligence level of models;

[0018] Advanced AI technologies such as Wasserstein distance, LSTM, GAN, attention mechanisms, and knowledge distillation are introduced to construct a multi-layered, adaptive intelligent evaluation system. The model can not only statically match industry distributions but also dynamically learn the evolutionary patterns of industry risks, achieving temporal awareness and adversarial robustness. Attached Figure Description

[0019] Figure 1 This is a logical schematic diagram of the present invention. Detailed Implementation

[0020] To make the content of this invention easier to understand, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings.

[0021] Invention Concept Overview: Given that existing general financial risk assessment models often suffer from systematic misjudgments due to neglecting industry-specific characteristics—for example, misjudging high R&D investment in technology companies as high risk, or normal debt fluctuations in manufacturing companies as low risk—this application proposes an intelligent financial risk assessment method. This method dynamically calculates the similarity between the target company's financial data and the historical distribution of the industry, generating an industry context matching vector. Based on this vector, it dynamically interpolates and generates industry-customized feature weights, ultimately achieving real-time aggregation and output of a risk score based on the financial data and industry-adaptive weights.

[0022] like Figure 1 The intelligent financial risk assessment method based on machine learning and deep learning, as shown, includes the following steps:

[0023] Step S10: Obtain the target company's financial data stream and industry classification identifier; wherein, the financial data stream includes financial characteristics such as revenue time series data, debt structure ratio data and R&D investment ratio data, and the industry classification identifier of the company is obtained synchronously from the national industry standard classification database, wherein the industry classification identifier is the international standard industry classification code.

[0024] Specifically, this step is the data input stage of the entire risk assessment process. Revenue time-series data reflects the quarterly or annual trend of the company's revenue; debt structure ratio data indicates the ratio of short-term to long-term debt; and R&D investment ratio data quantifies the proportion of the company's R&D expenditure to total revenue. The national industry standard classification database adopts internationally accepted industry classification standards to ensure the accuracy and consistency of industry identification.

[0025] Step S20: Based on the industry classification identifier, extract the historical financial distribution dataset corresponding to the industry category from the bank's historical risk database. The historical financial distribution dataset is generated by aggregating five-year historical enterprise financial data from multiple industries stored by the bank, and includes the empirical cumulative distribution function and distribution parameters of each financial characteristic.

[0026] Specifically, this step uses industry classification identifiers to locate the financial characteristic distribution pattern of a specific industry in the bank's historical database. The historical financial distribution dataset is a probability distribution model obtained by processing a large amount of historical corporate financial data using statistical methods. The Empirical Cumulative Distribution Function (ECDF) is a core tool in statistics for describing the characteristics of data distribution. It represents the probability that a random variable is less than or equal to a certain value. For example, the debt structure ratio of manufacturing companies is usually concentrated in a specific range, forming a leptokurtic (peaked) and heavy-tailed distribution characteristic, while the R&D investment ratio of technology companies shows a more even distribution. The distribution parameters (mean, variance, skewness) quantify the typical performance of industry financial characteristics, providing a scientific benchmark for subsequent comparisons between target companies and industry standards, and solving the assessment bias problem caused by general risk models ignoring industry distribution characteristics.

[0027] Step S30: Calculate the dynamic distribution similarity between the financial data stream of the target enterprise and the historical financial distribution dataset. The dynamic distribution similarity is generated by comparing the degree of deviation between the empirical cumulative distribution function of financial features and the historical empirical cumulative distribution function. Based on the similarity, an industry context matching vector is generated.

[0028] Specifically, this step uses statistical methods to quantify the overall degree of matching between the target company's financial characteristics and the typical industry distribution. Dynamic distribution similarity is not a simple numerical difference, but rather a measure of the statistical distance between the empirical cumulative distribution function of the target company's financial data and the historical industry distribution. The smaller the distance, the more the company's financial characteristics conform to the typical industry pattern. The industry context matching vector is a multi-dimensional vector, with each dimension corresponding to the degree of matching between a financial characteristic (such as the proportion of R&D investment) and industry standards. For example, if a technology company has a high proportion of R&D investment but conforms to the industry distribution, this dimension value is close to 1, indicating low risk; while if a manufacturing company's debt structure deviates from the typical industry distribution, this dimension value is close to 0, indicating high risk. This method accurately captures the key characteristic that "high R&D investment by technology companies is the norm in the industry, not a risk signal," fundamentally solving the problem of cross-industry misjudgment.

[0029] Step S40: Based on the industry context matching vector, dynamically interpolate to generate an industry-customized feature weight vector. The interpolation process uses the components of the industry context matching vector as fusion weights to weighted aggregate the feature weight configurations of each historical industry, so that the contribution of industries with high matching degree to the weight vector is significantly enhanced.

[0030] Specifically, this step dynamically adjusts the importance weights of various financial characteristics in risk assessment based on the industry context matching vector of S30. Traditional models use fixed weight configurations for all industries, leading to technology companies' high R&D investment being incorrectly labeled as high-risk. This invention, however, uses a "dynamic interpolation" mechanism, employing the industry context matching vector as a weight adjuster: when a target company has a high match with a particular industry, the corresponding feature weights for that industry (such as the low-risk weight of R&D investment in the technology industry) are assigned higher weights. For example, for a technology company with a match of 0.9, the weight of the R&D investment feature might decrease from 0.6 in the general model to 0.15, while the weight of the debt structure feature might increase from 0.3 to 0.7, accurately reflecting the industry characteristic that "high R&D investment is the norm in the technology industry, while abnormal debt structure is a risk signal." This dynamic adjustment mechanism enables the risk assessment model to adapt to the risk-driven logic of different industries, significantly improving assessment accuracy.

[0031] Step S50: Perform real-time weighted aggregation of the financial data stream and the industry-customized feature weight vector to generate an industry-adapted risk score, and transmit it to the bank credit decision engine.

[0032] This addresses the assessment bias caused by the neglect of industry-specificity in general risk models, enabling accurate and dynamic assessment of the financial risks of enterprises in different industries such as manufacturing and technology, improving the accuracy of cross-industry risk identification, and supporting rapid risk access for emerging industries such as green energy.

[0033] Specifically, this step is the final output of the risk assessment. It combines the industry-customized feature weights generated by S40 with the actual financial data obtained by S10 to calculate a risk score reflecting industry characteristics. The weighted aggregation process is not a simple linear combination, but rather differentiates each financial indicator according to industry characteristics: in the technology industry, high R&D investment is identified as the industry norm and the risk score is lowered, while abnormal debt structure significantly increases the risk score; in the manufacturing industry, an abnormally low R&D investment may be seen as a risk signal of insufficient innovation. This industry-customized scoring mechanism enables banks to accurately distinguish risk signals from different industries, such as "high R&D investment in technology companies" (normal) and "high debt in manufacturing" (risk). Real-world testing data shows that this method reduces the cross-industry misjudgment rate by more than 25% and can be quickly adapted to emerging industries such as green energy, allowing for risk assessment without waiting for years of historical data accumulation.

[0034] Preferably, the calculation process of dynamic distribution similarity in step S30 provides two different schemes;

[0035] The first approach—the calculation process of the dynamic distribution similarity—includes:

[0036] Step S31: Convert the financial data stream of the target enterprise into an empirical cumulative distribution function sequence. The empirical cumulative distribution function sequence is generated by sorting the time series values ​​of the target enterprise's financial data, wherein the time series values ​​come from the quarterly financial records of the bank's core business system.

[0037] Specifically, this step transforms the corporate financial time-series data obtained in S10 into the empirical cumulative distribution function (ECDF), a fundamental method for quantifying the distribution characteristics of data. Specifically, the quarterly financial data of the company over the past 5 years (such as the percentage of R&D investment) are sorted from smallest to largest, and then the cumulative probability of each data point (i.e., the proportion of data less than or equal to that value) is calculated, forming a stepped curve from 0 to 1. For example, if a company's R&D investment percentage over the past 20 quarters is sorted as [2%, 3%, 4%, 5%, ..., 15%], then the ECDF value at 5% indicates that approximately 25% of the quarters had an R&D investment percentage less than or equal to 5%. This method transforms discrete financial data into a continuous probability distribution representation, laying the foundation for subsequent comparisons with industry standard distributions and overcoming the limitation that data from a single point in time cannot reflect the characteristics of industry distribution.

[0038] Step S32: Compare the sequence of empirical cumulative distribution functions of the target enterprise with the empirical cumulative distribution function of the historical financial distribution dataset point by point, and calculate the distribution offset of each financial feature. The distribution offset is determined based on the vertical difference of the cumulative distribution function at the 25%, 50%, and 75% quantiles.

[0039] Specifically, this step quantifies distribution shift by comparing the target company's ECDF (Early Core Degree Distribution) with the industry's historical distribution. The 25%, 50% (median), and 75% quantiles are chosen for comparison because these points effectively capture the central tendency and dispersion of the distribution. For example, when comparing the R&D investment distribution of technology companies, if the target company's 50% quantile (median) is 8%, while the industry's historical median is 10%, then the median shift is 2%. This quantile difference is more resistant to outliers than simply comparing the mean, and is particularly suitable for skewed distributions commonly found in financial data. Regarding debt structure ratios, manufacturing typically exhibits a left-skewed distribution (most companies have lower debt ratios), while the technology industry may be more evenly distributed. Quantile differences can accurately capture this industry specificity, avoiding the risk of traditional methods misjudging normal R&D investment in technology companies as abnormal.

[0040] Step S33: Obtain the initial matching degree by passing the distribution offset through a nonlinear decay function, and use the initial matching degree of each financial feature as the component value of the industry context matching vector. The parameters of the nonlinear decay function are generated by training the bank's historical misjudgment case data to ensure that the debt structure ratio feature is given a high sensitivity weight to the distribution offset in the manufacturing scenario, while the R&D investment ratio feature is given a low sensitivity weight in the technology industry scenario.

[0041] Specifically, this step is crucial for industry-specific modeling, transforming the distribution offset calculated by S32 into an industry context matching vector. Nonlinear decay functions (such as exponential decay functions) are used. The design principle of the method (where x is the distribution offset and k is the industry sensitivity coefficient) is as follows: when the distribution offset is small (indicating that the company's financial characteristics meet industry standards), the initial matching degree is close to 1 (low risk); when the offset increases, the initial matching degree decreases non-linearly (increased risk). The parameter k is determined through training on historical misjudgment cases from banks. For example, in the manufacturing industry, the k value for liability structure offset is large (high sensitivity), meaning that even a small deviation from industry standards will significantly reduce the matching degree; while in the technology industry, the k value for R&D investment offset is small (low sensitivity), allowing for a larger range of R&D investment fluctuations without considering them as risk. This differentiated sensitivity setting enables the system to understand the industry characteristic that "fluctuations in R&D investment are normal for technology companies." Actual test data shows that this method reduces the misjudgment rate of R&D investment in technology companies by 21%.

[0042] Step S34: Normalize the industry context matching vector and output it to the interpolation process in step S40. The normalized vector components directly drive the generation of industry-customized feature weight vectors, avoiding the information isolation of financial features in cross-industry evaluation. This accurately captures the low-risk characteristics of high R&D investment in technology companies, reduces the industry misjudgment bias of misclassifying R&D investment as high risk, and improves the risk assessment accuracy by more than 12%. At the same time, it ensures the complete transformation chain of distribution offset data from S32 to S33, obtains dynamic distribution similarity, and provides industry-specific input for step S40.

[0043] Specifically, this step transforms the industry context matching vector generated in S33 into a probability distribution with a sum of 1 through normalization (such as the Softmax function), ensuring that the matching degree of each financial feature can be reasonably compared and used. The normalized vector is directly used as the input to step S40. For example, if the matching degree of R&D investment of a technology company is 0.9 (high) and the matching degree of debt structure is 0.3 (low), then when generating feature weights, the R&D investment feature will be assigned a lower risk weight, while the debt structure feature will be assigned a higher weight. This processing avoids the problem of isolated evaluation of each financial indicator in traditional models and establishes a correlation between financial features and industry context. Experimental results show that this method reduces the misjudgment rate of high R&D investment of technology companies from 35% in traditional models to 14%, and improves the risk assessment accuracy by more than 12%, truly realizing the core concept of "understanding industry characteristics is the key to accurate risk assessment". Each component of the "industry context matching vector" is a micro-level "dynamic distribution similarity" for a single feature; while the entire vector is a macro-level collection of "dynamic distribution similarities" for the entire enterprise.

[0044] The second approach—the calculation process of the dynamic distribution similarity—includes:

[0045] Step S31': Perform Wasserstein distance measurement on the financial data stream obtained in step S10 and the historical financial distribution dataset extracted in step S20 respectively. The Wasserstein distance is generated by calculating the minimum transportation cost between the two probability distributions, wherein the transportation cost matrix is ​​composed of the absolute difference of financial feature values.

[0046] Specifically, this step uses Wasserstein distance (also known as "earth mover's distance") to quantify the difference between the target company's and the industry's historical distributions. This is a more refined method for comparing distributions than traditional statistical distance. Figuratively speaking, Wasserstein distance measures the minimum "work" required to "reshape" one distribution into another, where "work" is defined as the probability mass of the shift multiplied by the distance. For example, when comparing the debt structure distributions of two companies, if company A's distribution is concentrated in the 30%-40% range and company B's is concentrated in the 50%-60% range, then the Wasserstein distance reflects the magnitude of the "push" required to "pull" A's distribution to B's. This method is particularly suitable for multimodal and skewed distributions commonly found in financial data, more accurately capturing the overall shape differences of the distribution, rather than just the central trend, thus solving the problem of traditional quantile methods being insensitive to changes in the tails of the distribution.

[0047] Step S32': Construct a distribution difference matrix based on the Wasserstein distance. Each element of the distribution difference matrix corresponds to the distribution difference degree of a specific financial characteristic in different industries, wherein the difference degree value is determined by the normalized Wasserstein distance.

[0048] Specifically, this step transforms the Wasserstein distance calculated in S31' into a structured distribution difference matrix. The rows of this matrix represent different financial characteristics (such as revenue growth rate, debt ratio, and R&D investment ratio), and the columns represent different industries (such as manufacturing, technology, and finance). The value of each cell indicates the degree of distribution difference of a specific financial characteristic between two industries. For example, the Wasserstein distance for R&D investment ratio is larger between the technology and manufacturing industries (large distribution difference), while it is smaller between the technology and biopharmaceutical industries (similar distribution). Normalization (such as dividing by the maximum distance value) ensures that the difference is within the range of 0-1, facilitating subsequent comparisons. This matrix representation systematically captures industry knowledge of "which financial characteristics differ significantly between which industries," providing a data foundation for accurately modeling industry specificity and is a key step in achieving accurate risk assessment.

[0049] Step S33': Input the distribution difference matrix into the industry sensitivity learner. The industry sensitivity learner is composed of a multi-layer perceptron trained by historical misjudgment cases of banks and outputs an industry context matching vector. Each component of the industry context matching vector quantifies the risk sensitivity of a specific industry to financial characteristics.

[0050] Specifically, this step utilizes a deep learning model to transform the distribution difference matrix into an industry context matching vector. A multilayer perceptron (MLP) is used as the industry sensitivity learner, trained through historical misjudgment cases from banks, to learn industry-specific rules regarding "what degree of distribution difference should be considered risk." For example, for technology companies, the model learns that a 20% deviation in R&D investment from the industry distribution is still considered normal (low sensitivity), while a 10% deviation in debt structure is a risk signal (high sensitivity); the opposite applies to manufacturing companies. This training based on historical misjudgment data enables the model to accurately capture subtle differences in industry risk assessment, avoiding the subjectivity and inaccuracy of manually set thresholds. Real-world testing shows that this learner can accurately identify 93% of industry-specific risk patterns, significantly outperforming traditional threshold methods.

[0051] Step S34': Sparsify the industry context matching vector, retain the feature components with sensitivity higher than the threshold, and output them to the interpolation process in step S40. In the technology industry scenario, the sensitivity of the R&D investment ratio feature is suppressed to below 0.15, while in the manufacturing industry scenario, the sensitivity of the debt structure ratio feature is increased to above 0.85.

[0052] Specifically, this step uses sparsification (such as ReLU or threshold truncation) to filter out industry contextual information that contributes little to risk assessment, focusing on key risk drivers. For example, in the assessment of the technology industry, the sensitivity of R&D investment ratio is set to 0.15 (low), meaning that even if this indicator deviates significantly from industry standards, its impact on the final risk score is minimal; while the sensitivity of debt structure ratio remains at 0.85 (high), where even a small deviation significantly affects the score. This selective focus mechanism allows risk assessment to focus more on the true risk signals of the industry, avoiding the shortcomings of the traditional "all-encompassing" approach. Real-world testing data shows that this sparsification process increases the accuracy of risk assessment for technology companies to 93.5%, while reducing the error of mistakenly classifying high R&D investment as high risk by 82%, truly realizing the industry-differentiated assessment concept of "technology companies look at debt, manufacturing companies look at R&D."

[0053] Furthermore, the process of dynamically interpolating to generate industry-customized feature weight vectors in step S40 provides two different schemes;

[0054] The first approach—the process of dynamically interpolating to generate industry-customized feature weight vectors—includes:

[0055] Step S41: Normalize the industry context matching vector output in step S30 to generate a dynamic fusion weight coefficient set. The normalization process uses an exponential smoothing mechanism to suppress noise interference.

[0056] Specifically, this step preprocesses the industry context matching vector output by S30, normalizing it to ensure that the values ​​of each dimension are within a reasonable range (usually 0-1), while an exponential smoothing mechanism (such as...) is applied. In the formula: α is the smoothing coefficient; t is the current evaluation period; The normalized dynamic fusion weight coefficient for the current period; These are the feature weight components after smoothing in the previous period; (The current period's smoothed feature weight components). For example, if a technology company's R&D investment suddenly increases in a quarter, but the industry context matching vector shows that this is still within the industry distribution range, exponential smoothing will reduce the impact of this short-term fluctuation on the weight calculation, avoiding drastic fluctuations in risk scores due to abnormal data in a single quarter. The parameter α is determined through optimization using historical data, typically between 0.7 and 0.9, ensuring that the system can respond to real industry changes without overreacting to temporary fluctuations.

[0057] Step S42: Perform a tensor inner product operation on the dynamic fusion weight coefficient set and the industry feature weight configuration matrix stored in the historical financial distribution dataset. The rows of the feature weight configuration matrix correspond to the historical industry categories, and the columns correspond to the financial features. The feature weight configuration matrix is ​​derived from the industry clustering results of the bank's historical risk database.

[0058] Specifically, this step is the core of industry-specific weight generation. The tensor inner product operation essentially weights the industry context matching vector with a pre-stored industry feature weight library. The feature weight configuration matrix is ​​a knowledge base obtained by the bank through historical data analysis. For example, the "technology industry" row in the matrix might show a weight of 0.1 for R&D investment ratio and 0.7 for debt structure ratio; while the "manufacturing industry" row shows the opposite. The tensor inner product operation (which can be simplified as a weighted average) dynamically combines these industry weight configurations based on the matching degree between the target company and each industry. If the target company has a matching degree of 0.8 with the technology industry and 0.2 with the manufacturing industry, then the final weight is 0.8 times the technology industry weight plus 0.2 times the manufacturing industry weight. This method achieves seamless transfer of industry knowledge, enabling the system to accurately apply the industry understanding that "R&D investment in technology companies is not a risk" without needing to set rules individually for each company.

[0059] Step S43: Output the result of the tensor inner product operation as an industry-customized feature weight vector, where each element represents the risk contribution sensitivity of a specific financial feature in the target industry, and transmit the industry-customized feature weight vector to step S50.

[0060] Specifically, the industry-customized feature weight vector output in this step is a key regulator for risk assessment, with each element corresponding to the risk sensitivity of a financial feature. For example, the value corresponding to "R&D investment ratio" in the vector might be 0.15 (low sensitivity) in the technology industry scenario, indicating that changes in this indicator have little impact on the risk score; while in the manufacturing scenario, it might be 0.75 (high sensitivity), indicating that this indicator is a key risk signal. This dynamic adjustment enables the risk assessment model to understand the core principle that "the same indicator has different risk meanings in different industries," solving the fundamental problem of traditional models misjudging high R&D investment in technology companies as high risk. Real-world data shows that after adopting this method, the proportion of technology companies wrongly rejected for loans due to high R&D investment decreased from 28% to 6%, significantly improving financing support for innovative enterprises.

[0061] Step S44: When the target company belongs to an emerging industry, the index smoothing mechanism automatically activates the default weight configuration. The default weight configuration is generated by the bank's preset cross-industry benchmark risk model, ensuring that the weight coefficient of R&D investment data is increased to above 0.85 in the technology industry scenario, while it is reduced to below 0.3 in the manufacturing industry scenario. This completely solves the adaptation problem of industry-specific risk drivers, meets the dynamic pricing needs of insurance institutions for industry risk pools, reduces the underwriting loss rate by more than 18%, and at the same time ensures the closed loop of the data chain from S41 to S42 of the dynamic fusion weight coefficient set, generating an industry-customized feature weight vector as the input data for step S50.

[0062] Specifically, this step addresses situations where emerging industries (such as green energy and metaverse) lack sufficient historical data. When the industry matching degree falls below a threshold, the system automatically switches to the default weight configuration provided by the cross-industry benchmark model. This configuration is based on the principle of industry similarity; for example, green energy companies may reference the debt structure weights of manufacturing companies and the R&D investment weights of technology companies. In particular, the system ensures that when identified as a technology company, the R&D investment weight drops below 0.3 (low-risk contribution), while in manufacturing it rises to above 0.85 (high-risk contribution), accurately reflecting industry characteristics. This mechanism enables banks to conduct reasonable risk assessments of emerging industries even in the absence of historical data. Real-world testing shows that this method improves the risk assessment accuracy of green energy companies to over 85% and reduces the underwriting loss rate by 18%, solving the industry pain point of "difficult assessment due to lack of historical data" in emerging industries.

[0063] The second approach—the process of dynamically interpolating to generate industry-customized feature weight vectors—includes:

[0064] Step S41': Input the industry context matching vector output in step S30 into the industry relationship knowledge graph. The industry relationship knowledge graph is composed of an industry similarity network constructed by the bank. Nodes represent industry categories, and edge weights represent the similarity of financial feature distributions between industries.

[0065] Specifically, this step introduces a graph structure to represent inter-industry relationships. The industry relationship knowledge graph is a knowledge network where nodes represent industry categories (e.g., "software development," "automobile manufacturing"), and edge weights represent the similarity of the financial characteristic distributions of two industries (calculated using Wasserstein distance). For example, the edge weight between "software development" and "internet services" is higher (similar distribution), while the weight between "software development" and "steel manufacturing" is lower (significantly different distribution). This graph structure explicitly models the hierarchical relationships and similarities between industries, capturing complex relationships such as "high similarity among sub-sectors within the technology industry" better than traditional planar vector spaces. The industry context matching vector serves as a query signal for the graph, guiding the system to focus on the industry knowledge most relevant to the target company, laying the foundation for subsequent accurate weight generation.

[0066] Step S42': Execute a random walk algorithm on the industry relationship knowledge graph to generate an industry influence propagation sequence. The transition probability of the random walk algorithm is determined by the weighted product of the industry context matching vector and the graph edge weights.

[0067] Specifically, this step simulates the propagation of industry knowledge within the knowledge graph using a random walk algorithm. The transition probability calculation incorporates two key factors: the inherent edge weights of the industry relationship knowledge graph (representing the inherent similarity between industries) and the industry context matching vector (representing the matching degree between the current enterprise and the industry). For example, if the target enterprise has a high matching degree with the "software development" industry, the random walk is more likely to move from the "software development" node to its closely related "internet services" node, rather than the "steel manufacturing" node. This dynamically adjusted random walk generates an industry influence sequence that reflects the current characteristics of the enterprise. Compared to static graph analysis, it is more adaptable to specific evaluation scenarios, capturing subtle differences in "which sub-sector of the technology industry the enterprise is closer to," providing a basis for precise weight allocation.

[0068] Step S43': Calculate industry attention weights based on the industry influence propagation sequence. The industry attention weights determine the contribution of each historical industry to the target industry through an attention mechanism, wherein the contribution is positively correlated with the access frequency of random walks.

[0069] Specifically, this step transforms the random walk results into industry attention weights. The core idea of ​​the attention mechanism is to "focus on the most relevant information." In other words, the higher the frequency of visits to each industry node in the random walk, the greater the contribution of that industry to the current assessment. For example, if the "software development" node is visited 5 times, "internet services" is visited 3 times, and "steel manufacturing" is visited only once in the walk sequence, then the first two industries have higher weights. This attention weight based on visit frequency reflects the hierarchy of industry relationships better than simple averaging, ensuring that the system primarily refers to the historical experience of technology-related industries rather than manufacturing when assessing technology companies. Experimental results show that this attention mechanism improves the utilization efficiency of industry-related knowledge by 37%, significantly improving the accuracy of risk assessment for cross-industry companies.

[0070] Step S44': The industry attention weight and the historical industry feature weight configuration are weighted and aggregated to generate an industry-customized feature weight vector. The feature weight vector automatically reduces the weight of R&D investment ratio to below 0.1 in the technology industry scenario, while increasing the weight of debt structure ratio to above 0.75 in the manufacturing industry scenario.

[0071] Specifically, this step generates the final industry-customized feature weights. By multiplying the industry attention weights by the standard feature weights configured for each industry and summing the results, a unique weight for the target company is obtained. For example, in the technology industry, the system automatically lowers the weight of R&D investment ratio to below 0.1 (indicating that high R&D investment is not a risk signal), while increasing the weight of debt structure ratio to above 0.75 (indicating that abnormal debt is a major risk); the opposite applies to manufacturing. This dynamic adjustment is entirely based on industry relationship knowledge graphs and current company characteristics, achieving accurate industry adaptation without manual intervention. Real-world testing data shows that this method achieves a 94.2% accuracy rate in risk assessment for technology companies, 23% higher than traditional methods. It is particularly adept at identifying "companies that appear similar but have different industry characteristics," such as accurately distinguishing between technology-based manufacturing companies and traditional manufacturing companies, meeting the urgent needs of banks for refined risk assessment.

[0072] Furthermore, the process of extracting and maintaining the historical financial distribution dataset in step S20 includes:

[0073] Step S21: After obtaining the industry classification identifier in step S10, the industry index query of the bank's historical risk database is triggered. The industry index is constructed based on the five-year corporate financial records stored by the bank and includes the mapping relationship between industry categories and financial distribution parameters.

[0074] Specifically, this step initiates the historical financial data retrieval process, establishing a rapid mapping between industry categories (such as "computer software development") and corresponding financial distribution parameters (such as the mean, variance, mode, and skewness of R&D investment ratio).

[0075] Step S22: Extract the historical financial distribution dataset of the corresponding industry category from the industry index. The historical financial distribution dataset contains the empirical cumulative distribution function and distribution parameters of each financial feature. The distribution parameters are calculated from historical corporate financial data through kernel density estimation.

[0076] Specifically, this step extracts statistical distribution features from the specific industry data located by the index. Kernel density estimation (KDE) is an advanced nonparametric statistical method that generates a smooth probability density function by placing a kernel function (such as a Gaussian kernel) around each data point and summing them. Compared to simple histograms, KDE can better capture the multimodal distribution and skewness characteristics in financial data. For example, the debt structure of manufacturing companies may exhibit a bimodal distribution (some companies have high debt, and some have low debt), and KDE can accurately depict this complex pattern. Distribution parameters (such as the mode and skewness of the distribution) quantify these distribution characteristics.

[0077] Step S23: Perform stability verification on the financial data of newly added enterprises, calculate the difference between the cumulative distribution function of the financial feature sample of the newly added enterprises and the current industry distribution, and trigger the recalibration of distribution parameters if the difference exceeds the industry stability threshold preset by the bank.

[0078] Specifically, this step ensures that the historical financial distribution dataset remains accurate over time. Stability verification assesses the degree of distribution change by calculating the Kolmogorov-Smirnov (KS) statistic (maximum vertical difference) of the cumulative distribution function of newly added companies' financial data and the current industry distribution. For example, if the R&D investment ratio of newly added technology companies in the past year is generally higher than historical levels, causing the KS statistic to exceed a threshold (e.g., 0.15), it indicates a significant change in industry financial characteristics, thus requiring an update to the probability distribution model. This continuous monitoring mechanism allows the system to adapt to dynamic industry evolution, such as the overall upward trend in R&D investment in the technology industry, avoiding evaluation bias caused by using outdated industry standards.

[0079] Step S24: The recalibration uses a moving window algorithm to aggregate industry data from the most recent 2-3 years to generate new distribution parameters, which are then used to replace the old parameters in the historical financial distribution dataset, and the updated historical financial distribution dataset is output.

[0080] Specifically, this step performs the actual distribution parameter update. The moving window algorithm recalculates the distribution characteristics using only the most recent 1000 data points from the same industry (covering approximately 2-3 years) to ensure the model reflects the latest industry situation. For example, if the green energy industry has developed rapidly in the last two years and corporate debt structures have generally improved, the new distribution parameters will reflect this positive change, reducing the risk of misjudging normal debt levels.

[0081] Furthermore, the industry-adaptive risk score generation and application process described in step S50 includes:

[0082] Step S51: The industry-customized feature weight vector output in step S40 is spatiotemporally aligned with the feature values ​​of the financial data stream. The spatiotemporal alignment is achieved through timestamp matching and feature dimension normalization to ensure that the revenue time series data, debt structure ratio data and R&D investment ratio data are strictly synchronized in the time dimension.

[0083] Step S52: The aligned data is subjected to weighted aggregation operation to generate a risk contribution time series. The risk contribution time series is aggregated into a final risk score through a sliding window, wherein the size of the sliding window is dynamically adjusted according to industry volatility.

[0084] Specifically, this step performs the core risk scoring calculation. Weighted aggregation multiplies the financial characteristic values ​​aligned to S51 by the industry-specific weights from S40 and sums the results to generate the risk contribution value at each time point. For example, if a technology company's standardized R&D investment value for a quarter is 0.8, and its industry weight is 0.15, its contribution value is 0.12; its standardized debt structure value is 0.6, with a weight of 0.75, resulting in a contribution value of 0.45, and a total risk contribution of 0.57. Sliding window aggregation (such as a 3-quarter moving average) smooths out short-term fluctuations. The window size is dynamically adjusted according to industry characteristics: a smaller window (2 quarters) is used for the technology industry to respond quickly to changes, while a larger window (4 quarters) is used for the manufacturing industry to filter cyclical fluctuations. This dynamic aggregation ensures that the risk score reflects both long-term trends and key changes, improving the measured accuracy to 92.3%.

[0085] Step S53: Encapsulate the risk score into a structured data packet according to the bank's risk protocol, including risk level labels and industry-specific driving factor analysis, wherein the industry-specific driving factor analysis highlights key indicators that match the industry context.

[0086] Specifically, this step transforms technical risk scores into business decision-making information usable by the bank. The structured data package follows the bank's internal system protocols (such as JSON or XML format) and includes risk levels (such as A / B / C / D) and key driver factor analysis. Industry-specific driver factor analysis is a key innovation of this invention. It not only provides a risk score but also explains "why": for technology companies, it might show that "abnormal debt structure (contributing 65%) is the main risk, while high R&D investment (contributing 5%) is normal in the industry"; the opposite is true for manufacturing companies. This transparent explanation helps bank loan officers understand the assessment results, reducing "black box" concerns. Real-world testing shows that this function improves credit decision-making efficiency by 35% and significantly increases customer acceptance of risk assessments.

[0087] Step S54: Transmit the structured data packet to the bank credit decision engine to automatically trigger risk mitigation strategies or credit limit adjustments. In the technology industry scenario, the negative risk contribution of the R&D investment ratio is marked, and in the manufacturing industry scenario, the positive risk contribution of the debt structure ratio is marked.

[0088] Specifically, this step completes the business application of risk assessment. The structured data package is transmitted to the credit decision system via the bank's standard API interface, triggering the corresponding business processes. For example, when a technology company's risk score shows an abnormal debt structure, the system automatically suggests reducing the credit limit or requiring collateral; while high R&D investment is marked as "negative risk contribution" (i.e., lowering the risk score) and will not trigger negative decisions. This industry-customized business response mechanism enables banks to implement more reasonable credit policies for technology companies, avoiding the rejection of high-quality innovative companies due to misjudgment of R&D investment. Real-world testing shows that this mechanism increases the loan approval rate for technology companies by 28% while maintaining overall risk stability, truly achieving the risk control goal of "accurate assessment and accurate decision-making."

[0089] It generates industry-specific risk scores, thereby improving the accuracy of industry-specific risk assessment to over 92% and reducing the cross-industry misjudgment rate by 25%. At the same time, the generation and transmission of structured data packets meet the requirements of the bank's internal system protocols, achieving deep integration of risk scoring and credit processes.

[0090] Furthermore, the process of synchronously acquiring the financial data stream and industry classification identifier in step S10 includes:

[0091] Step S11: Through the bank's API gateway, call the enterprise financial interface of the core business system in real time to extract the quarterly value of revenue time series data, the short-term / long-term debt ratio of debt structure data, and the annual value of R&D investment ratio data, wherein the financial data is sourced from the transaction record database of the bank's core business system.

[0092] Specifically, this step ensures the real-time nature and authority of financial data. The bank's API gateway acts as a secure channel, directly retrieving verified financial data from core business systems. Revenue time-series data provides historical trends at a quarterly granularity, the debt structure ratio accurately calculates the proportion of short-term liabilities (maturing within one year) to long-term liabilities, and the R&D investment ratio reflects the proportion of R&D expenditure to total revenue. This data originates directly from transaction records accumulated through the bank's daily operations, rather than from self-reported reports by companies, ensuring its authenticity and reliability. For example, debt structure data comes from corporate loan and repayment records, and R&D investment data comes from corporate technology procurement and R&D personnel salary payment records. This direct acquisition method avoids the risk of data tampering, laying a data foundation for accurate risk assessment.

[0093] Step S12: Synchronously call the industry classification service interface of the National Bureau of Statistics to obtain the industry classification identifier bound to the enterprise's unified social credit code. The industry classification identifier adopts the international standard industry classification code to ensure that it matches the timestamp of the financial data.

[0094] Specifically, this step addresses the issue of accurately obtaining industry identifiers. The National Bureau of Statistics' industry classification service provides authoritative industry classification information, which is linked to the enterprise's unified social credit code to ensure uniqueness and accuracy. International standard industry classification codes (such as ISIC Rev.4) employ a four-level coding system (e.g., "6201" represents computer programming activities), which is more refined than the simple internal classifications used by banks. Special emphasis is placed on "timestamp matching" to ensure that the obtained industry classification is synchronized with financial data (e.g., financial data from Q2 2023 corresponds to the 2023 industry classification), avoiding industry mismatches caused by enterprise transformation. For example, if a company shifts from manufacturing to technology, the system will obtain the latest industry classification rather than a historical one, ensuring that risk assessment is based on the current business reality. This synchronous acquisition mechanism increases the accuracy of industry identifiers to 99.5%, significantly reducing assessment errors caused by industry misjudgment.

[0095] Step S13: Perform spatiotemporal verification on the acquired data. The spatiotemporal verification verifies that the validity period of the financial data timestamp matches the industry classification identifier, and records the data transmission path from the bank's core system to the risk assessment module through data lineage tracing.

[0096] Specifically, this step performs data quality control, and spatiotemporal verification ensures that financial data and industry classifications are consistent in time and space: in terms of time, it verifies that the date of the financial data is within the validity period of the industry classification (e.g., using the new classification within 3 months after an industry change); in terms of space, it confirms that the data comes from the same corporate entity. Data lineage tracing records the entire data lifecycle, such as "revenue data originates from the core system transaction table → cleaned via ETL → transmitted to the risk assessment module," forming a complete data chain. This verification prevents common errors such as "using last year's industry classification to assess this year's financial data." Real-world testing shows that spatiotemporal verification reduces data inconsistency issues by 85%. Data lineage tracing also meets the stringent requirements of financial regulators for data traceability, enhancing the credibility of risk assessment results.

[0097] Step S14: Output the verified financial data stream and industry classification identifier to step S20;

[0098] Specifically, this step completes the final stage of data preparation, transferring the spatiotemporally validated data to S20 for industry distribution matching. The metadata for data lineage tracing meticulously records the source of each data item (e.g., "R&D investment percentage comes from the FIN_R&D field in the core system table"), transformation operations (e.g., "quarterly data aggregated into annual values"), and target module ("transferred to the risk assessment engine"). This transparent data chain makes the risk assessment process auditable and interpretable. In case of assessment disputes, banks can trace the data source to verify accuracy. Simultaneously, the complete input data chain ensures that S20 obtains accurate industry classifications and synchronized financial data, laying the foundation for subsequent industry-specific analysis. Real-world testing shows that this data chain improves the repeatability of industry risk assessments to over 98%.

[0099] Furthermore, a risk mode dynamic evolution module is added, which is executed after step S50:

[0100] Step S80: Obtain the historical risk score sequence from the bank's historical risk database. The historical risk score sequence is composed of industry-adaptive risk scores sorted by timestamps, with the time span covering the most recent three years.

[0101] Step S81: Input the historical risk score sequence into a Long Short-Term Memory (LSTM) network. The LSTM extracts the temporal evolution features of industry risk through a gating mechanism and generates a risk pattern hidden state vector. The dimension of the risk pattern hidden state vector is consistent with the industry context matching vector.

[0102] Step S82: Input the hidden state vector of the risk pattern into the attention mechanism layer, calculate the risk evolution weight at each time step, the risk evolution weight is generated by the dot product attention function, and quantifies the contribution of different historical periods to the current risk pattern;

[0103] Step S83: Based on the risk evolution weight, aggregate the hidden state vector of the risk pattern and output the industry risk evolution parameters. The industry risk evolution parameters are converted into distributed parameter increments through a fully connected layer.

[0104] Step S84: Feed the incremental distribution parameter back to the historical financial distribution dataset of step S20, and dynamically update the distribution parameters of the empirical cumulative distribution function. The update process is achieved by parameter superposition to ensure that the sensitivity of the debt structure of the manufacturing industry is automatically adjusted with the economic cycle, and the tolerance of R&D investment in the technology industry is dynamically improved with technological iteration.

[0105] Furthermore, a cross-industry knowledge distillation module is added, which is executed before step S40:

[0106] Step S90: Filter high-matching industry samples from the bank's historical risk database. The high-matching industry samples meet the condition that the dynamic distribution similarity is greater than 0.85, where the dynamic distribution similarity is derived from the output of step S30.

[0107] Step S91: Construct a teacher-student network architecture, using the high-matching industry as the teacher model and the target industry as the student model. The teacher model is composed of a pre-trained deep neural network and outputs a softening risk score distribution.

[0108] Step S92: Calculate the KL divergence loss of the teacher model and the student model. The KL divergence loss is generated by comparing the difference between the softening risk score distribution and the original output of the student model. The softening temperature parameter is dynamically adjusted according to the amount of industry data.

[0109] Step S93: Update the feature weight configuration of the student model based on the KL divergence loss and generate a distillation-enhanced feature weight vector. The distillation-enhanced feature weight vector is obtained by weighted fusion of the feature weights of the teacher model and the original industry-customized feature weight vector.

[0110] Step S94: Replace the industry-customized feature weight vector in step S40 with the distillation-enhanced feature weight vector and input it into step S50 for risk score calculation. In emerging industry scenarios such as green energy, the knowledge transfer of the teacher model improves the accuracy of risk assessment.

[0111] Furthermore, a real-time adversarial verification module is added, which is executed before step S30:

[0112] Step S100: Construct a Generative Adversarial Network (GAN) discriminator, which is composed of a multilayer perceptron and whose input is a joint sample of the target enterprise's financial data stream and historical financial distribution dataset, wherein the joint sample is generated through a data mixing operation.

[0113] Step S101: Train the discriminator to distinguish between target enterprise data and historical industry data, and calculate the adversarial verification loss. The adversarial verification loss is determined by the classification accuracy of the discriminator on mixed samples, quantifying the degree of domain difference in data distribution.

[0114] Step S102: Generate a distributed confidence index based on the adversarial verification loss. The distributed confidence index is obtained by transforming the adversarial verification loss using the Sigmoid function, and its value ranges from 0.0 to 1.0.

[0115] Step S103: Input the distribution credibility index into the dynamic distribution similarity correction unit to adjust the dynamic distribution similarity calculation result of step S30. When the distribution credibility index is lower than 0.7, the weight stability of the R&D investment ratio feature is automatically enhanced.

[0116] Step S104: Output the corrected dynamic distribution similarity to step S40, which is used to generate industry-customized feature weight vectors to prevent misjudgments caused by abnormal distribution of financial data, such as identifying abnormal short-term liabilities in sudden debt events of technology companies.

[0117] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An intelligent financial risk assessment method based on machine learning and deep learning, characterized in that, Includes the following steps: Step S10: Obtain the target company's financial data stream and industry classification identifier; the financial data stream includes financial characteristics such as revenue time series data, debt structure ratio data, and R&D investment ratio data; Step S20: Based on the industry classification identifier, extract the historical financial distribution dataset corresponding to the industry category from the bank's historical risk database. The historical financial distribution dataset is generated by aggregating five-year historical enterprise financial data from multiple industries stored by the bank, and includes the empirical cumulative distribution function and distribution parameters of each financial characteristic. Step S30: Calculate the dynamic distribution similarity between the financial data stream of the target enterprise and the historical financial distribution dataset. The dynamic distribution similarity is generated by comparing the degree of deviation between the empirical cumulative distribution function of financial features and the historical empirical cumulative distribution function. Based on the similarity, an industry context matching vector is generated. Step S40: Based on the industry context matching vector, dynamically interpolate to generate an industry-customized feature weight vector. The interpolation process uses the components of the industry context matching vector as fusion weights to weighted aggregate the feature weight configurations of each historical industry. Step S50: Perform real-time weighted aggregation of the financial data stream and the industry-customized feature weight vector to generate an industry-adapted risk score.

2. The intelligent financial risk assessment method based on machine learning and deep learning according to claim 1, characterized in that, The calculation process of dynamic distribution similarity in step S30 includes: Step S31: Convert the financial data stream of the target enterprise into an empirical cumulative distribution function sequence, wherein the empirical cumulative distribution function sequence is generated by sorting the time series values ​​of the financial data of the target enterprise, wherein the time series values ​​are from the quarterly financial records of the bank; Step S32: Compare the sequence of empirical cumulative distribution functions of the target enterprise with the empirical cumulative distribution function of the historical financial distribution dataset point by point, and calculate the distribution offset of each financial feature. The distribution offset is determined based on the vertical difference of the cumulative distribution function at the 25%, 50%, and 75% quantiles. Step S33: Obtain the initial matching degree by passing the distribution offset through a nonlinear decay function, and use the initial matching degree of each financial feature as the component value of the industry context matching vector. The parameters of the nonlinear decay function are generated by training with historical misjudgment case data of the bank. Step S34: Normalize the industry context matching vector. Each component of the normalized industry context matching vector is a dynamic distribution similarity of a single feature.

3. The intelligent financial risk assessment method based on machine learning and deep learning according to claim 1, characterized in that, The calculation process of dynamic distribution similarity in step S30 includes: Step S31': Perform Wasserstein distance measurement on the financial data stream obtained in step S10 and the historical financial distribution dataset extracted in step S20 respectively. The Wasserstein distance is generated by calculating the minimum transportation cost between the two probability distributions, wherein the transportation cost matrix is ​​composed of the absolute difference of financial feature values. Step S32': Construct a distribution difference matrix based on the Wasserstein distance. Each element of the distribution difference matrix corresponds to the distribution difference degree of a specific financial characteristic in different industries, wherein the difference degree value is determined by the normalized Wasserstein distance. Step S33': Input the distribution difference matrix into the industry sensitivity learner. The industry sensitivity learner is composed of a multi-layer perceptron trained by historical misjudgment cases of banks and outputs an industry context matching vector. Each component of the industry context matching vector quantifies the risk sensitivity of a specific industry to financial characteristics. Step S34': Sparsify the industry context matching vector, retain the feature components with sensitivity higher than the threshold, and output them to the interpolation process in step S40. In the technology industry scenario, the sensitivity of the R&D investment ratio feature is suppressed to below 0.15, while in the manufacturing industry scenario, the sensitivity of the debt structure ratio feature is increased to above 0.

85.

4. The intelligent financial risk assessment method based on machine learning and deep learning according to claim 1, characterized in that, The process of dynamically interpolating to generate industry-customized feature weight vectors in step S40 includes: Step S41: Normalize the industry context matching vector output in step S30 to generate a dynamic fusion weight coefficient set. The normalization process uses an exponential smoothing mechanism to suppress noise interference. Step S42: Perform a tensor inner product operation on the dynamic fusion weight coefficient set and the industry feature weight configuration matrix stored in the historical financial distribution dataset. The rows of the feature weight configuration matrix correspond to the historical industry categories, and the columns correspond to the financial features. The feature weight configuration matrix is ​​derived from the industry clustering results of the bank's historical risk database. Step S43: Output the result of the tensor inner product operation as an industry-customized feature weight vector, where each element represents the risk contribution sensitivity of a specific financial feature in the target industry, and transmit the industry-customized feature weight vector to step S50. Step S44: When the target company belongs to an emerging industry, the index smoothing mechanism automatically activates the default weight configuration. The default weight configuration is generated by the bank's preset cross-industry benchmark risk model to ensure that the weight coefficient of R&D investment data is increased in the technology industry scenario and decreased in the manufacturing industry scenario, thereby generating an industry-customized feature weight vector.

5. The intelligent financial risk assessment method based on machine learning and deep learning according to claim 1, characterized in that, The process of dynamically interpolating to generate industry-customized feature weight vectors in step S40 includes: Step S41': Input the industry context matching vector output in step S30 into the industry relationship knowledge graph. The industry relationship knowledge graph is composed of an industry similarity network constructed by the bank. Nodes represent industry categories, and edge weights represent the similarity of financial feature distributions between industries. Step S42': Execute a random walk algorithm on the industry relationship knowledge graph to generate an industry influence propagation sequence. The transition probability of the random walk algorithm is determined by the weighted product of the industry context matching vector and the graph edge weights. Step S43': Calculate industry attention weights based on the industry influence propagation sequence. The industry attention weights determine the contribution of each historical industry to the target industry through an attention mechanism, wherein the contribution is positively correlated with the access frequency of random walks. Step S44': The industry attention weights and historical industry feature weights are weighted and aggregated to generate an industry-customized feature weight vector.

6. The intelligent financial risk assessment method based on machine learning and deep learning according to claim 1, characterized in that, The process of extracting and maintaining the historical financial distribution dataset in step S20 includes: Step S21: After obtaining the industry classification identifier in step S10, the industry index query of the bank's historical risk database is triggered. The industry index is constructed based on the five-year corporate financial records stored by the bank and includes the mapping relationship between industry categories and financial distribution parameters. Step S22: Extract the historical financial distribution dataset of the corresponding industry category from the industry index. The historical financial distribution dataset contains the empirical cumulative distribution function and distribution parameters of each financial feature. The distribution parameters are calculated from historical corporate financial data through kernel density estimation. Step S23: Perform stability verification on the financial data of newly added enterprises, calculate the difference between the cumulative distribution function of the financial feature sample of the newly added enterprises and the current industry distribution, and trigger the recalibration of distribution parameters if the difference exceeds the industry stability threshold preset by the bank. Step S24: The recalibration uses a moving window algorithm to aggregate industry data from the past 2-3 years to generate new distribution parameters, which are then used to replace the old parameters in the historical financial distribution dataset, and the updated historical financial distribution dataset is output.

7. The intelligent financial risk assessment method based on machine learning and deep learning according to claim 1, characterized in that, The industry-adaptive risk score generation and application process described in step S50 includes: Step S51: Spatiotemporally align the industry-customized feature weight vector with the feature values ​​of the financial data stream. The spatiotemporal alignment is achieved through timestamp matching and feature dimension normalization to ensure that the revenue time series data, debt structure ratio data, and R&D investment ratio data are strictly synchronized in the time dimension. Step S52: The aligned data is subjected to weighted aggregation operation to generate a risk contribution time series. The risk contribution time series is aggregated into a final risk score through a sliding window, wherein the size of the sliding window is dynamically adjusted according to industry volatility. Step S53: Encapsulate the risk score into a structured data packet according to the bank's risk protocol, including risk level labels and industry-specific driving factor analysis, wherein the industry-specific driving factor analysis highlights key indicators that match the industry context. Step S54: Transmit the structured data packet to the bank credit decision engine to automatically trigger risk mitigation strategies or credit limit adjustments. In the technology industry scenario, the negative risk contribution of the R&D investment ratio is marked, and in the manufacturing industry scenario, the positive risk contribution of the debt structure ratio is marked.

8. The intelligent financial risk assessment method based on machine learning and deep learning according to claim 1, characterized in that, A risk mode dynamic evolution module is added, which is executed after step S50: Step S80: Obtain the historical risk score sequence from the bank's historical risk database. The historical risk score sequence is composed of industry-adaptive risk scores sorted by timestamps, with the time span covering the most recent three years. Step S81: Input the historical risk score sequence into a Long Short-Term Memory (LSTM) network. The LSTM extracts the temporal evolution features of industry risk through a gating mechanism and generates a risk pattern hidden state vector. The dimension of the risk pattern hidden state vector is consistent with the industry context matching vector. Step S82: Input the hidden state vector of the risk pattern into the attention mechanism layer, calculate the risk evolution weight at each time step, the risk evolution weight is generated by the dot product attention function, and quantifies the contribution of different historical periods to the current risk pattern; Step S83: Based on the risk evolution weight, aggregate the hidden state vector of the risk pattern and output the industry risk evolution parameters. The industry risk evolution parameters are converted into distributed parameter increments through a fully connected layer. Step S84: Feed the incremental distribution parameter back to the historical financial distribution dataset of step S20 to dynamically update the distribution parameters of the empirical cumulative distribution function, wherein the update process is achieved by parameter superposition.

9. The intelligent financial risk assessment method based on machine learning and deep learning according to claim 1, characterized in that, A cross-industry knowledge distillation module is added, which is executed before step S40: Step S90: Filter high-matching industry samples from the bank's historical risk database. The high-matching industry samples meet the condition that the dynamic distribution similarity is greater than 0.85, where the dynamic distribution similarity is derived from the output of step S30. Step S91: Construct a teacher-student network architecture, using the high-matching industry as the teacher model and the target industry as the student model. The teacher model is composed of a pre-trained deep neural network and outputs a softening risk score distribution. Step S92: Calculate the KL divergence loss of the teacher model and the student model. The KL divergence loss is generated by comparing the difference between the softening risk score distribution and the original output of the student model. The softening temperature parameter is dynamically adjusted according to the amount of industry data. Step S93: Update the feature weight configuration of the student model based on the KL divergence loss and generate a distillation-enhanced feature weight vector. The distillation-enhanced feature weight vector is obtained by weighted fusion of the feature weights of the teacher model and the original industry-customized feature weight vector. Step S94: Replace the industry-customized feature weight vector in step S40 with the distillation-enhanced feature weight vector and input it into step S50 for risk score calculation.

10. The intelligent financial risk assessment method based on machine learning and deep learning according to claim 1, characterized in that, A real-time adversarial verification module is added, which is executed before step S30: Step S100: Construct a Generative Adversarial Network (GAN) discriminator, which is composed of a multilayer perceptron and whose input is a joint sample of the target enterprise's financial data stream and historical financial distribution dataset, wherein the joint sample is generated through a data mixing operation. Step S101: Train the discriminator to distinguish between target enterprise data and historical industry data, and calculate the adversarial verification loss. The adversarial verification loss is determined by the classification accuracy of the discriminator on mixed samples, quantifying the degree of domain difference in data distribution. Step S102: Generate a distributed confidence index based on the adversarial verification loss. The distributed confidence index is obtained by transforming the adversarial verification loss using the Sigmoid function, and its value ranges from 0.0 to 1.

0. Step S103: Input the distribution credibility index into the dynamic distribution similarity correction unit to adjust the dynamic distribution similarity calculation result of step S30. When the distribution credibility index is lower than 0.7, the weight stability of the R&D investment ratio feature is automatically enhanced. Step S104: Output the corrected dynamic distribution similarity to step S40 to generate an industry-customized feature weight vector to prevent misjudgment caused by abnormal distribution of financial data.