Enterprise loan willingness prediction method and device, computer equipment and storage medium
By building a deep neural network model under the multi-task learning framework, combining multi-dimensional data and the dynamic influence of core enterprises, the problem of insufficient prediction accuracy in traditional methods is solved, and high-precision prediction of corporate loan intention is achieved.
Patent Information
- Application Number
- CN202510098409.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional corporate loan intention prediction methods rely too much on limited financial data and static models, ignore the dynamic impact of multi-dimensional data and core enterprises, resulting in insufficient prediction accuracy and inability to fully reflect the actual loan needs and potential risks of the company.
By obtaining the company's basic information, supply chain roles, transaction records, industrial and commercial information change records and public information, performing data preprocessing, cleaning and feature engineering, building a deep neural network model under the multi-task learning framework, and combining the dynamic influence of core enterprises, loan intention prediction is carried out.
It has achieved high-precision prediction of corporate loan intentions, overcomes the shortcomings of traditional methods, can more comprehensively reflect the company's loan needs and potential risks, and improves the financial institutions' refined management capabilities of loan risks.
Smart Images

Figure CN120046777A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computers, and more specifically to a method, device, computer equipment, and storage medium for predicting the loan willingness of enterprises. Background Art
[0002] At present, the prediction of the loan willingness of enterprises has been studied and applied to a certain extent in the financial field. Traditional loan willingness prediction methods usually rely on some simple statistical models, such as linear regression, decision trees, etc. These methods mainly rely on limited financial data, credit scores, and enterprise historical loan records.
[0003] However, traditional prediction methods usually only rely on basic information such as enterprise financial statements, credit scores, and historical loan records, ignoring multi-dimensional data such as industrial and commercial registration information, industry trends, and transaction relationships between enterprises. This single data source and narrow feature dimension limit the accuracy of the model and are difficult to comprehensively reflect the actual loan demand and potential risks of enterprises. Traditional machine learning algorithms used in many current methods, such as logistic regression, random forest, etc., are often difficult to effectively integrate different types of data when facing multi-source heterogeneous data, resulting in low prediction accuracy of the model and unable to meet the needs of financial institutions for refined management of loan risks. Many prediction systems rely on static models and are unable to flexibly respond to dynamic changes in the market environment or enterprise operating conditions. For example, the credit rating, transaction scale, etc. of core enterprises may change, and these factors may have an important impact on the loan willingness, but existing prediction systems usually cannot update and reflect these changes in real time. In traditional methods, the influence of core enterprises is often simplified or ignored, resulting in the inability to accurately predict the loan willingness of upstream and downstream enterprises in the supply chain. Factors such as the transaction relationship and financial stability between enterprises and core enterprises should be regarded as important prediction variables, but many existing technologies have failed to effectively capture and model this key relationship.
[0004] Therefore, it is necessary to design a new method to solve the technical problems that traditional enterprise loan willingness prediction methods rely too much on limited financial data and static models, ignore multi-dimensional data and the dynamic influence of core enterprises, resulting in insufficient prediction accuracy and unable to comprehensively reflect the actual loan demand and potential risks of enterprises. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects of the prior art and provide a method, device, computer equipment, and storage medium for predicting the loan willingness of enterprises.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A method for predicting the loan willingness of enterprises, including:
[0007] Obtain the basic information of the enterprise to be predicted, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the data of the enterprise's industrial and commercial information change records, and the public information of the enterprise to obtain the initial data;
[0008] Preprocess and clean the initial data, and perform feature engineering processing based on the cleaned data to obtain the processing result;
[0009] Input the processing result into the loan willingness prediction model to predict the loan willingness of the enterprise, so as to obtain the loan willingness score and the application rate score;
[0010] Output the loan willingness score and the application rate score.
[0011] Its further technical solution is: The preprocessing and cleaning of the initial data, and the feature engineering processing based on the cleaned data to obtain the processing result include:
[0012] Perform missing value processing, outlier detection, and data normalization processing on the initial data to obtain the cleaned data;
[0013] Construct new features based on the cleaned data, and cross-process the new features to obtain combined features;
[0014] Use dimensionality reduction technology to process the combined features to obtain the processing result.
[0015] Its further technical solution is: The loan willingness prediction model is obtained by acquiring historical data, where the historical data includes the basic information of the enterprise, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the data of the enterprise's industrial and commercial information change records, and the public information of the enterprise, and forming a sample set through preprocessing, cleaning, and feature engineering processing of the historical data, and then training a deep neural network under a multi-task learning framework.
[0016] Its further technical solution is: The loan willingness prediction model is obtained by acquiring the basic information of the enterprise, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the data of the enterprise's industrial and commercial information change records, and the public information of the enterprise, and forming a sample set through preprocessing, cleaning, and feature engineering processing, and then training a deep neural network under a multi-task learning framework, including:
[0017] Obtain the basic information of the enterprise, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the data of the enterprise's industrial and commercial information change records, and the public information of the enterprise, and perform preprocessing, cleaning, and feature engineering processing to obtain a sample set;
[0018] Construct a deep neural network under a multi-task learning framework;
[0019] Define the loss function;
[0020] Use the sample set in combination with the loss function and adopt the gradient descent optimization algorithm to train the deep neural network to obtain a loan willingness prediction model.
[0021] Its further technical solution is that the deep neural network under the multi-task learning framework includes an input layer, a feature fusion layer, a shared deep learning network, and an output layer. Among them, the shared deep learning network includes multiple fully connected layers, batch normalization, activation functions, and Dropout layers.
[0022] Its further technical solution is that the loss function is a function obtained by summing the loan willingness loss function and the application behavior loss function with weighted coefficients.
[0023] Its further technical solution is that the use of the sample set in combination with the loss function and the adoption of the gradient descent optimization algorithm to train the deep neural network to obtain a loan willingness prediction model includes:
[0024] Define positive samples and negative samples according to the sample set, and process the positive samples and negative samples using undersampling, oversampling, or weighted loss techniques to obtain a processed sample set;
[0025] Use the processed sample set in combination with the loss function and adopt the gradient descent optimization algorithm to train the deep neural network to obtain a loan willingness prediction model.
[0026] The present invention also provides an enterprise loan willingness prediction device, including:
[0027] A data acquisition unit for acquiring basic information of the enterprise to be predicted, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the enterprise's industrial and commercial information change record data, and the public information of the enterprise to obtain initial data;
[0028] A processing unit for preprocessing and cleaning the initial data and performing feature engineering processing based on the cleaned data to obtain a processing result;
[0029] A prediction unit for inputting the processing result into the loan willingness prediction model to predict the enterprise loan willingness to obtain a loan willingness score and an application rate score;
[0030] An output unit for outputting the loan willingness score and the application rate score.
[0031] The present invention also provides a computer device, which includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program, the above-mentioned method is implemented.
[0032] The present invention also provides a storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.
[0033] The beneficial effects of the present invention compared with the prior art are as follows: By integrating multi-source data such as supply chain roles, transaction records, and industrial and commercial changes, the present invention breaks through the limitation of traditional methods that only rely on financial data. First, obtain the basic information of the enterprise to be predicted and its role in the supply chain, and combine the enterprise's transaction history and industrial and commercial change records to form comprehensive initial data. Then, preprocess and clean these initial data to remove noise and ensure data quality. Next, perform feature engineering based on the cleaned data to extract multi-dimensional features that have an important impact on the loan willingness. Subsequently, input these features into the loan willingness prediction model, and comprehensively evaluate the enterprise's loan demand and potential risks in combination with the dynamic influence of the core enterprise. Finally, output the loan willingness score and the application acceptance rate score of the enterprise to achieve high-precision loan prediction and overcome the deficiencies of traditional static models.
[0034] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 It is a schematic diagram of the application scenario of the enterprise loan willingness prediction method provided by the embodiment of the present invention;
[0037] Figure 2 It is a schematic flowchart of the enterprise loan willingness prediction method provided by the embodiment of the present invention;
[0038] Figure 3 It is a schematic diagram of the sub-process of the enterprise loan willingness prediction method provided by the embodiment of the present invention Figure 1 ;
[0039] Figure 4 It is a schematic diagram of the sub-process of the enterprise loan willingness prediction method provided by the embodiment of the present invention Figure 2 ;
[0040] Figure 5Schematic block diagram of the enterprise loan willingness prediction device provided by the embodiment of the present invention;
[0041] Figure 6 Schematic block diagram of the computer device provided by the embodiment of the present invention. Detailed implementation manners
[0042] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0043] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0044] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0045] It should be further understood that the term " / and / " used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.
[0046] Please refer to Figure 1 and Figure 2 , Figure 1 Schematic diagram of the application scenario of the enterprise loan willingness prediction method provided by the embodiment of the present invention. Figure 2Schematic flowchart of the enterprise loan willingness prediction method provided by the embodiments of the present invention. This enterprise loan willingness prediction method is applied to a server. The server conducts data interaction with a terminal. By combining multi-dimensional data such as the basic information of an enterprise, industry trends, transaction records, and changes in industrial and commercial information, it transcends the traditional single data source that relies on financial statements and credit scores, comprehensively improving the accuracy of prediction. Using a deep neural network under a multi-task learning framework, by integrating different types of heterogeneous data, it overcomes the deficiencies of traditional machine learning algorithms in multi-source data fusion, thereby improving the prediction accuracy of the model. Constructing a flexible model architecture that can reflect the changes in the market environment and the operating conditions of enterprises in real time, adapting to the impacts of fluctuations in the credit ratings and transaction scales of core enterprises, and enhancing the timeliness and accuracy of prediction. Through meticulous cleaning, processing, and feature engineering of the data, combined with dimensionality reduction techniques and feature crossing, combined features are constructed to effectively capture and model the transaction relationships between enterprises and their impacts on loan willingness. Designing a comprehensive loss function, combining the dual objectives of loan willingness and application behavior, and optimizing model training through weighted summation to further improve the accuracy and reliability of loan willingness prediction.
[0047] Figure 2 is a schematic flowchart of the enterprise loan willingness prediction method provided by the embodiments of the present invention. As Figure 2 shown, this method includes the following steps S110 to S140.
[0048] S110. Obtain the basic information of the enterprise to be predicted, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the data of the changes in the industrial and commercial information of the enterprise, and the public information of the enterprise to obtain initial data.
[0049] In this embodiment, a T+1 update mechanism is adopted to ensure that the data can timely reflect the latest enterprise situation. The field collection covers basic information such as the enterprise name, registered capital, registration type (such as limited liability company), registered province, city, district or county, tax registration number, industry classification (from the first level to the fourth level), creation time, and update time, as specifically shown in Table 1.
[0050] Table 1. Basic information of the enterprise
[0051]
[0052] Collect the transaction records between the enterprise and the core enterprise and its role in the supply chain, including financial information such as the transaction scale level in the past 12 months, the number of investment enterprises, the number of shareholders, and whether it is an A-level taxpayer. The data types are for reference as shown in Table 2.
[0053] Table 2. Transaction records
[0054]
[0055]
[0056] In this embodiment, the core enterprise plays a crucial role in the supply chain, and its influence has a direct impact on the loan demands and credit status of upstream and downstream enterprises. During the prediction process of loan willingness, the core enterprise is not only an important reference factor for the risk control and credit granting decisions of financial institutions, but also the manifestation of its influence at different levels and dimensions is crucial for the prediction accuracy of the model. In order to ensure that the influence of the core enterprise can be accurately reflected in the loan willingness prediction model, the present invention conducts a detailed analysis and innovation on the definition and processing logic of the core enterprise.
[0057] Core enterprises are usually the leaders in the supply chain, having direct and indirect impacts on upstream and downstream enterprises. The core enterprise is defined from the following dimensions:
[0058] Transaction scale: Core enterprises usually conduct frequent large-scale transactions with multiple upstream and downstream enterprises, and the transaction volume is usually higher than a certain threshold. By analyzing information such as the transaction scale and frequency between an enterprise and other enterprises, core enterprises can be effectively identified. If the transaction volume between an enterprise and multiple suppliers and customers is at a high level in the past 12 months, it can be determined as a core enterprise.
[0059] Industry ranking: Core enterprises usually occupy a leading position in their respective industries, and may be the industry leaders or enterprises with a large market share. By analyzing data such as the industry ranking, market share, and historical status of an enterprise in the industry, its importance in the industry can be measured. The industry ranking is an important indicator for measuring the status of core enterprises, and the influence of core enterprises in the supply chain can be further quantified by analyzing their competitiveness, market share, etc.
[0060] Credit rating: Core enterprises usually have a high credit rating, enabling them to obtain funds at a lower financing cost. Therefore, the credit rating of an enterprise (such as the AAA level) is an important sign for judging its core status. If an enterprise has a high credit rating, it indicates a high level of trust in the supply chain and is a potential core enterprise.
[0061] Enterprise stability: The operating state of core enterprises is usually relatively stable, and they can maintain a good financial condition for a long time. Their stability is manifested by indicators such as continuous profitability, low debt ratio, and healthy cash flow. These factors are important bases for judging whether an enterprise is a core enterprise.
[0062] In order to overcome the limitation of traditional methods that rely only on a single data dimension to identify core enterprises, this embodiment proposes an innovative judgment method based on technologies such as multi-source heterogeneous data fusion, network analysis and deep learning. By integrating data from multiple dimensions, core enterprises can be identified more accurately and their impact on loan willingness prediction can be evaluated. Specifically, the method of this embodiment integrates data from different sources and types (such as transaction records, industry rankings, credit ratings, financial information, etc.) for the first time to comprehensively identify core enterprises. This multi-dimensional data fusion can effectively avoid the identification errors that may be caused by relying on a single data source, thereby improving the accuracy of core enterprise identification. For example, by comprehensively analyzing information such as the company's transaction history, industry rankings, and financial status, it can be more accurately determined whether the company plays a core role in the supply chain.
[0063] The method of this embodiment introduces a transaction network analysis method, combined with graph analysis algorithms (such as degree centrality, betweenness centrality, etc.), to quantify the core position of enterprises in the supply chain. These methods can identify the importance of enterprises in upstream and downstream transaction networks and assign weights according to their positions in the network, thereby further improving the accuracy of the loan willingness prediction model.
[0064] The method of this embodiment adopts a dynamic update mechanism to respond to market changes, regularly track the business registration information, transaction data and industry data of enterprises, and adjust the influence score of enterprises based on these data changes. For example, changes in the transaction scale or credit rating of an enterprise will directly affect its core position and weight distribution in the model, thereby optimizing the accuracy of loan willingness prediction.
[0065] In the feature extraction process, the method of this embodiment adopts a multi-dimensional feature cross method to combine multi-dimensional data such as transaction size and financial health status to generate new composite features. Combined with deep learning technology, the system can automatically mine the complex nonlinear relationship between enterprise data, provide a more accurate core enterprise assessment for the model, and thus improve the prediction ability of loan willingness.
[0066] According to the definition of core enterprises above, the system needs to process and process relevant data to identify core enterprises in the supply chain and measure their impact on loan willingness. The specific processing logic is as follows:
[0067] Based on indicators such as transaction size, industry ranking, and credit rating, the data analysis system identifies core enterprises in the supply chain. These core enterprises will be marked as special categories and processed separately in subsequent data analysis. The basis for identifying core enterprises includes data such as the transaction behavior of the enterprise with upstream and downstream enterprises, industry position, and financial robustness.
[0068] When evaluating the loan willingness of enterprises, special attention should be paid to the transaction relationships between core enterprises and their upstream and downstream enterprises. By analyzing data such as the transaction scale, payment behavior, and contract performance between enterprises and core enterprises, the financial support and credit impact of core enterprises on upstream and downstream enterprises can be deeply understood. For example, the larger the transaction scale of the core enterprise and the more stable the payment behavior, the stronger the loan willingness of downstream enterprises is usually.
[0069] By combining multi-dimensional information such as the transaction behavior, credit rating, and financial status of core enterprises, calculate their influence on upstream and downstream enterprises. The influence of core enterprises may be reflected in aspects such as the degree of financial support for upstream and downstream enterprises, the improvement of credit ratings, or the enhancement of loan willingness. Through comprehensive analysis of these data, the system assigns appropriate weights to core enterprises in the loan willingness prediction model, so as to more accurately predict the loan demand and risks of upstream and downstream enterprises. The calculation process is based on the following key data:
[0070] The transaction frequency and amount between core enterprises and upstream and downstream enterprises;
[0071] The credit rating and debt situation of core enterprises;
[0072] The stability (such as operating years, profitability) and financial health status of core enterprises.
[0073] Comprehensively evaluate the stability and credit rating of core enterprises to quantify their impact on upstream and downstream enterprises. The evaluation indicators include the tax payment records of enterprises, historical loan repayment situations, debt levels, and shareholder changes. Through this multi-dimensional information, the system can dynamically adjust the influence of core enterprises to ensure the accuracy of loan willingness prediction.
[0074] The status of core enterprises may change due to factors such as industry cycles, economic environments, and the development of the enterprises themselves. To cope with this change, the system needs to regularly update the relevant data of core enterprises, especially the transaction scale, credit rating, and financial status, and dynamically adjust the weights in the loan willingness prediction model according to this information. This dynamic adjustment mechanism can improve the adaptability and accuracy of the model.
[0075] The role of core enterprises in loan willingness prediction is not only to provide data support, but also to greatly influence the financing decisions of upstream and downstream enterprises. Specifically, factors such as the stability, credit rating, and industry ranking of core enterprises will significantly affect the loan demand of other enterprises in the supply chain. Through comprehensive analysis of these key factors, the present invention can provide more accurate risk assessments for financial institutions, optimize the loan approval process, improve the credit granting efficiency, and reduce the bad debt risk. Based on the precise definition and innovative processing logic of core enterprises, the system can effectively extract the influence of enterprises in the supply chain, thereby optimizing the loan willingness prediction model and further enhancing the risk control ability and business decision-making efficiency of financial institutions.
[0076] In addition, collect the change records of the enterprise in the industrial and commercial information, including the number of changes in registered capital, the number of changes in shareholders, the number of changes in legal persons, the number of changes in business addresses, etc., and record the change frequency in the past 24 months to evaluate the stability of the enterprise. The reference of data types is shown in Table 3.
[0077] Table 3. Industrial and Commercial Change Records
[0078]
[0079]
[0080] Extract multiple features from the enterprise's public data, such as registered capital, registered province, number of invested enterprises, number of shareholders, etc. In addition, it also includes features such as the transaction relationship between the enterprise and the core enterprise, and the change rate of registered capital in the past 24 months. The reference of data types is shown in Table 4.
[0081] Table 4. Features Corresponding to Enterprise Public Data
[0082]
[0083]
[0084] S120. Preprocess and clean the initial data, and perform feature engineering processing based on the cleaned data to obtain a processing result.
[0085] In this embodiment, the processing result refers to the cleaned data set obtained after a series of processing steps such as missing value processing, outlier detection, data normalization, feature construction, feature crossing, and feature dimensionality reduction; this data set has been optimized to be able to reflect more potential patterns and relationships between features, while maintaining a low dimension to avoid the problem of data overfitting; this data set will be used as the basis for subsequent modeling for machine learning algorithms or statistical models to perform training, prediction, and analysis.
[0086] In one embodiment, please refer to Figure 3 , the above step S120 may include steps S121 to S123.
[0087] S121. Perform missing value processing, outlier detection, and data normalization processing on the initial data to obtain the cleaned data.
[0088] In this embodiment, for the missing values in the initial data set, use appropriate processing methods to fill or delete. For example, mean filling, interpolation method, or directly deleting the samples containing missing values.
[0089] The goal of missing value handling is to ensure data integrity and avoid negative impacts of missing values on model training.
[0090] Use statistical methods (such as box plots, z-scores, IQR, etc.) to detect and identify outliers in the data.
[0091] Outliers may indicate data entry errors or genuine extreme data values, and need to be processed according to business requirements (such as removal or replacement).
[0092] The purpose of outlier handling is to improve data quality and avoid interference of these values on model training.
[0093] Normalize numerical features to ensure that features are within the same scale range.
[0094] For example, use standardization (z-score standardization) or Min-Max normalization methods.
[0095] Data normalization helps to eliminate differences in different feature dimensions and scales, ensuring that each feature has a comparable influence on the model, especially important in models sensitive to distance metrics (such as KNN, SVM, etc.).
[0096] The result of these operations is the cleaned data, that is, a version where missing values, outliers have been processed, feature values have been standardized, and data quality has been improved.
[0097] S122. Construct new features based on the cleaned data, and cross-process the new features to obtain combined features.
[0098] In this embodiment, new derivative features are constructed based on the cleaned data. These features are obtained by combining or calculating existing features, aiming to provide more business insights or reflect potential patterns. For example: transaction frequency, capital expansibility, business stability, industry relevance, etc. These new features involve aggregation, calculation or transformation of existing data, and can help the model better understand the business background of the data.
[0099] Construct new combined features by crossing existing features. For example, cross the enterprise scale with the industry type, and cross the core enterprise transaction scale with financial indicators to capture more hidden relationships. This feature crossing helps the model to mine complex non-linear relationships, thereby improving the prediction accuracy.
[0100] The goal of these operations is to construct combined features that can provide more valuable information for subsequent modeling by enhancing the data representation ability.
[0101] S123. Process the combined features using dimensionality reduction techniques to obtain a processing result.
[0102] In this embodiment, for the high-dimensional feature space, dimensionality reduction techniques (such as principal component analysis PCA, linear discriminant analysis LDA, etc.) are used to reduce the feature dimension and reduce the complexity of the feature space. The purpose of dimensionality reduction is to remove redundant features, reduce the risk of overfitting, and improve the training efficiency of the model. After dimensionality reduction, a new, lower-dimensional feature set is usually obtained, which can effectively retain the important information of the original data.
[0103] S130. Input the processing result into the loan willingness prediction model to predict the enterprise's loan willingness, so as to obtain the loan willingness score and the application rate score.
[0104] In this embodiment, the data obtained and processed in the early stage is input into the pre-trained loan willingness prediction model to obtain two key scores: the loan willingness score and the application rate score. These two scores are important bases for financial institutions to evaluate the loan demand and behavior of enterprises. They not only reflect the loan intention of enterprises, but also reflect the possibility of enterprises submitting loan applications in actual operations.
[0105] The loan willingness score refers to the intensity of the enterprise's demand for loans predicted by the model, usually expressed as a probability value between 0 and 1. The higher this score, the more likely the enterprise is to have a loan demand.
[0106] The application rate score refers to the predicted score of the possibility of the enterprise submitting a loan application, which is also a probability value between 0 and 1. A higher score means that the enterprise is more likely to formally submit a loan application.
[0107] In order to comprehensively determine the loan willingness of the enterprise, financial institutions need to consider both the loan willingness score and the application rate score at the same time. This is because the simple loan willingness score can only reflect whether the enterprise hopes to obtain a loan, but it may not accurately predict whether it will take actual actions to apply for a loan. Similarly, it is difficult to fully understand the real loan demand of the enterprise only based on the application rate score, because some enterprises may show strong application intentions for various reasons (such as market environment, internal decision-making, etc.), but actually do not have real capital needs.
[0108] Therefore, combining these two scores can help financial institutions more accurately evaluate the loan willingness of enterprises. When the loan willingness score is high and the application rate score is also high, it indicates that the enterprise not only has a loan demand but is also very likely to take action. Such enterprises can be given priority consideration. If the loan willingness score is relatively high while the application rate score is relatively low, it may imply that although the enterprise has an interest in loans, it is not eager to apply for loans under the current circumstances, which may be due to waiting for a better opportunity or immature conditions. Conversely, if the application rate score is high while the loan willingness score is low, this situation is more complex and requires further investigation into whether the enterprise has hidden financial pressures or other special circumstances.
[0109] Among them, the loan willingness prediction model is obtained by acquiring historical data, where the historical data includes enterprise basic information, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the enterprise's industrial and commercial information change record data, and the enterprise's public information, and forming a sample set through preprocessing, cleaning, and feature engineering processing of the historical data, and then training a deep neural network under a multi-task learning framework.
[0110] In one embodiment, please refer to Figure 4 , the above-mentioned loan willingness prediction model is obtained by acquiring enterprise basic information, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the enterprise's industrial and commercial information change record data, and the enterprise's public information, and forming a sample set through preprocessing, cleaning, and feature engineering processing, and then training a deep neural network under a multi-task learning framework, and may include steps S131 to S134.
[0111] S131. Acquire enterprise basic information, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the enterprise's industrial and commercial information change record data, and the enterprise's public information, and perform preprocessing, cleaning, and feature engineering processing to obtain a sample set.
[0112] In this embodiment, the above content is acquired from historical data. For specific preprocessing, cleaning, and feature engineering processing, etc., steps S110 to S120 can be referred to and will not be elaborated here.
[0113] S132. Construct a deep neural network under a multi-task learning framework.
[0114] In this embodiment, the deep neural network under the multi-task learning framework includes an input layer, a feature fusion layer, a shared deep learning network, and an output layer, where the shared deep learning network includes multiple fully connected layers, batch normalization, activation functions, and Dropout layers.
[0115] Specifically, the input layer: The input layer receives various feature data of the enterprise, including numerical features (such as financial data, enterprise scale, industry category, etc.) and text features (such as industrial and commercial information, transaction records, etc.). These features first need to be appropriately preprocessed, such as normalization and categorical encoding (one-hot encoding or embedding encoding, etc.), and then passed into the network.
[0116] Feature fusion layer: At this layer, numerical features and text features will be fused to form a unified feature vector. The goal of the feature fusion layer is to effectively combine different types of input data, enhance the learning ability of the model, so that the subsequent deep learning network can comprehensively process.
[0117] Shared deep learning network: The shared deep learning network contains multiple fully connected layers, batch normalization, activation functions (ReLU or other non-linear activation functions), and Dropout layers. The network learns and extracts features through these layers, capturing the deep patterns of the data. Importantly, multiple tasks (loan willingness prediction and application rate prediction) share the hidden layers of this part of the network, so as to be able to share and transfer valuable feature information.
[0118] Output layer: The output layer of the model is divided into two branches:
[0119] Loan willingness prediction layer: Through the Sigmoid activation function, it outputs the probability value of the enterprise's loan willingness, indicating whether the enterprise has a loan demand and the intensity of the demand.
[0120] Application rate prediction layer: Similarly using the Sigmoid activation function, it outputs the probability of the enterprise's application, indicating the possibility of the enterprise submitting a loan application.
[0121] Secondly, the mapping relationship between features and network layers is shown in Table 5.
[0122] Table 5. Mapping relationship between features and network layers
[0123]
[0124]
[0125] The core advantage of this multi-task learning framework is that through the shared deep learning network, it can optimize multiple objectives (loan willingness and application rate) simultaneously, improving the overall performance and generalization ability of the model.
[0126] S133. Define the loss function.
[0127] In this embodiment, the loss function includes a function obtained by summing the loan willingness loss function and the application submission behavior loss function with weighted coefficients.
[0128] In this embodiment, the loan willingness loss represents the error between the loan willingness predicted by the model (i.e., whether the enterprise has a loan demand, the intensity of the loan demand, etc.) and the actual situation. The loan willingness loss is calculated using binary cross-entropy loss, and the goal is to minimize the difference between the predicted loan willingness and the actual loan willingness. The specific calculation formula is: LoanWillingness Loss = -(y true log(y pred ) + (1 - y true ) log(1 - y pred )); where y true is the actual label (loan willingness, taking values of 0 or 1), and y pred is the probability of the loan willingness predicted by the model (the value range is from 0 to 1).
[0129] The application submission behavior loss measures the difference between the application submission behavior predicted by the model (i.e., whether the enterprise submits a loan application) and the actual situation. This goal is calculated using binary cross-entropy loss, and the goal is to minimize the difference between the predicted application submission behavior and the actual application submission behavior. The specific calculation formula is as follows: SubmissionBehaviorLoss = -(b true log(b pred ) + (1 - b true ) log(1 - b pred )); where b true is the actual label (application submission behavior, taking values of 0 or 1), and b pred is the probability of the application submission behavior predicted by the model (the value range is from 0 to 1).
[0130] Since the two target tasks have different importance, a hyperparameter a is used to balance the weights of the loan willingness loss and the application submission behavior loss. The weighted combined loss function is as follows: Loss = a · Loan Willingness Loss + (1 - a) · Submission Behavior Loss; where a is a hyperparameter that controls the weight of the loan willingness loss in the total loss. By adjusting the value of a, the importance of the two goals can be flexibly adjusted according to the actual application scenario. If more attention is paid to the prediction accuracy of the loan willingness, the value of a is larger; if more emphasis is placed on the prediction of the application submission behavior, the value of a is smaller.
[0131] Specifically, in practical applications, the value of a usually needs to be adjusted through experiments and verification. Common adjustment methods include:
[0132] Grid search: Traverse different values of a (e.g., 0.1, 0.2, 0.5, 0.7, 1.0), evaluate the performance of the model under different weights, and select the optimal a.
[0133] Cross-validation: Select the optimal value of a through cross-validation and evaluate the effect of the model under different combinations of loss weighting.
[0134] By flexibly adjusting a, the model can achieve good results in both loan willingness prediction and application behavior prediction, thereby optimizing the loan approval process and business efficiency.
[0135] S134. Use the sample set and the loss function to train the deep neural network using the gradient descent optimization algorithm to obtain a loan willingness prediction model.
[0136] In one embodiment, the above step S134 may include steps S1341 to S1342.
[0137] S1341. Define positive samples and negative samples according to the sample set, and process the positive samples and negative samples using undersampling, oversampling, or weighted loss techniques to obtain a processed sample set.
[0138] S1342. Use the processed sample set and the loss function to train the deep neural network using the gradient descent optimization algorithm to obtain a loan willingness prediction model.
[0139] In this embodiment, the training step includes the following two sub-steps:
[0140] During the training process, it is first necessary to define positive and negative samples and process the data. The definition and processing methods of positive and negative samples are crucial for the training effect of the model, especially when dealing with imbalanced data. Common sample definitions are as follows:
[0141] Positive and negative samples for the loan willingness task: Positive samples refer to those enterprises that have expressed a clear loan demand (e.g., through an outbound call connection, the enterprise indicates a loan demand); negative samples are those situations where the enterprise clearly states that it has no loan demand.
[0142] Positive and negative samples for the application behavior task: Positive samples are those enterprises that have entered the loan application process (i.e., submitted a loan application); negative samples are those enterprises that have not submitted a loan application.
[0143] In addition, during the processing of the sample set, it is necessary to apply undersampling, oversampling, or weighted loss techniques to the data to address the problem of imbalance between positive and negative samples in the dataset. This helps to improve the performance of the model on imbalanced data and reduce the bias of the model towards negative samples.
[0144] Specifically, before model training, it is necessary to preprocess the data to ensure data quality. The preprocessing steps include handling missing values, data normalization, encoding categorical data, etc.
[0145] Use a gradient descent optimization algorithm (such as the Adam optimizer) to train the model, optimize the model's parameters by minimizing the loss function, and perform cross-validation.
[0146] In model training, the definition and handling of positive and negative samples are crucial. Positive and negative samples involve the following aspects:
[0147] Whether the outbound call is connected: Whether the enterprise has successfully connected the outbound call is usually used as an indicator to judge whether the enterprise has a need for a loan.
[0148] Whether there is a clear feedback of no need: Whether the enterprise clearly indicates that it does not need a loan. If there is such feedback, it is usually regarded as a negative sample.
[0149] Whether the application is submitted: Whether the enterprise has entered the loan application process. Enterprises that have submitted applications are usually regarded as positive samples.
[0150] Whether there is a need but the application is not submitted: For enterprises that have a need for a loan but have not submitted an application, it may be necessary to further analyze their potential needs to determine whether they can be regarded as positive samples.
[0151] The definition and handling methods of these positive and negative samples will directly affect the training effect of the model. Especially when building a classification model, how to balance the proportion and quality of positive and negative samples is the key to improving the accuracy of the model. For an imbalanced dataset, techniques such as undersampling, oversampling, or weighted loss can be used for processing to ensure that the model has high accuracy in predicting the loan intention.
[0152] After processing the dataset, use a gradient descent optimization algorithm (such as the Adam optimizer) for model training. In the training process, by minimizing the loss function, update the weights of the model, and gradually improve the prediction accuracy of the loan intention and the application behavior. In each training step, calculate the gradient through the backpropagation algorithm and adjust the weights in the network to finally obtain an optimized prediction model for the loan intention and the application behavior.
[0153] Since the model involves multiple objectives, different metrics will be used to evaluate the performance of each objective during the evaluation process:
[0154] Loan intention evaluation metrics: Use AUC (Area Under the Curve) and accuracy to evaluate the prediction effect of the loan intention.
[0155] Application rate evaluation metrics: Use AUC and F1-score to evaluate the prediction effect of the application rate.
[0156] During the training process, methods such as cross-validation can also be used to evaluate the generalization ability of the model, and hyperparameters (such as learning rate, number of network layers, weighting coefficient a, etc.) can be adjusted to optimize the model performance.
[0157] The multi-task learning framework of this embodiment can optimize the two objectives of loan willingness and application behavior simultaneously through a shared deep neural network structure. By flexibly setting the weighting coefficient a in the loss function, the relative importance of the two tasks can be adjusted according to the actual business needs, and the model can be trained through the gradient descent optimization algorithm to improve the overall prediction effect.
[0158] In actual application, the latest financial data, business registration information, transaction records, etc. of the enterprise are input into the model in real time through the API interface. After the input data of the enterprise is predicted by the model, a loan willingness score and the application rate score are output.
[0159] S140. Output the loan willingness score and the application rate score.
[0160] In this embodiment, the loan willingness score and the application rate score are output to the terminal for display.
[0161] The method of this embodiment comprehensively considers data from multiple dimensions such as the business registration information, financial status, transaction records, and business changes of the enterprise, and can comprehensively evaluate the loan willingness of the enterprise. Especially on the basis of considering the influence of the core enterprise, the method of this embodiment can accurately identify the key factors that have an important impact on the loan willingness, thereby significantly improving the accuracy of the prediction.
[0162] The method of this embodiment can adjust the parameters of the loan willingness prediction model in real time through a dynamic update tracking mechanism of the core enterprise influence to adapt to the changes in the market and the enterprise itself. This dynamic adaptability ensures that the system can respond in a timely manner to the changes in the supply chain structure and industry cycle, and improves the reliability and effectiveness of the model in actual application.
[0163] The method of this embodiment can identify potential loan risks. Especially by evaluating the influence of the core enterprise, it helps financial institutions better identify potential credit risks, optimize the credit granting decision-making, and reduce the bad debt risk. Through the accurate loan willingness score, financial institutions can make accurate decisions in a shorter time.
[0164] The method of this embodiment uses automated feature extraction and machine learning models, which significantly improves the efficiency of loan willingness prediction. Financial institutions can reduce the time for manual review and data analysis, respond more quickly to loan applications, thereby improving customer satisfaction and reducing labor costs.
[0165] Specifically, the method of this embodiment significantly improves the accuracy of the loan willingness prediction model by integrating multi-source heterogeneous data such as the industrial and commercial registration information, financial status, transaction records, and shareholder information of enterprises. Traditional technologies mostly rely on a single data source, while the method of this embodiment can comprehensively evaluate the multi-dimensional characteristics of enterprises by innovatively integrating data from different fields, thereby improving the prediction accuracy.
[0166] The method of this embodiment solves the defect in the prior art that the influence of core enterprises fails to be fully utilized by accurately identifying core enterprises and their influence on the loan willingness of upstream and downstream enterprises. In the method of this embodiment, the definition and evaluation method of core enterprises are innovated, and factors such as transaction scale, industry status, and credit rating are comprehensively considered, making the evaluation of the influence of core enterprises more accurate.
[0167] Compared with the existing static model, the method of this embodiment proposes a tracking mechanism for the influence of core enterprises based on dynamic update. By regularly updating information such as the transaction data and financial status of enterprises, it can ensure timely response to market changes and flexibly adjust the parameters of the loan willingness prediction model, thereby improving the practicality and adaptability of the model.
[0168] The method of this embodiment innovatively combines the multi-dimensional feature crossing method and deep learning technology to construct a more complex prediction model. By extracting potential features through feature crossing and combining with a deep learning model (such as MLP), complex non-linear relationships are captured, thereby significantly improving the accuracy of loan willingness prediction.
[0169] In another embodiment, for scenarios with a small amount of data, the data processing flow can be considered simplified, and more simple data cleaning and feature engineering methods can be adopted. For example, using traditional statistical models (such as logistic regression) to replace the deep learning model. Although this alternative solution may sacrifice a certain amount of prediction accuracy, it can still provide a relatively simple solution in the case of limited data.
[0170] Although this embodiment uses a deep learning model (such as MLP) for loan willingness prediction, in some application scenarios, other machine learning models (such as support vector machines, random forests, etc.) can be used. These models may not be able to fully capture complex non-linear relationships, but they can still provide a relatively reasonable prediction effect with a short training time in the case of small data volume or limited computing resources.
[0171] For some application scenarios with low requirements for real-time updates, a rule-based method can be used to judge the influence of core enterprises. For example, the influence of core enterprises can be manually evaluated according to preset rules (such as transaction frequency, industry status, etc.), reducing the dependence on complex models. Although this method may sacrifice a certain degree of dynamic adaptability, it can still provide relatively accurate prediction results in some fixed scenarios.
[0172] The method of this embodiment improves the credit-granting efficiency and risk control ability of financial institutions.
[0173] The above enterprise loan willingness prediction method integrates multi-source data such as supply chain roles, transaction records, and industrial and commercial changes, breaking through the limitation of traditional methods that only rely on financial data; First, obtain the basic information of the enterprise to be predicted and its role in the supply chain, and combine the enterprise's transaction history and industrial and commercial change records to form comprehensive initial data; Then, preprocess and clean these initial data to remove noise and ensure data quality; Then, perform feature engineering based on the cleaned data to extract multi-dimensional features that have an important impact on loan willingness; Subsequently, input these features into the loan willingness prediction model, and comprehensively evaluate the loan demand and potential risks of the enterprise in combination with the dynamic influence of the core enterprise; Finally, output the loan willingness score and the application rate score of the enterprise to achieve high-precision loan prediction, overcoming the deficiencies of traditional static models.
[0174] Figure 5 It is a schematic block diagram of an enterprise loan willingness prediction device 300 provided by an embodiment of the present invention. As Figure 5 shown, corresponding to the above enterprise loan willingness prediction method, the present invention also provides an enterprise loan willingness prediction device 300. The enterprise loan willingness prediction device 300 includes units for executing the above enterprise loan willingness prediction method, and the device can be configured in a server. Specifically, please refer to Figure 5 . The enterprise loan willingness prediction device 300 includes a data acquisition unit 301, a processing unit 302, a prediction unit 303, and an output unit 304.
[0175] The data acquisition unit 301 is used to obtain the basic information of the enterprise to be predicted, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the industrial and commercial information change record data of the enterprise, and the public information of the enterprise to obtain initial data; the processing unit 302 is used to preprocess and clean the initial data, and perform feature engineering processing according to the cleaned data to obtain a processing result; the prediction unit 303 is used to input the processing result into the loan willingness prediction model to predict the enterprise loan willingness to obtain a loan willingness score and an application rate score; the output unit 304 is used to output the loan willingness score and the application rate score.
[0176] In one embodiment, the processing unit 302 includes:
[0177] A cleaning subunit, configured to perform missing value processing, outlier detection, and data normalization processing on the initial data to obtain cleaned data; a combining subunit, configured to construct new features based on the cleaned data and perform feature crossing on the new features to obtain combined features; and a processing subunit, configured to process the combined features by using a dimensionality reduction technique to obtain a processing result.
[0178] In one embodiment, the apparatus further includes a training unit, configured to:
[0179] Obtain enterprise basic information, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the industrial and commercial information change record data of the enterprise, and the public information of the enterprise, and perform preprocessing, cleaning, and feature engineering processing to obtain a sample set; construct a deep neural network under a multi-task learning framework; define a loss function; and use the sample set in combination with the loss function to train the deep neural network by using a gradient descent optimization algorithm to obtain a loan willingness prediction model.
[0180] In one embodiment, the training unit is further configured to:
[0181] Define positive samples and negative samples according to the sample set, and perform processing on the positive samples and negative samples by using undersampling, oversampling, or weighted loss techniques to obtain a processed sample set; and use the processed sample set in combination with the loss function to train the deep neural network by using a gradient descent optimization algorithm to obtain a loan willingness prediction model.
[0182] It should be noted that those skilled in the art can clearly understand that the specific implementation processes of the above enterprise loan willingness prediction apparatus and each unit can refer to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and brevity of description, they will not be elaborated herein.
[0183] The above enterprise loan willingness prediction apparatus 300 can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 6 shown.
[0184] Please refer to Figure 6 , Figure 6 which is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 is a server. Among them, the server can be an independent server or a server cluster composed of multiple servers.
[0185] Refer to Figure 6, the computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. Among them, the memory may include a non-volatile storage medium 503 and an internal memory 504.
[0186] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, can cause the processor 502 to execute an enterprise loan willingness prediction method.
[0187] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0188] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, it can cause the processor 502 to execute an enterprise loan willingness prediction method.
[0189] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device 500 to which the solution of this application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0190] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the following steps:
[0191] Obtain the basic information of the enterprise to be predicted, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the enterprise's industrial and commercial information change record data, and the public information of the enterprise to obtain initial data; preprocess and clean the initial data, and perform feature engineering processing based on the cleaned data to obtain a processing result; input the processing result into a loan willingness prediction model to predict the enterprise's loan willingness to obtain a loan willingness score and an acceptance rate score; output the loan willingness score and the acceptance rate score.
[0192] Among them, the loan willingness prediction model is obtained by acquiring historical data, where the historical data includes the basic information of the enterprise, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the enterprise's industrial and commercial information change record data, and the public information of the enterprise, and forming a sample set through preprocessing, cleaning, and feature engineering processing of the historical data and training a deep neural network under a multi-task learning framework.
[0193] In one embodiment, when the processor 502 implements the step of preprocessing and cleaning the initial data and performing feature engineering on the cleaned data to obtain a processing result, the specific implementation is as follows:
[0194] Perform missing value processing, outlier detection, and data normalization on the initial data to obtain the cleaned data; construct new features based on the cleaned data, and perform feature crossing on the new features to obtain combined features; use dimensionality reduction technology to process the combined features to obtain a processing result.
[0195] In one embodiment, when the processor 502 implements the step that the loan willingness prediction model is obtained by acquiring historical data, where the historical data includes enterprise basic information, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the enterprise's industrial and commercial information change record data, and the enterprise's public information, and forming a sample set through preprocessing, cleaning, and feature engineering on the historical data and training a deep neural network under a multi-task learning framework, the specific implementation is as follows:
[0196] Acquire enterprise basic information, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the enterprise's industrial and commercial information change record data, and the enterprise's public information, and perform preprocessing, cleaning, and feature engineering to obtain a sample set; construct a deep neural network under a multi-task learning framework; define a loss function; use the sample set in combination with the loss function and adopt a gradient descent optimization algorithm to train the deep neural network to obtain a loan willingness prediction model.
[0197] Among them, the deep neural network under the multi-task learning framework includes an input layer, a feature fusion layer, a shared deep learning network, and an output layer, where the shared deep learning network includes multiple fully connected layers, batch normalization, activation functions, and Dropout layers.
[0198] The loss function is a function obtained by summing the loan willingness loss function and the application behavior loss function with weighted coefficients.
[0199] In one embodiment, when the processor 502 implements the step of using the sample set in combination with the loss function and adopting a gradient descent optimization algorithm to train the deep neural network to obtain a loan willingness prediction model, the specific implementation is as follows:
[0200] Define positive samples and negative samples according to the sample set, and perform processing on the positive samples and negative samples using undersampling, oversampling, or weighted loss techniques to obtain a processed sample set; use the processed sample set in combination with the loss function and adopt a gradient descent optimization algorithm to train the deep neural network to obtain a loan willingness prediction model.
[0201] It should be understood that in the embodiments of the present application, the processor 502 may be a central processing unit 302 (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0202] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0203] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the following steps:
[0204] Obtain the basic information of the enterprise to be predicted, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the industrial and commercial information change record data of the enterprise, and the public information of the enterprise to obtain initial data; preprocess and clean the initial data, and perform feature engineering processing according to the cleaned data to obtain a processing result; input the processing result into a loan willingness prediction model to predict the loan willingness of the enterprise to obtain a loan willingness score and an application acceptance rate score; output the loan willingness score and the application acceptance rate score.
[0205] Among them, the loan willingness prediction model is obtained by acquiring historical data, where the historical data includes the basic information of the enterprise, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the industrial and commercial information change record data of the enterprise, and the public information of the enterprise, and preprocessing, cleaning, and feature engineering processing the historical data to form a sample set and training a deep neural network under a multi-task learning framework.
[0206] In one embodiment, when the processor executes the computer program to implement the steps of preprocessing and cleaning the initial data, and performing feature engineering processing based on the cleaned data to obtain a processing result, the specific implementation is as follows:
[0207] Perform missing value processing, outlier detection, and data normalization on the initial data to obtain cleaned data; construct new features based on the cleaned data, and perform feature crossing on the new features to obtain combined features; use dimensionality reduction technology to process the combined features to obtain a processing result.
[0208] In one embodiment, when the processor executes the computer program to implement the step of obtaining the loan willingness prediction model by acquiring enterprise basic information, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the enterprise's industrial and commercial information change record data, and the enterprise's public information, and forming a sample set through preprocessing, cleaning, and feature engineering processing, and then training a deep neural network under a multi-task learning framework, the specific implementation is as follows:
[0209] Acquire enterprise basic information, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the enterprise's industrial and commercial information change record data, and the enterprise's public information, and perform preprocessing, cleaning, and feature engineering processing to obtain a sample set; construct a deep neural network under a multi-task learning framework; define a loss function; use the sample set in combination with the loss function to train the deep neural network using a gradient descent optimization algorithm to obtain a loan willingness prediction model.
[0210] Among them, the deep neural network under the multi-task learning framework includes an input layer, a feature fusion layer, a shared deep learning network, and an output layer. Among them, the shared deep learning network includes multiple fully connected layers, batch normalization, activation functions, and Dropout layers.
[0211] The loss function is a function obtained by summing the loan willingness loss function and the application behavior loss function with weighted coefficients.
[0212] In one embodiment, when the processor executes the computer program to implement the step of using the sample set in combination with the loss function to train the deep neural network using a gradient descent optimization algorithm to obtain a loan willingness prediction model, the specific implementation is as follows:
[0213] Define positive samples and negative samples according to the sample set, and perform processing on the positive samples and negative samples using undersampling, oversampling, or weighted loss techniques to obtain a processed sample set; use the processed sample set in combination with the loss function to train the deep neural network using a gradient descent optimization algorithm to obtain a loan willingness prediction model.
[0214] The storage medium may be a variety of computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc, etc., which can store program codes.
[0215] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0216] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0217] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into a processing unit 302, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0218] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.
[0219] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. The method for predicting corporate loan willingness is characterized by: include: Obtain the basic information of the enterprise to be predicted, the role of the enterprise in the supply chain, the transaction record of the enterprise within the set time period, the record data of the change of the enterprise's industrial and commercial information, and the public information of the enterprise to obtain the initial data; Preprocessing and cleaning the initial data, and performing feature engineering processing on the cleaned data to obtain processing results; Input the processing results into a loan willingness prediction model to predict the enterprise's loan willingness, so as to obtain a loan willingness score and an application acceptance rate score; The loan willingness score and the application acceptance rate score are output.
2. The method for predicting corporate loan willingness according to claim 1, characterized in that: The preprocessing and cleaning of the initial data, and performing feature engineering processing on the cleaned data to obtain a processing result, include: Performing missing value processing, outlier detection and data normalization processing on the initial data to obtain cleaned data; Constructing new features according to the cleaned data, and processing the new features by feature cross-processing to obtain combined features; The combined features are processed using a dimension reduction technique to obtain a processing result.
3. The method for predicting corporate loan willingness according to claim 1, characterized in that: The loan willingness prediction model is obtained by obtaining historical data, wherein the historical data includes basic information of the enterprise, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the record data of changes in the business information of the enterprise, and the public information of the enterprise, and preprocessing, cleaning and feature engineering the historical data to form a sample set to train a deep neural network under a multi-task learning framework.
4. The method for predicting corporate loan willingness according to claim 3, characterized in that: The loan willingness prediction model is obtained by obtaining basic information of the enterprise, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the record data of the change of the enterprise's industrial and commercial information, and the public information of the enterprise, and then performing preprocessing, cleaning, and feature engineering to form a sample set to train a deep neural network under a multi-task learning framework, including: Obtain basic information about the enterprise, the role of the enterprise in the supply chain, the transaction records of the enterprise within a set time period, the record data of changes in the enterprise's industrial and commercial information, and the public information of the enterprise, and perform preprocessing, cleaning, and feature engineering to obtain a sample set; Construct a deep neural network under a multi-task learning framework; Define the loss function; The deep neural network is trained using the sample set in combination with the loss function using a gradient descent optimization algorithm to obtain a loan willingness prediction model.
5. The method for predicting corporate loan willingness according to claim 4, characterized in that: The deep neural network under the multi-task learning framework includes an input layer, a feature fusion layer, a shared deep learning network and an output layer, wherein the shared deep learning network contains multiple layers of fully connected layers, batch normalization, activation functions, and Dropout layers.
6. The method for predicting corporate loan willingness according to claim 4, characterized in that: The loss function includes a function obtained by summing a loan willingness loss function and an application behavior loss function using weighted coefficients.
7. The method for predicting corporate loan willingness according to claim 4, characterized in that: The step of using the sample set in combination with the loss function to train the deep neural network using a gradient descent optimization algorithm to obtain a loan willingness prediction model includes: Defining positive samples and negative samples according to the sample set, and processing the positive samples and negative samples by using undersampling, oversampling or weighted loss technology to obtain a processed sample set; The processed sample set is combined with the loss function to train the deep neural network using a gradient descent optimization algorithm to obtain a loan willingness prediction model.
8. The enterprise loan willingness prediction device is characterized by: include: The data acquisition unit is used to obtain the basic information of the enterprise to be predicted, the role of the enterprise in the supply chain, the transaction record of the enterprise within a set time period, the record data of the change of the enterprise's industrial and commercial information, and the public information of the enterprise to obtain the initial data; A processing unit, used to preprocess and clean the initial data, and perform feature engineering processing based on the cleaned data to obtain a processing result; A prediction unit, used for inputting the processing result into a loan willingness prediction model to predict the enterprise's loan willingness, so as to obtain a loan willingness score and a submission rate score; An output unit is used to output the loan willingness score and the application acceptance rate score.
9. A computer device, characterized in that: The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.