Credit risk determination method and device, electronic equipment and storage medium

By training the credit student model and utilizing the legal person credit data and labels of the credit teacher model, the problem of insufficient corporate credit information is solved, accurate risk assessment of enterprises without corporate credit is achieved, and the coverage of credit assessment is expanded.

CN120634715AInactive Publication Date: 2025-09-12ZHEJIANG E COMMERCE BANK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511127119.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Many companies, especially start-ups and small and medium-sized enterprises, lack available credit information and are unable to conduct credit risk assessments, resulting in an inability to obtain sufficient credit information and risk assessments.

Method used

By training the credit student model, using legal person credit data to assess corporate credit risk, and combining the corporate credit risk labels provided by the credit teacher model, a credit student model is constructed, and distillation learning is used to improve assessment accuracy.

Benefits of technology

It has achieved the accuracy of credit risk assessment for enterprises without corporate credit information, expanded the coverage of credit assessment, and enabled enterprises with personal credit of legal persons but no corporate credit to share the corporate credit value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634715A_ABST
    Figure CN120634715A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a credit risk determination method and device, electronic equipment and a storage medium. The method comprises the steps of performing loss calculation on a first sample credit risk result generated by an initial credit student model based on sample legal person credit data through a first preset loss function, a first sample credit risk label and a second sample credit risk result, the second sample credit risk result is obtained by the credit teacher model based on the sample enterprise credit data of the first sample enterprise; the initial credit student model is iteratively trained based on the loss value, a trained credit student model is obtained, and the trained credit student model is used for generating a target credit risk result of the target enterprise based on the legal person credit data of the target enterprise.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a credit risk determination method, device, electronic device, and storage medium. Background Art

[0002] Corporate credit reporting plays a vital role in the modern economy, particularly in improving the efficiency of credit transactions and reducing information asymmetry and transaction costs. First, by establishing a robust credit assessment mechanism, corporate credit reporting helps reduce information asymmetry, enabling banks and financial institutions to obtain more comprehensive and accurate business data, thereby effectively reducing the risk of loan defaults. This transparent credit reporting system helps provide businesses, especially small and medium-sized enterprises, with more financing opportunities, thereby promoting their healthy development. Furthermore, corporate credit reporting promotes inter-business credit transactions and, through voluntary disclosure of business information, establishes a risk-constraining mechanism, thereby enhancing the market's credit environment and the stability of the overall financial system.

[0003] However, despite the obvious advantages of the corporate credit reporting system, current technology applications do not ensure that every enterprise has full access to complete credit information. Due to the limitations of information collection, processing, and sharing among various credit reporting agencies and platforms, many enterprises, especially start-ups and small and medium-sized enterprises, often lack available credit information, making it difficult to conduct risk assessments. Summary of the Invention

[0004] The main purpose of this specification is to provide a credit risk determination method, device, electronic device, and storage medium, aiming to achieve accurate assessment of corporate credit risk. The technical solution is as follows: In a first aspect, embodiments of this specification provide a credit risk determination method, including: Obtaining a first sample data set; the first sample data set includes sample legal person credit data and sample credit risk labels of each first sample enterprise; Inputting the sample legal person credit data into an initial credit student model to obtain a first sample credit risk result output by the initial credit student model; Determining, based on a first preset loss function, a first loss value between the first sample credit risk result and the second sample credit risk result, and a second loss value between the first sample credit risk result and the first sample credit risk label; the second sample credit risk result is obtained by the credit teacher model based on the sample enterprise credit data of the first sample enterprise; The initial credit student model is iteratively trained based on the first loss value and the second loss value until the first loss value and the second loss value meet the convergence condition, thereby obtaining a trained credit student model; the credit student model is used to generate the target credit risk result of the target enterprise based on the legal person credit data of the target enterprise.

[0005] In a second aspect, an embodiment of this specification provides a risk determination device, including: A sample acquisition unit is configured to acquire a first sample data set, wherein the first sample data set includes sample legal person credit data and sample credit risk labels of each first sample enterprise; a prediction unit, configured to input the sample legal person credit data into an initial credit student model to obtain a first sample credit risk result output by the initial credit student model; a loss calculation unit, configured to determine, based on a first preset loss function, a first loss value between the first sample credit risk result and the second sample credit risk result, and a second loss value between the first sample credit risk result and the first sample credit risk label; the second sample credit risk result being obtained by a credit teacher model based on the sample enterprise credit data of the first sample enterprise; A training unit is used to iteratively train the initial credit student model based on the first loss value and the second loss value until the first loss value and the second loss value meet the convergence condition, thereby obtaining a trained credit student model; the credit student model is used to generate the target credit risk result of the target enterprise based on the legal person credit data of the target enterprise.

[0006] In a third aspect, an embodiment of this specification provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the above method when executed by the processor.

[0007] In a fourth aspect, an embodiment of this specification provides a storage medium having a computer program stored thereon, and the computer program implements the steps of the above method when executed by a processor.

[0008] In a fifth aspect, an embodiment of this specification provides a computer program product, comprising: a computer program, which, when executed by a processor of an electronic device, enables the processor to at least implement the method described in the first aspect.

[0009] In the embodiments of this specification, a credit student model is trained to assess corporate credit risk based on corporate legal person credit data. This allows companies with personal legal person credit but no corporate credit to share the value of corporate credit. The credit student model is trained based on corporate credit risk labels provided by a credit teacher model. These corporate credit risk labels are derived by the credit teacher model based on corporate credit data. Leveraging the high similarity between corporate and personal credit data, the credit student model is constructed. The corporate credit learned by the credit teacher model guides the credit student model's learning of personal credit. This distillation-based credit student model improves the accuracy of credit risk assessments for companies without corporate credit information. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0011] Figure 1 This is an example diagram of a credit risk determination method provided in the embodiments of this specification; Figure 2 This is an example diagram of a credit risk determination method provided in the embodiments of this specification; Figure 3 This is a flow chart of a credit risk determination method provided in an embodiment of this specification; Figure 4 This is a flow chart of a credit risk determination method provided in an embodiment of this specification; Figure 5 This is a flow chart of a credit risk determination method provided in an embodiment of this specification; Figure 6 This is an example diagram of a credit risk determination method provided in the embodiments of this specification; Figure 7 This is a flow chart of a credit risk determination method provided in an embodiment of this specification; Figure 8 This is a flow chart of a credit risk determination method provided in an embodiment of this specification; Figure 9 This is a schematic diagram of the structure of a credit risk determination device provided in an embodiment of this specification; Figure 10 This is a schematic diagram of the structure of a credit risk determination device provided in an embodiment of this specification; Figure 11This is a structural diagram of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0012] The following will be combined with the drawings in the embodiments of this specification to clearly and completely describe the technical solutions in the embodiments of this specification. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.

[0013] See Figure 1 , Figure 1 An example schematic diagram of a credit risk determination method is provided for the embodiments of this specification. In related technologies, when a risk assessment of an enterprise is required, it is necessary to determine it based on the enterprise's credit data. However, not all enterprises have credit data (such as enterprise credit reports), and risk assessment cannot be performed for such enterprises.

[0014] Based on the above problems, this specification provides a credit risk determination method. Figure 2 , provides an example schematic diagram of a credit risk determination method for the embodiments of this specification. Corporate credit reporting and personal credit reporting are indeed highly similar in many aspects, mainly because they have similarities in basic principles, assessment methods, and purposes. The common point between the two is that both rely on historical behavioral records for credit scoring, and both are intended to help creditors (banks, investors, etc.) assess the borrower's repayment ability and default risk. Since a corporate legal person refers to an individual who exercises rights and assumes obligations on behalf of the company in accordance with the law, usually the company's legal representative, chairman, or general manager, the legal person's personal credit data and corporate credit data are both of great reference value for assessing corporate risks. By obtaining the target company's corporate credit data and conducting corporate risk estimation based on the corporate credit data, it is possible to solve the risk assessment of companies without corporate credit, thereby expanding the customer base and improving the coverage of services.

[0015] Specifically, embodiments of this specification also provide a credit risk determination method for training a credit student model (risk assessment model) for credit risk assessment. This method obtains sample legal person credit data and sample credit risk labels for sample enterprises, inputs the sample legal person credit data into an initial credit student model, and outputs a first sample credit risk result from the initial credit student model. A second sample credit risk result is obtained from a credit teacher model. The second sample credit risk result is obtained by the credit teacher model based on the sample corporate credit data of the sample enterprises. Based on a first preset loss function, a first loss value is determined between the first sample credit risk result and the second sample credit risk result, as well as a second loss value between the first sample credit risk result and the sample credit risk label. The initial credit student model is iteratively trained based on the first and second loss values ​​until the first and second loss values ​​meet convergence conditions, thereby obtaining a trained credit student model. After the credit teacher model is trained on a customer base with corporate credit data, the corporate credit risk labels generated by the credit teacher model are used to learn the credit student model. Specifically, the corporate credit learned by the credit teacher model guides the personal credit learning of the credit student model, enabling the credit student model to accurately predict corporate risk based on the legal person's personal credit data.

[0016] It should be noted that for enterprises with corporate credit data, corporate risk prediction can be carried out through the credit teacher model.

[0017] It is understandable that the credit risk determination device provided in the embodiments of this specification may be a terminal device such as a mobile phone, computer, tablet computer or vehicle-mounted device, or may be a module in the terminal device for implementing the credit risk determination method.

[0018] The credit risk determination method provided in this specification is described in detail below with reference to specific embodiments.

[0019] See Figure 3 , provides a flow chart of a credit risk determination method according to the embodiment of this specification. Figure 3 As shown, the method of the embodiment of this specification may include the following steps S102-S104.

[0020] S102, obtaining the legal person credit data of the target enterprise; In one embodiment of this specification, the target enterprise is an enterprise requiring risk assessment. Legal person credit data refers to various information related to the personal credit behavior of the legal person of the target enterprise. Legal person credit data may include basic information, credit accounts (such as credit limits, number of non-commercial accounts, etc.), credit behavior (such as loan application records, credit limit utilization rate, revolving overdraft ratio), debt status (such as overdue amounts), repayment records (such as repayment information), etc. For example, a big data platform can be used to collect information publicly released by credit institutions, judicial institutions, public institutions, and private enterprises to collect massive amounts of credit and non-credit transaction information, civil and criminal case judgment information, public announcement information, and other massive amounts of information. Personal information can then be extracted and analyzed to obtain legal person credit data.

[0021] S104: Using a pre-trained credit student model, generate a target credit risk result for the target enterprise based on the legal person credit data.

[0022] In one embodiment of the present specification, the pre-trained credit student model is trained based on the enterprise credit risk label provided by the credit teacher model, and the enterprise credit risk label is obtained by the credit teacher model based on the enterprise credit data. By inputting the legal person credit data into the pre-trained credit student model, the target credit risk result of the target enterprise output by the credit student model can be obtained. Credit risk refers to the risk that an enterprise may face when fulfilling its financial obligations (such as repaying debts, paying interest, etc.). The target credit risk result can be whether the target enterprise has credit risk, or more specifically, it can be the risk level of the target enterprise. If the target credit risk result is that there is credit risk, then the target enterprise has a weak ability to fulfill its financial obligations and has a higher potential risk of not being able to fulfill its financial responsibilities. Furthermore, financial institutions can provide personalized financial products for target enterprises based on different credit risk results.

[0023] The Credit Teacher model is trained on a sample of companies with valid corporate credit data. It can learn and identify the correlation between corporate credit data and corporate risk. Based on this corporate credit data, the Credit Teacher model can predict a company's credit risk outcomes (such as risk probability and risk level). After training, the corporate credit learned by the Credit Teacher model is used to guide the Credit Student model's individual credit learning. Since legal entity credit data is generally readily available, for companies with corporate credit data, after the Credit Teacher model predicts their risk based on this corporate credit data, the Credit Student model can then further obtain their legal entity credit data and use this data to predict their risk. For the same company, the risk predictions from the Credit Student model based on the legal entity credit data are identical or similar to those from the Credit Teacher model based on the corporate credit data.

[0024] Optionally, in one embodiment of this specification, if the target enterprise's corporate credit data is available, a pre-trained credit teacher model is used to generate a target credit risk result for the target enterprise based on the corporate credit data. It will be appreciated that corporate credit data can more accurately reflect the enterprise's risk. If the target enterprise's corporate credit data is available, a risk prediction for the target enterprise is performed based on the corporate credit data and the pre-trained credit teacher model. Corporate credit data may include the enterprise's financial information, operating conditions, credit history, etc.

[0025] In the embodiments of this specification, the target enterprise's legal person credit data is obtained, and a pre-trained credit student model is used to generate a target credit risk result for the target enterprise based on the legal person credit data. The legal person's personal credit is used to perform risk assessments for enterprises without corporate credit, allowing enterprises with legal person personal credit but no corporate credit to share the value of corporate credit. The credit student model is trained based on the corporate credit risk labels provided by the credit teacher model. The credit teacher model is capable of performing corporate risk assessments based on corporate credit data. The corporate credit risk labels are obtained by the credit teacher model based on corporate credit data. By leveraging the high similarity between corporate and personal credit data, a credit teacher model and a credit student model are constructed, using corporate credit to guide the learning of personal credit. The use of a distillation-based learning model can improve the accuracy of risk assessments for enterprises without corporate credit information.

[0026] See Figure 4 , provides a flow chart of a credit risk determination method according to the embodiment of this specification. Figure 4 As shown, the method of the embodiment of this specification may include the following steps S202-S208.

[0027] S202, obtaining a first sample data set; In one embodiment of the present specification, a first sample dataset includes multiple sample enterprises, as well as sample legal person credit data and a first sample credit risk label for each sample enterprise. Enterprises with personal credit data are obtained as sample enterprises, and whether the sample enterprises are at risk is used as the first sample credit risk label. The first sample credit risk label can be a hard label indicating whether the sample enterprise is at risk (e.g., represented by 1) or not (e.g., represented by 0). The sample legal person credit data is the legal person credit data of the sample enterprise over a period of time (e.g., the past year). The sample legal person credit data may include basic information about the legal person of the sample enterprise, credit accounts (e.g., credit limit, number of non-commercial accounts, etc.), credit behavior (e.g., loan application records, credit limit utilization rate, revolving overdraft ratio), debt status (e.g., overdue amount), repayment records (e.g., repayment information), etc.

[0028] S204, inputting the sample legal person credit data into an initial credit student model to obtain a first sample credit risk result output by the initial credit student model; In one embodiment of this specification, the initial credit student model is obtained by preparing and configuring the model according to training requirements. Exemplarily, the initial credit student model configuration process may include: selecting an appropriate model structure (such as logistic regression, decision tree, neural network, etc.), determining model parameters (such as the number of neural network layers and the number of nodes per layer), then selecting appropriate hyperparameters (such as learning rate, batch size, optimization algorithm, etc.). Based on the nature of the task (regression, classification, etc.), selecting an appropriate loss function (such as mean squared error (MSE) or cross entropy), and selecting an appropriate optimizer (such as gradient descent, Adam, SGD, etc.) based on the task requirements, and configuring parameters such as the learning rate. Furthermore, for neural network models, the weights of the network layers typically need to be initialized. The obtained sample legal entity credit data is then input into the initial credit student model, which outputs a first sample credit risk result. This first sample credit risk result can be the probability of the enterprise being at risk, and can range from 0 to 1.

[0029] For example, the model structure of the initial credit student model can be selected as MLP (Multilayer Perceptron). MLP is a feedforward artificial neural network model consisting of an input layer, a hidden layer, and an output layer.

[0030] S206: Determine, based on a first preset loss function, a first loss value between the first sample credit risk result and the second sample credit risk result, and a second loss value between the first sample credit risk result and the first sample credit risk label; In one embodiment of the present specification, the second sample credit risk result is obtained by the trained credit teacher model based on the sample enterprise credit data of the sample enterprise. The sample enterprise credit data includes the financial information, operating conditions, credit record, etc. of the sample enterprise. The first sample credit risk result generated by the initial credit student model is compared with the second sample credit risk result predicted by the credit teacher model. The credit teacher model guides the credit student model to be consistent with it to alleviate the impact of changes in the enterprise credit magnitude on the model, while allowing enterprises with legal person personal credit but no corporate credit to share the corporate credit value. In addition, the first sample credit risk result can also be compared with the first sample credit risk label of the sample enterprise. The first sample credit risk label is the real result of whether the sample enterprise is risky. By calculating the second loss value, the result of the credit student model can be close to the real result. Exemplarily, the first preset loss function can be a cross-entropy loss (CE Loss).

[0031] Optionally, in one embodiment of the present specification, if the sample enterprise credit data of the sample enterprise can be obtained, then the second sample credit risk result output by the credit teacher model is obtained, and the first loss value between the first sample credit risk result and the second sample credit risk result and the second loss value between the first sample credit risk result and the first sample credit risk label are determined based on the first preset loss function; if the sample enterprise credit data of the sample enterprise cannot be obtained, then the second loss value between the first sample credit risk result and the first sample credit risk label is confirmed based on a third loss function. It is understandable that some of the sample enterprises in the first sample data set may be able to obtain sample enterprise credit data, while some may not be able to obtain sample enterprise credit data. For the sample enterprises for which the credit data of the sample enterprises cannot be obtained, it is only necessary to calculate the second loss value between the first sample credit risk result and the sample credit risk label.

[0032] S208, iteratively training the initial credit student model based on the first loss value and the second loss value until the first loss value and the second loss value meet a convergence condition, thereby obtaining a trained credit student model.

[0033] In one embodiment of the present specification, a first loss value and a second loss value are used to measure the model's prediction error. These loss values ​​are then used to update the model parameters of the initial credit student model via a backpropagation algorithm. After each iteration, the first and second loss values ​​are checked to see if they meet a preset convergence condition (e.g., the error change is within a certain threshold or a certain number of training rounds has been reached). When the convergence condition is met, training stops, ultimately resulting in a trained credit student model. For example, the first and second loss values ​​can be weighted and summed to obtain a total loss. If the total loss meets the convergence condition, the credit student model training is complete.

[0034] In an embodiment of this specification, a first sample data set is obtained, sample legal person credit data is input into an initial credit student model, and a first sample credit risk result is output by the initial credit student model. Based on a first preset loss function, a first loss value is determined between the first sample credit risk result and a second sample credit risk result, as well as a second loss value between the first sample credit risk result and a sample credit risk label. The second sample credit risk result is obtained by the credit teacher model based on the sample enterprise credit data of the sample enterprise. The initial credit student model is iteratively trained based on the first and second loss values ​​until the first and second loss values ​​meet convergence conditions, thereby obtaining a trained credit student model. This risk determination model training method allows the credit student model trained based on personal credit data to maximize the value of corporate credit data when the volume of personal credit data queries is very high. Specifically, although some corporate legal persons do not have corporate credit data, they may have personal credit data. After learning information about their personal credit through the credit student model, the value of corporate credit data can still be shared and utilized to some extent. This is because, through distillation, the model of personal credit data can indirectly reflect some characteristics and patterns of corporate credit, thereby providing a similar basis for judging corporate credit for enterprises without corporate credit.

[0035] See Figure 5 , provides a flow chart of a credit risk determination method according to the embodiment of this specification. Figure 5 As shown, the method of the embodiment of this specification may include the following steps S302-S310.

[0036] S302, obtaining a second sample data set; In one embodiment of this specification, a method for training a credit teacher model is also provided. A second sample dataset is obtained for training the credit teacher model. The second sample dataset includes sample enterprise credit data and second sample credit risk labels for each sample enterprise. When training the credit teacher model, the sample enterprises in the second sample dataset may be the same as or different from the sample enterprises in the first sample dataset. It is understood that because the first sample dataset contains sample enterprises with only sample legal person credit data, the sample enterprises contained in the first and second sample datasets will not be exactly the same.

[0037] S304: Input the sample enterprise credit data into the initial credit teacher model to obtain a third sample credit risk result output by the initial credit teacher model; In one embodiment of the present specification, the model structures of the initial credit teacher model and the initial credit student model are generally consistent. Therefore, MLP can be selected as the model structure of the initial credit teacher model. The initial model parameters (such as the number of layers of the neural network, the number of nodes in each layer, etc.) are set accordingly. Afterwards, appropriate hyperparameters (such as learning rate, batch size, optimization algorithm, etc.) are selected. The sample legal person credit data is input into the initial credit teacher model to obtain a third sample credit risk result output by the model. The third sample credit risk result can be the probability that the enterprise is at risk, and the value range can be between 0 and 1.

[0038] S306, determining a third loss value between the third sample credit risk result and the second sample credit risk label based on a second preset loss function; In one embodiment of the present specification, a third loss value is obtained by determining the difference between the third sample credit risk result and the second sample credit risk label according to a second preset loss function. The third loss value is used to measure the deviation between the result of the initial credit teacher model and the true label. Exemplarily, the second preset loss function can be a cross-entropy loss function.

[0039] S308, iteratively training the initial credit teacher model based on the third loss value until the third loss value meets a convergence condition, thereby obtaining a trained credit teacher model; In one embodiment of this specification, the third loss value measures the error between the model's prediction and the true value. An optimization algorithm (such as gradient descent) is then used to update the model's parameters through backpropagation, calculating and optimizing the third loss value with each iteration. The training process continues until the third loss value meets a preset convergence condition (for example, the change in the loss value is sufficiently small or the maximum number of iterations is reached). When the convergence condition is met, training is terminated, ultimately resulting in a trained credit teacher model.

[0040] S310: Input the sample enterprise credit data of the first sample enterprise into the credit teacher model to obtain a second sample credit risk result output by the credit teacher model.

[0041] In one embodiment of the present specification, after training the credit teacher model, the obtained sample enterprise credit data of the first sample enterprise is input into the credit teacher model during the training of the credit student model. This generates a second sample credit risk result for the credit teacher model. This second sample credit risk result is then used to guide the credit student model in generating a result distribution based on the sample legal person credit data of the first sample enterprise. This second risk result can be the probability that the enterprise is at risk, and can have a value between 0 and 1.

[0042] See Figure 6 , Figure 6This is an example diagram of a credit risk determination method provided in an embodiment of this specification. In the embodiment of this specification, when training the credit student model, the first sample data set can include two types of sample enterprises, one is a sample with enterprise credit, and the other is a sample without enterprise credit. For the sample with enterprise credit, its sample data includes sample legal person credit data and sample enterprise credit data. The sample legal person credit data is input into the credit student model to obtain the first enterprise result. In addition, the sample enterprise credit data is input into the trained credit teacher model to obtain the second enterprise risk result. The first loss value is determined based on the first enterprise risk result and the second enterprise risk result. The second loss value is determined based on the first enterprise risk result and the first sample credit risk label, and then the credit student model is iteratively trained based on the first loss value and the second loss value. That is, in the sample with enterprise credit characteristics, the hard label (first sample credit risk label) and the soft label (second enterprise risk result) are learned simultaneously. For samples without corporate credit characteristics, the sample legal person credit data of these companies is input into the credit student model to obtain the first corporate risk result ' (where "'" is used to distinguish the results obtained from different types of sample companies). The second loss value ' is determined based on the first corporate risk result ' and the first sample credit risk label. Simultaneously, the credit student model is iteratively trained based on the second loss value '. In other words, the hard label is directly learned in samples without corporate credit characteristics. Because personal credit data is larger than corporate credit data, unlike traditional distillation learning methods, the credit student model uses a larger "student" model to imitate a smaller "teacher" model, achieving better results.

[0043] In an embodiment of this specification, a second sample data set is obtained, sample enterprise credit data from the second sample data set is input into an initial credit teacher model, and a third sample credit risk result output by the initial credit teacher model is obtained. Based on a second preset loss function, a third loss value is determined between the third sample credit risk result and the second sample credit risk label. The initial credit teacher model is iteratively trained based on the third loss value until the third loss value meets a convergence condition, thereby obtaining a trained credit teacher model. The credit teacher model learns enterprise credit data, enabling it to generate accurate risk result soft labels based on the sample enterprise credit data for learning by the credit student model.

[0044] See Figure 7 , provides a flow chart of a credit risk determination method according to the embodiment of this specification. Figure 7 As shown, the method of the embodiment of this specification may include the following steps S402-S404.

[0045] S402, inputting the sample legal person credit data into the feature processing module to obtain a sample individual feature vector corresponding to the sample legal person credit data; In one embodiment of the present specification, the initial credit student model includes a feature processing module and a prediction module. The feature processing module is used to process the sample legal person credit data into corresponding sample personal feature vectors, and make the sample personal feature vectors adapt to the input requirements of the model. Exemplarily, the feature processing module can extract useful features from the sample legal person credit data. Common methods include: (1) Numerical feature extraction: directly extracting numerical features (such as age, income, credit score, etc.) from the data. (2) Categorical feature processing: discrete features such as sex and occupation can be converted into numerical representations using methods such as One-Hot encoding, Label encoding, and Embedding. (3) Time feature processing: For example, a user's loan history or recent consumption records, features such as "day difference" and "time of the last loan" can be extracted based on the timestamp.

[0046] Optionally, feature scaling is often performed to ensure that different features can be compared and processed at the same scale. A common approach is to scale feature values ​​to a specific interval (such as [0, 1]). The number of features can then be reduced by selecting the most important features or using dimensionality reduction methods to reduce computational complexity and avoid overfitting. For certain high-dimensional, sparse features (such as user IDs and product IDs), embedding methods can be used to convert them into low-dimensional, dense vectors.

[0047] S404: Input the sample individual feature vector into a prediction module to obtain a sample credit risk probability of the sample enterprise, and use the sample credit risk probability as the first sample credit risk result.

[0048] In one embodiment of the present specification, the sample individual feature vector is input into a prediction module, which outputs a sample credit risk probability of the sample enterprise and uses the sample credit risk probability as the first sample credit risk result. Exemplarily, the prediction module may be a neural network.

[0049] Optionally, in one embodiment of the present specification, the feature processing module includes a preprocessing module and a vector conversion module. Inputting the sample legal person credit data into the feature processing module to obtain a sample individual feature vector corresponding to the sample legal person credit data includes the following steps S4022-S4026: S4022: Input the sample legal person credit data into the pre-processing module, perform binning processing on the sample legal person credit data based on the pre-processing module, and obtain serialization features corresponding to the sample legal person credit data; and / or, S4024: Input the sample legal person credit data into the preprocessing module, determine a branch path of the sample individual credit data in a decision tree based on the preprocessing module, and generate a tree coding feature corresponding to the sample legal person credit data based on the branch path; In one embodiment of this specification, sample legal entity credit data may undergo one or two preprocessing steps. Specifically, this includes at least one of binning, serialization, and tree encoding to obtain serialized features, tree-encoded features, or both. During binning, continuous features are divided into several intervals (bins), and the values ​​within each interval are uniformly labeled as a category (or number). For example, the age feature may be divided into intervals such as [0-20), [20-40), [40-60), and [60,∞). The original continuous values ​​are then converted to corresponding interval labels. Because continuous features are not well learned in neural learning, binning can convert continuous features into discrete features, thereby improving learning effectiveness. Additionally, some features are preprocessed using a Light Gradient Boosting Machine (LGBM), an efficient and scalable machine learning algorithm based on gradient boosted decision trees (GBDT). Each tree is constructed during training based on a feature partitioning rule (by partitioning the dataset based on different feature values). The core idea of ​​tree encoding is to create a feature representation based on the structure of these trees, which not only preserves the partitioning information of the tree but also generates new features in this way, thereby improving the expressiveness and performance of the model.

[0050] S4026: Perform embedding processing on the serialization feature and / or the tree encoding feature based on the vector conversion module to obtain a sample legal person credit data vector corresponding to the sample legal person credit data.

[0051] In one embodiment of this specification, these features need to be embedded, for example using Embedding Lookup. Embedding Lookup is a technique commonly used in natural language processing (NLP) and deep learning. It aims to map discrete words or symbols (such as words, characters, user IDs, etc.) into a continuous vector space, allowing computers to more efficiently process and understand this discrete data. Specifically, Embedding Lookup is implemented using a pre-trained embedding matrix.

[0052] See Figure 8 , Figure 8This is a flowchart of a credit risk determination method provided in an embodiment of this specification. Taking an MLP prediction module as an example, the training process is explained. In both the teacher and student credit models, the preprocessing modules include a binning component and an lgbm (LightGBM) component. lgbm is a gradient boosting framework based on a decision tree algorithm. For samples with corporate credit data (corporate credit feature), the sample legal person credit data (personal credit feature) is input into the preprocessing module, and the serialized features are obtained based on binning and serialization processing; the branch path of the sample personal credit data in the decision tree is determined based on the lgbm part of the preprocessing module, and the tree encoding (encoding) features corresponding to the sample legal person credit data are generated based on the branch path. The sample corporate credit data (corporate credit feature) is input into the preprocessing module, and the serialized features are obtained based on binning and serialization processing; the branch path of the sample personal credit data in the decision tree is determined based on the lgbm part of the preprocessing module, and the tree encoding (encoding) features corresponding to the sample legal person credit data are generated based on the branch path, and then input into the MLP in each model to obtain Logits refers to the raw output value that has not been normalized (such as the output of the fully connected layer) These values ​​are converted into probability distributions by applying an activation function, for example, the activation function selected is the sigmoid function. Generate soft labels . By crediting the student model based on the original output Generate soft probability results The difference between the generated credit risk probability of the credit student model and the label of the credit teacher model is calculated by cross-entropy loss (CE Loss) f( , ), and calculate the credit risk probability of the credit student model With sample credit risk label The difference between f( , ). The superscript “ˇ” is used to distinguish between two different types of results: soft label (specific probability) and hard label (yes / no). For samples without corporate credit data, the sample legal person credit data is directly input into the student model to obtain the result , calculate the credit risk probability generated by the credit student model through CE Loss With sample credit risk label The loss between , ). No corporate credit sample ( ) is the loss of , where i is the sample number, N is the training sample set, and there are enterprise credit samples ( ) is the loss of + (1- ) .in, is the weight.

[0053] In the embodiments of this specification, sample legal person credit data is input into a feature processing module to obtain a sample individual feature vector corresponding to the sample legal person credit data. This sample individual feature vector is then input into a prediction module to obtain a sample credit risk probability for the sample enterprise, which is then used as the first sample credit risk result. The stability of model training is improved by preprocessing the input data, including binning, serialization, and tree encoding.

[0054] The following will be combined with the Figure 9 , the credit risk determination device provided in the embodiment of this specification is introduced in detail. Figure 9 The credit risk determination device in the Figure 2-Figure 8 For the convenience of explanation, only the part related to the embodiment of this specification is shown. For the specific technical details not disclosed, please refer to this specification. Figure 2-Figure 8 The embodiment shown.

[0055] See Figure 9 , which shows a schematic diagram of the structure of a credit risk determination device provided in one embodiment of this specification. The credit risk determination device can be implemented as all or part of a device through software, hardware, or a combination of both. The device 1 includes an acquisition unit 11 and a generation unit 12.

[0056] An acquisition unit 11 is used to acquire the legal person credit data of the target enterprise; Generation unit 12 is used to use a pre-trained credit student model to generate the target credit risk result of the target enterprise based on the legal person credit data; the credit student model is trained based on the enterprise credit risk label provided by the credit teacher model, and the enterprise credit risk label is obtained by the credit teacher model based on the enterprise credit data.

[0057] Optionally, the generating unit 12 is further configured to, if the corporate credit data of the target enterprise is acquired, use a pre-trained credit teacher model to generate a target credit risk result for the target enterprise based on the corporate credit data.

[0058] See Figure 10, which shows a schematic diagram of the structure of a credit risk determination device provided in one embodiment of this specification. The credit risk determination device can be implemented as all or part of a device through software, hardware, or a combination of both. The device 2 includes a sample acquisition unit 21, a prediction unit 22, a loss calculation unit 23, and a training unit 24.

[0059] The sample acquisition unit 21 is configured to acquire a first sample data set; the first sample data set includes sample legal person credit data and sample credit risk labels of each first sample enterprise; A prediction unit 22 is configured to input the sample legal person credit data into an initial credit student model to obtain a first sample credit risk result output by the initial credit student model; A loss calculation unit 23 is configured to determine, based on a first preset loss function, a first loss value between the first sample credit risk result and the second sample credit risk result, and a second loss value between the first sample credit risk result and the first sample credit risk label; the second sample credit risk result is obtained by the credit teacher model based on the sample enterprise credit data of the first sample enterprise; The training unit 24 is used to iteratively train the initial credit student model based on the first loss value and the second loss value until the first loss value and the second loss value meet a convergence condition, thereby obtaining a trained credit student model.

[0060] Optionally, the sample acquisition unit 21 is further configured to acquire a second sample data set; the second sample data set includes sample enterprise credit data and second sample credit risk labels of each second sample enterprise; Optionally, the prediction unit 22 is specifically configured to input the sample enterprise credit data into an initial credit teacher model to obtain a third sample credit risk result output by the initial credit teacher model; Optionally, the loss calculation unit 23 is further configured to determine a third loss value between the third sample credit risk result and the second sample credit risk label based on a second preset loss function; Optionally, the training unit 24 is further configured to iteratively train the initial credit teacher model based on the third loss value until the third loss value satisfies a convergence condition, thereby obtaining a trained credit teacher model.

[0061] Optionally, the prediction unit 22 is further configured to input the sample enterprise credit data of the first sample enterprise into the credit teacher model to obtain a second sample credit risk result output by the credit teacher model.

[0062] Optionally, the initial credit student model includes a feature processing module and a prediction module, wherein the prediction unit 22 is specifically configured to input the sample legal person credit data into the feature processing module to obtain a sample individual feature vector corresponding to the sample legal person credit data; The sample individual feature vector is input into a prediction module to obtain a sample credit risk probability of the sample enterprise, and the sample credit risk probability is used as the first sample credit risk result.

[0063] Optionally, the feature processing module includes a preprocessing module and a vector conversion module, and the prediction unit 22 is specifically configured to input the sample legal person credit data into the preprocessing module, perform binning processing on the sample legal person credit data based on the preprocessing module, and obtain serialized features corresponding to the sample legal person credit data; and / or, Inputting the sample legal person credit data into the preprocessing module, determining a branch path of the sample individual credit data in a decision tree based on the preprocessing module, and generating a tree coding feature corresponding to the sample legal person credit data based on the branch path; Based on the vector conversion module, the serialization feature and / or the tree encoding feature are embedded to obtain a sample legal person credit data vector corresponding to the sample legal person credit data.

[0064] It should be noted that the credit risk determination device provided in the above-mentioned embodiments, when executing the credit risk determination method, is illustrated only by the division of the aforementioned functional modules. In actual applications, the aforementioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the credit risk determination device provided in the above-mentioned embodiments and the credit risk determination method embodiment are based on the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.

[0065] The serial numbers of the embodiments in this specification are for descriptive purposes only and do not represent the merits of the embodiments. In some cases, the actions or steps recited in the claims may be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0066] The embodiment of this specification also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the above Figure 2-Figure 8 The credit risk determination method of the embodiment shown in the figure can be found in the specific implementation process. Figure 2-Figure 8The detailed description of the illustrated embodiment will not be repeated here.

[0067] Please refer to Figure 11 , which shows a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of this specification. The electronic device described in this specification may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected via the bus 150.

[0068] The processor 110 may include one or more processing cores. Using various interfaces and circuits, the processor 110 connects to various components within the electronic device. It executes instructions, programs, code sets, or instruction sets stored in the memory 120, as well as accesses data stored in the memory 120, to perform various functions and process data. Optionally, the processor 110 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interfaces, and applications; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 110 and may instead be implemented via a separate communications chip.

[0069] The memory 120 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium (Non-Transitory Computer-Readable Storage Medium). The memory 120 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc. The operating system may be an Android system, including a system deeply developed based on the Android system, an iOS system developed by Apple, including a system deeply developed based on the iOS system, or other systems.

[0070] The memory 120 can be divided into an operating system space and a user space. The operating system runs in the operating system space, and native and third-party applications run in the user space. In order to ensure that different third-party applications can achieve better operating results, the operating system allocates corresponding system resources to different third-party applications. However, the requirements for system resources in different application scenarios in the same third-party application are also different. For example, in the local resource loading scenario, the third-party application has higher requirements for disk reading speed; in the animation rendering scenario, the third-party application has higher requirements for GPU performance. The operating system and the third-party application are independent of each other, and the operating system often cannot perceive the current application scenario of the third-party application in a timely manner, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application.

[0071] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to open up data communication between third-party applications and the operating system so that the operating system can obtain the current scenario information of third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.

[0072] The input device 130 is used to receive input commands or data and includes, but is not limited to, a keyboard, a mouse, a camera, a microphone, or a touch-sensitive device. The output device 140 is used to output commands or data and includes, but is not limited to, a display device and a speaker. In one example, the input device 130 and the output device 140 may be combined, and the input device 130 and the output device 140 may be a touch-sensitive display.

[0073] The touch display screen can be designed as a full screen, a curved screen or a special-shaped screen. The touch display screen can also be designed as a combination of a full screen and a curved screen, or a combination of a special-shaped screen and a curved screen, which is not limited in the embodiments of this specification.

[0074] In addition, those skilled in the art will understand that the structures of the electronic devices shown in the above figures do not limit the electronic devices. The electronic devices may include more or fewer components than shown, or may combine certain components or arrange the components differently. For example, the electronic devices may also include radio frequency circuits, input units, sensors, audio circuits, WiFi modules, power supplies, Bluetooth modules, and other components, which will not be described in detail here.

[0075] exist Figure 11 In the electronic device shown, the processor 110 may be configured to call a computer application stored in the memory 120 and specifically perform the following operations: Obtain the legal person credit data of the target enterprise; A pre-trained credit student model is used to generate the target credit risk result of the target enterprise based on the legal person credit data; the credit student model is trained based on the enterprise credit risk label provided by the credit teacher model, and the enterprise credit risk label is obtained by the credit teacher model based on the enterprise credit data.

[0076] In one embodiment, the processor 110 is further configured to perform the following operations: if the corporate credit data of the target enterprise is obtained, a pre-trained credit teacher model is used to generate a target credit risk result of the target enterprise based on the corporate credit data.

[0077] In one embodiment, the processor 110 may be configured to call a computer application stored in the memory 120 and specifically perform the following operations: Obtain a first sample data set; the first sample data set includes sample legal person credit data and first sample credit risk labels of each first sample enterprise; Inputting the sample legal person credit data into an initial credit student model to obtain a first sample credit risk result output by the initial credit student model; Determining, based on a first preset loss function, a first loss value between the first sample credit risk result and the second sample credit risk result, and a second loss value between the first sample credit risk result and the first sample credit risk label; the second sample credit risk result is obtained by the credit teacher model based on the sample enterprise credit data of the first sample enterprise; The initial credit student model is iteratively trained based on the first loss value and the second loss value until the first loss value and the second loss value meet a convergence condition, thereby obtaining a trained credit student model.

[0078] In one embodiment, the processor 110 is further configured to perform the following operations: Obtain a second sample data set; the second sample data set includes sample enterprise credit data of each second sample enterprise and a second sample credit risk label; Inputting the sample enterprise credit data into the initial credit teacher model to obtain a third sample credit risk result output by the initial credit teacher model; determining, based on a second preset loss function, a third loss value between the third sample credit risk result and the second sample credit risk label; The initial credit teacher model is iteratively trained based on the third loss value until the third loss value meets the convergence condition, thereby obtaining a trained credit teacher model.

[0079] In one embodiment, the processor 110 is further configured to perform the following operations: The sample enterprise credit data of the first sample enterprise is input into the credit teacher model to obtain a second sample credit risk result output by the credit teacher model.

[0080] In one embodiment, the initial credit student model includes a feature processing module and a prediction module. When the processor 110 inputs the sample legal person credit data into the initial credit student model and obtains the first sample credit risk result output by the initial credit student model, the processor 110 specifically performs the following operations: Inputting the sample legal person credit data into the feature processing module to obtain a sample individual feature vector corresponding to the sample legal person credit data; The sample individual feature vector is input into a prediction module to obtain a sample credit risk probability of the sample enterprise, and the sample credit risk probability is used as the first sample credit risk result.

[0081] In one embodiment, the feature processing module includes a preprocessing module and a vector conversion module. When the processor 110 inputs the sample legal person credit data into the feature processing module to obtain the sample individual feature vector corresponding to the sample legal person credit data, the processor 110 specifically performs the following operations: Inputting the sample legal person credit data into the pre-processing module, performing binning processing on the sample legal person credit data based on the pre-processing module to obtain serialization features corresponding to the sample legal person credit data; and / or, Inputting the sample legal person credit data into the preprocessing module, determining a branch path of the sample individual credit data in a decision tree based on the preprocessing module, and generating a tree coding feature corresponding to the sample legal person credit data based on the branch path; Based on the vector conversion module, the serialization feature and / or the tree encoding feature are embedded to obtain a sample legal person credit data vector corresponding to the sample legal person credit data.

[0082] In the embodiments of this specification, by obtaining the legal person credit data of the target enterprise, a pre-trained credit student model is used to generate the target credit risk results for the target enterprise based on the legal person credit data. Risk assessment is performed on enterprises without corporate credit using the legal person's personal credit, allowing enterprises with legal person personal credit but no corporate credit to share the value of corporate credit. The credit student model is trained based on the corporate credit risk labels provided by the credit teacher model. The credit teacher model is capable of performing corporate risk assessment based on corporate credit data. The corporate credit risk labels are obtained by the credit teacher model based on corporate credit data. By leveraging the high similarity between corporate credit data and personal credit data, a credit teacher model and a credit student model are constructed, using corporate credit to guide the learning of personal credit. The use of a distillation learning-based model can improve the accuracy of risk assessments for enterprises without corporate credit information.

[0083] Furthermore, a first sample dataset is obtained and the sample legal person credit data is input into an initial credit student model. A first sample credit risk result is output by the initial credit student model. Based on a first preset loss function, a first loss value is determined between the first sample credit risk result and the second sample credit risk result, as well as a second loss value between the first sample credit risk result and the sample credit risk label. The second sample credit risk result is obtained by the credit teacher model based on the sample enterprise credit data of the sample enterprise. The initial credit student model is iteratively trained based on the first and second loss values ​​until the first and second loss values ​​meet convergence conditions, thereby obtaining a trained credit student model. This model training method allows the credit student model trained on personal credit data to maximize the value of corporate credit data, even when the volume of personal credit data queries is very high. Specifically, while some corporate legal persons may not have corporate credit data, they may have personal credit data. By learning information about their personal credit through the credit student model, the value of corporate credit data can still be shared and utilized to some extent. This is because, through distillation, the model based on personal credit data can indirectly reflect some characteristics and patterns of corporate credit, thereby providing a similar basis for judging corporate credit for enterprises without corporate credit data.

[0084] Furthermore, a second sample data set is obtained and the sample enterprise credit data in the second sample data set is input into the initial credit teacher model to obtain a third sample credit risk result output by the initial credit teacher model. Based on a second preset loss function, a third loss value between the third sample credit risk result and the second sample credit risk label is determined. The initial credit teacher model is iteratively trained based on the third loss value until the third loss value meets the convergence condition, thereby obtaining a trained credit teacher model. The credit teacher model learns from the enterprise credit data, enabling it to generate accurate risk result soft labels based on the sample enterprise credit data for the credit student model to learn from.

[0085] Furthermore, by inputting sample legal person credit data into the feature processing module, a sample individual feature vector corresponding to the sample legal person credit data is obtained. The sample individual feature vector is then input into the prediction module to obtain a sample credit risk probability for the sample enterprise, which is used as the first sample credit risk result. The stability of model training is improved by preprocessing the input data, including binning, serialization, and tree encoding.

[0086] In addition, an embodiment of this specification provides a computer program product, which includes a computer program. When the computer program is executed by a processor of an electronic device, the processor can at least implement the above-mentioned Figures 2 to 8 The methods provided in the illustrated embodiments.

[0087] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0088] The above disclosure is only a preferred embodiment of this specification, and certainly cannot be used to limit the scope of rights of this specification. Therefore, equivalent changes made according to the claims of this specification are still within the scope covered by this specification.

Claims

1. A method for determining credit risk, comprising: Obtaining a first sample data set; The first sample data set includes sample legal person credit data and first sample credit risk labels of each first sample enterprise; Inputting the sample legal person credit data into an initial credit student model to obtain a first sample credit risk result output by the initial credit student model; Determining, based on a first preset loss function, a first loss value between the first sample credit risk result and the second sample credit risk result, and a second loss value between the first sample credit risk result and the first sample credit risk label; The second sample credit risk result is obtained by the credit teacher model based on the sample enterprise credit data of the first sample enterprise; Iteratively training the initial credit student model based on the first loss value and the second loss value until the first loss value and the second loss value meet a convergence condition, thereby obtaining a trained credit student model; The credit student model is used to generate a target credit risk result of the target enterprise based on the legal person credit data of the target enterprise.

2. The method according to claim 1, further comprising: Obtaining a second sample data set; The second sample data set includes sample enterprise credit data and second sample credit risk labels of each second sample enterprise; Inputting the sample enterprise credit data into the initial credit teacher model to obtain a third sample credit risk result output by the initial credit teacher model; determining, based on a second preset loss function, a third loss value between the third sample credit risk result and the second sample credit risk label; The initial credit teacher model is iteratively trained based on the third loss value until the third loss value meets the convergence condition, thereby obtaining a trained credit teacher model.

3. The method of claim 2, further comprising: The sample enterprise credit data of the first sample enterprise is input into the credit teacher model to obtain a second sample credit risk result output by the credit teacher model.

4. The method of claim 1, wherein the initial credit student model comprises a feature processing module and a prediction module, and wherein inputting the sample legal person credit data into the initial credit student model to obtain a first sample credit risk result output by the initial credit student model comprises: Inputting the sample legal person credit data into the feature processing module to obtain a sample individual feature vector corresponding to the sample legal person credit data; The sample individual feature vector is input into a prediction module to obtain a sample credit risk probability of the sample enterprise, and the sample credit risk probability is used as the first sample credit risk result.

5. The method according to claim 4, wherein the feature processing module comprises a preprocessing module and a vector conversion module, and wherein inputting the sample legal person credit data into the feature processing module to obtain a sample individual feature vector corresponding to the sample legal person credit data comprises: Inputting the sample legal person credit data into the preprocessing module, performing binning processing on the sample legal person credit data based on the preprocessing module, and obtaining serialization features corresponding to the sample legal person credit data; and / or, Inputting the sample legal person credit data into the preprocessing module, determining a branch path of the sample individual credit data in a decision tree based on the preprocessing module, and generating a tree coding feature corresponding to the sample legal person credit data based on the branch path; Based on the vector conversion module, the serialization feature and / or the tree encoding feature is embedded to obtain a sample legal person credit data vector corresponding to the sample legal person credit data.

6. The method of claim 2, further comprising: Obtain the legal person credit data of the target enterprise; The credit student model is used to generate a target credit risk result of the target enterprise based on the legal person credit data.

7. The method of claim 6, further comprising: If the corporate credit data of the target enterprise is obtained, the credit teacher model is used to generate the target credit risk result of the target enterprise based on the corporate credit data.

8. A training device for a risk determination model, comprising: A sample acquisition unit, configured to acquire a first sample data set; The first sample data set includes sample legal person credit data and sample credit risk labels of each first sample enterprise; a prediction unit, configured to input the sample legal person credit data into an initial credit student model to obtain a first sample credit risk result output by the initial credit student model; a loss calculation unit, configured to determine, based on a first preset loss function, a first loss value between the first sample credit risk result and the second sample credit risk result, and a second loss value between the first sample credit risk result and the first sample credit risk label; The second sample credit risk result is obtained by the credit teacher model based on the sample enterprise credit data of the first sample enterprise; a training unit, configured to iteratively train the initial credit student model based on the first loss value and the second loss value until the first loss value and the second loss value meet a convergence condition, thereby obtaining a trained credit student model; The credit student model is used to generate a target credit risk result of the target enterprise based on the legal person credit data of the target enterprise.

9. An electronic device comprising: processor and memory; The memory stores a computer program, which is suitable for being loaded by the processor and executing the steps of the method according to any one of claims 1 to 7.

10. A storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.

11. A computer program product comprising: A computer program, when executed by a processor of an electronic device, causes the processor to perform the steps of the method according to any one of claims 1 to 7.