Risk prediction method, device and electronic device

Through the risk prediction model of multi-perspective boundary discrimination constraints, the classification hyperplane is optimized, which solves the problem of identification of high-risk customers in loan business, and improves the accuracy of risk prediction and the profitability of financial institutions.

CN113052512BActive Publication Date: 2025-07-29INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110516028.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-12
Publication Date
2025-07-29
Estimated Expiration
2041-05-12

AI Technical Summary

Technical Problem

In complex loan business scenarios, it is difficult for existing technology to effectively identify high-risk customers, resulting in an increase in non-performing loans and affecting the profits and reputation of financial institutions.

Method used

A risk prediction model of multi-view boundary discrimination constraints is adopted. By obtaining multi-dimensional features and using boundary samples to distinguish constraint terms and inter-view discrimination constraint terms, the classification hyperplane is optimized, so that similar boundary samples are approached in the output space and heterogeneous boundary samples are far away, improving the generalization performance of the model.

Benefits of technology

It improves the accuracy and generalization ability of risk prediction, reduces the occurrence of non-performing loans, and enhances the competitiveness of financial institutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113052512B_ABST
    Figure CN113052512B_ABST
Patent Text Reader

Abstract

The present disclosure provides a risk prediction method, apparatus, and electronic device, which can be used in the fields of artificial intelligence, finance, etc. The risk prediction method includes: obtaining data to be predicted; obtaining risk features of the data to be predicted; and processing the risk features by using a trained risk prediction model to obtain a risk prediction result, where the risk prediction result includes the category to which the data to be predicted belongs; wherein the sample data of each category includes boundary sample data, and the objective function of the risk prediction model includes a boundary sample discrimination constraint term, and the boundary sample discrimination constraint term makes the first distance of the boundary sample data belonging to different categories in the output space greater than the second distance of the boundary sample data belonging to the same category in the output space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of artificial intelligence technology and finance, and more particularly, to a risk prediction method, apparatus, and electronic device. Background Art

[0002] For various types of institutions, risk prediction is a hot topic. For example, with the development of the financial industry, the proportion of corporate loan business in financial institutions is increasing, and the risk prediction of corporate loan business has become an increasingly important matter.

[0003] In the process of implementing the concept of the present disclosure, the applicant found that there are at least the following problems in the related art. Some business scenarios (such as loan business scenarios) have the characteristic of high complexity, making it difficult to identify high-risk customers before processing the business, resulting in abnormal business processing, such as the increasing severity of non-performing loans, which will have an adverse impact on the institution. Summary of the Invention

[0004] In view of this, the present disclosure provides a risk prediction method, apparatus, and electronic device to at least partially solve the problem of high-risk prediction.

[0005] One aspect of the present disclosure provides a risk prediction method, including: obtaining data to be predicted; obtaining risk characteristics of the data to be predicted; and using a trained risk prediction model to process the risk characteristics to obtain a risk prediction result, where the risk prediction result includes the category to which the data to be predicted belongs; wherein, the sample data of each category includes boundary sample data, and the objective function of the risk prediction model includes a boundary sample discrimination constraint term, and the boundary sample discrimination constraint term makes the first distance in the output space between boundary sample data belonging to different categories greater than the second distance in the output space between boundary sample data belonging to the same category.

[0006] According to an embodiment of the present disclosure, the category includes a positive sample category and a negative sample category; and the boundary sample discrimination constraint term includes: a positive sample sub-constraint term, a negative sample sub-constraint term, and a cross term, wherein the output of the positive sample sub-constraint term is related to the difference between the sub-result of the model processing the positive sample data and the sub-result of the model processing the mean of the boundary positive sample subset, the output of the negative sample sub-constraint term is related to the difference between the sub-result of the model processing the negative sample data and the sub-result of the model processing the mean of the boundary negative sample subset, and the output of the cross term is related to the product of the sub-result of the model processing the positive sample data and the sub-result of the model processing the negative sample data.

[0007] According to an embodiment of the present disclosure, obtaining the risk characteristics of the data to be predicted includes: obtaining at least one of the sub-characteristics of the data to be predicted from the perspective of basic information, obtaining the sub-characteristics of the data to be predicted from the perspective of business operation information, and obtaining the sub-characteristics of the data to be predicted from the perspective of behavioral information.

[0008] According to an embodiment of the present disclosure, the boundary sample discrimination constraint items include at least one of the following: positive sample sub-constraint items, negative sample sub-constraint items and cross items from the basic information perspective; positive sample sub-constraint items, negative sample sub-constraint items and cross items from the business information perspective; or positive sample sub-constraint items, negative sample sub-constraint items and cross items from the behavioral information perspective.

[0009] According to an embodiment of the present disclosure, the objective function of the risk prediction model also includes an inter-view boundary sample discrimination constraint term, which makes the output of the mean of boundary samples of the same category in different viewpoints tend to be consistent.

[0010] According to an embodiment of the present disclosure, the inter-view boundary sample discrimination constraint item includes at least one of the following: positive sample sub-constraint items and negative sample sub-constraint items for different view angles; or positive sample sub-constraint items and negative sample sub-constraint items for different view angles.

[0011] According to an embodiment of the present disclosure, training a risk prediction model includes: obtaining a training sample data set, the training sample data set including a positive training sample data subset and a negative training sample data subset; determining a boundary positive sample subset in the positive training sample data subset based on the heterogeneous neighbor information of the sample, and determining a boundary negative sample subset in the negative training sample data subset based on the heterogeneous neighbor information of the sample; and inputting the positive training sample data subset and / or the negative training sample data subset into the risk prediction model, and adjusting the parameters of the risk prediction model until a preset number of iterations is reached or the difference in the loss function of the objective function during two iterations is less than a preset threshold.

[0012] According to an embodiment of the present disclosure, determining a boundary positive sample subset in a positive training sample data subset based on heterogeneous neighbor information of the sample includes: for any negative sample, adding a first specified number of positive class samples that are neighbors of the negative sample to the boundary positive sample subset; and determining a boundary negative sample subset in a negative training sample data subset based on heterogeneous neighbor information of the sample includes: for any positive sample, adding a second specified number of negative class samples that are neighbors of the positive sample to the boundary negative sample subset.

[0013] According to an embodiment of the present disclosure, the above method further includes: testing the trained risk prediction model using a test sample set to obtain a test accuracy of the risk prediction result.

[0014] According to an embodiment of the present disclosure, obtaining risk features of the data to be predicted includes at least one of the following: performing one-hot encoding on the category data in the data to be predicted to obtain category features; calculating the associated data of the business information and / or behavior information in the data to be predicted to obtain derived features.

[0015] According to an embodiment of the present disclosure, the objective function further includes an empirical loss constraint term and a regularization constraint term.

[0016] According to an embodiment of the present disclosure, the sample data of each category further includes non-boundary sample data, and the third distance between the boundary sample data and the class center in the same category is greater than the fourth distance between the non-boundary sample data and the class center in the same category.

[0017] One aspect of the present disclosure provides a risk prediction device, including: a data acquisition module, a risk feature acquisition module, and a risk feature processing module. Among them, the data acquisition module is used to acquire data to be predicted; the risk feature acquisition module is used to acquire the risk features of the data to be predicted; and the risk feature processing module is used to process the risk features by using a trained risk prediction model to obtain a risk prediction result, and the risk prediction result includes the category to which the data to be predicted belongs; wherein, the sample data of each category includes boundary sample data, and the objective function of the risk prediction model includes a boundary sample discrimination constraint term, and the boundary sample discrimination constraint term makes the first distance between the boundary sample data belonging to different categories in the output space greater than the second distance between the boundary sample data belonging to the same category in the output space.

[0018] Another aspect of the present disclosure provides an electronic device, including one or more processors and a storage device, wherein the storage device is used to store executable instructions, and when the executable instructions are executed by the processor, the above-mentioned risk prediction method is implemented.

[0019] Another aspect of the present disclosure provides a computer-readable storage medium, storing computer-executable instructions, and the instructions are used to implement the above-mentioned risk prediction method when executed.

[0020] Another aspect of the present disclosure provides a computer program, which includes computer-executable instructions, and the instructions are used to implement the above-mentioned risk prediction method when executed.

[0021] The risk prediction method, device and electronic device provided by the embodiments of the present disclosure calculate the means of two types of boundary sample data sets through the boundary sample set, so as to mine and learn the distribution information unique to the boundary samples to optimize the classification hyperplane. The boundary sample discrimination constraint term provided by the embodiments of the present disclosure makes the same-class boundary sample data as close as possible in the output space, and makes the different-class boundary sample data as far away as possible in the output space, so that the classification hyperplane passes through the middle region of the two types of boundary sample data as much as possible, so as to improve the generalization performance of the risk prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above and other objects, features and advantages of the present disclosure will become clearer. In the drawings:

[0023] Figure 1 Schematically shows an exemplary system architecture to which the risk prediction method, apparatus, and electronic device according to embodiments of the present disclosure can be applied;

[0024] Figure 2 Schematically shows a flowchart of the risk prediction method according to embodiments of the present disclosure;

[0025] Figure 3 Schematically shows a logic diagram of the risk prediction method according to embodiments of the present disclosure;

[0026] Figure 4 Schematically shows a flowchart of the method for training a risk prediction model according to embodiments of the present disclosure;

[0027] Figure 5 Schematically shows a schematic diagram of the training data according to embodiments of the present disclosure;

[0028] Figure 6 Schematically shows a logic diagram of training a risk prediction model according to embodiments of the present disclosure;

[0029] Figure 7 Schematically shows a schematic diagram of a subset of boundary sample data according to embodiments of the present disclosure;

[0030] Figure 8 Schematically shows a flowchart of the risk prediction method according to another embodiment of the present disclosure;

[0031] Figure 9 Schematically shows a block diagram of the risk prediction apparatus according to embodiments of the present disclosure; and

[0032] Figure 10 Schematically shows a block diagram of the electronic device according to embodiments of the present disclosure. Detailed Description of the Embodiments

[0033] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.

[0034] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components. All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0035] In the case of using expressions such as "at least one of A, B, or C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). The terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features.

[0036] With the development of the financial industry, the proportion of corporate loan business in financial institutions is increasing. Due to the complexity of this business scenario, it is difficult to detect high-risk customers in advance. If non-performing loans become more and more serious, it will have an adverse impact on financial institutions, resulting in a decline in the reputation of financial institutions and a reduction in profits, etc.

[0037] To better conduct risk prediction, multi-dimensional features can be extracted and risk prediction can be performed based on the multi-dimensional features. Taking the risk prediction of corporate loans as an example for illustration, there are still deficiencies in the risk prediction of corporate loans in the related art. For example, through a large amount of analysis and research, the applicant found that the machine learning methods in the related art treat all training samples equally and do not distinguish the different importance of boundary samples and non-boundary samples for classification. In fact, samples with different spatial distributions also have different importance for the classification hyperplane. If all training samples are treated equally and the unique distribution information of boundary samples is not mined and learned to optimize the classification hyperplane, it may lead to the model not achieving the expected effect.

[0038] The development of big data technology has made it increasingly easy to accumulate customer-related feature information. Corporate loan risk prediction involves a wide range of feature categories, which can be used to construct different perspectives on a sample. For example, feature extraction can be performed based on basic information, operational information, and behavioral information. Multi-perspective learning techniques can leverage these perspectives to model and learn from the different perspectives of a sample, and make predictions about unknown samples. Therefore, applying multi-perspective learning techniques to corporate loan risk prediction is a worthwhile approach.

[0039] During the trial process, the applicant found that although the multi-perspective model contains information from multiple perspectives, if it is trained using the model training method in the relevant technology, since the learning processes between multiple perspectives are relatively independent, only the sub-results of multiple perspectives are integrated into the final risk prediction model. If the information contained in a certain perspective is not sufficient to provide the category information of the sample, then the existence of this perspective will instead reduce the classification effect of the final model.

[0040] The risk prediction method, device and electronic device provided by the embodiments of the present disclosure include a feature acquisition process and a prediction result output process, wherein, in the feature acquisition process, first, the data to be predicted is acquired, and then the risk features of the data to be predicted are acquired. After completing the feature acquisition process, the prediction result output process is entered, and the risk features are processed using a trained risk prediction model to obtain a risk prediction result, which includes the category to which the data to be predicted belongs. The sample data of each category includes boundary sample data, and the objective function of the risk prediction model includes a boundary sample discriminant constraint item. The boundary sample discriminant constraint item makes the first distance of the boundary sample data belonging to different categories in the output space greater than the second distance of the boundary sample data belonging to the same category in the output space.

[0041] The present disclosure provides a risk prediction model based on multi-perspective boundary discrimination constraints, such as a corporate loan risk prediction model. On the one hand, a boundary sample set is selected within the perspective through the heterogeneous neighbor information of the sample, and the mean of the two types of boundary sample sets is calculated through the boundary sample set. Through the boundary sample discrimination constraint item designed by this patent, the boundary samples of the same type are made as close as possible in the output space, and the heterogeneous boundary samples are made as far away as possible in the output space, so that the classification hyperplane passes through the middle area of the two types of boundary samples as much as possible to improve the generalization performance of the model. On the one hand, between the basic information perspective, the business information perspective and the behavioral information perspective, the consistency of the boundary samples in different output spaces is mined. Through the boundary sample discrimination constraint item between perspectives provided by this embodiment, the output of the mean of the boundary samples of the same type in different perspectives is made as consistent as possible. The purpose is to optimize each other between multiple perspectives to improve the accuracy of the classification boundary.

[0042] The method, apparatus, and electronic device for predicting risks provided by the embodiments of the present disclosure can be used in the field of artificial intelligence for aspects related to risk prediction, and can also be used in various fields other than the field of artificial intelligence, such as the financial field. The application fields of the risk prediction method, apparatus, and electronic device provided by the embodiments of the present disclosure are not limited.

[0043] Figure 1 Schematically shows an exemplary system architecture to which the method, apparatus, and electronic device for predicting risks according to the embodiments of the present disclosure can be applied. It should be noted that, Figure 1 What is shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios.

[0044] Such as Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and servers 105, 106, 107. The network 104 may include multiple gateways, routers, hubs, network cables, etc., and is used as a medium to provide a communication link between the terminal devices 101, 102, 103 and the servers 105, 106, 107. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0045] Users can use the terminal devices 101, 102, 103 to interact with other terminal devices and servers 105, 106, 107 through the network 104 to receive or send information, such as sending risk prediction requests, training model requests, model maintenance requests, and receiving processing results, etc. The terminal devices 101, 102, 103 may be installed with various communication client applications, such as risk prediction applications, software development applications, banking applications, government affairs applications, monitoring applications, web browser applications, search applications, office applications, instant messaging tools, email clients, social platform software, etc. (only as examples). For example, users can use the terminal device 101 to view the risk prediction results and related processing suggestions, automatic processing results, etc. feedback from the server side. For example, users can request the server to perform model iterative training, etc.

[0046] The terminal devices 101, 102, 103 include but are not limited to smart phones, virtual reality devices, augmented reality devices, tablet computers, laptop portable computers, desktop computers, and the like.

[0047] Servers 105, 106, and 107 can receive requests and process them, which can specifically be storage servers, background management servers, server clusters, etc. For example, server 105 can store a risk prediction model, server 106 can be used as a model training server to optimize model parameters, etc., and server 107 can store business data, training databases, etc.

[0048] It should be noted that the method for predicting risk provided in the embodiments of the present disclosure can generally be executed by a server. Correspondingly, the device for predicting risk provided in the embodiments of the present disclosure can generally be arranged in a server. The method for predicting risk provided in the embodiments of the present disclosure can also be executed by a server or a server cluster capable of communicating with terminal devices 101, 102, 103, and / or servers 105, 106, 107.

[0049] It should be understood that the numbers of terminal devices, networks, and servers are merely illustrative. According to actual needs, there can be any number of terminal devices, networks, and servers.

[0050] Figure 2 A flowchart of a risk prediction method according to an embodiment of the present disclosure is schematically shown.

[0051] As Figure 2 shown, the above method includes operation S210 to operation S230.

[0052] In operation S210, obtain data to be predicted. The data to be predicted can be various business data. For example, data generated during the process of customers handling business. Taking a financial institution as an example, the data to be predicted can be data generated during the process of users applying for loans, applying for credit cards, applying for credit limits, etc. Specifically, the data to be predicted can include operation information and behavior information, etc., and can also include customer attribute information. Customer attribute information can include at least one of the following: name, unit name, address, annual income, etc. In addition, the data to be predicted can also include statistical information, such as the user's historical consumption amount, credit information, consumption habits, consumption preferences, etc.

[0053] In operation S220, obtain the risk characteristics of the data to be predicted.

[0054] In this embodiment, the risk characteristics can be used to characterize the potential risk level of the object to be predicted. For example, for the same loan amount, the risk of default for people with a high annual income is lower than that of people with a low income.

[0055] In some embodiments, obtaining the risk characteristics of the data to be predicted includes at least one of the following: performing one-hot encoding on the categorical data in the data to be predicted to obtain categorical characteristics. Calculating the associated data of the operation information and / or behavior information in the data to be predicted to obtain derivative characteristics.

[0056] Specifically, during the process of extracting risk features, for categorical features such as the company's industry category and the company's economic nature, one-hot encoding is performed on them.

[0057] For the processing of derived indicators, relevant original features of business information and behavioral information are used to form derived features, such as the mean, standard deviation, maximum value, and minimum value.

[0058] For example, for a certain legal person, first, the features related to the risk prediction of the legal person's loan are divided into three categories: basic information, business information, and behavioral information. The data range can be determined according to the category, and thus the data table involved can be determined.

[0059] Then, features are taken from the data table, including basic information, business information, and behavioral information.

[0060] Next, the categorical features are transformed, the derived features are processed, and the labels are constructed.

[0061] In operation S230, the risk features are processed using the trained risk prediction model to obtain a risk prediction result, and the risk prediction result includes the category to which the data to be predicted belongs.

[0062] In this embodiment, the sample data of each category includes boundary sample data, and the objective function of the risk prediction model includes a boundary sample discrimination constraint term. The boundary sample discrimination constraint term makes the first distance in the output space between the boundary sample data belonging to different categories greater than the second distance in the output space between the boundary sample data belonging to the same category.

[0063] Among them, the risk prediction model can be various machine learning models, such as neural networks, etc. The risk prediction model can be trained through algorithms such as backpropagation to improve the prediction accuracy.

[0064] The following gives an exemplary description of the risk prediction model.

[0065] In some embodiments, the categories include a positive sample category and a negative sample category.

[0066] The boundary sample discrimination constraint term includes: a positive sample sub-constraint term, a negative sample sub-constraint term, and a cross term. Among them, the output of the positive sample sub-constraint term is related to the difference between the sub-result of the model processing the positive sample data and the sub-result of the model processing the mean of the boundary positive sample subset. The output of the negative sample sub-constraint term is related to the difference between the sub-result of the model processing the negative sample data and the sub-result of the model processing the mean of the boundary negative sample subset. The output of the cross term is related to the product of the sub-result of the model processing the positive sample data and the sub-result of the model processing the negative sample data.

[0067] For example, the boundary sample discrimination constraint term can be as shown in Equation (1).

[0068] Equation (1)

[0069] Wherein, is the negative-class boundary sample set, is the positive-class boundary sample set, and are the means of the two types of boundary sample sets. is the classifier. It should be noted that the above formula is only shown exemplarily. For example, taking the square in the formula can also be the first power or the third power, etc. In addition, there are also settings of coefficients, constant offsets, etc., which are not limited herein.

[0070] It should be noted that the above only shows the processing process of the boundary sample set, and the non-boundary sample set can also be processed. Specifically, the sample data of each category also includes non-boundary sample data. The third distance between the boundary sample data and the class center in the same category is greater than the fourth distance between the non-boundary sample data and the class center.

[0071] The boundary sample discrimination constraint term provided by the embodiments of the present disclosure makes the same-class boundary samples as close as possible in the output space, makes the different-class boundary samples as far away as possible in the output space, and makes the classification hyperplane pass through the middle region of the two types of boundary samples as much as possible, so as to improve the generalization performance of the model.

[0072] In some embodiments, in order to further improve the model prediction accuracy, risk prediction can also be performed separately from multiple perspectives, and then the final risk prediction result is determined based on the prediction sub-results of each perspective.

[0073] For example, obtaining the risk features of the data to be predicted includes at least one of: obtaining the sub-features of the data to be predicted from the perspective of basic information, obtaining the sub-features of the data to be predicted from the perspective of business information, and obtaining the sub-features of the data to be predicted from the perspective of behavior information. Among them, the extraction process of the sub-features for different perspectives can be as shown above and will not be elaborated herein.

[0074] In some embodiments, the boundary sample discrimination constraint term can include at least one of the following: the positive sample sub-constraint term, the negative sample sub-constraint term, and the cross term for the basic information perspective. The positive sample sub-constraint term, the negative sample sub-constraint term, and the cross term for the business information perspective. The positive sample sub-constraint term, the negative sample sub-constraint term, and the cross term for the behavior information perspective.

[0075] For example, the boundary sample discrimination constraint term can be as shown in Equation (2).

[0076] Equation (2)

[0077] Among them, represents the perspective serial number, is the negative class boundary sample set, is the positive class boundary sample set, , is the mean of the two-class boundary sample sets. is the sub-classifier for the th perspective. It should be noted that the above formula is only shown exemplarily. For example, taking the square in the formula can also be the first power or the third power, etc. In addition, there are also settings such as coefficients and constant offsets, which are not limited here.

[0078] In some embodiments, the objective function of the risk prediction model may further include an inter-perspective boundary sample discrimination constraint term, which makes the means of the boundary samples of the same class tend to be consistent in the respective outputs of different perspectives.

[0079] The inter-perspective boundary sample discrimination constraint term provided in this embodiment makes the outputs of the means of the same-class boundary samples as consistent as possible in different perspectives. The purpose is to optimize among multiple perspectives to improve the accuracy of the classification boundary.

[0080] Specifically, the inter-perspective boundary sample discrimination constraint term may include at least one of the following: a positive sample sub-constraint term and a negative sample sub-constraint term for different perspectives; or a positive sample sub-constraint term and a negative sample sub-constraint term for different perspectives.

[0081] For example, the inter-perspective boundary sample discrimination constraint term can be as shown in Equation (3).

[0082] Equation (3)

[0083] Among them, , respectively represent the perspective serial numbers.

[0084] In some embodiments, the objective function may further include an empirical loss constraint term and a regularization constraint term.

[0085] For example, the expression of the objective function can be as shown in Equation (4).

[0086] Equation (4)

[0087] Among them, is the empirical loss, is the regularization term, , , are hyperparameters used to adjust the weights of the above items. For example, and are respectively as shown in Equation (5) and Equation (6).

[0088] Equation (5)

[0089] Equation (6)

[0090] Wherein, is the label of the sample, is the feature weight parameter of the sub - model and T represents transpose.

[0091] Figure 3 Schematically shows a logic diagram of a risk prediction method according to an embodiment of the present disclosure.

[0092] As Figure 3 shown, first, obtain feature information related to corporate loan risk prediction from a data warehouse, such as basic information, business information, and behavior information. The basic information includes the industry category to which the enterprise belongs, economic nature, credit rating, enterprise popularity, etc. The business information includes the amount and number of account operating inflows and outflows of the enterprise in the past year. The behavior information includes transaction amount, transaction currency, capital flow direction, etc. Perform data pre - processing and feature engineering processing on the samples to construct the basic information perspective, business information perspective, and behavior information perspective of the samples. Use the features of the data to be predicted to construct test samples. Input the test samples into a corporate loan risk prediction model based on multi - perspective boundary discrimination constraints to obtain a prediction result.

[0093] Among them, the pre - processing process can be as follows.

[0094] First, data selection can be performed. For example, positive samples can be high - quality customers, and the selection criteria for high - quality customers can be that the customer has no problems with previous repayment records and is a stable - operating corporate customer. Classify the features related to corporate loan risk prediction into three categories: basic information, business information, and behavior information. The data range can be determined according to the category, thereby determining the relevant data tables.

[0095] Then, perform data pre - processing. Such as the data columns related to basic information, business information, and behavior information in the data table. Concatenate the relevant data columns in different tables according to the customer identifier (id) to form the original features. For columns with missing values, complete them in a certain way. For example, for missing values of numerical features, complete them with the column mean, and for missing values of non - numerical features, complete them with "unknown".

[0096] The following gives an exemplary description of the training process of the risk prediction model.

[0097] Figure 4 Schematically shows a flowchart of a method for training a risk prediction model according to an embodiment of the present disclosure.

[0098] As Figure 4As shown, the training risk prediction model may include operations S410 to S430.

[0099] In operation S410, a training sample data set is obtained, and the training sample data set includes a positive training sample data subset and a negative training sample data subset.

[0100] First, it is necessary to construct training samples.

[0101] Figure 5 A schematic diagram of training data according to an embodiment of the present disclosure is schematically shown.

[0102] As Figure 5 shown, the training data can be taken from a sample set. The sample set may include a negative sample subset and a positive sample subset. The negative sample subset may include a boundary negative sample subset and a non-boundary negative sample subset. The positive sample subset may include a boundary positive sample subset and a non-boundary positive sample subset. Each sample can be respectively subjected to feature extraction from three perspectives.

[0103] The labels of the labeled samples are 1 ( ) and -1 ( ) which can respectively represent corporate loan risk customers and risk-free customers. Each sample consists of three perspectives, namely the basic information perspective, the business information perspective, and the behavior information perspective.

[0104] In operation S420, based on the heterogeneous nearest neighbor information of the samples, the boundary positive sample subset in the positive training sample data subset is determined, and / or, based on the heterogeneous nearest neighbor information of the samples, the boundary negative sample subset in the negative training sample data subset is determined. Among them, in order to facilitate the determination of boundary samples, one or more samples belonging to other categories that are closest to any sample of the current category can be used as boundary samples.

[0105] In operation S430, the positive training sample data subset and / or the negative training sample data subset are input into the risk prediction model, and the parameters of the risk prediction model are adjusted until a preset number of iterations is reached or the difference in the loss function of the objective function between two iterations is less than a preset threshold.

[0106] Specifically, the gradient descent method is used to solve this optimization problem until a preset number of iterations is reached or the difference between the loss values of the two loss functions is less than a preset threshold. For example, the gradient descent method is used to minimize the objective function to obtain a sub-classification model in each perspective, such as obtaining the final classification model , and the algorithm formula is as shown in Equation (7).

[0107] Equation (7)

[0108] Among them, the objective function of the risk prediction model can be as shown above, which is not limited here. The risk prediction model is trained by the gradient descent method.

[0109] Figure 6 Schematically shows a logic diagram for training a risk prediction model according to an embodiment of the present disclosure.

[0110] As Figure 6 shown, after data preprocessing, training samples are obtained, which consist of three perspectives, namely the basic information perspective, the business information perspective, and the behavior information perspective. First, within the perspective, boundary sample sets are selected through the heterogeneous nearest neighbor information of the samples, and the means of the two types of boundary sample sets are calculated through the boundary sample subset. Through the boundary sample discrimination constraint term designed in this patent, the same-class boundary samples are made as close as possible in the output space, and the different-class boundary samples are made as far away as possible in the output space, so that the classification hyperplane passes through the middle region of the two types of boundary samples as much as possible to improve the generalization performance of the model. Between different perspectives, the consistency of boundary samples in different output spaces is mined. Through the inter-perspective boundary sample discrimination constraint term provided in this embodiment, the means of the same-class boundary samples are made as consistent as possible in the outputs of different perspectives, aiming to optimize among multiple perspectives to improve the accuracy of the classification boundary. By minimizing the empirical loss of the model, the boundary sample discrimination constraint term, and the inter-perspective boundary sample discrimination constraint term, three sub-classifiers are obtained. Finally, the results of the three sub-classifiers are integrated to classify and predict the test samples.

[0111] In some embodiments, determining the boundary positive sample subset in the positive training sample data subset based on the heterogeneous nearest neighbor information of the samples includes: for any negative sample, adding the first specified number of positive class samples adjacent to the negative sample to the boundary positive sample subset. Wherein, the first specified number includes but is not limited to: 1, 2, 3, 4, 5, 7, 8, 10, 11 or more.

[0112] Determining the boundary negative sample subset in the negative training sample data subset based on the heterogeneous nearest neighbor information of the samples includes: for any positive sample, adding the second specified number of negative class samples adjacent to the positive sample to the boundary negative sample subset. Wherein, the second specified number includes but is not limited to: 1, 2, 3, 4, 5, 7, 8, 10, 11 or more. The first specified number and the second specified number can be the same or different.

[0113] Figure 7 Schematically shows a schematic diagram of the boundary sample data subset according to an embodiment of the present disclosure.

[0114] As Figure 7 shown, screen the boundary sample set. The boundary sample set is selected through the heterogeneous nearest neighbor information of the samples. For any positive class sample , add its 5 nearest negative class samples to the negative class boundary sample set ; For any positive class sample , add its 5 nearest negative class samples to the positive class boundary sample set , and calculate the means of the two types of boundary sample sets through the boundary sample subsets , . The calculation formulas are shown in Formulas (8) to (10).

[0115] Formula (8)

[0116] Among them, represents 5 positive samples near the negative sample x, represents 5 negative samples near the positive sample x.

[0117] Formula (9)

[0118] Formula (10)

[0119] Among them, Mean() represents calculating the average value.

[0120] In some embodiments, after completing the model training, the above method may further include a verification operation.

[0121] Figure 8 Schematically shows a flowchart of a risk prediction method according to another embodiment of the present disclosure.

[0122] As Figure 8 shown, after performing operation S530 for model training in the above method, operation S810 may further be included.

[0123] In operation S810, use the test sample set to test the trained risk prediction model to obtain the test accuracy of the risk prediction result.

[0124] For the test sample , input the discriminant function of the classifier to obtain the discriminant result of the model.

[0125] Formula (11)

[0126] The risk prediction method provided by the embodiments of the present disclosure takes the legal person loan risk prediction based on multi-perspective boundary discrimination constraints as an example. Its samples consist of three perspectives, namely the basic information perspective, the business information perspective, and the behavior information perspective. The labels of the samples are 1 and -1, representing risky users and general users respectively. For the training samples, first, the features are divided into three subsets according to the feature categories, namely basic information, business information, and behavior information, corresponding to the three perspectives of the samples respectively. During the model training process, boundary sample sets are selected through the heterogeneous nearest neighbor information of the samples within the perspective, and the means of the two types of boundary sample sets are calculated through the boundary sample sets. Through the boundary sample discrimination constraint term designed in this patent, the same-class boundary samples are made as close as possible in the output space, and the different-class boundary samples are made as far away as possible in the output space, so that the classification hyperplane passes through the middle area between the two types of boundary samples as much as possible to improve the generalization performance of the model. Between different perspectives, the consistency of boundary samples in different output spaces is mined. Through the inter-perspective boundary sample discrimination constraint term designed in this patent, the means of the same-class boundary samples are made as consistent as possible in the outputs of different perspectives, aiming to optimize each other among multiple perspectives to improve the accuracy of the classification boundary. By minimizing the empirical loss of the model, the boundary sample discrimination constraint term, and the inter-perspective boundary sample discrimination constraint term, three sub-classifiers are obtained. Finally, the results of the three sub-classifiers are integrated to perform classification prediction on the test samples. Through optimizing the learning process of the model within and between perspectives, the same-class boundary samples are made as close as possible in the output space, the different-class boundary samples are made as far away as possible in the output space, and the classification hyperplane passes through the middle area between the two types of boundary samples as much as possible. The final model integrating the results of the three sub-classifiers can improve the generalization ability of the model.

[0127] The embodiments of the present disclosure have better effects than traditional machine learning algorithms in terms of the precision, recall rate, and comprehensive evaluation value of risk prediction (such as legal person loan risk prediction) classification, and can predict business risk situations more accurately. For example, applying this risk prediction model to financial institutions such as banks for accurate prediction before user loans, the customer manager can refer to the model prediction results for corresponding processing, reduce the issuance of non-performing loans, reduce losses, and enhance the competitiveness of the institution in the same industry.

[0128] The embodiments of the present disclosure also provide a risk prediction device.

[0129] Figure 9 A block diagram of the risk prediction device according to the embodiments of the present disclosure is schematically shown.

[0130] As Figure 9 shown, the risk prediction device 900 may include: a data acquisition module 910, a risk feature acquisition module 920, and a risk feature processing module 930.

[0131] The data acquisition module 910 is used to acquire the data to be predicted.

[0132] The risk feature acquisition module 920 is used to acquire the risk features of the data to be predicted.

[0133] The risk feature processing module 930 is used to process the risk features by using the trained risk prediction model to obtain a risk prediction result, where the risk prediction result includes the category to which the data to be predicted belongs.

[0134] Among them, the sample data of each category includes boundary sample data, and the objective function of the risk prediction model includes a boundary sample discrimination constraint term, and the boundary sample discrimination constraint term makes the first distance of the boundary sample data belonging to different categories in the output space greater than the second distance of the boundary sample data belonging to the same category in the output space.

[0135] It should be noted that the implementation manners, the technical problems solved, the functions achieved, and the technical effects achieved by each module / unit and the like in the device partial embodiments are respectively the same as or similar to those of the corresponding steps in the method partial embodiments, and will not be elaborated herein one by one.

[0136] According to the embodiments of the present disclosure, any plurality of modules or units, or at least part of the functions of any plurality of them can be implemented in one module. Any one or more of the modules or units according to the embodiments of the present disclosure can be split into multiple modules for implementation. Any one or more of the modules or units according to the embodiments of the present disclosure can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable manner of integrating or packaging the circuit in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware or in an appropriate combination of any several of them. Alternatively, one or more of the modules or units according to the embodiments of the present disclosure can be at least partially implemented as a computer program module, and when the computer program module runs, the corresponding functions can be executed.

[0137] For example, any combination of the data acquisition module 910, the risk feature acquisition module 920, and the risk feature processing module 930 can be integrated into one module, or any one of them can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the data acquisition module 910, the risk feature acquisition module 920, and the risk feature processing module 930 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or any other reasonable manner of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any suitable combination of several of them. Alternatively, at least one of the data acquisition module 910, the risk feature acquisition module 920, and the risk feature processing module 930 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0138] Figure 10 Schematically shows a block diagram of an electronic device according to an embodiment of the present disclosure. Figure 10 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0139] As Figure 10 shown, the electronic device 1000 according to an embodiment of the present disclosure includes a processor 1001, which can perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 1002 or a program loaded from a storage section 1008 into a random access memory (RAM) 1003. The processor 1001 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 1001 can also include on-board memory for caching purposes. The processor 1001 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure. The multiple processing units can be integrated in one processor or distributed in multiple processors, which is not limited herein.

[0140] In the RAM 1003, various programs and data required for the operation of the electronic device 1000 are stored. The processor 1001, the ROM 1002, and the RAM 1003 are communicatively connected to each other via the bus 1004. The processor 1001 performs various operations of the method flow according to the embodiments of the present disclosure by executing programs in the ROM 1002 and / or the RAM 1003. It should be noted that the programs can also be stored in one or more memories other than the ROM 1002 and the RAM 1003. The processor 1001 can also perform various operations of the method flow according to the embodiments of the present disclosure by executing programs stored in one or more memories.

[0141] According to an embodiment of the present disclosure, the electronic device 1000 may further include an input / output (I / O) interface 1005, and the input / output (I / O) interface 1005 is also connected to the bus 1004. The electronic device 1000 may further include one or more of the following components connected to the I / O interface 1005: an input portion 1006 including a keyboard, a mouse, etc.; an output portion 1007 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 1008 including a hard disk, etc.; and a communication portion 1009 including a network interface card such as a LAN card, a modem, etc. The communication portion 1009 performs communication processing via a network such as the Internet. The drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 1010 as needed so that a computer program read therefrom can be installed into the storage portion 1008 as needed.

[0142] According to an embodiment of the present disclosure, the method flow according to the embodiments of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication portion 1009, and / or installed from the removable medium 1011. When the computer program is executed by the processor 1001, the above functions defined in the system according to the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described system, device, apparatus, module, unit, etc. can be implemented by computer program modules.

[0143] The present disclosure also provides a computer-readable storage medium.

[0144] Reference Figure 10As shown, the computer-readable storage medium may be included in the device / device / system described in the above embodiments; or it may exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.

[0145] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, device, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the ROM 1002 and / or RAM 1003 described above and / or one or more memories other than the ROM 1002 and RAM 1003.

[0146] Embodiments of the present disclosure also include a computer program product, which includes a computer program that contains program code for executing the method provided by the embodiments of the present disclosure. When the computer program product runs on an electronic device, the program code is used to cause the electronic device to implement the image model training method or risk prediction method provided by the embodiments of the present disclosure.

[0147] When the computer program is executed by the processor 1001, the above functions defined in the system / device of the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, devices, modules, units, etc. may be implemented by computer program modules.

[0148] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and is downloaded and installed through the communication part 1009, and / or installed from the removable medium 1011. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0149] According to embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, programming languages such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0150] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. These embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the various embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.

Claims

1. A risk prediction method, comprising: Obtain the data to be predicted; Obtaining risk characteristics of the data to be predicted; as well as Processing the risk features using a trained risk prediction model to obtain a risk prediction result, wherein the risk prediction result includes the category to which the data to be predicted belongs; The sample data of each category includes boundary sample data, and the objective function of the risk prediction model includes a boundary sample discrimination constraint item, wherein the boundary sample discrimination constraint item makes the first distance of the boundary sample data belonging to different categories in the output space greater than the second distance of the boundary sample data belonging to the same category in the output space. The objective function of the risk prediction model also includes an inter-view boundary sample discrimination constraint item. The calculation formula of the boundary sample discrimination constraint term includes: Among them, is the negative class boundary sample set, is the positive class boundary sample set, , are the means of the two types of boundary sample sets, is the classifier, The calculation formula for the boundary sample discrimination constraint term between view angles includes: Among them, , is the mean of two types of boundary sample sets, and are perspective serial numbers respectively, is a classifier.

2. The method according to claim 1, wherein The categories include positive sample categories and negative sample categories; as well as The boundary sample discrimination constraint item includes: a positive sample sub-constraint item, a negative sample sub-constraint item and a cross item, wherein the output of the positive sample sub-constraint item is related to the difference between the sub-result of the model processing the positive sample data and the sub-result of the model processing the mean of the boundary positive sample subset, the output of the negative sample sub-constraint item is related to the difference between the sub-result of the model processing the negative sample data and the sub-result of the model processing the mean of the boundary negative sample subset, and the output of the cross item is related to the product of the sub-result of the model processing the positive sample data and the sub-result of the model processing the negative sample data.

3. The method according to claim 2, wherein, The obtaining of risk characteristics of the data to be predicted includes: obtaining at least one of sub-features of the data to be predicted from a basic information perspective, obtaining sub-features of the data to be predicted from an operating information perspective, and obtaining sub-features of the data to be predicted from a behavioral information perspective.

4. The method according to claim 3, wherein, The boundary sample discrimination constraint item includes at least one of the following: Positive sample sub-constraints, negative sample sub-constraints, and cross-terms from the perspective of basic information; Positive sample sub-constraints, negative sample sub-constraints, and cross-terms from the perspective of business information; or Positive sample sub-constraints, negative sample sub-constraints, and cross terms from the perspective of behavioral information.

5. The method according to claim 3, wherein the boundary sample discrimination constraint between views makes the outputs of the mean values of boundary samples of the same category in different view angles tend to be consistent.

6. The method according to claim 5, wherein, The inter-view boundary sample discrimination constraint item includes at least one of the following: Positive sample constraints for different viewpoints; or Negative sample sub-constraints for different viewpoints.

7. The method according to claim 1, wherein, Training the risk prediction model includes: Acquire a training sample data set, where the training sample data set includes a positive training sample data subset and / or a negative training sample data subset; Determining a subset of boundary positive samples in the positive training sample data subset based on the heterogeneous neighbor information of the samples, and / or determining a subset of boundary negative samples in the negative training sample data subset based on the heterogeneous neighbor information of the samples; and Input the positive training sample data subset and / or the negative training sample data subset into the risk prediction model, and adjust the parameters of the risk prediction model until the preset number of iterations is reached or the difference in the loss function of the objective function between two iterations is less than the preset threshold.

8. The method according to claim 7, wherein: Determining the boundary positive sample subset in the positive training sample data subset based on the heterogeneous neighbor information of the samples includes: for any negative sample, adding the first specified number of positive class samples adjacent to the negative sample to the boundary positive sample subset; and Determining the boundary negative sample subset in the negative training sample data subset based on the heterogeneous neighbor information of the samples includes: for any positive sample, adding the second specified number of negative class samples adjacent to the positive sample to the boundary negative sample subset.

9. The method according to claim 7, further comprising: Testing the trained risk prediction model using the test sample set to obtain the test accuracy of the risk prediction result.

10. The method according to any one of claims 1 to 9, wherein, Obtaining the risk features of the data to be predicted includes at least one of the following: Performing one-hot encoding on the categorical data in the data to be predicted to obtain categorical features; or Calculating the associated data of the business information and / or behavioral information in the data to be predicted to obtain derived features.

11. According to the method described in any one of claims 1 to 9, wherein, The objective function further includes an empirical loss constraint term and a regularization constraint term.

12. The method according to any one of claims 1 to 9, wherein The sample data of each category further includes non-boundary sample data, and the third distance between the boundary sample data and the class center in the same category is greater than the fourth distance between the non-boundary sample data and the class center.

13. A risk prediction device, comprising: A data acquisition module for acquiring data to be predicted; A risk feature acquisition module for acquiring the risk features of the data to be predicted; And A risk feature processing module for processing the risk features using the trained risk prediction model to obtain a risk prediction result, where the risk prediction result includes the category to which the data to be predicted belongs; wherein the sample data of each category includes boundary sample data, the objective function of the risk prediction model includes a boundary sample discrimination constraint term, the boundary sample discrimination constraint term makes the first distance between the boundary sample data belonging to different categories in the output space greater than the second distance between the boundary sample data belonging to the same category in the output space, and the objective function of the risk prediction model further includes an inter-perspective boundary sample discrimination constraint term, The calculation formula of the boundary sample discrimination constraint term includes: Among them, is the negative class boundary sample set, is the positive class boundary sample set, , are the means of the two types of boundary sample sets, is the classifier, The calculation formula of the inter-perspective boundary sample discrimination constraint term includes: Among them, , is the mean of two types of boundary sample sets, , are perspective serial numbers respectively, is a classifier.

14. An electronic device, comprising: One or more processors; A storage device for storing executable instructions, which when executed by the processor, implement the risk prediction method according to any one of claims 1 to 12.

15. A computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, implement the risk prediction method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Method for constructing neural network with boundary condition constraint

    CN106203618A

  • Semi-supervised classification method of modified clustering assumption combined with pairwise constraints

    CN108038511A