A training method, device and equipment for an illegal fund-raising risk prediction model
By using point-state manifold regularization terms in the training process of the illegal fundraising risk prediction model, the problem of poor learning effect of prediction model in the existing technology is solved, and the accuracy of prediction of the risk of illegal fundraising in enterprises is improved.
Patent Information
- Application Number
- CN202110551911.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-20
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-05-20
AI Technical Summary
The illegal fundraising risk prediction model trained by semi-supervised machine learning methods in the prior art cannot achieve good learning results, resulting in low accuracy of predicting illegal fundraising risk in enterprises.
By obtaining a training data set containing label samples and unlabeled samples, the local density value and cluster membership of each sample are calculated, and the point-state manifold regularization constraint term is constructed to constrain the smooth relationship between each sample and its nearest neighbor samples, thereby determining the loss function of the preset classifier and training, and obtaining an illegal fundraising risk prediction model.
It improves the generalization effect and accuracy of the model and enhances the accuracy of predicting the risk of illegal fundraising in enterprises.
Smart Images

Figure CN113516550B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of machine learning, and particularly relates to a method, device, and equipment for training an illegal fund-raising risk prediction model. Background Art
[0002] With the continuous development of technology, illegal fund-raising has seriously disrupted the normal economic and financial order, causing economic losses to participants and even plunging them into difficult living situations. It is extremely likely to trigger social instability and a large number of social security problems, and even lead to social unrest in some local areas. Therefore, how to establish a prediction model based on a large amount of enterprise information and determine whether an enterprise has the risk of illegal fund-raising is of great importance to regulatory authorities, enterprise partners, and investors.
[0003] In the prior art, when using a semi-supervised machine learning method to train a prediction model to predict whether an enterprise has the risk of illegal fund-raising, it is mainly based on the idea of pairwise constraints to preserve the smoothness between any two sample points. However, smoothness can be pointwise in nature, that is, smoothness can occur "anywhere", not just between two points. Therefore, the prediction model trained in this way cannot achieve a good learning effect, resulting in a low accuracy in predicting the illegal fund-raising risk of enterprises.
[0004] Therefore, there is an urgent need in the industry for a technical solution that can solve the above technical problems. Summary of the Invention
[0005] Embodiments of this specification provide a method, device, and equipment for training an illegal fund-raising risk prediction model, which can improve the accuracy of predicting the illegal fund-raising risk of enterprises.
[0006] A method, device, and equipment for training an illegal fund-raising risk prediction model provided in this specification are implemented in the following manner.
[0007] A method for training an illegal fund-raising risk prediction model includes: obtaining a training data set related to the enterprise's fund-raising risk; wherein, the training data set includes labeled samples and unlabeled samples; the labels of the labeled samples are determined according to whether there is an illegal fund-raising risk; calculating the local density value and clustering membership degree of each sample in the training data set; constructing a pointwise manifold regularization constraint term according to the local density and clustering membership degree of each sample and the discriminant function of a preset classifier; wherein, the pointwise manifold regularization constraint term is used to constrain the relationship between each sample and its neighboring samples; determining the loss function of the preset classifier based on the pointwise manifold regularization constraint term; training the preset classifier based on the training data set and the loss function to obtain an illegal fund-raising risk prediction model.
[0008] A training device for an illegal fund-raising risk prediction model, comprising: an acquisition module for acquiring a training data set related to the enterprise fund-raising risk; wherein, the training data set includes labeled samples and unlabeled samples; the labels of the labeled samples are determined according to whether there is illegal fund-raising risk; a calculation module for calculating the local density value and clustering membership degree of each sample in the training data set; a construction module for constructing a point-wise manifold regularization constraint term according to the local density and clustering membership degree of each sample and the discriminant function of a preset classifier; wherein, the point-wise manifold regularization constraint term is used to constrain the relationship between each sample and its neighboring samples; a determination module for determining the loss function of the preset classifier based on the point-wise manifold regularization constraint term; an obtaining module for training the preset classifier based on the training data set and the loss function to obtain an illegal fund-raising risk prediction model.
[0009] A training device for an illegal fund-raising risk prediction model, comprising at least one processor and a memory storing computer-executable instructions, and when the processor executes the instructions, the steps of any one of the method embodiments in the embodiments of the present specification are implemented.
[0010] A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed, the steps of any one of the method embodiments in the embodiments of the present specification are implemented.
[0011] A training method, device and device for an illegal fund-raising risk prediction model provided in the present specification. In some embodiments, a training data set related to the enterprise fund-raising risk can be acquired; wherein, the training data set includes labeled samples and unlabeled samples; the labels of the labeled samples are determined according to whether there is illegal fund-raising risk; the local density value and clustering membership degree of each sample in the training data set are calculated. It is also possible to construct a point-wise manifold regularization constraint term according to the local density and clustering membership degree of each sample and the discriminant function of a preset classifier; determine the loss function of the preset classifier based on the point-wise manifold regularization constraint term; train the preset classifier based on the training data set and the loss function to obtain an illegal fund-raising risk prediction model. Since in order to fully learn the spatial distribution information of the labeled samples and unlabeled samples, a point-wise manifold regularization term that constrains the smooth relationship between each sample and its neighboring samples is constructed in the loss function of the preset classifier, thereby improving the generalization effect and accuracy of the model and improving the accuracy of predicting the illegal fund-raising risk of enterprises. Description of the Drawings
[0012] The drawings described herein are used to provide a further understanding of the present specification, form a part of the present specification, and do not constitute a limitation to the present specification. In the drawings:
[0013] Figure 1It is a schematic flowchart of an embodiment of a method for training an illegal fund-raising risk prediction model provided in this specification;
[0014] Figure 2 It is a schematic module structure diagram of an embodiment of a device for training an illegal fund-raising risk prediction model provided in this specification;
[0015] Figure 3 It is a schematic hardware structure block diagram of an embodiment of a server for training an illegal fund-raising risk prediction model provided in this specification. Detailed implementation manners
[0016] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only some of the embodiments in this specification, rather than all the embodiments. Based on one or more embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the embodiments of this specification.
[0017] The implementation solutions of this specification will be described below by taking a specific application scenario as an example. Specifically, Figure 1 It is a schematic flowchart of an embodiment of a method for training an illegal fund-raising risk prediction model provided in this specification. Although the method operation steps or device structures shown in the following embodiments or drawings are provided in this specification, based on routine or non-creative labor, more or fewer operation steps or module units may be included in the method or device after partial combination.
[0018] An implementation solution provided in this specification can be applied to a client, a server, etc. The client may include terminal devices, such as smart phones, tablet computers, etc. The server may include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed system, etc.
[0019] It should be noted that the following embodiments do not limit the technical solutions in other extensible application scenarios based on this specification. A specific example is Figure 1 As shown, in an embodiment of a method for training an illegal fund-raising risk prediction model provided in this specification, the method may include the following steps.
[0020] S0: Obtain a training data set related to the enterprise's fund-raising risk; wherein, the training data set includes labeled samples and unlabeled samples; the labels of the labeled samples are determined according to whether there is an illegal fund-raising risk.
[0021] In some embodiments, before obtaining the training data set related to the risk of corporate illegal fund-raising, it may include: obtaining feature data related to the risk of corporate fund-raising; wherein, the feature data includes labeled feature data and unlabeled feature data; preprocessing the feature data to obtain a training data set related to the risk of corporate fund-raising; wherein, the preprocessing includes missing value processing and feature engineering processing.
[0022] In some implementation scenarios, feature data related to the risk of corporate illegal fund-raising can be obtained from a data warehouse. Among them, the data warehouse can pre-store data such as basic information of enterprises, annual reports of enterprises, and tax payment situations of enterprises. In some implementation scenarios, the data in the data warehouse can include data of various data types such as numerical type, character type, and date type.
[0023] In some implementation scenarios, the feature data related to the risk of corporate illegal fund-raising that can be obtained from the data warehouse can be divided into five categories. For example, it can be divided into basic information of the enterprise, basic annual report information of the enterprise, tax information of the enterprise, change information of the enterprise, and news public opinion information. Of course, the above is only an exemplary description, and the feature data is not limited to the above examples. Those skilled in the art may make other changes under the inspiration of the technical essence of this application, but as long as the functions and effects achieved are the same or similar to those of this application, they should all be covered within the protection scope of this application.
[0024] In some implementation scenarios, after determining the category of the feature data, the data range can be determined according to the category and the corresponding data can be obtained. In some implementation scenarios, after obtaining the feature data, the data can be stored in a data table. Among them, each row in the data table can represent a sample, and each column can represent the data corresponding to a feature. Of course, the above is only an exemplary description, and this specification does not limit this.
[0025] In some implementation scenarios, after storing the feature data in the data table, labels can be artificially assigned to a part of the samples. In some implementation scenarios, for labeled samples, the sample label of those with the risk of illegal fund-raising can be set to 1, representing the first type of sample ω 1 , and the sample label of those without the risk of illegal fund-raising can be set to 0, representing the first type of sample ω 2 ; for unlabeled samples, no label needs to be assigned. Of course, the above is only an exemplary description, and the way of assigning labels to samples is not limited to the above examples. Those skilled in the art may make other changes under the inspiration of the technical essence of this application, but as long as the functions and effects achieved are the same or similar to those of this application, they should all be covered within the protection scope of this application.
[0026] In some implementation scenarios, after storing the feature data in a data table, the samples in the data table can be preprocessed to obtain a training data set related to the enterprise financing risk.
[0027] In some implementation scenarios, preprocessing the samples in the data table may include missing value processing, feature engineering processing, etc.
[0028] For example, in some implementation scenarios, for columns with missing values, they can be filled in a certain way. For example, the missing values of numerical features can be filled with the column "0" value, and the missing values of non-numerical features can be filled with "unknown". For columns with particularly severe missing values, the field (feature) can be directly deleted. Of course, the above is only an illustrative description, and this specification does not limit it.
[0029] For example, in some implementation scenarios, feature engineering processing may include feature combination, feature expansion, etc. For example, statistical information (maximum value, minimum value, mean value, variance, etc.) of numerical features is grouped and counted according to categorical features, deviation value features of numerical features (the difference between the original feature and the minimum value, maximum value, mean value of this column, etc.), cross features between numerical features (new columns obtained by performing relevant addition, subtraction, multiplication, and division operations between numerical features), etc. In this way, by evolving the features, more features can be derived, thereby improving the accuracy of the subsequent training model. Of course, the above is only an illustrative description, and this specification does not limit it.
[0030] In some implementation scenarios, the preprocessed data can be used as a training data set related to the enterprise financing risk. Among them, the training data set may include labeled samples and unlabeled samples, and the labels of the labeled samples can be determined according to whether there is an illegal fund-raising risk.
[0031] S2: Calculate the local density value and clustering membership degree of each sample in the training data set.
[0032] In the embodiments of this specification, after obtaining the training data set related to the enterprise financing risk, the local density value and clustering membership degree of each sample in the training data set can be calculated.
[0033] In some embodiments, calculating the local density value and clustering membership degree of each sample in the training dataset may include: determining the set of nearest neighbor samples of the first sample; calculating the sum of distances between the first sample and each sample in the corresponding set of nearest neighbor samples to obtain the sum of nearest neighbor distances of the first sample; determining the total sum of nearest neighbor distances of all samples in the training dataset based on the sum of nearest neighbor distances of the first sample; determining the local density value of the first sample based on the sum of nearest neighbor distances of the first sample and the total sum of nearest neighbor distances of all samples; clustering the samples in the training dataset to obtain the first cluster and the second cluster; calculating the first probability that the first sample belongs to the first cluster and the second probability that the first sample belongs to the second cluster; and taking the maximum value of the first probability and the second probability as the clustering membership degree of the first sample.
[0034] In some implementation scenarios, the set of nearest neighbor samples of each sample in the training dataset can be determined according to the K-nearest neighbor algorithm. Among them, the set of nearest neighbor samples may include a preset number of nearest neighbor samples, such as 3, 5, 7, etc. The number of nearest neighbor samples included in the set of nearest neighbor samples corresponding to each sample can be the same or different, and this specification does not limit this. Of course, the above is only an illustrative description, and the method for determining the set of nearest neighbor samples of each sample is not limited to the above examples. Those skilled in the art may make other changes under the inspiration of the technical essence of this application, but as long as the functions and effects achieved are the same or similar to those of this application, they should all be covered within the protection scope of this application.
[0035] In some implementation scenarios, the local density value of the first sample can be determined according to the following formula:
[0036]
[0037] where d(x i ) represents the local density value of sample x i , N(x i ) represents the set of nearest neighbor samples of sample x i , x j represents a sample in the set of nearest neighbor samples of sample x i , d(x i , x j ) represents the distance between sample x i and sample x j , N(x s ) represents the set of nearest neighbor samples of sample x s , x t represents a sample in the set of nearest neighbor samples of sample x s , d(x s , x t ) represents the distance between sample x s and sample x tThe distance between them, n represents the total number of samples, and s represents the serial number.
[0038] In some implementation scenarios, the larger the value of d(x i ), it can indicate that the distribution around the sample is denser, and the possibility that the sample x i belongs to a noise sample is smaller. At this time, a relatively large weight should be assigned to the sample x i .
[0039] In some implementation scenarios, the clustering membership degree of the first sample can be determined according to the following formula:
[0040] u(x i ) = max(u 1i , u 2i )
[0041] Among them, u(x i ) represents the clustering membership degree of the sample x i , u 1i represents the probability that the sample x i belongs to the first cluster, and u 2i represents the probability that the sample x i belongs to the second cluster.
[0042] In some implementation scenarios, the above u 1i and u 2i can be calculated through an unsupervised learning method (such as fuzzy C-means clustering).
[0043] In some implementation scenarios, the larger the clustering membership degree of the sample x i , it can indicate that the possibility that x i is a non-boundary sample is greater, and the certainty of its classification should also be greater. At this time, a relatively large weight can be assigned to the sample x i .
[0044] S4: Construct a pointwise manifold regularization constraint term according to the local density and clustering membership degree of each sample and the discriminant function of the preset classifier; wherein, the pointwise manifold regularization constraint term is used to constrain the relationship between each sample and its neighboring samples.
[0045] In the embodiments of this specification, after obtaining the local density value and clustering membership of each sample in the training dataset, a pointwise manifold regularization constraint term can be constructed according to the local density and clustering membership of each sample and the discriminant function of a preset classifier. Among them, the pointwise manifold regularization constraint term can be used to constrain the relationship between each sample and its neighboring samples. Among them, the relationship between each sample and its neighboring samples can be a smooth relationship. The smooth relationship between each sample and its neighboring samples can be understood as that the relationship between the training samples and the neighboring samples in the output space can be retained in the original feature space. For example, samples that are close in the feature space are also close in the output space. The output space can be understood as the output result of the samples in the model (for example, if a sample outputs one result, it is a one-dimensional output space), and the original feature space can be understood as the features of the samples (for example, if there are m features, it is an m-dimensional feature space).
[0046] In some embodiments, the constructing a pointwise manifold regularization constraint term according to the local density and clustering membership of each sample and the discriminant function of a preset classifier may include: obtaining the sample weight of each sample according to the local density and clustering membership of each sample; and constructing a pointwise manifold regularization constraint term based on the sample weights of all samples and the discriminant function of the preset classifier. Among them, the preset classifier can be a semi-supervised classifier.
[0047] In some implementation scenarios, after obtaining the local density value and clustering membership of each sample, the sample weight of each sample can be calculated according to the following formula:
[0048]
[0049] where p(x i ) represents the weight of sample x i , N(x i ) represents the set of neighboring samples of sample x i , x j represents a sample in the set of neighboring samples of sample x i , d(x i , x j ) represents the distance between sample x i and sample x j , N(x s ) represents the set of neighboring samples of sample x s , x t represents a sample in the set of neighboring samples of sample x s , d(x s , x t ) represents the distance between sample x s and sample x t , n represents the total number of samples, s represents the serial number, u 1iDenote the sample as \(x\). i The probability that \(x\) belongs to the first cluster, \(u\). 2i Denote the sample as \(x\). i The probability that \(x\) belongs to the second cluster.
[0050] In the embodiments of this specification, by comprehensively considering the local density value of the sample and the clustering membership of the sample, a sample weight based on the local density value and the clustering membership is designed, which can not only reduce the influence of noise samples, but also ensure the importance of non-boundary samples.
[0051] In some implementation scenarios, after obtaining the sample weight of each sample, in order to fully learn the spatial distribution information of the labeled samples and unlabeled samples, a pointwise manifold regularization term can be constructed by constraining the smooth relationship between each sample and its neighboring samples. Among them, the spatial distribution information can be understood as the information contained in the features, such as the similarity relationship and distance relationship in the feature space.
[0052] In some implementation scenarios, the pointwise manifold regularization constraint term constructed based on the sample weights of all samples and the discriminant function of the preset classifier can be:
[0053]
[0054] where \(R\) pm represents the pointwise manifold regularization constraint term, \(n\) represents the total number of samples, \(i\) represents the serial number, \(p(x\) i ) represents the weight of the sample \(x\) i , \(f(·)\) is the discriminant function of the preset classifier, \(x\) i represents the \(i\)-th sample, \(N(x\) i ) represents the set of neighboring samples of the sample \(x\) i , \(x\) j represents a sample in the set of neighboring samples of the sample \(x\) i , \(w\) i,j represents the similarity between the sample \(x\) i and the sample \(x\) j .
[0055] In the embodiments of this specification, by constructing a pointwise manifold regularization term by constraining the smooth relationship between each sample and its neighboring samples, the training samples can retain the smooth relationship in the original feature space between the output space and the neighboring samples, thereby improving the generalization effect of the model (preset classifier).
[0056] S6: Determine the loss function of the preset classifier based on the pointwise manifold regularization constraint term.
[0057] In the embodiments of this specification, after constructing the pointwise manifold regularization constraint term, the loss function of the preset classifier can be determined based on the pointwise manifold regularization constraint term.
[0058] In some embodiments, determining the loss function of the preset classifier based on the pointwise manifold regularization constraint term may include: determining the loss function of the preset classifier according to the following formula:
[0059] L = R emp + αR pm + βR reg
[0060] where L represents the loss function of the preset classifier, and R emp represents the minimized empirical loss function constructed based on the labeled samples, and R pm represents the pointwise manifold regularization constraint term, and R reg represents the l 2 regularization loss function, and α and β represent hyperparameters used to adjust the weights of each term.
[0061] In some implementation scenarios, the minimized empirical loss function R emp can include various types. For example, the squared loss function the absolute loss function and so on. Among them, f(x) represents the predicted value, and y represents the true value. The l 2 regularization loss function R reg can be used to prevent overfitting, and it can also include various types. For example, R reg = ||w|| 2 , where w is the parameter of f(·).
[0062] S8: Training the preset classifier based on the training data set and the loss function to obtain an illegal fund-raising risk prediction model.
[0063] In the embodiments of this specification, after determining the loss function of the preset classifier, the preset classifier can be trained based on the training data set and the loss function to obtain an illegal fund-raising risk prediction model. Among them, the illegal fund-raising risk prediction model can be used to determine the probability that an enterprise has an illegal fund-raising risk.
[0064] Since the principle of training the preset classifier using the training samples is to solve an optimization problem, and this optimization problem is to minimize the value of the loss function, and then the model corresponding to the minimum result of the loss function is used as the optimal model, that is, the illegal fund-raising risk prediction model.
[0065] In some implementation scenarios, the loss function obtained by using the gradient descent method can be utilized. By minimizing the loss function of a preset classifier until the preset number of iterations is reached or the difference between the loss values of two loss functions is less than a preset threshold, the training is stopped, and a non - public fund - raising risk prediction model f(·) is obtained. Among them, the non - public fund - raising risk prediction model f(·) can also be referred to as the discriminant function of the preset classifier.
[0066] In some embodiments, after obtaining the non - public fund - raising risk prediction model, it may further include: obtaining feature data related to the fund - raising risk in the target enterprise; using the non - public fund - raising risk prediction model to process the feature data to obtain the probability of the non - public fund - raising risk of the target enterprise; and determining whether the target enterprise has a non - public fund - raising risk based on the relationship between the probability of the non - public fund - raising risk of the target enterprise and a preset risk value.
[0067] In some implementation scenarios, after obtaining the non - public fund - raising risk prediction model, feature data related to the fund - raising risk in the target enterprise can be obtained, and the feature data is input into the non - public fund - raising risk prediction model to obtain the probability of the non - public fund - raising risk of the target enterprise. Then, it is judged whether the probability is greater than or equal to a preset risk value. When it is determined that it is greater than or equal to, it can be stated that the target enterprise has a non - public fund - raising risk, that is, it belongs to the first - type sample ω in the training sample set. 1 In some implementation scenarios, when the probability of the non - public fund - raising risk of the target enterprise is less than the preset risk value, it can be stated that the target enterprise does not have a non - public fund - raising risk, that is, it belongs to the first - type sample ω in the training sample set. 2 Among them, the preset risk value can be set according to the actual scenario. For example, it can be 0.5, 0.6, etc.
[0068] In the embodiments of this specification, by using a point - wise manifold regularization term in the process of training the sample model, the learning effect of the model can be optimized, so that the accuracy of the obtained model is higher.
[0069] In the embodiments of this specification, after obtaining the non - public fund - raising risk prediction model, by applying it and the traditional semi - supervised learning algorithm to the actual scenario at the same time, it can be known that the solution of this application is better than the traditional semi - supervised learning algorithm in terms of the precision rate, recall rate, and comprehensive evaluation value of the non - public fund - raising risk prediction classification of enterprises, and can more accurately predict the non - public fund - raising risk of enterprises.
[0070] In the embodiments of this specification, after obtaining the non - public fund - raising risk prediction model, the model can be applied to financial institutions such as banks.
[0071] In the embodiments of this specification, first, taking advantage of the high cost of obtaining sample labels in the enterprise illegal fund-raising risk prediction scenario, a small number of labeled samples and a large number of unlabeled samples are constructed for modeling and learning. Second, in order to fully learn the spatial distribution information of the labeled samples and the unlabeled samples, a pointwise manifold regularization term based on semi-supervised learning is developed by constraining the smooth relationship between each sample and its neighboring samples, so that the training samples can retain the smooth relationship in the original feature space between the output space and the neighboring samples, thereby improving the generalization effect of the model. In addition, the importance of the sample is described by the local density value and the clustering membership degree of a single sample. Samples with relatively higher local density values are assigned relatively larger weights, and samples with relatively lower local density values are assigned relatively smaller weights, which can prevent the learning process of the model from being affected when there are abnormal samples. At the same time, according to the clustering results, samples with higher clustering membership degrees are given higher weights, which can improve the accuracy of the model. The classifier is iteratively optimized by minimizing the empirical loss and dynamic pairwise constraints.
[0072] Of course, the above is only an exemplary illustration. The embodiments of this specification are not limited to the above examples. Those skilled in the art may make other changes under the inspiration of the technical essence of this application. However, as long as the functions and effects achieved are the same or similar to those of this application, they should all be covered within the protection scope of this application.
[0073] From the above description, it can be seen that the embodiments of this application can obtain a training data set related to the enterprise fund-raising risk; among them, the training data set includes labeled samples and unlabeled samples; the labels of the labeled samples are determined according to whether there is an illegal fund-raising risk; calculate the local density value and the clustering membership degree of each sample in the training data set. It is also possible to construct a pointwise manifold regularization constraint term according to the local density and clustering membership degree of each sample and the discriminant function of the preset classifier; based on the pointwise manifold regularization constraint term, determine the loss function of the preset classifier; train the preset classifier based on the training data set and the loss function to obtain an illegal fund-raising risk prediction model. Since, in order to fully learn the spatial distribution information of the labeled samples and the unlabeled samples, a pointwise manifold regularization term that constrains the smooth relationship between each sample and its neighboring samples is constructed in the loss function of the preset classifier, the generalization effect and accuracy of the model can be improved, and the accuracy of predicting the enterprise illegal fund-raising risk can be improved.
[0074] The various embodiments of the above methods in this specification are all described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. For the relevant parts, refer to the partial description of the method embodiments.
[0075] Based on the above-mentioned training method of an illegal fund-raising risk prediction model, one or more embodiments of this specification also provide a training device for an illegal fund-raising risk prediction model. The described device may include a system (including a distributed system), software (application), module, component, server, client, etc. that uses the method described in the embodiments of this specification and combines the necessary implementation hardware. Based on the same innovative concept, the devices in one or more embodiments provided by the embodiments of this specification are as described in the following embodiments. Since the implementation solutions for the device to solve problems are similar to those of the method, the implementation of the specific device in the embodiments of this specification can refer to the implementation of the foregoing method, and the repeated parts will not be elaborated here. As used hereinafter, the term "unit" or "module" may be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0076] Specifically, Figure 2 is a schematic diagram of the module structure of an embodiment of a training device for an illegal fund-raising risk prediction model provided by this specification. As Figure 2 shown, a training device for an illegal fund-raising risk prediction model provided by this specification may include: an acquisition module 120, a calculation module 122, a construction module 124, a determination module 126, and an acquisition module 128.
[0077] The acquisition module 120 can be used to acquire a training data set related to the enterprise's fund-raising risk; wherein, the training data set includes labeled samples and unlabeled samples; the labels of the labeled samples are determined according to whether there is an illegal fund-raising risk.
[0078] The calculation module 122 can be used to calculate the local density value and clustering membership degree of each sample in the training data set.
[0079] The construction module 124 can be used to construct a point-wise manifold regularization constraint term according to the local density and clustering membership degree of each sample and the discriminant function of a preset classifier; wherein, the point-wise manifold regularization constraint term is used to constrain the relationship between each sample and its neighboring samples.
[0080] The determination module 126 can be used to determine the loss function of the preset classifier based on the point-wise manifold regularization constraint term.
[0081] The acquisition module 128 can be used to train the preset classifier based on the training data set and the loss function to obtain an illegal fund-raising risk prediction model. It should be noted that the above-mentioned device may also include other implementation manners according to the description of the method embodiments, and the specific implementation manners may refer to the description of the relevant method embodiments and will not be elaborated here one by one.
[0082] This specification also provides an embodiment of a training device for an illegal fund-raising risk prediction model, including a processor and a memory for storing processor-executable instructions. When the instructions are executed by the processor, the following steps are implemented: obtaining a training data set related to the enterprise's fund-raising risk; wherein, the training data set includes labeled samples and unlabeled samples; the labels of the labeled samples are determined according to whether there is an illegal fund-raising risk; calculating the local density value and clustering membership degree of each sample in the training data set; constructing a point-wise manifold regularization constraint term according to the local density and clustering membership degree of each sample and the discriminant function of a preset classifier; wherein, the point-wise manifold regularization constraint term is used to constrain the relationship between each sample and its neighboring samples; determining the loss function of the preset classifier based on the point-wise manifold regularization constraint term; training the preset classifier based on the training data set and the loss function to obtain an illegal fund-raising risk prediction model.
[0083] It should be noted that the above-described device may also include other implementation manners according to the description of the method or apparatus embodiments. The specific implementation manner may refer to the description of the relevant method embodiments and will not be elaborated herein one by one.
[0084] The method embodiments provided in this specification may be executed on a mobile terminal, a computer terminal, a server, or a similar computing device. Taking running on a server as an example, Figure 3 is a hardware structure block diagram of an embodiment of a training server for an illegal fund-raising risk prediction model provided in this specification. The server may be the training device or training equipment for the illegal fund-raising risk prediction model in the above embodiments. As Figure 3 shown, the server 10 may include one or more (only one is shown in the figure) processors 100 (the processor 100 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 200 for storing data, and a transmission module 300 for communication functions. Those of ordinary skill in the art can understand that Figure 3 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the server 10 may further include more or fewer components than those shown in Figure 3 such as other processing hardware, such as a database or a multi-level cache, a GPU, or have a different configuration from that shown in Figure 3 shown.
[0085] The memory 200 can be used to store software programs and modules of application software, such as the program instructions / modules corresponding to the training method of the illegal fund-raising risk prediction model in the embodiments of this specification. The processor 100 executes various functional applications and data processing by running the software programs and modules stored in the memory 200. The memory 200 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 200 may further include a memory remotely disposed relative to the processor 100, and these remote memories can be connected to the computer terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0086] The transmission module 300 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of a computer terminal. In one instance, the transmission module 300 includes a network interface controller (NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one instance, the transmission module 300 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0087] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0088] The method or device described in the above embodiments provided in this specification can implement the business logic through a computer program and record it on a storage medium. The storage medium can be read and executed by a computer to achieve the effects of the solutions described in the embodiments of this specification. The storage medium may include a physical device for storing information, usually by digitizing the information and then storing it using media such as electricity, magnetism, or optics. The storage medium may include: devices that store information using electrical energy, such as various memories, such as RAM, ROM, etc.; devices that store information using magnetic energy, such as hard disks, floppy disks, magnetic tapes, magnetic core memories, bubble memories, USB flash drives; devices that store information using optical means, such as CDs or DVDs. Of course, there are also other ways of readable storage media, such as quantum memories, graphene memories, and so on.
[0089] The training method or device embodiments of the above-mentioned illegal fund-raising risk prediction model provided in this specification can be implemented by a processor executing corresponding program instructions in a computer. For example, it can be implemented in C++ language on a PC using the Windows operating system, in a Linux system, or in other ways, such as using programming languages for Android or iOS systems on intelligent terminals, and implemented based on the processing logic of quantum computers, etc.
[0090] It should be noted that the devices, equipment, and systems described above in the specification may also include other implementation manners according to the descriptions of the relevant method embodiments. The specific implementation manners can refer to the descriptions of the corresponding method embodiments and will not be elaborated here one by one.
[0091] The embodiments in this application are all described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the hardware + program type embodiments, since they are basically similar to the method embodiments, the descriptions are relatively simple, and the relevant parts can refer to the partial descriptions of the method embodiments.
[0092] For the convenience of description, when describing the above device, it is divided into various modules according to functions for description. Of course, when implementing one or more of this specification, the functions of some modules can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be implemented by a combination of multiple sub-modules or sub-units, etc.
[0093] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices, equipment, and systems according to the embodiments of the present invention. It should be understood that it can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate a device for implementing the specified functions. These computer program instructions can also be stored in a computer-readable memory that can guide the computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or block Figure 1 diagrams or multiple block diagrams.
[0094] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.
[0095] The above description is only for the embodiments of one or more embodiments of this specification, and is not intended to limit one or more embodiments of this specification. For those skilled in the art, various modifications and variations can be made to one or more embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims.
Claims
1. A training method for an illegal fund-raising risk prediction model, characterized in that, it includes: Obtain a training data set related to the enterprise's fund-raising risk; wherein, the training data set includes labeled samples and unlabeled samples; the labels of the labeled samples are determined according to whether there is an illegal fund-raising risk; Calculate the local density value and clustering membership degree of each sample in the training data set; Construct a pointwise manifold regularization constraint term according to the local density and clustering membership degree of each sample and the discriminant function of a preset classifier; wherein, the pointwise manifold regularization constraint term is used to constrain the relationship between each sample and its neighboring samples; Based on the pointwise manifold regularization constraint term, determine the loss function of the preset classifier; Train the preset classifier based on the training data set and the loss function to obtain an illegal fund-raising risk prediction model; Among them, the determining the loss function of the preset classifier based on the pointwise manifold regularization constraint term includes: Determine the loss function of the preset classifier according to the following formula: L = R emp + αR pm + βR reg Among them, \(L\) represents the loss function of the preset classifier, and \(R\) emp represents the minimized empirical loss function constructed according to the labeled samples, and \(R\) pm represents the pointwise manifold regularization constraint term, and \(R\) reg represents the \(l\) 2 regularization loss function, \(\alpha\) and \(\beta\) represent hyperparameters, and the regularization loss function is used to prevent overfitting; The minimization of the empirical loss function includes an absolute loss function, which is determined according to the following formula: The regularization loss function is determined according to the following formula: R reg = ||w|| 2 wherein, f(x) represents the predicted value, y represents the true value, and w is the parameter in the discriminant function f(·) of the preset classifier; Among them, the constructing a pointwise manifold regularization constraint term according to the local density and clustering membership degree of each sample and the discriminant function of a preset classifier includes: Obtain the sample weight of each sample according to the local density and clustering membership degree of each sample, wherein, a large weight is assigned to a sample with a high local density value, a small weight is assigned to a sample with a low local density value, and a large weight is assigned to a sample with a high clustering membership degree; Construct a pointwise manifold regularization constraint term based on the sample weights of all samples and the discriminant function of the preset classifier; The pointwise manifold regularization constraint term constructed based on the sample weights of all samples and the discriminant function of the preset classifier is: Among them, R pm represents the pointwise manifold regularization constraint term, n represents the total number of samples, i represents the serial number, p(x i ) represents the weight of the sample x i , f(·) is the discriminant function of the preset classifier, x i represents the i-th sample, N(x i ) represents the set of neighboring samples of the sample x i , x j represents a sample in the set of neighboring samples of the sample x i , w i,j represents the similarity between the sample x i and the sample x j .
2. The method according to claim 1, characterized in that, before obtaining the training data set related to the enterprise's illegal fund-raising risk, it includes: Obtain feature data related to the enterprise's fund-raising risk; wherein, the feature data includes labeled feature data and unlabeled feature data; Preprocess the feature data to obtain a training data set related to the enterprise's fund-raising risk; wherein, the preprocessing includes missing value processing and feature engineering processing.
3. The method according to claim 1, characterized in that, the calculating the local density value and clustering membership degree of each sample in the training data set includes: Determine the neighboring sample set of the first sample; Calculate the sum of the distances between the first sample and each sample in the corresponding neighboring sample set to obtain the neighboring distance sum of the first sample; According to the neighboring distance sum of the first sample, determine the total neighboring distance sum of all samples in the training data set; Based on the neighboring distance sum of the first sample and the total neighboring distance sum of all samples, determine the local density value of the first sample; Cluster the samples in the training data set to obtain the first type of clusters and the second type of clusters; Calculate the first probability that the first sample belongs to the first type of cluster and the second probability that it belongs to the second type of cluster; Take the maximum value of the first probability and the second probability as the clustering membership degree of the first sample.
4. The method according to claim 3, characterized in that, determine the local density value of the first sample according to the following formula: Among them, d(x i ) represents the local density value of the sample x i , N(x i ) represents the set of nearest neighbor samples of the sample x i , x j represents a sample in the set of nearest neighbor samples of the sample x i , d(x i , x j ) represents the distance between the sample x i and the sample x j , N(x s ) represents the set of nearest neighbor samples of the sample x s , x t represents a sample in the set of nearest neighbor samples of the sample x s , d(x s , x t ) represents the distance between the sample x s and the sample x t , n represents the total number of samples, and s represents the serial number.
5. The method according to claim 1, characterized in that, calculate the sample weight of each sample according to the following formula: Among them, p(x i ) represents the weight of the sample x i , N(x i ) represents the set of nearest neighbor samples of the sample x i , x j represents a sample in the set of nearest neighbor samples of the sample x i , d(x i , x j ) represents the distance between the sample x i and the sample x j , N(x s ) represents the set of nearest neighbor samples of the sample x s , x t represents a sample in the set of nearest neighbor samples of the sample x s , d(x s , x t ) represents the distance between the sample x s and the sample x t , n represents the total number of samples, s represents the serial number, u 1i represents the probability that the sample x i belongs to the first type of cluster, u 2i represents the probability that the sample x i belongs to the second type of cluster.
6. The method according to claim 1, characterized in that, further comprising: obtain the feature data related to the fund-raising risk in the target enterprise; process the feature data by using the illegal fund-raising risk prediction model to obtain the probability of the illegal fund-raising risk of the target enterprise; determine whether the target enterprise has an illegal fund-raising risk based on the relationship between the probability of the illegal fund-raising risk of the target enterprise and the preset risk value.
7. A training device for an illegal fund-raising risk prediction model, characterized in that, comprising: an acquisition module, configured to acquire a training data set related to the enterprise fund-raising risk; wherein, the training data set includes labeled samples and unlabeled samples; the labels of the labeled samples are determined according to whether there is an illegal fund-raising risk; a calculation module, configured to calculate the local density value and the clustering membership degree of each sample in the training data set; a construction module, configured to construct a point-wise manifold regularization constraint term according to the local density and clustering membership degree of each sample and the discriminant function of a preset classifier; wherein, the point-wise manifold regularization constraint term is used to constrain the relationship between each sample and its neighboring samples; a determination module, configured to determine the loss function of the preset classifier based on the point-wise manifold regularization constraint term; an obtaining module, configured to train the preset classifier based on the training data set and the loss function to obtain an illegal fund-raising risk prediction model; wherein, the determining the loss function of the preset classifier based on the point-wise manifold regularization constraint term includes: determine the loss function of the preset classifier according to the following formula: L = R emp + αR pm + βR reg Among them, L represents the loss function of the preset classifier, and R emp represents the minimized empirical loss function constructed according to the labeled samples, and R pm represents the pointwise manifold regularization constraint term, and R reg represents the l 2 regularization loss function, α and β represent hyperparameters, and the regularization loss function is used to prevent overfitting; The minimization of the empirical loss function includes an absolute loss function, which is determined according to the following formula: The regularization loss function is determined according to the following formula: R reg = ||w|| 2 wherein, f(x) represents the predicted value, y represents the true value, and w is a parameter in the discriminant function f(·) of the preset classifier; wherein, the constructing the point-wise manifold regularization constraint term according to the local density and clustering membership degree of each sample and the discriminant function of the preset classifier includes: obtain the sample weight of each sample according to the local density and clustering membership degree of each sample, wherein, assign a large weight to the sample with a high local density value, assign a small weight to the sample with a low local density value, and assign a large weight to the sample with a high clustering membership degree; construct a point-wise manifold regularization constraint term based on the sample weights of all samples and the discriminant function of the preset classifier; The point-wise manifold regularization constraint term constructed based on the sample weights of all samples and the discriminant function of the preset classifier is: Among them, R pm represents the pointwise manifold regularization constraint term, n represents the total number of samples, i represents the serial number, p(x i ) represents the weight of sample x i , f(·) is the discriminant function of the preset classifier, x i represents the i-th sample, N(x i ) represents the set of neighboring samples of sample x i , x j represents a sample in the set of neighboring samples of sample x i , w i,j represents the similarity between sample x i and sample x j .
8. A training device for an illegal fund-raising risk prediction model, characterized in that, Comprising at least one processor and a memory storing computer-executable instructions, when the processor executes the instructions, the steps of the method according to any one of claims 1-6 are implemented.
9. A computer-readable storage medium, characterized in that computer instructions are stored thereon, and when the instructions are executed, the steps of the method according to any one of claims 1-6 are implemented.