System Resource Data Allocation Method and Device
By building and optimizing the classifier, using user risk characteristic data to allocate resource data, the accuracy and efficiency of user risk prediction and resource data allocation in the prior art are solved, and more efficient and accurate user risk prediction and resource allocation are achieved.
Patent Information
- Application Number
- CN202110238542.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-04
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-03-04
AI Technical Summary
The prior art has problems with accuracy and efficiency in user risk prediction and resource data allocation, especially when there are fewer user risk samples and no accurate characterization of new service types.
By obtaining the specified information set with characteristic data that characterizes user risk characteristics, a classifier is built using the tagged sample set and tag set, and the nearest tagged samples with no tags are extracted, the information entropy of the tagless samples is calculated to optimize the classifier, and finally allocate system resource data to the target user.
It improves the accuracy and efficiency of system resource data allocation, enhances the ability to predict user risk characteristics, improves user experience and reduces resource data losses.
Smart Images

Figure CN113011722B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of machine learning, and particularly to a system resource data allocation method and apparatus. Background Art
[0002] With the rapid development of big data service platform technology, the types of financial resource data services and the selectable service channels are becoming more and more diverse and convenient, and user risk prediction has become increasingly important for financial institutions. For example, for some online loan services with relatively convenient service channels, since the manual intervention is relatively small, if the user risk prediction is not accurate enough, there may be problems such as unreasonable resource data allocation and poor user experience.
[0003] Currently, the commonly used user risk prediction methods are mainly classification methods based on supervised learning models. By modeling based on known customer risk data, the trained model is used to predict the user risk of new samples to determine the risk level of users. However, when using the classification method of supervised learning models, information on known user risks needs to be utilized. However, usually, the samples with known user risks are few, and the sample data with known user risks may not accurately represent the risk characteristics of users under new service types, thus affecting the accuracy of user risk prediction and the accuracy of resource data allocated to corresponding users, and further possibly leading to unreasonable resource data allocation and reducing the user experience.
[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] The purpose of the embodiments of this specification is to provide a system resource data allocation method and apparatus, which can improve the accuracy and efficiency of system resource data allocation.
[0006] The embodiments of this application provide a system resource data allocation method, including: obtaining a specified information set and a label set with feature data for characterizing user risk characteristics, where the specified information set includes a labeled sample set and an unlabeled sample set, and the label set includes the risk categories corresponding to each labeled sample in the labeled sample set; constructing a classifier using the labeled sample set and the label set; extracting the nearest neighbor labeled samples of each unlabeled sample in the unlabeled sample set, and calculating the information entropy corresponding to each unlabeled sample according to the distribution of the risk categories of the nearest neighbor labeled samples of each unlabeled sample, where the nearest neighbor labeled samples include the labeled samples in the labeled sample set whose proximity degree to the corresponding unlabeled sample in the user risk feature space meets a preset condition; optimizing the classifier based on the information entropy corresponding to each unlabeled sample in the unlabeled sample set to obtain an optimized classifier, so as to allocate system resource data to the target user based on the risk prediction result of the target user by the optimized classifier.
[0007] An embodiment of the present application further provides a system resource data allocation device, including: an acquisition module, configured to acquire a specified information set and a label set having feature data for characterizing user risk characteristics, where the specified information set includes a labeled sample set and an unlabeled sample set, and the label set includes risk categories corresponding to each labeled sample in the labeled sample set; a construction module, configured to construct a classifier by using the labeled sample set and the label set; a calculation module, configured to extract the neighboring labeled samples of each unlabeled sample in the unlabeled sample set, and calculate the information entropy corresponding to each unlabeled sample according to the distribution of the risk categories of the neighboring labeled samples of each unlabeled sample, where the neighboring labeled samples include the labeled samples in the labeled sample set whose proximity degree to the corresponding unlabeled sample in the user risk feature space meets a preset condition; an optimization module, configured to optimize the classifier based on the information entropy corresponding to each unlabeled sample in the unlabeled sample set to obtain an optimized classifier, so as to allocate system resource data to the target user based on the risk prediction result of the target user by the optimized classifier.
[0008] An embodiment of the present application further provides a computer device, including a processor and a memory for storing processor-executable instructions, where the processor, when executing the instructions, implements the steps of the system resource data allocation method in any of the above embodiments.
[0009] An embodiment of the present application further provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed, the steps of the system resource data allocation method in any of the above embodiments are implemented.
[0010] In an embodiment of the present application, a method for allocating system resource data is provided. A labeled sample set, an unlabeled sample set, and a label set with feature data for characterizing user risk characteristics can be obtained. The label set includes the risk categories corresponding to each labeled sample in the labeled sample set. First, an empirical loss can be used to initialize a classifier with the labeled sample set and the label set to maximize the fitting degree of the classifier on the labeled data. After that, the information entropy of each unlabeled sample can be calculated using the distribution of the risk categories of the labeled samples neighboring the unlabeled sample. The greater the information entropy, the greater the possibility that the unlabeled sample is in the risk classification boundary region in the risk feature space and the greater the contribution ratio to the classification boundary. The smaller the information entropy, the smaller the possibility that the unlabeled sample is in the classification boundary region and the smaller the contribution to the classification boundary. The classifier can be optimized based on the information entropy corresponding to each unlabeled sample in the unlabeled sample set, which can make the optimized classifier more accurate in risk classification of the feature data near the classification boundary, thereby improving the accuracy and efficiency of system resource data allocation. In addition, the classifier can be used to perform risk classification on each unlabeled sample in the unlabeled sample set to obtain the pseudo-labels corresponding to each unlabeled sample. According to the similarities and differences between the pseudo-labels of each unlabeled sample and the risk categories of the neighboring labeled samples of the unlabeled sample, a neighboring discriminant matrix corresponding to the unlabeled sample set can be calculated, making full use of the spatial distribution information between the unlabeled samples and the neighboring labeled samples. After that, the classifier can be optimized using the neighboring discriminant matrix, making the output of the optimized classifier for the unlabeled samples as close as possible to the output of the neighboring labeled samples of the same class and as opposite as possible to the output of the neighboring labeled samples of different classes, thereby improving the accuracy of classification. Description of the Drawings
[0011] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:
[0012] Figure 1 It is a schematic flowchart of an embodiment of a method for allocating system resource data provided by this specification;
[0013] Figure 2 It is a schematic flowchart of the construction process of a risk prediction model in an embodiment provided by this specification;
[0014] Figure 3 It is a schematic flowchart of a method for allocating system resource data in an embodiment provided by this specification;
[0015] Figure 4 Schematic diagram of the module structure of a system resource data allocation device provided in this specification;
[0016] Figure 5 Schematic diagram of a computer device provided in this specification. Specific embodiments
[0017] The principles and spirit of this application will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement this application, and do not limit the scope of this application in any way. On the contrary, these embodiments are provided to make the disclosure of this application more thorough and complete, and to be able to convey the scope of this disclosure fully to those skilled in the art.
[0018] Those skilled in the art know that the embodiments of this application can be implemented as a system, device, equipment, method, or computer program product. Therefore, the disclosure of this application can be specifically implemented in the following forms, namely: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0019] In a scenario example provided by an embodiment of this specification, the system resource data allocation method can be applied to a device that executes system resource data allocation. The device can include a single server or a server cluster composed of multiple servers. For a target user, the server can extract feature data from multiple types of information of the target user as the feature data of the target user, and then use a pre-configured algorithm or model, etc. to perform risk prediction on the target user to obtain the risk prediction result of the target user, so as to allocate resource data to the target user based on this risk prediction result. The resource data can include data resources such as services and products provided or recommended to users. For example, in the loan business scenario, the resource data can be the loan amount and / or loan type allocated to the target user. In the cloud platform data service business scenario, the resource data can be the system resource data allocated to the target user, etc. By accurately identifying the riskiness of users, resource data can be allocated to users more accurately and reasonably, improving the user experience and effectively reducing the loss of resource data of the institution.
[0020] Figure 1It is a schematic flowchart of an embodiment of the system resource data allocation method provided in this specification. Although this specification provides method operation steps or device structures as shown in the following embodiments or drawings, more or fewer operation steps or module units may be included in the method or device based on routine or non-creative labor. In steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure shown in the embodiments or drawings of this specification. When the method or module structure is applied to an actual device, server, or terminal product, it can be executed sequentially or in parallel according to the method or module structure shown in the embodiments or drawings (for example, in an environment of parallel processors or multi-threaded processing, and even including an implementation environment of distributed processing or server clusters). A specific example is as follows Figure 1 As shown, in an embodiment of the system resource data allocation method provided in this specification, the method can be applied to the data processing device, and the method may include the following steps.
[0021] Step S101, obtain a specified information set and a label set having feature data for characterizing user risk characteristics, where the specified information set includes a labeled sample set and an unlabeled sample set, and the label set includes the risk categories corresponding to each labeled sample in the labeled sample set.
[0022] The server can obtain the specified information set and the label set. The specified information set may include multiple sample data. The sample data may include feature data for characterizing user risk characteristics. Correspondingly, the specified information set may be a data set composed of feature data for characterizing user risk characteristics. The feature data may be, for example, feature data extracted based on the business data of users stored in the business system of a financial institution. Feature extraction can be performed through feature engineering. The extraction method and feature type of the feature data can be set according to the actual application scenario and are not limited here. Of course, it may also include feature data extracted from user information obtained by the server from an associated platform of a financial institution.
[0023] In some embodiments, the sample data can be labeled samples or unlabeled samples. The labeled samples can be sample data with the risk categories of the configured users. The unlabeled samples can refer to the sample data without the configured risk categories of the users. The risk categories can be "at risk", "risk-free", etc., or can be risk levels such as "high risk", "medium risk", "low risk", etc., which can be configured according to the actual application scenarios. For example, for the convenience of processing, a single sample data can be set to correspond to a single user. Correspondingly, the feature data of the users with known risk categories and the feature data of the users with unknown risk categories can be extracted respectively to construct labeled samples and unlabeled samples. After associating the feature data corresponding to the labeled samples and unlabeled samples with the user identifiers, they are respectively stored in the labeled sample set and the unlabeled sample set of the specified information set. After associating the risk categories corresponding to the labeled samples with the user identifiers, they are stored in the label set.
[0024] The pre-constructed specified information set and label set can be stored locally or stored in a database. The server can extract the specified information set and label set when allocating system resource data or constructing a prediction model. If the constructed specified information set is an information set composed of the information of the users corresponding to the specified product or the specified service scenario, an information set identifier can be set for each specified information set. Correspondingly, the server can obtain the specified information set and label set corresponding to the corresponding information set identifier according to the needs of the current test scenario for the allocation of system resource data in the current test scenario. At present, most of the business data in the business system is updated relatively fast. Correspondingly, the feature data of the specified information set and the label set can be dynamically updated at intervals to ensure the accuracy of the information in the information set and thus improve the accuracy of prediction.
[0025] Step S102: Construct a classifier using the labeled sample set and the label set.
[0026] During the process of constructing the classifier, the server can select a classification algorithm according to its needs, such as Bayesian, support vector machine, neural network, etc. Then, based on the selected classification algorithm, the labeled sample set and the label set are respectively used for model construction to obtain a classifier.
[0027] Step S103: Extract the neighboring labeled samples of each unlabeled sample in the unlabeled sample set, and calculate the information entropy corresponding to each unlabeled sample according to the distribution of the risk categories of the neighboring labeled samples of each unlabeled sample, where the neighboring labeled samples include the labeled samples in the labeled sample set whose proximity degree in the user risk feature space to the corresponding unlabeled sample meets the preset conditions.
[0028] The server can extract the neighboring labeled samples corresponding to each unlabeled sample in the unlabeled sample set. The neighboring labeled samples of each unlabeled sample are the labeled samples in the labeled sample set that satisfy a preset condition in terms of the proximity degree in the user risk feature space. Among them, the user risk feature space can be a space composed of feature data used to characterize user risk features. For example, the feature data can be a feature vector, and a multi-dimensional feature vector can form a multi-dimensional feature space. The labeled samples that satisfy the preset condition in terms of the proximity degree can be the top n (where n is a positive integer) labeled samples closest to the unlabeled sample, or the labeled samples whose distance from the unlabeled sample is less than a preset distance.
[0029] In one embodiment, the k-nearest neighbor algorithm can be used to determine the neighboring labeled samples corresponding to each unlabeled sample in the unlabeled sample set. Among them, the k-nearest neighbor algorithm can determine the k labeled samples closest to the unlabeled sample as the neighboring labeled samples corresponding to the unlabeled sample.
[0030] After extracting the neighboring labeled samples corresponding to each unlabeled sample in the unlabeled sample set, the information entropy corresponding to each unlabeled sample can be calculated according to the distribution of the risk categories of the neighboring labeled samples of each unlabeled sample. In one embodiment, the distribution of the risk categories of the neighboring labeled samples can refer to the proportion of the labeled samples belonging to each risk category among the neighboring labeled samples. The risk category of the neighboring labeled sample can be obtained according to the corresponding label set. The more disordered the distribution of the risk categories of the neighboring labeled samples is, the higher the information entropy of the corresponding unlabeled sample is, and the more ordered the distribution of the risk categories of the neighboring labeled samples is, the lower the information entropy of the corresponding unlabeled sample is. The information entropy can be used to characterize the degree of disorder in the distribution of the neighboring labeled samples of the unlabeled sample. The higher the information entropy of the unlabeled sample is, the greater the possibility that the unlabeled sample is in the classification boundary region and the greater the contribution to the classification boundary. The lower the information entropy of the unlabeled sample is, the smaller the possibility that the unlabeled sample is in the classification boundary region and the smaller the contribution to the classification boundary.
[0031] Step S104: Optimize the classifier based on the information entropy corresponding to each unlabeled sample in the unlabeled sample set to obtain an optimized classifier, so as to allocate system resource data to the target user based on the risk prediction result of the target user by the optimized classifier.
[0032] The server can optimize the classifier based on the information entropy corresponding to each unlabeled sample in the unlabeled sample set to obtain an optimized classifier. In one embodiment, when optimizing the classifier, different weights are assigned to unlabeled samples with different information entropy values. The larger the information entropy of an unlabeled sample, the greater the weight assigned to it, so that the optimized classifier can classify the risk of feature data near the classification boundary more accurately, thereby improving the accuracy and efficiency of system resource data allocation.
[0033] The method in the above embodiment can obtain a labeled sample set, an unlabeled sample set, and a label set with feature data for characterizing user risk features. First, an empirical loss can be used to initialize the classifier with the labeled sample set and the label set to maximize the fitting degree of the classifier on the labeled data. Then, the information entropy of each unlabeled sample can be calculated using the distribution of the risk categories of the labeled samples neighboring the unlabeled samples. The larger the information entropy, the greater the likelihood that the unlabeled sample is in the risk classification boundary region of the risk feature space and the greater its contribution to the classification boundary. The smaller the information entropy, the smaller the likelihood that the unlabeled sample is in the classification boundary region and the smaller its contribution. The classifier can be optimized based on the information entropy corresponding to each unlabeled sample in the unlabeled sample set, so that the optimized classifier can classify the risk of feature data near the classification boundary more accurately, thereby improving the accuracy and efficiency of system resource data allocation.
[0034] In some embodiments of the present application, after constructing the classifier using the labeled sample set and the label set, it may further include: using the classifier to perform risk classification on each unlabeled sample in the unlabeled sample set to obtain the pseudo-label corresponding to each unlabeled sample; calculating the neighbor discrimination matrix corresponding to the unlabeled sample set according to the similarities and differences between the pseudo-labels of each unlabeled sample in the unlabeled sample set and the risk categories of the neighboring labeled samples; correspondingly, optimizing the classifier based on the information entropy corresponding to each unlabeled sample in the unlabeled sample set to obtain an optimized classifier may include: optimizing the classifier based on the neighbor discrimination matrix corresponding to the unlabeled sample set and the information entropy corresponding to each unlabeled sample to obtain an optimized classifier.
[0035] Specifically, after constructing a classifier using the labeled sample set and the label set, the classifier can be used to perform risk classification on each unlabeled sample in the unlabeled sample set to obtain the pseudo-labels corresponding to each unlabeled sample. That is, each unlabeled sample is input into the classifier, and the pseudo-labels of each unlabeled sample are output. After obtaining the pseudo-labels of each unlabeled sample, the proximity discrimination matrix corresponding to the unlabeled sample set can be calculated based on the similarities and differences between the pseudo-labels of each unlabeled sample and the risk categories of the neighboring labeled samples of the unlabeled sample. The proximity discrimination matrix can characterize the data distribution information between the unlabeled sample set and the labeled sample set. After obtaining the proximity discrimination matrix corresponding to the unlabeled sample set, the classifier can be optimized based on the proximity discrimination matrix corresponding to the unlabeled sample set and the information corresponding to each unlabeled sample to obtain an optimized classifier. In the above embodiment, by optimizing the classifier using the proximity discrimination matrix, the output of the optimized classifier for unlabeled samples is made as close as possible to the output of neighboring labeled samples of the same class and as opposite as possible to the output of neighboring labeled samples of different classes, thereby improving the accuracy of classification.
[0036] In some embodiments of the present application, optimizing the classifier based on the proximity discrimination matrix corresponding to the unlabeled sample set and the information entropy corresponding to each unlabeled sample to obtain an optimized classifier may include: using the information entropy and pseudo-labels corresponding to each unlabeled sample to construct a boundary enhancement constraint for the classifier to predict the user risk characteristics of the unlabeled samples in the unlabeled sample set; using the elements in the proximity discrimination matrix corresponding to the unlabeled sample set to construct a proximity discrimination constraint for the classifier to predict the user risk characteristics of the unlabeled samples in the unlabeled sample set and the labeled samples in the labeled sample set; and optimizing the classifier based on the boundary enhancement constraint and the proximity discrimination constraint to obtain an optimized classifier.
[0037] Specifically, the information entropy and pseudo-labels corresponding to each unlabeled sample can be used to construct a boundary enhancement constraint for the classifier to predict the user risk characteristics of the unlabeled samples in the unlabeled sample set. Among them, the boundary enhancement constraint means that a larger weight is assigned to the unlabeled samples near the classification boundary, and a smaller weight is assigned to the unlabeled samples far from the classification boundary, thereby achieving boundary enhancement. By optimizing the classifier through the boundary enhancement constraint, the optimized classifier can more accurately predict the risk categories of the feature data near the classification.
[0038] The elements in the nearest neighbor discrimination matrix corresponding to the unlabeled samples can be utilized to construct a classifier for performing a nearest neighbor discrimination constraint on the unlabeled samples in the unlabeled sample set and the labeled samples in the labeled sample set for user risk feature prediction. Among them, the nearest neighbor discrimination constraint can characterize information such as whether the unlabeled samples in the unlabeled sample set are near neighbors with the labeled samples in the labeled sample set, and in the case of being near neighbors, whether the pseudo-labels of the unlabeled samples are the same as the risk categories of the labeled samples. By optimizing the classifier through the nearest neighbor discrimination constraint, the spatial distribution information between the unlabeled samples and the nearest neighbor labeled samples can be fully utilized, so that the output of the optimized classifier for the unlabeled samples is as close as possible to the output of the nearest neighbor labeled samples of the same class and as opposite as possible to the output of the nearest neighbor labeled samples of different classes, thereby improving the classification accuracy.
[0039] In some embodiments of the present application, the risk categories include positive classes and negative classes. Correspondingly, according to the distribution of the risk categories of the nearest neighbor labeled samples of each unlabeled sample, calculating the information entropy corresponding to each unlabeled sample may include: calculating the information entropy of each unlabeled sample according to the following formula:
[0040]
[0041] Among them, is the i-th unlabeled sample, H i is the information entropy of, N is the number of nearest neighbor labeled samples of, N + is the number of positive class samples among the N nearest neighbor labeled samples of, N - is the number of negative class samples among the N nearest neighbor labeled samples of. Among them, can also represent the probability that the label of the unlabeled sample is a positive class, and can also represent the probability that the unlabeled sample is a negative class. Through the above method, the information entropy corresponding to each unlabeled sample can be calculated based on the distribution of the risk categories of the nearest neighbor labeled samples of each unlabeled sample.
[0042] In some embodiments of the present application, using the information entropy and pseudo-labels corresponding to each unlabeled sample to construct a boundary enhancement constraint for the classifier to perform user risk feature prediction on the unlabeled samples in the unlabeled sample set may include: constructing the boundary enhancement constraint according to the following formula:
[0043]
[0044] weight i = exp((H i - μ) / σ);
[0045] Among them, Rbe is the boundary enhancement constraint, X U is the unlabeled sample set, |X U | is the number of unlabeled samples in X U ; is the i-th unlabeled sample in X U , f(·) is the discriminant function of the classifier is the corresponding pseudo-label, weight i is the boundary enhancement coefficient of i is the information entropy of U μ is the mean of the information entropies of multiple unlabeled samples in X U σ is the standard deviation of the information entropies of multiple unlabeled samples in X. In the above manner, the boundary enhancement constraint can be constructed based on the information entropy and pseudo-label corresponding to each unlabeled sample.
[0046] In some embodiments of the present application, according to the similarities and differences between the pseudo-labels of the unlabeled samples in the unlabeled sample set and the risk categories of the neighboring labeled samples, the neighboring discriminant matrix corresponding to the unlabeled sample set can be calculated, which may include: determining the neighboring discriminant matrix corresponding to the unlabeled sample set according to the following formula:
[0047]
[0048]
[0049] where S is the neighboring discriminant matrix corresponding to the unlabeled sample set, and its dimension is X U |×|X|, X U is the unlabeled sample set, |X U | is the number of unlabeled samples in X U X is the labeled sample set, |X| is the number of labeled samples in X, s i,j is an element in the neighboring discriminant matrix is the i-th unlabeled sample in X U , x j is the j-th labeled sample in X represents the neighboring labeled samples whose risk categories are the same as the pseudo-label of , represents the neighboring labeled samples whose risk categories are different from the pseudo-label of . Among them, other means that x j does not belong to the unlabeled sample The neighboring samples with labels. Through the above method, based on the similarities and differences between the pseudo-labels of each unlabeled sample in the unlabeled sample set and the risk categories of the neighboring labeled samples, the neighboring discrimination matrix corresponding to the unlabeled sample set can be calculated.
[0050] In some embodiments of the present application, using the elements in the neighboring discrimination matrix corresponding to the unlabeled sample set to construct a classifier for predicting the user risk characteristics of the unlabeled samples in the unlabeled sample set and the labeled samples in the labeled sample set, the neighboring discrimination constraint may include: constructing the neighboring discrimination constraint according to the following formula:
[0051]
[0052] where R nd is the boundary enhancement constraint, X is the labeled sample set, |X| is the number of labeled samples in X, X U is the unlabeled sample set, |X U | is the number of unlabeled samples in X U , s i,j is an element in the neighboring discrimination matrix, is the i-th unlabeled sample in X U , x j is the j-th labeled sample in X, and f(·) is the discrimination function of the classifier. Through the above method, the neighboring discrimination constraint can be constructed using the elements in the neighboring discrimination matrix corresponding to the unlabeled sample set.
[0053] In some embodiments of the present application, based on the boundary enhancement constraint and the neighboring discrimination constraint, optimizing the classifier to obtain the optimized classifier may include: optimizing the classifier according to the following formula:
[0054] f * = argmin f L(f, X, Y, X U );
[0055] L = R emp + αR be + βR nd ;
[0056]
[0057] where f * is the discrimination function corresponding to the optimized classifier, L(·) is the objective function, f is the discrimination function corresponding to the classifier, X is the labeled sample set, Y is the label set, X U is the unlabeled sample set, R emp is the empirical loss of the classifier for predicting the user risk characteristics of the labeled samples in the labeled sample set, R beis the boundary enhancement constraint, R nd is the boundary enhancement constraint, |X| is the number of labeled samples in X, and x j is the j-th labeled sample in X, and y j is the j-th label in Y, that is, x j corresponding risk category, and α and β are hyperparameters. Through the above method, the classifier can be optimized based on the boundary enhancement constraint and the nearest neighbor discrimination constraint.
[0058] In some embodiments of the present application, the feature data may include time series aggregation features and time series historical features; wherein, the time series aggregation features may refer to data obtained by extracting features from the specified information of the user based on different time dimensions and time series feature extraction algorithms; the time series historical features may include time series distribution data statistically obtained from the specified information of the user based on different time dimensions. The time dimensions may include, for example, the previous month, the previous two months, the previous three months, etc., and the second previous month, the third previous month, the fourth previous month, etc. The time series feature extraction algorithms may include, for example, mean, variance, standard deviation, etc. By further combining time series feature information to construct feature data, the features of users of different churn types can be characterized more accurately, thereby improving the accuracy of system resource data allocation. By performing time series feature analysis on the information in the user's information that fluctuates significantly over time, horizontal analysis of user features can be achieved, thus greatly improving the accuracy of user risk category prediction.
[0059] In some embodiments, the time series aggregation feature F agg can be extracted in the following manner:
[0060] F agg =[f(feature) time , time = 1 - 3, 1 - 6, 1 - 9, 1 - 12]
[0061] f() respectively takes the Mean() average value, Max() maximum value, Min() minimum value, and Std() standard deviation, and the time periods are respectively the previous month, the previous three months, the previous six months, the previous ninth month, and the previous twelfth month.
[0062] The time series historical feature F his can be extracted in the following manner:
[0063] F his =[feature time , time = 1, 2, 3, 4, 5, 6]
[0064] The time periods are respectively the first previous month, the second previous month, the third previous month, the fourth previous month, the fifth previous month, and the sixth previous month.
[0065] In the above manner, according to the feature information at different time nodes, temporal feature information is constructed, enabling the model to better consider the feature information of the past when learning the features of the current time node, thereby improving the accuracy of the model.
[0066] The above method will be described below in conjunction with a specific embodiment. However, it should be noted that this specific embodiment is only for better explaining the present application and does not constitute an improper limitation of the present application.
[0067] Please refer to Figure 2 and Figure 3 , Figure 2 , which is a schematic diagram of the construction process of a risk prediction model in an embodiment provided in this specification; Figure 3 , which is a schematic diagram of the process of a system resource data allocation method in an embodiment provided in this specification. As Figure 2 shown, the construction process of the risk prediction model may include the following steps:
[0068] Step 1: Obtain labeled samples and unlabeled samples, where the labeled samples are set with labels corresponding to the risk categories of each sample, and the unlabeled samples do not include the labels corresponding to the risk categories of each sample.
[0069] Step 2: An empirical loss can be used to initialize the classification model based on the labeled samples to maximize the fitting degree of the classifier on the labeled data, and an empirical loss for the classification model to predict the user risk features of the labeled samples in the labeled sample set can be constructed.
[0070] Step 3: Calculate the information entropy of the unlabeled samples using the categories of the labeled samples neighboring the unlabeled samples. The larger the entropy, the more the sample is in the classification boundary region and the greater the contribution to the classification boundary; the smaller the entropy, the farther the sample is from the classification boundary and the smaller the contribution to the classification boundary.
[0071] Step 4: Use the information entropy to construct a boundary enhancement coefficient for the unlabeled samples.
[0072] Step 5: Use the empirical loss to initialize the classifier to assign pseudo-labels to the unlabeled samples to obtain pseudo-labeled samples.
[0073] Step 6: Use the boundary enhancement coefficient and the pseudo-labeled samples to construct a boundary enhancement constraint.
[0074] Step 7: Calculate the nearest neighbor discrimination matrix using the pseudo-labeled samples and the labeled samples neighboring the pseudo-labeled samples.
[0075] Step 8: Construct a nearest neighbor discrimination constraint based on the nearest neighbor discrimination matrix.
[0076] Step 9: Optimize the classifier based on the empirical loss, boundary enhancement constraint, and neighbor discrimination constraint to obtain a semi-supervised classification model based on boundary enhancement and neighbor discrimination constraints.
[0077] As Figure 3 shown, taking individual users as an example, the system resource data allocation method may include the following.
[0078] Obtain user-related feature information from the data warehouse.
[0079] Data preprocessing. Classify the features related to personal loan risk prediction into three categories: customer basic information, customer asset information, and customer transaction information. The data range can be determined according to the category, and thus the data tables involved can be determined. Observe the data columns related to customer basic information, customer asset information, and customer transaction information in the data tables. Concatenate the relevant data columns in different tables according to the customer ID to form the original features. For columns with missing values, complete them in a certain way. For example, for missing values of numerical features, fill them with the column mean, and for missing values of non-numerical features, fill them with "unknown".
[0080] Feature engineering. For categorical features such as education level and gender, perform One-Hot encoding. The basic features include personal basic information, personal asset information, and personal transaction information, and derivative features are constructed based on these information, including time series aggregation features and time series historical features. Then, construct training samples, which include labeled samples and unlabeled samples. The labels of the labeled samples are 1 (ω 1 ) and -1 (ω 2 ) representing personal loan risk customers and risk-free customers respectively. Unlabeled training samples do not need to construct labels.
[0081] Among them, the time series aggregation feature F agg can be extracted in the following way,
[0082] F agg = [f(feature) time , time = 1-3, 1-6, 1-9, 1-12]
[0083] f() takes the Mean() average, Max() maximum, Min() minimum, and Std() standard deviation respectively, and the time periods are the previous month, the previous three months, the previous six months, the previous ninth month, and the previous twelfth month.
[0084] The time series historical feature F his can be extracted in the following way,
[0085] F his = [feature time, where time = 1, 2, 3, 4, 5, 6
[0086] The time periods are respectively the first month before, the second month before, the third month before, the fourth month before, the fifth month before, and the sixth month before.
[0087] In some embodiments, the feature data may include time series aggregation features and time series historical features. Among them, the time series aggregation features may refer to the data obtained by extracting features from the specified information of the user based on different time dimensions and time series feature extraction algorithms. The time series historical features may include the time series distribution data obtained by statistically analyzing the specified information of the user based on different time dimensions. The time dimensions may include, for example, the previous month, the previous two months, the previous three months, etc., and the second month before, the third month before, the fourth month before, etc. The time series feature extraction algorithms may include, for example, mean, variance, standard deviation, etc. By further combining time series feature information to construct feature data, the features of users of different churn types can be more accurately characterized, thereby improving the accuracy of system resource data allocation. By performing time series feature analysis on the information in the user's information that fluctuates significantly over time, horizontal analysis of user features can be achieved, thus greatly improving the accuracy of user stability prediction.
[0088] In some embodiments, the time series aggregation feature F agg can be extracted in the following manner.
[0089] F agg = [f(feature) time , where time = 1, 2, 3, 4, 5, 6, 1 - 2, 1 - 3, 1 - 4, 1 - 5, 1 - 6
[0090] f() takes Mean() average value, Max() maximum value, Min() minimum value, and Std() standard deviation respectively, and the time periods are respectively the previous month, the previous two months, the previous three months, the previous four months, the previous five months, the previous six months, the second month before, the third month before, the fourth month before, the fifth month before, and the sixth month before. Correspondingly, each deposit and loan feature respectively derives 44 - dimensional time series aggregation features.
[0091] The time series historical feature F his can be extracted in the following manner.
[0092] F his = [feature time , where time = 1, 2, 3, 4, 5, 6
[0093] The time periods are respectively the first month before, the second month before, the third month before, the fourth month before, the fifth month before, and the sixth month before.
[0094] Model training. Use the labeled samples to construct an empirical loss to initialize the classification model, so as to maximize the fitting degree of the classification model on the labeled data. Secondly, calculate the information entropy of the unlabeled samples using the classes of the labeled samples that are neighbors of the unlabeled samples. The larger the entropy, the more the sample is in the classification boundary region and the greater its contribution to the classification boundary; the smaller the entropy, the farther the sample is from the classification boundary and the smaller its contribution to the classification boundary. Use the information entropy to construct a boundary enhancement coefficient for the unlabeled samples. During the optimization process of the model, the classifier assigns pseudo-labels to the unlabeled samples, and uses the boundary enhancement coefficient and the pseudo-labels to construct a boundary enhancement constraint. Finally, use the labeled samples to construct a neighbor discrimination matrix for the unlabeled samples to make full use of the data distribution information between the unlabeled samples and the neighboring labeled samples. Design a neighbor discrimination constraint using the neighbor discrimination matrix, so that the outputs of the pseudo-labeled samples and the neighboring labeled samples of the same class are as close as possible, and the outputs of the pseudo-labeled samples and the neighboring labeled samples of different classes are as opposite as possible, to improve the accuracy of the pseudo-labels. Iteratively optimize the classification model by minimizing the empirical loss, the boundary enhancement constraint, and the neighbor discrimination constraint to obtain a semi-supervised classification model based on boundary enhancement and neighbor discrimination constraints.
[0095] Input the test samples to be predicted into the semi-supervised classification model based on boundary enhancement and neighbor discrimination constraints, and output the prediction results, that is, the risk classes of the test samples.
[0096] Based on the same inventive concept, an embodiment of the present application also provides a system resource data allocation device as described in the following embodiments. Since the principle of the system resource data allocation device for solving problems is similar to that of the system resource data allocation method, the implementation of the system resource data allocation device can refer to the implementation of the system resource data allocation method, and the repeated parts will not be described again. As used below, the term "unit" or "module" can be a combination of software and / or hardware that can implement a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated. Figure 4 is a structural block diagram of the system resource data allocation device according to an embodiment of the present application, as Figure 4 shown, including: an acquisition module 401, a construction module 402, a calculation module 403, and an optimization module 404. The following describes this structure.
[0097] The acquisition module 401 is used to acquire a specified information set having feature data for characterizing user risk characteristics and a label set, where the specified information set includes a labeled sample set and an unlabeled sample set, and the label set includes the risk classes corresponding to each labeled sample in the labeled sample set.
[0098] The construction module 402 is used to construct a classifier using the labeled sample set and the label set.
[0099] The calculation module 403 is configured to extract the neighboring labeled samples of each unlabeled sample in the unlabeled sample set, and calculate the information entropy corresponding to each unlabeled sample according to the distribution of the risk categories of the neighboring labeled samples of each unlabeled sample, where the neighboring labeled samples include the labeled samples in the labeled sample set whose proximity to the corresponding unlabeled sample in the user risk feature space meets a preset condition.
[0100] The optimization module 404 is configured to optimize the classifier based on the information entropy corresponding to each unlabeled sample in the unlabeled sample set, so as to obtain an optimized classifier, and allocate system resource data to the target user based on the risk prediction result of the target user by the optimized classifier.
[0101] In some embodiments of the present application, the device further includes a matrix calculation module, which can be used to: after constructing a classifier using the labeled sample set and the label set, use the classifier to perform risk classification on each unlabeled sample in the unlabeled sample set to obtain the pseudo-labels corresponding to each unlabeled sample; calculate the neighboring discriminant matrix corresponding to the unlabeled sample set according to the similarities and differences between the pseudo-labels of each unlabeled sample in the unlabeled sample set and the risk categories of the neighboring labeled samples; correspondingly, the optimization module can be specifically configured to: optimize the classifier based on the neighboring discriminant matrix corresponding to the unlabeled sample set and the information entropy corresponding to each unlabeled sample, so as to obtain an optimized classifier.
[0102] From the above description, it can be seen that the embodiments of the present application achieve the following technical effects: It is possible to obtain a labeled sample set, an unlabeled sample set, and a label set with feature data for characterizing user risk characteristics. The label set includes the risk categories corresponding to each labeled sample in the labeled sample set. First, an empirical loss can be used to initialize the classifier with the labeled sample set and the label set to maximize the fitting degree of the classifier on the labeled data. After that, the information entropy of each unlabeled sample can be calculated using the distribution of the risk categories of the labeled samples neighboring the unlabeled samples. The greater the information entropy, the greater the possibility that the unlabeled sample is in the risk classification boundary region in the risk feature space and the greater the contribution ratio to the classification boundary. The smaller the information entropy, the smaller the possibility that the unlabeled sample is in the classification boundary region and the smaller the contribution to the classification boundary. The classifier can be optimized based on the information entropy corresponding to each unlabeled sample in the unlabeled sample set, which can make the optimized classifier more accurate in risk classification of the feature data near the classification boundary, thereby improving the accuracy and efficiency of system resource data allocation. In addition, the classifier can also be used to perform risk classification on each unlabeled sample in the unlabeled sample set to obtain the pseudo-labels corresponding to each unlabeled sample. According to the similarities and differences between the pseudo-labels of each unlabeled sample and the risk categories of the neighboring labeled samples of the unlabeled sample, the neighboring discriminant matrix corresponding to the unlabeled sample set can be calculated, which can make full use of the spatial distribution information between the unlabeled samples and the neighboring labeled samples. After that, the classifier can be optimized using the neighboring discriminant matrix, so that the output of the optimized classifier for the unlabeled samples is as close as possible to the output of the neighboring labeled samples of the same class and as opposite as possible to the output of the neighboring labeled samples of different classes, thereby improving the accuracy of classification.
[0103] The embodiments of the present application also provide a computer device, which can be specifically referred to Figure 5 the schematic structural diagram of the computer device composition based on the system resource data allocation method provided by the embodiments of the present application shown in the figure. The computer device may specifically include an input device 51, a processor 52, and a memory 53. Among them, the memory 53 is used to store instructions executable by the processor. When the processor 52 executes the instructions, the steps of the system resource data allocation method described in any of the above embodiments are implemented.
[0104] In this embodiment, the input device may specifically be one of the main devices for information exchange between a user and a computer system. The input device may include a keyboard, a mouse, a camera, a scanner, a light pen, a handwriting input board, a voice input device, etc.; the input device is used to input raw data and programs for processing these data into the computer. The input device may also acquire and receive data transmitted from other modules, units, and devices. The processor may be implemented in any suitable manner. For example, the processor may take the form of, for example, a microprocessor or a processor, a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an Application Specific Integrated Circuit (ASIC), a programmable logic controller, and a form embedded microcontroller, and so on. The memory may specifically be a memory device for storing information in modern information technology. The memory may include multiple levels. In a digital system, anything that can store binary data can be a memory; in an integrated circuit, a circuit without a physical form but with a storage function is also called a memory, such as RAM, FIFO, etc.; in a system, a storage device with a physical form is also called a memory, such as a memory module, a TF card, etc.
[0105] In this embodiment, the functions and effects specifically implemented by this computer device may be explained by comparison with other embodiments and will not be elaborated here.
[0106] In an embodiment of the present application, there is also provided a computer storage medium based on a system resource data allocation method. The computer storage medium stores computer program instructions, and when the computer program instructions are executed, the steps of the system resource data allocation method described in any of the above embodiments are implemented.
[0107] In this embodiment, the above storage medium includes but is not limited to a Random Access Memory (RAM), a Read-Only Memory (ROM), a Cache, a Hard Disk Drive (HDD), or a Memory Card. The memory may be used to store computer program instructions. The network communication unit may be set according to the standards specified by the communication protocol and is used as an interface for network connection communication.
[0108] In this embodiment, the functions and effects specifically implemented by the program instructions stored in this computer storage medium may be explained by comparison with other embodiments and will not be elaborated here.
[0109] Obviously, those skilled in the art should understand that the various modules or steps of the embodiments of the present application described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0110] It should be understood that the above description is for illustrative purposes rather than for limitation. Many embodiments and many applications other than the examples provided will be apparent to those skilled in the art upon reading the above description. Therefore, the scope of the present application should not be determined with reference to the above description, but should be determined with reference to the full scope of the foregoing claims and the equivalents thereof.
[0111] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A method for allocating system resource data, characterized in that, applied to a server, the method includes: obtaining a specified information set and a label set with feature data for characterizing user risk characteristics, wherein the specified information set includes a labeled sample set and an unlabeled sample set, and the label set includes the risk categories corresponding to each labeled sample in the labeled sample set; constructing a classifier using the labeled sample set and the label set; extracting the neighboring labeled samples of each unlabeled sample in the unlabeled sample set, and calculating the information entropy corresponding to each unlabeled sample according to the distribution of the risk categories of the neighboring labeled samples of each unlabeled sample, wherein the neighboring labeled samples include the labeled samples in the labeled sample set whose proximity to the corresponding unlabeled sample in the user risk feature space meets a preset condition; the labeled samples whose proximity meets the preset condition include: the preset number of labeled samples closest to the unlabeled sample, or the labeled samples whose distance from the unlabeled sample is less than a preset distance; optimizing the classifier based on the information entropy corresponding to each unlabeled sample in the unlabeled sample set to obtain an optimized classifier, so as to allocate system resource data to the target user based on the risk prediction result of the target user by the optimized classifier; wherein, optimizing the classifier based on the information entropy corresponding to each unlabeled sample in the unlabeled sample set to obtain an optimized classifier includes: optimizing the classifier based on the neighboring discrimination matrix corresponding to the unlabeled sample set and the information entropy corresponding to each unlabeled sample to obtain an optimized classifier; wherein the risk categories include positive classes and negative classes, and correspondingly, calculating the information entropy corresponding to each unlabeled sample according to the distribution of the risk categories of the neighboring labeled samples of each unlabeled sample includes: calculating the information entropy of each unlabeled sample according to the following formula: Among them, is the i-th unlabeled sample, and H i is 's information entropy. N is the number of labeled samples that are the nearest neighbors of + is the number of positive-class samples among the N labeled samples that are the nearest neighbors of - is the number of negative-class samples among the N labeled samples that are the nearest neighbors of.
2. The method according to claim 1, characterized in that, after constructing the classifier using the labeled sample set and the label set, it further includes: performing risk classification on each unlabeled sample in the unlabeled sample set using the classifier to obtain the pseudo-labels corresponding to each unlabeled sample; calculating the neighboring discrimination matrix corresponding to the unlabeled sample set according to the similarities and differences between the pseudo-labels of each unlabeled sample in the unlabeled sample set and the risk categories of the neighboring labeled samples.
3. The method according to claim 2, characterized in that, optimizing the classifier based on the neighboring discrimination matrix corresponding to the unlabeled sample set and the information entropy corresponding to each unlabeled sample to obtain an optimized classifier includes: constructing a boundary enhancement constraint for the classifier to predict the user risk characteristics of the unlabeled samples in the unlabeled sample set using the information entropy and pseudo-labels corresponding to each unlabeled sample. Construct a nearest neighbor discrimination constraint for the classifier to predict user risk features for the unlabeled samples in the unlabeled sample set and the labeled samples in the labeled sample set by using the elements in the nearest neighbor discrimination matrix corresponding to the unlabeled sample set. Optimize the classifier based on the boundary enhancement constraint and the nearest neighbor discrimination constraint to obtain an optimized classifier.
4. The method according to claim 3, wherein, Constructing a boundary enhancement constraint for the classifier to predict user risk features for the unlabeled samples in the unlabeled sample set by using the information entropy and pseudo-labels corresponding to each unlabeled sample includes: Construct the boundary enhancement constraint according to the following formula: weight i = exp((H i - μ) / σ); Among them, R be is the boundary enhancement constraint, X U is the unlabeled sample set, |X U | is the number of unlabeled samples in X U . is the i-th unlabeled sample in X U , f(g) is the discriminant function of the classifier, is 's corresponding pseudo-label, weight i is 's boundary enhancement coefficient, H i is 's information entropy, μ is the mean of the information entropies of multiple unlabeled samples in X U , and σ is the standard deviation of the information entropies of multiple unlabeled samples in X U .
5. The method according to claim 2, wherein, Calculating the nearest neighbor discrimination matrix corresponding to the unlabeled sample set according to the similarities and differences between the pseudo-labels of the unlabeled samples in the unlabeled sample set and the risk categories of the nearest neighbor labeled samples includes: Determine the nearest neighbor discrimination matrix corresponding to the unlabeled sample set according to the following formula: Among them, S is the nearest neighbor discrimination matrix corresponding to the unlabeled sample set, and its dimension is |X U |×|X|, where X U is the unlabeled sample set, |X U | is the number of unlabeled samples in X U , X is the labeled sample set, |X| is the number of labeled samples in X, and s i,j is an element in the nearest neighbor discrimination matrix. is the i-th unlabeled sample in X U , x j is the j-th labeled sample in X. represents the neighboring labeled samples whose risk classes are the same as the pseudo-labels. represents the neighboring labeled samples whose risk classes are different from the pseudo-labels.
6. The method according to claim 3, wherein, Constructing a nearest neighbor discrimination constraint for the classifier to predict user risk features for the unlabeled samples in the unlabeled sample set and the labeled samples in the labeled sample set by using the elements in the nearest neighbor discrimination matrix corresponding to the unlabeled sample set includes: Construct the nearest neighbor discrimination constraint according to the following formula: Among them, R nd is the boundary enhancement constraint, X is the labeled sample set, |X| is the number of labeled samples in X, X U is the unlabeled sample set, |X U | is the number of unlabeled samples in X U , s i,j is an element in the nearest neighbor discrimination matrix, is the i-th unlabeled sample in X U , x j is the j-th labeled sample in X, and f(g) is the discrimination function of the classifier.
7. The method according to claim 3, wherein, Optimizing the classifier based on the boundary enhancement constraint and the nearest neighbor discrimination constraint to obtain an optimized classifier includes: Optimize the classifier according to the following formula: f * = argmin f L(f, X, Y, X U ) L = R emp + αR be + βR nd ; Among them, f * is the discriminant function corresponding to the optimized classifier, L(g) is the objective function, f is the discriminant function corresponding to the classifier, X is the labeled sample set, Y is the label set, X U is the unlabeled sample set, R emp is the empirical loss of the classifier for predicting user risk characteristics of the labeled samples in the labeled sample set, R be is the boundary enhancement constraint, R nd is the boundary enhancement constraint, |X| is the number of labeled samples in X, x j is the j-th labeled sample in X, y j is the j-th label in Y, that is, the risk category corresponding to x j corresponding to, α and β are hyperparameters.
8. The method according to claim 1, wherein, The feature data includes time series aggregation features and time series historical features; wherein, the time series aggregation features refer to data obtained by extracting features of the specified information of the user based on different time dimensions and time series feature extraction algorithms; the time series historical features include time series distribution data statistically obtained based on different time dimensions of the specified information of the user.
9. A system resource data allocation device, wherein, Applied to a server, the device includes: An acquisition module, configured to acquire a specified information set having feature data for characterizing user risk features and a label set, wherein the specified information set includes a labeled sample set and an unlabeled sample set, and the label set includes the risk categories corresponding to each labeled sample in the labeled sample set; A construction module, configured to construct a classifier by using the labeled sample set and the label set; A calculation module, configured to extract the neighboring labeled samples of each unlabeled sample in the unlabeled sample set, and calculate the information entropy corresponding to each unlabeled sample according to the distribution of the risk categories of the neighboring labeled samples of each unlabeled sample, where the neighboring labeled samples include the labeled samples in the labeled sample set whose proximity to the corresponding unlabeled sample in the user risk feature space meets a preset condition; the labeled samples whose proximity meets the preset condition include: the preset number of labeled samples closest to the unlabeled sample, or the labeled samples whose distance from the unlabeled sample is less than a preset distance; An optimization module, configured to optimize the classifier based on the information entropy corresponding to each unlabeled sample in the unlabeled sample set, so as to obtain an optimized classifier, and allocate system resource data to the target user based on the risk prediction result of the target user by the optimized classifier; Wherein, the optimization module is specifically configured to: optimize the classifier based on the neighboring discrimination matrix corresponding to the unlabeled sample set and the information entropy corresponding to each unlabeled sample, so as to obtain an optimized classifier; Wherein, the calculation module is specifically configured to: calculate the information entropy of each unlabeled sample according to the following formula: Among them, is the i-th unlabeled sample, and H i is 's information entropy. N is 's number of neighboring labeled samples. N + is 's number of positive-class samples among the N neighboring labeled samples. N - is 's number of negative-class samples among the N neighboring labeled samples.
10. A computer device, characterized in that, it includes a processor and a memory for storing processor-executable instructions, and when the processor executes the instructions, the steps of the method according to any one of claims 1 to 8 are implemented.
11. A computer-readable storage medium, on which computer instructions are stored, characterized in that, when the instructions are executed, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Object recognition method and system
CN109636430A
Multi-task supervised learning model training method and device and multi-task supervised learning model prediction method and device
CN109657696A