System resource adjustment method, device and equipment

Through clustering processing and label weight calculation methods, the problem of user stability prediction under the diversified financial service types is solved, the accuracy and efficiency of system resource adjustment are improved, and the overall performance of financial institutions system is enhanced.

CN112836749BActive Publication Date: 2025-05-16INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110147724.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-03
Publication Date
2025-05-16
Estimated Expiration
2041-02-03

AI Technical Summary

Technical Problem

In the case of diversified financial service types, it is difficult to efficiently and accurately predict user stability, which affects the accuracy and efficiency of system resource adjustment.

Method used

By obtaining the specified information set and tag set with feature data that characterizes user churn characteristics, performing clustering processing, counting the number of benchmark samples for user churn results in the cluster, determining the reference sample, calculating the tag weight based on the user churn results of the reference sample and the similarity of the predicted samples, and then determining the predicted tag value and adjusting the system resources.

Benefits of technology

It improves the accuracy and efficiency of system resource adjustment, enhances the overall performance of financial institution systems, and reduces user churn and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112836749B_ABST
    Figure CN112836749B_ABST
Patent Text Reader

Abstract

The embodiments of this specification relate to the field of artificial intelligence technology, and disclose a system resource adjustment method, device and equipment, the method comprising obtaining a specified information set and a label set having feature data for characterizing user churn characteristics; the specified information set comprises at least a plurality of prediction samples and a benchmark sample; the label set comprises the user churn results corresponding to the benchmark sample; clustering is performed on each prediction sample and the benchmark sample based on the feature data corresponding to the prediction sample and the benchmark sample to obtain a plurality of clusters; for any cluster, the benchmark sample corresponding to the user churn result whose number of benchmark samples meets the preset requirements is used as the reference sample of the corresponding cluster; the prediction label value of the corresponding prediction sample is determined according to the reference sample. When the stability value obtained by evaluating the target user based on the benchmark sample and the prediction sample associated with the prediction label value is lower than the preset stability value, the system resources provided to the target user are adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present specification relates to the field of artificial intelligence technology, and in particular, to a system resource adjustment method, device and equipment. Background Art

[0002] With the rapid development of big data service platform technology, financial service types and optional service channels are becoming more and more diverse and convenient, giving users more and more choices. Correspondingly, the users of financial institutions are becoming more and more unstable. In order to effectively retain users, the service systems of financial institutions usually need to perform a large amount of data analysis and resource adjustments to make the resources provided to users more in line with their needs. On this basis, user stability prediction in various application scenarios is very important for the accuracy of system resource adjustment.

[0003] At present, the commonly used user stability assessment method is mainly a classification method based on a supervised learning model. By modeling and analyzing the existing customer churn information, the trained model is used to predict the churn of new samples to determine the stability of each user. However, the classification method using a supervised learning model requires the use of information on known user churn results. However, with the diversification of financial service types, it is difficult to clearly define the churn results of users in many cases, resulting in the difficulty in efficiently and accurately selecting the sample data based on the prediction, which affects the accuracy of user stability prediction, and further affects the accuracy and efficiency of system resource adjustment. Therefore, a more accurate and efficient system resource adjustment method is urgently needed. Summary of the invention

[0004] The purpose of the embodiments of this specification is to provide a system resource adjustment method, device and equipment, which can improve the accuracy and efficiency of system resource adjustment.

[0005] This specification provides a system resource adjustment method, device and equipment which are implemented in the following ways:

[0006] A system resource adjustment method is applied to a server, the method comprising the following steps: obtaining a specified information set having feature data for characterizing user churn characteristics and a label set; the specified information set at least comprises a plurality of prediction samples and a benchmark sample; the label set comprises a user churn result corresponding to the benchmark sample; clustering the prediction samples and the benchmark sample in the specified information set based on the feature data corresponding to the prediction samples and the benchmark sample to obtain a plurality of clusters; for any cluster, counting the number of benchmark samples corresponding to the user churn result in the cluster, and selecting the user churn result whose number of benchmark samples meets the preset requirements. The corresponding benchmark sample is used as the reference sample of the corresponding cluster; the representative user churn result of the corresponding cluster is determined according to the user churn result corresponding to the reference sample of the corresponding cluster; for any predicted sample, the label weight of the corresponding predicted sample is determined according to the similarity between the predicted sample and the reference sample in the corresponding cluster; the predicted label value of the corresponding predicted sample is determined according to the representative user churn result of the cluster where the predicted sample is located and the label weight corresponding to the predicted sample; when the stability value obtained by evaluating the target user based on the benchmark sample and the predicted sample associated with the predicted label value is lower than the preset stability value, the system resources provided to the target user are adjusted.

[0007] In some other embodiments of the method provided in this specification, the use of benchmark samples corresponding to user churn results whose number of benchmark samples meets preset requirements as reference samples for corresponding clusters includes: using benchmark samples corresponding to user churn results with the largest number of benchmark samples within the cluster as reference samples for the corresponding clusters.

[0008] In some other embodiments of the method provided in this specification, determining the label weight of the corresponding predicted sample according to the similarity between the predicted sample and the reference sample in the corresponding cluster includes:

[0009]

[0010] Among them, s(x u ) is the label weight, x u is the prediction sample, x i is the i-th reference sample in the cluster, N c is the x in the corresponding cluster u The number of corresponding reference samples, γ is a hyperparameter used to adjust the similarity calculation.

[0011] In some other embodiments of the method provided in this specification, a user churn prediction model is constructed based on the benchmark samples and the prediction samples associated with the prediction label values; churn prediction is performed on the target user according to the user churn prediction model; and stability evaluation is performed on the target user using the churn prediction result of the target user to obtain a stability value of the target user.

[0012] In some other embodiments of the method provided in this specification, the user churn prediction model is constructed based on the following objective function, including:

[0013] L(f)=R emp (Y L ,f(X L ))+αR pemp (Y U ,S,f(X U ))+λR reg

[0014] Among them, L(f) is the objective function of the user churn prediction model, R emp (Y L ,f(X L )) represents the first loss function, Y L represents the set of user churn results corresponding to each benchmark sample in the specified information set, X L represents a set of feature data corresponding to each benchmark sample in the specified information set, R pemp (Y U ,S,f(X U )) is the second loss function, S represents the weight set composed of the label weights corresponding to each predicted sample in the specified information set, and Y U represents a set of user churn results corresponding to each prediction sample in the specified information set, X U represents a set of feature data corresponding to each prediction sample in the specified information set, R reg is the L2 regularization loss, f(·) is the discriminant function, and α and λ are hyperparameters.

[0015] In some other embodiments of the method provided in this specification, the feature data includes time series aggregate features and time series historical features. The time series aggregate features refer to data obtained by extracting features from the user's specified information based on different time dimensions and time series feature extraction algorithms; the time series historical features include time series distribution data obtained by statistics of the user's specified information based on different time dimensions.

[0016] In some other embodiments of the method provided in this specification, the designated information includes loan information and deposit information.

[0017] On the other hand, the embodiments of the present specification also provide a system resource adjustment device, which is applied to a server, and the device includes: an information acquisition module, which is used to obtain a specified information set with feature data for characterizing user churn characteristics, and a label set; the specified information set includes at least a plurality of prediction samples and a benchmark sample; the label set includes the user churn results corresponding to the benchmark samples; a clustering processing module, which performs clustering processing on each prediction sample and benchmark sample in the specified information set based on the feature data corresponding to the prediction sample and the benchmark sample to obtain a plurality of clusters; a reference sample determination module, which is used to count the number of benchmark samples corresponding to the user churn results in any cluster, and to determine the number of benchmark samples corresponding to the user churn results that meet the preset requirements; as a reference sample of the corresponding cluster; a first prediction module, used to determine the representative user churn result of the corresponding cluster according to the user churn result corresponding to the reference sample of the corresponding cluster; a weight determination module, used to determine the label weight of the corresponding prediction sample according to the similarity between the prediction sample and the reference sample in the corresponding cluster for any prediction sample; a label determination module, used to determine the predicted label value of the corresponding prediction sample according to the representative user churn result of the cluster where the prediction sample is located and the label weight corresponding to the prediction sample; a resource adjustment module, used to adjust the system resources provided to the target user when the stability value obtained by evaluating the target user based on the benchmark sample and the prediction sample associated with the prediction label value is lower than the preset stability value.

[0018] In some other embodiments of the device provided in this specification, the weight determination module is further used to determine the label weight of the prediction sample in the following manner:

[0019]

[0020] Among them, s(x u ) is the label weight, x u is the prediction sample, x i is the i-th reference sample in the cluster, N c is the x in the corresponding cluster u The number of corresponding reference samples, γ is a hyperparameter used to adjust the similarity calculation.

[0021] On the other hand, an embodiment of the present specification also provides a system resource adjustment device, which is applied to a server. The device includes at least one processor and a memory for storing processor executable instructions. When the instructions are executed by the processor, the steps of the method described in any one or more of the above embodiments are implemented.

[0022] The system resource adjustment method, device and equipment provided by one or more embodiments of the present specification can only determine the churn results of some users with obvious churn results, and the user churn results of other users who are difficult to determine whether they have churned can be determined first. Then, the characteristic data corresponding to the users whose churn results have been determined are used to construct the benchmark samples, and the samples corresponding to the users whose churn results have not been determined are used to construct the prediction samples. Clustering is performed on each prediction sample and benchmark sample in the specified information set, and the number of benchmark samples corresponding to each user churn result in the cluster is counted, and the benchmark samples corresponding to the user churn results whose number of benchmark samples meets the preset requirements are used as reference samples of the corresponding cluster, so as to determine the representative user churn results of the corresponding cluster according to the user churn results corresponding to the reference samples of the cluster. And the label weight of the corresponding prediction sample can be determined according to the similarity between the prediction sample and the reference sample in the corresponding cluster, so as to determine the predicted label value of the corresponding prediction sample by using the representative user churn results of the cluster where the prediction sample is located and the label weight corresponding to the prediction sample. Then, user stability evaluation is performed based on the benchmark sample and the prediction sample associated with the prediction label value, and then system resource adjustment is performed based on the stability evaluation result, so as to improve the efficiency and accuracy of system resource adjustment. At the same time, the overall performance of the financial institution's system can be further improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art description. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. In the drawings:

[0024] Figure 1 A flowchart of an embodiment of a system resource adjustment method provided in this specification;

[0025] Figure 2 A schematic diagram of a system resource adjustment method flow in one embodiment provided in this specification;

[0026] Figure 3 This is a schematic diagram of the module structure of another system resource adjustment device provided in this specification. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of the specification, not all of the embodiments. Based on one or more embodiments of the specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the embodiments of this specification.

[0028] In a scenario example provided in an embodiment of this specification, the system resource adjustment method can be applied to a device for performing system resource adjustment, and the device may include a server or a server cluster composed of multiple servers. For the target user, the server can extract feature data from various information of the target user as the test data of the target user, and then use a pre-configured algorithm or model to perform a stability assessment on the target user to obtain a stability assessment result of the target user, so as to adaptively adjust the resources of the financial institution based on the stability assessment result. The system resources may include data resources such as services and products provided or recommended to users. Usually, the above-mentioned data resources associated with each user will also occupy certain system hardware resources. By reasonably allocating the data resources associated with the user, the rationality of data resource allocation can be further improved. While retaining users, the overall performance of the financial institution service system can also be further improved.

[0029] Figure 1 It is a flowchart of an embodiment of the system resource adjustment method provided in this specification. Although this specification provides the method operation steps or device structure shown in the following embodiments or drawings, the method or device may include more or fewer operation steps or module units after partial merger based on routine or no creative labor. In the steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure shown in the embodiments or drawings of this specification. When the method or module structure is applied in an actual device, server or terminal product, it can be executed sequentially or in parallel according to the method or module structure shown in the embodiments or drawings (for example, in an environment of parallel processors or multi-threaded processing, or even including distributed processing, server cluster implementation environment). A specific implementation example Figure 1 As shown, in one embodiment of the system resource adjustment method provided in this specification, the method can be applied to the data processing device, and the method may include the following steps:

[0030] S20: Acquire a specified information set having feature data for characterizing user churn characteristics and a label set; the specified information set includes at least a plurality of prediction samples and a benchmark sample; the label set includes the user churn results corresponding to the benchmark sample.

[0031] The server can obtain a specified information set and a label set. The specified information set may include multiple sample data. The sample data may include feature data for characterizing user churn characteristics. Accordingly, the specified information set may be a data set composed of feature data for characterizing user churn characteristics. The feature data may be feature data extracted from business data of users stored in a business system of a financial institution. Feature extraction may be performed through feature engineering. The feature data extraction method and feature type may be set according to the actual application scenario, which is not limited here. Of course, it may also include feature data extracted from user information obtained by the server from a platform associated with a financial institution.

[0032] In some embodiments, the sample data may be a prediction sample or a benchmark sample. The prediction sample may be sample data of unknown user churn results. The benchmark sample may be sample data of known user churn results. Accordingly, the specified information set may include at least a plurality of prediction samples and benchmark samples. The label set may include user churn results corresponding to each benchmark sample. The user churn results may include, for example, user churn, user not churn, etc. For ease of processing, a single sample data may be set to correspond to a single user. Accordingly, feature data of users with known churn results and feature data of users with unknown churn results may be extracted respectively to construct benchmark samples and prediction samples. The feature data corresponding to the benchmark sample and the prediction sample are associated with the user identifier and stored in the specified information set. The user churn results corresponding to the benchmark sample are associated with the user identifier and stored in the label set.

[0033] In actual application scenarios, users may have applied for more than one business product or business service in the current financial institution. The types of business products or business services are also diverse, some of which may be continuous, such as deposits, etc., while some may be one-time applications, such as loans, financial management, etc. For different business products or business services, different user churn result determination methods may need to be formulated. For example, for deposits, if the balance in the user's account is lower than the preset balance threshold, and there is no capital flow in the user's account after a specified length of time, it can be considered that for deposit business, the user is a lost user. If the frequency of capital flow in the user's account is greater than the frequency threshold or the balance is greater than the balance threshold, the user can be considered as a non-lost user. For financial products, if the user's financial products in the current financial institution have all expired, and the user has not applied for any financial products after a period of time, it can be considered that for financial products, the user is a lost user. Alternatively, if the user's financial products in the current financial institution have not expired, the user can be considered as a non-lost user. Alternatively, the user's churn results can also be determined by combining multiple business products or business scenarios. Of course, the above-mentioned churn result determination method is only an example of preferred examples. In actual application scenarios, it can be flexibly configured as needed, and is not limited here.

[0034] The information characteristics of users corresponding to different products or services usually have great differences. It is also possible to construct a designated information set by distinguishing products or services, and then perform user stability prediction and system resource adjustment based on the corresponding designated information set, so that the prediction results can be more in line with the actual application scenario, thereby improving the prediction accuracy. For some new products or services, the number of corresponding users may be small. Accordingly, products or services with similar characteristics to the product or service can be obtained as designated products or designated services, and then the information of users corresponding to the designated product or designated service is obtained to construct a designated information set. Of course, the above implementation is only an example of preferred examples, and other methods of constructing designated information sets can also be used in specific implementations.

[0035] The pre-built information set can be stored locally or in a database. The server can extract the specified information set when adjusting system resources or building a prediction model. If the constructed information set refers to an information set composed of information about users corresponding to a specified product or a specified service scenario, an information set identifier can be set for each specified information set. Accordingly, the server can obtain the specified information set corresponding to the corresponding information set identifier according to the needs of the current test scenario for use in system resource adjustment under the current test scenario. Most of the business data in the business system is updated at a faster speed. Accordingly, the specified information set and label set can be dynamically updated at intervals to ensure the accuracy of the information in the information set.

[0036] S22: performing clustering processing on each prediction sample and reference sample in the specified information set based on feature data corresponding to the prediction sample and the reference sample to obtain a plurality of clusters.

[0037] The server can perform clustering processing on each prediction sample and benchmark sample in the specified information set based on the feature data corresponding to the prediction sample and the benchmark sample, and cluster each sample data in the specified information set into multiple clusters. For example, clustering algorithms such as K-means algorithm and DBSCAN (density-based clustering method) can be used to perform clustering processing on the feature data of each prediction sample and benchmark sample in the specified information set. For example, the spatial distance between the feature data of each sample can be calculated, and the degree of proximity of each sample in the user churn feature space can be determined based on the spatial distance, and multiple prediction samples and benchmark samples with certain similar user churn features are clustered as one cluster. The specific implementation method of clustering processing will not be described here.

[0038] S24: For any cluster, count the number of benchmark samples corresponding to the user churn results in the cluster, and use the benchmark samples corresponding to the user churn results whose number of benchmark samples meets a preset requirement as reference samples of the corresponding cluster.

[0039] For any cluster, the number of benchmark samples corresponding to each user churn result in the cluster can be counted first. For example, the number of benchmark samples corresponding to lost users and the number of benchmark samples corresponding to non-lost users. Then, the reference sample can be determined based on the number of benchmark samples corresponding to each user churn result. For example, the benchmark sample corresponding to the user churn result with the largest number of samples can be used as the reference sample. If there are more than two forms of user churn results, the benchmark samples corresponding to the two or more user churn results with the highest number of samples can also be used as the reference sample.

[0040] S26: Determine the representative user churn result of the corresponding cluster according to the user churn result corresponding to the reference sample of the corresponding cluster.

[0041] Then, the server may determine the representative user churn results of the corresponding clusters according to the user churn results corresponding to the reference samples.

[0042] In some embodiments, the benchmark sample corresponding to the user churn result with the largest number of samples can be used as a reference sample. Then, the representative user churn result of the corresponding cluster can be determined based on the user churn result with the largest number of benchmark samples. For example, assume that the user churn result includes two types: user churn and user not churn, which are marked as 1 and -1 respectively. Then, the number of benchmark samples marked as 1 and -1 in the cluster can be counted. If the number of benchmark samples marked as 1 is the largest, then the mark 1 can be used as the representative user churn result of the corresponding cluster.

[0043] Of course, if the reference sample is composed of benchmark samples corresponding to two or more user churn results ranked at the top in terms of sample quantity, the two or more user churn results can be combined to determine the representative user churn results of the corresponding cluster.

[0044] S28: For any predicted sample, determine a label weight of the corresponding predicted sample according to the similarity between the predicted sample and the reference sample in the corresponding cluster.

[0045] For any prediction sample, the similarity between the prediction sample and the reference sample can be calculated as the label weight of the corresponding prediction sample. In some embodiments, the similarity between the prediction sample and the reference sample can be calculated as the label weight of the corresponding prediction sample in the following manner:

[0046]

[0047] Among them, s(x u ) is the label weight, x u is the prediction sample, x i is the i-th reference sample in the cluster, N c is the x in the corresponding cluster u The number of corresponding reference samples, γ is a hyperparameter used to adjust the similarity calculation.

[0048] Of course, the above calculation method is only a preferred method. Other methods can also be used in practical applications. For example, the central value of each reference sample in any cluster can be counted, and then the distance between each predicted sample and the central value is calculated as the similarity, which is then used as the label weight of the corresponding predicted sample.

[0049] S210: Determine a prediction label value of a corresponding prediction sample according to the result representing user churn of the cluster where the prediction sample is located and the label weight corresponding to the prediction sample.

[0050] The server can further determine the predicted label value of the corresponding predicted sample based on the representative user churn result of the cluster where the predicted sample is located and the label weight corresponding to the predicted sample. Preferably, the product of the representative user churn result of the cluster where the predicted sample is located and the label weight can be used as the predicted label value of the corresponding predicted sample. Alternatively, the ratio of the representative user churn result of the cluster where the predicted sample is located to the label weight can also be used as the predicted label value of the corresponding predicted sample.

[0051] S212: When a stability value obtained by evaluating the target user based on the benchmark sample and the prediction sample associated with the prediction label value is lower than a preset stability value, adjust system resources provided to the target user.

[0052] The server may associate and store the prediction samples with the corresponding prediction label values. When predicting the stability of the target user, the user stability assessment may be performed based on the reference samples in the specified information set and the prediction samples associated with the prediction label values.

[0053] The server may adjust the system resources provided to the target user when the stability value obtained by evaluating the target user is lower than the preset stability value. The preset stability value may be preset according to the actual application scenario. The system resources may include data resources such as services and products provided or recommended to users. Usually, the above-mentioned data resources associated with each user will also occupy certain system hardware resources. By reasonably allocating the data resources associated with the user, the rationality of data resource allocation can be further improved. While retaining users, the overall performance of the financial institution service system can also be further improved.

[0054] In some other embodiments, the feature data may include time series aggregation features and time series historical features. The time series aggregation features may refer to data obtained by extracting features from the user's specified information based on different time dimensions and time series feature extraction algorithms. The time series historical features may include time series distribution data obtained by statistics of the user's specified information based on different time dimensions. The time dimensions may include, for example, the previous month, the previous two months, the previous three months, etc., as well as the previous second month, the previous third month, the previous fourth month, etc. The time series feature extraction algorithm may include, for example, the mean value, variance, standard deviation, etc. By further combining the time series feature information to construct the user's feature data, the characteristics of users of different churn types can be more accurately characterized, thereby improving the accuracy of system resource adjustment.

[0055] In some embodiments, the specified information may refer to loan information and / or deposit information, etc. Of course, the specified information may also refer to credit information of the user, etc. By performing time series feature analysis on information of the user that fluctuates significantly over time, horizontal analysis of user features can be achieved, thereby greatly improving the accuracy of user stability prediction.

[0056] In some implementations, the time series aggregation feature F agg It can be extracted in the following way:

[0057] F agg =[f(feature) time ,time=1,2,3,4,5,6,1-2,1-3,1-4,1-5,1-6]

[0058] f() takes Mean() average, Max() maximum, Min() minimum, Std() standard deviation respectively, and the time periods are the previous month, previous two months, previous three months, previous four months, previous five months, previous six months, previous second month, previous third month, previous fourth month, previous fifth month, previous sixth month. Accordingly, each deposit and loan feature derives 44-dimensional time series aggregation features.

[0059] Time series historical features F his It can be extracted in the following way:

[0060] F his =[feature time ,time=1,2,3,4,5,6]

[0061] The time periods are the first month, the second month, the third month, the fourth month, the fifth month, and the sixth month. Accordingly, each deposit and loan feature can derive a 6-dimensional time series historical feature.

[0062] Of course, other information features of the user can be further improved, and after being associated with the above-mentioned time series aggregation features and time series historical features, they can be used together as the user's feature data to predict user stability.

[0063] like Figure 2 As shown, in some other embodiments, the following method can also be used to predict user stability. A user churn prediction model is constructed based on the benchmark samples in the specified information set and the prediction samples associated with the prediction label value, and churn prediction is performed on the target user according to the user churn prediction model. The model construction algorithm can adopt a deep neural network, a convolutional network, etc.

[0064] Among them, the user churn prediction model can be constructed based on the following objective function:

[0065] L(f)=R emp (Y L ,f(X L ))+αR pemp (Y U ,S,f(X U ))+λR reg

[0066] Among them, L(f) is the objective function of the user churn prediction model, R emp (Y L ,f(X L )) represents the first loss function, Y L represents the set of user churn results corresponding to each benchmark sample in the specified information set, X Lrepresents a set of feature data corresponding to each benchmark sample in the specified information set, R pemp (Y U ,S,f(X U )) is the second loss function, S represents the weight set composed of the label weights corresponding to each predicted sample in the specified information set, and Y U represents a set of user churn results corresponding to each prediction sample in the specified information set, X U represents a set of feature data corresponding to each prediction sample in the specified information set, R reg is the L2 regularization loss, f(·) is the discriminant function, and correspondingly, f(X L ) is to process each benchmark sample based on the discriminant function, f(X U ) is to process each prediction sample based on the discriminant function, and α and λ are hyperparameters. By pre-building the model in the above manner, the user's churn probability can be predicted more quantitatively, thereby improving the accuracy of the user's stability prediction.

[0067] With the development of Internet finance, the cost for corporate clients to reselect financial service institutions is getting lower and lower. If the loss of corporate clients becomes more serious, it will have an adverse impact on financial institutions, resulting in a decline in the reputation of financial institutions and reduced profits. At the same time, financial institution systems may need to conduct large-scale service and product analysis and adjust data resources to obtain strategies that can retain users, which will further lead to a waste of hardware resources and costs of financial institution systems. Accordingly, in an implementation scenario provided in the embodiments of this specification, taking corporate clients as an example, the solution provided in the above embodiments is described as follows.

[0068] First, characteristic information related to legal person customer churn prediction is obtained from the data warehouse, including legal person basic information, legal person asset information, legal person loan information, and legal person transaction information. Data preprocessing and feature extraction are performed on the test samples, and the specified information set is constructed using the legal person's basic information characteristics and deposit and loan time series information characteristics.

[0069] Data selection. The relevant characteristics of corporate customer company deposits can be divided into four categories: basic information of legal persons, legal person asset information, legal person loan information, and legal person transaction information. The data scope can be determined by category, thereby determining the data table involved.

[0070] Data preprocessing. Observe the data columns in the data table that involve corporate client company deposit and loan information. Concatenate the data columns in different tables that involve corporate client company deposit information according to corporate client ID and time to form the original features. For columns with incorrect data types, convert them to the correct data types first. For example, the data type should be numeric, but a pseudo-string type is set in the data table. You can determine whether it is wrong based on the meaning of the data column name and convert the wrong ones. For columns with missing values, fill them in a certain way, such as missing values ​​of numerical features, fill them with "0", and missing values ​​of non-numeric features, fill them with "-1".

[0071] Then, feature extraction can be performed as follows.

[0072] Feature conversion: For categorical features, such as economic nature and enterprise size, they are One-Hot encoded, and some numerical features with particularly large ranges are bucketed.

[0073] Mining time series aggregate features. Use legal person asset information and legal person loan information to construct deposit and loan related time series features, including time series aggregate features and time series historical features. Among them, the time series aggregate feature F agg The construction method is as follows:

[0074] F agg =[f(feature) time ,time=1,2,3,4,5,6,1-2,1-3,1-4,1-5,1-6]

[0075] f() takes Mean() average, Max() maximum, Min() minimum, Std() standard deviation respectively, and the time periods are the previous month, previous two months, previous three months, previous four months, previous five months, previous six months, previous second month, previous third month, previous fourth month, previous fifth month, previous sixth month. Each deposit and loan feature derives a 44-dimensional time series aggregation feature.

[0076] Mining historical features of time series. Mining historical features of time series his The construction method is as follows,

[0077] F his =[feature time ,time=1,2,3,4,5,6]

[0078] The time periods are the first month, the second month, the third month, the fourth month, the fifth month, and the sixth month. Each deposit and loan feature derives a 6-dimensional time series historical feature.

[0079] The characteristic data of users with known churn results and the characteristic data of users with unknown churn results can be extracted respectively to construct a benchmark sample and a prediction sample. The characteristic data corresponding to the benchmark sample and the prediction sample are associated with the user identifier and stored in the specified information set. The user churn result corresponding to the benchmark sample is associated with the user identifier and stored in the label set.

[0080] Assume that the user churn results are divided into two types: user churned and user not churned, which are marked with "1" (positive) or "-1" (negative). In determining the user churn results, for the corporate customer deposit application scenario, "1" can be set to represent the corporate customer's deposit diversion inflow in the next month, and "-1" can be set to represent the corporate customer's deposit outflow. Based on the above setting rules, the churn results of some users can be determined in advance.

[0081] Given the number of clusters k, perform k-means clustering on the feature data of the benchmark samples and the predicted samples. Determine the corresponding cluster's representative user churn results based on the user churn results of the benchmark samples in the cluster. For example, if the number of positive samples in the cluster is greater than the number of negative samples, the positive samples in the cluster are used as reference samples, and the corresponding cluster "1" is used as the representative user churn result of the corresponding cluster. Then, the representative user churn result of the cluster can be used as the initial label y of each predicted sample in the corresponding cluster. u Otherwise, the negative samples in the cluster are used as reference samples, and the predicted samples are assigned "-1" as the initial label y u .

[0082]

[0083] Calculate label weights. For the predicted samples in the cluster, calculate the similarity between the predicted samples and the reference samples in the cluster as their corresponding label weights.

[0084] The predicted label value of the predicted sample can be determined according to the label weight and the initial label, and the predicted label value can be associated with the corresponding predicted sample. Then, the model is constructed based on the benchmark sample and the predicted sample associated with the predicted label value.

[0085] For the test data x corresponding to the target user, the test data can be input into the above-constructed model to obtain the output result. The result of "1" represents that the customer's balance will flow in next month, and the result of "-1" represents that the customer's balance will flow out next month.

[0086]

[0087] In the above way, based on the feature information at different time nodes, the time series feature information is constructed, so that the model can better take into account the previous feature information when learning the features of the current time node. Secondly, for samples with relatively vague user churn results, the churn result distribution of such samples can be accurately determined by fully mining the spatial distribution information between such samples and samples with known churn results. After that, the model is constructed by combining the two types of samples, which can improve the generalization performance of the model and make the model more accurate in predicting corporate customer churn. Then, using the model results, resources can be adjusted before corporate customers churn, reducing user churn and losses. At the same time, the overall performance of the financial system can also be improved.

[0088] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. For details, please refer to the description of the above-mentioned related processing related embodiments, and no further description is given here.

[0089] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0090] The system resource adjustment method provided by one or more embodiments of this specification can determine the churn results of some users with more obvious churn results, and for other users whose churn is more difficult to determine, the user churn results can be determined first. Then, the benchmark sample is constructed using the feature data corresponding to the users with determined churn results, and the prediction sample is constructed using the samples corresponding to the users with undetermined churn results. Then, the churn results of each user in the prediction sample can be estimated as the prediction label value of each prediction sample. Then, user stability assessment is performed based on the benchmark sample and the prediction sample associated with the prediction label value, and then system resource adjustment is performed based on the stability assessment result, thereby improving the efficiency and accuracy of system resource adjustment. At the same time, the overall performance of the financial institution system can be further improved.

[0091] Based on the system resource adjustment method described above, one or more embodiments of this specification also provide a system resource adjustment device. The device may include a system, software (application), module, component, server, etc. that uses the method described in the embodiments of this specification and is combined with necessary implementation hardware. Based on the same innovative concept, the device in one or more embodiments provided in the embodiments of this specification is as described in the following embodiments. Since the implementation scheme of the device to solve the problem is similar to the method, the implementation of the specific device in the embodiments of this specification can refer to the implementation of the aforementioned method, and the repetitions will not be repeated. As used below, the term "unit" or "module" can be a combination of software and / or hardware that implements predetermined functions. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived. Specifically, Figure 3 A schematic diagram of the module structure of an embodiment of a system resource adjustment device provided in the specification is shown as follows: Figure 3 As shown, applied to a server, the device may include:

[0092] The information acquisition module 302 can be used to obtain a specified information set and a label set having feature data for characterizing user churn characteristics; the specified information set includes at least a plurality of prediction samples and a benchmark sample; the label set includes the user churn results corresponding to the benchmark sample.

[0093] The clustering processing module 304 may be used to perform clustering processing on each prediction sample and reference sample in the specified information set based on feature data corresponding to the prediction sample and the reference sample to obtain a plurality of clusters.

[0094] The reference sample determination module 306 can be used to count the number of benchmark samples corresponding to the user churn results in any cluster, and use the benchmark samples corresponding to the user churn results whose number of benchmark samples meets preset requirements as reference samples of the corresponding cluster.

[0095] The first prediction module 308 may be used to determine the representative user churn result of the corresponding cluster according to the user churn result corresponding to the reference sample of the corresponding cluster.

[0096] The weight determination module 310 can be used to determine the label weight of the corresponding prediction sample according to the similarity between the prediction sample and the reference sample in the corresponding cluster for any prediction sample.

[0097] The label determination module 312 may be used to determine the predicted label value of the corresponding predicted sample according to the representative user churn result of the cluster where the predicted sample is located and the label weight corresponding to the predicted sample.

[0098] The resource adjustment module 314 may be configured to adjust the system resources provided to the target user when a stability value obtained by evaluating the target user based on the benchmark sample and the prediction sample associated with the prediction label value is lower than a preset stability value.

[0099] It should be noted that the above-mentioned device may also include other implementations according to the description of the method embodiment. The specific implementation methods can refer to the description of the relevant method embodiment, and will not be described one by one here.

[0100] This specification also provides a system resource adjustment device, which can be applied to a separate system resource adjustment system or to a variety of computer data processing systems. The system can be a separate server, or it can include a server cluster, system (including distributed system), software (application), actual operation device, logic gate circuit device, quantum computer, etc. that uses one or more of the methods or one or more embodiments of this specification and a terminal device combined with necessary implementation hardware. In some embodiments, the device may include at least one processor and a memory for storing processor executable instructions, and when the instructions are executed by the processor, the steps of the method described in any one or more of the above embodiments are implemented.

[0101] The memory may include a physical device for storing information, which is usually to digitize the information and then store it in a medium using electrical, magnetic or optical means. The storage medium may include: a device that uses electrical energy to store information, such as various memories, such as RAM, ROM, etc.; a device that uses magnetic energy to store information, such as a hard disk, a floppy disk, a magnetic tape, a magnetic core memory, a magnetic bubble memory, a USB flash drive; a device that uses optical means to store information, such as a CD or a DVD. Of course, there are other types of readable storage media, such as quantum memory, graphene memory, etc.

[0102] It should be noted that the above-mentioned device may also include other implementation methods according to the description of the method or device embodiment. The specific implementation method can refer to the description of the relevant method embodiment, which will not be described one by one here.

[0103] It should be noted that the embodiments of this specification are not limited to the situations that must comply with the standard data model / template or the embodiments of this specification. Certain industry standards or slightly modified implementation plans based on the implementation described in the custom method or the embodiment can also achieve the same, equivalent or similar, or predictable implementation effects after deformation of the above-mentioned embodiments. The embodiments obtained by using these modified or deformed data acquisition, storage, judgment, processing methods, etc. can still fall within the scope of the optional implementation plans of this specification.

[0104] Each embodiment in this specification is described in a progressive manner, and the same and similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. In the description of this specification, the description of the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representation of the above terms does not necessarily target the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, in the absence of contradiction, a person skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0105] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.

Claims

1. A system resource adjustment method, characterized in that: Applied to a server, the method comprises: Acquire a specified information set having characteristic data for characterizing user churn characteristics and a label set; the specified information set includes at least a plurality of prediction samples and a reference sample; the label set includes the user churn results corresponding to the reference sample; Performing clustering processing on each prediction sample and the reference sample in the specified information set based on feature data corresponding to the prediction sample and the reference sample to obtain a plurality of clusters; For any cluster, the number of benchmark samples corresponding to the user churn results in the cluster is counted, and the benchmark samples corresponding to the user churn results whose number of benchmark samples meets the preset requirements are used as reference samples of the corresponding cluster; Determine the representative user churn result of the corresponding cluster according to the user churn result corresponding to the reference sample of the corresponding cluster; For any predicted sample, determining the label weight of the corresponding predicted sample according to the similarity between the predicted sample and the reference sample in the corresponding cluster; Determine the predicted label value of the corresponding predicted sample according to the representative user churn result of the cluster where the predicted sample is located and the label weight corresponding to the predicted sample, so as to adjust the system resources provided to the target user when the stable value obtained by evaluating the target user based on the benchmark sample and the predicted sample associated with the predicted label value is lower than the preset stable value; The method further includes: constructing a user churn prediction model based on the benchmark sample and the prediction sample associated with the prediction label value; performing churn prediction on the target user according to the user churn prediction model; performing stability evaluation on the target user using the churn prediction result of the target user to obtain a stability value of the target user; Specifically, the user churn prediction model is constructed based on the following objective function, including: L(f)=R emp (Y L ,f(X L ))+αR pemp (Y U ,S,f(X U ))+λR reg Among them, L(f) is the objective function of the user churn prediction model, R emp (Y L ,f(X L )) represents the first loss function, Y L represents the set of user churn results corresponding to each benchmark sample in the specified information set, X L represents a set of feature data corresponding to each benchmark sample in the specified information set, R pemp (Y U ,S,f(X U )) is the second loss function, S represents the weight set composed of the label weights corresponding to each predicted sample in the specified information set, and Y U represents a set of user churn results corresponding to each prediction sample in the specified information set, X U represents a set of feature data corresponding to each prediction sample in the specified information set, R reg is the L2 regularization loss, f(·) is the discriminant function, and α and λ are hyperparameters.

2. The method according to claim 1, characterized in that The step of using the benchmark samples corresponding to the user churn results whose number of benchmark samples meets the preset requirement as reference samples for the corresponding clusters includes: The benchmark sample corresponding to the user churn result with the largest number of benchmark samples in the cluster is used as the reference sample of the corresponding cluster.

3. The method according to claim 1, characterized in that The step of determining the label weight of the corresponding predicted sample according to the similarity between the predicted sample and the reference sample in the corresponding cluster includes: Among them, s(x u ) is the label weight, x u is the prediction sample, x i is the i-th reference sample in the cluster, N c is the x in the corresponding cluster u The number of corresponding reference samples, γ is a hyperparameter used to adjust the similarity calculation.

4. The method according to claim 1, characterized in that: The feature data includes time series aggregation features and time series historical features; wherein the time series aggregation features refer to data obtained by extracting features from the user's specified information based on different time dimensions and a time series feature extraction algorithm; the time series historical features include time series distribution data obtained by statistics of the user's specified information based on different time dimensions.

5. The method according to claim 4, characterized in that The designated information includes loan information and deposit information.

6. A system resource adjustment device, characterized in that: Applied to a server, the device comprises: An information acquisition module, used to acquire a specified information set having characteristic data for characterizing user churn characteristics, and a label set; the specified information set includes at least a plurality of prediction samples and a benchmark sample; the label set includes the user churn results corresponding to the benchmark sample; A clustering processing module, used for performing clustering processing on each prediction sample and the reference sample in the specified information set based on the feature data corresponding to the prediction sample and the reference sample to obtain a plurality of clusters; A reference sample determination module is used to count the number of benchmark samples corresponding to the user churn results in any cluster, and use the benchmark samples corresponding to the user churn results whose number of benchmark samples meets the preset requirements as reference samples of the corresponding cluster; A first prediction module, used to determine a representative user churn result of a corresponding cluster according to a user churn result corresponding to a reference sample of the corresponding cluster; A weight determination module, used for determining, for any predicted sample, a label weight of the corresponding predicted sample according to the similarity between the predicted sample and the reference sample in the corresponding cluster; A label determination module, used to determine the predicted label value of the corresponding predicted sample according to the representative user churn result of the cluster where the predicted sample is located and the label weight corresponding to the predicted sample; A resource adjustment module, configured to adjust system resources provided to the target user when a stability value obtained by evaluating the target user based on the benchmark sample and the prediction sample associated with the prediction label value is lower than a preset stability value; The device is further used to construct a user churn prediction model based on the benchmark sample and the prediction sample associated with the prediction label value; perform churn prediction on the target user according to the user churn prediction model; perform stability evaluation on the target user using the churn prediction result of the target user to obtain the stability value of the target user; Specifically, the device constructs the user churn prediction model based on the following objective function: L(f)=R emp (Y L ,f(X L ))+αR pemp (Y U ,S,f(X U ))+λR reg Among them, L(f) is the objective function of the user churn prediction model, R emp (Y L ,f(X L )) represents the first loss function, Y L represents the set of user churn results corresponding to each benchmark sample in the specified information set, X L represents a set of feature data corresponding to each benchmark sample in the specified information set, R pemp (Y U ,S,f(X U )) is the second loss function, S represents the weight set composed of the label weights corresponding to each predicted sample in the specified information set, and Y U represents a set of user churn results corresponding to each prediction sample in the specified information set, X U represents a set of feature data corresponding to each prediction sample in the specified information set, R reg is the L2 regularization loss, f(·) is the discriminant function, and α and λ are hyperparameters.

7. The device according to claim 6, characterized in that The weight determination module is also used to determine the label weight of the predicted sample in the following manner: Among them, s(x u ) is the label weight, x u is the prediction sample, x i is the i-th reference sample in the cluster, N c is the x in the corresponding cluster u The number of corresponding reference samples, γ is a hyperparameter used to adjust the similarity calculation.

8. A system resource adjustment device, characterized in that: Applied to a server, the device includes at least one processor and a memory for storing processor executable instructions, and the instructions, when executed by the processor, implement the steps of the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Track predication method based on Gauss mixture time series model

    CN107610464A

  • Method and apparatus for predicting customer stability, computer device and storage medium

    CN109376237A