Model training method for improving stability, application and initial parameter determination method

By constructing a mapping between recognition performance and generalization performance, the training process of machine learning models is optimized, solving the problems of long training time, high resource consumption and poor stability in existing technologies, and achieving efficient and stable training of models and improved generalization ability.

CN120910499APending Publication Date: 2025-11-07BAIRONG ZHIXIN (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510890061.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-07

Smart Images

  • Figure CN120910499A_ABST
    Figure CN120910499A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the field of machine learning, and particularly relates to a stability-improved model training method and application and an initial parameter determination method, and the training method comprises the steps: generating a training set and a verification set based on historical business data; obtaining each hyper-parameter group in a preset hyper-parameter group set of an initial model, performing iterative training on the initial model by using each hyper-parameter group, the training set and the verification set, and obtaining a test result of each hyper-parameter group after the iterative training meets a preset condition; and determining an optimal test result based on the test result of each hyper-parameter group, and determining the trained tuning model corresponding to the optimal test result as a target model. According to the embodiment of the invention, the problems of how to train a model with stable performance, how to apply the model to carry out risk prediction and how to provide initial parameters for the model are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present specification relate to the field of machine learning, and in particular to a model training method for improving stability, an application method and an initial parameter determination method. BACKGROUND

[0002] With the continuous progress of machine learning technology, machine learning is widely used in various industries to mine and process the internal correlation between data in various industries. In order to improve the generalization ability of machine learning and accurately capture the internal correlation (such as complex nonlinear relationship) in the training data, it is usually necessary to continuously adjust the parameters of the machine learning model. However, some existing parameter adjustment methods often rely on statistical analysis and expert experience, which is very time-consuming and prone to errors, and it is difficult to guarantee to find a global optimal solution, and it is easy to appear under-fitting or over-fitting problem, resulting in a long and inefficient training process of the model, high time cost and resource cost of training, and poor stability of the model. SUMMARY

[0003] Embodiments of the present specification provide a model training method and device for improving stability, which are used to solve or at least partially solve the problem of how to train a model with stable performance.

[0004] Embodiments of the present specification provide a model application method and device, which are used to solve or at least partially solve the problem of how to apply the model for risk prediction.

[0005] Embodiments of the present specification provide a model initial parameter determination method and device, which are used to solve or at least partially solve the problem of how to provide initial parameters for the model.

[0006] In order to solve the above technical problems, the first aspect of the embodiments of the present specification provides a model training method for improving stability, the training method comprising:

[0007] generating a training set and a validation set based on historical business data;

[0008] obtaining each hyperparameter group in a preset hyperparameter group set of an initial model, and using each hyperparameter group, the training set and the validation set to iteratively train the initial model, and obtaining a test result of each hyperparameter group after the iterative training meets a preset condition, wherein after each iteration of the initial model, a corresponding optimized model is obtained and a performance parameter is calculated, the performance parameter is used to adjust the model parameters of the optimized model, the performance parameter is determined based on a first performance value and a first difference between the first performance value and a second performance value, the first performance value and the second performance value respectively represent performance results of the optimized model for the validation set and the training set at the current training iteration, and the performance parameter represents the identification performance and the generalization performance of the optimized model.

[0009] Determine the best test result based on the test results of each of the super parameter groups, and determine the trained and optimized model corresponding to the best test result as the target model.

[0010] Further, the calculation performance parameter comprises:

[0011] When the first difference is greater than the preset threshold, the first difference is first adjusted to obtain a second difference;

[0012] When the first difference is not greater than the preset threshold, the first difference is second adjusted to obtain a second difference;

[0013] Combine the numerical value corresponding to the first performance value and the second difference to obtain the performance parameter.

[0014] Further, the first adjustment of the first difference comprises:

[0015] Determine the difference between the first difference and the preset threshold.

[0016] Further, the second adjustment of the first difference comprises:

[0017] Adjust the first difference to a default difference.

[0018] Further, the preset threshold is determined according to the following manner:

[0019] Determine the decay coefficient according to the current iteration number when training the initial model;

[0020] Adjust the initial threshold according to the decay coefficient to obtain the preset threshold.

[0021] Further, the preset threshold is determined according to the following formula:

[0022] overfit_th(t)=overfit_th0·e -kt ;

[0023] Wherein, overfit_th(t) represents the preset threshold at the tth iteration in the training process of the initial model, overfit_th0 represents the initial threshold, e -kt represents the decay coefficient, and k represents the self-defined coefficient.

[0024] Further, the test result of the super parameter group is obtained according to the following manner:

[0025] Obtain the first performance value corresponding to each training iteration of the initial model;

[0026] determine an optimal first performance value from all the first performance values, and determine the optimal first performance value as a test result of the hyperparameter group.

[0027] Further, the test result of the hyperparameter group is obtained according to the following manner:

[0028] obtain the first performance value corresponding to each training iteration of the initial model, and a first difference between the first performance value and the second performance value;

[0029] According to each first difference, filter all first performance values that meet the condition;

[0030] determine an optimal first performance value from all the first performance values that meet the condition, and determine the optimal first performance value and the corresponding first difference as the test result of the hyperparameter group.

[0031] Further, determining an optimal test result based on the test result of each hyperparameter group comprises:

[0032] determine a first mapping value corresponding to the first performance value and a second mapping value corresponding to the first difference of each hyperparameter group based on a preset rule;

[0033] determine a comprehensive mapping value of each hyperparameter group according to the first mapping value and the second mapping value;

[0034] determine the test result corresponding to the highest comprehensive mapping value of each hyperparameter group as the optimal test result.

[0035] The second aspect of the embodiments of the present specification provides a model application method, and the application method comprises:

[0036] obtain real-time business data of a target user;

[0037] input the real-time business data into a target model to obtain a risk prediction result.

[0038] The third aspect of the embodiments of the present specification provides a model initial parameter determination method, and the determination method comprises:

[0039] obtain a hyperparameter group corresponding to the optimal test result;

[0040] determine the hyperparameter group as an initial hyperparameter of a to-be-trained risk control model, and use the initial hyperparameter as a reference value when the to-be-trained risk control model is trained and fine-tuned.

[0041] The fourth aspect of the embodiments of the present specification provides a model training device for improving stability, and the training device comprises:

[0042] a generation module configured to generate a training set and a verification set based on historical business data;

[0043] The first determining module is configured to obtain each hyperparameter group in the preset hyperparameter group set of the initial model, and perform iterative training on the initial model by using each hyperparameter group, the training set and the verification set. After the iterative training meets a preset condition, a test result of each hyperparameter group is obtained. After each iteration of the initial model, a corresponding fine-tuned model is obtained, and a performance parameter is calculated. The performance parameter is used to adjust the model parameter of the fine-tuned model. The performance parameter is determined based on a first performance value and a first difference between the first performance value and a second performance value. The first performance value and the second performance value respectively represent performance results of the fine-tuned model in a current training iteration with respect to the verification set and the training set. The performance parameter represents the identification performance and the generalization performance of the fine-tuned model.

[0044] The second determining module is configured to determine a best test result based on the test results of each hyperparameter group, and determine a fine-tuned model trained according to the best test result as a target model.

[0045] The fifth aspect of the embodiments of the present specification provides a model application device, and the model application device comprises:

[0046] The first obtaining module is configured to obtain real-time service data of a target user.

[0047] The prediction module is configured to input the real-time service data into a target model to obtain a risk prediction result.

[0048] The sixth aspect of the embodiments of the present specification provides a model initial parameter determination device, and the determination device comprises:

[0049] The second obtaining module is configured to obtain a hyperparameter group corresponding to the best test result.

[0050] The third determining module is configured to determine the hyperparameter group as an initial hyperparameter of a to-be-trained risk control model. The initial hyperparameter is used as a reference value when the to-be-trained risk control model is fine-tuned.

[0051] The seventh aspect of the embodiments of the present specification provides a computer device, which comprises a memory, a processor and a computer program stored in the memory. When the computer program is run by the processor, instructions of the training method, the application and the initial parameter determination method for improving the stability of the model according to any one of the preceding embodiments are executed.

[0052] An eighth aspect of the embodiments of the present specification provides a computer storage medium, which stores a computer program. When the computer program is run by a processor of a computer device, instructions of the improved stability model training method, the application of the model, and the initial parameter determination method according to any one of the preceding embodiments are executed.

[0053] A ninth aspect of the embodiments of the present specification provides a computer program product, which comprises a computer program. When the computer program is run by a processor of a computer device, instructions of the improved stability model training method, the application of the model, and the initial parameter determination method according to any one of the preceding embodiments are executed.

[0054] The improved stability model training method and device provided by the embodiments of the present specification map the recognition performance and the generalization performance in the model training process by constructing the first performance value and the difference between the first performance value and the second performance value, and use the loss reference in the next iteration in the model training as a reference, so as to train a model that takes into account the recognition and generalization capabilities.

[0055] The application method and device of the model provided by the embodiments of the present specification input the real-time service data of the target user into the trained target model, and realize the risk control prediction of the target user behavior.

[0056] The initial parameter determination method and device of the model provided by the embodiments of the present specification obtain the hyperparameter group corresponding to the best test result, and use the hyperparameter group as the initial hyperparameter group of the risk control model for various risk control tasks, so as to accelerate the model convergence efficiency and improve the comprehensive performance of the model.

[0057] In order to make the above and other objects, features and advantages of the embodiments of the present specification more obvious and easy to understand, the following preferred embodiments are described in detail below, and the accompanying drawings are described as follows. BRIEF DESCRIPTION OF DRAWINGS

[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present specification or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present specification, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0059] Figure 1 An implementation system schematic diagram of the improved stability model training method of the embodiments of the present specification is shown;

[0060] Figure 2 A flowchart of the improved stability model training method of the embodiments of the present specification is shown;

[0061] Figure 3A flowchart showing the determination of performance parameters by embodiments of the present specification is shown;

[0062] Figure 4 A flowchart showing the determination process of the preset threshold by embodiments of the present specification is shown;

[0063] Figure 5 A first flowchart showing the determination of test results of the hyperparameter group by embodiments of the present specification is shown;

[0064] Figure 6 A second flowchart showing the determination of test results of the hyperparameter group by embodiments of the present specification is shown;

[0065] Figure 7 A flowchart showing the determination of the best test result by embodiments of the present specification is shown;

[0066] Figure 8 A flowchart showing the application method of the model by embodiments of the present specification is shown;

[0067] Figure 9 A flowchart showing the initial parameter determination method of the model by embodiments of the present specification is shown;

[0068] Figure 10 A structural diagram of the model training device for improving stability by embodiments of the present specification is shown;

[0069] Figure 11 A structural diagram of the application device of the model by embodiments of the present specification is shown;

[0070] Figure 12 A structural diagram of the initial parameter determination device of the model by embodiments of the present specification is shown;

[0071] Figure 13 A structural diagram of the computer equipment by embodiments of the present specification is shown.

[0072] Explanation of the drawing symbols:

[0073] 101, terminal;

[0074] 102, server;

[0075] 1010, generation module;

[0076] 1020, first determination module;

[0077] 1030, second determination module;

[0078] 1110, first acquisition module;

[0079] 1120, prediction module;

[0080] 1210, second acquisition module;

[0081] 1220. a third determining module;

[0082] 1302. a computer device;

[0083] 1304. a processor;

[0084] 1306. a memory;

[0085] 1308. a driving mechanism;

[0086] 1310. an input / output module;

[0087] 1312. an input device;

[0088] 1314. an output device;

[0089] 1316. a presentation device;

[0090] 1318. a graphical user interface;

[0091] 1320. a network interface;

[0092] 1322. a communication link;

[0093] 1324. a communication bus. DETAILED DESCRIPTION

[0094] The technical solutions in the embodiments of the present specification will be described clearly and completely in combination with the drawings in the embodiments of the present specification. Obviously, the described embodiments are only some of the embodiments of the present specification, but not all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the embodiments of the present specification.

[0095] It should be noted that the terms "first", "second", and the like in the present specification and claims and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present specification described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, device, product, or apparatus that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products, or apparatuses.

[0096] The specification provides method operation steps as described in the embodiments or flowcharts, but can include more or fewer operation steps based on conventional or non-creative labor. The order of steps listed in the embodiments is only one of the many step execution orders, and does not represent the only execution order. In actual system or device product execution, the method order shown in the embodiments or the drawings can be executed in sequence or in parallel.

[0097] It should be noted that the technical solutions of the embodiments of the present specification obtain the authorization of the relevant parties and comply with the relevant provisions of the national laws and regulations in the acquisition, storage, use, processing, etc. of data.

[0098] It should be noted that in the embodiments of the present specification, some software, components, models, etc. may be mentioned, which should be considered as exemplary, and the purpose is only to illustrate the feasibility of the implementation of the technical solutions of the embodiments of the present specification, but it does not mean that the applicant has or will necessarily use the scheme.

[0099] It should be noted that the embodiments of the present specification provide a model training method, application and initial parameter determination method for improving stability, and a device, taking the field of risk control as an example, but not only can be applied to the field of risk control, but also can be applied to target detection in autonomous driving, error analysis of log systems, security prediction of device runtime, behavior analysis of monitoring scenarios, etc. The present specification is not limited.

[0100] As Figure 1 The embodiment of the present application is a model training method, application and initial parameter determination method for improving stability, which can include: a terminal 101 and a server 102, the terminal 101 and the server 102 communicate through a network, the network can include a local area network (Local Area Network, LAN), a wide area network (Wide Area Network, WAN), the Internet or a combination thereof, and is connected to a website, a user device (such as a computing device) and a backend system. The staff can send the training, application, initial parameter determination request of the model to the server 102 through the terminal 101, the server 102 receives the training, application, initial parameter determination request of the model, calls the historical business data or real-time business data in the database for calculation and processing, obtains the training, application, initial parameter determination result, and sends the result to the terminal 101, so that the staff can process the business according to the result.

[0101] In the embodiments of the present disclosure, the server 102 can be a stand-alone physical server, a server cluster composed of multiple physical servers, or a distributed system, and can also be a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms.

[0102] In an optional embodiment, the terminal 101 can include, but is not limited to, a self-service terminal device, a desktop computer, a tablet computer, a notebook computer, a smart wearable device, and the like. Optionally, the operating system running on the electronic device can include, but is not limited to, an Android system, an IOS system, Linux, Windows, and the like. Of course, the terminal 101 is not limited to the above-mentioned electronic devices with certain entities, and can also be software running in the above-mentioned electronic devices.

[0103] In addition, it should be noted that, Figure 1 The above-mentioned is only one application environment provided by the present disclosure, and in actual application, a plurality of terminals 101 can be included, which is not limited in the present disclosure.

[0104] In an embodiment of the present disclosure, an improved stability model training method is provided to solve the problem of how to train a model with stable performance.

[0105] Specifically, as Figure 2 indicated, applied to the server side, the training method comprises:

[0106] Step 210, generating a training set and a verification set based on historical business data;

[0107] Step 220, obtaining each hyperparameter group in a preset hyperparameter group set of an initial model, and using each hyperparameter group, the training set and the verification set to iteratively train the initial model, and obtaining a test result of each hyperparameter group after the iterative training meets a preset condition;

[0108] Wherein, after each iteration of the initial model, an optimized model corresponding to the iteration is obtained and a performance parameter is calculated, the performance parameter is used to adjust the model parameters of the optimized model, the performance parameter is determined based on a first performance value and a first difference between the first performance value and a second performance value, the first performance value and the second performance value respectively represent performance results of the optimized model in the current training iteration for the verification set and the training set, and the performance parameter represents the identification performance and the generalization performance of the optimized model.

[0109] In step 230, the best test result is determined based on the test results of each of the super parameter groups, and the trained and optimized model corresponding to the best test result is determined as the target model.

[0110] The model training method for improving stability provided in this embodiment maps the recognition performance in the model training process by constructing a first performance value, maps the generalization performance in the model training process by constructing a difference between the first performance value and a second performance value, and optimizes the model based on the performance parameters representing the recognition performance and the generalization performance, so that the model considers both the loss of the recognition performance and the loss of the generalization performance in the training process, thereby ensuring that the model has stronger applicability in a complex business environment. This high applicability reflects the high recognition ability of the model for general features of multiple business scenarios (such as consumption, operation, and investment scenarios). Since the recognition of the model depends more on general features, the resistance to noise interference (non-general business features) can be improved, and higher stability is exhibited, thereby reducing the possibility of training overfitting. Taking the recognition performance and the generalization performance as two indicators for constraining model training can help the model balance in two directions in the training process, avoid falling into a local optimal solution of a single indicator, thereby facilitating the model to search for more directions, jump out of underfitting and find a global optimal solution, more greatly tap the potential of the super parameter group, greatly reduce the time cost of searching for super parameters meeting the requirements, shorten the computing power consumption of the trained model, and prolong the service life of the hardware devices (such as CPU, GPU, etc.) required for training the model.

[0111] In the embodiments of the present specification, the historical business data is divided into white sample data (i.e., data corresponding to normal business) and black sample data (i.e., data corresponding to abnormal business), and the historical business data includes account attributes, account assets, associated account information, and transaction feature information. The account attributes include user features such as nickname, professional attribute, identification code, and communication method. The account assets include the number of virtual asset products held (such as the number of loan products and the number of financial products). The associated account information includes accounts associated with the account (such as accounts having a kinship relationship or a frequent transaction relationship). The transaction feature information includes transaction frequency, number of transaction counterparts, transaction time, transaction channel, and transaction mode (such as installment and lump sum payment). The information, data, and signals involved in this embodiment are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data comply with relevant laws, regulations, and standards of relevant countries and regions.

[0112] The business risk data set can be constructed through historical business data. After setting the hyperparameters, the initial model is iteratively trained using the business risk data set to obtain the optimized model and the iteration result (i.e., the performance parameter) of each round of training. The first performance value in the performance parameter represents the performance of the optimized model on the validation set, which can reflect the current recognition performance of the optimized model. The second performance value in the performance parameter represents the performance of the optimized model on the training set, which can reflect the learning effect of the optimized model. The first difference (such as the difference) between the first performance value and the second performance value can reflect the generalization ability of the optimized model. Therefore, using the performance parameter as a reference for the next training iteration of the optimized model can constrain the model to consider both recognition ability and generalization ability.

[0113] The initial model can be various typical machine learning models. In some embodiments of the present specification, the XGBoost model is taken as an example, and the hyperparameters involved include learning rate (Learning_rate), number of trees (num_leaves), tree depth (max_depth), and minimum leaf sample size (min_data_in_leaf), etc. By taking the values of the hyperparameters and grouping them, a preset hyperparameter group set containing various hyperparameter groups can be obtained, and then a hyperparameter group is selected from the set as a to-be-tested hyperparameter group to constrain the training of the initial model.

[0114] Based on different hyperparameter groups, there are trained optimized models and test results corresponding to the trained optimized models. The test results include performance parameters, performance results of the training set, performance results of the validation set, etc. According to the test results corresponding to different hyperparameter groups, the best test result can be determined, and the optimized model corresponding to the best test result can be used as a target model.

[0115] In an embodiment of the present specification, as shown in Figure 3 The calculation of the performance parameter includes:

[0116] In step 310, when the first difference is greater than a preset threshold, the first difference is first adjusted to obtain a second difference.

[0117] In step 320, when the first difference is not greater than the preset threshold, the first difference is second adjusted to obtain a second difference.

[0118] In step 330, the performance parameter is obtained by combining the value corresponding to the first performance value and the second difference.

[0119] The embodiment determines the adjustment mode of the first difference according to the preset threshold value, so that the evaluation of the generalization ability of the optimized model is more flexible, and then the value corresponding to the first performance value is combined with the second difference to obtain a performance parameter for evaluating the comprehensive recognition ability and generalization ability.

[0120] When the first difference is greater than the preset threshold value, it indicates that the generalization ability of the optimized model is weak, and overfitting may occur, and the greater the first difference, the more unreliable the model (that is, there is a large difference between the training performance and the actual performance of the model), at this time, the first adjustment can be to magnify the first difference by a magnification factor, and the magnification factor can be determined according to different range intervals. The farther away from the preset threshold value, the greater the magnification factor, so that the penalty on the model is heavier. In some embodiments of the present specification, considering that when the first difference is greater than the preset threshold value, the difference between the first difference and the preset threshold value can be regarded as an intolerable error, and the part of the first difference that does not reach the preset threshold value can be regarded as a tolerable error, and then a first adjustment mode is provided: adjusting the difference between the first difference and the preset threshold value to make the difference between them smaller and smaller, that is, the first adjustment can be to determine the difference between the first difference and the preset threshold value as the second difference, and the part exceeding the preset threshold value is used as a loss to guide the model training.

[0121] When the first difference is less than or equal to the preset threshold value, it indicates that the generalization ability of the optimized model is strong, and the model is more reliable, and the smaller the first difference, the higher the stability of the model, at this time, the first adjustment can be to reduce the first difference by a reduction factor, and the reduction factor can be determined according to different range intervals. The farther away from the preset threshold value, the smaller the reduction factor, so that the penalty on the model is lighter. In some embodiments of the present specification, a second adjustment mode is also provided, considering that when the first difference does not reach the preset threshold value, the difference between the first difference and the preset threshold value can be regarded as a tolerable error, and then the first difference is directly adjusted to a default difference to obtain the second difference. The default difference can be a fixed loss or zero, and no corresponding loss is output.

[0122] By using the part exceeding the preset threshold value as a loss, the model can pay more attention to the improvement of the generalization performance in the training based on this part of the loss, at the same time, the part not exceeding the preset threshold value is counted into a constant loss or not counted into a loss, which provides a certain buffer space for the model and improves the training stability of the model.

[0123] In an embodiment of the present specification, the determination method of the preset threshold value is: attenuating (or shrinking) the initial threshold value of the initial model according to the current iteration number of the initial model, so that the preset threshold value gradually decreases with the increase of the iteration number. This design allows a larger overfitting difference in the early stage of optimization, so that the model has more freedom to explore the parameter space. With the progress of optimization, the threshold value is gradually tightened to strengthen the control of overfitting. For example,Figure 4 In the embodiment shown, the preset threshold is determined according to the following manner:

[0124] In step 410, the decay coefficient is determined according to the current iteration number during the initial model training;

[0125] In step 420, the initial threshold is adjusted according to the decay coefficient to obtain the preset threshold.

[0126] In this embodiment, the decay coefficient of the initial threshold is determined based on the current iteration number during the initial model training, which can flexibly adjust the initial threshold with the depth of training iteration, and improve the effectiveness of the performance parameter.

[0127] In some embodiments of the present specification, the preset threshold can be determined by the following formula:

[0128] overfit_th(t) = overfit_th0 e -kt ;

[0129] Wherein, overfit_th(t) represents the preset threshold at the tth iteration during the initial model training, overfit_th0 represents the initial threshold, e -kt represents the decay coefficient, k represents the custom coefficient, which can be customized as overfit_th0 ∈ [0.01, 0.2], k ∈ [0.005, 0.1] according to business requirements.

[0130] With the increase of each iteration t, the threshold gradually decreases. This design allows a larger overfitting difference in the early stage of optimization, so that the model has more freedom to explore the parameter space. With the progress of optimization, the threshold is gradually tightened to strengthen the control of overfitting, forcing the model to converge to a more generalized parameter group. Exponential decay can avoid optimization shock caused by sudden changes in threshold, and is suitable for parameters within the overfitting allowed range, allowing the model to freely explore high-performance parameters. When the threshold is adjusted automatically when it exceeds the limit, it avoids falling into the overfitting or underfitting state.

[0131] In some embodiments of the present specification, KS index is used as the performance result of the tuning model. The KS index can reflect the ability of the tuning model to distinguish between positive and negative samples, and therefore is a suitable index for measuring the identification performance of the tuning model. The KS index of the tuning model for the validation set, i.e. the first performance value, can represent the performance result of the tuning model for the validation set at the current training iteration. The KS index of the tuning model for the training set, i.e. the second performance value, can represent the performance result of the tuning model for the training set at the current training iteration.

[0132] The performance parameter of the tuning model can be represented in the form of a loss function as follows:

[0133] loss = (1-KSval )+(max(KS diff ,overfit_th)-overfit_th);

[0134] KS diff =KS train -KS val ;

[0135] wherein loss represents a performance parameter, KS train represents the KS index of the tuned model for the training set (i.e., the second performance value), KS val represents the KS index of the tuned model for the validation set (i.e., the first performance value), KS diff represents the first difference, and overfit_th represents a preset threshold.

[0136] The performance parameter can be considered as two parts, the first part focuses on investigating KS on the validation set. The larger the KS value is, the better the model performs in distinguishing positive and negative samples. Therefore, in the objective function, we convert the KS value by 1-KS val , so that the loss is smaller when the KS value is larger, thereby encouraging the model to exhibit stronger discrimination ability on the validation set.

[0137] The second part max(KS diff ,overfit_th)-overfit_th is a fine control of overfitting problem, where overfit_th represents the performance difference between the training set and the validation set, which reflects the overfitting degree of the model. By presetting an acceptable overfitting difference threshold, when the actual overfitting degree is lower than overfit_th, it indicates that the model is not overfitting or only slightly overfitting, so this part of loss is 0 and does not contribute to the total loss. However, when the actual overfitting degree exceeds the threshold, the excess part will be added as a regular term to the total loss as a punishment for the overfitting behavior of the model, which helps to automatically adjust the parameters in the optimization process to reduce the risk of overfitting and make the model more robust.

[0138] In an embodiment of the present specification, as shown in Figure 5 , the test result of the hyperparameter group is obtained in the following manner:

[0139] Step 510, obtaining a first performance value corresponding to each training iteration of the initial model;

[0140] Step 520, determining the best first performance value from all first performance values, and determining the best first performance value as the test result of the hyperparameter group.

[0141] The best first performance value in the first performance values of each iteration of the target model is determined as the test result of the hyperparameter group, and the test result is used to represent that the hyperparameter group can provide the best identification performance and generalization performance, thereby serving as a source for subsequent hyperparameter group screening.

[0142] In another embodiment of the present specification, as shown in Figure 6 The test result of the hyperparameter group is obtained in the following manner:

[0143] In step 610, the first performance value corresponding to each training iteration of the initial model is obtained, and the first difference between the first performance value and the second performance value is obtained.

[0144] In step 620, all first performance values meeting the condition are screened according to the first difference.

[0145] In step 630, the best first performance value is determined from all first performance values meeting the condition, and the best first performance value and the corresponding first difference are determined as the test result of the hyperparameter group.

[0146] In this embodiment, the performance result meeting the generalization performance requirement is screened according to the first difference, the best first performance value is screened from the first performance values meeting the requirement, the best first performance value and the corresponding first difference are determined as the test result of the hyperparameter group, and the test result capable of comprehensively representing the identification performance and the generalization performance is obtained.

[0147] In some embodiments of the present specification, by using the business risk data set and the performance parameter, the preset threshold is set to 0.03, and the following test results of the hyperparameter group are obtained:

[0148] Table 1

[0149]

[0150]

[0151] Based on the results in Table 1, one screening rule is to use the hyperparameters of the hyperparameter group 4 as the best hyperparameter group, and the corresponding test result is the best test result. In this case, although the KS of the validation set of the hyperparameter groups 5, 6 and 7 is higher than that of the hyperparameter group 4, the three groups of parameters will cause overfitting of the training set and the validation set, i.e., exceeding the threshold 0.03. The model trained by using the parameters of the hyperparameter groups 5, 6 and 7 will easily have a too much decreased effect when applied to real data in the future, thereby causing inconsistency between offline modeling and online actual application, and therefore the parameters of the groups are not selected. Meanwhile, in the case that the hyperparameter groups 1, 2, 3 and 4 are within the threshold range, the KS of the hyperparameter group 4 on the validation set is excellent, which fully verifies the stability and prediction ability of the model.

[0152] In another embodiment of the present specification, as shown in Figure 7 The method for determining the optimal test result based on the test results of each of the hyperparameter groups comprises:

[0153] In step 710, the first mapping value corresponding to the first performance value and the second mapping value corresponding to the first difference of each of the hyperparameter groups are determined based on the preset rule.

[0154] In step 720, the comprehensive mapping value of each hyperparameter group is determined according to the first mapping value and the second mapping value.

[0155] In step 730, the test result corresponding to the highest comprehensive mapping value among the comprehensive mapping values of each hyperparameter group is determined as the optimal test result.

[0156] In this embodiment, the performance result and the first difference of each hyperparameter group are mapped based on the preset rule, and the comprehensive mapping value of each hyperparameter group is determined according to the mapping value of the performance result and the first difference mapping value, and then the optimal test result is determined from the comprehensive mapping value that can reflect the identification performance and the generalization performance.

[0157] The preset rule is, for example, the mapping coefficient for the performance result and the first difference respectively (such as the mapping coefficient a for the performance result and the mapping coefficient b for the first difference). For a task with a relatively single application scenario, more attention is paid to the performance result, and the mapping coefficient for the performance result can be greater than the mapping coefficient for the first difference (i.e., a > b). For a task with a relatively complex and diverse application scenario, more attention is paid to the first difference, and the mapping coefficient for the first difference can be greater than the mapping coefficient for the performance result (i.e., b > a). For a scene without obvious inclination, the mapping coefficient for the performance result or the mapping coefficient for the first difference can be the same (i.e., a = b), so that the model can guarantee stable performance for simple or complex scenes in the actual use process.

[0158] In an embodiment of the present specification, a model application method is also provided, which is applied to the server side as described above, as shown in Figure 8 The application method comprises:

[0159] In step 810, real-time business data of a target user is obtained.

[0160] In step 820, the real-time business data is input into the target model to obtain a risk prediction result.

[0161] In this embodiment, the real-time business data of the target user is input into the trained target model to realize the risk control prediction of the target user behavior.

[0162] In some embodiments, the account attributes (user features such as nickname, professional attributes, identification codes, communication methods, etc.), account assets (the number of virtual asset products held (such as the number of loan products, the number of financial products, etc.)), associated account information (accounts having a kinship or frequent transaction relationship with the user to be predicted), and transaction feature information (transaction frequency, number of transaction counterparts, transaction time, transaction channel, transaction mode (such as installment, lump sum, etc.), etc.) of the user to be predicted in different scenarios (such as consumption, operation, and investment scenarios) can be input into the target model. The target model extracts relevant features according to the input business data and outputs stable risk prediction results based on the previously learned general features of each scenario.

[0163] Through the aforementioned trained target model, the internal correlation between real-time business data and risk prediction results can be obtained, and more accurate risk prediction results can be obtained. The reason why the trained target model can obtain more accurate risk prediction results is that the trained target model considers both the recognition performance in the training process of the first performance value mapping model and the generalization performance in the training process of the difference mapping model between the first performance value and the second performance value, and optimizes the model based on the performance parameters representing the recognition performance and the generalization performance. Therefore, the model considers both the loss of recognition performance and the loss of generalization performance during the training process, thereby ensuring that the model has stronger applicability in complex business environments. This high applicability reflects the high recognition ability of the model for general features in multiple business scenarios (such as consumption, operation, and investment scenarios). Since the recognition of the model relies more on general features, it can improve the resistance to noise interference (non-general business features) and exhibit high stability, thereby reducing the possibility of overfitting during training. By taking the recognition performance and the generalization performance as two indicators for constraining the model training, the model can balance the two directions during training and avoid falling into a local optimal solution of a single indicator, thereby facilitating the model to search for more directions, jump out of underfitting, and find a global optimal solution. This can tap the potential of the hyperparameter set to a greater extent, greatly reduce the time cost of searching for hyperparameters that meet the requirements, shorten the computational power consumption of the trained model, and prolong the service life of the hardware devices (such as CPUs and GPUs) required for training the model.

[0164] In an embodiment of the present specification, an initial parameter determination method of a model is also provided, which is applied to the server side as described above. Figure 9 As shown in the figure, the determination method comprises:

[0165] Step 910: obtaining the hyperparameter set corresponding to the best test result;

[0166] Step 920: determining the hyperparameter set as the initial hyperparameters of the to-be-trained risk control model, and using the initial hyperparameters as the reference value when fine-tuning the to-be-trained risk control model.

[0167] The embodiment sets the hyperparameter group corresponding to the best test result as the initial hyperparameter group of the to-be-trained model when a specific model needs to be trained and applied to other risk control tasks, and fine-tunes the initial hyperparameter group based on the training result of the to-be-trained model, thereby accelerating the convergence efficiency of the model, and further obtaining a risk control prediction model corresponding to the risk control task.

[0168] The specific model can be a model with similar or identical structure to the target model, so that the previously obtained hyperparameter group can be migrated to the specific model as the initial hyperparameter group, and the initial hyperparameter group is fine-tuned in the training process of the specific model until the specific model meets the performance requirement.

[0169] It should be noted that, in order to facilitate understanding of the outstanding effects of the present application, several mainstream hyperparameter tuning methods in model training are compared as follows:

[0170] (1) Manual search: Manual search requires developers to try different combinations of hyperparameters and select the best-performing model. This method is very time-consuming and prone to errors, because developers need to rely on their experience and intuition to tune, and it is difficult to guarantee finding the global optimal solution.

[0171] (2) Grid search: Grid search performs exhaustive search on the user-specified hyperparameter set, which will result in huge computation, especially in the case of large hyperparameter space. In addition, grid search is prone to local optimal solution because it tries all possible combinations, but may not be able to jump out of the current search range limit.

[0172] (3) Random search: Although random search is faster than grid search, it is still a computationally intensive method. In addition, the result of random search is difficult to guarantee global optimality because it relies on the random sampling process and may not cover all important hyperparameter combinations.

[0173] (4) Bayesian optimization: Bayesian optimization relies on learning the shape of the objective function, which requires some understanding of the model. If the objective function is very complex or difficult to model, the effectiveness of Bayesian optimization may be affected.

[0174] The model training and parameter determination methods described in the embodiments of this specification help the model balance recognition and generalization performance during training. By constraining the model to search for more possible directions through indicators in two directions, it avoids getting trapped in local optima due to statistical or expert experience, thus overcoming the limitations of manual search or grid search. Simultaneously, the parameter tuning methods described in the embodiments of this specification can fully explore the potential of hyperparameters to guide model training, avoiding the defect of random search failing to train the target model due to incomplete coverage of target hyperparameters. Furthermore, the performance parameter design in the parameter tuning methods described in the embodiments of this specification has high reproducibility, high portability, and high versatility. Compared to Bayesian optimization, which requires complex prior knowledge, it has stronger usability. Based on the same inventive concept, the embodiments of this specification also provide a model training device with improved stability, as described in the following embodiments. Since the principle of the model training device with improved stability is similar to that of the model training method with improved stability, the implementation of the model training device with improved stability can refer to the model training method with improved stability, and repeated details will not be repeated.

[0175] Specifically, such as Figure 10 As shown, the improved stability model training apparatus is applied to the aforementioned server side, including:

[0176] The generation module 1010 is used to generate training and validation sets based on historical business data.

[0177] The first determining module 1020 is used to obtain each hyperparameter group in the preset hyperparameter group set of the initial model, and to perform iterative training on the initial model using each hyperparameter group, the training set and the validation set. After the iterative training meets the preset conditions, the test results of each hyperparameter group are obtained.

[0178] Wherein, after each iteration of the initial model, the corresponding optimized model is obtained and performance parameters are calculated. The performance parameters are used to adjust the model parameters of the optimized model. The performance parameters are determined based on a first performance value and a first difference between the first performance value and a second performance value. The first performance value and the second performance value represent the performance results of the optimized model for the validation set and the training set during the current training iteration, respectively. The performance parameters represent the recognition performance and generalization performance of the optimized model.

[0179] The second determining module 1030 is used to determine the best test result based on the test results of each hyperparameter group, and to determine the trained and tuned model corresponding to the best test result as the target model.

[0180] Based on the same inventive concept, the embodiments of the present specification also provide a model application device, as described in the following embodiments. Since the principle of the model application device to solve the problem is similar to the model application method, the implementation of the model application device can be referred to the model application method, and the repeated parts will not be described herein.

[0181] Specifically, as shown in the following embodiments, the model application device is applied to the server side, and includes: Figure 11

[0182] The first acquisition module 1110 is configured to acquire real-time service data of a target user.

[0183] The prediction module 1120 is configured to input the real-time service data into a target model to obtain a risk prediction result.

[0184] Based on the same inventive concept, the embodiments of the present specification also provide a model initial parameter determination device, as described in the following embodiments. Since the principle of the model initial parameter determination device to solve the problem is similar to the model initial parameter determination method, the implementation of the model initial parameter determination device can be referred to the model initial parameter determination method, and the repeated parts will not be described herein.

[0185] Specifically, as shown in the following embodiments, the model initial parameter determination device is applied to the server side, and includes: Figure 12

[0186] The second acquisition module 1210 is configured to acquire a hyperparameter group corresponding to the best test result.

[0187] The third determination module 1220 is configured to determine the hyperparameter group as initial hyperparameters of a to-be-trained risk control model, and use the initial hyperparameters as a reference value when the to-be-trained risk control model is trained and fine-tuned.

[0188] The model training method and device for improving stability provided by the embodiments of the present specification map the recognition performance and the generalization performance in the model training process by constructing the first performance value and the difference between the first performance value and the second performance value, and use the loss as a reference for the next iteration in the model training, so as to train a model that takes into account the recognition and generalization capabilities.

[0189] The model application method and device provided by the embodiments of the present specification input the real-time service data of the target user into the trained target model, and achieve the risk control prediction of the target user behavior.

[0190] The model initial parameter determination method and device provided by the embodiments of the present specification acquire the hyperparameter group corresponding to the best test result, and use the hyperparameter group as the initial hyperparameter group of the risk control model for various risk control tasks, which can accelerate the model convergence efficiency and improve the comprehensive performance of the model. ​​

[0191] An embodiment of the present specification further provides a computer device for implementing the method described in any of the above embodiments. Figure 13 As shown in FIG. 13, a structure schematic diagram of a computer device in an embodiment of the present specification is shown, the computer device 1302 can include one or more processors 1304, such as one or more central processing units (CPUs), each of which can implement one or more hardware threads. The computer device 1302 can also include any memory 1306 for storing any kind of information, such as code, settings, data, etc. Without limitation, for example, the memory 1306 can include any one or a combination of the following: any type of RAM, any type of ROM, a flash memory device, a hard disk, an optical disk, etc. More generally, any memory can store information using any technology. Further, any memory can provide volatile or non-volatile retention of information. Further, any memory can represent a fixed or removable component of the computer device 1302. In one case, the computer device 1302 can perform any operation of the associated instructions when the processor 1304 executes the associated instructions stored in any memory or combination of memories. The computer device 1302 also includes one or more drive mechanisms 1308 for interacting with any memory, such as a hard disk drive mechanism, an optical disk drive mechanism, etc.

[0192] The computer device 1302 can also include an input / output module 1310 (I / O) for receiving various inputs (via input devices 1312) and for providing various outputs (via output devices 1314). One particular output mechanism can include a presentation device 1316 and an associated graphical user interface (GUI) 1318. In other embodiments, the input / output module 1310 (I / O), the input devices 1312, and the output devices 1314 can also not be included, just as a computer device in a network. The computer device 1302 can also include one or more network interfaces 1320 for exchanging data with other devices via one or more communication links 1322. One or more communication buses 1324 couple the above-described components together.

[0193] The communication links 1322 can be implemented in any manner, for example, through a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication links 1322 can include any combination of hardwired links, wireless links, routers, gateway functionality, name servers, etc., governed by any protocol or combination of protocols.

[0194] Corresponding to Figures 2 to 9The method is executed by the processor, and the method comprises the steps of: receiving a first image and a second image; determining a first feature point in the first image and a second feature point in the second image; determining a first feature point descriptor of the first feature point and a second feature point descriptor of the second feature point; determining a first feature point descriptor distance between the first feature point descriptor and the second feature point descriptor; determining a first feature point descriptor distance threshold; determining a first feature point descriptor distance ratio between the first feature point descriptor distance and the first feature point descriptor distance threshold; determining a first feature point descriptor distance ratio threshold; determining whether the first feature point descriptor distance ratio is greater than the first feature point descriptor distance ratio threshold; and determining a first feature point descriptor distance ratio threshold.

[0195] The computer readable storage medium stores a computer program, and the computer program is executed by the processor to perform the steps of the method. Figures 2 to 9 The computer readable storage medium stores a computer program, and the computer program is executed by the processor to perform the steps of the method.

[0196] It should be understood that the size of the sequence number of the above processes in various embodiments of the present specification does not mean the order of execution, and the execution order of the processes should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present specification.

[0197] It should also be understood that in the embodiments of the present specification, the term "and / or" is only a description of the association relationship of the associated objects, and can represent three relationships. For example, A and / or B can represent three cases of A alone, A and B together, and B alone. In addition, the character " / " in the present specification generally represents an "or" relationship between the front and rear associated objects.

[0198] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in the present specification can be realized by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present embodiments.

[0199] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0200] In several embodiments provided in the specification, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the above-described device embodiments are only illustrative, and for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can also be electrical, mechanical or other forms of connection.

[0201] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., can be located in one place or can be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments of the specification.

[0202] In addition, the functional units in each embodiment of the specification can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0203] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the embodiments of the specification essentially or say the part of the prior art that contributes to the technical solutions, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the specification. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0204] The principles and implementation manners of the specification are described in the specific embodiments in the specification, and the above embodiment description is only to help understand the method and core idea of the embodiments of the specification; at the same time, for those skilled in the art, according to the idea of the embodiments of the specification, the specific implementation manner and application range will be changed, and the above-mentioned content of the specification should not be understood as a limitation of the embodiments of the specification.

Claims

1. A model training method for improving stability, characterized by, The training method comprises: generating a training set and a verification set based on historical business data; obtaining each hyperparameter group in a preset hyperparameter group set of an initial model, and using each hyperparameter group, the training set and the verification set to iteratively train the initial model, and obtaining a test result of each hyperparameter group after the iterative training meets a preset condition, wherein, after each iteration of the initial model, a corresponding optimized model is obtained and a performance parameter is calculated, the performance parameter is used to adjust model parameters of the optimized model, the performance parameter is determined based on a first performance value, a first difference between the first performance value and a second performance value, the first performance value and the second performance value respectively represent performance results of the optimized model in a current training iteration for the verification set and the training set, and the performance parameter represents identification performance and generalization performance of the optimized model; determining a best test result based on the test results of each hyperparameter group, and determining an optimized model trained according to the best test result as a target model.

2. The method of claim 1, wherein, The calculation of the performance parameter comprises: when the first difference is greater than a preset threshold, first adjusting the first difference to obtain a second difference; when the first difference is not greater than the preset threshold, second adjusting the first difference to obtain a second difference; combining a numerical value corresponding to the first performance value and the second difference to obtain the performance parameter.

3. The method of claim 2, wherein, The first adjustment of the first difference comprises: determining a difference between the first difference and the preset threshold.

4. The method of claim 2, wherein, The second adjustment of the first difference comprises: adjusting the first difference to a default difference.

5. The method of claim 2 or 3, wherein, The preset threshold is determined according to the following manner: determining a decay coefficient according to a current iteration number in initial model training; adjusting an initial threshold according to the decay coefficient to obtain the preset threshold.

6. The method of claim 2, wherein, The preset threshold is determined according to the following formula: overfit_th(t) = overfit_th0 e -kt ; where overfit_th(t) represents a preset threshold at the tth iteration in the initial model training process, overfit_th0 represents an initial threshold, e -kt represents an attenuation coefficient, and k represents a self-defined coefficient.

7. The method of claim 1, wherein, The test result of the hyperparameter group is obtained according to the following manner: obtaining a first performance value corresponding to each training iteration of the initial model; determining a best first performance value from all first performance values, and determining the best first performance value as the test result of the hyperparameter group.

8. The method of claim 2, wherein, The test result of the hyperparameter group is obtained according to the following manner: obtaining a first performance value corresponding to each training iteration of the initial model, and a first difference between the first performance value and a second performance value; screening all first performance values meeting a condition according to each first difference; determining a best first performance value from all first performance values meeting the condition, and determining the best first performance value and a corresponding first difference as the test result of the hyperparameter group.

9. The method of claim 8, wherein, The determination of the best test result based on the test results of each hyperparameter group comprises: determining a first mapping value corresponding to a first performance value and a second mapping value corresponding to a first difference of each hyperparameter group based on a preset rule; determining a comprehensive mapping value of each hyperparameter group according to the first mapping value and the second mapping value; determining a test result corresponding to a highest comprehensive mapping value in the comprehensive mapping values of each hyperparameter group as the best test result.

10. A method of applying a model, characterized by, The application method comprises: obtaining real-time business data of a target user; Input the real-time business data into the target model determined based on the method in any one of claims 1 to 9 to obtain a risk prediction result.

11. A method of determining initial parameters of a model, characterized by, The determination method comprises: acquiring a hyperparameter group corresponding to the optimal test result determined based on the method in any one of claims 1 to 9; determining the hyperparameter group as initial hyperparameters of a to-be-trained risk control model, the initial hyperparameters being used as a benchmark value when the to-be-trained risk control model is fine-tuned.

12. A model training device for improving stability, characterized by, The training device comprises: a generation module configured to generate a training set and a verification set based on historical business data; a first determination module configured to acquire each hyperparameter group in a preset hyperparameter group set of an initial model, and perform iterative training on the initial model by using the each hyperparameter group, the training set and the verification set, to obtain a test result of each hyperparameter group after the iterative training meets a preset condition, wherein after each iteration of the initial model, a fine-tuned model corresponding to the iteration is acquired and a performance parameter is calculated, the performance parameter is used to adjust model parameters of the fine-tuned model, the performance parameter is determined based on a first performance value and a first difference between the first performance value and a second performance value, the first performance value and the second performance value respectively represent performance results of the fine-tuned model in a current training iteration with respect to the verification set and the training set, and the performance parameter represents identification performance and generalization performance of the fine-tuned model; a second determination module configured to determine an optimal test result based on the test results of the hyperparameter groups, and determine a fine-tuned model trained based on the optimal test result as a target model.

13. A model application apparatus characterized by comprising: The application device comprises: a first acquisition module configured to acquire real-time business data of a target user; a prediction module configured to input the real-time business data into the target model determined based on the method in any one of claims 1 to 9 to obtain a risk prediction result.

14. An initial parameter determination device of a model characterized by comprising: The determination device comprises: a second acquisition module configured to acquire a hyperparameter group corresponding to an optimal test result determined based on the method in any one of claims 1 to 9; a third determination module configured to determine the hyperparameter group as initial hyperparameters of a to-be-trained risk control model, the initial hyperparameters being used as a benchmark value when the to-be-trained risk control model is fine-tuned.

15. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method in any one of claims 1 to 11.

16. A computer-readable storage medium, the computer-readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor of the computer device to implement the method in any one of claims 1 to 11.

17. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor of the computer device to implement the method in any one of claims 1 to 11.