Interactive logistic regression modeling method for guaranteeing reasonable weight coefficient

Through logistic regression algorithm, the weight exception sub-model is identified and processed, and the training and construction strategies of the credit model are optimized, which solves the problem of weight imbalance in multi-model fusion, and improves the stability and generation efficiency of the credit model.

CN120337181AActive Publication Date: 2025-07-18HANGYIN CONSUMER FINANCE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510812270.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

In credit business scenarios, the difference in weight values of different sub-models when multi-models are fusion leads to excessive impact of a certain data source, affecting the stability and accuracy of the credit model.

Method used

The logistic regression algorithm determines the weight coefficient interval of the target submodel, identify the weight anomaly submodel, determines the exception handling strategy based on the correlation status and weight deviation of the credit data source, updates the credit model dynamically, optimizes the training sequence and construction strategy, eliminates the abnormal data source, and improves training efficiency.

Benefits of technology

It effectively reduces the impact of a single data source on the credit model, improves the efficiency and accuracy of the credit model generation, avoids the inefficiency problems caused by repeated training, dynamically updates the sub-model composition, and improves the stability of credit evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337181A_ABST
    Figure CN120337181A_ABST
Patent Text Reader

Abstract

The invention provides an interactive logistic regression modeling method for guaranteeing a reasonable weight coefficient, and belongs to the technical field of financial model training, and the method specifically comprises the steps: taking a sub-model which does not have an abnormal credit data source as a to-be-incorporated model; and determining an inclusion processing sequence of the to-be-included model according to the association condition of the to-be-included model and the credit data sources of different target sub-models and the deviation condition of the weight coefficients of the target sub-models and the weight coefficient intervals, and performing retraining processing on the credit model according to the inclusion processing sequence and the exception processing strategy. According to the method, the reconstruction strategy of the credit model is determined on the basis of the abnormal credit data source and the weight abnormal sub-model, and training is stopped until the training result meets the requirement or the reconstruction strategy is reached, so that the training processing efficiency is improved, and the influence degree of a certain single data source in the credit model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of financial model training, and particularly relates to an interactive logistic regression modeling method for ensuring reasonable weight coefficients. Background Art

[0002] In the credit business scenario, the model team will synchronously build and launch a large number of models to meet the needs of the strategy for multi-angle analysis of different business scenarios. The explosive growth of model scores makes the model fusion method of cross-validation no longer applicable, because it will cause dimensional explosion and the interpretability of the fusion logic is poor. For multi-model fusion, the common method in the industry at present is that the strategy analyst combines the business scenario to do target design and evaluation, and constructs a logistic regression model for fusion.

[0003] However, the following technical problems exist in the above technical solutions. The model scores of different sub-models are often constructed based on multiple credit data sources. Therefore, when building a credit model based on multi-model fusion, there will be differences in the weight values for different sub-models. As a result, it is possible that the number of sub-models constructed from a certain data source is relatively large and the sum of weights is large, which makes the entire credit model highly affected by a single data source, thus affecting the stability of the overall credit model.

[0004] To solve the above technical problems, the present application provides an interactive logistic regression modeling method for ensuring reasonable weight coefficients. Summary of the Invention

[0005] To achieve the object of the present invention, the present invention adopts the following technical solutions: Specifically, the present application provides an interactive logistic regression modeling method for ensuring reasonable weight coefficients, which specifically includes: S1 Determine the weight coefficient intervals of different target sub-models based on the coincidence of the credit data sources referred to by the target sub-models; S2 Determine the weight coefficients of different target sub-models in the credit model through the logistic regression algorithm. When the deviation of the weight coefficient from the weight coefficient interval does not meet the requirements, the target sub-model with the weight coefficient not within the weight coefficient interval is used as the weight abnormal sub-model; S3 Determine the abnormal processing strategy of the abnormal credit data source according to the composition data of the weight abnormal sub-models associated with different credit data sources and the deviation of the weight coefficient from the weight coefficient interval; S4 uses the sub - models without abnormal credit data sources as the sub - models to be incorporated. According to the association situation between the sub - models to be incorporated and the credit data sources of different target sub - models, as well as the deviation situation between the weight coefficient of the target sub - model and the weight coefficient interval, determine the incorporation processing order of the sub - models to be incorporated. Perform retraining processing on the credit model according to the incorporation processing order and the abnormal processing strategy. Determine the reconstruction strategy of the credit model based on the abnormal credit data sources and the weight - abnormal sub - models, and stop training until the training result meets the requirements or reaches the reconstruction strategy.

[0006] The beneficial effects of the present invention are as follows: According to the composition data of the weight - abnormal sub - models associated with different credit data sources and the deviation situation between the weight coefficient and the weight coefficient interval, determine the abnormal processing strategy for abnormal credit data sources, thus avoiding the technical problem that when the number of weight - abnormal sub - models associated with credit data sources is large and the deviation situation is serious, if no elimination processing is performed, it may lead to a low generation processing efficiency of the credit model. By specifically determining the elimination processing strategy, the dynamic update of the composition of the sub - models of the credit model is realized, and the generation processing efficiency is improved.

[0007] Determine the reconstruction strategy of the credit model based on the abnormal credit data sources and the weight - abnormal sub - models, thus avoiding the technical problem that when the number of abnormal credit data sources or the number of weight - abnormal sub - models is large, continuing to perform repeated training of the credit model may lead to a low update processing efficiency of the credit model. And specifically adopt the credit data sources without weight - abnormal sub - models, thereby realizing the reconstruction processing of the credit model, laying a foundation for further improving the generation processing efficiency, and at the same time reducing the influence degree of a single data source in the credit model, and avoiding the technical problem that when the data of a certain data source is missing, it has a greater impact on the accuracy of the overall credit assessment result.

[0008] A further technical solution is that the target sub - models are determined according to the random selection results of the training personnel, and specifically, a preset number of target sub - models are selected.

[0009] A further technical solution is that the coincidence situation of the credit data sources referred to by the target sub - models includes the number of coincidences of the credit data sources referred to by the target sub - models, where the referred credit data sources are the credit data sources called by the target sub - models when conducting credit risk assessment.

[0010] It can be understood that the credit data sources include the credit risk data of users, specifically including position, credit utilization record, overdue record, age, marital status, housing data, housing loan data, vehicle data, and vehicle loan data.

[0011] A further technical solution lies in that the method for determining the weight coefficient interval of the target sub-model is as follows: Based on the coincidence of the credit data sources referred to by the target sub-model, determine the number of coincidences between the credit data sources referred to by the target sub-model in different cases and the credit data sources of other target sub-models; Based on the number of coincidences in the credit data sources referred to in different cases, determine the total number of coincidences of the target sub-model; Determine the weight coefficient interval of the target sub-model according to the total number of coincidences.

[0012] A further technical solution lies in that the method for determining the reconstruction strategy of the credit model is as follows: Based on the number of abnormal credit data sources of the credit model, determine the proportion of the number of abnormal credit data sources in the credit data sources referred to by the sub-models of the credit model, and use it as the abnormal data source proportion; According to the weight abnormal sub-models corresponding to different abnormal credit data sources, determine the number of weight abnormal sub-models corresponding to different abnormal credit data sources; Based on the abnormal data source proportion and the number of weight abnormal sub-models corresponding to different abnormal credit data sources, determine the reconstruction strategy of the credit model.

[0013] A further technical solution lies in that the training is stopped until the training result meets the requirements or the reconstruction strategy is reached, which specifically includes: When the recognition accuracy of the credit risk of the credit model in different customer groups is greater than the preset accuracy threshold, it is determined that the training result meets the requirements; If the recognition accuracy of the credit risk of the credit model in different customer groups is not greater than the preset accuracy threshold and the credit model reaches the condition for reconstructing the credit model, the training process is stopped and the credit model is reconstructed.

[0014] Other features and advantages will be described in the following specification. The objectives and other advantages of the present invention are realized and obtained by the structures specifically pointed out in the specification and the drawings.

[0015] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following preferred embodiments are specifically given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] By referring to the accompanying drawings and describing its exemplary embodiments in detail, the above and other features and advantages of the present invention will become more obvious.

[0017] Figure 1It is a flowchart of an interactive logistic regression modeling method for ensuring reasonable weight coefficients; Figure 2 It is a flowchart of a method for determining the weight coefficient interval of the target sub-model; Figure 3 It is a flowchart of a method for determining that the deviation between the weight coefficient and the weight coefficient interval does not meet the requirements; Figure 4 It is a flowchart of a method for determining the abnormal handling strategy of the abnormal credit data source; Figure 5 It is a flowchart of a method for determining the inclusion processing order of the data to be included in the model. Detailed implementation mode

[0018] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.

[0019] Embodiment 1 As Figure 1 shown, this application provides an interactive logistic regression modeling method for ensuring reasonable weight coefficients, specifically including: S1 Based on the coincidence of the credit data sources cited by the target sub-model, determine the weight coefficient intervals of different target sub-models; Furthermore, the target sub-model is determined according to the random selection results of the training personnel, specifically selected according to a preset number of target sub-models.

[0020] Specifically, the coincidence of the credit data sources cited by the target sub-model includes the number of coincidences of the credit data sources cited by the target sub-model, where the cited credit data source is the credit data source called by the target sub-model when performing credit risk assessment.

[0021] It can be understood that the credit data source includes the credit risk data of users, specifically including position, credit usage record, overdue record, age, marital status, housing data, housing loan data, vehicle data, and vehicle loan data.

[0022] Specifically, as Figure 2 shown, the method for determining the weight coefficient interval of the target sub-model is: Determine the number of overlaps between the credit data sources cited by the target sub-model and the credit data sources of other target sub-models based on the overlap situation of the cited credit data sources of the target sub-model; Determine the total number of overlaps of the target sub-model based on the number of overlaps in different cited credit data sources; Determine the weight coefficient interval of the target sub-model according to the total number of overlaps.

[0023] It can be understood that based on the total number of overlaps of different target sub-models, perform normalization processing to obtain the recommended weight coefficients of different target sub-models. Centered on the recommended weight coefficients, determine the weight coefficient interval according to the preset value. For example, if the recommended weight coefficient is 0.1, the weight coefficient interval is 0.08 - 0.12.

[0024] Optionally, the method for determining the weight coefficient interval of the target sub-model is as follows: Determine the number of overlaps between the credit data sources cited by the target sub-model and the credit data sources of other target sub-models based on the overlap situation of the cited credit data sources of the target sub-model; Determine the recommended weight coefficients for different cited credit data sources based on the number of overlaps in different cited credit data sources; The recommended weight coefficients are obtained after normalization processing according to the number of overlaps in different cited credit data sources.

[0025] Normalize the sum of the recommended weight coefficients for different cited credit data sources, and then determine the weight coefficient interval of the target sub-model in combination with the preset value.

[0026] S2 determines the weight coefficients of different target sub-models in the credit model through the logistic regression algorithm. When the deviation between the weight coefficient and the weight coefficient interval does not meet the requirements, the target sub-model with the weight coefficient outside the weight coefficient interval is regarded as a weight abnormal sub-model; Furthermore, determining the weight coefficients of different target sub-models in the credit model through the logistic regression algorithm specifically includes: Set the initial weight coefficients of different target sub-models, iterate and update the weights in the training set, and adjust the parameters along the negative gradient direction of the loss function until the target iteration number is reached or the loss function meets the requirements, and then determine the weight coefficients of different target sub-models.

[0027] Specifically, as Figure 3 shown, determining that the deviation between the weight coefficient and the weight coefficient interval does not meet the requirements specifically includes: Regarding the target sub-model with the weight coefficient outside the weight coefficient interval as a weight abnormal sub-model; Determine the number of weight anomaly sub-models in different credit data sources according to the credit data sources cited by the weight anomaly sub-model; Determine whether the deviation of the weight coefficient from the weight coefficient interval meets the requirements according to the number of weight anomaly sub-models in different credit data sources.

[0028] It should be noted that if there is a credit data source where the number of weight anomaly sub-models is greater than the preset threshold of the number of anomaly sub-models, it is determined that the deviation of the weight coefficient from the weight coefficient interval does not meet the requirements.

[0029] In another embodiment, if the total number of weight anomaly sub-models in different credit data sources is more than 3 or the number of credit data sources with weight anomaly sub-models is more than 2, it is determined that the deviation of the weight coefficient from the weight coefficient interval does not meet the requirements, and the specific threshold number needs to be dynamically adjusted according to the number of target sub-models.

[0030] It can be understood that when the deviation of the weight coefficient from the weight coefficient interval meets the requirements, the credit model is used as the final credit model.

[0031] Optionally, determining that the deviation of the weight coefficient from the weight coefficient interval does not meet the requirements specifically includes: Regarding the target sub-model whose weight coefficient is not within the weight coefficient interval as a weight anomaly sub-model; Determine the number of weight anomaly sub-models in different credit data sources according to the credit data sources cited by the weight anomaly sub-model; Determine whether the deviation of the weight coefficient from the weight coefficient interval meets the requirements according to the number of credit data sources with weight anomaly sub-models.

[0032] S3 Determine the abnormal processing strategy for the abnormal credit data source according to the composition data of the weight anomaly sub-model associated with different credit data sources and the deviation of the weight coefficient from the weight coefficient interval; Specifically, as Figure 4 shown, the method for determining the abnormal processing strategy for the abnormal credit data source is: Based on the composition data of the weight anomaly sub-model associated with different credit data sources, determine the weight anomaly sub-model associated with the credit data source and use it as the associated abnormal sub-model; Determine the deviation amount from the adjacent endpoint of the weight coefficient according to the deviation of the weight coefficient of different associated abnormal sub-models from the weight coefficient interval; Determine the abnormal processing strategy for the credit data source through the deviation amounts of different associated abnormal sub-models.

[0033] It is understandable that when the number of associated abnormal sub - models with a deviation amount greater than the preset deviation amount threshold is more than 3, it is determined that the credit data source is an abnormal credit data source.

[0034] In another embodiment, the abnormal processing strategy of the credit data source is determined by the deviation amounts of different associated abnormal sub - models, specifically including: When the proportion of the number of sub - models associated with the credit data source by the associated abnormal sub - models is above the preset proportion threshold, the abnormal processing strategy of the credit data source is determined by using the first preset strategy; When the proportion of the number of sub - models associated with the credit data source by the associated abnormal sub - models is not above the preset proportion threshold but within a preset range, the abnormal processing strategy of the credit data source is determined by using the second preset strategy; When the proportion of the number of sub - models associated with the credit data source by the associated abnormal sub - models is not within the preset range, the abnormal processing strategy of the credit data source is determined by using the third preset strategy.

[0035] The first preset strategy is that after retraining, when there are associated abnormal sub - models in more than a target number of consecutive training processing times, all sub - models associated with the credit data source are removed from the credit model; The second preset strategy is that after retraining, when there are associated abnormal sub - models in more than a second target number of consecutive training processing times, all associated abnormal sub - models associated with the credit data source are removed from the credit model, and the third prediction strategy is that no abnormal processing is required.

[0036] In another embodiment, the method for determining the abnormal processing strategy of the abnormal credit data source is: Based on the composition data of the weight - abnormal sub - models associated with different credit data sources, the weight - abnormal sub - models associated with the credit data source are determined and used as associated abnormal sub - models. According to the deviation situation between the weight coefficients of different associated abnormal sub - models and the weight - coefficient interval, the deviation amount from the adjacent endpoint of the weight coefficient is determined; It is understandable that in the above steps, if the number of associated abnormal sub - models is large or the proportion of the number of associated abnormal sub - models in the associated sub - models is large, that is, greater than the threshold, at this time, it can be directly determined that the credit data source is an abnormal credit data source, and the abnormal processing strategy of the credit data source is determined by using the first preset strategy; In addition, it should be noted that when the number of associated abnormal sub-models is small or the proportion of associated abnormal sub-models in the associated sub-models is not large, it is necessary to determine the deviation of different associated abnormal sub-models at this time. When the number of associated abnormal sub-models with a deviation greater than the preset deviation does not meet the requirement, that is, when it is greater than the preset quantity threshold, the credit data source can be directly determined to be an abnormal credit data source, and the abnormal processing strategy of the credit data source can be determined through the first preset strategy; In addition, it can be understood that even when the number of associated abnormal sub-models with a deviation greater than the preset deviation meets the requirement, it is still necessary to further determine whether the average value of the proportion of the number of associated abnormal sub-models with a deviation greater than the preset deviation in the associated sub-models and the proportion of the number of associated abnormal sub-models in the associated sub-models is greater than the proportion threshold. When it is greater than the proportion threshold, the credit data source can be directly determined to be an abnormal credit data source, and the abnormal processing strategy of the credit data source can be determined through the first preset strategy; When the number of associated abnormal sub-models with a deviation greater than the preset deviation threshold is not more than 3, it is determined that the credit data source does not belong to an abnormal credit data source.

[0037] Take the other associated sub-models except the weight abnormal sub-models as associated sub-models, and take the associated sub-models with a deviation within the preset range from the adjacent endpoints of the adjacent weight coefficient intervals as critical sub-models; In addition, it can be understood that if there is no critical sub-model in the above steps, the abnormal processing strategy of the credit data source can be directly determined by using the second preset strategy at this time.

[0038] In addition, it should be noted that if there is a critical sub-model in the above steps, and the proportion of the number of critical sub-models and associated abnormal sub-models is above 0.6, the credit data source can be directly determined to be an abnormal credit data source, and the abnormal processing strategy of the credit data source can be determined through the first preset strategy; In addition, it should be noted that if the proportion of the number of critical sub-models and associated abnormal sub-models is above 0.6, for example, less than a certain threshold, the credit data source can be directly determined to be an abnormal credit data source, and the abnormal processing strategy of the credit data source can be determined through the second preset strategy; Determine the abnormal processing strategy of the credit data source through the deviation of different associated abnormal sub-models and the number of critical sub-models, and in combination with the number of sub-models of the credit data source.

[0039] In a possible embodiment, based on the deviation amounts of different associated anomaly sub-models, preset deviation weight coefficients corresponding to the deviation amounts of different associated anomaly sub-models are determined. An influence value is determined according to the sum of the preset deviation weight coefficients of different associated anomaly sub-models and the average value of the proportion of the number of the critical sub-model and the associated anomaly sub-models in the number of sub-models in the credit model.

[0040] Where when the influence value does not meet the requirement, that is, is greater than the preset influence threshold, the abnormal processing strategy for the credit data source at this time is the first preset strategy, and in other cases it is the second preset strategy.

[0041] S4 uses the sub-models without abnormal credit data sources as the models to be incorporated. According to the association situation between the models to be incorporated and the credit data sources of different target sub-models and the deviation situation between the weight coefficients of the target sub-models and the weight coefficient intervals, the incorporation processing order of the models to be incorporated is determined. The credit model is retrained according to the incorporation processing order and the abnormal processing strategy. A strategy for reconstructing the credit model is determined based on the abnormal credit data sources and the weight-abnormal sub-models, and the training is stopped until the training result meets the requirement or the reconstruction strategy is reached.

[0042] Specifically, as Figure 5 shown, the method for determining the incorporation processing order of the models to be incorporated is as follows: Based on the association situation between the models to be incorporated and the credit data sources of different target sub-models, the coincidence situation of the credit data sources cited by the models to be incorporated and different target sub-models is determined; The target sub-models with coincident cited credit data sources are determined using the coincidence situation and are used as the data-source coincidence models; According to the deviation situation between the weight coefficients of different data-source coincidence models and the weight coefficient intervals, the incorporation processing order of the models to be incorporated is determined.

[0043] It can be understood that determining the incorporation processing order of the models to be incorporated according to the deviation situation between the weight coefficients of the credit data sources of different data-source coincidence models specifically includes: Determining the incorporation processing order of the models to be incorporated in ascending order of the number of data-source coincidence models with weight coefficients not within the weight coefficient intervals.

[0044] Furthermore, the method for determining the strategy for reconstructing the credit model is as follows: Based on the number of abnormal credit data sources of the credit model, the proportion of the number of abnormal credit data sources in the credit data sources cited by the sub-models of the credit model is determined and used as the abnormal data-source proportion; Determine the number of weight anomaly sub-models corresponding to different abnormal credit data sources according to the weight anomaly sub-models corresponding to different abnormal credit data sources; Based on the proportion of the abnormal data sources and the number of weight anomaly sub-models corresponding to different abnormal credit data sources, determine the reconstruction strategy of the credit model.

[0045] It can be understood that determining the reconstruction strategy of the credit model based on the proportion of the abnormal data sources and the number of weight anomaly sub-models corresponding to different abnormal credit data sources specifically includes: When the proportion of the abnormal data sources is greater than the preset abnormal data source proportion threshold, it indicates that the number of abnormal data sources at this time is relatively large. Therefore, on this basis, if the number of abnormal data sources excluded in the credit model is above the preset data source number and there are still weight anomaly sub-models among the remaining abnormal data sources, then perform reconstruction processing on the credit model; When the proportion of the abnormal data sources is not greater than the preset abnormal data source proportion threshold, it indicates that the number of abnormal data sources at this time is relatively small. Therefore, on this basis, it is also necessary to determine whether the sum of the number of weight anomaly sub-models corresponding to different abnormal credit data sources is greater than the preset abnormal sub-model number threshold. If so, it indicates that the number of sub-models with abnormal weight coefficients is relatively large. Therefore, on this basis, if the sum of the number of weight anomaly sub-models corresponding to different abnormal credit data sources is still greater than the preset abnormal sub-model number threshold after retraining more than 3 times, then perform reconstruction processing on the credit model; and if there are still weight anomaly sub-models after retraining more than 10 times, then perform reconstruction processing on the credit model.

[0046] If the sum of the number of weight anomaly sub-models corresponding to different abnormal credit data sources is not greater than the preset abnormal sub-model number threshold, at this time, it is necessary to determine the retraining times threshold according to the proportion of the number of weight anomaly sub-models. When the number of retraining times is greater than the retraining times threshold and there are still weight anomaly sub-models, then perform reconstruction processing on the credit model.

[0047] The retraining times threshold is determined according to the preset times threshold corresponding to the proportion of the number of weight anomaly sub-models, or can also be determined according to the ratio of the preset value to the proportion of the number of weight anomaly sub-models. The larger the proportion of the number of weight anomaly sub-models, the smaller the retraining times threshold, and its value range is between 15 times and 20 times.

[0048] In another possible embodiment, the method for determining the reconstruction strategy of the credit model is: S41 determines the proportion of the number of abnormal credit data sources in the credit data sources referred to by the sub-models of the credit model based on the number of abnormal credit data sources of the credit model, and uses it as the proportion of abnormal data sources; It can be understood that in the above steps, it is necessary to determine whether the number of abnormal credit data sources and the proportion of abnormal data sources meet the requirements, that is, whether they are greater than the preset threshold. When any one of them does not meet the requirements, it means that the number of abnormal credit data sources in the credit model is relatively large. Therefore, on this basis, in order to improve the training and processing efficiency of the credit model, if the number of abnormal data sources excluded from the credit model is above the preset data source number and there are still weight-abnormal sub-models among the remaining abnormal data sources, the credit model will be reconstructed; In addition, it should be noted that in the above steps, when the number of abnormal credit data sources is within the preset data source number and the proportion of abnormal data sources is less than 0.1, it means that the number of abnormal credit data sources is relatively small. Therefore, in order to avoid excessive repeated construction of the credit model, only when the number of retraining times is greater than the retraining times threshold and there are still weight-abnormal sub-models, the credit model will be reconstructed, and in other cases, it will proceed to the next step.

[0049] S42 determines the number of weight-abnormal sub-models corresponding to different abnormal credit data sources according to the weight-abnormal sub-models corresponding to different abnormal credit data sources, and determines the attention data sources in the abnormal credit data sources based on the number of the weight-abnormal sub-models; It should be noted that in one of the embodiments, the attention data sources are abnormal credit data sources with the number of weight-abnormal sub-models being more than 5.

[0050] It can be understood that in the above steps, it is necessary to determine whether the sum of the numbers of weight-abnormal sub-models corresponding to different abnormal credit data sources is greater than the preset abnormal sub-model number threshold. If so, it means that the number of sub-models with abnormal weight coefficients is relatively large. Therefore, on this basis, if the sum of the numbers of weight-abnormal sub-models corresponding to different abnormal credit data sources is still greater than the preset abnormal sub-model number threshold after retraining more than 3 times, the credit model will be reconstructed; and if there are still weight-abnormal sub-models after retraining more than 10 times, the credit model will be reconstructed.

[0051] It can be further understood that if the sum of the numbers of weight anomaly sub-models corresponding to different abnormal credit data sources is not greater than a preset abnormal sub-model number threshold, it is also necessary to determine whether there is an attention data source at this time. When there is an attention data source, if the number of attention data sources is greater than a preset attention data source number threshold at this time, it indicates that the number of sub-models with abnormal weight coefficients is relatively large. Therefore, on this basis, if the sum of the numbers of weight anomaly sub-models corresponding to different abnormal credit data sources is still greater than the preset abnormal sub-model number threshold after retraining more than 3 times, the credit model will be reconstructed; if there are still weight anomaly sub-models after retraining more than 10 times, the credit model will be reconstructed, and in other cases, it will proceed to step S43.

[0052] S43 Determine the reconstruction strategy of the credit model based on the proportion of the abnormal data source, the proportion of the number of weight anomaly sub-models, and the composition data of the attention data source.

[0053] It should be noted that the weight anomaly value of the credit model is determined by the average of the proportion of the abnormal data source, the proportion of the number of weight anomaly sub-models, and the proportion of the attention data source in the data sources cited by the credit model.

[0054] It can be understood that when the weight anomaly value of the credit model is greater than the preset anomaly threshold, if the number of abnormal data sources excluded from the credit model is above the preset data source number and there are still weight anomaly sub-models among the remaining abnormal data sources, the credit model will be reconstructed; in other cases, it is necessary to determine whether the sum of the numbers of weight anomaly sub-models corresponding to different abnormal credit data sources is greater than the preset abnormal sub-model number threshold. If so, it indicates that the number of sub-models with abnormal weight coefficients is relatively large. Therefore, on this basis, if the sum of the numbers of weight anomaly sub-models corresponding to different abnormal credit data sources is still greater than the preset abnormal sub-model number threshold after retraining more than 3 times, the credit model will be reconstructed; if there are still weight anomaly sub-models after retraining more than 10 times, the credit model will be reconstructed.

[0055] If the sum of the numbers of weight anomaly sub-models corresponding to different abnormal credit data sources is not greater than the preset abnormal sub-model number threshold, it is necessary to determine the retraining times threshold according to the proportion of the number of weight anomaly sub-models at this time. When the retraining times are greater than the retraining times threshold and there are still weight anomaly sub-models, the credit model will be reconstructed.

[0056] It should be noted that reconstructing the credit model specifically includes: Based on a credit data source without a weight anomaly sub-model in the credit model, and using the sub-model corresponding to the credit data source without a weight anomaly sub-model as the basis, a re-construction process of the credit model is carried out.

[0057] Further, the training is stopped until the training result meets the requirements or reaches the re-construction strategy, which specifically includes: When the recognition accuracy of the credit risk of the credit model in different customer groups is greater than the preset accuracy threshold, it is determined that the training result meets the requirements; If the recognition accuracy of the credit risk of the credit model in different customer groups is not all greater than the preset accuracy threshold, and when the credit model meets the conditions for re-building the credit model, the training process is stopped and the credit model is re-built.

[0058] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the description of the method embodiments.

[0059] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multi-tasking and parallel processing are also possible or may be advantageous.

[0060] The above is only one or more embodiments of this specification and is not used to limit this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.

Claims

1. An interactive logistic regression modeling method for ensuring reasonable weight coefficients, characterized in that, Specifically, it includes: Determine the weight coefficient intervals of different target sub-models based on the coincidence of credit data sources referenced by the target sub-models; Determine the weight coefficients of different target sub-models in the credit model through the logistic regression algorithm. When the deviation of the weight coefficient from the weight coefficient interval does not meet the requirements, regard the target sub-model whose weight coefficient is not within the weight coefficient interval as a weight abnormal sub-model; Determine the abnormal processing strategy of the abnormal credit data source according to the composition data of the weight abnormal sub-model associated with different credit data sources and the deviation of the weight coefficient from the weight coefficient interval; Regard the sub-model without abnormal credit data source as the sub-model to be incorporated. Determine the incorporation processing order of the sub-model to be incorporated according to the association between the sub-model to be incorporated and the credit data sources of different target sub-models and the deviation of the weight coefficient of the target sub-model from the weight coefficient interval. Perform retraining processing on the credit model according to the incorporation processing order and the abnormal processing strategy. Determine the reconstruction strategy of the credit model based on the abnormal credit data source and the weight abnormal sub-model, and stop training until the training result meets the requirements or reaches the reconstruction strategy.

2. The interactive logistic regression modeling method for ensuring a reasonable weight coefficient as described in claim 1, wherein The target sub-model is determined according to the random selection result of the trainer.

3. The interactive logistic regression modeling method for ensuring a reasonable weight coefficient as described in claim 1, characterized in that, The coincidence of the credit data sources referenced by the target sub-model includes the number of coincidences of the credit data sources referenced by the target sub-model.

4. The interactive logistic regression modeling method for ensuring a reasonable guarantee weight coefficient according to claim 1, characterized in that, The credit data source includes the credit risk data of the user, specifically including position, credit utilization record, overdue record, age, marital status, housing data, housing loan data, vehicle data, and vehicle loan data.

5. The interactive logistic regression modeling method for ensuring a reasonable weight coefficient as described in claim 1, characterized in that, The method for determining the weight coefficient interval of the target sub-model is as follows: Based on the coincidence of the credit data sources referenced by the target sub-model, determine the number of coincidences of the target sub-model in different referenced credit data sources and the credit data sources of other target sub-models; Based on the number of coincidences in different referenced credit data sources, determine the total number of coincidences of the target sub-model; Determine the weight coefficient interval of the target sub-model according to the total number of coincidences.

6. The interactive logistic regression modeling method for ensuring a reasonable weight coefficient according to claim 1, characterized in that, Determine the weight coefficients of different target sub-models in the credit model through the logistic regression algorithm, specifically including: Set the initial weight coefficients of different target sub-models, iterate and update the weights in the training set, and adjust the parameters along the negative gradient direction of the loss function until the target iteration number is reached or the loss function meets the requirements, and then determine the weight coefficients of different target sub-models.

7. The interactive logistic regression modeling method for ensuring a reasonable weight coefficient as claimed in claim 1, characterized in that Determine that the deviation of the weight coefficient from the weight coefficient interval does not meet the requirements, specifically including: Regard the target sub-model whose weight coefficient is not within the weight coefficient interval as a weight abnormal sub-model; According to the credit data source referenced by the weight abnormal sub-model, determine the number of weight abnormal sub-models in different credit data sources; According to the number of weight abnormal sub-models in different credit data sources, determine whether the deviation of the weight coefficient from the weight coefficient interval meets the requirements.

8. The interactive logistic regression modeling method for ensuring a reasonable weight coefficient as claimed in claim 7, wherein, If there is a credit data source with the number of weight abnormal sub-models greater than the preset abnormal sub-model number threshold, it is determined that the deviation of the weight coefficient from the weight coefficient interval does not meet the requirements.

9. The interactive logistic regression modeling method for ensuring a reasonable weight coefficient as described in claim 1, characterized in that The method for determining the inclusion processing order of the model to be included is as follows: Based on the association between the model to be included and the credit data sources of different target sub-models, determine the coincidence of the credit data sources cited by the model to be included and different target sub-models; Use the coincidence situation to determine the target sub-models of the cited credit data sources with coincidence, and take them as the data source coincidence models; According to the deviation between the weight coefficients of different data source coincidence models and the weight coefficient intervals, determine the inclusion processing order of the model to be included.

10. The interactive logistic regression modeling method for ensuring a reasonable weight coefficient according to claim 1, characterized in that, The method for determining the reconstruction strategy of the credit model is as follows: Based on the number of abnormal credit data sources of the credit model, determine the proportion of the number of abnormal credit data sources in the credit data sources cited by the sub-models of the credit model, and take it as the abnormal data source proportion; According to the weight abnormal sub-models corresponding to different abnormal credit data sources, determine the number of weight abnormal sub-models corresponding to different abnormal credit data sources; Based on the abnormal data source proportion and the number of weight abnormal sub-models corresponding to different abnormal credit data sources, determine the reconstruction strategy of the credit model.

Citation Information

Patent Citations

  • Telecommunication fraud victim identification method and system and electronic equipment

    CN114548243A

  • Longitudinal logic regression modeling method based on anonymized data

    CN114662156A

  • Method for carrying out overdue probability modeling based on heterogeneous integration of different labels

    CN117391836A

  • Credit risk abnormity inspection attribution early warning method and system

    CN117853232A

  • Elm- and deep-forest-based hybrid model traffic anomaly detection system and method

    WO2024000944A1