Enterprise credit model optimization method and system based on big data
By optimizing the training data of the enterprise credit model and selecting adaptive analysis methods based on coverage completeness and data uniformity, the problem of imbalanced training data was solved, the stability and accuracy of the model were improved, and more accurate credit assessment was achieved.
Patent Information
- Application Number
- CN202510318109.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-03-18
AI Technical Summary
Existing corporate credit models exhibit poor stability and reliability when training data is imbalanced, and cannot adaptively adjust data supplementation methods, leading to unstable evaluation results.
By acquiring training data for the enterprise credit model, the optimization analysis method is determined based on the coverage completeness and data uniformity. Data supplementation or elimination analysis is adopted, and appropriate search and elimination strategies are selected by combining indicators such as the proportion of combined data, data richness and relevance to optimize the training data and improve the stability and accuracy of the model.
It improves the balance and redundancy management of training data for enterprise credit models, enhances the generalization ability and evaluation accuracy of the models, ensures the relevance and efficiency of data supplementation, and improves the stability and reliability of credit assessment.
Smart Images

Figure CN120162594B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of enterprise credit evaluation, and in particular relates to an enterprise credit model optimization method and system based on big data. BACKGROUND
[0002] In the era of big data, more and more data are publicly released by enterprises, which provides rich data for building and optimizing enterprise credit models. However, the training data of the current enterprise credit model generally faces the serious challenges of data incompleteness and data imbalance, resulting in significant volatility of the model in evaluating enterprise credit, and the stability of the prediction results is greatly discounted when facing new data, and thus the real credit status of the enterprise cannot be accurately reflected. Therefore, how to optimize the training data of the enterprise training model to improve the stability and reliability of the enterprise credit model is a technical problem to be solved by those skilled in the art.
[0003] Chinese Patent Publication No. CN114219562A discloses a model training method, an enterprise credit evaluation method and device, equipment, and medium, comprising: obtaining enterprise sample data, preprocessing the enterprise sample data to obtain training data, optimizing the enterprise sample data to improve the accuracy of model training, then calculating the first objective function of the original training model according to the training data to obtain the first objective value; according to the first objective value, fine-tuning the to-be-adjusted parameters of the second objective function of the original training model, taking the fine-tuned second objective function as the model parameters of the original training model to update the original training model to obtain an enterprise credit evaluation model; wherein the enterprise credit evaluation model is used for evaluating enterprise credit. It can be seen that the above technical solution has the following problems: when optimizing the enterprise sample data, the problem of unstable model training results caused by unbalanced training data is not considered, and the adaptive adjustment of the data supplement method according to the actual state of the training data is also not considered, resulting in poor stability and reliability of the enterprise credit model. SUMMARY
[0004] Therefore, the present application provides an enterprise credit model optimization method and system based on big data to overcome the problem that the imbalance of the training data is not considered in the prior art, and the adaptive adjustment of the data supplement method according to the actual state of the training data is not considered, resulting in poor stability and reliability of the enterprise credit model.
[0005] To achieve the above purpose, the present application provides an enterprise credit model optimization method based on big data, comprising:
[0006] The training data corresponding to the target enterprise credit model is acquired, the model data state is determined according to the coverage completeness and the data single degree of the training data, and the optimization analysis mode is determined according to the model data state, and the optimization analysis mode is data supplement analysis or data elimination analysis;
[0007] In the data supplement analysis, the search keywords are determined according to the combined data proportion, the search mode corresponding to each search keyword is determined according to the data richness and the effective evaluation index to obtain search data, and the supplement mode is determined according to the number of effective search data of the associated search data combination and the error motion coefficient;
[0008] The adjustment mode is determined according to the correlation coefficient and the comparison coefficient of the supplement data, and the adjustment mode is adjustment of search depth or backtracking range;
[0009] The search mode is matching search according to the keyword correlation degree and the link evaluation value or individual search according to the link effectiveness, and the supplement mode is selecting supplement data according to the rule representative degree and the data error motion value or determining whether to take the search data as supplement data according to the abnormal index;
[0010] In the data elimination analysis, noise indicators are eliminated according to the index fluctuation coefficient and the index bias rate;
[0011] Optimized data is used to train the target enterprise credit model.
[0012] Further, the optimization analysis mode is determined according to the model data state, including:
[0013] If the model data state is that the coverage completeness is less than the preset coverage completeness or the data single degree is greater than or equal to the preset data single degree, the optimization analysis mode is data supplement analysis;
[0014] If the model data state is that the coverage completeness is greater than or equal to the preset coverage completeness and the data single degree is less than the preset data single degree, the optimization analysis mode is data elimination analysis.
[0015] Further, the confirmation mode of the data single degree includes:
[0016] If the data quantity reference value is greater than or equal to the preset data quantity reference value, the data single degree is determined according to the ratio of the balance coefficient to the data quantity reference value;
[0017] If the data quantity reference value is less than the preset data quantity reference value, the data single degree is determined according to the data correlation degree.
[0018] Further, the search keywords are determined according to the combined data proportion, including:
[0019] According to the industry keywords, the associated data combination is determined, and the search keyword selection analysis is performed for each associated data combination, when the search keyword selection analysis is performed for a single associated data combination,
[0020] If the combined data proportion is greater than or equal to the preset combined data proportion, the search keyword is determined according to the index comparison coefficient and the index distribution proportion;
[0021] If the combined data proportion is less than the preset combined data proportion, the search keyword is determined according to the level jitter coefficient.
[0022] Further, the search mode corresponding to each search keyword is determined according to the data richness and the effective evaluation index, including:
[0023] For a single search keyword,
[0024] If the data richness is less than the preset data richness or the effective evaluation index is less than the preset effective evaluation index, the search mode is matching search according to the keyword correlation degree and the link evaluation value;
[0025] If the data richness is greater than or equal to the preset data richness and the effective evaluation index is greater than or equal to the preset effective evaluation index, the search mode is separate search according to the link effectiveness.
[0026] Further, the supplement mode is determined according to the effective search data number of the associated search data combination and the jitter coefficient, including:
[0027] For a single associated search data combination,
[0028] If the effective search data number is greater than the standard number and the jitter coefficient is greater than or equal to the preset jitter coefficient, the supplement mode is to select the supplement data according to the regular representative degree and the data jitter value;
[0029] If the effective search data number is equal to the standard number or the jitter coefficient is less than the preset jitter coefficient, the supplement mode is to determine whether the search data is used as the supplement data according to the abnormal index.
[0030] Further, the confirmation mode of the associated search data combination includes:
[0031] For a single enterprise name, each search data corresponding to the enterprise name in the backtracking range of the enterprise name is used as the associated search data combination corresponding to the enterprise name;
[0032] The backtracking range corresponding to a single enterprise name is determined according to the dynamic comparison coefficient corresponding to the enterprise name and the domain correlation.
[0033] Further, the adjustment mode is determined according to the correlation coefficient and the comparison coefficient of the supplement data, including:
[0034] If the correlation coefficient is greater than the preset correlation coefficient and the comparison coefficient is less than the preset comparison coefficient, the adjustment mode is to increase the search depth;
[0035] If the correlation coefficient is greater than the preset correlation coefficient and the comparison coefficient is greater than or equal to the preset comparison coefficient, the adjustment mode is to increase the backtracking range.
[0036] Further, according to the index fluctuation coefficient and the index bias rate, the noise index is removed, including:
[0037] The special index keyword with an index fluctuation coefficient less than a preset index fluctuation coefficient and an index bias rate less than a preset index bias rate is recorded as a noise index, and the noise index and each data keyword corresponding to the noise index are removed.
[0038] The application also provides an enterprise credit model optimization system based on big data, comprising:
[0039] A data acquisition module is used to obtain training data corresponding to a target enterprise credit model;
[0040] An optimization analysis module is connected with the data acquisition module, and is used to determine the model data state according to the coverage completeness and the data single degree of the training data, and determine the optimization analysis mode according to the model data state, wherein the optimization analysis mode is data supplement analysis or data removal analysis;
[0041] A data supplement module is connected with the data acquisition module and the optimization analysis module respectively, and is used to determine the search keyword according to the combined data proportion in the data supplement analysis, determine the search mode corresponding to each search keyword according to the data richness and the effective evaluation index to obtain search data, and determine the supplement mode according to the effective search data quantity of the associated search data combination and the jitter coefficient;
[0042] A supplement adjustment module is connected with the data acquisition module and the data supplement module respectively, and is used to determine the adjustment mode according to the correlation coefficient and the comparison coefficient of the supplement data, wherein the adjustment mode is to adjust the search depth or the backtracking range;
[0043] A data removal module is connected with the optimization analysis module, and is used to remove the noise index according to the index fluctuation coefficient and the index bias rate in the data removal analysis;
[0044] A model optimization module is connected with the data supplement module, the supplement adjustment module and the data removal module respectively, and is used to train the target enterprise credit model by using the optimization data.
[0045] Compared with the prior art, the beneficial effects of the present application are that, in the technical scheme of the present application, the model data state is determined according to the coverage completeness of the training data and the data single degree, the coverage completeness and the data single degree effectively reflect the completeness of the training data, and then different optimization analysis modes are adaptively selected according to the model data state, so that the selection of the optimization analysis mode is more in line with the actual application scene, when the balance degree of the training data is poor, more data can be introduced to improve the completeness of the training data, when the redundancy degree of the data is high, the features with low redundancy or correlation can be removed, the generalization ability of the model is improved, and then the model can more comprehensively and accurately capture the credit status of the enterprise.
[0046] Further, in the present application, the search keyword is determined according to the combination data proportion, the combination data proportion effectively reflects the completeness of the training data corresponding to the associated data combination, and then different confirmation modes of the search keyword are adaptively selected according to the combination data proportion, which can improve the accuracy of the search keyword determination, and then more accurately locate the training data that needs to be supplemented, not only improve the efficiency of data supplement, but also ensure the high correlation between the supplemented data and the existing data, and then improve the overall performance of the model.
[0047] Further, in the present application, the data richness and the effective evaluation index effectively reflect the richness of the training data corresponding to the search keyword and the richness of the information in the search page corresponding to the search keyword, and then different search modes are adaptively selected according to the data richness and the effective evaluation index, so that the selection of the search mode is more in line with the actual application scene, the search strategy can be flexibly adjusted, and then the search efficiency and the accuracy of the search are improved.
[0048] Further, in the present application, the dynamic comparison coefficient and the intra-domain correlation degree can reflect the reliability of the search data and the change of the search data over time, and then the backtracking range is determined through the dynamic comparison coefficient and the intra-domain correlation degree, which helps to filter out irrelevant or low-quality search data, and then the effectiveness of the search data in the associated search data combination can be ensured, and then the difference degree of the search data in the associated search data combination is reflected through the number of effective search data and the displacement coefficient, and then different supplement modes are adaptively selected according to the number of effective search data and the displacement coefficient, so that the supplement mode can select search data with a larger representative degree, and then the accuracy of information supplement is improved.
[0049] Further, in the present application, the correlation coefficient and the comparison coefficient effectively reflect the richness of the supplemented search data, and then different adjustment modes are adaptively selected according to the correlation coefficient and the comparison coefficient of the supplemented data, which can mine more related information through the selected adjustment mode, and help to improve the efficiency and effect of information retrieval. BRIEF DESCRIPTION OF DRAWINGS
[0050] Figure 1 a flowchart for determining an optimization analysis mode according to a model data state of the enterprise credit model optimization method based on big data;
[0051] Figure 2 a flowchart for determining an optimization analysis mode according to a model data state of the enterprise credit model optimization method based on big data;
[0052] Figure 3 a flowchart for determining an optimization analysis mode according to a model data state of the enterprise credit model optimization method based on big data;
[0053] Figure 4 a flowchart for determining an optimization analysis mode according to a model data state of the enterprise credit model optimization method based on big data; DETAILED DESCRIPTION
[0054] In order to make the objects and advantages of the present application clearer, the following further describes the present application with reference to examples; it should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.
[0055] The preferred embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the protection scope of the present application.
[0056] It should be noted that, in the description of the present application, the terms "upper", "lower", "left", "right", "inner", "outer" and the like indicate the direction or positional relationship terms based on the direction or positional relationship shown in the drawings, which are only for the convenience of description and do not indicate or imply that the device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.
[0057] In addition, it should be noted that, in the description of the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connecting" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two elements. Those skilled in the art can understand the specific meaning of the above terms in the present application according to the specific circumstances.
[0058] Please refer to Figures 1 to 3 The present application provides an enterprise credit model optimization method based on big data, which comprises:
[0059] The training data corresponding to the target enterprise credit model is acquired, the model data state is determined according to the coverage completeness and the data single degree of the training data, and the optimization analysis mode is determined according to the model data state, and the optimization analysis mode is data supplement analysis or data elimination analysis;
[0060] In the data supplement analysis, the search keywords are determined according to the combined data proportion, the search mode corresponding to each search keyword is determined according to the data richness and the effective evaluation index to obtain search data, and the supplement mode is determined according to the effective search data quantity and the error motion coefficient of the associated search data combination;
[0061] The adjustment mode is determined according to the correlation coefficient and the comparison coefficient of the supplement data, and the adjustment mode is adjusting the search depth or the backtracking range;
[0062] The search mode is matching search according to the keyword correlation degree and the link evaluation value or individual search according to the link effectiveness, and the supplement mode is selecting supplement data according to the regularity degree and the data error motion value or determining whether the search data is used as supplement data according to the abnormal index;
[0063] In the data elimination analysis, the noise indicators are eliminated according to the index fluctuation coefficient and the index bias rate.
[0064] The optimization data is used to train the target enterprise credit model.
[0065] In the training process of the target enterprise credit model, the training data is optimized, a single training data includes data keywords corresponding to a plurality of index keywords, the target enterprise credit model is an enterprise credit model that needs to be optimized, the index keywords include but are not limited to establishment time, annual turnover, net profit rate, employee number, industry category and credit level, the data keywords corresponding to the establishment time include but are not limited to 1 year, 5 years and 10 years, the data keywords corresponding to the annual turnover include but are not limited to 1 million yuan and 20 million yuan, the data keywords corresponding to the net profit rate include but are not limited to 10% and 12%, the data keywords corresponding to the employee number include but are not limited to 50 people and 100 people, the data keywords corresponding to the industry category include but are not limited to manufacturing industry, service industry and information technology, and the data keywords corresponding to the credit level include AAA level, AA level, A level, BBB level and BB level, which is easily understood by those skilled in the art, and specific details are not described.
[0066] The present application corresponds to several historical records, any one of which records at least one data single degree, data quantity reference value, index comparison coefficient, level error coefficient, data richness and effective evaluation index in the historical process of enterprise credit model optimization, and each historical record corresponds to a qualified mark, which records whether the enterprise credit model optimization process meets the user's demand. The qualified mark can be recorded manually. It can be understood that the user can determine whether the enterprise credit model optimization process meets the demand according to the self-set index, which can be but not limited to the misjudgment rate, which will not be described here. The misjudgment rate is the number of times of the target enterprise credit model error evaluation credit level after training with optimized training data.
[0067] The optimization data is the supplementary data obtained by data supplement analysis and the training data, or the training data remaining after the noise indicators and the corresponding data keywords of the noise indicators are removed by data removal analysis; the present application uses the optimization data for subsequent model training to obtain an optimized target enterprise credit model, so that the enterprise credit rating result is more accurate. This is easily understood by those skilled in the art, and will not be described in detail.
[0068] Specifically, the optimization analysis mode is determined according to the model data state, including:
[0069] If the model data state is that the coverage completeness is less than the preset coverage completeness or the data single degree is greater than or equal to the preset data single degree, the optimization analysis mode is data supplement analysis;
[0070] If the model data state is that the coverage completeness is greater than or equal to the preset coverage completeness and the data single degree is less than the preset data single degree, the optimization analysis mode is data removal analysis.
[0071] Among them, the model data state includes the first model data state and the second model data state, the first model data state is that the coverage completeness is less than the preset coverage completeness or the data single degree is greater than or equal to the preset data single degree, and the second model data state is that the coverage completeness is greater than or equal to the preset coverage completeness and the data single degree is less than the preset data single degree;
[0072] The coverage completeness=(the minimum value in the training data quantity corresponding to each associated data combination) / (the maximum value in the training data quantity corresponding to each associated data combination), and the training data quantity corresponding to a single associated data combination is the number of training data contained in the associated data combination;
[0073] The preset coverage completeness and the preset data single degree of value can be determined by the user according to the actual application scene. The smaller the preset coverage completeness value is, the larger the preset data single degree of value is, and the greater the user's demand for data supplement analysis is. A preset coverage completeness and a preset data single degree of value are provided. The preset coverage completeness is 50%, the historical record of data supplement analysis is detected, and the average value of the data single degree corresponding to the historical record that can meet the user's demand is recorded as the preset data single degree.
[0074] Specifically, the confirmation method of the data single degree includes:
[0075] If the data quantity reference value is greater than or equal to the preset data quantity reference value, the data single degree is determined according to the ratio of the balance coefficient to the data quantity reference value;
[0076] If the data quantity reference value is less than the preset data quantity reference value, the data single degree is determined according to the data correlation degree.
[0077] If the data quantity reference value is greater than or equal to the preset data quantity reference value, the data single degree = the balance coefficient / data quantity reference value;
[0078] If the data quantity reference value is less than the preset data quantity reference value, the data single degree = 1 / data correlation degree;
[0079] The data quantity reference value is the total amount of the training data corresponding to the target enterprise credit model. The value of the preset data quantity reference value can be determined by the user according to the actual application scene. The smaller the value of the preset data quantity reference value is, the greater the user's demand for determining the data single degree according to the ratio of the balance coefficient to the data quantity reference value is. A value of the preset data quantity reference value is provided. The historical record of the user's determination of the data single degree according to the ratio of the balance coefficient to the data quantity reference value is detected. The average value of the data quantity reference value corresponding to the historical record that can meet the user's demand is recorded as the preset data quantity reference value.
[0080] The balance coefficient is the average value of the sub-balance coefficients corresponding to the special index keywords. The special index keyword is an index keyword whose data keyword in the corresponding training data contains a number. The sub-balance coefficient corresponding to a single special index keyword is the standard deviation of the values of the data keywords corresponding to the special index keyword. The value corresponding to a single data keyword is the number contained in the data keyword.
[0081] The data correlation degree is the average value of the sub-correlation means corresponding to the special index keywords.
[0082] For a single special index keyword, the special index keyword is denoted as a target special index keyword, other special index keywords except the target special index keyword are denoted as reference special index keywords, and the sub-correlation mean corresponding to the target special index keyword is the average of the correlation coefficients corresponding to the target special index keyword and each reference special index keyword;
[0083] The calculation formula of the correlation coefficient r corresponding to any two special index keywords is:
[0084]
[0085] The two special index keywords are denoted as a first special index keyword and a second special index keyword, n is the number of training data; x i is the value of the data keyword corresponding to the first special index keyword of the i-th training data, y i is the value of the data keyword corresponding to the second special index keyword of the i-th training data, is the average of the values of the data keywords corresponding to the first special index keyword, is the average of the values of the data keywords corresponding to the second special index keyword, i = 1, 2, 3, …, n.
[0086] Specifically, determining the search keyword according to the combination data proportion includes:
[0087] According to the industry keyword, the associated data combination is determined, and search keyword selection analysis is performed on each associated data combination. When search keyword selection analysis is performed on a single associated data combination,
[0088] If the combination data proportion is greater than or equal to the preset combination data proportion, the search keyword is determined according to the index comparison coefficient and the index distribution proportion;
[0089] If the combination data proportion is less than the preset combination data proportion, the search keyword is determined according to the level jitter coefficient.
[0090] According to the industry keyword, the associated data combination is determined, including: the data keywords in each training data corresponding to the industry category are denoted as industry keywords, and the set of each training data corresponding to the same industry keyword is denoted as an associated data combination;
[0091] The confirmation method of the combination data proportion is that, for a single associated data combination, the combination data proportion corresponding to the associated data combination = (the number of training data contained in the associated data combination) / (the total amount of training data corresponding to the target enterprise credit model);
[0092] The preset combined data proportion value can be determined by the user according to the actual application scene. The smaller the preset combined data proportion value is, the greater the demand of the user for determining the search keyword according to the index comparison coefficient and the index distribution proportion is. A preset combined data proportion value is provided, and the preset combined data proportion is 20%;
[0093] The search keyword is determined according to the index comparison coefficient and the index distribution proportion, including: for a single associated data combination, taking the industry keyword corresponding to the associated data combination and each analysis word group as a search keyword;
[0094] The analysis word group is a single analysis word and a single missing data keyword corresponding to the analysis word. Each missing data keyword corresponds to an analysis word group.
[0095] The analysis word is a special index keyword corresponding to the associated data combination, and the index comparison coefficient of the special index keyword is less than a preset index comparison coefficient or the index distribution proportion of the special index keyword is less than a preset index distribution proportion.
[0096] The confirmation method of the missing data keyword is that, for a single analysis word, the analysis word is recorded as a target analysis word, the data keyword in each training data corresponding to the target analysis word is recorded as a target data keyword, and the target data keyword with a frequency coefficient less than a preset frequency coefficient is recorded as a missing data keyword corresponding to the target analysis word.
[0097] For a single associated data combination, the associated data combination is recorded as a target combination, and other associated data combinations outside the target combination are recorded as reference combinations. The confirmation method of the index comparison coefficient is that, for a single special index keyword, the value of each data keyword in the single associated data combination corresponding to the special index keyword is recorded as a first value, the numerical analysis coefficient of the single associated data combination is equal to the maximum first value corresponding to the associated data combination minus the minimum first value corresponding to the associated data combination, and the index comparison coefficient corresponding to the special index keyword is the average value of the index difference degrees corresponding to the target combination and each reference combination. The index difference degree is the absolute value of the difference between the numerical analysis coefficients corresponding to two associated data combinations.
[0098] The confirmation method of the index distribution proportion is that the number of different data keywords in each training data in the target combination corresponding to a single special index keyword is recorded as a, the number of different data keywords in all training data corresponding to the special index keyword is recorded as b, and the index distribution proportion is equal to a / b.
[0099] The confirmation manner of the frequency coefficient is that the data keywords in each training data corresponding to a single special index keyword are recorded as first reference data keywords, for a single first reference data keyword, the first reference data keyword is recorded as a target word, and the frequency coefficient corresponding to the target word is equal to (the number of times that the target word appears in each training data corresponding to the target combination) / (the number of times that the target word appears in each training data corresponding to the target enterprise credit model).
[0100] The values of the preset index comparison coefficient, the preset index distribution proportion and the preset frequency coefficient can be determined by the user according to the actual application scene. The greater the accuracy of the user to the model optimization is, the greater the values of the preset index comparison coefficient, the preset index distribution proportion and the preset frequency coefficient are. A value of a preset index comparison coefficient, a preset index distribution proportion and a preset frequency coefficient is provided. The historical records of the search keywords determined according to the index comparison coefficient and the index distribution proportion are detected. The average value of the index comparison coefficients of the historical records that can meet the user's demand is recorded as the preset index comparison coefficient. The preset index distribution proportion is 70%, and the preset frequency coefficient is 60%.
[0101] The search keywords are determined according to the level shift coefficient, including: for a single associated data combination, the level keywords whose level shift coefficients of the industry keywords and the level keywords corresponding to the associated data combination are less than a preset level shift coefficient are taken as the search keywords.
[0102] The level keyword is a data keyword corresponding to a credit level.
[0103] For a single level keyword in a single associated data combination, the level shift coefficient is equal to the number of times that the level keyword appears in the training data corresponding to the associated data combination / (the total amount of the training data corresponding to the associated data combination / 5).
[0104] The value of the preset level shift coefficient can be determined by the user according to the actual application scene. The greater the accuracy of the user to the model optimization is, the greater the value of the level shift coefficient is. A value of a preset level shift coefficient is provided. The historical records of the search keywords determined according to the level shift coefficient are detected. The average value of the level shift coefficients of the historical records that meet the user's demand is recorded as the preset level shift coefficient.
[0105] Specifically, the search manner corresponding to each search keyword is determined according to the data richness and the effective evaluation index, including:
[0106] For a single search keyword,
[0107] If the data richness is less than the preset data richness or the effective evaluation index is less than the preset effective evaluation index, the search mode is a matching search according to the keyword correlation degree and the link evaluation value;
[0108] If the data richness is greater than or equal to the preset data richness and the effective evaluation index is greater than or equal to the preset effective evaluation index, the search mode is a separate search according to the link effectiveness.
[0109] The data richness is determined by counting the number of enterprise names in the search page corresponding to the single search keyword, which is taken as a target search keyword.
[0110] In the present application, each search keyword corresponds to a search page, which is a network page obtained by searching the search keyword. The search page contains a plurality of links, each of which corresponds to a subpage. The present application extracts search information contained in each search page and subpage by using network crawler technology and text analysis technology, and each search information corresponds to a publication time. The publication time of a single search information is the time when the search information is uploaded to the network. The enterprise name includes but is not limited to XX company, XX factory and XX center, which is easily understood by those skilled in the art and will not be described in detail. The search information contains data keywords corresponding to each index keyword of the training data.
[0111] The effective evaluation index is determined by:
[0112] If the number of links is less than the preset number of links, the effective evaluation index is determined according to the link depth value, wherein the effective evaluation index and the link depth value are positively correlated.
[0113] If the number of links is greater than or equal to the preset number of links, the effective evaluation index is determined according to the area ratio value, wherein the effective evaluation index and the area ratio value are positively correlated.
[0114] The number of links is the total number of links in the search page corresponding to the target search keyword. The value of the preset number of links can be determined by the user according to the actual application scenario. The greater the value of the preset number of links, the greater the user's demand for determining the effective evaluation index according to the link depth value. A value of the preset number of links is provided, which is 10.
[0115] The search page corresponding to the target search keyword is recorded as a target page, a tree structure diagram is established based on the target page, node analysis is performed on the target page, all subpages corresponding to the target page are extracted, a plurality of new child nodes are created for each subpage, each child node is connected with the root node, node analysis is performed on each subpage, and each subpage is connected with the corresponding child node until the page threshold value of the subpage is 0, and the node analysis is stopped. Finally, a tree structure diagram is established with the target page as the root node and each subpage as the child node, which is easily understood by those skilled in the art and will not be described in detail.
[0116] The link depth value is the number of all child nodes in the tree structure diagram established based on the target page, and the area ratio value is the area of the overlapping part of the first area and the second area. The first area is the smallest rectangular area that can contain each enterprise name in the search page corresponding to the target search keyword, and the second area is the smallest rectangular area that can contain each link in the search page corresponding to the target search keyword.
[0117] The values of the preset data richness and the preset effective evaluation index can be determined by the user according to the actual application scenario. The greater the values of the preset data richness and the preset effective evaluation index, the greater the user's demand for matching search based on the keyword correlation degree and the link evaluation value. A value of a preset data richness and a preset effective evaluation index is provided. The historical records of matching search based on the keyword correlation degree and the link evaluation value are detected. The average value of the data richness corresponding to the historical records that meet the user's demand is recorded as the preset data richness. The average value of the effective evaluation index corresponding to the historical records that meet the user's demand is recorded as the preset effective evaluation index.
[0118] Matching search based on the keyword correlation degree and the link evaluation value includes: for a single search keyword, recording the search keyword as a target keyword, recording other search keywords except the target keyword as reference keywords, selecting the reference keywords with the largest keyword correlation degree with the target keyword and the target keyword to search simultaneously to obtain a target search page, and determining the priority coefficient of each link according to the link evaluation value of the target search page. The priority coefficient of a single link has a positive correlation with the link evaluation value corresponding to the link.
[0119] The link evaluation value corresponding to a single link is the number of enterprise names contained in the subpage corresponding to the link.
[0120] Separate search according to the link effectiveness includes: selecting a link with a link effectiveness greater than a preset link effectiveness for search.
[0121] The link validity degree is (a connectivity coefficient / preset connectivity coefficient) + (a link evaluation value / preset link evaluation value). For a single link in a single target search page, the link is recorded as a target link, and other links in the target search page except the target link are recorded as reference links. The connectivity coefficient corresponding to the target link is the average of the shortest distances from the center point corresponding to the target link to the center points corresponding to the reference links. The center point corresponding to the single link is the center of the circumscribed circle of the smallest rectangle that can contain the single link. The preset connectivity coefficient is the average of the connectivity coefficients corresponding to the links in the target search page. The preset link evaluation value is the average of the link evaluation values corresponding to the links in the target search page.
[0122] For any two search keywords, the keyword correlation degree is (the number of training data in which the two search keywords appear simultaneously / the number of training data corresponding to the target enterprise credit model) + a combination record reference value. The combination record reference value is determined by detecting historical records that match the search according to the keyword correlation degree and the link evaluation value and that can meet the user's demand, and recording the historical records as reference historical records. The combination record reference value is the number of reference historical records in which the two search keywords are selected simultaneously for search / the total amount of reference historical records.
[0123] The value of the preset link validity degree can be determined by the user according to the actual application scenario. The greater the user's demand for improving search efficiency, the smaller the value of the preset link validity degree. A value of the preset link validity degree is provided. Historical records that can meet the user's demand and that are searched according to the link validity degree are recorded as reference records. The average value of the link validity degrees corresponding to the links selected for search in the reference records is recorded as the preset link validity degree.
[0124] Specifically, the supplement method is determined according to the number of effective search data in the associated search data combination and the displacement coefficient, and includes:
[0125] For a single associated search data combination,
[0126] If the number of effective search data is greater than the standard number and the displacement coefficient is greater than or equal to the preset displacement coefficient, the supplement method is to select supplement data according to the regularity degree and the data displacement value.
[0127] If the number of effective search data is equal to the standard number or the displacement coefficient is less than the preset displacement coefficient, the supplement method is to determine whether to take the search data as supplement data according to the abnormality index.
[0128] The effective search data quantity is the number of search data in a single associated search data combination, the displacement coefficient is the maximum value in the sub-displacement coefficients corresponding to the special index keywords in the single associated search data combination, the sub-displacement coefficient is the standard deviation of the numerical values of the data keywords corresponding to the special index keywords in the associated search data combination, and the standard quantity is 1;
[0129] The value of the preset displacement coefficient can be determined by the user according to the actual application scenario. The smaller the value of the preset displacement coefficient is, the greater the demand of the user for selecting supplementary data according to the regularity representative degree and the data displacement value is. A value of a preset displacement coefficient is provided. Historical records of selecting supplementary data according to the regularity representative degree and the data displacement value are detected. The average value of the displacement coefficients corresponding to the historical records that can meet the user's demand is recorded as the preset displacement coefficient.
[0130] The supplementary data is selected according to the regularity representative degree and the data displacement value, including: search data with a regularity representative degree greater than a preset regularity representative degree and a data displacement value less than a preset data displacement value is selected as the supplementary data.
[0131] The search data is determined as the supplementary data according to the anomaly index, including:
[0132] If the anomaly index is greater than or equal to a preset anomaly index, the search data is not selected as the supplementary data.
[0133] If the anomaly index is less than the preset anomaly index, the search data is selected as the supplementary data.
[0134] The confirmation method of the regularity representative degree is that, for a single search data, the average value of the sub-representative degrees corresponding to the data keywords contained in the search data is recorded as the regularity representative degree corresponding to the search data. The sub-representative degree corresponding to a single data keyword is the number of training data containing the data keyword.
[0135] The confirmation method of the data displacement value is that, for a single search data, the numerical values of the data keywords corresponding to the special index keywords in the search data are recorded as analysis values. The data displacement value is the maximum value in the mutation rates corresponding to the analysis values. For a single analysis value, the mutation rate corresponding to the analysis value is |analysis value-average value of the analysis value| / analysis value. Other search data located in the same associated search data combination as the search data is recorded as analysis search data. For a single analysis value, the average value of the numerical values of the data keywords corresponding to the special index keywords corresponding to the analysis value in each analysis search data is recorded as the average value of the analysis value.
[0136] The abnormal index is determined by detecting historical records that cannot meet the user's demand by using the search data as supplementary data, and recording as abnormal historical records. For each data keyword corresponding to a single search data, the average value of the abnormal rate corresponding to each data keyword is recorded as the abnormal index. The abnormal rate corresponding to a single data keyword is the number of abnormal historical records that exist for the data keyword / the total number of abnormal historical records.
[0137] The values of the preset regularity representative degree, the preset data shift value and the preset abnormal index can be determined by the user according to the actual application scenario. The greater the user's demand for improving the optimization accuracy of the model, the greater the values of the preset regularity representative degree and the preset abnormal index, and the smaller the value of the preset data shift value. A value of a preset regularity representative degree, a preset data shift value and a preset abnormal index is provided. The historical records that can meet the user's demand are selected from the supplementary data selected from the historical records selected according to the regularity representative degree and the data shift value. The supplementary data selected from the historical records that can meet the user's demand is recorded as reference supplementary data. The average value of the regularity representative degree corresponding to each reference supplementary data is recorded as the preset regularity representative degree. The preset data shift value is 20%, and the preset abnormal index is 40%.
[0138] Specifically, the confirmation method of the associated search data combination includes:
[0139] For a single enterprise name, each search data corresponding to the enterprise name within the backtracking range of the enterprise name is taken as an associated search data combination corresponding to the enterprise name.
[0140] The backtracking range corresponding to a single enterprise name is determined according to the dynamic comparison coefficient corresponding to the enterprise name and the intra-domain correlation degree.
[0141] The confirmation method of the backtracking range is that, for a single enterprise name, the publication time corresponding to each search data corresponding to the enterprise name is detected. The earliest publication time is recorded as a target time point. The backtracking range is a time range between a backtracking time point and a current time point. The backtracking time point is any time point in the interval [target time point, current time point). The backtracking range and the backtracking reference value have a positive correlation. The backtracking reference value = dynamic comparison coefficient - intra-domain correlation degree.
[0142] The dynamic comparison coefficient = rating change degree + level comparison value.
[0143] The confirmation method of the rating change degree is that, for a single enterprise name, each search data corresponding to the enterprise name is recorded as a same-name data. The same-name data with the same level keyword is recorded as a same-level combination. The rating change degree = 1-(the maximum value in the number reference value corresponding to each same-level combination / the total amount of same-name data). The number reference value corresponding to a single same-level combination is the number of search data included in the same-level combination.
[0144] The level comparison value = (the maximum value in the level change rate corresponding to each same-level combination) - (the minimum value in the level change rate corresponding to each same-level combination); the level change rate corresponding to a single same-level combination is the average value of the fluctuation coefficients corresponding to each special index keyword, and the fluctuation coefficient corresponding to a single special index keyword is the standard deviation of the numerical values of the data keywords corresponding to the special index keyword in the same-level combination;
[0145] For a single search data, the numerical values of the data keywords corresponding to each special index keyword in the search data are recorded as analysis numerical values, and the data jitter value is the maximum value in the mutation rate corresponding to each analysis numerical value; for a single analysis numerical value, the mutation rate corresponding to the analysis numerical value = |analysis numerical value - average coefficient corresponding to the analysis numerical value| / analysis numerical value; other search data in the same associated search data combination as the search data are recorded as analysis search data; for a single analysis numerical value, the average coefficient corresponding to the analysis numerical value is the average value of the numerical values of the data keywords corresponding to the special index keyword corresponding to the analysis numerical value in each analysis search data;
[0146] The intra-domain correlation degree = 1 / (the maximum value in the fluctuation coefficient corresponding to each special index keyword - the minimum value in the fluctuation coefficient corresponding to each special index keyword), and the fluctuation coefficient corresponding to a single special index keyword = (the absolute value of the difference between the maximum value and the minimum value in the numerical values of the data keywords corresponding to the special index keyword in each same-name data) / (the maximum value in the numerical values of the data keywords corresponding to the special index keyword in each same-name data).
[0147] Specifically, the adjustment mode is determined according to the correlation coefficient and the comparison coefficient of the supplementary data, including:
[0148] If the correlation coefficient is greater than the preset correlation coefficient and the comparison coefficient is less than the preset comparison coefficient, the adjustment mode is to increase the search depth;
[0149] If the correlation coefficient is greater than the preset correlation coefficient and the comparison coefficient is greater than or equal to the preset comparison coefficient, the adjustment mode is to increase the backtracking range.
[0150] It should be noted that when the correlation coefficient is less than or equal to the preset correlation coefficient, the search depth and the backtracking range are not adjusted, the correlation coefficient is the average value of the sub-correlation degrees corresponding to each supplementary data, and the sub-correlation degree corresponding to a single supplementary data is the average value of the same word quantities corresponding to each reference supplementary data with respect to a single supplementary data, the reference supplementary data is recorded as the target supplementary data, and the other supplementary data except the target supplementary data is recorded as the reference supplementary data, and the sub-correlation degree is the average value of the same word quantities corresponding to each reference supplementary data;
[0151] The comparison coefficient = 1 / (average of the domain correlation of the enterprise name corresponding to each supplementary data) ;
[0152] The user can determine the values of the preset correlation coefficient and the preset comparison coefficient according to the actual application scene. The smaller the value of the preset correlation coefficient is, the larger the value of the preset comparison coefficient is, and the greater the user's demand for increasing adjustment of the search depth is. A value of a preset correlation coefficient and a preset comparison coefficient is provided, historical records of increasing adjustment of the search depth are detected, an average value of the correlation coefficients corresponding to the historical records that can meet the user's demand is recorded as the preset correlation coefficient, and an average value of the comparison coefficients corresponding to the historical records that can meet the user's demand is recorded as the preset comparison coefficient.
[0153] When the search depth is increased, the increase value of the search depth and the correlation coefficient are in a positive correlation;
[0154] When the search depth is increased, the increase value of the search depth and the correlation coefficient are in a positive correlation;
[0155] The search depth is the time of network crawler processing for the search page corresponding to each search keyword. It should be noted that in the present application, the search depth corresponding to a single search keyword and the data richness corresponding to the search keyword are in a negative correlation when the initial crawler processing is performed. When the search depth is increased, the interval between the selected backtracking time point and the target time point gradually decreases.
[0156] Specifically, the noise indicators are removed according to the index fluctuation coefficient and the index bias rate, including:
[0157] The special indicator keyword with an index fluctuation coefficient less than a preset index fluctuation coefficient and an index bias rate less than a preset index bias rate is recorded as a noise indicator, and the noise indicator and each data keyword corresponding to the noise indicator are removed.
[0158] The index fluctuation coefficient corresponding to a single special indicator keyword = (the sub-equilibrium coefficient corresponding to the special indicator keyword) / (the average value of the numerical values of each data keyword corresponding to the special indicator keyword).
[0159] For a single special indicator keyword, the index bias rate corresponding to the special indicator keyword = the sub-correlation average value corresponding to the special indicator keyword - the correlation reference value. The historical records that can meet the user's demand and determine the removal mode according to the data smoothing ratio are recorded as comparison historical records. The average value of the sub-correlation average values of the special indicator keywords removed in the comparison historical records is recorded as the correlation reference value.
[0160] The preset index fluctuation coefficient and the preset index bias rate can be determined by the user according to the actual application scene, and the greater the demand of the user for the model training accuracy, the greater the value of the preset index fluctuation coefficient and the preset index bias rate, and the application provides a value of the preset index fluctuation coefficient and the preset index bias rate, detects the historical records of the noise indicators and the data keywords corresponding to the noise indicators, and records the average value of the index fluctuation coefficient corresponding to each noise indicator in each historical record that can meet the user's demand as the preset index fluctuation coefficient, and records the average value of the index bias rate corresponding to each noise indicator in each historical record that can meet the user's demand as the preset index bias rate.
[0161] Please refer to Figure 4 The application also provides an enterprise credit model optimization system based on big data, which comprises:
[0162] A data acquisition module is used to acquire training data corresponding to a target enterprise credit model.
[0163] An optimization analysis module is connected with the data acquisition module and is used to determine a model data state according to the coverage completeness and the data single degree of the training data, and determine an optimization analysis mode according to the model data state, wherein the optimization analysis mode is data supplement analysis or data elimination analysis.
[0164] A data supplement module is connected with the data acquisition module and the optimization analysis module respectively, and is used to determine a search keyword according to a combined data proportion in the data supplement analysis, determine a search mode corresponding to each search keyword according to a data richness and an effective evaluation index to acquire search data, and determine a supplement mode according to the number of effective search data of the associated search data combination and a jitter coefficient.
[0165] A supplement adjustment module is connected with the data acquisition module and the data supplement module respectively, and is used to determine an adjustment mode according to a correlation coefficient and a comparison coefficient of the supplement data, wherein the adjustment mode is adjustment of the search depth or the backtracking range.
[0166] A data elimination module is connected with the optimization analysis module and is used to eliminate noise indicators according to an index fluctuation coefficient and an index bias rate in the data elimination analysis.
[0167] A model optimization module is connected with the data supplement module, the supplement adjustment module and the data elimination module respectively, and is used to train the target enterprise credit model by using the optimization data.
[0168] Thus far, the technical solutions of the present application have been described in combination with the preferred embodiments shown in the drawings, but it is readily understood by those skilled in the art that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the relevant technical features without departing from the principles of the present application, and the technical solutions after these changes or replacements will all fall within the protection scope of the present application.
[0169] The above merely describes the preferred embodiments of the present application and is not intended to limit the present application; the present application can have various changes and variations for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A big data-based enterprise credit model optimization method, characterized in that, The method comprises the following steps: acquiring training data corresponding to a target enterprise credit model, determining a model data state according to a coverage completeness of the training data and a data single degree, and determining an optimization analysis mode according to the model data state, the optimization analysis mode being data supplement analysis or data elimination analysis; in the data supplement analysis, determining a search keyword according to a combined data proportion, determining a search mode corresponding to each search keyword according to a data richness and an effective evaluation index to acquire search data, and determining a supplement mode according to an effective search data quantity of associated search data combinations and a jitter coefficient; determining an adjustment mode according to a correlation coefficient of the supplement data and a comparison coefficient, the adjustment mode being adjustment of a search depth or a backtracking range; the search mode is matching search according to a keyword correlation degree and a link evaluation value or individual search according to a link effectiveness, and the supplement mode is selection of supplement data according to a rule representation degree and a data jitter value or determination of whether to take the search data as supplement data according to an abnormal index; in the data elimination analysis, eliminating noise indexes according to an index fluctuation coefficient and an index bias rate; training the target enterprise credit model by using the optimized data; the combined data proportion is determined as follows: for a single associated data combination, the combined data proportion corresponding to the associated data combination = (the number of training data contained in the associated data combination) / (the total amount of training data corresponding to the target enterprise credit model); the correlation coefficient is an average value of sub-correlation degrees corresponding to each supplement data, the sub-correlation degree is an average value of the same word quantities corresponding to each reference supplement data, and the same word quantity corresponding to a single reference supplement data is the number of data keywords identical to a target keyword in the reference supplement data; the comparison coefficient = 1 / (an average value of domain-associated degrees of enterprise names corresponding to each supplement data); the jitter coefficient is a maximum value in sub-jitter coefficients corresponding to each special index keyword in a single associated search data combination, and the sub-jitter coefficient is a standard deviation of the numerical values of data keywords corresponding to the special index keyword in the associated search data combination; the data jitter value is determined as follows: for a single search data, the numerical values of data keywords corresponding to each special index keyword in the search data are recorded as analysis numerical values, the data jitter value is a maximum value in mutation rates corresponding to each analysis numerical value, for a single analysis numerical value, the mutation rate corresponding to the analysis numerical value = |analysis numerical value - mean coefficient corresponding to the analysis numerical value| / analysis numerical value, and other search data located in the same associated search data combination as the search data are recorded as analysis search data, for a single analysis numerical value, the mean coefficient corresponding to the analysis numerical value is an average value of the numerical values of data keywords corresponding to the special index keyword corresponding to the analysis numerical value in each analysis search data.
2. The big data based enterprise credit model optimization method of claim 1, wherein, determining the optimization analysis mode according to the model data state, comprising: if the model data state is that the coverage completeness is less than a preset coverage completeness or the data single degree is greater than or equal to a preset data single degree, the optimization analysis mode is data supplement analysis; If the model data state is that the coverage completeness is greater than or equal to the preset coverage completeness and the data single degree is less than the preset data single degree, the optimization analysis mode is data elimination analysis; The coverage completeness is (the minimum value in the training data amount corresponding to each associated data combination) / (the maximum value in the training data amount corresponding to each associated data combination), and the training data amount corresponding to a single associated data combination is the number of training data contained in the associated data combination; The confirmation mode of the data single degree includes: If the data amount reference value is greater than or equal to the preset data amount reference value, the data single degree is determined according to the ratio of the balance coefficient to the data amount reference value; If the data amount reference value is less than the preset data amount reference value, the data single degree is determined according to the data correlation degree.
3. The big data based enterprise credit model optimization method of claim 2, wherein, The search keyword is determined according to the combination data proportion, including: According to the industry keyword, the associated data combination is determined, and search keyword selection analysis is performed on each associated data combination. When the search keyword selection analysis is performed on a single associated data combination, If the combination data proportion is greater than or equal to the preset combination data proportion, the search keyword is determined according to the index comparison coefficient and the index distribution proportion; If the combination data proportion is less than the preset combination data proportion, the search keyword is determined according to the level error motion coefficient.
4. The big data based enterprise credit model optimization method of claim 3, wherein, The search mode corresponding to each search keyword is determined according to the data richness and the effective evaluation index, including: For a single search keyword, If the data richness is less than the preset data richness or the effective evaluation index is less than the preset effective evaluation index, the search mode is matching search according to the keyword correlation degree and the link evaluation value; If the data richness is greater than or equal to the preset data richness and the effective evaluation index is greater than or equal to the preset effective evaluation index, the search mode is individual search according to the link effectiveness; The data richness is the number of enterprise names appearing in the search page corresponding to a single search keyword.
5. The big data based enterprise credit model optimization method of claim 4, wherein, The supplement mode is determined according to the effective search data amount of the associated search data combination and the error motion coefficient, including: For a single associated search data combination, If the effective search data amount is greater than the standard amount and the error motion coefficient is greater than or equal to the preset error motion coefficient, the supplement mode is to select supplement data according to the rule representative degree and the data error motion value; If the effective search data amount is equal to the standard amount or the error motion coefficient is less than the preset error motion coefficient, the supplement mode is to determine whether to take the search data as supplement data according to the abnormal index.
6. The big data based enterprise credit model optimization method of claim 5, wherein, The confirmation mode of the associated search data combination includes: For a single enterprise name, each search data corresponding to the enterprise name within the backtracking range of the enterprise name is taken as the associated search data combination corresponding to the enterprise name; The backtracking range corresponding to a single enterprise name is determined according to the dynamic comparison coefficient corresponding to the enterprise name and the domain correlation degree.
7. The big data based enterprise credit model optimization method of claim 6, wherein, The adjustment mode is determined according to the correlation coefficient of the supplement data and the comparison coefficient, including: If the correlation coefficient is greater than the preset correlation coefficient and the comparison coefficient is less than the preset comparison coefficient, the adjustment mode is to increase the search depth; If the correlation coefficient is greater than the preset correlation coefficient and the comparison coefficient is greater than or equal to the preset comparison coefficient, the adjustment mode is to increase the backtracking range.
8. The big data based enterprise credit model optimization method of claim 2, wherein, According to the index fluctuation coefficient and the index bias rate, the noise index is removed, including: The special index keyword with the index fluctuation coefficient less than the preset index fluctuation coefficient and the index bias rate less than the preset index bias rate is recorded as the noise index, and the noise index and each data keyword corresponding to the noise index are removed.
9. An optimization system for applying the big data-based enterprise credit model optimization method of any one of claims 1 to 8, characterized in that, Including: The data acquisition module is used to acquire the training data corresponding to the target enterprise credit model; The optimization analysis module is connected with the data acquisition module, and is used to determine the model data state according to the coverage completeness and the data single degree of the training data, and determine the optimization analysis mode according to the model data state, the optimization analysis mode is data supplement analysis or data removal analysis; The data supplement module is connected with the data acquisition module and the optimization analysis module respectively, and is used to determine the search keyword according to the combined data proportion in the data supplement analysis, determine the search mode corresponding to each search keyword according to the data richness and the effective evaluation index to acquire the search data, and determine the supplement mode according to the effective search data quantity of the associated search data combination and the error motion coefficient; The supplement adjustment module is connected with the data acquisition module and the data supplement module respectively, and is used to determine the adjustment mode according to the correlation coefficient and the comparison coefficient of the supplement data, the adjustment mode is to adjust the search depth or the backtracking range; The data removal module is connected with the optimization analysis module, and is used to remove the noise index according to the index fluctuation coefficient and the index bias rate in the data removal analysis; The model optimization module is connected with the data supplement module, the supplement adjustment module and the data removal module respectively, and is used to train the target enterprise credit model by using the optimization data.
Citation Information
Patent Citations
Model training method, enterprise credit evaluation method and device, equipment and medium
CN114219562A
Enterprise credit assessment method based on big data mining technology
CN105787073A
Data processing method and device for optimizing credit evaluation model
CN107633265A