Data analysis method and device, electronic equipment and storage medium
By acquiring historical data performance characteristics from the same month and partial actual data from the data evaluation month, baseline indicators and calibration coefficients are calculated, solving the accuracy problem of predicting immature data in big data business and achieving higher prediction accuracy and timeliness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DUXIAOMAN TECH (BEIJING) CO LTD
- Filing Date
- 2026-03-20
- Publication Date
- 2026-06-05
Smart Images

Figure CN122152656A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a data analysis method, apparatus, electronic device, and storage medium. Background Technology
[0002] Big data business refers to a broad range of business models that utilize massive, diverse, timely, and low-value-density big data as their core production factor. Through technologies such as big data collection, storage, cleaning, analysis, mining, and visualization, data is deeply processed and its value extracted to ultimately provide data-driven decision support and refined operational services for enterprise operations, industry services, and public decision-making. Typical big data business scenarios include: finance (credit risk control, intelligent marketing), e-commerce (precision marketing, supply chain optimization), and logistics (intelligent logistics scheduling). Therefore, optimizing data analysis methods to improve the accuracy of analysis and prediction is crucial for enhancing decision-making accuracy across various business scenarios. Summary of the Invention
[0003] This application provides a data analysis method, apparatus, electronic device, and storage medium, which can improve the accuracy of data analysis and prediction. The technical solution is as follows: According to one aspect of this application, a data analysis method is provided, the method comprising: Based on the data observation date, determine the data evaluation month and the historical same period month to be analyzed; Based on the historical performance data of the same period in the past, baseline data indicators are determined; Based on the actual performance data of the month, the actual data indicators are determined. Based on the actual data indicators and the baseline data indicators, the prediction calibration coefficient for the data evaluation month is determined; Based on the predicted calibration coefficient, the baseline data index, and the actual data index, predict the target analysis value for the data evaluation month; Based on the target analysis values, a business management strategy for the target business is generated.
[0004] According to another aspect of this application, a data analysis apparatus is provided, the apparatus comprising: The first determination module is used to determine the data evaluation month and the historical same period month to be analyzed based on the data observation day; The second determining module is used to determine baseline data indicators based on the historical performance data of the same period in history; The third determining module is used to evaluate the actual performance data of the month based on the data and determine the actual data indicators; The fourth determining module is used to determine the prediction calibration coefficient for the data evaluation month based on the actual data indicators and the baseline data indicators. The first prediction module is used to predict the target analysis value of the data evaluation month based on the prediction calibration coefficient, the baseline data index, and the actual data index. The generation module is used to generate business management strategies for the target business based on the target analysis values.
[0005] According to one aspect of this application, an electronic device is provided, comprising: a processor and a memory storing a program, the program including instructions that, when executed by the processor, cause the processor to perform the data analysis method as described above.
[0006] According to another aspect of this application, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing the computer to perform the data analysis method described above.
[0007] According to another aspect of this application, a computer program product is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned data analysis method.
[0008] The beneficial effects of the technical solutions provided in this application include at least the following: The data analysis principle first acquires historical performance data for the same period in previous months and calculates baseline data indicators on a preset statistical dimension based on this data. Then, based on the actual performance data of the data evaluation month, it calculates actual data indicators on the same preset statistical dimension. By comparing the baseline and actual data indicators, a prediction calibration coefficient for the data evaluation month is determined. Finally, based on the prediction calibration coefficient, baseline data indicators, and actual data indicators, the target analysis value for the data evaluation month is predicted. In this entire data analysis principle, the performance characteristics of historical months provide baseline data indicators for prediction, and a portion (not all) of the actual data performance of the data evaluation month provides calibration indicators. This allows for the introduction of actual performance data to further improve the accuracy of the target analysis value prediction, building upon the accuracy of the previous month's predictions. Attached Figure Description
[0009] Further details, features, and advantages of this application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 A flowchart of a data analysis method according to an exemplary embodiment of this application is shown; Figure 2A flowchart of another data analysis method according to an exemplary embodiment of this application is shown; Figure 3 This is a schematic diagram of the structure of a data analysis device provided in an embodiment of this application; Figure 4 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of this application is shown. Detailed Implementation
[0010] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0011] It should be understood that the steps described in the method embodiments of this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.
[0012] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this application are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should be noted that the modifications "a" and "a plurality" mentioned in this application are illustrative and not restrictive, and those skilled in the art should understand that unless explicitly indicated in the context, they should be understood as "one or more". The names of messages or information exchanged between multiple devices in the embodiments of this application are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0013] The present invention will now be described with reference to the accompanying drawings. The technical solutions provided by the embodiments of this application will be explained in detail through specific examples and application scenarios.
[0014] In big data business scenarios, users often need to predict business metrics for a specific month or two in advance to manage operations precisely. For example, in e-commerce, if the current month is September, sales in October need to be predicted to adjust current inventory levels and avoid stockouts or overstocking. In manufacturing, if the current month is September, order volume in October needs to be predicted to adjust production line capacity or raw material procurement to avoid raw material overstocking or shortages. In lending, if the current month is September, the annualized risk of various financial products in October needs to be predicted in advance so that asset allocation can be adjusted by comparing annualized risks across different dimensions. Therefore, improving the accuracy of business metric predictions is crucial for precise business management.
[0015] Please refer to Figure 1 The diagram illustrates a flowchart of a data analysis method according to an exemplary embodiment of this application. The method is described using an example of its application in an electronic device. Figure 1 As shown, the method includes: Step 101: Determine the data evaluation month and historical concurrent month to be analyzed based on the data observation date; Step 102: Determine baseline data indicators based on historical performance data of the same month in the past. Step 103: Based on the actual performance data of the data evaluation month, determine the actual data indicators; Step 104: Based on the actual data indicators and baseline data indicators, determine the prediction calibration coefficient for the data evaluation month; Step 105: Based on the predicted calibration coefficient, baseline data indicators, and actual data indicators, predict the target analysis value for the data evaluation month.
[0016] Step 106: Generate business management strategies for the target business based on the target analysis values.
[0017] The target analysis value in this application embodiment is determined by the specific business scenario. The target analysis value can also be referred to as the core business indicator that is significant for decision-making in business management strategies under that business scenario. For example, if the business scenario is an e-commerce inventory management scenario, the target analysis value can be the target sales volume; if the business scenario is a manufacturing supply chain management scenario, the target analysis value can be the target capacity utilization rate; if the business scenario is a credit scenario, the target analysis value can be the target annualized risk value. This application embodiment does not limit the specific scenario of data analysis.
[0018] Generally, analyzing core business metrics for a given month requires obtaining all actual business data for that month. However, in business forecasting scenarios, because the business data is not yet mature, it is impossible to obtain all the actual business data for that month. If data analysis is conducted only after the actual business data is produced, the timeliness is obviously poor, and it is impossible to adapt to changes in the business scenario in a timely manner. Therefore, how to obtain reliable core business metric data when the business data is not yet mature is the technical problem that this application aims to solve.
[0019] This application embodiment takes into account the characteristics of small short-term fluctuations and similar performance of data in the same period. It provides baseline data indicators for business indicator prediction based on historical data of the same period. Moreover, based on the baseline data indicators, it also introduces some existing actual data for calibration, so that the prediction results are closer to the actual performance of the prediction month, thereby comprehensively improving the accuracy of data analysis.
[0020] In one possible implementation, when performing data analysis, a data observation date (including year, month, and day) is preset. The data observation date is a time when there is no actual business data performance. Based on the data observation date, the data evaluation month and the historical period month to be analyzed are determined. The data evaluation month is a time when there is no complete business data performance, but there may be some business data performance. The historical period month is a time when there is complete business performance data.
[0021] Taking month T as an example, the data observation date is month T or month T-1, and the historical contemporaneous months are month T-4, month T-3, and month T-2. For instance, if the data observation date is November 15, 2025, the historical contemporaneous months are July to September 2025, and the data evaluation months are October and November, with October being the rough estimate month.
[0022] The data analysis process begins by acquiring historical performance data for the same month in the same period of the previous year. Based on this historical data, baseline data indicators are calculated across preset statistical dimensions. Then, based on the actual performance data for the data evaluation month, actual data indicators are calculated across the preset statistical dimensions. By comparing the baseline and actual data indicators, a prediction calibration coefficient for the data evaluation month is determined. Finally, based on the prediction calibration coefficient, baseline data indicators, and actual data indicators, the target analysis value for the data evaluation month is predicted. In this overall data analysis principle, the performance characteristics of historical months in the same period provide baseline data indicators for prediction, while a portion (not all) of the actual data performance for the data evaluation month provides calibration indicators. This approach, building upon the accuracy of predictions from the same period, incorporates the assistance of actual performance data to further improve the accuracy of the target analysis value prediction.
[0023] Taking e-commerce inventory management as an example, with the target analysis value being the target sales value, one possible implementation is to obtain sales data from the same month in history to determine the baseline sales. Based on the partial sales data of the data evaluation month, the actual sales are determined. Based on the baseline sales and the actual sales, the sales calibration coefficient for the data evaluation month is determined. Then, based on the sales calibration coefficient, the baseline sales, and the actual sales, the target sales value for the data evaluation month is predicted, so as to generate an inventory management strategy for the e-commerce business based on the target sales value.
[0024] Taking the manufacturing supply chain management scenario as an example, with the target analysis value being the target capacity, one possible implementation involves acquiring historical capacity and order volume data for the same month to determine the baseline capacity utilization rate. Based on partial capacity and order volume data for the data evaluation month, the actual capacity utilization rate is determined. Based on the baseline and actual capacity utilization rates, a capacity utilization rate calibration coefficient for the data evaluation month is determined. Then, based on the capacity utilization rate calibration coefficient, the baseline capacity utilization rate, and the actual capacity utilization rate, the target capacity utilization rate for the data evaluation month is predicted, so as to generate a production plan for the manufacturing supply chain based on the target capacity utilization rate.
[0025] Taking a credit scenario as an example, and the target analysis value as the target annualized risk value; one possible implementation is to obtain historical loan disbursement data and historical delinquency data for the same month in the past to determine baseline risk indicators (baseline delinquency rate, baseline credit score, baseline risk value, etc.); based on the actual credit data (actual loan disbursement data, actual delinquency data, etc.) of the data assessment month, determine actual risk indicators (actual delinquency rate, actual credit score, etc.); then, based on the baseline risk indicators and actual risk indicators, determine the risk calibration coefficient for the data assessment month; and finally, based on the baseline risk indicators, actual risk indicators, and risk calibration coefficient, predict the annualized risk value for the data assessment month; so as to adjust the asset allocation of the financial product based on the annualized risk value.
[0026] After predicting the target analysis value for the data evaluation month—that is, after predicting the core business indicators for the data evaluation month under the business scenario—the target analysis value can guide the generation of business management strategies for the target business. For example, if the business scenario is e-commerce inventory management, the target analysis value can be the target sales volume, which can guide the generation of inventory management strategies for the e-commerce business; if the business scenario is manufacturing supply chain management, the target analysis value can be the target capacity utilization rate, which can be used to generate production plans for the manufacturing supply chain; if the business scenario is credit, the target analysis value can be the target annualized risk value, which can be used to adjust the asset allocation of the financial product.
[0027] For example, if the target sales value is much greater than the current inventory value, an inventory management strategy to increase the inventory level can be generated; conversely, if the target sales value is less than or equal to the current inventory value, an inventory management strategy to stop increasing the inventory level can be generated.
[0028] In summary, this application provides a data analysis method: first, historical performance data of the same month in the same period are obtained, and baseline data indicators on a preset statistical dimension are calculated based on this data; then, actual data indicators on a preset statistical dimension are calculated based on the actual performance data of the data evaluation month; by comparing the baseline data indicators and the actual data indicators, a prediction calibration coefficient for the data evaluation month is determined; and finally, based on the prediction calibration coefficient, the baseline data indicators, and the actual data indicators, the target analysis value for the data evaluation month is predicted. In this entire data analysis principle, the data performance characteristics of the same month in the same period provide baseline data indicators for prediction, and a portion (not all) of the actual data performance of the data evaluation month provides calibration indicators for prediction. This allows for the introduction of actual performance data to further improve the prediction accuracy of the target analysis value, building upon the accuracy of the predictions made during the same period.
[0029] The following examples primarily use credit scenarios within various data analysis contexts to describe the principles of data analysis in detail. Please refer to... Figure 2 This document illustrates a flowchart of another data analysis method according to an exemplary embodiment of this application. The method is described using an example of its application to an electronic device. Figure 2 As shown, the method includes: Step 201: Determine the data evaluation month and historical concurrent month to be analyzed based on the data observation date.
[0030] The data evaluation month includes the first evaluation month and the second evaluation month. The first evaluation month is T-1 month, and the second evaluation month is T month. The month of the second evaluation month is the same as the month of the data observation date.
[0031] For example, the data observation date is November 15, 2025, and the historical corresponding months are July, August, and September 2025. The data evaluation month includes the first evaluation month and the second evaluation month. The first evaluation month is October, and the second evaluation month is November. The first evaluation month (October) is the rough estimate month.
[0032] It should be noted that if month T is January of the current year, then month T-1 should be December of the previous year; for example, if month T is January 2025, then month T-1 should be December 2024.
[0033] Step 202: Based on historical performance data for the same period in history, determine the baseline delinquency rate, first baseline credit score, baseline initial credit score, and first baseline risk value for the same period in history.
[0034] The baseline data indicators mainly include the baseline delinquency rate, baseline credit score, baseline initial credit score, and baseline risk value. One possible implementation involves extracting baseline data indicators such as the baseline delinquency rate, first baseline credit score, baseline initial credit score, and first baseline risk value for each historical month (three months) based on historical performance data, such as delinquency amount, loan amount, initial loan amount, initial loan amount due, annualized risk, and duration for each historical month.
[0035] In an exemplary example, the formula for calculating the baseline delinquency rate can be as shown in formula (1): (1) Among them, the overdue amount of DPD3 (Days Past Due 3) in each month refers to the amount of loans disbursed in the same month in history that are overdue for more than 3 days; the amount of loans with performance in each month refers to the amount of loans disbursed in the same month in history that have repayment performance.
[0036] Taking November 15, 2025 as the data observation date, and July, August, and September 2025 as historical examples, Σ (overdue amount of DPD3 in each month) is the sum of overdue amounts from July 2025 (loans disbursed to August 15, 2025) to overdue amounts overdue by more than 3 days, from August 2025 (loans disbursed to September 15, 2025) to overdue amounts overdue by more than 3 days, and from September 2025 (loans disbursed to October 15, 2025) to overdue amounts overdue by more than 3 days; Σ (loan amount with repayment performance in each month) is the sum of loan amounts from July 2025 (loans disbursed to August 15, 2025) to loan amounts with repayment performance, from August 2025 (loans disbursed to September 15, 2025) to loan amounts with repayment performance, and from September 2025 (loans disbursed to October 15, 2025) to loan amounts with repayment performance.
[0037] In an exemplary example, the formula for calculating the first baseline credit score can be as shown in formula (2): (2) Among them, the baseline FICO (Fair Isaac Corporation) score can also be called the first baseline credit score. The loan amount for each month is the loan amount for the same period in history. The FICO score represents the credit score of each loan recipient in the same period in history.
[0038] Taking the data observation date as November 15, 2025, and the historical corresponding months as July, August, and September 2025 as examples, the loan amounts for each month are the loan amounts for July, August, and September 2025; FICO is divided into the credit scores of each loan recipient in July, August, and September 2025.
[0039] In an exemplary example, the formula for calculating the baseline initial credit score can be as shown in formula (3): (3) The baseline initial FICO score, also known as the baseline initial credit score, is calculated as follows: the initial loan amount for each month is the amount of the first loan disbursement to the borrower in the same historical month; the initial loan amount due for each month is the amount of the first loan disbursement to the borrower in the same historical month that has reached its repayment deadline. The FICO score is the credit score of the borrower corresponding to the initial loan disbursement amount in the same historical month.
[0040] In an exemplary example, the formula for calculating the first baseline risk value can be as shown in formula (4): (4) Among them, the annualized risk for each month refers to the annualized risk for each historical month in the same period, and the loan amount is the loan amount for each historical month in the same period. By substituting the historical performance data of the historical months (three months) such as overdue amount, loan amount, first loan amount, first loan amount due, annualized risk and duration for each historical month in the same period into formulas (1) to (4), the baseline data indicators such as the baseline overdue rate, first baseline credit score, baseline first loan amount and first baseline risk value for the historical months in the same period can be calculated.
[0041] Step 203: Based on the actual performance data of the data assessment month, determine the actual delinquency rate, actual credit score, actual first-payment credit score, and first actual due rate for the same period of the data assessment month.
[0042] The actual data indicators mainly include the actual delinquency rate, actual credit score, actual first-phase credit score, and actual maturity percentage for the same period. One possible implementation involves extracting actual data indicators such as the actual delinquency rate, actual credit score, actual first-phase credit score, and first-phase maturity percentage for the same period based on the actual performance data of the data evaluation month, such as delinquency amount, total loan amount, loan amount with good performance, first-phase loan amount, and first-phase maturity loan amount for the data evaluation month.
[0043] In an exemplary example, the formula for calculating the actual delinquency rate during the same period can be shown in formula (5): (5) As can be seen from formula (5), taking the data assessment month as the first assessment month (T-1 month) as an example, the actual delinquency rate (T-1 delinquency rate) of the data assessment month is determined by the amount of overdue loans exceeding 3 days in the same period of T-1 month compared with the amount of loans disbursed in the same period of T-1 month.
[0044] For example, taking November 15, 2025 as the data observation date and October as the data evaluation month, the overdue amount of DPD3 in the same period of T-1 month is the amount of loan disbursement from October to November 15 that is overdue for more than 3 days, and the overdue amount of DPD3 in the same period of T-1 month is the amount of loan disbursement from October to November 15 that has repayment performance.
[0045] In an exemplary example, the actual credit score can be calculated using the formula shown in formula (6): (6) As can be seen from formula (6), taking the data assessment month as the first assessment month (T-1 month) as an example, the T-1 FICO score represents the actual credit score of the data assessment month. The actual credit score (T-1 FICO score) is determined by the sum of the product of the loan amount in T-1 month and the credit score of the loan recipient, divided by the loan amount in T-1 month. In an exemplary example, the formula for calculating the actual initial credit score can be shown in formula (7): (7) As can be seen from formula (7), taking the data assessment month as the first assessment month (T-1 month) as an example, the first FICO score of T-1 represents the actual first credit score of the data assessment month. The actual first credit score (T-1 FICO score) is determined by the sum of the product of the first loan amount in T-1 month and the credit score of the loan recipient, divided by the first loan amount due in T-1 month.
[0046] In an exemplary example, the formula for calculating the first actual maturity percentage can be as shown in formula (8): (8) As can be seen from formula (8), taking the data assessment month as the first assessment month (T-1 month) as an example, the maturity ratio is the first actual maturity ratio of T-1 month. The first actual maturity ratio is determined by the loan amount with repayment performance in T-1 month divided by the total loan amount in T-1 month.
[0047] By substituting the actual performance data of the first assessment month (T-1 month) in the data assessment month, such as overdue amount, total loan amount, loan amount with performance, first loan amount, first loan due amount, etc. into formulas (5) to (8), the actual overdue rate, actual credit score, actual first loan credit score and first actual due rate of the first assessment month can be calculated.
[0048] Step 204: Determine the delinquency rate increase value based on the actual delinquency rate and the baseline delinquency rate for the same period.
[0049] In one possible implementation, when performing forecast calibration for the first assessment month (T-1) in the data assessment month, a dual boosting calibration mechanism is adopted, specifically including a delinquency rate boosting value and a model segment initial boosting value.
[0050] For example, the formula for calculating the delinquency rate increase can be shown in formula (9): (9) Among them, the delinquency rate Lift is the delinquency rate increase value, the T-1 month delinquency rate is the actual delinquency rate of the same period in the first assessment month, and the baseline delinquency rate is the baseline delinquency rate of the same period in the historical same month; the delinquency rate increase value is determined by the ratio of the actual delinquency rate of the same period in the first assessment month to the baseline delinquency rate of the same period in the historical same month.
[0051] Step 205: Determine the model score improvement value based on the first baseline credit score and the actual credit score.
[0052] In one possible implementation, the model score enhancement value and the model score initial enhancement value are derived from the default probability of the first maturity asset. Therefore, it is necessary to calculate the default probability first, and then calculate the model score enhancement value and the model score initial enhancement value.
[0053] For example, the formula for calculating the probability of default (FICO) can be shown in formula (10): (10) Among them, FICO is divided into the credit score of the loan recipient. Based on formula (10), the default probability of T-1 month can be calculated by substituting the T-1 month FICO score into formula (10), that is, substituting the T-1 month FICO score into formula (10) to replace the FICO score to calculate the default probability (T-1 month FICO); based on formula (10), the baseline FICO score can be calculated by substituting the baseline FICO score into formula (10) to replace the FICO score to calculate the default probability (baseline FICO); EXP(X) is the natural exponential function, that is, an exponential function with the real number e (e≈2.71828) as the base.
[0054] For example, the formula for calculating the model score lift (model score Lift) can be shown in formula (11): (11) As can be seen from formula (11), after calculating the default probability (T-1 month FICO) and the default probability (baseline FICO), the ratio of the default probability (T-1 month FICO) and the default probability (baseline FICO) can be determined as the model score lift (model score Lift).
[0055] Step 206: Determine the initial improvement value of the model score based on the baseline initial credit score and the actual initial credit score.
[0056] Similarly, based on formula (10), the default probability of the first period of T-1 can be calculated by substituting the FICO score of the first period of T-1 into formula (10), which means substituting the FICO score of the first period of T-1 into formula (10) to replace the FICO score and calculate the default probability (FICO score of the first period of T-1); based on formula (10), the default probability of the first period of baseline can be calculated by substituting the FICO score of the first period of baseline into formula (10), which means substituting the FICO score of the first period of baseline into formula (10) to replace the FICO score and calculate the default probability (FICO score of the first period of baseline).
[0057] For example, the formula for calculating the initial improvement value of the model can be shown in formula (12): (12) As can be seen from formula (12), after calculating the default probability (T-1 month first period FICO) and the default probability (baseline first period FICO), the ratio of the default probability (T-1 month first period FICO) and the default probability (baseline first period FICO) can be determined as the model first period lift value (model first period Lift).
[0058] Step 207: Based on the delinquency rate increase and the model's initial increase, determine the prediction calibration coefficient for the data assessment month.
[0059] In an exemplary example, the formula for calculating the predicted calibration coefficient can be as shown in formula (13): (13) As can be seen from formula (13), after calculating the overdue rate increase value and the model first-period increase value, the ratio of the overdue rate increase value (overdue rate Lift) to the model first-period increase value (overdue rate Lift) can be determined as the prediction calibration coefficient (calibration Lift).
[0060] Step 208: Determine the pre-calibration risk value based on the first baseline risk value and the model score improvement value.
[0061] When predicting the annualized risk value (target analysis value or first analysis value) for the first assessment month, it is achieved by fusing the pre-calibration risk value and the post-calibration risk value. The pre-calibration risk value is the annualized risk value without the inclusion of the predicted calibration coefficient, while the post-calibration risk value is the annualized risk value with the inclusion of the predicted calibration coefficient.
[0062] In an exemplary example, the formula for calculating the risk value before calibration can be as shown in formula (14): (14) As can be seen from formula (14), the pre-calibration risk is determined by the product of the first baseline risk value (baseline risk) and the model score lift value (model score Lift) of the same month in history.
[0063] Step 209: Determine the calibrated risk value based on the first baseline risk value, the initial improvement value of the model, and the predicted calibration coefficient.
[0064] For example, the formula for calculating the risk value after calibration can be shown in formula (15): (15) As can be seen from formula (15), the calibrated risk value is determined by the product of the first baseline risk value (baseline risk), the model initial lift value (model initial lift), and the predicted calibration coefficient (calibration lift).
[0065] Step 210: Based on the first actual maturity percentage, the calibrated risk value, and the pre-calibration risk value, determine the target analysis value for the data assessment month.
[0066] The target analytical value for the data assessment month is obtained by weighted fusion of the post-calibration risk and the pre-calibration risk. Optionally, the first analytical value for the first assessment month is determined based on the first actual due date percentage, the post-calibration risk value, and the pre-calibration risk value.
[0067] For example, the formula for calculating the first analytical value (the annualized risk value for T-1 months) can be shown in formula (16): (16) As can be seen from formula (16), the first actual maturity ratio (maturity ratio) of T-1 month is used as the weight of the calibrated risk, and 2-first actual maturity ratio / 2 is used as the weight of the pre-calibration risk. The two are then added together to determine the target analysis value (first analysis value) of the data assessment month (first assessment month), which is also the annualized risk value of T-1 month.
[0068] After calculating the first risk value for the first assessment month, the second risk value for the second assessment month can be calculated based on the first risk value. Then, by combining the first and second risk values, a business management strategy for the target business can be generated.
[0069] Optionally, the data analysis method may further include steps 301 to 304.
[0070] Step 301: Based on the historical analysis values, historical performance data, actual performance data of the first assessment month, the first analysis value, and the first actual maturity percentage of the same month, determine the second baseline risk value for the second assessment month; Step 302: Based on the first baseline credit score of the same period in history, historical performance data, and the first actual maturity ratio of the first assessment month, determine the second baseline credit score of the second assessment month; Step 303: Determine the model score adjustment value based on the second baseline credit score and the actual model score in the second assessment month; Step 304: Based on the second baseline risk value and the model sub-adjustment value, determine the second analysis value for the second assessment month.
[0071] For example, the calculation process for the second analysis value of the second assessment month can be shown in formulas (17) to (20): Baseline for month T = (Risk in month T-1 × Average balance in month T-1 × Percentage of maturing months in month T-1) + (T-2 month risk × T-2 month average balance) + (T-3 month risk × T-3 month average balance) + (T-4 month risk × T-4 month average balance × (1-T-1 maturity percentage)) (17) Baseline FICO for Month T = (FICO for Month T-1 × Average Balance for Month T-1 × Percentage of FICOs Matured for Month T-1) + (T-2 month FICO × T-2 month average balance) + (T-3 month FICO × T-3 month average balance) + (T-4 month FICO × T-4 month average balance × (1-T-1 maturity percentage)) (18) Model score adjustment = Default probability (T-month FICO score) / Default probability (T-month baseline FICO score) (19) Final risk in month T = Baseline risk in month T × Model score Lift (20) From formulas (17) to (20), it can be seen that when calculating the second analysis value for month T, the second baseline risk value (baseline for month T) for the second assessment month is first determined based on the historical analysis values (risk for month T-2, risk for month T-3, and risk for month T-4), historical performance data (average balance for month T-2, average balance for month T-3, and average balance for month T-4), actual performance data for the first assessment month (average balance for month T-1), the first analysis value (risk for month T-1), and the first actual maturity percentage (maturity percentage for month T-1). Then, based on the first baseline credit score (FICO for month T-2, risk for month T-4) for the same historical month, the second baseline risk value (baseline for month T) for the second assessment month is determined. Based on the FICO data for March and April (T-3 and T-4 months), historical performance data (average balance for T-2, T-3, and T-4 months), and the first actual maturity percentage for the first assessment month (T-1 maturity percentage), the second baseline credit score (T-month baseline FICO) for the second assessment month is determined. Then, based on the second baseline credit score (T-month baseline FICO) and the actual model score for the second assessment month (T-month FICO score), the model score adjustment value is determined. Finally, based on the product of the second baseline risk value (T-month baseline risk) and the model score adjustment value, the second analytical value (T-month final risk) for the second assessment month is determined.
[0072] Step 211: Generate business management strategies for the target business based on the target analysis values.
[0073] The implementation method of this step can be referred to the above embodiment, and will not be repeated here.
[0074] In other possible application scenarios, the target analysis value (annualized risk value) for each customer group in the data evaluation month can be calculated according to different dimensions. Then, by comparing different dimensions, business management strategies for different dimensions can be generated. For example, if the annualized risk of new customer groups is higher and the annualized risk of old customer groups is lower, more asset allocation can be increased to old customer groups.
[0075] In one exemplary example, a multi-dimensional hierarchical computing architecture can be shown in Table 1: Table 1
[0076] As shown in Table 1, users can be divided into new customers and returning customers based on whether they are new or returning customers of Mob6 (i.e., whether it is their first loan in the past 6 months); they can also be divided into multiple customer tiers based on their credit score (historical performance and model score); or, users can be divided into multiple customer tiers based on business lines such as product type and customer acquisition channels (e.g., in-app customers, out-of-app customers, innovative consumption customers, etc.).
[0077] After segmenting the customer groups into multiple dimensions, the annualized risk value (target analysis value) for each dimension can be calculated using the data analysis methods described above. Then, by comparing the annualized risk values of different dimensions, business management strategies for different dimensions can be generated.
[0078] In an exemplary example, the formula for calculating stratified risk can be as shown in formulas (21) to (23): (twenty one) (twenty two) (twenty three) From formulas (21) to (23), it can be seen that the risk of new customer A / 1 can be determined by the risk of new customer A layer, the loan amount of new customer A, the risk of old customer 1 and the loan amount of old customer 1; the risk of new customer can be determined by the risk of new customer ABCD layer and the loan amount of new customer ABCD; the overall risk can be determined by the risk of each layer and the loan amount of each layer.
[0079] Optionally, fallback strategies are also set up to address some anomalies that may occur during the data analysis process. For example, if the proportion of due dates in T-1 month is less than 10%, the risk before calibration can be used as the first analysis value for the first assessment month; if the baseline delinquency rate is 0, the historical average can be used instead of the baseline delinquency rate; if the annualized risk for a certain year is missing, the most recent available version can be taken.
[0080] In this embodiment, a two-layer Lift calibration system is used: the overdue Lift captures the actual risk fluctuations; the model-based Lift eliminates differences in customer distribution. The calibration coefficient = overdue Lift / model-based Lift, which can eliminate more than 70% of the structural bias of the customer group and improve the accuracy of risk assessment. Moreover, the calibration weight is intelligently adjusted according to the actual performance coverage. The innovative annualized risk calculation formula is: (proportion of due dates × risk after calibration + (2 - proportion of due dates) × risk before calibration) / 2. This solves the estimation bias problem under partial observation and achieves a smooth transition. In addition, a multi-dimensional hierarchical calculation architecture is adopted to support multi-dimensional fine-grained risk assessment of new and old customers, ABCD hierarchies, business lines, etc., so as to achieve independent calculation and summarization of 15 hierarchical dimensions and meet the needs of refined risk management.
[0081] Please refer to Figure 3 This is a schematic diagram of the structure of a data analysis device provided in an embodiment of this application. For example, as shown... Figure 3 As shown, the device 300 includes: The first determination module 301 is used to determine the data evaluation month and the historical same period month to be analyzed based on the data observation day; The second determining module 302 is used to determine baseline data indicators based on the historical performance data of the same period in history; The third determining module 303 is used to determine the actual data indicators based on the data to evaluate the actual performance data of the month. The fourth determining module 304 is used to determine the prediction calibration coefficient of the data evaluation month based on the actual data indicators and the baseline data indicators; The first prediction module 305 is used to predict the target analysis value of the data evaluation month based on the prediction calibration coefficient, the baseline data index and the actual data index. The generation module 306 is used to generate a business management strategy for the target business based on the target analysis value.
[0082] Optionally, the second determining module 302 is further configured to: Based on the historical performance data of the same historical month, the baseline delinquency rate, first baseline credit score, baseline initial credit score, and first baseline risk value of the same historical month are determined. The third determining module 303 is further configured to: Based on the actual performance data of the data assessment month, the actual delinquency rate, actual credit score, actual first-payment credit score, and first actual due date percentage for the same period of the data assessment month are determined.
[0083] Optionally, the fourth determining module 304 is further configured to: Based on the actual delinquency rate and the baseline delinquency rate for the same period, determine the delinquency rate increase value; Based on the first baseline credit score and the actual credit score, determine the model score improvement value; Based on the baseline initial credit score and the actual initial credit score, determine the initial improvement value of the model score; Based on the delinquency rate increase and the model's initial increase, the prediction calibration coefficient for the data evaluation month is determined.
[0084] Optionally, the first prediction module 305 is further configured to: Based on the first baseline risk value and the model score improvement value, the risk value before calibration is determined; Based on the first baseline risk value, the initial improvement value of the model, and the prediction calibration coefficient, the calibrated risk value is determined; Based on the first actual due rate, the post-calibration risk value, and the pre-calibration risk value, the target analysis value for the data evaluation month is determined.
[0085] Optionally, the data evaluation month includes a first evaluation month and a second evaluation month, where the first evaluation month is month T-1 and the second evaluation month is month T, and the month of the second evaluation month is the same as the month of the data observation date; The third determining module 303 is further configured to: Based on the actual performance data of the first assessment month, the actual data indicators are determined; The first prediction module 305 is further configured to: Based on the predicted calibration coefficient, the baseline data index, and the actual data index, predict the first analysis value for the first evaluation month.
[0086] Optionally, the device further includes: The fifth determining module is used to determine the second baseline risk value of the second assessment month based on the historical analysis value of the same historical month, the historical performance data, the actual performance data of the first assessment month, the first analysis value, and the first actual maturity percentage. The sixth determining module is used to determine the second baseline credit score for the second assessment month based on the first baseline credit score of the same historical month, the historical performance data, and the first actual maturity percentage of the first assessment month. The seventh determination module is used to determine the model score adjustment value based on the second baseline credit score and the actual model score of the second assessment month; The second prediction module is used to predict the second analysis value for the second assessment month based on the second baseline risk value and the model sub-adjustment value.
[0087] An exemplary embodiment of this application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the electronic device to perform a data analysis method according to an embodiment of this application.
[0088] An exemplary embodiment of this application also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a data analysis method according to an embodiment of this application.
[0089] An exemplary embodiment of this application also provides a computer program product, including a computer program, wherein, when executed by a computer's processor, the computer program is used to cause the computer to perform a data analysis method according to an embodiment of this application.
[0090] refer to Figure 4The present invention describes a structural block diagram of an electronic device 400 that can serve as a server or client of this application, which is an example of a hardware device that can be applied to various aspects of this application. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the application described and / or claimed herein.
[0091] like Figure 4 As shown, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. The RAM 403 may also store various programs and data required for the operation of the electronic device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0092] Multiple components in electronic device 400 are connected to I / O interface 405, including: input unit 406, output unit 407, storage unit 408, and communication unit 409. Input unit 406 can be any type of device capable of inputting information to electronic device 400. Input unit 406 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 407 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 408 may include, but is not limited to, disks and optical discs. Communication unit 409 allows electronic device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0093] Optionally, the electronic device 400 also includes a single-channel EEG signal acquisition module (not shown in the figure). This module is used to acquire EEG signals and transmit them to the signal processor of the electronic device 400 for EEG signal processing.
[0094] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above. For example, in some embodiments, the methods shown in the above embodiments can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 400 via ROM 402 and / or communication unit 409. In some embodiments, the computing unit 401 can be configured to perform the methods shown in the above embodiments by any other suitable means (e.g., by means of firmware).
[0095] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0096] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0097] As used in this application, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0098] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0099] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0100] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
Claims
1. A data analysis method, characterized in that, The method includes: Based on the data observation date, determine the data evaluation month and the historical same period month to be analyzed; Based on the historical performance data of the same period in the past, baseline data indicators are determined; Based on the actual performance data of the month, the actual data indicators are determined. Based on the actual data indicators and the baseline data indicators, the prediction calibration coefficient for the data evaluation month is determined; Based on the predicted calibration coefficient, the baseline data index, and the actual data index, predict the target analysis value for the data evaluation month; Based on the target analysis values, a business management strategy for the target business is generated.
2. The method according to claim 1, characterized in that, The determination of baseline data indicators based on historical performance data of the same historical month includes: Based on the historical performance data of the same historical month, the baseline delinquency rate, first baseline credit score, baseline initial credit score, and first baseline risk value of the same historical month are determined. The determination of actual data indicators based on the actual performance data of the month evaluated by the data includes: Based on the actual performance data of the data assessment month, the actual delinquency rate, actual credit score, actual first-payment credit score, and first actual due date percentage for the same period of the data assessment month are determined.
3. The method according to claim 2, characterized in that, The determination of the prediction calibration coefficient for the data evaluation month based on the actual data indicators and the baseline data indicators includes: Based on the actual delinquency rate and the baseline delinquency rate for the same period, determine the delinquency rate increase value; Based on the first baseline credit score and the actual credit score, determine the model score improvement value; Based on the baseline initial credit score and the actual initial credit score, determine the initial improvement value of the model score; Based on the delinquency rate increase and the model's initial increase, the prediction calibration coefficient for the data evaluation month is determined.
4. The method according to claim 3, characterized in that, The prediction of the target analysis value for the data evaluation month based on the predicted calibration coefficient, the baseline data index, and the actual data index includes: Based on the first baseline risk value and the model score improvement value, the risk value before calibration is determined; Based on the first baseline risk value, the initial improvement value of the model, and the prediction calibration coefficient, the calibrated risk value is determined; Based on the first actual due rate, the post-calibration risk value, and the pre-calibration risk value, the target analysis value for the data evaluation month is determined.
5. The method according to any one of claims 1 to 4, characterized in that, The data evaluation month includes a first evaluation month and a second evaluation month. The first evaluation month is month T-1, and the second evaluation month is month T. The month of the second evaluation month is the same as the month of the data observation date. The determination of actual data indicators based on the actual performance data of the month evaluated by the data includes: Based on the actual performance data of the first assessment month, the actual data indicators are determined; The prediction of the target analysis value for the data evaluation month based on the predicted calibration coefficient, the baseline data index, and the actual data index includes: Based on the predicted calibration coefficient, the baseline data index, and the actual data index, predict the first analysis value for the first evaluation month.
6. The method according to claim 5, characterized in that, The method further includes: Based on the historical analysis values of the same historical month, the historical performance data, the actual performance data of the first assessment month, the first analysis value, and the first actual maturity percentage, the second baseline risk value of the second assessment month is determined. Based on the first baseline credit score of the historical same month, the historical performance data, and the first actual maturity ratio of the first assessment month, the second baseline credit score of the second assessment month is determined; The model score adjustment value is determined based on the second baseline credit score and the actual model score for the second assessment month; Based on the second baseline risk value and the model sub-adjustment value, the second analysis value for the second assessment month is determined.
7. A data analysis device, characterized in that, The device includes: The first determination module is used to determine the data evaluation month and the historical same period month to be analyzed based on the data observation day; The second determining module is used to determine baseline data indicators based on the historical performance data of the same period in history; The third determining module is used to evaluate the actual performance data of the month based on the data and determine the actual data indicators; The fourth determining module is used to determine the prediction calibration coefficient for the data evaluation month based on the actual data indicators and the baseline data indicators. The first prediction module is used to predict the target analysis value of the data evaluation month based on the prediction calibration coefficient, the baseline data index, and the actual data index. The generation module is used to generate business management strategies for the target business based on the target analysis values.
8. The apparatus according to claim 7, characterized in that, The second determining module is further configured to: Based on the historical performance data of the same historical month, the baseline delinquency rate, first baseline credit score, baseline initial credit score, and first baseline risk value of the same historical month are determined. The third determining module is further configured to: Based on the actual performance data of the data assessment month, the actual delinquency rate, actual credit score, actual first-payment credit score, and first actual due date percentage for the same period of the data assessment month are determined.
9. An electronic device, comprising: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the data analysis method according to any one of claims 1-6.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the data analysis method according to any one of claims 1-6.