Historical behavior coupon brushing risk prediction method and system for saas platform designated driver
By comprehensively using data analysis and algorithm models on the designated driver's history and real-time risk scores, the platform's shortcomings in coupon risk control are solved, and more accurate risk prediction and more effective risk control strategies are achieved, ensuring the fairness and healthy development of the platform.
Patent Information
- Application Number
- CN202510132429.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-10
AI Technical Summary
The SaaS platform has shortcomings in coupon risk control, and it is difficult to effectively predict and prevent drivers from swiping coupons among multiple tenants, resulting in damage to the platform's economic interests and destruction of market fairness.
By comprehensively applying data analysis and algorithm models, using data warehouses and HBase to store and manage drivers' historical behavior data, generalized least squares, entropy weight method and D-S evidence theory algorithms are used to calculate drivers' historical risk scores, and the current risk score is calculated based on real-time data after order payment, and risk levels are divided to support risk control decisions.
It improves the accuracy and real-time risk of risk control, can more accurately predict the risk of drivers' coupon swiping, optimize risk control strategies, effectively prevent and crack down on coupon swiping behavior, and ensure the fairness and healthy development of the platform.
Smart Images

Figure CN120125013A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of SaaS (saas) online car-hailing platform, and in particular to a promotion management strategy for coupons or vouchers on a SaaS online car-hailing platform, and in particular to a method and system for predicting the risk of using coupons based on the historical behavior of designated drivers on the SaaS platform. Background Art
[0002] In all walks of life, coupons are widely used as a common promotional means to attract customers and increase consumer activity. However, the risk of coupon fraud has become a problem that all platforms have to face. Especially in the field of emerging designated driver SaaS platforms in recent years, this risk is particularly prominent. As a new service model, the designated driver SaaS platform has significant gaps in risk control compared to traditional non-SaaS designated driver platforms. Traditional designated driver companies, such as Didi Designated Driver and e Designated Driver, usually have a relatively complete risk control system, while the designated driver SaaS platform brings together many designated driver companies that provide transportation capacity (hereinafter referred to as "tenants"), forming a more complex market ecology.
[0003] In the designated driver SaaS platform, there may be multiple tenants in the same city, and drivers and passengers have the flexibility to freely accept and place orders between multiple platforms. This high degree of mobility facilitates the behavior of coupon swiping. In addition, the types of travel coupons are more complex than those in other industries, including but not limited to coupons for mileage deductions and coupons for time deductions. Although the refined design of these coupons enhances the user's consumption experience, it also brings challenges to the formulation of risk control strategies.
[0004] Moreover, the threshold for becoming a designated driver is relatively low, and there is no need to provide a vehicle of your own, and the process cost for drivers to register with multiple tenants is also low. This low threshold and low cost has led to the frequent movement of drivers between tenants, providing a breeding ground for coupon swiping. Drivers may switch between multiple tenants and exploit the risk control loopholes of different tenants to swipe orders, thereby making illegal profits. At present, the designated driver SaaS platform is still relatively lacking in methods for coupon risk control, and has not yet formed an effective risk control system. At the same time, compared with the risk control strategies of non-SaaS platforms, the designated driver SaaS platform also has obvious deficiencies in effectively controlling the behavior of drivers swiping coupons between multiple tenants. This not only damages the economic interests of the platform, but also destroys the fair competition environment in the market and affects the healthy development of the industry.
[0005] Therefore, it is urgent to conduct in-depth research and exploration on the coupon risk control problem of the designated driver SaaS platform and formulate more effective and targeted risk control strategies to cope with the challenges brought by the driver's coupon swiping behavior among multiple tenants. To this end, the present invention proposes a method and system for predicting the risk of coupon swiping based on the historical behavior of designated driver drivers on the SaaS platform. Summary of the Invention
[0006] In view of this, the present invention aims to provide a method and system for predicting the risk of coupon fraud by drivers of a SaaS platform for substitute drivers, so as to solve or alleviate the technical problems existing in the prior art, that is, how to provide a method for predicting whether a driver will fraudulently use coupons based on the driver's historical behavior and the characteristics of the current order, and fill the gap in relevant risk control. The technical solution of the present invention is realized as follows:
[0007] In a first aspect, a method and system for predicting the risk of coupon fraud by drivers of a SaaS platform for substitute drivers:
[0008] (1) Overview:
[0009] The present invention aims to accurately predict the risk of coupon fraud by drivers through the comprehensive application of data analysis and algorithm models. First, a data warehouse and HBase are used to store and manage the historical behavior data of drivers, including multi-dimensional information such as the number of completed orders, the number of coupon orders, and the total order amount. The generalized least squares method, the entropy weight method, and the D-S evidence theory algorithm are used to calculate and correct the historical risk scores of drivers. At the same time, real-time data after order payment, such as the amount of discounts and the amount of commission deducted from the driver, are combined to calculate the real-time risk score of the current order. According to the final risk score of the driver, different risk levels are divided, providing strong support for the risk control decision-making of the platform, thereby effectively preventing and combating coupon fraud behavior and ensuring the fairness and healthy development of the platform.
[0010] (2) Technical solution:
[0011] To achieve the above technical objectives, the present invention selects the following operating steps.
[0012] When triggered by scheduling or a specific event (new order generation or order payment completion), connect to the data warehouse and HBase to prepare for data reading and writing operations; check whether the unionId of a newly registered or existing driver exists, and if not, generate a new unionId; then, start to execute the following steps:
[0013] 2.1 Step S1, collection and preprocessing of historical behavior data:
[0014] Collect the historical behavior data of drivers in the past six months from the data warehouse; the data can be subjected to conventional existing cleaning and existing preprocessing techniques to ensure data accuracy.
[0015] 2.2 Step S2, calculation of historical risk scores:
[0016] According to the preprocessed historical data, including the number of completed orders a, the number of coupon orders b, the total order amount c, the total coupon amount d, and the number of orders marked as coupon fraud e, and the corresponding weight α for each of themi , calculate the historical risk score β of each driver's coupon swiping; use the generalized least squares method and the entropy weight method to adjust the weight α i ; apply the D-S evidence theory algorithm to correct the historical risk score to obtain β'; update the corrected historical risk score β' to the driver's index library.
[0017] 2.2.1 Step S200, calculate the initial coupon swiping risk score β of each driver:
[0018] β = α1*(b / a) + α2*(d / c) + α3*(e / b);
[0019] 2.2.2 Step S201, calculate the weight:
[0020] For each weight αi (i = 1, 2, 3), calculate the residual r of the corresponding feature term based on the generalized least squares method (for α1, the feature term is b / a; for α2, the feature term is d / c; for α3, the feature term is e / b) i .
[0021] The generalized least squares method is a regression analysis method that takes into account the heteroscedasticity of errors and can be used to estimate the relationship between weights and feature terms and calculate residuals. Substitute the residual ri into the calculation of the energy entropy ei, and the energy entropy reflects the information dispersion degree of the feature term corresponding to the weight αi. Based on the entropy weight method, reassign each weight αi according to the energy entropy ei to make the weight more in line with the actual distribution of the data.
[0022] S2010, for each weight αi and its feature term: Let the feature term be X and the weight be Y. The model based on the generalized least squares method is: Y = β0 + β1X + ∈;
[0023] Among them, β0 and β1 are parameters to be estimated, and ∈ is the error term; the generalized least squares method estimates the parameters by minimizing the weighted sum of squared residuals, that is:
[0024]
[0025] where wi is the weight inversely proportional to the variance of the error. Obtain the estimated values of the parameters by solving the above optimization problem, and then calculate the residuals
[0026] S2011, calculate the energy entropy: used to measure the dispersion degree of residuals. Based on the residuals ri (i = 1, 2,..., n), the energy entropy Ei is defined as:
[0027] Among them, pi is the probability (or relative frequency) of the occurrence of the residual ri, which is estimated by methods such as histogram or kernel density estimation, or approximated by the standard deviation, variance or other statistics of the residual to represent the energy entropy.
[0028] S2012, perform the entropy weight method: used to reassign weights according to the energy entropy. Let the original weight be αi and the energy entropy be Ei, then the new weight αi′ is:
[0029] 2.2.3 Step S202, correct the historical risk score:
[0030] Based on the driver's coupon - swiping historical risk score βt obtained at the current time step t and the driver's coupon - swiping historical risk score βti obtained at any previous time step ti, they are regarded as two pieces of evidence p1 and p2 respectively.
[0031] According to the D - S evidence theory algorithm, calculate the basic probability assignments BPA1 and BPA2 of the two pieces of evidence p1 and p2 under the recognition framework O through a preset recognition framework O. Based on the Dempster combination principle, merge BPA1 and BPA2 to obtain the corrected driver's coupon - swiping historical risk score β′;
[0032] 2.2.3.1 Step S2020, perform mapping:
[0033] Let the recognition framework be O = {low risk, medium risk, high risk}, representing three levels of the driver's coupon - swiping risk.
[0034] For the risk score βt at the current time step t, we map it onto the recognition framework O to obtain the basic probability assignment BPA1. Similarly, for the risk score βti at the previous time step ti, we also obtain the basic probability assignment BPA2. The specific mapping method is implemented through a rule set:
[0035] If βt < threshold 1, then BPA1(low risk) = 1, and the rest are 0;
[0036] If threshold 1 ≤ βt < threshold 2, then BPA1(medium risk) = 1, and the rest are 0;
[0037] If βt ≥ threshold 2, then BPA1(high risk) = 1, and the rest are 0;
[0038] Similarly, for βti, we can also obtain a similar BPA2.
[0039] 2.2.3.2 Step S2021, Dempster combination principle:
[0040] Calculate the new basic probability assignment BPA′ after combining BPA1 and BPA2, which is achieved through the Dempster combination principle:
[0041] Among them, and B ∩ C = A means that the intersection of B and C is equal to A. The term in the denominator is for normalization to avoid conflicts (i.e., the case when the intersection of B and C is an empty set at that time).
[0042] 2.2.3.3 Step S2022, correct the historical risk score β′ of the driver's coupon swiping:
[0043] According to the combined basic probability assignment BPA′, the corrected historical risk score β′ of the driver's coupon swiping can be obtained:
[0044] If BPA′(low risk) is the largest, then β′ = the score value corresponding to low risk;
[0045] If BPA′(medium risk) is the largest, then β′ = the score value corresponding to medium risk;
[0046] If BPA′(high risk) is the largest, then β′ = the score value corresponding to high risk.
[0047] Alternatively, we can also use the weighted average method to calculate β′ according to the probabilities of each risk level in BPA′, that is:
[0048] β′ = the score value corresponding to low risk × BPA′(low risk) + the score value corresponding to medium risk × BPA′(medium risk) + the score value corresponding to high risk × BPA′(high risk)
[0049] 2.2.4 Step S203, update the index library:
[0050] Update the corrected historical risk score β’ of the driver's coupon swiping into the driver's index library in HBase for subsequent risk control decisions and real-time risk score calculations.
[0051] 2.3 Step S3, post-order payment processing:
[0052] (If the current trigger event is the completion of the order payment) Check whether the order carries a coupon; if there is no coupon, set the predicted result of the driver's coupon swiping risk to 0 and end this prediction process; otherwise, analyze the coupon type (mileage deduction coupon or duration deduction coupon), and obtain relevant real-time data from the order system, including the discount amount f, the commission amount g deducted from the driver, and the actual mileage h, etc.
[0053] 2.3.1 Step S300, check whether the order carries a coupon:
[0054] Let the status of the order carrying a coupon be C, where C = 1 indicates carrying a coupon and C = 0 indicates not carrying a coupon; if C = 0, then the coupon - swiping risk of the driver in this order is 0; let the predicted result of the driver's coupon - swiping risk be R, and at this time set R to 0; end this prediction process.
[0055] Conversely, if the order carries a coupon C = 1, obtain the coupon type T, where T = 1 indicates a mileage - deduction coupon and T = 2 indicates a duration - deduction coupon;
[0056] 2.3.2 Step S301, obtain relevant real - time data from the order system:
[0057] Including the discount amount f, the commission amount g received by the driver, and the actual mileage h.
[0058] 2.4 Step S4, calculate the real - time risk score:
[0059] According to the real - time data and the preset calculation model, calculate the real - time risk score of the current order.
[0060] The basic risk factor = max(discount amount - commission amount received by the driver, 0);
[0061] Calculate the specific risk factor according to the coupon type:
[0062] (1) For the mileage - deduction coupon (T = 1): The specific risk factor is the sum of the absolute - value difference ratio of the actual mileage and the estimated mileage, and the ratio of the coupon - deducted mileage to the actual mileage.
[0063] (2) For the duration - deduction coupon (T = 2): The specific risk factor is the sum of the absolute - value difference ratio of the actual duration and the estimated duration, and the ratio of the coupon - deducted duration to the actual duration.
[0064] (3) Calculate the real - time risk score: Use the corrected historical risk score of the driver's coupon - swiping β′, the basic risk factor, and the specific risk factor to calculate the real - time risk score:
[0065] The real - time risk score of the driver's coupon - swiping = β′×basic risk factor×specific risk factor.
[0066] 2.5 Step S5, risk assessment and output:
[0067] Take the risk level as the external output index. According to the driver's final risk score, it is divided into 5 risk levels. [0, X1] is risk - free, [X1, X2] is low - risk, [2X, X3] is medium - risk, [X3, X4] is medium - high - risk, > X4 is high - risk.
[0068] (III) The mechanism for solving technical problems:
[0069] Using database systems such as data warehouses and HBase, collect and store the historical behavior data of drivers, such as order completion status, coupon usage, order amount, etc.
[0070] According to the preprocessed historical behavior data, use statistical and machine learning algorithms (generalized least squares method, entropy weight method) to calculate the historical risk scores of each driver. Use the D-S evidence theory algorithm to correct the risk scores to improve the accuracy and robustness of prediction.
[0071] After the order payment is completed, according to the real-time characteristics of the order (such as discount amount, driver's commission amount, actual mileage or duration, etc.), combined with the driver's historical risk score, calculate the real-time risk score of the current order. According to the driver's final risk score, divide them into different risk levels. Develop corresponding risk control strategies according to the risk levels, such as manual review, refusal of coupon usage, restriction of driver behavior, etc.
[0072] In the second aspect, a method and system for predicting the risk of coupon brushing for the historical behavior of drivers on the SaaS platform:
[0073] As Figure 5 shown, this system is used to implement the method and system for predicting the risk of coupon brushing for the historical behavior of drivers on the SaaS platform as described above, and it includes:
[0074] (1) Data layer:
[0075] (1.1) Data warehouse: Used to store the historical behavior data of drivers, such as order completion status, coupon usage, order amount, etc.
[0076] (1.2) HBase: As a distributed database, used to store and process large-scale data, providing efficient data reading and writing capabilities.
[0077] (2) Business logic layer:
[0078] (2.1) Data collection module: Responsible for obtaining the required data from the data layer, including historical behavior data and real-time order data.
[0079] (2.2) Data preprocessing module: Clean and preprocess the collected data, such as duplicate removal, filling missing values, data normalization, etc.
[0080] (2.3) Risk calculation module: According to the preprocessed data, use statistical and machine learning algorithms to calculate the historical risk score and real-time risk score of the driver.
[0081] (2.4) Risk assessment module: According to the results of the risk calculation module, conduct risk assessment on the driver and divide them into different risk levels.
[0082] (2.5) Strategy Execution Module: According to the risk assessment results, execute corresponding risk control strategies, such as manual review, rejecting coupon usage, restricting driver behavior, etc.
[0083] (3) Application Layer:
[0084] (3.1) User Interface: Provide an interactive interface for users to display the risk assessment results and the execution status of risk control strategies.
[0085] (3.2) API Interface: Provide data access and service interfaces for other systems or applications to achieve integration and interoperability between systems.
[0086] (4) Modules and Connection Modes between Modules:
[0087] Through the database connection pool and the Data Access Object (DAO) pattern, realize data interaction between the data layer and the business logic layer. The business logic layer accesses the database in the data layer through the DAO to perform operations such as addition, deletion, modification, and query. The business logic layer obtains the required data from the data layer, processes it, and then stores the results back to the data layer or passes them to the application layer for display.
[0088] The internal of the business logic layer adopts a microservices architecture or component-based design, and each module communicates and collaborates through interfaces. For example, the data collection module passes the collected data to the data preprocessing module, and the data preprocessing module passes the processed data to the risk calculation module, and so on.
[0089] Through communication protocols such as RESTful API or RPC (Remote Procedure Call), realize data interaction between the business logic layer and the application layer. The application layer obtains the required data and services by calling the API interfaces provided by the business logic layer. The application layer obtains the risk assessment results and the execution status of risk control strategies from the business logic layer, displays them to users or passes them to other systems for processing.
[0090] Compared with the prior art, the beneficial effects of the present invention are:
[0091] First, improve the accuracy of risk control: By comprehensively analyzing the historical behavior data of drivers and the characteristics of current orders, the present invention can more accurately predict the risk of drivers fraudulently using coupons, greatly improving the accuracy of risk identification compared with traditional risk control means.
[0092] Second, enhance the real-time performance of risk control: The real-time risk score calculation mechanism in the present invention enables the platform to immediately evaluate the driver's behavior after the order payment is completed, timely discover and handle potential risks. Statistical and machine learning algorithms (generalized least squares method, entropy weight method) are used to calculate the historical risk scores of each driver. The D-S evidence theory algorithm is used to correct the risk scores to improve the accuracy and robustness of prediction, effectively enhancing the real-time performance of risk control.
[0093] III. Optimize the risk control strategy: According to the risk levels of drivers, the present invention can formulate more accurate risk control strategies, such as taking different auditing, restricting or punishing measures for drivers with different risk levels, so as to optimize the utilization efficiency of risk control resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0095] Figure 1 It is a schematic diagram of the method flow of the present invention;
[0096] Figure 2 It is a schematic diagram of the implementation process of the present invention;
[0097] Figure 3 It is a schematic diagram of the process of step S2 of the present invention;
[0098] Figure 4 It is a schematic diagram of the sub-steps of step S2 of the present invention;
[0099] Figure 5 It is a schematic diagram of the system composition of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0100] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific embodiments of the present invention with reference to the drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below;
[0101] It should be noted that the embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0102] Explanation of related terms:
[0103] (1) Data warehouse and HBase: The data warehouse is used to store a large amount of structured data, and HBase is a distributed database used for efficient reading and writing of large-scale data.
[0104] (2) unionId: The unique identifier of the user, used to distinguish and identify different drivers.
[0105] (3) Conventional existing cleaning and existing preprocessing techniques: Process the data for denoising, filling missing values, normalization, etc., to improve the data quality.
[0106] (4) Historical behavior data: The past behavior records of the driver, such as orders, transactions, etc., used to analyze the driver's habits.
[0107] (5) The number of completed orders a: The number of orders completed by the driver.
[0108] (6) The number of coupon orders b: The number of orders completed by the driver using coupons.
[0109] (7) The total order amount c: The total amount of all orders of the driver.
[0110] (8) The total coupon amount d: The total amount of coupons used by the driver.
[0111] (9) The number of orders marked as coupon fraud e: The number of orders identified as coupon fraud behavior by the system.
[0112] (10) Generalized least squares method: A statistical method used to estimate parameters and consider the heteroscedasticity of errors.
[0113] (11) Entropy weight method: Determine the index weights according to information entropy, reflecting the degree of dispersion of index information.
[0114] (12) D-S evidence theory algorithm: Integrate multi-source information, process uncertainty and conflict, and improve the accuracy of decision-making.
[0115] (13) Coupon: A promotional method that provides discounts or offers to attract customers (drivers) to use the service.
[0116] Example 1: As Figures 1 to 4 shown, this example discloses the application of the historical behavior coupon fraud risk prediction method for drivers on the SaaS platform in the online car-hailing platform.
[0117] In the SaaS online car-hailing platform, in order to effectively prevent and control the risk of drivers fraudulently using coupons, we designed a set of risk prediction methods based on the driver's historical behavior and the characteristics of the current order. When a new order is generated or the order payment is completed, the system triggers the data processing process to evaluate the coupon fraud risk of the driver in real time.
[0118] In this example, regarding the trigger event and data preparation:
[0119] Trigger condition: A new order is generated or the order payment is completed.
[0120] Data connection: Connect the data warehouse and HBase to prepare for data reading and writing operations.
[0121] Driver identification: Check whether the unionId of newly registered or registered drivers exists. If not, generate a new unionId.
[0122] In this embodiment, regarding step S1: Historical behavior data collection and preprocessing
[0123] Data collection: Collect the historical behavior data of drivers in the past six months from the data warehouse, including the number of completed orders a, the number of coupon orders b, the total order amount c, the total coupon amount d, and the number of orders marked as coupon fraud e.
[0124] Data cleaning and preprocessing: Perform routine cleaning and preprocessing on the data, such as removing outliers and filling missing values, to ensure data accuracy.
[0125] In this embodiment, regarding step S2: Historical risk score calculation:
[0126] Step S200, calculate the initial coupon fraud risk score β for each driver:
[0127] β = α1*(b / a) + α2*(d / c) + α3*(e / b);
[0128] Step S201, calculate the weights: For each weight αi (i = 1, 2, 3), calculate the residual r of the corresponding feature term (for α1, the feature term is b / a; for α2, the feature term is d / c; for α3, the feature term is e / b) based on the generalized least squares method i .
[0129] S2010, for each weight αi and its feature term: Let the feature term be X and the weight be Y. The model representation based on the generalized least squares method is: Y = β0 + β1X + ∈;
[0130] where β0 and β1 are the parameters to be estimated, and ∈ is the error term; the generalized least squares method estimates the parameters by minimizing the weighted sum of squared residuals, that is:
[0131]
[0132] where wi is the weight inversely proportional to the variance of the error. By solving the above optimization problem, the estimated values of the parameters are obtained, and then the residuals are calculated
[0133] S2011, calculate the energy entropy: Used to measure the dispersion degree of the residuals. Based on the residuals ri (i = 1, 2,..., n), the energy entropy Ei is defined as:
[0134] Among them, pi is the probability (or relative frequency) of the occurrence of the residual ri, which is estimated by methods such as histogram or kernel density estimation, or approximated by the standard deviation, variance or other statistics of the residual to represent the energy entropy.
[0135] S2012, execute the entropy weight method: used to reassign weights according to the energy entropy. Let the original weight be αi and the energy entropy be Ei, then the new weight αi′ is:
[0136] Step S202, based on the driver's coupon swiping historical risk score βt obtained at the current time step t and the driver's coupon swiping historical risk score βti obtained at any previous time step ti, regard them as two evidences p1 and p2 respectively. According to the D-S evidence theory algorithm, calculate the basic probability assignments BPA1 and BPA2 of the two evidences p1 and p2 under the recognition framework O. Based on the Dempster combination principle, merge BPA1 and BPA2 to obtain the corrected driver's coupon swiping historical risk score β′;
[0137] Step S2020, let the recognition framework be O = {low risk, medium risk, high risk}, representing three levels of the driver's coupon swiping risk. For the risk score βt at the current time step t, we map it onto the recognition framework O to obtain the basic probability assignment BPA1. Similarly, for the risk score βti at the previous time step ti, we also obtain the basic probability assignment BPA2. The specific mapping method is implemented through a rule set:
[0138] If βt < threshold 1, then BPA1(low risk) = 1, and the rest are 0;
[0139] If threshold 1 ≤ βt < threshold 2, then BPA1(medium risk) = 1, and the rest are 0;
[0140] If βt ≥ threshold 2, then BPA1(high risk) = 1, and the rest are 0;
[0141] Similarly, for βti, we can also obtain a similar BPA2.
[0142] Step S2021, Dempster combination principle: Calculate the new basic probability assignment BPA′ after combining BPA1 and BPA2, which is realized through the Dempster combination principle:
[0143]
[0144] Among them, and B ∩ C = A means that the intersection of B and C is equal to A. The term in the denominator is to avoid conflicts (i.e., the intersection of B and C is an empty set Normalization is performed for the situation at
[0145] Step S2022, correct the historical risk score β' of the driver's coupon swiping: According to the combined basic probability assignment BPA', the corrected historical risk score β' of the driver's coupon swiping can be obtained:
[0146] If BPA'(low risk) is the largest, then β' = the score value corresponding to low risk;
[0147] If BPA'(medium risk) is the largest, then β' = the score value corresponding to medium risk;
[0148] If BPA'(high risk) is the largest, then β' = the score value corresponding to high risk.
[0149] Alternatively, we can also use the weighted average method to calculate β' according to the probabilities of each risk level in BPA', that is:
[0150] β' = the score value corresponding to low risk × BPA'(low risk) + the score value corresponding to medium risk × BPA'(medium risk) + the score value corresponding to high risk × BPA'(high risk)
[0151] Step S203, update the index library: Update the corrected historical risk score β' of the driver's coupon swiping to the index library of the driver in HBase for subsequent risk control decisions and real-time risk score calculations.
[0152] In this embodiment, regarding step S3: Post-order payment processing:
[0153] Step S300: Check whether the order carries a coupon: Let the status of the order carrying a coupon be C, where C = 1 indicates carrying a coupon and C = 0 indicates not carrying a coupon. If C = 0, then the driver's coupon swiping risk prediction result R = 0, and this prediction process ends. If C = 1, then obtain the coupon type T, where T = 1 indicates a mileage deduction coupon and T = 2 indicates a duration deduction coupon.
[0154] Step S301: Obtain relevant real-time data from the order system: including the discount amount f, the commission amount g deducted from the driver, and the actual mileage h (or actual duration, determined according to the coupon type).
[0155] In this embodiment, regarding step S4: Real-time risk score calculation:
[0156] (1) Basic risk factor: Basic risk factor = max(f - g, 0);
[0157] (2) Specific risk factor:
[0158] (2.1) For mileage deduction coupons (T = 1): Specific risk factor = Weight 2 × |Actual mileage - Estimated mileage| + Weight 3 × Actual mileage - Coupon deduction mileage
[0159] (2.2) For duration deduction coupons (T = 2): Specific risk factor = Weight 2 × |Actual duration - Estimated duration| + Weight 3 × Actual duration - Coupon deduction duration
[0160] (3) Real-time risk score: Driver's coupon-swiping real-time risk score = β′ × Basic risk factor × Specific risk factor
[0161] In this embodiment, regarding step S5: Risk assessment and output:
[0162] Use the risk level as the external output indicator. According to the driver's final risk score, it is divided into 5 risk levels. [0, X1] is risk-free, [X1, X2] is low risk, [2X, X3] is medium risk, [X3, X4] is medium-high risk, >X4 is high risk.
[0163] Embodiment 2: On the basis of the application provided in Embodiment 1, the following python execution program involved is further disclosed:
[0164]
[0165]
[0166]
[0167]
[0168]
[0169]
[0170]
[0171] In the above program:
[0172] Trigger event and data preparation: Trigger condition (new order generation or order payment completion). Data link (data warehouse and HBase). Check or generate driver identification (unionId).
[0173] Historical behavior data collection and preprocessing: Collect the driver's historical behavior data from the data warehouse. Perform data cleaning and preprocessing to ensure data accuracy.
[0174] Historical risk score calculation: Calculate the initial coupon-swiping risk score β for each driver. Keep the weights unchanged and directly return the initial risk score.
[0175] Post - payment processing of the order: Check whether the order carries a coupon. If it does, obtain the type of the coupon.
[0176] Real - time risk score calculation: Calculate the basic risk factors and specific risk factors. Combine the historical risk score β' to calculate the real - time risk score.
[0177] Risk assessment and output: Divide the risk levels according to the real - time risk score. Output the result of the risk level.
[0178] All of the above embodiments only express the implementation manners of the relevant practical applications of the present invention. The descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
[0179] For those skilled in the art, it can be further realized that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0180] Meanwhile, those skilled in the art can understand that all or part of the processes in the methods of all the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to the memory, storage, database or other media provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
Claims
1. A risk prediction method for the historical behavior of designated drivers on the SaaS platform. When a dispatch or a specific event is triggered, it checks whether the unionId of the newly registered or registered driver exists. The following features are used: Follow these steps: S1, collects the historical behavior data of recent drivers from the data warehouse; S2, based on the pre-processed historical data, including the number of completed orders a, the number of coupon orders b, the total order amount c, the total coupon amount d, and the number of orders marked as coupon swiping e, as well as the corresponding weights α i , calculate each driver's ticket swiping history risk score β; use generalized least squares method and entropy weight method to adjust the weight α i ; Apply the DS evidence theory algorithm to correct the historical risk score and obtain β'; Update the corrected historical risk score β' to the driver's indicator library; S3, check whether the order carries a coupon; if there is no coupon, set the driver's coupon risk prediction result to 0, and end the prediction process; otherwise, parse the coupon type and obtain relevant real-time data from the order system, including the discount amount f, the driver's commission amount g, and the actual mileage h; S4, calculates the real-time risk score of the current order based on the real-time data and the preset calculation model; S5, using risk level as external output indicator.
2. The historical behavior coupon risk prediction method according to claim 1 is characterized by: The execution method of S2 includes: Calculate each driver’s initial ticket risk score β: β=α1*(b / a)+α2*(d / c)+α3*(e / b); Among them, αi is the weight.
3. The historical behavior coupon risk prediction method according to claim 2 is characterized by: For each weight αi, the residual r of the corresponding feature item is calculated based on the generalized least squares method i ; Substitute the residual ri into the energy entropy ei to reflect the information discreteness of the corresponding feature item, and assign each weight αi according to the energy entropy ei based on the entropy weight method.
4. The historical behavior coupon risk prediction method according to claim 3 is characterized by: The method for copying the weight αi includes: S2010, for each weight αi and its feature item: let the feature item be X, the weight be Y, and the model representation based on generalized least squares method is: Y = β0 + β1X + ∈; Among them, β0 and β1 are the parameters to be estimated, ∈ is the error term; S2011, based on the residual ri, the energy entropy Ei is defined as: Among them, pi is the probability of occurrence of residual ri; S2012, execute the entropy weight method: let the original weight be αi, and the energy entropy be Ei, then the new weight αi′ is:
5. The historical behavior coupon risk prediction method according to claim 3 is characterized by: The execution method of the DS evidence theory algorithm includes: based on the driver's ticket swiping history risk score βt obtained at the current time step t and the driver's ticket swiping history risk score βti obtained at any previous time step ti, they are regarded as two pieces of evidence p1 and p2 respectively; according to the DS evidence theory algorithm, the basic probability distribution BPA1 and BPA2 of the two pieces of evidence p1 and p2 under the recognition framework O are calculated through the preset recognition framework O; Based on the Dempster combination principle, BPA1 and BPA2 are combined to obtain the revised driver's ticket swiping history risk score β'.
6. The method for predicting risk of historical behavior coupon swiping according to any one of claims 1 to 5, characterized in that: The implementation method of S3 includes: setting the status of the order carrying coupons to C, where C=1 indicates that the order carries coupons, and C=0 indicates that the order does not carry coupons; if C=0, then the driver's risk of swiping coupons in the order is 0; setting the driver's risk of swiping coupons to R, then setting R to 0; ending the prediction process; conversely, if the order carries coupons C=1, obtaining the coupon type T, where T=1 indicates a mileage deduction coupon, and T=2 indicates a duration deduction coupon.
7. The method for predicting risk of historical behavior coupon swiping according to any one of claims 1 to 5, characterized in that: The implementation method of S4 includes: Calculate the real-time risk score of the current order based on real-time data and the preset calculation model: Basic risk factor = max(discount amount - driver commission amount, 0); Calculate specific risk factors based on coupon type: For mileage discount coupons: The specific risk factor is the sum of the absolute difference ratio between actual mileage and estimated mileage, and the ratio between coupon discount mileage and actual mileage; For time discount coupons: the specific risk factor is the sum of the absolute difference between the actual time and the estimated time, and the ratio of the coupon discount time to the actual time; Calculate the real-time risk score: Use the modified driver's historical risk score β′, basic risk factor and specific risk factor to calculate the real-time risk score: The real-time risk score of the driver when swiping the coupon = β′×basic risk factor×specific risk factor.
8. The historical behavior coupon risk prediction method according to claim 7 is characterized by: In S5, [0, X1] means no risk, [X1, X2] means low risk, [2X, X3] means medium risk, [X3, X4] means medium-high risk, and >X4 means high risk.
9. A system for implementing the historical behavior ticket swiping risk prediction method as described in any one of claims 1 to 8, characterized in that: The system comprises: Data layer, including data warehouse and HBase; Business logic layer, including data collection module, data preprocessing module, risk calculation module, risk assessment module and strategy execution module; The application layer includes the user interface and API interface.
10. The system according to claim 9, characterized in that: The business logic layer accesses the database in the data layer through DAO; the business logic layer obtains the required data from the data layer, stores the results back to the data layer after processing, or passes them to the application layer for display.
Citation Information
Cited By
Cross-service category integration and promotion method and system based on user demand analysis
CN120410611A