A risk data detection method, device, equipment and medium
By analyzing ride-hailing drivers' marketing campaign behavior data, monitoring operational details in real time, and dynamically adjusting risk control strategies, the problems of high false alarm rates and missed alarm risks in existing technologies have been solved, achieving a more accurate risk control effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING BAIJU YIXING TECH CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-05-29
Smart Images

Figure CN122115023A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of ride-hailing risk monitoring, specifically to a method, device, equipment, and medium for detecting risk data. Background Technology
[0002] In the operation of ride-hailing platforms, various marketing campaigns are frequently launched to increase driver activity and order volume. Risk control during these campaigns is a crucial aspect of ensuring platform financial security and operational order. Currently, the mainstream risk control methods in the industry mostly employ rule-based risk control systems. These systems assess the risk of drivers' behavior during marketing campaigns by pre-setting fixed rules such as order quantity, reward amount, and withdrawal frequency. However, with the increasing complexity of ride-hailing services and the diversification of fraud methods, traditional rule-based risk control systems are gradually showing their inadequacy. Firstly, the rules are often quite simple, making it difficult to accurately distinguish between normal and risky behaviors, resulting in a high false alarm rate. For example, some drivers may have high-frequency orders and high reward collection records due to long working hours and active participation in platform activities, which may be misjudged as risky behaviors such as reward fraud. Secondly, static pre-set rules are insufficient to cover new cheating methods such as collusion between drivers and passengers and cashing out through illegal channels, resulting in a significant risk of underreporting and an inability to identify new risks in a timely manner. Third, the rules engine lacks dynamic adjustment capabilities and cannot perform real-time self-optimization based on market strategy adjustments, changes in the operating environment, etc., making it difficult to meet complex and ever-changing risk control needs. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a method, apparatus, device and medium for detecting risk data, in order to solve the problems of high false alarm rate, easy omission of new risks and lack of flexibility and adaptability in the prior art.
[0004] In a first aspect, embodiments of the present invention provide a method for detecting risk data, the method comprising: Acquire behavioral data generated by ride-hailing drivers during marketing activities, analyze the behavioral data, and determine the operational risk level of the ride-hailing drivers; If the operational risk level is greater than or equal to the preset risk level, then the target risk control mode configured for the marketing campaign is obtained; Real-time monitoring of whether the operational details of the ride-hailing driver in the marketing campaign meet the control conditions of the target risk control model; If the operational details data meets the control conditions, then risk control operations are performed on the operational details data according to the processing strategy corresponding to the target risk control mode.
[0005] Furthermore, the analysis of the behavioral data to determine the operational risk level of the ride-hailing driver includes: Extract key behavioral data corresponding to different dimensions from the behavioral data; The key behavioral data corresponding to each dimension are quantified to obtain the quantified values for each dimension. Calculate the probability value used to characterize the degree of risk based on the quantitative values corresponding to each dimension; Obtain the risk threshold range corresponding to the marketing activity from the decision tree, and compare the probability value with the risk threshold range to obtain the target risk threshold range in which the probability value falls; The risk level corresponding to the target risk threshold range is used as the operational risk level.
[0006] Furthermore, the calculation of probability values to characterize the degree of risk based on the quantified values corresponding to each dimension includes: Based on the characteristics of historical risk orders, the basic risk weights for each dimension are determined. Based on the activity type and risk control focus of the marketing campaign, determine the dynamic adjustment coefficients for each dimension, and use the basic risk weights and the dynamic adjustment coefficients to calculate the target weights for each dimension. Based on the target weights and corresponding quantified values of each dimension, the initial risk score is calculated, and the preliminary probability value after mapping the initial risk score is determined according to the optimization function parameters of the ride-hailing risk control scenario. The correction factor is determined based on the historical operating data of the ride-hailing driver, and the probability value is calculated based on the correction factor and the preliminary probability value.
[0007] Furthermore, the real-time detection of whether the ride-hailing driver's operational details data in the marketing activity meets the control conditions of the target risk control model includes: Determine the corresponding monitoring dimensions based on the type of the target risk control model; Obtain the characteristic distribution pattern of historical risk orders, and use the characteristic distribution pattern to determine the anomaly judgment rules corresponding to each monitoring dimension; Extract the corresponding dimension detail data from the operational detail data according to the monitoring dimensions; Analyze whether the detailed data of the dimension triggers the corresponding anomaly judgment rule of the monitoring dimension, and obtain the judgment result corresponding to each monitoring dimension; Based on the judgment results corresponding to each of the monitoring dimensions, it is determined whether the operational detail data meets the control conditions of the target risk control model.
[0008] Furthermore, the step of obtaining the characteristic distribution pattern of historical risk orders and using the characteristic distribution pattern to determine the anomaly judgment rules corresponding to each monitoring dimension includes: Based on the order association data of historical risk orders, determine the distribution range of feature values and the degree of data dispersion for each monitoring dimension; Obtain the extreme values of the feature value distribution interval, and use the extreme values and the data dispersion to determine the basic risk threshold for each monitoring dimension; The adjustment coefficient is determined based on the risk control level and business rules corresponding to the marketing activity; The basic risk threshold is adjusted using the adjustment coefficient to obtain the target risk threshold, and the anomaly determination rule is constructed using the target risk threshold.
[0009] Furthermore, the analysis of whether the detailed data of the dimension triggers the corresponding anomaly judgment rule of the monitoring dimension, and the determination result corresponding to each monitoring dimension, includes: If the target risk control mode is a real-time risk control mode, then determine the evaluation strategy suitable for the marketing campaign; If the evaluation strategy is a pre-evaluation strategy, then determine whether the dimension detail data of the first monitoring dimension corresponding to the pre-evaluation strategy exceeds the threshold; or, if the evaluation strategy is a post-evaluation strategy, then determine whether the dimension detail data of the second monitoring dimension corresponding to the post-evaluation strategy exceeds the threshold. Based on the threshold comparison results of detailed data in each dimension and the priority rules of the real-time risk control mode, the judgment results corresponding to each monitoring dimension under the pre-strategy or post-strategy are determined.
[0010] Furthermore, the analysis of whether the detailed data of the dimension triggers the corresponding anomaly judgment rule of the monitoring dimension, and the determination result corresponding to each monitoring dimension, includes: If the target risk control mode is an offline risk control mode, the anomaly judgment rule corresponding to the third monitoring dimension is determined according to the reward feature monitoring requirements of the offline risk control mode. Obtain the dimension detail data corresponding to the third monitoring dimension from the dimension detail data; The dimension detail data corresponding to the third monitoring dimension is matched with the anomaly judgment rule to obtain the judgment result corresponding to each third monitoring dimension.
[0011] Secondly, embodiments of the present invention provide a risk data detection device, the device comprising: The first acquisition module is used to acquire behavioral data generated by ride-hailing drivers during marketing activities, and analyze the behavioral data to determine the operational risk level of the ride-hailing drivers; The second acquisition module is used to acquire the target risk control mode configured for the marketing activity if the operational risk level is greater than or equal to the preset risk level. The detection module is used to detect in real time whether the operational details data of the ride-hailing driver in the marketing activity meet the risk threshold of the target risk control mode; The control module is used to perform risk control operations on the operational detail data according to the processing strategy corresponding to the target risk control mode if the operational detail data meets the risk threshold.
[0012] Thirdly, embodiments of the present invention provide a computer device, including: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the method described in the first aspect or any corresponding embodiment thereof.
[0013] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer instructions that cause a computer to perform the method described in the first aspect or any of its corresponding embodiments.
[0014] This application determines operational risk levels by analyzing driver marketing activity behavior data, moving away from relying on single static rules and combining multi-dimensional behavioral data for comprehensive judgment. This distinguishes between drivers' normal, high-frequency participation in activities and abnormal prize-grabbing behavior, reducing false alarm rates. Secondly, it matches corresponding target risk control models based on risk levels, adapting different monitoring logics to different risks, covering more risk scenarios, and identifying new risks such as driver-passenger collusion, reducing missed reports. Thirdly, it monitors detailed operational data in real time and matches the control conditions of corresponding risk control models, allowing for adjustments to monitoring dimensions and rules based on the configuration of marketing activities, providing flexibility. Finally, it matches corresponding processing strategies based on control conditions, employing differentiated treatment for different risks, further improving the accuracy of risk control and addressing the shortcomings of existing technologies. Attached Figure Description
[0015] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating a risk data detection method according to some embodiments of the present invention; Figure 2 This is a flowchart illustrating another risk data detection method according to some embodiments of the present invention; Figure 3 This is a structural block diagram of a risk data detection device according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] According to embodiments of the present invention, a method, apparatus, device, and medium for detecting risk data are provided. It should be noted that the steps shown in the flowcharts in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0019] This embodiment provides a method for detecting risk data. Figure 1 This is a flowchart of a risk data detection method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Obtain behavioral data generated by ride-hailing drivers during marketing activities, analyze the behavioral data, and determine the operational risk level of ride-hailing drivers.
[0020] In this embodiment, real-time behavioral data such as driver clicks on activities and reward claims are first obtained through on-device tracking. This is combined with order event monitoring to capture related data such as order acceptance and dispatch. Activity configuration data, driver activity matching data, and reward data are synchronized to form a complete dataset. The data is then cleaned, removing duplicate records, filling in null values, and correcting abnormal data. Effective features such as activity type and reward amount are extracted, and derived features such as reward claim ratio and average reward amount per order are constructed. After standardization, the data is stored in HBase. Finally, using these feature data as input, a probability score is calculated using the Sigmoid function. Combined with a decision tree model for correlation analysis, driver operational risk is classified into high, medium, and low levels based on features such as activity type and driver reward frequency.
[0021] In this embodiment of the application, behavioral data is analyzed to determine the operational risk level of ride-hailing drivers, including steps A1-A5: Step A1: Extract key behavioral data corresponding to different dimensions from the behavioral data.
[0022] Specifically, the system reads the marketing activity behavior dataset of ride-hailing drivers and loads a pre-defined list of risk-related dimensions, including dimensions such as order operations, reward collection, and activity participation. It then uses a feature filtering algorithm to traverse the behavioral data fields, matching key behavioral indicators corresponding to each dimension. For example, the order dimension is associated with order acceptance timeliness and dispatch distance, while the reward dimension is associated with the amount and frequency of reward collection. Finally, a field extraction function is used to extract these key behavioral data, which are then categorized and stored by dimension, forming a mapping table of "dimension-key behavioral data" to complete the extraction of multi-dimensional key behavioral data.
[0023] Step A2 involves quantifying the key behavioral data corresponding to each dimension to obtain the quantified values for each dimension.
[0024] Specifically, the system calls the corresponding quantitative rule library for each dimension. The rule library pre-sets quantitative methods for different key behavioral data: for numerical data such as dispatch distance, it uses min-max normalization to map to the [0,1] interval; for categorical data such as driver-passenger relationship, it assigns a quantitative score of 1-5 according to the frequency of association; for count data such as award frequency, it compresses the numerical range through logarithmic transformation; and performs the corresponding quantitative operation on the key behavioral data of each dimension one by one to generate a single quantitative value for each dimension, thus constructing a corresponding dataset of "dimension-quantitative value".
[0025] Step A3: Calculate the probability value used to characterize the degree of risk based on the quantitative values corresponding to each dimension.
[0026] Specifically, probability values representing the degree of risk are calculated based on the quantitative values corresponding to each dimension. This includes: determining the basic risk weights for each dimension based on the characteristic correlation analysis results of historical risk orders; determining the dynamic adjustment coefficients for each dimension based on the activity type and risk control focus of the marketing campaign, and calculating the target weights for each dimension using the basic risk weights and dynamic adjustment coefficients; calculating the initial risk score based on the target weights and corresponding quantitative values for each dimension, and determining the preliminary probability value after mapping the initial risk score based on the optimization function parameters of the ride-hailing risk control scenario; determining the correction factor based on the historical operating conditions of ride-hailing drivers, and calculating the probability value based on the correction factor and the preliminary probability value.
[0027] Understandably, the process begins by acquiring historical risk order datasets and feature correlation analysis results, extracting the correlation values between each dimension and risk orders from the analysis results; normalizing the correlation values, eliminating the influence of dimensions through a weighted summation algorithm, and obtaining the initial weight values for each dimension; using regularization to constrain the initial weight values to avoid overfitting; and storing the processed weight values as the basic risk weights for each dimension in the weight configuration library and marking the data version.
[0028] Read the marketing campaign type tags and risk control key parameters, match them with the preset dynamic adjustment coefficient matrix, where different campaign types and risk control key parameters correspond to unique coefficient values; extract the dynamic adjustment coefficients corresponding to each dimension, and multiply them with the basic risk weights determined in step 1; normalize the calculation results so that the sum of the target weights of all dimensions is 1; generate a relationship table of "dimension-basic weight-dynamic coefficient-target weight" to complete the calculation and storage of the target weights of each dimension.
[0029] Obtain the target weights and corresponding quantified values for each dimension, and calculate the initial risk score using the weighted summation formula, i.e., score = Σ(target weight × quantified value); call the optimization function for the ride-hailing risk control scenario, and load the preset function parameters (including the score interval and probability mapping coefficient); input the initial risk score into the optimization function, and convert the score into a preliminary probability value in the [0,1] interval through the Sigmoid mapping; round the preliminary probability value to three decimal places, output it, and temporarily store it in the cache.
[0030] Historical operational datasets of ride-hailing drivers were retrieved, and indicators such as order compliance rate, complaint rate, and reward claim compliance records for the past 30 days were extracted. After standardizing the indicators, the weights of each indicator were calculated using the analytic hierarchy process (AHP). The comprehensive operational performance score was obtained by weighted summation, and the score was mapped to a correction factor in the range of [0.8, 1.2]. The better the performance, the closer the factor was to 0.8, and vice versa. The correction factor was multiplied by the preliminary probability value to obtain the final risk probability value. After the calculation was completed, the result was updated to the risk control dataset.
[0031] The system analyzes the type identifier of the target risk control mode. If it is real-time risk control, the core monitoring dimensions are dispatch distance, driver-passenger relationship, order acceptance time, order mileage, and amount. If it is offline risk control, the monitoring dimensions are award frequency, participation frequency, average order amount, and periodic reward amount. The system initiates a second-level data collection process and standardizes the format of the collected data. The data is compared with preset control conditions. Real-time risk control checks whether the distance and relationship frequency exceed the limits, while offline risk control checks whether the reward amount exceeds the average value, and outputs the matching results.
[0032] Step A4: Obtain the risk threshold range corresponding to the marketing campaign from the decision tree, and compare the probability value with the risk threshold range to obtain the target risk threshold range in which the probability value falls.
[0033] Load the trained risk control decision tree model, traverse the decision tree nodes according to the type identifier of the current marketing activity, and locate the risk threshold configuration branch corresponding to the activity; extract the high, medium and low risk threshold interval data under the branch to form a threshold interval list; substitute the probability value into the interval comparison algorithm, and judge whether the probability value falls into each interval in turn, record the interval that is successfully matched for the first time, determine it as the target risk threshold interval, and output the interval identifier and the corresponding threshold range.
[0034] Step A5: Assign the risk level corresponding to the target risk threshold range as the operational risk level.
[0035] The system invokes a risk grading mapping table, which pre-defines the correspondence between "threshold range - risk level". It inputs the identifier of the target risk threshold range to obtain the corresponding risk grading, such as a high threshold range corresponding to a high risk level. Simultaneously, based on the target risk control mode type, it determines the dimensions for monitoring operational details: real-time risk control focuses on data such as order dispatch distance, while offline risk control focuses on data such as reward amount. Data is collected at a rate of seconds and compared with control conditions, outputting the data matching results. Finally, the risk grading is used as the operational risk level.
[0036] As an example, in real-time risk control mode, the first step is to determine the operational details data dimensions to be monitored, namely, the dispatch distance, driver-passenger relationship, and order acceptance time before order settlement, and the order mileage and order amount after settlement. The system synchronizes the corresponding data from the order system and driver-passenger relationship database at a rate of seconds. It performs range verification on the dispatch distance to determine whether it exceeds the normal range of 0.5-10km, performs frequency statistics on the driver-passenger relationship to determine whether it exceeds 5 times in a single day, and performs average comparison on the order amount to determine whether it exceeds 30% of the average order amount in the same area and at the same time. At the same time, it checks whether the order acceptance time exceeds the threshold of 3 minutes after dispatch. The data is compared with the control conditions one by one to confirm whether the data matches the anomaly judgment rules.
[0037] Step S102: If the operational risk level is greater than or equal to the preset risk level, then obtain the target risk control mode configured for the marketing campaign.
[0038] In this embodiment, a preset operational risk level threshold is established, and the determined driver operational risk level is compared with this threshold. If the risk level value is greater than or equal to the preset threshold, the marketing activity configuration database is retrieved, and the risk control mode identifier field in the activity configuration parameters is parsed. The corresponding risk control mode type is matched according to the identifier field. If the identifier is an immediate activity (such as commission-free), the real-time risk control mode is matched; if the identifier is a delayed activity (such as periodic settlement and reward), the offline risk control mode is matched. The control rule parameters corresponding to the target risk control mode are extracted, including the data dimensions to be monitored, the judgment threshold, etc., to complete the acquisition and parameter loading of the target risk control mode.
[0039] Step S103: Real-time monitoring of whether the operational details of ride-hailing drivers in marketing activities meet the control conditions of the target risk control model.
[0040] In this embodiment, under real-time risk control mode, the core monitoring dimensions are dispatch distance, driver-passenger relationship, order acceptance time, order mileage, and order amount. Under offline risk control mode, the core monitoring dimensions are driver award frequency, activity participation frequency, average reward amount per order, daily / weekly / monthly cumulative reward amount, and average reward amount per duration. Operational detail data for the corresponding dimensions are collected at a second-level frequency. After the data is standardized, it is compared item by item with the control condition thresholds preset by the target risk control mode. For example, real-time risk control checks whether the dispatch distance exceeds the normal range and whether the driver-passenger relationship frequency exceeds the standard, while offline risk control checks whether the reward amount exceeds the average range of drivers of the same level. The output results determine whether the data meets the control conditions.
[0041] In this embodiment of the application, real-time detection of whether the operational details data of ride-hailing drivers in marketing activities meet the control conditions of the target risk control model includes steps B1-B5: Step B1: Determine the corresponding monitoring dimensions based on the target risk control mode's mode type.
[0042] Specifically, the target risk control mode's mode type identifier is read and matched against the preset risk control mode-monitoring dimension mapping table: if the mode type is real-time risk control, the dispatch distance before order settlement, driver-passenger relationship, order acceptance time, and the order mileage and order amount after settlement are extracted as monitoring dimensions; if the mode type is offline risk control, the driver's award frequency, activity participation frequency, and daily / monthly / weekly cumulative reward amount are extracted as monitoring dimensions, and a monitoring dimension list is generated and stored.
[0043] Step B2: Obtain the characteristic distribution pattern of historical risk orders, and use the characteristic distribution pattern to determine the anomaly judgment rules corresponding to each monitoring dimension.
[0044] Specifically, the process involves obtaining the characteristic distribution patterns of historical risk orders and using these patterns to determine the anomaly judgment rules for each monitoring dimension. This includes: determining the characteristic value distribution range and data dispersion of each monitoring dimension based on the order association data of historical risk orders; obtaining the extreme values of the characteristic value distribution range and using the extreme values and data dispersion to determine the basic risk threshold for each monitoring dimension; determining the adjustment coefficient based on the risk control level and business rules corresponding to the marketing activity; adjusting the basic risk threshold using the adjustment coefficient to obtain the target risk threshold; and using the target risk threshold to construct anomaly judgment rules.
[0045] Understandably, the process begins by retrieving the order association dataset of historical risk orders, splitting the data by monitoring dimensions, performing frequency distribution statistics on the feature values of each dimension to obtain the feature value distribution interval, calculating the variance and standard deviation to determine the data dispersion of each dimension, and storing the upper and lower boundary values of the distribution interval and the dispersion values in the dimension feature parameter library, marking the corresponding time range and order sample size to ensure the timeliness and representativeness of the parameters.
[0046] Secondly, the maximum and minimum values of the feature value distribution interval of each monitoring dimension are extracted as extreme values. Combined with the standard deviation of the data dispersion, the basic risk threshold of each dimension is calculated by the formula "basic risk threshold = extreme value ± k × standard deviation", where k is the preset dispersion coefficient. Boundary constraints are applied to the calculated basic risk threshold to ensure that the threshold falls within a reasonable business range. The basic risk threshold is then stored in the threshold configuration library.
[0047] Then, read the risk control level label and business rule document of the marketing activity, and match it with the preset control level-adjustment coefficient mapping table: when the control level is high, the adjustment coefficient is 1.2; when the control level is medium, the adjustment coefficient is 1.0; when the control level is low, the adjustment coefficient is 0.8; extract the matched adjustment coefficient, perform a second verification with the business rules of the marketing activity to ensure that the coefficient meets the risk control requirements of the activity, and store the adjustment coefficient after confirmation.
[0048] Finally, the basic risk threshold and adjustment coefficient of each monitoring dimension are multiplied to obtain the target risk threshold; anomaly judgment rules are constructed based on the target risk threshold, and the rule content is "if the detailed data of the dimension exceeds the target risk threshold range, it is judged as abnormal"; then, the monitoring dimensions are determined according to the target risk control mode type: real-time risk control matches order-related dimensions such as dispatch distance, and offline risk control matches driver reward-related dimensions.
[0049] Step B3: Extract the corresponding dimension detail data from the operation detail data according to the monitoring dimensions.
[0050] Specifically, the monitoring dimension list is loaded, the field index of the operational detail data is traversed, and the field names corresponding to each monitoring dimension are matched; the values of the corresponding fields are extracted from the operational detail data through the field extraction function, and the data is sorted and organized according to the monitoring dimensions to form a dimension detail dataset; the dimension detail data is formatted uniformly, the categorized data is converted into numerical data, and missing values are filled with the mean of the same dimension to ensure that the data format meets the requirements of subsequent analysis.
[0051] Step B4: Analyze whether the detailed data of the analysis dimension triggers the corresponding anomaly judgment rules of the monitoring dimension, and obtain the judgment results corresponding to each monitoring dimension.
[0052] In this embodiment, the analysis of whether the detailed data of the analysis dimension triggers the corresponding anomaly judgment rule of the monitoring dimension is used to obtain the judgment result corresponding to each monitoring dimension, including: if the target risk control mode is a real-time risk control mode, then the evaluation strategy adapted to the marketing activity is determined; if the evaluation strategy is a pre-positional strategy, then whether the detailed data of the first monitoring dimension corresponding to the pre-positional strategy exceeds the threshold is determined; or, if the evaluation strategy is a post-positional strategy, then whether the detailed data of the second monitoring dimension corresponding to the post-positional strategy exceeds the threshold is determined; based on the threshold comparison results of the detailed data of each dimension and the priority rules of the real-time risk control mode, the judgment result corresponding to each monitoring dimension under the pre-positional strategy or the post-positional strategy is determined.
[0053] Understandably, the process begins by reading the target risk control mode identifier. If it's a real-time risk control mode, the system retrieves the marketing campaign's attribute parameters, matches them against a pre-defined evaluation strategy mapping table, and determines the appropriate pre- or post-strategy. If it's a pre-strategy, the system extracts detailed data for the first monitoring dimension (order dispatch distance, driver-passenger relationship, order acceptance time) and compares it with the corresponding thresholds. If it's a post-strategy, the system extracts detailed data for the second monitoring dimension (order mileage, order amount) and compares it with the corresponding thresholds. The system then reads the priority rules for the real-time risk control mode. If the pre-strategy has abnormal results, the pre-strategy judgment result is output first; otherwise, the post-strategy judgment result is output. The judgment results for each monitoring dimension are then summarized.
[0054] Secondly, read the type identifier of the target risk control mode and match it with the preset monitoring dimension configuration library: under real-time risk control mode, extract the dispatch distance, driver-passenger relationship, order acceptance time before order settlement, and the order mileage and amount after settlement as monitoring dimensions; under offline risk control mode, extract the driver award frequency, activity participation frequency, and daily / monthly / weekly cumulative reward amount as monitoring dimensions; start a second-level data collection task to synchronize the detailed data of the corresponding dimensions from the operation data flow channel, perform data format normalization processing, and compare it item by item with the control condition threshold of the target risk control mode, and output the matching result of each dimension.
[0055] As an example, the target risk control mode is identified as real-time risk control. The activity attribute parameters are retrieved and matched with the pre-assessment strategy. The first monitoring dimension is the dispatch distance, driver-passenger relationship, and order acceptance time, with corresponding thresholds of 0.5-10km, ≤5 associations per day, and order acceptance time ≤3 minutes after dispatch, respectively. The system synchronizes data from the order system and driver-passenger relationship database at a frequency of seconds. For a certain order of driver Zhang during the morning rush hour, the dispatch distance is 58km, the number of driver-passenger associations per day is 7, and the order acceptance time is 2 minutes. After comparing the data with the thresholds, the dispatch distance and driver-passenger relationship both exceed the thresholds, which is judged as an abnormality of the pre-assessment strategy. According to the real-time risk control priority rules, the pre-assessment result is output first. At the same time, the order mileage and amount data after the settlement of the order are collected simultaneously. The threshold comparison of the post-assessment dimension data is used as a supplement, and finally the risk control judgment result of the order is obtained by summarizing.
[0056] In this embodiment of the application, the analysis of whether the detailed data of the dimension triggers the corresponding anomaly judgment rule of the monitoring dimension is used to obtain the judgment result corresponding to each monitoring dimension, including: if the target risk control mode is the offline risk control mode, the anomaly judgment rule corresponding to the third monitoring dimension is determined according to the reward feature monitoring requirements of the offline risk control mode; the detailed data of the dimension corresponding to the third monitoring dimension is obtained from the detailed data of the dimension; the detailed data of the dimension corresponding to the third monitoring dimension is matched with the anomaly judgment rule to obtain the judgment result corresponding to each third monitoring dimension.
[0057] Understandably, firstly, the target risk control mode identifier is read. If it is an offline risk control mode, the reward feature monitoring requirement parameters of the offline risk control mode are retrieved and matched with the preset third monitoring dimension (driver award frequency, activity participation frequency, daily / monthly / weekly cumulative reward amount). Based on the reward feature distribution of historical risk orders, the anomaly judgment rules for each third monitoring dimension are determined. For example, if the award frequency exceeds the 75th percentile of drivers of the same level or the cumulative reward amount exceeds the average by 50%, it is judged as abnormal. The anomaly judgment rules are stored in the rule cache area.
[0058] Secondly, load the list of third monitoring dimensions, traverse the field indexes of the dimension detail dataset, and match the field names corresponding to each third monitoring dimension; extract the values of the corresponding fields from the dimension detail data using the field extraction function, classify and organize them according to the third monitoring dimension, and form a detailed dataset of the third monitoring dimension; preprocess the data, fill missing values with the mean of the same dimension, convert count data into standardized values, and ensure that the data format meets the matching requirements.
[0059] Then, the anomaly judgment rules of the third monitoring dimension are read, and the corresponding dimension details are substituted into the rules for matching in turn: for the award frequency dimension, it is judged whether the value exceeds the 75th percentile threshold; for the cumulative reward amount dimension, it is judged whether the value exceeds the 50% threshold of the average value of drivers of the same level; for the activity participation frequency dimension, it is judged whether the value exceeds the weekly participation threshold; after each dimension is matched, the judgment result (normal / abnormal) of the dimension is output, and the difference between the matched threshold and the actual value is recorded.
[0060] Finally, the target risk control mode type is read, and the corresponding operational detailed data monitoring dimensions are matched: under real-time risk control mode, the dispatch distance, driver-passenger relationship, order acceptance time, order mileage, and amount are determined as monitoring dimensions; under offline risk control mode, the driver's award frequency, participation frequency in activities, and daily / monthly / weekly cumulative reward amount are determined as monitoring dimensions; a second-level data synchronization thread is started to pull detailed data of the corresponding dimensions from the operational data link, the data is normalized, compared with the control condition threshold, and the matching results of each dimension are output, while the timestamp of the comparison and the data source are recorded synchronously.
[0061] Step B5: Based on the judgment results corresponding to each monitoring dimension, determine whether the operational details data meet the control conditions of the target risk control model.
[0062] Specifically, the judgment results of each monitoring dimension are summarized, and the control conditions configuration of the target risk control mode is read: In real-time risk control mode, if the judgment result of any dimension is abnormal, the control conditions are met; in offline risk control mode, if the judgment results of ≥2 dimensions are abnormal, or a single core dimension (such as the cumulative reward amount) is abnormal, the control conditions are met; the final judgment result is output, and the judgment details of each dimension are recorded for the basis of subsequent risk control operations.
[0063] Step S104: If the operational detail data meets the control conditions, then perform risk control operations on the operational detail data according to the processing strategy corresponding to the target risk control mode.
[0064] In this embodiment, under real-time risk control mode, pre- and post-risk control strategies are distinguished. Pre-settlement strategies target pre-settlement features such as dispatch distance and driver-passenger relationship, while post-settlement strategies target post-settlement features such as order mileage and amount. Under offline risk control mode, driver reward feature data is used as input to calculate probability values using the Sigmoid function, and high, medium, and low risks are classified according to training thresholds. Corresponding processing strategies are executed: real-time risk control triggers order interception or freezing operations, offline risk control automatically intercepts high-risk orders, pushes medium-risk orders for manual review, and releases low-risk orders. At the same time, risk control operation logs and data processing results are recorded.
[0065] As an example, such as Figure 2As shown, the multi-source data collection process is initiated: when the driver completes login, the platform generates an order event, and the activity configuration is completed and launched, the activity configuration data and prize pool data are acquired simultaneously. These five types of data (driver login data, order event data, activity configuration launch data, activity configuration data, and prize pool data) are aggregated to the data collection node to complete the centralized collection of data.
[0066] Next, duplicate records are removed using a deduplication algorithm, and empty values are filled with the average value of the same scenario. Abnormal data is corrected based on business rules to complete data cleaning. Then, the cleaned data is transformed according to preset feature dimensions to extract effective features related to orders, activities, and prize pools. Numerical features are standardized to obtain structured feature data.
[0067] Finally, the transformed and organized feature data is formatted according to the structure requirements of distributed storage and transmitted to the feature storage system through the data writing interface to achieve stable storage of feature data, providing a data foundation for subsequent risk control analysis, risk level determination and other processes.
[0068] This embodiment also provides a risk data detection device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0069] This embodiment provides a risk data detection device, such as... Figure 3 As shown, it includes: The first acquisition module 301 is used to acquire behavioral data generated by ride-hailing drivers in marketing activities, analyze the behavioral data, and determine the operational risk level of ride-hailing drivers. The second acquisition module 302 is used to acquire the target risk control mode configured for the marketing campaign if the operational risk level is greater than or equal to the preset risk level. The detection module 303 is used to detect in real time whether the operational details data of ride-hailing drivers in marketing activities meet the risk threshold of the target risk control model. The control module 304 is used to perform risk control operations on the operational detail data according to the processing strategy corresponding to the target risk control mode if the operational detail data meets the risk threshold.
[0070] In this embodiment, the first acquisition module 301 is used to extract key behavioral data corresponding to different dimensions from behavioral data; quantify the key behavioral data corresponding to each dimension to obtain the quantified value corresponding to each dimension; calculate the probability value used to characterize the degree of risk based on the quantified value corresponding to each dimension; obtain the risk threshold interval corresponding to the marketing activity from the decision tree, and compare the probability value with the risk threshold interval to obtain the target risk threshold interval into which the probability value falls; and classify the risk level corresponding to the target risk threshold interval as the operational risk level.
[0071] In this embodiment, the first acquisition module 301 is used to determine the basic risk weights of each dimension based on the characteristic correlation analysis results of historical risk orders; determine the dynamic adjustment coefficients of each dimension based on the activity type and risk control focus of the marketing activity, and calculate the target weights of each dimension using the basic risk weights and dynamic adjustment coefficients; calculate the initial risk score based on the target weights of each dimension and the corresponding quantitative values, and determine the preliminary probability value after mapping the initial risk score based on the optimization function parameters of the ride-hailing risk control scenario; determine the correction factor based on the historical operation of ride-hailing drivers, and calculate the probability value based on the correction factor and the preliminary probability value.
[0072] In this embodiment, the detection module 303 is used to determine the corresponding monitoring dimensions according to the mode type of the target risk control mode; obtain the feature distribution pattern of historical risk orders, and use the feature distribution pattern to determine the anomaly judgment rules corresponding to each monitoring dimension; extract the corresponding dimension detail data from the operational detail data according to the monitoring dimensions; analyze whether the dimension detail data triggers the corresponding anomaly judgment rules of the monitoring dimensions, and obtain the judgment results corresponding to each monitoring dimension; and determine whether the operational detail data meets the control conditions of the target risk control mode based on the judgment results corresponding to each monitoring dimension.
[0073] In this embodiment, the detection module 303 is used to determine the feature value distribution range and data dispersion of each monitoring dimension based on the order association data of historical risk orders; obtain the extreme values of the feature value distribution range, and use the extreme values and data dispersion to determine the basic risk threshold of each monitoring dimension; determine the adjustment coefficient according to the risk control level and business rules corresponding to the marketing activity; adjust the basic risk threshold using the adjustment coefficient to obtain the target risk threshold, and use the target risk threshold to construct anomaly judgment rules.
[0074] In this embodiment, the detection module 303 is used to determine the evaluation strategy suitable for the marketing activity if the target risk control mode is a real-time risk control mode; if the evaluation strategy is a pre-emptive strategy, determine whether the dimension details data of the first monitoring dimension corresponding to the pre-emptive strategy exceeds the threshold; or, if the evaluation strategy is a post-emptive strategy, determine whether the dimension details data of the second monitoring dimension corresponding to the post-emptive strategy exceeds the threshold; and determine the judgment result corresponding to each monitoring dimension under the pre-emptive strategy or the post-emptive strategy based on the threshold comparison results of each dimension details data and the priority rules of the real-time risk control mode.
[0075] In this embodiment, the detection module 303 is used to determine the anomaly judgment rule corresponding to the third monitoring dimension based on the reward feature monitoring requirements of the offline risk control mode if the target risk control mode is an offline risk control mode; obtain the dimension detail data corresponding to the third monitoring dimension from the dimension detail data; and match the dimension detail data corresponding to the third monitoring dimension with the anomaly judgment rule to obtain the judgment result corresponding to each third monitoring dimension.
[0076] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 4 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure X shows an example of a single processor 10.
[0077] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.
[0078] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.
[0079] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device as shown by a landing page for an app. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0080] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0081] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.
[0082] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0083] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for detecting risk data, characterized in that, The method includes: Acquire behavioral data generated by ride-hailing drivers during marketing activities, analyze the behavioral data, and determine the operational risk level of the ride-hailing drivers; If the operational risk level is greater than or equal to the preset risk level, then the target risk control mode configured for the marketing campaign is obtained; Real-time monitoring of whether the operational details of the ride-hailing driver in the marketing campaign meet the control conditions of the target risk control model; If the operational details data meets the control conditions, then risk control operations are performed on the operational details data according to the processing strategy corresponding to the target risk control mode.
2. The method according to claim 1, characterized in that, The analysis of the behavioral data to determine the operational risk level of the ride-hailing driver includes: Extract key behavioral data corresponding to different dimensions from the behavioral data; The key behavioral data corresponding to each dimension are quantified to obtain the quantified values for each dimension. Calculate the probability value used to characterize the degree of risk based on the quantitative values corresponding to each dimension; Obtain the risk threshold range corresponding to the marketing activity from the decision tree, and compare the probability value with the risk threshold range to obtain the target risk threshold range in which the probability value falls; The risk level corresponding to the target risk threshold range is used as the operational risk level.
3. The method according to claim 2, characterized in that, The calculation of probability values to characterize the degree of risk based on the quantitative values corresponding to each dimension includes: Based on the characteristics of historical risk orders, the basic risk weights for each dimension are determined. Based on the activity type and risk control focus of the marketing campaign, determine the dynamic adjustment coefficients for each dimension, and use the basic risk weights and the dynamic adjustment coefficients to calculate the target weights for each dimension. Based on the target weights and corresponding quantified values of each dimension, the initial risk score is calculated, and the preliminary probability value after mapping the initial risk score is determined according to the optimization function parameters of the ride-hailing risk control scenario. The correction factor is determined based on the historical operating data of the ride-hailing driver, and the probability value is calculated based on the correction factor and the preliminary probability value.
4. The method according to claim 1, characterized in that, The real-time detection of whether the ride-hailing driver's operational details data in the marketing campaign meets the control conditions of the target risk control model includes: Determine the corresponding monitoring dimensions based on the type of the target risk control model; Obtain the characteristic distribution pattern of historical risk orders, and use the characteristic distribution pattern to determine the anomaly judgment rules corresponding to each monitoring dimension; Extract the corresponding dimension detail data from the operational detail data according to the monitoring dimensions; Analyze whether the detailed data of the dimension triggers the corresponding anomaly judgment rule of the monitoring dimension, and obtain the judgment result corresponding to each monitoring dimension; Based on the judgment results corresponding to each of the monitoring dimensions, it is determined whether the operational detail data meets the control conditions of the target risk control model.
5. The method according to claim 4, characterized in that, The process of obtaining the feature distribution patterns of historical risk orders and using these patterns to determine the anomaly judgment rules for each monitoring dimension includes: Based on the order association data of historical risk orders, determine the distribution range of feature values and the degree of data dispersion for each monitoring dimension; Obtain the extreme values of the feature value distribution interval, and use the extreme values and the data dispersion to determine the basic risk threshold for each monitoring dimension; The adjustment coefficient is determined based on the risk control level and business rules corresponding to the marketing activity; The basic risk threshold is adjusted using the adjustment coefficient to obtain the target risk threshold, and the anomaly determination rule is constructed using the target risk threshold.
6. The method according to claim 4, characterized in that, The analysis determines whether the detailed data of the dimension triggers the corresponding anomaly judgment rule for the monitoring dimension, and obtains the judgment result for each monitoring dimension, including: If the target risk control mode is a real-time risk control mode, then determine the evaluation strategy suitable for the marketing campaign; If the evaluation strategy is a pre-evaluation strategy, then determine whether the dimension detail data of the first monitoring dimension corresponding to the pre-evaluation strategy exceeds the threshold; or, if the evaluation strategy is a post-evaluation strategy, then determine whether the dimension detail data of the second monitoring dimension corresponding to the post-evaluation strategy exceeds the threshold. Based on the threshold comparison results of detailed data in each dimension and the priority rules of the real-time risk control mode, the judgment results corresponding to each monitoring dimension under the pre-strategy or post-strategy are determined.
7. The method according to claim 4, characterized in that, The analysis determines whether the detailed data of the dimension triggers the corresponding anomaly judgment rule for the monitoring dimension, and obtains the judgment result for each monitoring dimension, including: If the target risk control mode is an offline risk control mode, the anomaly judgment rule corresponding to the third monitoring dimension is determined according to the reward feature monitoring requirements of the offline risk control mode. Obtain the dimension detail data corresponding to the third monitoring dimension from the dimension detail data; The dimension detail data corresponding to the third monitoring dimension is matched with the anomaly judgment rule to obtain the judgment result corresponding to each third monitoring dimension.
8. A risk data detection device, characterized in that, The device includes: The first acquisition module is used to acquire behavioral data generated by ride-hailing drivers during marketing activities, and analyze the behavioral data to determine the operational risk level of the ride-hailing drivers; The second acquisition module is used to acquire the target risk control mode configured for the marketing activity if the operational risk level is greater than or equal to the preset risk level. The detection module is used to detect in real time whether the operational details data of the ride-hailing driver in the marketing activity meet the risk threshold of the target risk control mode; The control module is used to perform risk control operations on the operational detail data according to the processing strategy corresponding to the target risk control mode if the operational detail data meets the risk threshold.
9. A computer device, characterized in that, include: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the method of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the method of any one of claims 1 to 7.