Online car-hailing risk assessment method and device, computer equipment and storage medium
By constructing a ride-hailing risk assessment method based on a multi-dimensional feature set and dynamic fusion parameters, the problems of insufficient scenario adaptability and high operational resource investment in the risk control system are solved, and the accurate identification and real-time prevention and control of ride-hailing risks are achieved.
Patent Information
- Application Number
- CN202510830077.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-21
AI Technical Summary
The existing online ride-hailing risk control system lacks adaptability to different scenarios, has delayed dynamic responses, and requires high investment in operational resources, making it difficult to accurately prevent and control false trips and new risks.
By constructing a target feature set, including basic features, cross features, composite features and embedded features, risk prediction is performed using multiple pre-trained scenario models, and a target risk score is generated through weighted fusion using dynamic fusion parameters.
It improves the accuracy and real-time performance of risk identification, reduces operating costs, and enhances the adaptability to ride-hailing-specific risks and the transparency of the model.
Smart Images

Figure CN120822822A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of online ride-hailing technology, and in particular to a method, device, computer equipment, and storage medium for assessing the risks of online ride-hailing. Background Art
[0002] With the deep integration of artificial intelligence and intelligent transportation systems, the online ride-hailing industry has become a vital pillar of urban transportation. Platforms have accumulated vast amounts of user, transaction, and device data during their operations, providing an information foundation for risk identification. However, current online ride-hailing risk control systems generally rely on third-party black-box models, which suffer from flaws such as opaque model logic, weak scenario adaptability, delayed rule updates, and high operating costs. This makes it difficult to accurately control unique risks such as fraudulent trips and abnormal GPS simulation behavior, hindering the industry's security and sustainable development.
[0003] In existing technologies, general models lack targeted modeling of online ride-hailing scenario characteristics (such as fatigue driving and device fingerprint forgery), resulting in insufficient recognition accuracy; static rules are difficult to dynamically respond to new risks such as AI face-changing attacks; at the same time, the high cost of interface calls increases the burden on the platform. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method, apparatus, computer equipment, and storage medium for assessing online ride-hailing risks to address the problems of insufficient scenario adaptability, delayed dynamic response, and high investment in operational resources in existing online ride-hailing risk control systems.
[0005] In a first aspect, an embodiment of the present invention provides a method for assessing the risk of online ride-hailing services, the method comprising:
[0006] Get the original business data associated with the current online car-hailing service;
[0007] Constructing a target feature set based on the original business data, wherein the target feature set includes: basic features, cross features, composite features and embedded features;
[0008] Inputting each feature in the target feature set into multiple pre-trained scenario models according to the scenario type to obtain a risk prediction result output by each pre-trained scenario model;
[0009] The risk prediction results are weighted and fused according to dynamic fusion parameters to generate a target risk score for the current online car-hailing service.
[0010] Furthermore, the constructing of a target feature set based on the original business data includes:
[0011] Perform binning conversion on the original business data based on preset business indicators to obtain basic features;
[0012] Identify the correlation between each of the basic features to obtain cross-features;
[0013] Combining the basic features according to the scene type to obtain a composite feature;
[0014] The basic features and / or the cross features are processed using an embedding algorithm to obtain embedded features.
[0015] Furthermore, the basic features are combined according to the scene type to obtain composite features, including:
[0016] Extracting a first feature and a second feature from the basic features according to the scene type;
[0017] calculating a first index based on the first feature, and calculating a second index based on the second feature;
[0018] The first index and the second index are combined into a composite feature through a preset weight formula.
[0019] Furthermore, the method of processing the basic features and / or the cross features using an embedding algorithm to obtain embedded features includes:
[0020] Identifying the data type of the original business data;
[0021] A corresponding embedding algorithm is determined according to the data type, and the basic features and / or the cross features are processed using the embedding algorithm to obtain embedded features.
[0022] Furthermore, the method for generating the dynamic fusion parameters includes:
[0023] Obtaining the prediction accuracy of each of the pre-trained scenario models in the historical risk score data;
[0024] The initial fusion weight of the pre-trained scene model is calculated based on the prediction accuracy, and the initial fusion weight is updated according to the performance index to generate the dynamic fusion parameter.
[0025] Furthermore, after generating the target risk score for the current online ride-hailing vehicle, the method further includes:
[0026] Obtaining the actual risk detection result of the current online ride-hailing service;
[0027] Analyzing the difference data between the risk prediction results output by each of the pre-trained scenario models and the actual risk detection results;
[0028] The performance indicators of each of the pre-trained scenario models are determined based on the difference data, and the model parameters of the pre-trained scenario models and the dynamic fusion parameters of the risk prediction results are updated using the performance indicators.
[0029] Furthermore, the method further comprises:
[0030] Comparing the target risk score with the interception threshold and the release threshold respectively to obtain a comparison result;
[0031] If the comparison result is that the target risk score is greater than the interception threshold, the order interception instruction is triggered and the corresponding review task is generated; or, if the comparison result is that the target risk score is less than the release threshold, the order release instruction is triggered and the risk status of the current online car-hailing vehicle is adjusted; or, if the comparison result is that the target risk score is less than or equal to the interception threshold and greater than or equal to the release threshold, the enhanced verification instruction is triggered and the behavior sequence of the current online car-hailing vehicle is dynamically monitored.
[0032] In a second aspect, an embodiment of the present invention provides a device for assessing the risk of online ride-hailing services, the device comprising:
[0033] The acquisition module is used to obtain the original business data associated with the current online car-hailing service;
[0034] A construction module, configured to construct a target feature set based on the original business data, wherein the target feature set includes: basic features, cross features, composite features, and embedded features;
[0035] An input module, configured to input each feature in the target feature set into a plurality of pre-trained scenario models according to the scenario type, and obtain a risk prediction result output by each pre-trained scenario model;
[0036] A fusion module is used to perform weighted fusion on the risk prediction results according to dynamic fusion parameters to generate a target risk score for the current online car-hailing service.
[0037] In a third aspect, an embodiment of the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, computer instructions being stored in the memory, and the processor executing the method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0038] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method of the first aspect or any corresponding embodiment thereof.
[0039] The method provided in the embodiments of the present application has the following beneficial effects:
[0040] The method provided in the embodiment of the present application avoids dependence on third-party data interfaces by directly calling the original business data in the platform, reduces the investment in data acquisition resources and improves real-time performance; at the same time, based on the full collection of the platform's own data, it ensures that the coverage dimensions of risk identification are complete, providing a highly reliable data source for subsequent feature extraction. Through the hierarchical construction of multi-dimensional features (basic features, cross-features, composite features, and embedded features), the feature interpretability and scene adaptability are significantly improved, solving the problem of insufficient modeling of online car-hailing scene features by general models. Pre-trained scenario models are introduced for different risk scenarios of online car-hailing, and the accuracy of risk identification is improved through scenario modeling. The fusion weights of each scenario model are adjusted in real time through dynamic migration strategies (such as weight distribution and feedback mechanism) to solve the problem that static rules cannot adapt to new attacks. The dynamic fusion mechanism takes into account both model stability and flexibility, reduces the misjudgment rate, and reduces the resource investment in rule maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0042] Figure 1 1 is a flow chart of a method for assessing online ride-hailing risk according to an embodiment of the present invention;
[0043] Figure 2 2. It is a schematic diagram of the architecture of the online car-hailing risk control system according to an embodiment of the present invention;
[0044] Figure 3 Schematic diagram of feature engineering in a risk control system for online ride-hailing services according to an embodiment of the present invention;
[0045] Figure 4 2. It is a schematic diagram of the model fusion training process in the online car-hailing risk control system according to an embodiment of the present invention;
[0046] Figure 5 2 is a structural block diagram of a device for assessing online car-hailing risk according to an embodiment of the present invention;
[0047] Figure 6 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0048] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.
[0049] According to an embodiment of the present invention, a method, apparatus, computer device and storage medium for assessing the risks of online ride-hailing are provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in an order different from that shown here.
[0050] In this embodiment, a method for assessing the risk of online car-hailing is provided. Figure 1 is a flow chart of a method for assessing the risk of online car-hailing according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0051] Step S101, obtain the original business data associated with the current online car-hailing service.
[0052] In an embodiment of the present application, by calling the platform's internal database and third-party data interface, the original business data associated with the current online car-hailing order is obtained, including but not limited to user behavior logs (such as login frequency, operation trajectory), order core information (such as starting and ending points, distance, time), driver driving behavior data (such as the number of sudden brakes, fatigue driving time), device fingerprints (such as device ID, model, jailbreak status), GPS trajectory (such as path deviation rate, speed fluctuation), payment records (such as abnormal amount, refund frequency) and vehicle-mounted sensor data (such as acceleration, gyroscope signal). During the data collection process, third-party device fingerprints, IP portraits and regulatory blacklist data are pulled in real time through the API interface, and the real-time streaming data of the platform's internal order system, GPS positioning service, payment gateway and complaint management system are synchronized and stored in the original data storage layer (ODS). This step ensures the integrity and real-time nature of the data source through an autonomous data collection mechanism (rather than relying on a third-party risk control interface), provides multi-dimensional, highly reliable underlying data support for subsequent feature engineering, and avoids the additional costs and delays caused by interface calls, meeting the risk identification requirements for timeliness and data coverage dimensions.
[0053] Step S102: construct a target feature set based on the original business data, wherein the target feature set includes: basic features, cross features, composite features and embedded features.
[0054] In the embodiment of the present application, the original business data is converted into a multi-dimensional interpretable target feature set through structured hierarchical feature engineering, specifically including: binning and transforming the original data based on preset business indicators to generate basic features to solve the problem of poor interpretability of continuous features; identifying the interaction between basic features through correlation analysis, constructing cross-features, and capturing the synergistic effect of multi-dimensional risk signals; extracting key features from basic features for different scenarios and generating composite features based on preset weight formulas to achieve scenario-based risk quantification; at the same time, using embedding algorithms to process unstructured data and generate embedded features to mine potential risk patterns. This step significantly improves the adaptability of features to the unique risks of online ride-hailing services through a hierarchical feature construction mechanism, integrating statistical interpretability, business logic, and deep learning representation capabilities, and provides highly discriminative and traceable input for multi-model fusion, namely the target feature set.
[0055] In step S103, each feature in the target feature set is input into a plurality of pre-trained scenario models according to the scenario type, and a risk prediction result output by each pre-trained scenario model is obtained.
[0056] In the embodiment of the present application, through the scenario-based model diversion mechanism, each feature in the target feature set is input into the corresponding pre-trained scenario model according to the scenario type. Specifically, the following steps are performed: filtering features that are strongly correlated with scenario risk based on the scenario type, and adapting the basic features, cross-features, composite features, and embedded features to the input requirements of each scenario model; each pre-trained scenario model independently infers the input features based on its training parameters and outputs the risk prediction results for that scenario. The scenario-based model parallel processing mechanism ensures that different risk types are specifically identified by dedicated models, solving the problem of insufficient adaptability of general models to online car-hailing risks and providing a multi-dimensional risk quantification basis for subsequent dynamic integration.
[0057] It should be noted that the scenario types include abnormal behavior detection, security warning, service evaluation, and compliance management. The corresponding pre-trained scenario models are abnormal behavior detection model, security warning model, service evaluation model, and compliance management model.
[0058] Specifically, the workflow of the abnormal behavior detection model is as follows: first, screen the features that are strongly correlated with abnormal behavior risks, such as the number of shared devices and the proportion of short-distance orders, perform WOE binning on the basic features, analyze the cross-features through SHAP interaction values, and then adapt the basic features, cross-features, composite features (such as abnormal behavior index) and embedded features (relationship vectors generated by GraphSAGE) to the model input requirements. The empirical model directly intercepts based on hard rules, the traditional model uses XGBoost to process structured features, and the neural network model uses GNN to analyze the device-order relationship graph. Finally, the fusion model uses Stacking or weighted voting methods to integrate the outputs of each sub-model, and realizes dynamic weight adjustment through Bayesian optimization to output the abnormal behavior risk prediction results.
[0059] Specifically, the safety warning model workflow is as follows: relevant features such as sudden braking frequency and fatigue driving duration are screened. The empirical model uses rule-based interception. The traditional model uses LightGBM to process time-series driving behavior features. The neural network model combines LSTM and CNN to process driving behavior and in-vehicle video data. The fusion model dynamically assigns weights based on Bayesian optimization and outputs safety risk prediction results.
[0060] Specifically, the service evaluation model workflow is as follows: Relevant features such as negative review rate and passenger complaint text are screened. The empirical model uses rule filtering. The traditional model processes passenger reviews and text sentiment features through logistic regression. The neural network model uses BERT to analyze passenger complaint text. The fusion model uses weighted voting to fuse multi-source scores and output service risk prediction results.
[0061] Specifically, the compliance management model workflow is as follows: screening relevant features such as driver qualification validity period and regulatory data; the empirical model uses rule verification; the traditional model uses random forest to process regulatory and regional data; the neural network model uses Rule2Vec to encode regulatory terms into vectors and match business behaviors through cosine similarity; the fusion model uses compliance reasoning based on knowledge graphs and multi-model fusion to output compliance risk prediction results.
[0062] Step S104: weighted fusion of the risk prediction results is performed according to the dynamic fusion parameters to generate a target risk score for the current online ride-hailing vehicle.
[0063] In an embodiment of the present application, the risk prediction results output by each pre-trained scenario model are weighted and fused through dynamic fusion parameters (including initial fusion weights and real-time weights updated based on performance indicators) to generate a target risk score for the current online car-hailing service. Specifically, the initial weights are calculated based on the prediction accuracy of each scenario model in the historical risk score data, and the weight distribution is dynamically adjusted through real-time monitoring of the model's performance indicators; finally, based on the adjusted dynamic fusion parameters, the risk prediction results output by each model are weighted and summed to generate a comprehensive target risk score. This step solves the problem of insufficient adaptability of static rules to new risks through a dynamic weight distribution mechanism, ensures the accuracy and real-time nature of risk scoring, and provides a reliable basis for subsequent risk decision-making.
[0064] In the embodiment of the present application, constructing a target feature set based on the original business data includes the following steps A1-A3:
[0065] Step A1: bin the original business data based on preset business indicators to obtain basic features.
[0066] Specifically, the continuous features in the original business data are converted into WOE bins by pre-set business indicators (such as driver's driving experience and order distance) to generate basic features. The basic features may include: discretizing continuous variables (such as driver's driving experience) into bin intervals with clear business meanings (such as 0-1 years, 1-5 years, and more than 5 years) based on business knowledge or statistical distribution (such as equal frequency and equal width bins), and calculating the weight of evidence (WOE value) of each bin. The formula is:
[0067]
[0068] Among them, WOE is the weight of evidence of the bin, the proportion of good samples is the proportion of the number of samples that meet the first specific standard (such as timely performance, no violations, etc.) in the bin to the total number of samples in the bin, and the proportion of bad samples is the proportion of the number of samples that meet the second specific standard (such as breach of contract, violation, etc.) in the bin to the total number of samples in the bin.
[0069] For example, after binning drivers by age, the WOE value for the 0-1 year bin is -0.8, indicating that the risk of drivers in this range is significantly higher than the average. By converting continuous features into discrete, interpretable bins, we can eliminate the impact of data noise, enhance feature stability, meet regulatory requirements for model transparency, and provide basic feature input that aligns with business logic for subsequent risk modeling.
[0070] Step A2: Identify the association between basic features and obtain cross features.
[0071] Specifically, through correlation analysis (such as Pearson correlation coefficient, mutual information or business rules), the interaction between basic features is identified to generate cross-features. Based on statistical methods or business logic (such as the combination rules of "night orders" and "high-risk areas"), the potential correlation between basic features (such as driver age classification and order distance) is explored to construct cross-features with business significance (such as "the proportion of orders in high-risk areas at night"), and verify their risk discrimination (for example, through IV value or SHAP interaction value evaluation). By capturing the synergistic effects of multi-dimensional features (such as the superimposed risks of equipment anomalies and high-risk time periods), the model's ability to identify complex risk patterns (such as GPS simulated abnormal behavior) is enhanced. At the same time, through business-interpretable combination logic (such as device sharing index = number of drivers logged in to the same device × number of active days of the device), it is ensured that the cross-features meet the actual risk control needs and improve the model's scenario adaptability and decision-making transparency.
[0072] Step A3: Combine basic features according to scene type to obtain composite features.
[0073] Specifically, according to the scenario type (such as abnormal behavior detection, safety warning), the first feature (such as the proportion of short-distance orders) and the second feature (such as the equipment sharing index) are selected from the basic features, and the first index (short-distance order risk value) and the second index (equipment sharing risk value) are calculated respectively. The two are combined into a composite feature through a preset weight formula (such as abnormal behavior index = α×short-distance order proportion + β×equipment sharing index, where α=0.4 and β=0.3 are determined by Bayesian optimization of historical data and manual verification).
[0074] The method provided in the embodiments of the present application quantifies scenario-based risk signals through a combination of features driven by business logic (for example, the superimposed effect of short-distance high-frequency orders and device sharing behaviors in abnormal behavior scenarios), thereby improving the model's recognition accuracy of risks specific to online ride-hailing (such as false trips), and at the same time ensuring the interpretability of composite features through weight verification, thereby meeting regulatory requirements for tracing the causes of risks.
[0075] Step A4: Process the basic features and / or cross features using an embedding algorithm to obtain embedded features.
[0076] Specifically, by identifying the type of original business data (such as relational data, time series data or text data), the corresponding embedding algorithm is selected for processing: GraphSAGE is used to generate node embedding vectors for graph structure data such as driver-device relationship diagrams to capture abnormal device sharing communities; LSTM is used to extract time series embedding features for time series data such as driving behavior sequences to identify abnormal patterns such as sudden acceleration and sudden braking; BERT is used to generate semantic embedding for text data (such as complaint records).
[0077] The method provided in the embodiments of the present application solves the problem that structured features are difficult to represent complex relationships and nonlinear patterns by converting basic features and cross-features into low-dimensional dense vectors (embedded features). For example, graph embedding can be used to discover potential abnormal behaviors of multiple drivers associated with the same device, thereby enhancing the model's ability to capture hidden risks (such as GPS simulated abnormal behavior). At the same time, the visual output of embedded features (such as community structure diagrams) provides an intuitive basis for explaining the causes of risks.
[0078] In an embodiment of the present application, basic features are combined according to scene types to obtain composite features, including: extracting the first feature and the second feature from the basic features according to the scene type; calculating the first index based on the first feature, and calculating the second index based on the second feature; and combining the first index and the second index into a composite feature through a preset weight formula.
[0079] Specifically, in the safety warning scenario, the "sudden braking frequency" of the basic features is extracted as the first feature, and the "fatigue driving duration" is extracted as the second feature; the emergency braking risk index is calculated based on the emergency braking frequency, and the emergency braking frequency is compared with the frequency thresholds corresponding to different risk levels in the historical data to perform standardized scoring to obtain an index between 0 and 1, and the fatigue risk index is calculated based on the fatigue driving duration. The corresponding risk index is set according to different intervals of fatigue driving duration, and the fatigue risk index is multiplied by the duration by the coefficient; the emergency braking risk index and the fatigue risk index are combined into a composite feature through the preset weight formula "safety risk index = 0.6 × emergency braking risk index + 0.4 × fatigue risk index" to comprehensively evaluate the safety risk status of the driver during driving.
[0080] Specifically, in the service evaluation scenario, the "bad review rate" of the basic features is extracted as the first feature, and the "passenger complaint response time" is extracted as the second feature; the bad review risk index is calculated based on the "bad review rate", and different levels are divided according to the bad review rate, and corresponding index values are assigned; the response efficiency index is calculated based on the "passenger complaint response time", and the response time is compared with the industry standard time, and reverse scoring is performed to obtain the index value; the bad review risk index and the response efficiency index are combined into a composite feature through the preset weight formula "service quality index = 0.5×bad review risk index + 0.5×response efficiency index", so as to comprehensively measure the driver's service quality level.
[0081] Specifically, in the abnormal behavior detection scenario, the "short-distance order ratio" in the basic features is extracted as the first feature, and the "abnormal device login times" is extracted as the second feature; the short-distance order risk index is calculated based on the "short-distance order ratio", and the short-distance order ratio is compared with the normal range. The excess part is assigned a risk index value in proportion, and the device login risk index is calculated based on the "abnormal device login times". Different risk gradients are set according to the number of logins, and corresponding indexes are assigned; the short-distance order risk index and the device login risk index are combined into a composite feature through the preset weight formula "abnormal behavior risk index = 0.7×short-distance order risk index + 0.3×device login risk index" to accurately identify potential abnormal behaviors.
[0082] Specifically, in the compliance management scenario, the "remaining days of driver qualification validity period" in the basic features is extracted as the first feature, and the "number of violation records" is extracted as the second feature; the qualification timeliness index is calculated based on the "remaining days of driver qualification validity period", and the decreasing index value is set according to the different intervals of the remaining days; the violation risk index is calculated based on the "number of violation records", and the corresponding index is assigned according to the number of violations; the qualification timeliness index is calculated through the preset weight formula "Compliance Risk Index = 0.6×Qualification Timeliness Index + 0.4×Violation Risk Index".
[0083] In an embodiment of the present application, the basic features and / or cross features are processed using an embedding algorithm to obtain embedded features, including: identifying the data type of the original business data; determining the corresponding embedding algorithm according to the data type, and using the embedding algorithm to process the basic features and / or cross features to obtain embedded features.
[0084] Specifically, for raw business data involving order interaction relationships between drivers and passengers, which falls into a graph-structured data type, the graph embedding algorithm GraphSAGE is used. Basic features such as the driver ID and passenger ID, as well as the combined features of order amount and order time period, are constructed into a graph structure, where nodes represent drivers and passengers, and edges represent order interaction relationships. The GraphSAGE algorithm is used to process this graph structure, aggregating feature information from neighboring nodes to generate an embedded feature vector for the driver and passenger order interaction relationship. This is then used to identify anomalous order interaction patterns.
[0085] Specifically, for raw business data involving drivers' driving behavior over a continuous period of time, which is a time series data type, LSTM is selected as the embedding algorithm. Basic features such as vehicle speed, braking frequency, and steering wheel angle are arranged in chronological order. These basic features, or the cross-features formed by their combination, are then input into the LSTM network. By learning from time series data, LSTM captures the long-term dependencies of driving behavior in the temporal dimension and outputs corresponding embedded features, effectively identifying abnormal driving behavior patterns.
[0086] Specifically, for raw business data containing textual information about regulations and descriptions of drivers' business behaviors, which fall into the text data type, BERT can be used as an embedding algorithm. The regulations and basic feature text describing drivers' business behaviors are encoded and fed into the BERT model. BERT generates high-quality text embedding representations through bidirectional encoding of the text, which serve as embedded features for subsequent compliance risk analysis and assessment based on textual semantics.
[0087] In an embodiment of the present application, a method for generating dynamic fusion parameters includes: obtaining the prediction accuracy of each pre-trained scenario model in the historical risk score data; calculating the initial fusion weight of the pre-trained scenario model based on the prediction accuracy, and updating the initial fusion weight according to the performance index to generate dynamic fusion parameters.
[0088] Specifically, first, the initial fusion weights are calculated based on the prediction accuracy (such as AUC and recall rate) of each pre-trained scenario model (such as the abnormal behavior detection model and the security warning model) in the historical risk scoring data (for example, the initial weight of the abnormal behavior model is 0.4, and the initial weight of the security model is 0.3). The principle of initial weight distribution is that the higher the accuracy, the greater the weight ratio; secondly, by real-time monitoring of model performance indicators (such as KS value to measure risk discrimination and PSI to detect feature offset), combined with online Bayesian update or reinforcement learning algorithm to dynamically adjust the weights (for example, when the KS value of the abnormal behavior model decreases, its weight is reduced and the decision ratio of other scenario models is increased). The dynamic fusion parameters finally generated can adapt to business changes (such as the frequent occurrence of new abnormal attacks), ensure the robustness of multi-model fusion and the accuracy of risk scoring, and at the same time meet the regulatory requirements for traceability of model decisions through transparent weight adjustment.
[0089] In this embodiment of the present application, after generating the target risk score for the current online ride-hailing service, the method further includes steps B1-B3:
[0090] Step B1: Obtain the actual risk detection results of the current online ride-hailing service.
[0091] Specifically, the actual risk detection results of the current online ride-hailing are obtained through multiple channels. On the one hand, it connects with the audit records in the business system, such as the manual judgment results of orders and driver behaviors by risk control personnel; on the other hand, it collects user complaint feedback data, such as passengers' complaints about abnormal driver behavior, service quality and other issues; at the same time, it accesses risk assessment data from third-party institutions, such as risk judgments from device fingerprint databases. After integrating these data, comprehensive and accurate actual risk detection results of the current online ride-hailing are obtained.
[0092] Step B2: Analyze the difference data between the risk prediction results output by each pre-trained scenario model and the actual risk detection results.
[0093] Specifically, the risk prediction results from each pre-trained scenario model, such as the abnormal behavior risk score from the abnormal behavior detection model and the security risk level from the security warning model, are compared with the actual risk detection results. Quantitative methods such as mean square error and absolute error are used to calculate the numerical deviation between the predicted and actual results. A confusion matrix is used to analyze the proportion of correct and incorrect predictions. Visualization tools are used to compare predicted and actual risk trends in different scenarios, resulting in a comprehensive and detailed analysis of the discrepancies between the two.
[0094] Step B3: Determine the performance indicators of each pre-trained scenario model based on the difference data, and use the performance indicators to update the model parameters of the pre-trained scenario model and the dynamic fusion parameters of the risk prediction results.
[0095] Specifically, performance metrics for each pre-trained scenario model, such as accuracy, recall, F1 value, and AUC, are calculated based on the difference data to measure the model's prediction accuracy, risk identification capabilities, and other performance indicators. Based on the performance indicators, an online Bayesian update algorithm is used to adjust model parameters and optimize the model's ability to learn features. A reinforcement learning algorithm is used to dynamically update the dynamic fusion parameters of risk prediction results. For example, the weight of the output results of scenario models such as abnormal behavior, security, service, and compliance in the final target risk score is adjusted. This allows the model to better adapt to business changes and improve the accuracy and robustness of overall risk identification.
[0096] In an embodiment of the present application, the method also includes: comparing the target risk score with the interception threshold and the release threshold respectively to obtain a comparison result; if the comparison result is that the target risk score is greater than the interception threshold, triggering the order interception instruction and generating the corresponding review task; or, if the comparison result is that the target risk score is less than the release threshold, triggering the order release instruction and adjusting the risk status of the current online car-hailing vehicle; or, if the comparison result is that the target risk score is less than or equal to the interception threshold and greater than or equal to the release threshold, triggering the enhanced verification instruction and dynamically monitoring the behavior sequence of the current online car-hailing vehicle.
[0097] Specifically, the output target risk score is numerically compared with the pre-set interception threshold and release threshold. When the target risk score is greater than the interception threshold, it indicates that the order has a high risk. The order interception instruction is immediately triggered, the order execution process is suspended, and the corresponding review task is automatically generated. For example, the order information is pushed to the manual review platform for further verification by risk control personnel.
[0098] When the target risk score is lower than the release threshold, it indicates that the order risk is low, triggering the order release instruction, allowing the order to proceed normally, and simultaneously adjusting the risk status of the current online ride-hailing vehicle to "low risk", reducing the frequency of subsequent monitoring;
[0099] When the target risk score is between the interception threshold and the release threshold, that is, less than or equal to the interception threshold and greater than or equal to the release threshold, it indicates that there is a certain degree of uncertainty in the order, triggering an enhanced verification instruction, requiring the driver or passenger to perform additional identity verification (such as facial recognition, SMS verification code, etc.), and at the same time starting a dynamic monitoring mechanism, using LSTM to monitor the current online car-hailing driving behavior sequence, order operation sequence, etc. in real time. Once an abnormal pattern is found, the risk level is immediately increased and further measures are taken. For example, when the target risk score of an order is 85 points, the interception threshold is set to 80 points, and the release threshold is 60 points, since 85 points is greater than 80 points, the system will intercept the order and generate an audit task; if the risk score of another order is only 50 points, which is less than the release threshold of 60 points, the order will be released directly; if the order risk score is 70 points, which is within the threshold range, the system will require the driver to perform a second identity verification and continue to track its subsequent trip data.
[0100] In the embodiment of the present application, the training method of the pre-trained scenario model includes the following core steps, covering the entire process of data processing, feature engineering, model training and fusion optimization:
[0101] First, multi-source data collection and layered processing. For the four major scenarios of abnormal behavior detection, security alerts, service assessment, and compliance management, the data collection layer acquires multi-dimensional raw business data: third-party data (device fingerprints, IP profiles, regulatory blacklists, in-vehicle videos, etc.), internal data (order records: start and end points, distance, time), GPS tracks, payment records, driver qualifications, passenger complaint texts, etc. After ETL cleansing (outlier processing and missing value filling), the data is stored in the raw data layer (ODS). It is then aggregated through distributed computing frameworks such as Spark to generate data at the detail layer (DWD), summary layer (DWS), and application layer (ADS), providing structured input for feature engineering.
[0102] Secondly, multi-dimensional feature engineering and interpretability are enhanced. In the feature engineering module, a target feature set is constructed based on hierarchically stored data, and feature quality is improved through interpretability techniques: Basic features: Continuous variables are binned using Word of Equality (WOE) and the weight of evidence for each bin is calculated, converting the original indicators into discrete, interpretable binned features; Cross features: Identify associations between features using the Pearson correlation coefficient or business rules, construct composite signals, and verify feature synergy using SHAP interaction values; Composite features: Combine basic features by scenario type and generate a scenario-based risk index using a preset weight formula. The weights are determined through Bayesian optimization and manual verification; Embedded features: GraphSAGE is used to generate relationship vectors for graph-structured data, LSTM is used to extract time series patterns for time series data, and BERT is used to generate semantic embeddings for text data, addressing the problem of expressing unstructured data features.
[0103] Secondly, scenario-based model training and sub-model construction. In the expert model module, three types of sub-models (empirical model, traditional model, neural network model) are trained for four major scenarios:
[0104] Abnormal behavior detection model: empirical model, based on hard rule interception; traditional model, XGBoost processes structured features and outputs feature importance; neural network model, GNN analyzes the device-order relationship graph and identifies abnormal communities;
[0105] Safety warning model: empirical model, rule-based interception; traditional model: LightGBM processes time-series driving behavior features; neural network model, LSTM+CNN integrates driving behavior data and in-vehicle video;
[0106] Service evaluation model: empirical model, rule filtering; traditional model, logistic regression analysis of passenger evaluation sentiment; neural network model: BERT extracts semantic features of complaint text;
[0107] Compliance management model: empirical model, rule verification; traditional model: random forest processes regulatory data and regional characteristics; neural network model: Rule2Vec encodes regulatory terms into vectors to match business behaviors.
[0108] Finally, multi-model fusion and dynamic weight optimization. A fusion model is constructed through stacking or weighted voting, using the true risk labels of historical data as supervisory signals. Sub-model weights are optimized based on performance indicators such as AUC and KS. Initial weights are assigned based on the sub-model's prediction accuracy in historical data (such as XGBoost's AUC value) (for example, in an abnormal behavior detection scenario, GNN weight accounts for 50%, XGBoost for 30%, and the empirical model for 20%). Through Bayesian optimization or reinforcement learning algorithms, model indicators are monitored in real time (such as PSI to detect feature offsets and KS value to assess distinguishing ability), and the weights of each sub-model are automatically adjusted to ensure the adaptability of the fusion model to new risks.
[0109] It should be noted that during the training process, WOE binning, SHAP / LIME and causal inference techniques are used to provide global and local explanations for the model: SHAP values are generated for key features (such as the device sharing index) to quantify their contribution to risk prediction; counterfactual analysis reports are generated through causal inference to meet regulatory requirements for model transparency.
[0110] Figure 2 Schematic diagram of the risk control system architecture for online car-hailing according to an embodiment of the present invention. Figure 2 As shown in the figure, the architecture includes: business layer, decision layer, expert model, feature engineering, data processing layer, and data collection layer;
[0111] The data collection layer is responsible for collecting multi-source data such as orders, drivers, itineraries, traffic, weather, and third-party risk control. The data processing layer processes the data through ODS, DWD, DWS, ADS, and other links. The feature engineering module generates basic, cross-cutting, composite, and embedding features based on the processed data. The expert model constructs empirical models, machine learning models, CNN models, and fusion models for risk analysis in four major scenarios: abnormal behavior detection, safety warnings, service evaluation, and compliance management. The decision-making layer integrates the output of the expert model through dynamic weight fusion and decision matching, and outputs it to the business layer. In addition, the system uses model monitoring (such as AUC, KS, and PSI indicators), online learning (Bayesian optimization and reinforcement learning), and automatic retraining to achieve adaptive learning to continuously improve risk control efficiency.
[0112] Figure 3 Schematic diagram of feature engineering in the online car-hailing risk control system according to an embodiment of the present invention. Figure 3As shown in the figure, the raw business data includes data on orders, drivers, itineraries, traffic, weather, and more. After data preprocessing (covering operations such as data cleaning, outlier processing, standardization, binning, and encoding), it enters the feature construction phase to generate basic features, statistics, combined features, vector features, time series features, and high-order features. Next, through the feature selection step, features are selected using methods such as feature filtering, tree model features, feature dimensionality reduction, feature evaluation, and regularization. The features are then stored as offline and real-time features and versioned. Throughout the entire process, interpretability techniques such as WOE binning, SHAP / LIME, and causal inference are used to improve feature interpretability.
[0113] Figure 4 : is a schematic diagram of the model fusion training process in the online car-hailing risk control system according to an embodiment of the present invention. Figure 4 As shown in the figure, the bottom layer is the feature storage, which contains various types of features such as basic features and statistics. It serves as the fundamental data source for model training. The next layer is the sub-model training layer. This layer uses a variety of models for training in different scenarios, including neural network models such as LSTM+CNN for processing time series and image data, CNN for image processing, GNN for analyzing graph structures, and BERT for text processing, as well as traditional models such as empirical rules, logistic regression, random forest, and XGBoost. The training process also includes model evaluation and parameter tuning, with either serial or parallel approaches. The next layer is the scenario model layer, covering four major scenarios: anomaly model detection, security alerts, service assessment, and compliance management. Each scenario has a corresponding sub-fusion model that further integrates the sub-model training results. The top layer is the fusion model layer. Through mechanisms such as predicted probability, dynamic weight allocation, and online feedback, using stacking or weighted / averaging methods, the outputs of the expert models are integrated to produce a comprehensive target risk score. Dynamic feedback is continuously used throughout the process to continuously optimize the model.
[0114] In this embodiment, a device for assessing the risk of online car-hailing is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments, and the details that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0115] This embodiment provides a device for assessing the risk of online car-hailing. Figure 5 Shown, including:
[0116] The acquisition module 51 is used to obtain the original business data associated with the current online car-hailing service;
[0117] A construction module 52 is used to construct a target feature set based on the original business data, wherein the target feature set includes: basic features, cross features, composite features and embedded features;
[0118] An input module 53 is used to input each feature in the target feature set into multiple pre-trained scenario models according to the scenario type, and obtain the risk prediction results output by each pre-trained scenario model;
[0119] The fusion module 54 is used to perform weighted fusion on the risk prediction results according to the dynamic fusion parameters to generate a target risk score for the current online car-hailing service.
[0120] Furthermore, a construction module 52 is used to perform binning conversion on the original business data based on preset business indicators to obtain basic features; identify the correlation between each basic feature to obtain cross-features; combine the basic features according to the scenario type to obtain composite features; and use the embedding algorithm to process the basic features and / or cross-features to obtain embedded features.
[0121] Furthermore, the construction module 52 further includes: a first processing submodule and a second processing submodule;
[0122] The first processing submodule is used to extract the first feature and the second feature from the basic features according to the scene type; calculate the first index based on the first feature, and calculate the second index based on the second feature; and combine the first index and the second index into a composite feature through a preset weight formula.
[0123] The second processing submodule is used to identify the data type of the original business data; determine the corresponding embedding algorithm according to the data type, and use the embedding algorithm to process the basic features and / or cross features to obtain embedded features.
[0124] Furthermore, the device also includes: a generation module for obtaining the prediction accuracy of each pre-trained scenario model in the historical risk score data; calculating the initial fusion weight of the pre-trained scenario model based on the prediction accuracy, and updating the initial fusion weight according to the performance index to generate dynamic fusion parameters.
[0125] Furthermore, the device also includes: an update module for obtaining the actual risk detection results of the current online car-hailing vehicle; analyzing the difference data between the risk prediction results output by each pre-trained scenario model and the actual risk detection results; determining the performance indicators of each pre-trained scenario model based on the difference data, and using the performance indicators to update the model parameters of the pre-trained scenario model and the dynamic fusion parameters of the risk prediction results.
[0126] Furthermore, the device also includes: a trigger module, which is used to compare the target risk score with the interception threshold and the release threshold respectively to obtain a comparison result; if the comparison result is that the target risk score is greater than the interception threshold, the order interception instruction is triggered and the corresponding review task is generated; or, if the comparison result is that the target risk score is less than the release threshold, the order release instruction is triggered and the risk status of the current online car-hailing vehicle is adjusted; or, if the comparison result is that the target risk score is less than or equal to the interception threshold and greater than or equal to the release threshold, the enhanced verification instruction is triggered and the behavior sequence of the current online car-hailing vehicle is dynamically monitored.
[0127] See also Figure 6 , Figure 6 is a structural diagram of a computer device provided by an optional embodiment of the present invention, such as Figure 6 As shown, the computer device includes: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system).
[0128] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0129] The memory 20 stores instructions that can be executed by at least one processor 10, so as to enable at least one processor 10 to execute the method shown in the above embodiment.
[0130] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created based on the use of a computer device for displaying a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0131] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0132] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0133] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0134] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for assessing the risk of online ride-hailing, characterized in that: The method comprises: Get the original business data associated with the current online car-hailing service; Constructing a target feature set based on the original business data, wherein the target feature set includes: basic features, cross features, composite features and embedded features; Inputting each feature in the target feature set into multiple pre-trained scenario models according to the scenario type to obtain the risk prediction results output by each pre-trained scenario model; The risk prediction results are weighted and fused according to dynamic fusion parameters to generate a target risk score for the current online car-hailing service.
2. The method according to claim 1, characterized in that The constructing of a target feature set based on the original business data includes: Perform binning conversion on the original business data based on preset business indicators to obtain basic features; Identify the correlation between each of the basic features to obtain cross-features; Combining the basic features according to the scene type to obtain a composite feature; The basic features and / or the cross features are processed using an embedding algorithm to obtain embedded features.
3. The method according to claim 2, characterized in that The basic features are combined according to the scene type to obtain composite features, including: Extracting a first feature and a second feature from the basic features according to the scene type; calculating a first index based on the first feature, and calculating a second index based on the second feature; The first index and the second index are combined into a composite feature through a preset weight formula.
4. The method according to claim 2, characterized in that The processing of the basic features and / or the cross features by using an embedding algorithm to obtain embedded features includes: Identifying the data type of the original business data; A corresponding embedding algorithm is determined according to the data type, and the basic features and / or the cross features are processed using the embedding algorithm to obtain embedded features.
5. The method according to claim 1, wherein The method for generating the dynamic fusion parameters includes: Obtaining the prediction accuracy of each of the pre-trained scenario models in the historical risk score data; The initial fusion weight of the pre-trained scene model is calculated based on the prediction accuracy, and the initial fusion weight is updated according to the performance index to generate the dynamic fusion parameter.
6. The method according to claim 1, characterized in that After generating the target risk score of the current online ride-hailing service, the method further includes: Obtaining the actual risk detection result of the current online ride-hailing service; Analyzing the difference data between the risk prediction results output by each of the pre-trained scenario models and the actual risk detection results; The performance indicators of each of the pre-trained scenario models are determined based on the difference data, and the model parameters of the pre-trained scenario models and the dynamic fusion parameters of the risk prediction results are updated using the performance indicators.
7. The method according to claim 1, characterized in that The method further comprises: Comparing the target risk score with the interception threshold and the release threshold respectively to obtain a comparison result; If the comparison result is that the target risk score is greater than the interception threshold, the order interception instruction is triggered and the corresponding review task is generated; or, if the comparison result is that the target risk score is less than the release threshold, the order release instruction is triggered and the risk status of the current online car-hailing vehicle is adjusted; or, if the comparison result is that the target risk score is less than or equal to the interception threshold and greater than or equal to the release threshold, the enhanced verification instruction is triggered and the behavior sequence of the current online car-hailing vehicle is dynamically monitored.
8. A device for assessing the risk of online car-hailing, characterized in that: The device comprises: The acquisition module is used to obtain the original business data associated with the current online car-hailing service; A construction module, configured to construct a target feature set based on the original business data, wherein the target feature set includes: basic features, cross features, composite features, and embedded features; An input module, configured to input each feature in the target feature set into a plurality of pre-trained scenario models according to the scenario type, and obtain a risk prediction result output by each pre-trained scenario model; A fusion module is used to perform weighted fusion on the risk prediction results according to dynamic fusion parameters to generate a target risk score for the current online car-hailing service.
9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
Index calculation method and device based on model fusion and medium
CN121326896A