A Method for Constructing Risk Profiles of Bus Drivers Based on Multi-Source Data Fusion

By integrating multi-source data to construct a risk profile of bus drivers, integrating multi-dimensional data, establishing a hierarchical indicator system and dynamically adjusting weights, the problems of insufficient data utilization and lack of targeted control in traditional assessment techniques are solved, thus achieving accurate assessment and effective management of bus driver risks.

CN122134194APending Publication Date: 2026-06-02YANCHENG INST OF TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YANCHENG INST OF TECH
Filing Date
2026-03-02
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Traditional driver risk assessment technologies suffer from insufficient data utilization, simplistic assessment logic, lack of systematic approach, non-standard data processing, and a lack of targeted control strategies, leading to one-sided risk identification and difficulties in optimizing management.

Method used

By constructing a risk profiling method for bus drivers that integrates multi-source data, we integrate data from onboard safety warnings, video surveillance, pre-job health, and vehicle operation to establish a hierarchical indicator system, dynamically adjust indicator weights, use a multi-dimensional risk coupling algorithm for precise quantitative assessment, and match differentiated control strategies to build a closed-loop management system.

Benefits of technology

It has enabled accurate identification and targeted control of driving safety risks, improved the comprehensiveness and timeliness of risk identification, and ensured the continuous optimization of public transportation operation safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134194A_ABST
    Figure CN122134194A_ABST
Patent Text Reader

Abstract

This invention discloses a method for constructing a risk profile of bus drivers based on multi-source data fusion, belonging to the field of public transportation safety management technology. The specific steps of this method are as follows: first, multi-source data collection is conducted to construct a system for collecting multiple types of data; then, data preprocessing and fusion are performed to construct a structured dataset through a series of processes; next, risk indicators are constructed, forming quantitative assessment dimensions according to principles; then, dynamic scoring modeling is used to calculate the driver's dynamic risk score; finally, risk control output is performed, determining the risk level and matching strategies, outputting results through multiple channels, and establishing a closed-loop control system. This invention constructs a multi-dimensional data source system, which, after processing, forms a structured risk feature dataset, achieving accurate quantitative risk assessment and breaking through the limitations of single methods; by dynamically adjusting indicator weights in conjunction with operational variables, risk levels are accurately classified and control strategies are matched, establishing a closed-loop control system to ensure that measures are practical, improving safety management efficiency, and guaranteeing the safety of public transportation operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of public transportation safety management technology, specifically a method for constructing risk profiles of public transportation drivers based on multi-source data fusion. Background Technology

[0002] Public transportation, as a core component of urban public transport, is directly related to the safety of passengers' lives and property and the orderly operation of urban traffic. Drivers, as key players in the operation process, have a decisive impact on operational safety due to their driving behavior, physical condition, and professional competence. With the improvement of the intelligence level of the transportation industry, various tools such as on-board monitoring equipment, video surveillance systems, and enterprise management platforms are widely used, which can collect various behavioral data, vehicle operation data, health monitoring data, and basic information data of drivers during the driving process. This multi-source data provides rich information support for driver risk assessment. At the same time, the reality of increased urban traffic flow, more complex road conditions, and diversified operating environments has placed higher demands on the comprehensiveness, timeliness, and accuracy of public transport driver risk identification. The industry urgently needs a systematic method that can integrate multi-dimensional data and scientifically assess risks to achieve effective prediction and control of driving safety risks.

[0003] Traditional driver risk assessment technologies often suffer from insufficient data utilization and simplistic assessment logic. Some assessment methods rely solely on a single data source, neglecting the correlation between various factors such as driving behavior, vehicle status, health status, and historical records, leading to one-sided risk identification. The assessment indicator system lacks systematic design, with unclear indicator divisions and fixed weightings, failing to dynamically adjust according to different operating scenarios, time periods, and individual driver characteristics, thus failing to reflect real-time risk changes. During data processing, the handling of missing and outlier values ​​is not standardized, and data fusion lacks scientific methods, affecting the accuracy of assessment results. Furthermore, traditional control strategies often adopt uniform standards, lacking specificity and failing to develop differentiated measures based on different driver risk levels. They also lack a complete closed loop from early warning to rectification and assessment, hindering continuous optimization of safety management and failing to effectively address the complex and ever-changing challenges of public transportation operation safety. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a method for constructing risk profiles of bus drivers based on multi-source data fusion. By integrating multi-dimensional data from vehicle safety warnings, video surveillance, pre-job health, vehicle operation, and enterprise management systems, and through standardized cleaning, standardization, feature engineering, and fusion processing, a structured dataset is formed. A hierarchical indicator system is built based on scientific principles, and the indicator weights are dynamically adjusted to adapt to operational scenarios. A multi-dimensional risk coupling algorithm is used to achieve accurate quantitative scoring, classify multiple risk levels, and match differentiated control strategies. A closed-loop system of early warning intervention, rectification tracking, effect evaluation, and continuous optimization is constructed. This method comprehensively covers key dimensions of driving safety, improves the pertinence and timeliness of risk identification and control, promotes the transformation of bus safety management towards intelligence and refinement, and provides strong protection for operational safety.

[0005] To address the aforementioned technical problems, this invention provides the following technical solution: a method for constructing a risk profile of bus drivers based on multi-source data fusion, the specific steps of which are as follows: S1, Multi-source data acquisition: Construct a multi-dimensional data source system to collect safety warning data from the vehicle safety warning platform, dangerous behavior data from the video surveillance platform, pre-job health data from offline equipment, vehicle operation data from vehicle equipment, and basic static data from the public transport enterprise management system. Real-time access and storage of the data are completed through a unified interface. S2, Data Preprocessing and Fusion: Extract the data from step S1, which is collected and stored from multiple sources, and process it according to the process of data cleaning, standardization, feature engineering, and fusion association to construct a structured risk feature dataset; S3, Risk Indicator Construction: Based on the structured risk feature dataset, and in accordance with the principles of MECE, measurability, importance and relevance, and operability, a hierarchical indicator system is constructed, indicator thresholds are determined, indicator weights are set, and coupling coefficients between indicators are calculated to form a quantitative assessment dimension for risk profiling. S4, Dynamic scoring modeling: Based on the risk indicator system built in step S3, the indicator weights are dynamically adjusted, and a multi-dimensional risk coupling scoring algorithm is used to calculate the driver's dynamic risk score, which serves as the core quantitative indicator for risk profiling. S5, Risk Management Output: Receive the score output from step S4, Dynamic Scoring Modeling, use the risk level dynamic correction algorithm to determine the risk level from level one to level five, match differentiated management strategies, output the risk profile results through multiple channels, and establish a closed-loop management system of early warning intervention - rectification tracking - effect evaluation - continuous optimization.

[0006] Furthermore, in S1, the multi-source data acquisition includes: vehicle safety warning platform data, including driver behavior warnings, vehicle forward risk warnings, and vehicle speed warnings, specifically covering data on making and receiving phone calls, obstructing cameras, distracted driving, fatigued driving, not wearing seat belts, forward collisions, lane departures, pedestrian collisions, frequent lane changes, speeding, abnormally low-speed driving, and violations of road speed limits; and video surveillance platform data on dangerous behaviors, including drivers yawning, not looking ahead, excessively fast when entering stations, chatting, and sudden deceleration. Data on drivers not slowing down at zebra crossings; offline pre-shift health data including heart rate, blood pressure, body temperature, and alcohol content test data before drivers start work; vehicle operation data from onboard equipment including vehicle speed, frequency of rapid acceleration, frequency of rapid deceleration, steering angle, engine speed, braking frequency, mileage, and road segment type data during vehicle operation; basic static data from the public transport company management system including driver employee number, name, mobile phone number, department, driving route, driving experience, age, historical accident records, historical violation records, safety training assessment results, monthly comprehensive score, and ranking changes.

[0007] Furthermore, in S2, during data preprocessing and fusion, the data cleaning stage supplements missing physiological health data by using median interpolation of data from drivers with the same age and department during the same period; missing vehicle operation data is filled using linear interpolation; missing basic static data is supplemented by manually verifying employee files; outliers are identified using the box plot IQR method; extreme outliers are replaced with the median; and general outliers are replaced with the modified median, reducing the data missing rate to below 3%.

[0008] Furthermore, in S2, during data preprocessing and fusion, the mathematical expression for the multi-source data confidence fusion algorithm is: ;in, For the first The core feature fusion value of each driver The total number of data sources. For the first Class data source for the first The basic weight of each driver For the first The real-time reliability coefficient of the data source. For the first The driver in The raw indicator values ​​in the data source are , For the first The global minimum and maximum values ​​of metrics for similar data sources. For the first Coefficient of variation of data source This is the fluctuation penalty coefficient.

[0009] Furthermore, in S3, the hierarchical indicator system constructed in the risk indicator construction includes: the MECE principle, meaning that each level of indicator is clearly defined and has a clear scope, with no overlap between them, and can comprehensively cover all key dimensions related to the risk profile of bus drivers, without omitting any core risk factors; the measurability principle, meaning that all indicators have clear data sources and can be quantitatively represented through specific values, frequency of occurrence, or classification labels, without vague qualitative descriptions; and the importance and relevance principle, meaning that the selected indicators are all directly related to the driving safety and accident probability of bus drivers, and that each indicator is relevant to driving safety. The degree of impact varies, so differentiated selection and settings are made accordingly. The principle of operability, namely the division of indicators, the determination of thresholds, and the setting of weights, are all combined with the actual scenarios of safety management in public transportation companies. The results of indicator analysis can be directly applied to the construction of risk profiles for bus drivers and related safety management work, which is convenient for implementation. The hierarchical indicator system is divided into primary indicators, secondary indicators, and tertiary indicators. Primary indicators include basic labels, pre-job health indicators, safety early warning indicators, behavioral assessment indicators, and dangerous behavior indicators. Secondary indicators are further subdivided based on primary indicators. Tertiary indicators further clarify the specific quantitative direction and classification standards corresponding to secondary indicators.

[0010] Furthermore, in S3, the specific content of determining indicator thresholds, setting indicator weights, and calculating coupling coefficients between indicators in the risk indicator construction is as follows: When determining indicator thresholds, continuous indicators are set with reference to industry standards and medical norms; blood pressure indicators are set with reference to medical diagnostic standards, with systolic blood pressure <120 mmHg and diastolic blood pressure <80 mmHg considered normal; systolic blood pressure 120-139 mmHg and / or diastolic blood pressure 80-89 mmHg considered high normal; and systolic blood pressure ≥140 mmHg and / or diastolic blood pressure ≥90 mmHg considered hypertension. Frequency-based behavioral indicators are set using the quartile method, with yawning behavior divided into three levels based on a daily average of 1 or 3 times, and fatigue driving divided into three levels based on a monthly average of 3 or 7 times. Complex risk indicators use a decision tree algorithm to select the optimal split node as the threshold. When setting indicator weights, the basic weights of the primary indicators are first determined using the Analytic Hierarchy Process (AHP): basic label weight 0.1, pre-employment health indicator weight 0.2, safety early warning indicator weight 0.4, behavioral assessment indicator weight 0.15, and hazardous behavior indicator weight 0.15. Then, the weights are adjusted using the entropy weight method combined with the actual data distribution over the past six months, with the adjustment range not exceeding ±20% of the basic weights. The final weights are determined after a consistency test (CR) of less than 0.1. When calculating the coupling coefficient between indicators, the characteristic values ​​of any two indicators are first discretized into 3 to 5 intervals. Then, the joint entropy and individual marginal entropy of the two indicators are calculated, and the coupling coefficient is obtained using the coupling coefficient formula. The coupling coefficient reflects the amplification effect of simultaneous anomalies in two indicators on the risk profile. The formula used to calculate the coupling coefficient between indicators is: ;in, Let be the coupling coefficient between the j-th indicator and the 1st indicator. Let the marginal entropy of the j-th index be , The marginal entropy of the first indicator, Let the joint entropy of the j-th index and the 1st index be... It is a local minimum constant.

[0011] Furthermore, in S4, the dynamic weight adjustment in the dynamic scoring modeling specifically involves increasing the weight of the safety warning indicator from 0.4 to 0.45 and the weight of the road environment-related indicator from 0.1 to 0.15 during peak hours (7:00-9:00 and 17:00-19:00); increasing the weight of the vehicle operation risk indicator from 0.2 to 0.25 during severe weather; increasing the weight of the historical safety risk indicator from 0.15 to 0.2 for drivers with less than 3 years of driving experience; and temporarily increasing the weight of the corresponding warning indicator by 50% for 48 hours when a driver has received one or more severe warnings within the past 24 hours.

[0012] Furthermore, in S4, the mathematical expression for the multidimensional risk-coupled scoring algorithm in dynamic scoring modeling is: ;in, For the first Dynamic risk score for each driver The total number of core risk indicators For the first The driver The fusion characteristic value of the indicators For the first The dynamically adjusted weights of the indicators For the first Item and the The coupling coefficient of the item index, These are the linear main effect weights.

[0013] Furthermore, in S5, the mathematical expression for the dynamic risk level correction algorithm in the risk control output is: ;in, Let i be the final risk level of the i-th driver. This is the floor function. Let i be the original risk score for the i-th driver. Here, T represents the rectification effect coefficient, and T represents the historical rectification tracking period. Let be the rectification effect value for the i-th driver in week t. The time decay coefficient, This represents the sum of rectification effect values ​​within the i-th driver's historical rectification tracking period. To obtain the maximum value between the sum of rectification effect values ​​and the baseline value v within the i-th driver's historical rectification tracking period, To correct for constants; dynamic risk scores of 0 to 20 correspond to Level 1, 21 to 40 to Level 2, 41 to 60 to Level 3, 61 to 80 to Level 4, and 81 to 100 to Level 5.

[0014] Furthermore, in the S5 risk management output, the differentiated management strategy specifically includes: Level 1 risk drivers generating a risk profile report once per quarter; Level 2 risk drivers receiving safe driving guidance materials once per month and participating in online safety learning once every six months; Level 3 risk drivers receiving specialized safety training, with each training session lasting no less than 4 hours, and having their operational schedules adjusted; Level 4 risk drivers having their operational qualifications suspended for no less than 3 days, and being required to participate in offline safety training and psychological counseling; and Level 5 risk drivers having their on-duty qualifications immediately suspended, being reported to the industry regulatory authority, and having their driving qualifications re-examined. The multi-channel output includes: displaying the driver's personal risk profile core indicators and targeted improvement suggestions on the driver's terminal.

[0015] Beneficial effects Compared with existing technologies, this method for constructing risk profiles of bus drivers based on multi-source data fusion has the following advantages: I. This invention constructs a multi-dimensional data source system, integrating various safety-related data during the driving process. Through systematic data cleaning, standardization, feature engineering, and fusion correlation processing, a structured risk feature dataset is formed. Based on scientific principles, a hierarchical indicator system is built, fully considering data correlation and practical application scenarios. The indicator threshold setting, weight allocation, and calculation of coupling relationships between indicators are optimized. A multi-dimensional risk coupling scoring algorithm is used to achieve accurate quantitative assessment of risk. This approach breaks the limitations of single-data assessment, deeply explores the potential value between different types of data, and allows risk assessment to cover the key dimensions of the entire driving process. It effectively avoids the bias caused by one-sided data and simplistic logic in traditional assessments, improves the comprehensiveness and accuracy of risk identification, and provides public transportation companies with detailed and reliable data on driver risk status, helping companies control driving safety hazards from the source.

[0016] Second, this invention dynamically adjusts indicator weights by combining variables such as time of day, weather, and driving experience in the operational scenario. It uses a dynamic risk level correction algorithm to accurately classify risk levels and matches differentiated control strategies to different levels. At the same time, it outputs risk profile results and improvement suggestions through multiple channels, establishing a closed-loop control system for the entire process. This achieves a complete management link from risk warning, intervention and rectification to effect evaluation and continuous optimization. This dynamic adaptation and closed-loop operation mode can respond to various changes in the driving process in real time, ensuring that control measures are in line with actual operational needs, specifically addressing the safety issues of drivers at different risk levels, reducing ineffective management costs, improving the pertinence and efficiency of safety management, and helping drivers clarify their own improvement direction, gradually reducing driving risks, and providing comprehensive and sustainable protection for public transportation operation safety.

[0017] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0019] Figure 1 A flowchart illustrating a method for constructing risk profiles of bus drivers based on multi-source data fusion; Figure 2 This diagram illustrates the data transmission between the steps in the method for constructing a risk profile of a bus driver based on multi-source data fusion. Detailed Implementation

[0020] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0021] Example 1: Risk profiles of bus drivers on major urban arterial roads during weekday morning rush hour.

[0022] S1, Multi-source data acquisition: Construct a multi-dimensional data source system to comprehensively collect safety warning data from the vehicle safety warning platform, dangerous behavior data from the video surveillance platform, offline pre-job health data, vehicle operation data from vehicle equipment, and basic static data from the public transport company management system. The vehicle-mounted safety warning platform's safety warning data covers behavioral warnings such as making phone calls, obstructing cameras, distracted driving, fatigued driving, and not wearing seat belts; forward risk warnings such as forward collisions, lane departures, pedestrian collisions, and frequent lane changes; and speed warnings such as speeding, abnormally low-speed driving, and violations of road speed limits. The video surveillance platform's dangerous behavior data includes data on drivers yawning, not looking ahead, excessive speed when entering stations, chatting, sudden deceleration, and failure to slow down at crosswalks. Offline pre-shift health data collection includes heart rate, blood pressure, body temperature, and alcohol content tests before drivers start work. Vehicle operation data from onboard equipment includes vehicle speed, frequency of rapid acceleration and deceleration, steering angle, engine speed, braking frequency, mileage, and the type of main road in the city. The public transport company management system's basic static data includes driver employee number, name, mobile phone number, department, driving route, driving experience, age, historical accident records, historical violation records, safety training assessment results, monthly comprehensive score, and ranking changes. All data is accessed and stored in real time through a unified interface, comprehensively covering driver behavior, vehicle operating status, basic physical conditions, and key information related to enterprise management. This provides complete and comprehensive data support for subsequent risk profiling. Figure 1 As shown.

[0023] S2, Data Preprocessing and Fusion: This involves extracting stored multi-source data and processing it step-by-step according to the data cleaning, standardization, feature engineering, and fusion correlation processes. During data cleaning, missing physiological health data is supplemented using median interpolation from data of drivers with the same experience and driving experience within the same department. Missing vehicle operation data is filled using linear interpolation. Missing basic static data is supplemented by manual verification of employee records. Box plot IQR is used to identify outliers; extreme outliers are replaced with the median, while general outliers are replaced with a modified median. After processing, the data missing rate is reduced to below 3%, effectively avoiding interference from missing and outlier data on subsequent analysis results and ensuring data quality. Subsequently, data standardization is performed to eliminate dimensional differences between different data sources. Then, feature engineering is used to extract key information. Finally, a multi-source data confidence fusion algorithm is applied to deeply integrate the effective information from various data sources. The mathematical expression of the multi-source data confidence fusion algorithm is: ;in, For the first The core feature fusion value of each driver The total number of data sources. For the first Class data source for the first The basic weight of each driver For the first The real-time reliability coefficient of the data source. For the first The driver in The raw indicator values ​​in the data source are , For the first The global minimum and maximum values ​​of metrics for similar data sources. For the first Coefficient of variation of data source To establish a volatility penalty coefficient, a structured risk feature dataset is constructed, transforming scattered data into usable features with clear risk orientations.

[0024] S3, Risk Indicator Construction: Based on a structured risk feature dataset, a hierarchical indicator system is constructed according to the principles of MECE (Mean Equivalence, Collective Evidence), measurability, importance and relevance, and operability. This system is divided into primary, secondary, and tertiary indicators. Primary indicators include basic labels, pre-job health indicators, safety warning indicators, behavioral assessment indicators, and hazardous behavior indicators. Secondary indicators are further refined based on primary indicators, and tertiary indicators further clarify the specific quantitative direction and classification standards corresponding to secondary indicators. The MECE principle ensures that the definitions of each level of indicators are clear, the scope is well-defined, and there is no overlap, comprehensively covering all core risk factors. The measurability principle ensures that all indicators have clear data sources and can be quantitatively represented through specific values, frequency of occurrence, or classification labels. The importance and relevance principle ensures that the selected indicators are directly related to driving safety and the probability of accident occurrence, and are set differently according to the degree of impact. The operability principle ensures that the indicator division, threshold determination, and weight setting are in line with the actual safety management of public transportation companies, facilitating implementation. When determining the thresholds for indicators, continuous indicators are referenced to industry standards and medical guidelines. Blood pressure indicators are set as follows: systolic blood pressure <120 mmHg and diastolic blood pressure <80 mmHg is considered normal; systolic blood pressure 120-139 mmHg and / or diastolic blood pressure 80-89 mmHg is considered high normal; and systolic blood pressure ≥140 mmHg and / or diastolic blood pressure ≥90 mmHg is considered hypertension. For behavioral frequency indicators, the quartile method is used to set three-level thresholds for yawning behavior (1-3 times per day) and fatigue driving (3-7 times per month). For complex risk indicators, a decision tree algorithm is used to select the optimal split node as the threshold, so that the risk definition of different types of indicators has a clear basis. When setting indicator weights, the basic weights of the primary indicators are first determined using the Analytic Hierarchy Process (AHP): 0.1 for basic label indicators, 0.2 for pre-employment health indicators, 0.4 for safety warning indicators, 0.15 for behavioral assessment indicators, and 0.15 for hazardous behavior indicators. Then, the weights are adjusted using the entropy weight method combined with the actual data distribution over the past six months, with the adjustment range not exceeding ±20% of the basic weights. The final weights are determined after a consistency test (CR) of less than 0.1, highlighting the differences in the impact of different indicators on driving risks. When calculating the coupling coefficient between indicators, the eigenvalues ​​of any two indicators are first discretized into 3 to 5 intervals. Then, the joint entropy and individual marginal entropy of the two indicators are calculated. The formula used to calculate the coupling coefficient between indicators is: ;in, Let be the coupling coefficient between the j-th indicator and the 1st indicator. Let the marginal entropy of the j-th index be , The marginal entropy of the first indicator, Let the joint entropy of the j-th index and the 1st index be... As a minimum constant, the result is calculated using the coupling coefficient formula, which accurately reflects the amplification effect on the risk profile when two indicators are abnormal at the same time.

[0025] S4, Dynamic Scoring Modeling: Based on the constructed hierarchical indicator system, dynamic weight adjustments are made. During the morning rush hour (7:00-9:00), the weight of the safety warning indicator is increased from 0.4 to 0.45, and the weight of road environment-related indicators is increased from 0.1 to 0.15, reflecting the reality that traffic and pedestrian flow are dense during peak hours, and safety warnings and road environment have a more significant impact on driving risks. For drivers with less than 3 years of driving experience, the weight of the historical safety risk indicator is increased from 0.15 to 0.2, strengthening the consideration of historical risk factors for novice drivers with insufficient safety experience. When a driver has received one or more severe warnings within the past 24 hours, the corresponding warning indicator weight is temporarily increased by 50% and lasts for 48 hours, promptly responding to safety hazards caused by recent high-risk behaviors. After the weight adjustments are completed, a multi-dimensional risk coupling scoring algorithm is used to calculate the driver's dynamic risk score. The mathematical expression of the multi-dimensional risk coupling scoring algorithm is: ;in, For the first Dynamic risk score for each driver The total number of core risk indicators For the first The driver The fusion characteristic value of the indicators For the first The dynamically adjusted weights of the indicators For the first Item and the The coupling coefficient of the item index, The weighting is linear main effect weighting. It comprehensively considers the risk contribution of individual indicators and the coupling effect between indicators, making the scoring results more consistent with the risk situation of real-time driving scenarios.

[0026] S5, Risk Management Output: Receives dynamic risk scoring results, uses a dynamic risk level correction algorithm to determine the risk level, and the mathematical expression for the dynamic risk level correction algorithm is: ;in, Let i be the final risk level of the i-th driver. This is the floor function. Let i be the original risk score for the i-th driver. Here, T represents the rectification effect coefficient, and T represents the historical rectification tracking period. Let be the rectification effect value for the i-th driver in week t. The time decay coefficient, This represents the sum of rectification effect values ​​within the i-th driver's historical rectification tracking period. To obtain the maximum value between the sum of rectification effect values ​​and the baseline value v within the i-th driver's historical rectification tracking period, To correct for constants, dynamic risk scores are assigned as follows: 0-20 points correspond to Level 1, 21-40 points to Level 2, 41-60 points to Level 3, 61-80 points to Level 4, and 81-100 points to Level 5. Risk levels are precisely adjusted based on drivers' historical rectification efforts to ensure objectivity. Differentiated management strategies are matched to different risk levels. Level 1 risk drivers receive a risk profile report quarterly, facilitating drivers and companies to understand long-term risk trends. Level 2 risk drivers receive monthly safe driving guidance materials and participate in online safety training every six months, gently guiding them to improve safety awareness. Level 3 risk drivers undergo at least four hours of specialized safety training each time, with adjustments to operational schedules to specifically address driving risk points. Level 4 risk drivers have their operational qualifications suspended for at least three days and are required to participate in offline safety training and psychological counseling to quickly reduce their high-risk status. Level 5 risk drivers are immediately disqualified from work, reported to the industry regulatory authority, and undergo a re-evaluation of their driving qualifications to eliminate serious safety hazards. By disseminating risk profile results through multiple channels, the driver's terminal displays the core indicators of the individual risk profile and targeted improvement suggestions, enabling drivers to clearly understand their shortcomings. This establishes a closed-loop management system of early warning intervention, rectification tracking, effect evaluation, and continuous optimization, achieving dynamic prevention and continuous reduction of driving risks.

[0027] In summary, this embodiment focuses on the scenario of urban core arterial roads during weekday morning rush hour, and implements the entire process of constructing a risk profile for bus drivers. It comprehensively collects various data, including driving behavior, vehicle operation, health status, and enterprise management data, through a multi-dimensional data source system. After data cleaning, standardization, feature engineering, and multi-source data confidence fusion algorithms, a high-quality structured risk feature dataset is formed. A hierarchical indicator system is constructed based on four core principles, with clearly defined thresholds, reasonable weight settings, and calculated indicator coupling coefficients. After dynamically adjusting indicator weights based on peak-hour characteristics, a multi-dimensional risk coupling scoring algorithm is used to obtain a dynamic risk score. Finally, a risk level dynamic correction algorithm is used to determine the risk level, match differentiated control strategies, and output results through multiple channels, constructing a closed-loop control system that accurately adapts to the high-traffic, high-risk driving environment during peak hours, achieving dynamic monitoring and effective reduction of driving risks.

[0028] Example 2: Risk profiles of bus drivers in suburban and urban-rural fringe areas during heavy rain.

[0029] S1, Multi-source data acquisition: Establish a multi-dimensional data source system to widely collect safety warning data from vehicle safety warning platforms, dangerous behavior data from video surveillance platforms, offline pre-job health data, vehicle operation data from vehicle equipment, and basic static data from public transport company management systems. The vehicle-mounted safety warning platform's safety warning data includes behavioral warnings such as drivers making and receiving phone calls, obstructing cameras, distracted driving, fatigued driving, and not wearing seat belts; forward risk warnings such as forward collisions, lane departures, pedestrian collisions, and frequent lane changes; and speed warnings such as speeding, abnormally low-speed driving, and violations of speed limits. The video surveillance platform's dangerous behavior data covers drivers yawning, not looking ahead, excessive speed when entering stations, chatting, sudden deceleration, and failure to slow down at crosswalks. Offline pre-shift health data collection includes heart rate, blood pressure, body temperature, and alcohol content tests before drivers start work. Vehicle operation data from onboard equipment includes vehicle speed, frequency of rapid acceleration and deceleration, steering angle, engine speed, braking frequency, mileage, and the type of road in the suburban / urban-rural fringe area. The public transport company management system's basic static data includes driver employee ID, name, mobile phone number, department, driving route, driving experience, age, historical accident records, historical violation records, safety training assessment results, monthly comprehensive score, and ranking changes. All data is accessed and stored in real time through a unified interface, fully capturing key information such as the characteristics of suburban and urban-rural fringe roads, vehicle operation patterns, and driver status during heavy rain, providing a data foundation tailored to specific scenarios for risk profiling, such as... Figure 2 As shown.

[0030] S2, Data Preprocessing and Fusion: Extracting stored multi-source data and processing it according to the workflow of data cleaning, standardization, feature engineering, and fusion correlation. During data cleaning, missing physiological health data is supplemented by median interpolation of data from drivers of the same department and driving experience during the same period; missing vehicle operation data due to heavy rain is filled by linear interpolation; missing basic static data is supplemented by manual verification of employee files; outliers are identified using box plot IQR; extreme outliers are replaced by the median, and general outliers are replaced by a modified median. After processing, the data missing rate is controlled below 3%, effectively eliminating the interference caused by heavy rain on data collection and ensuring data integrity and accuracy. After data cleaning, data standardization is performed to eliminate the dimensional differences between different data sources. Then, feature engineering is used to extract key risk features in the heavy rain scenario. Finally, a multi-source data confidence fusion algorithm is applied to deeply fuse data from different sources. The mathematical expression of the multi-source data confidence fusion algorithm is: ;in, For the first The core feature fusion value of each driver The total number of data sources. For the first Class data source for the first The basic weight of each driver For the first The real-time reliability coefficient of the data source. For the first The driver in The raw indicator values ​​in the data source are , For the first The global minimum and maximum values ​​of metrics for similar data sources. For the first Coefficient of variation of data source To establish a volatility penalty coefficient, a structured risk feature dataset was constructed to more accurately reflect the driving risk characteristics of suburban and urban-rural fringe areas during rainstorms.

[0031] S3, Risk Indicator Construction: Based on a structured risk feature dataset, a hierarchical indicator system is constructed according to the principles of MECE (Mean Equivalence, Collective Evidence), measurability, importance and relevance, and operability. This system is divided into primary, secondary, and tertiary indicators. Primary indicators include basic labels, pre-job health indicators, safety warning indicators, behavioral assessment indicators, and hazardous behavior indicators. Secondary indicators further refine the primary indicators, and tertiary indicators correspond to the secondary indicators, specifying concrete quantitative directions and classification standards. The MECE principle ensures that indicators do not overlap and comprehensively cover core risk factors. The measurability principle ensures that all indicators have clear data sources and can be quantified. The importance and relevance principle guarantees that indicators are directly related to driving safety and the probability of accidents, and are set differently according to the degree of impact. The operability principle ensures that the indicator settings are aligned with the actual safety management scenarios of public transportation companies, facilitating implementation and application. When determining the thresholds for indicators, continuous indicators are referenced to industry standards and medical guidelines. Blood pressure indicators are categorized into normal range, high-normal value, and hypertension level according to medical diagnostic standards. For behavioral frequency indicators, the quartile method is used to set three-level thresholds for yawning (1-3 times per day) and fatigue driving (3-7 times per month). Complex risk indicators are selected using a decision tree algorithm to select the optimal split node as the threshold, providing clear standards for risk assessment. When setting indicator weights, the basic weights of the first-level indicators are first determined using the analytic hierarchy process: 0.1 for basic label, 0.2 for pre-employment health indicators, 0.4 for safety warning indicators, 0.15 for behavioral assessment indicators, and 0.15 for hazardous behavior indicators. Then, the weights are adjusted using the entropy weight method combined with the actual data distribution of the past 6 months, with the adjustment range not exceeding ±20% of the basic weights. The final weights are determined after a consistency test (CR) of less than 0.1, reasonably allocating the risk contribution ratio of different indicators. When calculating the coupling coefficient between indicators, the eigenvalues ​​of any two indicators are first discretized into 3 to 5 intervals. Then, the joint entropy and individual marginal entropy of the two indicators are calculated. The formula used to calculate the coupling coefficient between indicators is as follows: ;in, Let be the coupling coefficient between the j-th indicator and the 1st indicator. Let the marginal entropy of the j-th index be , The marginal entropy of the first indicator, Let the joint entropy of the j-th index and the 1st index be... As a minimum constant, the result is calculated using the coupling coefficient formula, accurately reflecting the amplifying effect of simultaneous anomalies in two indicators on risk.

[0032] S4, Dynamic Scoring Modeling: Based on the established hierarchical indicator system, dynamic weight adjustments are implemented. During severe rainstorms, the weight of the vehicle operation risk indicator is increased from 0.2 to 0.25 to reflect the increased impact of rainstorms on vehicle braking, steering, and other operational states. For drivers with less than 3 years of driving experience, the weight of the historical safety risk indicator is increased from 0.15 to 0.2 to fully consider the limited ability of novice drivers to cope with complex weather conditions. If a driver experiences one or more severe warnings within the past 24 hours, the corresponding warning indicator weight is temporarily increased by 50% for 48 hours to promptly strengthen the risk assessment of recent high-risk behaviors. After the weight adjustments are complete, a multi-dimensional risk coupling scoring algorithm is used to calculate the driver's dynamic risk score. The mathematical expression of the multi-dimensional risk coupling scoring algorithm is: ;in, For the first Dynamic risk score for each driver The total number of core risk indicators For the first The driver The fusion characteristic value of the indicators For the first The dynamically adjusted weights of the indicators For the first Item and the The coupling coefficient of the item index, The weights are linear main effects. By combining the risk contribution of individual indicators with the coupling effect between indicators, the score accurately reflects the real-time driving risk in suburban and urban-rural fringe areas during rainstorms.

[0033] S5, Risk Management Output: Receives dynamic risk scoring results, uses a dynamic risk level correction algorithm to determine the risk level, and the mathematical expression for the dynamic risk level correction algorithm is: ;in, Let i be the final risk level of the i-th driver. This is the floor function. Let i be the original risk score for the i-th driver. Here, T represents the rectification effect coefficient, and T represents the historical rectification tracking period. Let be the rectification effect value for the i-th driver in week t. The time decay coefficient, This represents the sum of rectification effect values ​​within the i-th driver's historical rectification tracking period. To obtain the maximum value between the sum of rectification effect values ​​and the baseline value v within the i-th driver's historical rectification tracking period, To correct for constants, dynamic risk scores are assigned as follows: 0-20 points correspond to Level 1, 21-40 points to Level 2, 41-60 points to Level 3, 61-80 points to Level 4, and 81-100 points to Level 5. The rating system is optimized by considering the driver's historical rectification efforts, ensuring the results are objective and accurate. A differentiated management strategy is implemented based on risk level. Level 1 risk drivers receive a risk profile report quarterly to facilitate long-term risk monitoring; Level 2 risk drivers receive monthly safe driving guidance materials and participate in online safety training every six months to gradually improve their safe driving skills; Level 3 risk drivers undergo at least four hours of specialized safety training each time, and their operational schedules are adjusted to specifically address potential risks; Level 4 risk drivers have their operational qualifications suspended for at least three days and are required to participate in offline safety training and psychological counseling to quickly reverse the high-risk situation; Level 5 risk drivers are immediately disqualified from driving, reported to the industry regulatory authority, and undergo a new driving qualification review to resolutely prevent major safety accidents. By disseminating risk profile results through multiple channels, the system displays core indicators of individual risk profiles and targeted improvement suggestions on driver terminals, helping drivers accurately identify problems and establish a closed-loop management system of early warning intervention, rectification tracking, effect evaluation, and continuous optimization, thereby achieving effective management of driving risks in suburban and urban-rural fringe areas during rainstorms.

[0034] In summary, this embodiment systematically advances the construction of risk profiles for bus drivers in suburban and urban-rural fringe areas during heavy rain. It comprehensively collects multi-source data adapted to this scenario, eliminates interference from heavy rain through targeted data preprocessing, and integrates the data using a multi-source data confidence fusion algorithm to form a structured risk feature dataset. A hierarchical indicator system is constructed based on four principles, scientifically defining indicator thresholds, weights, and coupling coefficients. The weights of relevant indicators such as vehicle operation risk are dynamically adjusted based on heavy rain and road characteristics. A multi-dimensional risk coupling scoring algorithm is used to obtain accurate dynamic risk scores. After determining the risk level through a dynamic risk level correction algorithm, differentiated control strategies are implemented, and results are output through multiple channels, establishing a closed-loop control system. This effectively addresses the driving risks arising from the superposition of severe weather and complex road conditions, providing strong protection for bus driving safety in special scenarios.

[0035] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A method for constructing a risk profile of bus drivers based on multi-source data fusion, characterized in that, The specific steps of this method are as follows: S1, Multi-source data acquisition: Construct a multi-dimensional data source system to collect safety warning data from the vehicle safety warning platform, dangerous behavior data from the video surveillance platform, pre-job health data from offline equipment, vehicle operation data from vehicle equipment, and basic static data from the public transport enterprise management system. Real-time access and storage of the data are completed through a unified interface. S2, Data Preprocessing and Fusion: Extract the data from step S1, which is collected and stored from multiple sources, and process it according to the process of data cleaning, standardization, feature engineering, and fusion association to construct a structured risk feature dataset; S3, Risk Indicator Construction: Based on the structured risk feature dataset, and in accordance with the principles of MECE, measurability, importance and relevance, and operability, a hierarchical indicator system is constructed, indicator thresholds are determined, indicator weights are set, and coupling coefficients between indicators are calculated to form a quantitative assessment dimension for risk profiling. S4, Dynamic scoring modeling: Based on the risk indicator system built in step S3, the indicator weights are dynamically adjusted, and a multi-dimensional risk coupling scoring algorithm is used to calculate the driver's dynamic risk score, which serves as the core quantitative indicator for risk profiling. S5, Risk Management Output: Receive the score output from step S4, Dynamic Scoring Modeling, use the risk level dynamic correction algorithm to determine the risk level from level one to level five, match differentiated management strategies, output the risk profile results through multiple channels, and establish a closed-loop management system of early warning intervention - rectification tracking - effect evaluation - continuous optimization.

2. The method for constructing a risk profile of a bus driver based on multi-source data fusion according to claim 1, characterized in that, In S1, the multi-source data collection includes: vehicle safety warning platform data, including driver behavior warnings, vehicle forward risk warnings, and vehicle speed warnings, specifically covering data on making and receiving phone calls, obstructing cameras, distracted driving, fatigued driving, not wearing seat belts, forward collisions, lane departures, pedestrian collisions, frequent lane changes, speeding, abnormally low-speed driving, and violations of road speed limits; video surveillance platform data on dangerous behaviors, including drivers yawning, not looking ahead, excessive speed when entering stations, chatting, sudden deceleration, and failure to slow down at crosswalks; offline pre-shift health data, including driver's heart rate, blood pressure, body temperature, and alcohol content test data before starting work; vehicle operation data from onboard equipment, including vehicle speed, frequency of sudden acceleration, frequency of sudden deceleration, steering angle, engine speed, braking frequency, mileage, and road segment type data during vehicle operation; and basic static data from the public transport company management system, including driver employee number, name, mobile phone number, department, driving route, driving experience, age, historical accident records, historical violation records, safety training assessment results, monthly comprehensive score, and ranking changes.

3. The method for constructing a risk profile of a bus driver based on multi-source data fusion according to claim 1, characterized in that, In S2, during data preprocessing and fusion, the data cleaning process supplements missing physiological health data by using median interpolation of data from drivers with the same age and department during the same period. Missing vehicle operation data is filled using linear interpolation. Missing basic static data is supplemented by manually checking employee files. Outliers are identified using the box plot IQR method. Extreme outliers are replaced with the median, while general outliers are replaced with the modified median. The data missing rate is reduced to below 3%.

4. The method for constructing a risk profile of a bus driver based on multi-source data fusion according to claim 1, characterized in that, In step S2, during data preprocessing and fusion, the mathematical expression for the multi-source data confidence fusion algorithm is: ;in, For the first The core feature fusion value of each driver The total number of data sources. For the first Class data source for the first The basic weight of each driver For the first The real-time reliability coefficient of the data source. For the first The driver in The raw indicator values ​​in the data source are , For the first The global minimum and maximum values ​​of metrics for similar data sources. For the first Coefficient of variation of data source This is the fluctuation penalty coefficient.

5. The method for constructing a risk profile of a bus driver based on multi-source data fusion according to claim 1, characterized in that, The S3 risk indicator construction includes a hierarchical indicator system comprising: the MECE principle, meaning that each level of indicator is clearly defined and has a clear scope, with no overlap between them, and can comprehensively cover all key dimensions related to the risk profile of bus drivers, without omitting any core risk factors; the measurability principle, meaning that all indicators have clear data sources and can be quantitatively represented through specific values, frequency of occurrence, or classification labels, without vague qualitative descriptions; the importance and relevance principle, meaning that the selected indicators are directly related to the driving safety and accident probability of bus drivers, and are differentiated according to the degree of influence of each indicator on driving safety; and the operability principle, meaning that the division of indicators, the determination of thresholds, and the setting of weights are all combined with the actual safety management scenarios of bus companies, and the indicator analysis results can be directly applied to the construction of risk profiles of bus drivers and related safety management work, facilitating implementation; the hierarchical indicator system is divided into first-level indicators, second-level indicators, and third-level indicators. First-level indicators include basic labels, pre-job health indicators, safety warning indicators, behavioral assessment indicators, and dangerous behavior indicators. Second-level indicators are further subdivided based on first-level indicators, and third-level indicators further clarify the specific quantitative direction and classification standards corresponding to the second-level indicators.

6. The method for constructing a risk profile of a bus driver based on multi-source data fusion according to claim 1, characterized in that, In S3, the specific content of determining indicator thresholds, setting indicator weights, and calculating the coupling coefficients between indicators in the risk indicator construction process is as follows: When determining indicator thresholds, continuous indicators are set with reference to industry standards and medical norms; blood pressure indicators are set with reference to medical diagnostic standards, with systolic blood pressure <120 mmHg and diastolic blood pressure <80 mmHg considered normal; systolic blood pressure 120-139 mmHg and / or diastolic blood pressure 80-89 mmHg considered high normal; and systolic blood pressure ≥140 mmHg and / or diastolic blood pressure ≥90 mmHg considered hypertension. Frequency-based behavioral indicators are set using the quartile method, with yawning behavior divided into three levels based on a daily average of 1 or 3 times, and fatigue driving divided into three levels based on a monthly average of 3 or 7 times. Complex risk indicators use a decision tree algorithm to select the optimal split node as the threshold. When setting indicator weights, the basic weights of the primary indicators are first determined using the analytic hierarchy process (AHP): basic label weight 0.1, pre-job health indicator weight 0.2, safety warning indicator weight 0.4, behavioral assessment indicator weight 0.15, and hazardous behavior indicator weight 0.

15. Then, the weights are adjusted using the entropy weight method combined with the actual data distribution over the past 6 months, with the adjustment range not exceeding ±20% of the basic weight. The final weights are determined after the consistency test CR is less than 0.

1. When calculating the coupling coefficient between indicators, the feature values ​​of any two indicators are first discretized and divided into 3 to 5 intervals. Then, the joint entropy and the marginal entropy of the two indicators are calculated and obtained according to the coupling coefficient formula. The coupling coefficient reflects the amplification effect of the simultaneous occurrence of anomalies in two indicators on the risk profile.

7. The method for constructing a risk profile of bus drivers based on multi-source data fusion according to claim 1, characterized in that, In the dynamic scoring modeling described in S4, the dynamic weight adjustment specifically involves increasing the weight of the safety warning indicator from 0.4 to 0.45 and the weight of the road environment-related indicator from 0.1 to 0.15 during peak hours (7:00-9:00 and 17:00-19:00); increasing the weight of the vehicle operation risk indicator from 0.2 to 0.25 during severe weather; increasing the weight of the historical safety risk indicator from 0.15 to 0.2 for drivers with less than 3 years of driving experience; and temporarily increasing the weight of the corresponding warning indicator by 50% for 48 hours when a driver has received one or more severe warnings within the past 24 hours.

8. The method for constructing a risk profile of a bus driver based on multi-source data fusion according to claim 1, characterized in that, In S4, the mathematical expression for the multidimensional risk-coupled scoring algorithm in dynamic scoring modeling is: ;in, For the first Dynamic risk score for each driver The total number of core risk indicators For the first The driver The fusion characteristic value of the indicators For the first The dynamically adjusted weights of the indicators For the first Item and the The coupling coefficient of the item index, These are the linear main effect weights.

9. The method for constructing a risk profile of a bus driver based on multi-source data fusion according to claim 1, characterized in that, In S5, the mathematical expression for the dynamic risk level correction algorithm in the risk control output is: ;in, Let i be the final risk level of the i-th driver. This is the floor function. Let i be the original risk score for the i-th driver. Here, T represents the rectification effect coefficient, and T represents the historical rectification tracking period. Let be the rectification effect value for the i-th driver in week t. The time decay coefficient, This represents the sum of rectification effect values ​​within the i-th driver's historical rectification tracking period. To obtain the maximum value between the sum of rectification effect values ​​and the baseline value v within the i-th driver's historical rectification tracking period, To correct for constants; dynamic risk scores of 0 to 20 correspond to Level 1, 21 to 40 to Level 2, 41 to 60 to Level 3, 61 to 80 to Level 4, and 81 to 100 to Level 5.

10. The method for constructing a risk profile of a bus driver based on multi-source data fusion according to claim 1, characterized in that, In the S5 risk management output, the differentiated management strategy specifically includes generating a risk profile report for Level 1 risk drivers once per quarter; and sending safe driving guidance materials once per month to Level 2 risk drivers and having them participate in online safety training once every six months. Drivers at level 3 risk will receive specialized safety training, with each training session lasting no less than 4 hours, and their operational schedules will be adjusted. Drivers at level 4 risk will have their operational qualifications suspended for no less than 3 days and will be required to participate in offline safety training and psychological counseling. Drivers at level 5 risk will have their driving qualifications immediately suspended, and the matter will be reported to the relevant industry authorities for a new driving qualification review. The multi-channel output includes: displaying the driver's personal risk profile core indicators and targeted improvement suggestions on the driver's terminal.