Method and device for assessing traffic accidents
The pairwise difference learning model addresses the precision and reliability issues in traffic accident injury assessment by analyzing data pair differences with weighted anchor points, enhancing prediction accuracy and reliability even with limited data.
Patent Information
- Application Number
- DE102024116397
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-12
- Publication Date
- 2025-12-18
AI Technical Summary
Existing methods for assessing traffic accident injury severity are not sufficiently precise or reliable, especially with limited data sets, prone to overfitting, and lack nuanced probability assessments.
A pairwise difference learning (PDL) model is employed to analyze differences between accident data pairs, using regression analysis with weighted anchor points to improve prediction accuracy and reliability, even with scarce data, by focusing on causal relationships and physical principles.
The PDL model enhances prediction precision and reliability, providing finer granularity in injury severity assessments, supporting critical decision-making with improved generalizability and interpretability.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The present invention relates to a method for assessing traffic accidents. The present invention further relates to a corresponding device, a corresponding computer program, and a storage medium. State of the art
[0002] In the assessment of traffic accidents, it is crucial to be able to accurately and reliably estimate the severity of injuries sustained by vehicle occupants. This assessment is important not only for immediate medical care, but also for accident analysis, insurance, and, most importantly, industrial vehicle development.
[0003] Traditionally, the assessment of injury risk is based on statistical methods and biomechanical knowledge. Data from past accidents are analyzed to identify patterns and influencing factors on the severity of injuries. These factors include, for example, the speed and angle of impact, the type and mass of the vehicle, and the physiological characteristics of the occupants, such as height, weight, and age.
[0004] The simplified injury scale (AIS) is frequently used to classify injuries. This internationally recognized scale allows for a standardized assessment of injury severity by categorizing injuries from minor to life-threatening. Besides this scale, other categorizations exist that distinguish, for example, between "not injured," "slightly injured," "severely injured," and "fatally injured."
[0005] In recent years, the application of machine learning (ML) in accident analysis has gained importance. Machine learning offers the possibility of recognizing complex patterns in large datasets and, based on this, making predictions about the outcome of future events. Various learning models, from logistic regression to complex neural networks, are trained to predict the severity of injuries based on input data describing accident and occupant characteristics.
[0006] DE102008027509A1 describes a method of incremental simulation for predicting the severity of accident damage and pedestrian injuries. The categorization of vehicle damage is also explained. A statistical distribution function can be used within the procedure.
[0007] DE102009029955A1 describes a method for mapping test datasets to real-world accident data. Pedestrian injuries can be taken into account. A tolerance range is defined for classification.
[0008] CN116415181A describes a method for implementing a sorting algorithm by pairwise comparison.
[0009] CN108446713A describes a multi-label classification method based on logistic regression, which is intended to improve its error tolerance.
[0010] WO202463765A1 proposes a ranking procedure by pairwise comparison, which should enable a (graduated) logistic multiclass regression in which the labels of the classes are meaningfully ordered.
[0011] CN107766873A describes a supervised learning-based ranking method for sorting hit lists in information retrieval (IR), recommendation systems, or machine translation.
[0012] US2019197414A1 describes a system for predicting injuries, comprising a variety of sensors. It also provides a device with an input-output interface to which a processor is connected to make a prediction about the probability of an injury. The system predicts the types of injuries and also describes the severity of head trauma (whiplash).
[0013] US2020334928A1 describes a method to investigate a collision, in which the probability of injury can also be predicted.
[0014] CN116680552A also sheds light on various methods, devices and a motor vehicle to be able to predict the probability of injury to an occupant.
[0015] Finally, the as-yet-unpublished patent application DE10 2024 103 687.7 discloses a method for assessing traffic accidents based on a classifier. The task of assigning a given or simulated accident to one of several severity levels is ultimately reduced to a binary classification task. Disclosure of the invention
[0016] One problem is that existing methods for assessing the severity of injuries sustained in traffic accidents are not sufficiently precise or reliable, especially when limited data sets are available. Traditional statistical approaches and machine learning models typically rely on large and diverse datasets to make accurate predictions. In reality, however, such datasets are often unavailable, either due to privacy concerns, the rarity of certain accident types, or the inaccessibility of detailed accident data.
[0017] It also follows that state-of-the-art machine learning models are prone to overfitting when the amount of training data is small relative to the model's complexity. Overfitting leads to a model that performs well on training data but generalizes poorly and makes inaccurate predictions on new, unknown data.
[0018] Furthermore, existing approaches are often limited in their predictive capability because they only directly assign input data to injury categories. They do not consider the possibility of improving predictive accuracy by examining differences between similar cases. This leads to inadequate results in situations where a graded probability assessment of injury severity would be desirable.
[0019] Furthermore, conventional approaches do not provide sufficient information about the reliability or confidence of the prediction. In practice, this means that decision-makers may not have the full picture of the uncertainty of a forecast, which is particularly important in critical decision-making situations such as the immediate response to traffic accidents.
[0020] The problem described is solved by a method for assessing traffic accidents, a corresponding device, a corresponding computer program and a corresponding storage medium according to the independent claims.
[0021] This approach offers the advantage of increasing the precision and reliability of accident impact assessments, even with limited datasets. By training a model that focuses on predicting differences between similar cases rather than direct classification, the information content of each dataset is maximized. This leads to more effective use of existing data, which is beneficial given the scarcity of data in this field.
[0022] Pairwise difference learning (PDL) enables deeper understanding and more accurate predictions from a limited set of accident data by leveraging the relationships between data points. This approach can reduce the likelihood of overfitting because the model is trained to recognize the underlying pattern of differences rather than focusing on specific instances. This improves the model's generalizability to new, unknown accident data.
[0023] The proposed approach differs from the method known from DE10 2024 103 687.7 in that it reinterprets the present classification task as a regression analysis with several dependent variables, thus opening up the possibility of using a regression function (hereinafter referred to as "regressor") as a basic model whose anchor points are efficiently weighted.
[0024] To illustrate the advantages of this PDL model, consider four "internal" regressors. Each regressor is assigned to one of the four AIS severity levels and assigns to an accident scenario—explained in more detail below—the probability with which that accident reaches the respective severity level. The accident data set used as a predictor contains, for example, the combination of the passive safety features of the vehicle involved in the accident.
[0025] Under these assumptions, the inventive "decomposition" of the overall assessment into sub-problems enables the designer or safety engineer to understand whether and to what extent a particular safety feature influences the degree of occupant injury. For example, the regressor assigned to the "minor injury" category might depend heavily on features related to the knee airbag. This could lead to the interpretation that knee airbags primarily protect against minor injuries. Conversely, the regressor assigned to the "fatal injury" category might depend to a greater extent on features related to the seat belts.
[0026] A key advantage of PDL with regressors lies in its interpretability. Although mechanical engineers may be familiar with the meaning of the individual passive safety features, a detailed understanding of how the ML model functions is necessary before it can be used for important decisions in the development process. This ensures that the model has been trained correctly and is based on causal relationships and physical principles.
[0027] Pairwise difference learning thus offers finer granularity in prediction by providing probability distributions across severity levels. This additional information enables a more nuanced assessment of injury severity and therefore provides a more detailed basis for decisions in rescue services, accident analysis, and safety-related aspects of vehicle development.
[0028] Further advantageous embodiments of the invention are specified in the dependent patent claims.
[0029] Brief description of the drawings Fig. Figure 1 shows the percentage improvement in the pairwise difference compared to the underlying regressor. Fig. Figure 2 shows the percentage improvement of the weighted over the unweighted pairwise difference. Embodiments of the invention
[0030] The PDL method, well-known from regression, is used according to the invention to classify the severity of injuries in traffic accidents. The PDL method modifies the conventional task of supervised learning methods, in which results are inferred from individual input instances, by learning the differences between the results of input pairs.
[0031] Specifically, the model is trained on a dataset consisting of pairs of accident data. For each pair of instances, the difference in injury severity is used as the target variable. This target variable is expressed as a numerical value representing the difference between the injury severity levels assigned to the two instances of the pair. The process of learning this difference function can be mathematically represented as follows: It was D={(xi,yi)}i=1N⊂ℝd×Y the training dataset, consisting of N accident data instances, where x i the input characteristics (for example, road conditions, weather, vehicle type, vehicle mass, occupant characteristics) and y i represents the corresponding severity of the injuries. The severity is coded on a scale, such as the simplified injury scale AIS.
[0032] For the classification, where the result y is not a real number, but a discrete label from the set Y={1,…,K} If this is the case, the principle of PDL is adapted as follows: A model is trained that predicts whether two instances x and x' belong to the same class. For this purpose, the original training dataset is used. D into a new training dataset Pair converted, which consists of pairs of instances: Dpair={(zi,j,yi,j)|1≤i,j≤N}, where z i,j the concatenation of the feature vectors x i and x j and y i,j is defined as: yi,j=yi−yi, so that y i,j ∈ {-1,0,1} K A regressor f: ℝ 2d → ℝ K The regressor is trained on this dataset. It should be noted that the regressor accepts values outside of [-1,1] KThis may be assumed. To account for this circumstance, a corresponding truncation can be applied. Furthermore, assume that f is a probabilistic classifier such that f(z) ∈ [0,1] gives the probability with which i belongs to the positive class more often than j.
[0033] For the use according to the invention, the predictor is symmetrized: fsym(xi,xj)=f(xi,xj)+f(xj,xi)2.
[0034] When predicting a new query instance x q can the probability of belonging to the class be determined? yq∈Y now quantify as ppost,i(y)=yi+gsym(xq,xi). and average across all training examples to obtain the final probability: ppost(y)=1N∑i=1Nppost,i(y).
[0035] The class ŷ q The highest estimated probability is ultimately used as the prediction for the query instance x. qchosen: y^q=arg maxy∈Yppost(y).
[0036] This PDL classifier already achieves a significantly better test result than the underlying regressor in 66% of the classification tasks (see Fig. 1) Its prediction accuracy can be further increased by introducing weights for labeled training points—so-called anchor points. The variant of the k-nearest neighbor algorithm referred to below as "KNN Shapley" according to JIA, Ruoxi, et al. Efficient task specific data valuation for nearest neighbor algorithms. Proceedings of the VLDB Endowment, 2018, Vol. 12, No. 11, pp. 1610–1623.
[0037] While conventional PDL assigns equal weight to all training points, a preferred embodiment of the invention uses KNN Shapley to redistribute the weights among the training points. KNN Shapley assigns a scalar value to each training point. If a point tends to negatively impact the validation result, this data value is likely to be detrimental. One technique for dealing with such negative weights is to remove these points from the weighted average. Therefore, by weighting the PDL algorithm with KNN Shapley, only a portion of the training points are used for prediction.
[0038] This extension allows the predictions of the PDL method to be refined by weighting anchor points that provide more reliable information more heavily, while simultaneously giving less weight to those with lower predictive power. By training the weights on a validation dataset, the generalizability of the model is improved, and better performance can be achieved on new, unknown data. A corresponding implementation therefore outperforms the underlying PDL regressor for most classification tasks (see Fig. 2). QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] DE 102008027509A1
[0006] DE 102009029955A1
[0007] CN 116415181A
[0008] CN 108446713A
[0009] WO 202463765A1
[0010] CN 107766873A
[0011] US 2019197414A1
[0012] US 2020334928A1
[0013] CN 116680552A
[0014] DE 10 2024 103 687.7 [0015, 0023] Cited non-patent literature
[0000] Proceedings of the VLDB Endowment, 2018, Volume 12, No. 11, pp. 1610-1623
[0036]
Claims
Methods for assessing traffic accidents, characterized by the following features: - past traffic accidents are recorded, each with a known severity level, and - a model is trained by pairwise differential learning on the past traffic accidents and their respective severity levels to determine, for each severity level, the probability with which a simulated traffic accident reaches this severity level. Method according to claim 1, characterized by the following feature: - the past traffic accidents are described during recording based on external circumstances such as road conditions or weather. Method according to claim 1 or 2, characterized by the following feature: - the past traffic accidents are described during recording based on any impact, for example by impact speed or angle. Method according to one of claims 1 to 3, characterized by the following feature: - the past traffic accidents are described during recording based on the vehicles involved, for example by vehicle type or mass. Method according to claim 4, characterized by the following feature: - the past traffic accidents are described during recording based on any occupants of the vehicles, for example by height, weight or age. Method according to claim 5, characterized by the following feature: - the severity is measured according to any injuries of the occupants. Method according to claim 6, characterized by the following feature: - the severity is indicated on a simplified injury scale. Device characterized by the following features: - the device is set up to carry out a method according to one of claims 1 to 7. Computer program which is configured to perform all steps of a method according to any one of claims 1 to 7. Machine-readable storage medium with a computer program stored thereon according to claim 9.
Citation Information
Patent Citations
Method, system, test system and computer program product for the predictive determination of test results for a technical test object
DE102020132042A1
Method for determining a health effect, method for training a predictive function, monitoring device, vehicle and training device
DE102022104129A1
Collision injury prediction model creation method, collision injury prediction method, collision injury prediction system, and advanced automatic accident notification system
JP2020061088A
JP002020061088A