Methods for Mining Typical Risk Scenarios and Assessing Risks Based on Accident Big Data and Large Models

CN122736330APending Publication Date: 2026-09-11ROAD TRAFFIC SAFETY RES CENT THE MINIST OF PUBLIC SECURITY OF THE PEOPLES REPUBLIC OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610941485.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-28
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

(一)事故过程信息获取维度不足,难以还原碰撞机理

Benefits of technology

(1)本发明通过将结构化事故记录与非结构化事故文本共同作为分析输入,既利用了结构化字段的规范性,又挖掘了事故文本中蕴含的参与者运动状态、运动方向、碰撞形态等深层语义信息,能够更完整地还原风险场景中各方交通参与者的运动过程和相互之间的冲突关系,克服了传统方法在信息粒度与处理效率之间难以兼顾的技术难题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122736330A_ABST
    Figure CN122736330A_ABST
Patent Text Reader

Abstract

This invention relates to the fields of traffic safety and artificial intelligence technology, specifically a method for mining and assessing typical risk scenarios based on big data and large models of accidents. The method includes: acquiring road traffic accident data, including structured accident records and unstructured accident text; inputting both structured and unstructured accident text into a large model and outputting standardized results; performing consistency verification and inferring the positional relationships between the main accident scenarios; using basic accident scenarios as risk assessment units, calculating the frequency index and hazard index of each basic accident scenario within the same road space type to obtain a comprehensive risk index; and superimposing environmental factors on the basic accident scenarios to form enhanced accident scenarios. This invention achieves automated deep analysis of road traffic accident data, standardized risk scenario construction, and multi-dimensional risk assessment, and is applicable to the construction of a hazardous test scenario library, as well as urban traffic risk insight, road safety evaluation, and governance decision support.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of traffic safety and artificial intelligence technology, specifically a method for mining and assessing typical risk scenarios based on big data and large models of accidents. Background Technology

[0002] Road traffic accident data analysis is a crucial foundational task for traffic safety research and autonomous driving technology development. Currently, technological advancements in this field primarily revolve around accident statistics, accident type classification, injury and fatality analysis, and road environmental factor assessment. Existing methods for road traffic accident data analysis mainly include the following: (I) Structured Field Analysis Method Based on Accident Statistics Database In existing technologies, traffic management departments and research institutions have generally established accident statistics databases to record and manage structured fields such as accident time, location, weather, road type, accident type, and consequences of injuries and fatalities. Based on this, conventional statistical methods such as frequency statistics, cross-analysis, time series analysis, and trend analysis are used to analyze the macroscopic distribution characteristics of accident data to serve the assessment of regional traffic safety situations and governance decisions. These methods are relatively mature and have been widely used at the government management level.

[0003] (II) Accident Scenario Induction Method Based on Human Experience In the construction of autonomous driving test scenarios and traffic safety research, a common approach is to invite traffic safety experts or senior accident handlers to manually summarize and categorize several common accident types from a large number of accident cases based on their professional knowledge and practical experience. These types include left-turn conflicts at intersections, rear-end collisions, and pedestrian crossings. This method can utilize the experts' deep semantic understanding and domain knowledge to reconstruct the typical process patterns of accidents to a certain extent.

[0004] (III) Text information extraction methods based on keywords or regular expressions To automatically extract useful information from accident texts, some existing technologies employ keyword matching or regular expression matching to extract predefined categories such as "going straight," "turning left," "rear-end collision," "rainy weather," and "highway" from unstructured texts such as brief case descriptions. This method is simple to implement, has low computational overhead, and is easy to deploy in practical systems.

[0005] (iv) Accident text classification methods based on machine learning or natural language processing In recent years, with the advancement of natural language processing technology, some studies have begun to employ text classification models or sequence labeling models to process accident texts, including accident type classification, responsibility type identification, and participant type extraction. Compared to rule-based methods, these methods offer improved generalization ability and automation.

[0006] (v) Methods for constructing scenario libraries for autonomous driving testing In the field of autonomous driving technology research and development, building test scenario libraries based on regulatory requirements, industry standards, expert experience, or real-world accident cases has become a common industry practice. These scenario libraries are used for simulation testing, field testing, and open road testing to verify the safety performance of autonomous driving systems in various traffic scenarios.

[0007] (vi) Traffic safety risk assessment methods Existing traffic safety risk assessment methods typically employ statistical indicators such as the number of accidents, casualties, and serious accident rates, or methods like risk matrices, to evaluate the safety risk level of specific roads, areas, or accident types. These methods have relatively mature applications in macro-level traffic safety governance.

[0008] While the aforementioned existing technologies have some effectiveness in their respective application scenarios, they still have the following shortcomings when facing comprehensive application needs such as automated processing of large-scale accident data, objective identification of typical risk scenarios, and risk scenario recommendations for the development of autonomous driving systems: (i) Insufficient dimensions of information obtained about the accident process make it difficult to reconstruct the collision mechanism. The granularity of coded fields in accident databases is too coarse, making it difficult to reconstruct the specific movements and conflict relationships among traffic participants in an accident. While manual reading of accident texts can yield relatively rich semantic information, it is inefficient and inconsistent, making it difficult to support automated analysis of large-scale accident databases. Existing methods struggle to strike a balance between the granularity and efficiency of accident information acquisition.

[0009] (ii) Insufficient stability of semantic extraction from unstructured accident texts Traditional information extraction methods based on keywords or regular expressions struggle to reliably identify complex semantic relationships such as "turning left from north to south," "driving in the wrong direction," "stationary vehicle ahead," "oncoming vehicle," and "entering a ramp." Differences in accident text description habits among different regions and law enforcement personnel further exacerbate the instability of rule matching. While machine learning methods have improved generalization capabilities to some extent, their ability to model deeper semantic information such as movement direction and relative position implicit in accident texts remains limited.

[0010] (iii) Lack of a standardized risk scenario description framework and failure to differentiate between road space types. Current accident statistics often focus on accident type labels or casualty counts, lacking a unified method that organically combines multi-dimensional information such as participant type, motion state, relative positional relationships, and collision patterns into standardized descriptive units. Different researchers use varying standards when defining and classifying scenarios. Furthermore, existing methods typically fail to clearly distinguish between intersection scenarios and road segment scenarios, leading to the analysis of accident modes with different risk mechanisms and governance strategies, such as intersection conflicts, road segment rear-end collisions, road segment lane changes, and road segment speeding, all mixed together. This affects the relevance of subsequent governance recommendations and the accuracy of constructing autonomous driving test scenarios.

[0011] (iv) Risk assessment has a single dimension and lacks quantitative means to assess the impact of environmental factors. Existing risk assessment methods often rely solely on statistical analysis of accident frequency or the severity of accident consequences, making it difficult to identify accident types with low frequency but severe consequences. Furthermore, they lack a comprehensive risk ranking mechanism that weights and integrates frequency and hazard dimensions. In addition, analyses of the impact of environmental factors on risk often remain at the level of independent statistics for single fields, failing to quantify the conditional influence of specific environmental factors on the frequency and severity of specific accident types.

[0012] (v) The output of large models lacks structured constraints and has insufficient ability to analyze multiple participating entities. Existing intelligent accident text analysis methods, if relying directly on large models for free text output, are prone to problems such as inconsistent output categories, unstable output formats, and untraceable reasoning processes. This is especially true when extracting directional information, as the free output of large models often lacks unified coordinate standards and mechanisms for completing missing information. Furthermore, in complex accidents involving multiple traffic participants, traditional two-subject analysis frameworks are prone to overlooking other key risk sources in the accident chain, making it difficult to comprehensively depict the complete risk landscape of the accident scenario. Summary of the Invention

[0013] To address the shortcomings of existing technologies, this invention aims to provide a method for mining and assessing typical risk scenarios based on accident big data and large models. This method aims to achieve automated in-depth analysis of accident data, standardized risk scenario description, and comprehensive risk assessment, thereby overcoming or mitigating at least some of the deficiencies of existing technologies.

[0014] To achieve the above objectives, the technical solution adopted by this invention is as follows: A method for mining and assessing typical risk scenarios based on accident big data and large models, comprising the following steps: Step 1: Obtain road traffic accident data, which includes structured accident records and unstructured accident text; extract risk scenarios from the structured accident records to construct core elements, which at least include the accident time, weather, road type, intersection / road segment type, road surface condition, road surface condition, traffic signal method, number of participants, participant type, and consequences of personal injury or death; Step 2: Input the structured accident records and the unstructured accident text into the locally deployed large model, and use a fixed prompt word template to constrain the output of the large model to a standardized result. The standardized result includes a subject list, a subject relationship list, analysis basis, and uncertainty explanation; each subject in the subject list is identified in order of its degree of direct relevance to the occurrence of the collision, risk formation, or liability determination. Step 3: For any two entities in the entity list output in Step 2, perform consistency verification and mutual positional relationship inference: convert the movement direction text of each entity into a two-dimensional direction vector, and determine the relative positional relationship between the entities based on the same direction, opposite direction or perpendicular relationship between the direction vectors. The relative positional relationship includes same-direction road, opposite road, same-side perpendicular and opposite perpendicular. Step 4: Using basic accident scenarios as risk assessment units, each basic accident scenario consists of combinations of road space types, participant types, participant motion states, relative positional relationships, relative motion relationships or collision patterns, and accident consequences. Basic accident scenario sets are constructed according to intersections and road segments. The frequency index and hazard index of each basic accident scenario within the same road space type are calculated. The frequency index and hazard index are normalized separately, and then weighted and fused to obtain a comprehensive risk index. The hazard index has a greater weight than the frequency index. Step 5: Based on the basic accident scenario, environmental factors are superimposed to form an enhanced accident scenario; the conditional frequency increase and conditional hazard increase of a specific basic accident scenario relative to the overall level under each environmental factor condition are calculated to quantify the impact of each environmental factor on the occurrence frequency and hazard degree of each basic accident scenario.

[0015] As a limitation of the present invention: In step two, each subject in the subject list is identified as subject A, subject B and extended subject in order of its degree of direct relevance to the occurrence of the collision, the formation of the risk or the determination of liability; for accidents involving three or more traffic participants, the same accident is split into multiple subject pairs composed of different subjects, and each subject pair records the subject type, motion state, motion direction, relative position relationship and relative motion relationship, and the multiple subject pairs are associated as a multi-subject composite accident scenario through the same accident event number.

[0016] As a limitation of the present invention: in step two, the subject type, motion state, motion direction and relative motion relationship are all selected and output by the large model from a preset category set; The preset categories of the subject types include cars, buses, trucks, motorcycles, non-motorized vehicles, pedestrians, and others; The preset categories of the motion states include going straight, turning left, turning right, standing still, reversing, making a U-turn, driving along a roundabout, entering a roundabout, exiting a roundabout, inside a ramp, entering a ramp, exiting a ramp, and others; The preset categories of the direction of motion include south to north, north to south, east to west, west to east, and stationary; The preset categories of relative motion relationships include rear-end collision, lane change, overtaking, stationary object, U-turn, perpendicular, side collision, and others.

[0017] As a limitation of the present invention: in step three, converting the motion direction text of each subject into a two-dimensional direction vector includes: Convert south to north into a positive longitudinal direction vector, north to south into a negative longitudinal direction vector, west to east into a positive lateral direction vector, and east to west into a negative lateral direction vector. Set the direction vector of a stationary subject to an empty vector or inherit it from the direction of the road it is located on. When two motion direction vectors are in the same direction, they are determined to be roads in the same direction; when two motion direction vectors are opposite, they are determined to be roads in opposite directions; when two motion direction vectors are perpendicular to each other, the motion state of turning left, turning right, or going straight is combined to determine whether they are perpendicular on the same side or perpendicular in opposite directions. When a relative position description exists in the unstructured accident text, the relative position description takes precedence over the direction vector determination result and is used to correct the inference result of the mutual positional relationship.

[0018] As a limitation of the present invention: in step four, the frequency index F(s,r) is calculated in the following manner: F(s,r) = n(s,r) / N(r) Where r represents the road space type, s represents the basic accident scenario, n(s,r) represents the number of accidents in scenario s within road space type r, and N(r) represents the total number of accidents under road space type r. When traffic exposure data is available, the frequency index is replaced by the accident rate per unit of traffic exposure.

[0019] As a limitation of the present invention, the hazard index S(s,r) is calculated in the following manner: S(s,r) = [w0×n0(s,r) + w1×n1(s,r) + w2×n2(s,r)] / n(s,r) Where n0(s,r), n1(s,r), and n2(s,r) represent the number of accidents without injuries, the number of accidents with injuries but no deaths, and the number of accidents with deaths in scene s of road space type r, respectively, and w0, w1, and w2 are the corresponding weights, with w2>w1>w0.

[0020] As a limitation of the present invention: in step four, the normalization process adopts minimum-maximum normalization, logarithmic normalization or quantile normalization to map the frequency index and the hazard index to the [0,1] interval respectively; The comprehensive risk index R(s,r) is calculated as follows: R(s,r) = α×F'(s,r) +β×S'(s,r) Where F'(s,r) is the normalized frequency index, S'(s,r) is the normalized hazard index, α+β=1, and β>α.

[0021] As a limitation of the present invention, step four further includes: For all basic accident scenarios within the same statistical range, calculate the 75th percentile F of the frequency index set. 75 The 75th percentile S of the set of hazard indices 75 ; With the frequency exponent F≥F 75 The scene is determined to be a high-frequency scene, and F < F 75 The scenario is classified as a low-frequency scenario; the hazard index S≥S 75 The scenario is classified as a high-risk scenario, and S < S 75 The scenario was determined to be a low-risk scenario; Based on the assessment results, the basic accident scenarios are divided into the following four categories: Low-frequency high-risk scenarios: F < F 75 And S≥S 75 ; High-frequency, high-risk scenarios: F≥F 75 And S≥S 75 ; High-frequency, low-risk scenarios: F≥F 75 And S < S 75 ; Low-frequency, low-risk scenarios: F < F 75 And S < S 75 .

[0022] As a limitation of the present invention: in step five, the conditional frequency boost is calculated in the following manner: Conditional frequency boost = P(base scenario | environmental factors) / P(base scenario) Wherein, P(basic scenario | environmental factors) represents the conditional probability of the basic accident scenario occurring under specific environmental factors, and P(basic scenario) represents the overall probability of the basic accident scenario occurring in all samples; The conditional hazard elevation is calculated as follows: Conditional hazard enhancement = Hazard index under environmental factors / Overall hazard index; When the conditional frequency increase or conditional hazard increase is greater than 1, it is determined that the environmental factor has an increasing effect on the occurrence frequency or hazard level of the basic accident scenario.

[0023] As a limitation of the present invention: the environmental factors include one or more of the following: the time of the accident, weather conditions, road type, intersection and road segment type, road surface condition, road surface condition, and traffic signal method; The enhanced accident scenario is formed by superimposing the basic accident scenario and the environmental factors.

[0024] By adopting the above technical solution, the beneficial effects achieved by the present invention compared with the prior art are as follows: (1) By using both structured accident records and unstructured accident texts as analysis inputs, this invention not only utilizes the standardization of structured fields, but also mines the deep semantic information contained in the accident text, such as the movement state, movement direction, and collision pattern of the participants. It can more completely restore the movement process of all traffic participants in the risk scenario and the conflict relationship between them, thus overcoming the technical problem of traditional methods that are difficult to balance information granularity and processing efficiency.

[0025] (2) This invention completes the semantic understanding and information extraction of accident text through a locally deployed large model inference service, avoiding direct reliance on external cloud services. It can be deployed and used in application scenarios with high requirements for traffic accident data security, effectively reducing the security risks of data leakage. At the same time, a fixed prompt word template is used to structurally constrain the output of the large model, requiring it to output standardized results according to a preset category set, overcoming the defects that are easy to occur when the large model outputs freely, such as category inconsistency, unstable format, and untraceable inference process.

[0026] (3) In this invention, a large model is responsible for extracting basic information such as subject type, motion state, motion direction, and relative motion relationship from accident text. Then, the motion direction text is uniformly converted into two-dimensional direction vectors through the rule reasoning module. Consistency verification and mutual positional relationship inference are performed based on the same-direction, opposite-direction, or perpendicular relationships between direction vectors. The judgment results are preferentially corrected using relative positional descriptions such as "oncoming vehicle" and "vehicle ahead" in the text. The collaborative work of the large model and rule reasoning not only gives full play to the large model's advantage in understanding complex text semantics, but also reduces the instability of the large model's free text output through rule reasoning, thereby improving the consistency and reliability of accident text analysis across regions and expression habits.

[0027] (4) In actual road traffic accidents, there are often complex situations involving multiple traffic participants participating in the same accident. Traditional two-subject analysis frameworks are prone to overlooking other key risk sources in the accident chain when dealing with multi-subject accidents. This invention supports adding extended subjects in accidents involving three or more subjects, splitting the same accident into multiple subject pairs composed of different subjects for separate recording and analysis, and re-associating multiple subject pairs into a multi-subject composite accident scenario through the same accident event number, which can more comprehensively depict the complete risk picture of complex accidents.

[0028] (5) The present invention constructs a set of basic accident scenarios according to intersections and road segments. The intersection scenario focuses on representing the multi-flow interaction relationship such as straight, left turn, right turn, and intersection conflict. The road segment scenario focuses on representing the continuous road driving risks such as rear-end collision, lane change, overtaking, speeding, and single vehicle collision with stationary objects. This effectively avoids the problem of mixed accident mechanisms under different road space types, and enables the risk scenarios discovered to serve traffic management and autonomous driving test scenario construction in a more targeted manner.

[0029] (6) This invention calculates the frequency index and hazard index of each basic accident scenario, normalizes them separately, and then weights and merges them according to different weights to obtain a comprehensive risk index. By comparing it with the 75th percentile of the frequency index set and the hazard index set, each scenario is divided into four categories: low frequency high risk, high frequency high risk, high frequency low risk, and low frequency low risk. This can not only identify the common risk scenarios that occur frequently, but also effectively identify the extreme risk scenarios that occur infrequently but have serious consequences. This provides an objective basis for prioritizing traffic management and selecting extreme test scenarios for autonomous driving.

[0030] (7) This invention creates enhanced accident scenarios by superimposing environmental factors such as the time of the accident, weather, road type, road surface condition, road surface condition, and traffic signal mode on the basis of basic accident scenarios. It calculates the frequency increase and hazard increase of specific scenarios relative to the overall level under environmental factors and presents the degree of influence of each environmental factor on each basic accident scenario in a standardized quantitative index form. This provides data support for traffic control strategies under severe weather and the setting of test priorities for autonomous driving systems under extreme environmental conditions.

[0031] (8) The output of this invention can be used as a basis for constructing a dangerous test scenario library for the development of autonomous driving systems, and for prioritizing typical risk scenarios in simulation testing, field testing and open road testing; it can also be used for urban traffic risk insight, road safety evaluation and traffic management report preparation, serving the safety management decision-making of road traffic management departments, and has good industrial application prospects. In summary, this invention achieves automated deep analysis of road traffic accident data, standardized risk scenario construction, and multi-dimensional risk assessment. It effectively overcomes several shortcomings of existing technologies in areas such as accident process information reconstruction, semantic extraction stability, scenario standardization description, comprehensive risk assessment, and quantitative analysis of environmental factors. It is applicable to the construction of a hazardous test scenario library for autonomous driving system development and to urban traffic risk insight, road safety evaluation, and governance decision support for traffic management departments. Attached Figure Description

[0032] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0033] Figure 1 This is a schematic diagram of a system module architecture supporting multi-entity expansion in an embodiment of the present invention; Figure 2 This is a flowchart illustrating the process of determining the relationship between entities based on direction vectors and rules in an embodiment of the present invention. Figure 3 This is a schematic diagram of the basic scenario frequency-hazard four-quadrant classification in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the impact of environmental factors on scene frequency and hazards in an embodiment of the present invention. Detailed Implementation

[0034] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the method for mining and assessing typical risk scenarios based on accident big data and large models described herein is a preferred embodiment and is only used to illustrate and explain the present invention, and does not constitute a limitation thereof.

[0035] I. Definitions of Abbreviations and Key Terms To facilitate understanding of the technical solutions of this embodiment, some abbreviations and key terms used in the text are defined below. It should be noted that these terms are only used to assist in describing the technical solutions of this application and do not constitute the sole limitation on the scope of protection of this application.

[0036] Road traffic accident information investigation field system: refers to the standardized field system used in the investigation, registration and statistics of road traffic accident information, including but not limited to fields such as accident time, weather, road type, road surface condition, participant type, and injury / death consequences.

[0037] Structured accident data refers to accident records stored in the form of fields, such as accident time, weather, road type, traffic signal type, number of motor vehicles, number of non-motor vehicles, number of pedestrians, and casualties.

[0038] Unstructured accident text: refers to natural language text generated during the accident handling process, including brief case summary, accident liability determination letter, on-site description, etc., which usually contains semantic information such as the direction of movement of participants, turning intention, collision pattern, avoidance behavior, and liability determination.

[0039] Large models refer to natural language processing models capable of semantic understanding, information extraction, and structured output of accident texts. This application preferably uses locally deployed large model inference services, but is not limited to specific model names or deployment tools.

[0040] Key stakeholders: These refer to traffic participants directly involved in the collision, risk formation, liability determination, or consequences of injury or death. By default, the stakeholder who appears first or is first mentioned in the description of primary liability is designated as Stake A, the second key stakeholder as Stake B, etc. In multi-stake accidents, Stake C, Stake D, etc., are added following the same logic.

[0041] Subject pair: refers to an analysis unit consisting of two key subjects, such as subject A-subject B, subject A-subject C, or subject B-subject C. Multi-subject accidents can be broken down into multiple subject pairs for analysis of their relative positions and motion relationships.

[0042] Basic accident scenario: refers to the smallest statistical unit composed of road space type, combination of participant type, combination of motion state, relative position relationship, collision or relative motion relationship, and accident consequences.

[0043] Enhanced accident scenarios: These are scenarios created by overlaying environmental factors such as the time of the accident, weather, road type, road surface condition, road surface condition, and traffic signal type onto a basic accident scenario.

[0044] Frequency index: refers to the proportion or normalized occurrence level of a certain type of basic accident scenario in the same road space type.

[0045] Hazard Index: This refers to the intensity of casualties corresponding to a certain type of basic accident scenario, calculated by weighting consequences such as no injury, injury, and death.

[0046] Comprehensive Risk Index: This refers to the risk ranking index obtained by normalizing and weighting the frequency index and the hazard index.

[0047] Condition elevation: refers to the degree to which the frequency or severity of a particular scenario increases relative to the overall level when a specific environmental factor occurs.

[0048] II. System Architecture Overview This embodiment uses big data on road traffic accidents in my country as input, with data sources including structured accident records, brief case descriptions, and accident liability determination documents. The system employs a locally deployed large-scale model inference service to parse the accident text and outputs typical risk scenarios through rule-based inference, scenario construction, risk assessment, and environmental factor analysis modules.

[0049] like Figure 1 As shown, the system architecture consists of five layers from bottom to top: Data layer: Contains structured accident data, accident code descriptions, brief case details, and accident liability determination documents, serving as the system's raw data input.

[0050] Standardization processing layer: Includes field cleaning, enumeration mapping, missing value handling, and information survey element alignment, used to transform multi-source heterogeneous raw data into a unified format.

[0051] Large model analysis layer: includes prompt word templates, subject and multi-subject recognition, motion state and direction recognition, JSON structured output and confidence and basis records, used to extract key semantic information from unstructured accident text.

[0052] Rule-based reasoning layer: Includes subject-to-subject relationship generation, relative position inference, intersection / segment traffic diversion, collision pattern recognition, and standardized verification and conflict resolution, used for structured verification and relationship inference of model output.

[0053] The assessment output layer includes scenario frequency index, hazard index, comprehensive risk index, environmental factor enhancement degree, and a risk scenario library and report, which are used to generate the final risk assessment results.

[0054] The following provides a detailed explanation of each processing step.

[0055] III. Step 1: Extracting core elements for constructing risk scenarios (a) Obtaining road traffic accident data The system acquires road traffic accident data, which includes structured accident records and unstructured accident text. Structured accident records include at least the following fields: accident time, weather, road type, intersection / road segment type, road surface condition, road surface condition, traffic signal type, number of participants, participant type, and consequences of injuries or fatalities. Unstructured accident text includes a brief case summary and an accident liability determination letter. (ii) Standardize the processing of accident data Before inputting data into the large model, the system first standardizes the acquired multi-source heterogeneous data to ensure the accuracy and consistency of subsequent semantic extraction and statistical analysis. The standardization process specifically includes the following sub-steps: Field cleaning and enumeration mapping. The system performs validity checks on each field in the structured accident records, removing invalid, duplicate, or obviously erroneous records. Simultaneously, it maps codes or expressions from different sources and periods in the original data to the system's preset standard enumeration values. For example, "passenger car," "sedan," and "car" are uniformly mapped to "sedan," and "1 / 2 / 3 / 4" branch intersections are uniformly mapped to the corresponding intersection type enumerations.

[0056] Missing value handling and information investigation element alignment. The system detects missing fields (such as accident time, road type, and casualty consequences); for fields with a small number of missing fields but which can be inferred from the accident text, a mark is reserved for extraction by the larger model; for records with severely missing key fields that cannot be inferred, they are marked as low confidence for manual review or removed. Simultaneously, the system aligns the field system of the original accident database with the information investigation elements defined in this invention to ensure semantic consistency of each element in subsequent scenario construction.

[0057] Through the above standardization process, the adverse effects of common problems in the original accident data, such as missing fields, inconsistent coding, and mixed expressions, on the accuracy of subsequent large-scale model analysis and risk quantification assessment can be effectively reduced.

[0058] Based on this, the system extracts risk scenarios and constructs core elements according to the road traffic accident information investigation field system. The optional categories and technical applications of each core element are shown in Table 1.

[0059] Table 1. Optional Categories and Technical Applications of Core Elements

[0060] IV. Step Two: Parsing Accident Text Based on Local Large Model The system inputs structured fields along with the accident text (brief case summary and accident liability determination letter) into a locally deployed large-scale model inference service. To avoid issues such as inconsistent categories, unstable formats, and untraceable inference processes caused by the model's free output, the system uses fixed prompt word templates, requiring the model to output standardized results, including a list of subjects, a list of subject relationships, analysis basis, and uncertainty explanations. This embodiment does not limit the specific name, parameter size, or deployment tool of the large-scale model, only requiring that it can complete text understanding and structured output in a local environment.

[0061] (a) Subject Identification Rules By default, the first traffic participant appearing in the description of primary responsibility is designated as Subject A, and the second key traffic participant as Subject B. When there are three or more subjects in the text, the logic of Subject A and Subject B is followed, and Subject C, Subject D, etc., are added sequentially according to their direct relevance to the collision occurrence, risk formation, or liability determination. For multi-subject cases, the system splits the same accident into multiple subject pairs composed of different subjects. Each subject pair records its type, movement state, movement direction, relative position relationship, and relative movement relationship. Multiple subject pairs are associated into a multi-subject complex accident scenario through the same accident event number.

[0062] (ii) Output category constraints of large models Large models must select outputs from the following preset category set: Main vehicle types: cars, buses, trucks, motorcycles, non-motorized vehicles, pedestrians, and others; Movement states: straight, left turn, right turn, stationary, driving along the roundabout, entering the roundabout, exiting the roundabout, inside the ramp, entering the ramp, exiting the ramp, reversing, making a U-turn, others; Movement direction: South to North, North to South, East to West, West to East, Stationary; If the text does not specify a direction, the default direction can be used or it can be marked as low confidence pending review; Relative motion relationships: rear-end collision, lane change, overtaking, stationary object, U-turn, perpendicular, side collision, others.

[0063] Large models output structured JSON or equivalent machine-readable format, while retaining the basis for analysis, confidence markers, and fields requiring manual review.

[0064] (III) Prompt Word Template The prompt template can be summarized as follows: comprehensively analyze structured data, accident code descriptions, brief case details, and liability determination texts to identify key traffic participants involved in the accident, their movement status, direction of movement, and basic relative positional relationships; the subject type, movement status, direction of movement, and relative movement relationship must be selected from preset categories; the output includes a list of subjects, a list of subject relationships, relative movement relationships, analysis basis, and uncertainty explanations.

[0065] In this embodiment, the large model uses a fixed prompt word template for parsing, requiring it to output standard JSON to avoid inconsistencies in fields caused by free text. The prompt words include structured accident quantity information, accident code descriptions, brief case details, and accident liability determination text.

[0066] Here is an example of a prompt word template: Please analyze the following traffic accident information to identify the types of entities involved in the accident, their state of motion, direction of motion, and relative positions of their bases.

[0067] Structured data: Number of motor vehicles: {vehicle_count} Number of non-motorized vehicles: {non_motor_count} Number of pedestrians: {pedestrian_count} Accident code description: {accident_code_desc} Brief case description: {case_description} Accident Liability Determination Letter: {responsibility_text} Analysis rules: Subject identification: The first key subject is A, and the second key subject is B; Type selection is limited to: cars, buses, trucks, motorcycles, non-motorized vehicles, pedestrians, and others.

[0068] Movement status selection: straight, left turn, right turn, stationary, reversing, U-turn, driving along the roundabout, entering the roundabout, exiting the roundabout, inside the ramp, entering the ramp, exiting the ramp, other.

[0069] Movement direction is limited to: south to north, north to south, east to west, west to east, and stationary; turning records the initial direction before turning; missing directions can be inferred from semantics such as ahead and opposite.

[0070] Relative motion relationships are limited to: rear-end collision, lane change, overtaking, stationary object, U-turn, perpendicular, side collision, and others.

[0071] Output JSON, without any other irrelevant content, and provide the basis for the analysis.

[0072] The following is an example of the standard output of a large model: { "Type A of Main Body": "Sedan" "Subject Type B": "Non-motorized vehicles" "Subject A's motion state": "Straight ahead" "Main Body B's Movement Status": "Turn Left" "Direction of movement of subject A": "South to North" "Direction of movement of subject B": "West to East" "Relative Motion Relationship": "P6" Analysis Basis: "Main body A travels straight from south to north, while main body B turns left from west to east, creating a perpendicular intersection." } (iv) Case Study of Multi-Party Accident Analysis The following is a virtual example of multi-subject accident analysis, which is only used to illustrate the input and output method of the present invention and does not involve any real accident data.

[0073] Table 2 Examples of Multi-Agent Accident Analysis

[0074] V. Step Three: Rule-based reasoning to infer the relationships between subjects like Figure 2 As shown, after the large model is output, the system uses the rule reasoning module to perform consistency verification and infer the mutual positional relationship of any two subjects in the subject list.

[0075] (a) Direction vector transformation For subject pairs with sufficient directional information, the system converts the directional text into a two-dimensional directional vector: south to north represents the positive longitudinal direction, north to south represents the negative longitudinal direction, west to east represents the positive lateral direction, and east to west represents the negative lateral direction. The directional vector of a stationary subject can be set to an empty vector or inherited from the direction of the road it is located on.

[0076] (II) Rules for Determining Relative Positional Relationships When two motion direction vectors are in the same direction, they are determined to be on the same road; when two motion direction vectors are in opposite directions, they are determined to be on opposite roads; when two motion direction vectors are perpendicular to each other, the direction is determined by combining the motion states such as left turn, right turn, and straight ahead, and is determined to be perpendicular on the same side or perpendicular to opposite sides. If the text contains relative descriptions such as "oncoming vehicle," "vehicle ahead," or "non-motorized vehicle to the left front," these descriptions are used first to correct the direction inference results.

[0077] (III) Example of pseudocode for the judgment process Input: The direction and state of motion of subject i; the direction and state of motion of subject j; and a description of the relative position of the text.

[0078] ① Convert “south to north, north to south, west to east, east to west” into discrete two-dimensional direction vectors.

[0079] ② If both directions are valid, calculate the vector dot product: if the dot product is positive and the directions are consistent, output "same direction road"; if the dot product is negative and the directions are opposite, output "opposite direction road"; if the dot product is close to zero, proceed to the perpendicular relationship judgment.

[0080] ③ If one party is going straight and the other is turning right, or both are turning right, prioritize outputting "same-side perpendicular"; if one party is going straight and the other is turning left, or both are turning left, prioritize outputting "opposite perpendicular".

[0081] ④ If there is only one direction, make supplementary inferences based on the other direction's left turn, right turn, or straight-ahead status; if it still cannot be determined, output "Other" or mark it as low confidence for manual review.

[0082] ⑤ For multi-entity accidents, repeat the above process for each entity pair and output the entity pair relationship table.

[0083] Step 4: Construct basic accident scenarios and conduct risk assessments The system uses each basic accident scenario as a risk assessment unit. A basic accident scenario is formed by combining elements such as road space type, participant type, movement state, relative positional relationship, relative motion relationship or collision pattern, and accident consequences. For cases involving three or more traffic participants, the system decomposes and analyzes them according to subject pairs, and re-associates multiple subject pairs into multi-subject composite accident scenarios using the same accident event number.

[0084] (a) Standardized representation of basic scenarios Basic scenario = {Road space type, combination of participant types, combination of participant motion states, relative positional relationships, relative motion relationships or collision patterns, accident consequences} The system constructs basic scenario sets and statistical criteria for intersections and road segments respectively. Intersection scenarios focus on representing multi-directional interactions such as straight-ahead, left-turn, right-turn, and intersection conflicts; road segment scenarios focus on representing continuous road driving risks such as rear-end collisions, lane changes, overtaking, speeding, and single-vehicle collisions with stationary objects, thereby avoiding the mixing of accident mechanisms in different road spaces.

[0085] (ii) Calculation of frequency index Let r represent road space type, s represent a basic accident scenario, n(s,r) represent the number of accidents in scenario s within road space type r, and N(r) represent the total number of accidents under road space type r. When exposure data such as traffic flow, road mileage, or traffic volume are lacking, the frequency index can be represented by the relative frequency in the accident sample: F(s,r) = n(s,r) / N(r) When traffic exposure data is available, the frequency index can also be corrected to the accident rate per unit of traffic exposure, for example, by dividing the number of accidents by exposures such as vehicle kilometers, traffic flow, number of vehicles passing through, or road mileage, thereby reducing the bias caused by differences in traffic scale under different road, time, and weather conditions.

[0086] (III) Calculation of Hazard Index Accident consequences can be categorized into three levels: no injury, injury without death, and death, each assigned a different weight; the weight of death is greater than that of injury, and the weight of injury is greater than that of no injury. The hazard index S(s,r) is calculated as follows: S(s,r) = [w0×n0(s,r) + w1×n1(s,r) + w2×n2(s,r)] / n(s,r) Where n0(s,r), n1(s,r), and n2(s,r) represent the number of accidents without injuries, the number of accidents with injuries but no fatalities, and the number of accidents with fatalities in scenario s within road space type r, respectively. w0, w1, and w2 are the corresponding weights, with w2>w1>w0. If the accident database contains specific numbers of deaths, injuries, and property damage, a severity index can be further constructed based on the intensity of personal injury and the intensity of property damage.

[0087] (iv) Normalization processing To eliminate the influence of dimensions and sample size, the system maps the frequency index F and hazard index S to the [0,1] interval, which can be achieved using min-max normalization, logarithmic normalization, or quantile normalization. F'(s,r) = [F(s,r) - min(F)] / [max(F) - min(F)] S'(s,r) = [S(s,r) - min(S)] / [max(S) - min(S)] Where F'(s,r) is the normalized frequency index, and S'(s,r) is the normalized hazard index.

[0088] (v) Calculation of Comprehensive Risk Index The comprehensive risk index is calculated by weighting the standardized frequency index and the hazard index: R(s,r) = α×F'(s,r) +β×S'(s,r) Where α + β = 1. In this embodiment, α is 0.4 and β is 0.6, that is, the frequency weight is 0.4 and the hazard weight is 0.6. These values ​​are determined based on expert experience and application objectives, reflecting the method's priority focus on accident scenarios that cause serious consequences such as injury or death.

[0089] (vi) Four-quadrant classification The system further classifies the basic accident scenarios into four quadrants based on their frequency and hazard indices. Specifically, for all basic accident scenarios within the same statistical range, the 75th percentile F of the frequency index set is calculated. 75 The 75th percentile S of the set of hazard indices75 And use them as high-frequency thresholds and high-risk thresholds.

[0090] Let F be the frequency index and S be the hazard index for a given basic accident scenario. If F ≥ F 75 If F <F 75 If S ≥ S 75 If S 75 If the threshold is met, the scenario is classified as low-risk. For scenarios that are exactly equal to the threshold, they are preferentially classified into high-frequency or high-risk categories to avoid missing critical risk scenarios.

[0091] Based on the above rules, the system classifies basic accident scenarios into four categories, as shown in Table 3. Figure 3 The diagram shown is a schematic representation of the basic scenario frequency-hazard four-quadrant classification.

[0092] Table 3. Four-quadrant classification of basic accident scenarios

[0093] VII. Step Five: Analysis of the Impact of Environmental Factors on Scene Frequency and Hazards After obtaining the basic accident scenario, the system further superimposes environmental factors such as the accident occurrence time, weather conditions, road type, intersection and road segment type, road surface condition, road surface condition, and traffic signal method to form an environment-enhanced accident scenario, which is used to analyze the impact of environmental conditions on the frequency, severity, and distribution of accident types.

[0094] When traffic exposure data such as traffic flow, road mileage, vehicle kilometers, or vehicle traffic volume are available, the system can calculate the risk of accidents per unit of traffic exposure. When exposure data is lacking, the system can calculate the proportion of accidents, the proportion of serious consequences, and the relative escalation under different environmental conditions, and use these as relative comparison results within the accident sample.

[0095] Specifically, Figure 4 This is a flowchart illustrating the analysis of the environmental factors' impact on scene frequency and hazards corresponding to this step. The following is combined with... Figure 4 Each processing sub-step is described in detail.

[0096] (a) Input and Environment Grouping like Figure 4 As shown, the analysis process begins in the input phase: the system takes the basic scenario library generated in step four as input. The basic scenario library contains the intersection / road segment labels and their consequences data (including the number of accidents with no injuries, injuries but no deaths, and deaths) corresponding to each basic accident scenario.

[0097] ​The system then proceeds to the environmental grouping stage: Accident data in the basic scenario database is grouped according to environmental factors, including weather (sunny, cloudy, rainy, snowy, foggy, etc.), road type (highway, urban road, highway, etc.), road surface condition (intact, under construction, uneven, etc.), road surface condition (dry, wet, waterlogged, icy, etc.), traffic signal type (no control, traffic lights, signs and markings, etc.), and the time of accident occurrence (morning rush hour, daytime, evening rush hour, nighttime). Through environmental grouping, the system divides the overall accident data into subsamples under multiple environmental factor conditions, providing a data foundation for subsequent calculations of conditional frequency and conditional hazard.

[0098] (ii) Calculation of conditional frequency lift After the environment grouping is completed, the system calculates the conditional frequency increase of a specific basic accident scenario under each environmental factor condition according to the following formula: Frequency increase = P(base scenario | environmental factors) / P(base scenario) Where P(Base Scenario | Environmental Factor) represents the conditional probability of the occurrence of the base accident scenario under specific environmental factor conditions, and P(Base Scenario) represents the overall probability of the occurrence of the base accident scenario in all samples. If exposure data is available, the numerator and denominator can be replaced with the accident rate per unit exposure under the corresponding conditions. The frequency increase is used to quantify whether the occurrence of a certain environmental factor increases the frequency of a specific scenario relative to the overall level.

[0099] (III) Calculation of the hazard enhancement degree Subsequently, the system calculates the conditional hazard enhancement degree of a specific basic accident scenario under various environmental factors using the following formula: Hazard amplification = Hazard index under environmental factors / Overall hazard index The hazard escalation score quantifies whether the occurrence of a particular environmental factor increases the severity of the consequences in a specific scenario relative to the overall level. If the hazard escalation score is greater than 1, it indicates that the consequences of the scenario are more severe under that environmental factor; if the hazard escalation score is less than 1, it indicates that the severity of the scenario is lower than the overall level under that environmental factor.

[0100] (iv) Statistical verification After calculating the frequency boost and hazard boost, the system enters the statistical verification phase to assess the statistical reliability of the boost results. The statistical verification specifically includes the following sub-steps: ① Sample size threshold verification. The system checks whether the sample size of the basic accident scenario under specific environmental conditions reaches the preset minimum sample size threshold. If the sample size is insufficient, the lift result is marked as low confidence, indicating that more data needs to be accumulated before re-evaluation.

[0101] ② Interval estimation. The system uses bootstrapping to construct confidence intervals for frequency elevation and hazard elevation. If the confidence interval does not contain 1, the elevation is considered statistically significant.

[0102] ③ Permutation Test. The system further verifies the statistical significance of the lift by using a permutation test: the environmental factor labels are randomly shuffled in the sample, the lift distribution under random conditions is calculated, the actual observed lift is compared with this null distribution, and the p-value is calculated. If the p-value is less than the preset significance level, the lift is determined to be statistically significant.

[0103] Through the above statistical verification, misjudgments caused by random fluctuations in the sample can be effectively avoided, ensuring that the output environmental factor enhancement results have statistical reliability.

[0104] (v) Factor contribution output After statistical verification, the system enters the factor contribution output stage, which includes the following: ① Identification of high-impact factors. The system filters out environmental factors whose frequency increase or hazard increase is significantly greater than 1 and passes statistical verification, and marks them as high-impact factors.

[0105] ② Combination Factor Analysis. The system further analyzes the cumulative effect when multiple environmental factors occur simultaneously. For example, it analyzes whether the frequency increase of a specific scenario under the combination of "rainy day + night" is significantly higher than the increase when "rainy day" or "night" occurs alone, in order to identify combinations of environmental factors with synergistic effects.

[0106] ③ Explanatory rule generation. The system transforms the lift analysis results into explanatory rules described in natural language, such as "Under slippery road conditions, the fatality rate of speeding accidents on road sections increases by 16% relative to the overall level."

[0107] (vi) Final Results Convergence like Figure 4 As shown, the three intermediate results—hazard amplification calculation, significance and confidence verification, and factor contribution output—are ultimately fed into the final step to form scenario enhancement suggestions. Specifically, the system uses the statistically verified environmental factor amplification results as environmental modification parameters, which are then superimposed onto the corresponding basic accident scenario to form an environmentally enhanced risk scenario. The output format is a typical risk scenario + environmental modification parameters, for example: Road section - speeding - fatal consequences + slippery road surface (hazard escalation 1.16, p<0.05) Intersection - Left turn conflict - Injury consequences + Rainy weather (Frequency increase 1.32, p<0.05) The aforementioned enhanced environmental risk scenarios can be directly used to construct extreme environmental conditions in autonomous driving simulation testing, as well as to address key risks under severe weather conditions in urban traffic management.

[0108] (vii) Identification of long-tail risk scenarios For low-frequency but high-risk scenarios such as single-vehicle accidents, collisions with stationary objects, construction obstacles, road debris, and temporary traffic facilities, the system can extract the category of the collided object and abnormal road conditions from the accident case text, providing a reference for extreme testing scenarios of autonomous driving and urban traffic risk management.

[0109] (viii) Virtual Example of Environmental Factor Enhancement Calculation The following is a set of statistical samples of speeding accidents on completely virtual road sections, which is only used to demonstrate the calculation process of this invention and does not represent the conclusions of real accidents.

[0110] Table 4. Statistical Sample of Speeding Accidents on Virtual Road Sections

[0111] In the aforementioned virtual sample, the proportion of fatal accidents due to speeding on slippery road surfaces was 11.6%, and the overall proportion of fatal accidents due to speeding on all road sections was 10.0%. Therefore, the hazard enhancement level is: Hazard escalation rate = 11.6% / 10.0% = 1.16 Specifically, under slippery road conditions, the proportion of fatal accidents caused by speeding on road sections increases by 16% relative to the overall level. This result can be used to prompt urban traffic risk insight systems to pay attention to the scenario of "ordinary road section - speeding - slippery road surface - fatal consequences", and can be transformed into the risk scenario of high-speed driving on low-adhesion roads in autonomous driving simulation tests.

[0112] VIII. Implementation Results Output After the above steps, this embodiment outputs the following results: Table 5 Output Results and Explanations .

Claims

1. A method for mining and assessing typical risk scenarios based on big data and large models of accidents, characterized in that, Includes the following steps: Step 1: Obtain road traffic accident data, which includes structured accident records and unstructured accident text; extract risk scenarios from the structured accident records to construct core elements, which at least include the accident time, weather, road type, intersection / road segment type, road surface condition, road surface condition, traffic signal method, number of participants, participant type, and consequences of personal injury or death; Step 2: Input the structured accident records and the unstructured accident text into the locally deployed large model, and use a fixed prompt word template to constrain the output of the large model to a standardized result. The standardized result includes a subject list, a subject relationship list, analysis basis, and uncertainty explanation; each subject in the subject list is identified in order of its degree of direct relevance to the occurrence of the collision, risk formation, or liability determination. Step 3: For any two entities in the entity list output in Step 2, perform consistency verification and mutual positional relationship inference: convert the movement direction text of each entity into a two-dimensional direction vector, and determine the relative positional relationship between the entities based on the same direction, opposite direction or perpendicular relationship between the direction vectors. The relative positional relationship includes same-direction road, opposite road, same-side perpendicular and opposite perpendicular. Step 4: Using basic accident scenarios as risk assessment units, the basic accident scenarios consist of road space types, combinations of participant types, combinations of participant motion states, relative positional relationships, relative motion relationships or collision patterns, and combinations of accident consequences; construct basic accident scenario sets according to intersections and road segments respectively; Calculate the frequency index and hazard index of each basic accident scenario within the same road space type, normalize the frequency index and hazard index respectively, and then perform a weighted fusion of the normalized frequency index and hazard index to obtain a comprehensive risk index; wherein the weight of the hazard index is greater than the weight of the frequency index. Step 5: Add environmental factors to the basic accident scenario to create an enhanced accident scenario; Calculate the conditional frequency increase and conditional hazard increase of specific basic accident scenarios relative to the overall level under various environmental factors, in order to quantify the impact of each environmental factor on the occurrence frequency and hazard level of each basic accident scenario.

2. The method for mining and assessing typical risk scenarios based on accident big data and large models as described in claim 1, characterized in that, In step two, each entity in the entity list is identified as entity A, entity B, and extended entity in order of their direct relevance to the occurrence of the collision, the formation of the risk, or the determination of liability. For accidents involving three or more traffic participants, the same accident is split into multiple entity pairs composed of different entities. Each entity pair records the entity type, motion state, motion direction, relative position relationship, and relative motion relationship. The multiple entity pairs are associated as a multi-entity composite accident scenario through the same accident event number.

3. The method for mining and assessing typical risk scenarios based on accident big data and large models as described in claim 2, characterized in that, In step two, the subject type, motion state, motion direction, and relative motion relationship are all selected and output by the large model from a preset category set; The preset categories of the subject types include cars, buses, trucks, motorcycles, non-motorized vehicles, pedestrians, and others; The preset categories of the motion states include going straight, turning left, turning right, standing still, reversing, making a U-turn, driving along a roundabout, entering a roundabout, exiting a roundabout, inside a ramp, entering a ramp, exiting a ramp, and others; The preset categories of the direction of motion include south to north, north to south, east to west, west to east, and stationary; The preset categories of relative motion relationships include rear-end collision, lane change, overtaking, stationary object, U-turn, perpendicular, side collision, and others.

4. The method for mining and assessing typical risk scenarios based on accident big data and large models as described in claim 3, characterized in that, Step three, converting the motion direction text of each subject into a two-dimensional direction vector, includes: Convert south to north into a positive longitudinal direction vector, north to south into a negative longitudinal direction vector, west to east into a positive lateral direction vector, and east to west into a negative lateral direction vector. Set the direction vector of a stationary subject to an empty vector or inherit it from the direction of the road it is located on. When two motion direction vectors are in the same direction, they are determined to be roads in the same direction; when two motion direction vectors are opposite, they are determined to be roads in opposite directions; when two motion direction vectors are perpendicular to each other, the motion state of turning left, turning right, or going straight is combined to determine whether they are perpendicular on the same side or perpendicular in opposite directions. When a relative position description exists in the unstructured accident text, the relative position description takes precedence over the direction vector determination result and is used to correct the inference result of the mutual positional relationship.

5. The method for mining and assessing typical risk scenarios based on accident big data and large models according to any one of claims 1 to 4, characterized in that, In step four, the frequency exponent F(s,r) is calculated as follows: F(s,r) = n(s,r) / N(r) Where r represents the road space type, s represents the basic accident scenario, n(s,r) represents the number of accidents in scenario s within road space type r, and N(r) represents the total number of accidents under road space type r. When traffic exposure data is available, the frequency index is replaced by the accident rate per unit of traffic exposure.

6. The method for mining and assessing typical risk scenarios based on accident big data and large models according to claim 5, characterized in that, The hazard index S(s,r) is calculated as follows: S(s,r) = [w0×n0(s,r) + w1×n1(s,r) + w2×n2(s,r)] / n(s,r) Where n0(s,r), n1(s,r), and n2(s,r) represent the number of accidents without injuries, the number of accidents with injuries but no deaths, and the number of accidents with deaths in scene s of road space type r, respectively, and w0, w1, and w2 are the corresponding weights, with w2 > w1 > w0.

7. The method for mining and assessing typical risk scenarios based on accident big data and large models as described in claim 6, characterized in that, In step four, the normalization process uses minimum-maximum normalization, logarithmic normalization, or quantile normalization to map the frequency index and the hazard index to the [0,1] interval, respectively. The comprehensive risk index R(s,r) is calculated as follows: R(s,r) = α×F'(s,r) +β×S'(s,r) Where F'(s,r) is the normalized frequency index, S'(s,r) is the normalized hazard index, α+β=1, and β>α.

8. The method for mining and assessing typical risk scenarios based on accident big data and large models as described in claim 7, characterized in that, Step four also includes: For all basic accident scenarios within the same statistical range, calculate the 75th percentile F of the frequency index set. 75 The 75th percentile S of the set of hazard indices 75 ; With the frequency exponent F≥F 75 The scene is determined to be a high-frequency scene, and F < F 75 The scenario is classified as a low-frequency scenario; the hazard index S≥S 75 The scenario is classified as a high-risk scenario, and S < S 75 The scenario was determined to be a low-risk scenario; Based on the assessment results, the basic accident scenarios are divided into the following four categories: Low-frequency high-risk scenarios: F < F 75 And S≥S 75 ; High-frequency, high-risk scenarios: F≥F 75 And S≥S 75 ; High-frequency, low-risk scenarios: F≥F 75 And S < S 75 ; Low-frequency, low-risk scenarios: F < F 75 And S < S 75 .

9. The method for mining and assessing typical risk scenarios based on accident big data and large models according to claim 8, characterized in that, In step five, the conditional frequency boost is calculated as follows: Conditional frequency boost = P(base scenario | environmental factors) / P(base scenario) Wherein, P(basic scenario | environmental factors) represents the conditional probability of the basic accident scenario occurring under specific environmental factors, and P(basic scenario) represents the overall probability of the basic accident scenario occurring in all samples; The conditional hazard elevation is calculated as follows: Conditional hazard enhancement = Hazard index under environmental factors / Overall hazard index; When the conditional frequency increase or conditional hazard increase is greater than 1, it is determined that the environmental factor has an increasing effect on the occurrence frequency or hazard level of the basic accident scenario.

10. The method for mining and assessing typical risk scenarios based on accident big data and large models according to any one of claims 1-4 and 6-9, characterized in that, The environmental factors include one or more of the following: time of the accident, weather conditions, road type, intersection / road segment type, road surface condition, road surface condition, and traffic signal method; The enhanced accident scenario is formed by superimposing the basic accident scenario and the environmental factors.