A risk assessment method for signal-free intersections based on semantic analysis
Patent Information
- Application Number
- CN202610927552.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-22
AI Technical Summary
第一,传统主题聚类模型(如LDA)基于词袋模型运算,输出结果为离散词汇组合,逻辑关联性弱,仍需人工大量标注分析,智能化程度低;
(1)本发明通过将高维轨迹数据转化为纯客观运动学语义文本,并结合大语言模型的主题挖掘与归类能力,实现了对无信号交叉口危险场景的降维解析与自然语言聚类,能够直接、明确地揭示触发交通冲突的深层物理与空间致因,为路权优化和自动驾驶策略改进提供了高解释性的决策依据;
Smart Images

Figure CN122799622A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of traffic safety and intelligent traffic control, specifically relating to a risk assessment method for unsignalized intersections based on semantic analysis. Background Technology
[0002] As autonomous driving technology moves from closed testing to open road deployment, intersections without traffic signals become high-risk areas with frequent traffic accidents and low traffic efficiency due to the lack of traffic lights to allocate right-of-way.
[0003] Currently, intersection risk assessment in the industry mainly relies on micro-traffic flow models and uses kinematic alternatives such as Time to Collision (TTC), Time After Collision (PET), and Deceleration Requirement (DRAC) to conduct quantitative assessments. However, this type of purely numerical assessment method has inherent limitations: it only outputs a one-dimensional scalar to represent the degree of danger, losing key information such as the spatiotemporal topology of the scene and the dynamic evolution of vehicles; the risk causes corresponding to the same PET value are completely different (such as delayed response in blind spots, rapid acceleration / deceleration of vehicles, and extreme following), and the value can only trigger risk warnings but cannot explain the physical root cause of the danger.
[0004] To compensate for the shortcomings of numerical methods, the industry has begun to introduce natural language processing technology to analyze traffic scenarios, but existing solutions still have two major bottlenecks: First, traditional topic clustering models (such as LDA) are based on bag-of-words models, and the output results are discrete word combinations with weak logical connections. They still require a lot of manual annotation and analysis, and have a low level of intelligence. Second, when describing scenarios manually or analyzing scenarios in early models, it is easy to add inferences about the driver's psychology and subjective intentions, such as "attempting to cut in, being distracted, or hesitating," which violates the objectivity required for autonomous driving engineering analysis, and the analysis results cannot be directly used for autonomous driving algorithm iteration.
[0005] In summary, existing technologies cannot simultaneously address the four key requirements of quantifying risk, semantic interpretation, objectivity and neutrality, and automatic clustering. Therefore, a new technical solution is urgently needed to transform vehicle trajectory data into purely kinematic objective semantic text, and combine it with a large language model to automatically mine, cluster, and trace the source of risk topics, thereby achieving high-precision and highly interpretable automatic assessment of risks at unsignalized intersections. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention provides the following technical solution: This solution provides a risk assessment method for unsignaled intersections based on semantic analysis, comprising the following steps: S1. Use roadside detection equipment or drone aerial photography to obtain high-precision datasets of vehicle interaction trajectory data at unsignalized intersections. Calculate and filter the interaction trajectories based on the post-collision time PET threshold to extract numerical data of dangerous scenarios at unsignalized intersections. S2. Set the system roles and prompt word constraint rules of the large language model, and transform the extracted numerical data of dangerous scenes into natural language text that is a pure kinematic objective description, and construct a dangerous scene text set. S3. Extract a predetermined number of sample texts from the dangerous scenario text set, and use the context learning capability of the large language model to mine potential risks in the sample texts to obtain an initial risk topic set. S4. Vectorize the initial risk topics, calculate the cosine similarity between the vectors, use a large language model to semantically merge the highly similar topics, and remove low-frequency marginal topics to obtain the optimized standard risk topic set. S5. Input the optimized standard risk topic set as candidate labels into the large language model. Prompt the large language model to evaluate all scenario texts in the dangerous scenario text set, assign the most relevant risk topic to each scenario text, and extract the original text fragments that support the assignment result from the corresponding scenario text as the basis. The output format of the large language model must strictly follow the triple structure: [topic label]: [topic description] (quoted original text: "..."). If the output format of the large language model is incorrect or a label that does not exist in the standard risk topic set is assigned, the self-correction mechanism is triggered, and the error type and the original text are re-input into the model for secondary assignment. S6. Calculate the frequency and proportion of each standard risk subject at the specified unsignaled intersection, construct the unsignaled intersection risk assessment matrix and generate a safety risk assessment report.
[0007] Preferred technical solution one: In step S1, the extraction of numerical data on hazardous scenarios at unsignalized intersections based on the post-collision time (PET) index includes: (1) Identify any two vehicles with spatial interaction trajectories within an unsignalized intersection. and Extract its time point Micro kinematic state vector ,in These are the vehicle's length and width, respectively; (2) Based on the geometric center coordinates of the vehicle and heading angle Construct a dynamic rectangular projection surface of the vehicle in a two-dimensional plane. The two vehicles were within the preset observation time window. The intersection of the sweeping trajectories within the area constitutes the conflict zone. Its expression is:
[0008] make and Representing vehicles Distance from conflict zone during driving The distance to the boundary and the distance to the boundary. (3) Extract the vehicles that have passed through first. The absolute moment when the tail completely crossed the boundary of the conflict zone :
[0009] After extraction, through the vehicle The absolute moment when the leading end reaches the boundary of the conflict zone. : ; (4) Based on the traffic flow conflict theory, the post-collision time PET is calculated using the following formula: ; (5) Set a high-risk interaction time threshold Filter out those that meet the requirements The interaction sequence is used as numerical data for hazardous scenarios.
[0010] Preferred technical solution two: In step S2, the conversion of the natural language text into a purely kinematic objective description must follow the following refined instruction constraint system: (1) Cognitive isolation and the rule prohibiting subjective inference When constructing the system instructions for the large language model, a "subjective filter" module is preset, which forces the large language model to strictly prohibit the use of words that characterize the driver's psychological state, anticipated intentions, or non-observable motivations when analyzing the trajectory. The prohibited words include at least: "intention", "want", "attempt", "distracted", and "hesitate". All action descriptions must correspond to specific changes in physical quantities, such as replacing "attempt to overtake" with "increased longitudinal acceleration accompanied by left lateral displacement". (3) Semantic mapping matrix of kinematic parameters The text description must contain structured kinematic information in the following five dimensions: (2.1) Spatial entry topology: describes the absolute entry position of a vehicle relative to the center point of the intersection (e.g., entering from the south entrance, entering from the northwest entrance). (2.2) Instantaneous dynamic state: describes the instantaneous velocity of the vehicle before and after the critical conflict point. Tangential acceleration and centripetal acceleration ; (2.3) Trajectory geometric evolution: describing the sequence of vehicle center point coordinates The change in the radius of curvature formed by it is used to characterize the turning, straight-going or avoidance path; (2.4) Spatiotemporal occupancy derivative: describes the rate of change of vehicle speed, used to characterize braking intensity or acceleration urgency; (2.5) Relative kinematic interaction sequence: describes the dynamic Euclidean distance between the interacting vehicle pairs. Relative velocity And the rate of change of azimuth angle.
[0011] Preferred technical solution three: In step S4, the initial risk topic is vectorized, the cosine similarity between vectors is calculated, a large language model is used to semantically merge highly similar topics, and low-frequency marginal topics are removed, including: (1) Use a text vectorization model (such as Sentence-Transformer) to generate the initial risk topic set from the large language model. Mapped to A high-dimensional dense set of semantic vectors A single topic vector is represented as ; (2) For any two initial risk themes and The semantic vector is calculated using the vector dot product and vector norm. and Cosine similarity between The calculation formula is expanded as follows: ; (3) Set a topic similarity threshold When judged At that time, the topic was applied to The data was determined to be highly semantically overlapping, and its joint input to the large language model was merged into a single new topic encompassing the features of both. Let the total number of non-redundant topics after merging and optimization be . Introducing Indicator Functions When the first Sample text for each scenario Generate topics The value is 1 if the condition is met, and 0 otherwise. The combined result is the [number]th [item]. Global frequency of occurrence of each topic : ; (4) Set the rejection threshold ,like If so, the incidental marginal topic is removed from the set, ultimately forming a standard risk topic set containing the core causes of the intersection.
[0012] Preferred technical solution four: In step S6, the statistical analysis of the distribution frequency and proportion of each standard risk subject at a specified unsignalized intersection, the construction of an unsignalized intersection risk assessment matrix, and the generation of a safety risk assessment report include: (1) Suppose that the final set of standard risk topics determined after optimization and elimination includes One valid risk theme, denoted as ; (2) Statistics on current unsignalized intersections In the dangerous scenario text, it was assigned to the first by the large language model. One risk theme Total number of scenes Calculate the first The probability of occurrence of each risk theme : ; (3) Extracting topics assigned to the topic All scene clips Extreme value set , (4) Introduce an average severity index based on exponential decay. The formula for quantifying the intensity of conflict under this semantic topic is as follows: ; (5) Combining the probability of occurrence Severity Index Calculate the first The overall risk coefficient of each risk subject
[0013]
[0014] in, and Assign weights to the preset probability weights and severity, and
[0015] (6) Based on the calculated indicators, construct a multi-dimensional risk assessment matrix for the unsignalized intersection. :
[0016] Extracting the matrix The main risk characteristics are based on the comprehensive risk coefficient. Arrange the data in descending order to generate a comprehensive safety risk assessment report for the intersection.
[0017] A semantic analysis-based risk assessment method for unsignalized intersections also includes an assessment system, which comprises: roadside / drone trajectory acquisition equipment, a data processing and PET conflict calculation server, a large language model semantic analysis engine, a risk topic clustering database, and a terminal assessment report visualization system.
[0018] The present invention proposes a risk assessment method for unsignaled intersections based on semantic analysis. The beneficial effects achieved by using the above structure are as follows: (1) This invention transforms high-dimensional trajectory data into pure objective kinematic semantic text and combines the topic mining and classification capabilities of large language models to achieve dimensionality reduction analysis and natural language clustering of dangerous scenarios at unsignalized intersections. It can directly and clearly reveal the deep physical and spatial causes that trigger traffic conflicts, and provide highly interpretable decision-making basis for right-of-way optimization and autonomous driving strategy improvement. (2) This invention designs a rigorous prompt word constraint system, which forces the large language model to perform pure kinematic description based solely on the input numerical trajectory, and strictly prohibits the use of any words that represent the driver's psychological state or subjective intention. At the same time, it introduces a mechanism for extracting original text fragments and forcing format verification in the topic allocation stage, so that the classification results of each risk topic are verifiable. It eliminates the unavoidable subjective assumptions in traditional manual analysis from the source, and also effectively suppresses the "illusion" output that generative artificial intelligence may produce, significantly improving the engineering reliability and traceability of risk assessment results. (3) End-to-end risk analysis from numerical trajectory to semantic topic can be completed solely through context learning and prompt word engineering. Compared with supervised learning methods that require a large amount of labeled data and computing resources, this invention has a significantly lower implementation threshold and deployment cost, and can be easily transferred to unsignalized intersections with different structures and traffic flow characteristics, demonstrating excellent scenario generalization ability; (4) This invention introduces a vector semantic similarity calculation and a large language model-assisted merging mechanism, which can automatically identify and merge initial topics with highly overlapping semantics, while eliminating low-frequency and occasional topics to form a set of standard risk topics with high cohesion and strong mutual exclusivity. On this basis, a multi-dimensional risk assessment matrix is constructed by combining the occurrence probability of each topic with the severity of PET (Predicted Risk Level) quantification, realizing the calculation of comprehensive risk coefficients that combine qualitative and quantitative methods, and providing an objective numerical basis for prioritizing safety hazards at intersections. Attached Figure Description
[0019] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1This is a flowchart illustrating the overall control process of a semantic analysis-based risk assessment method for unsignaled intersections proposed in this invention. Figure 2 This is a schematic diagram illustrating the interaction between the semantic clustering and evidence tracing and allocation modules of the Large Language Model (LLM) for a risk assessment method for unsignaled intersections based on semantic analysis proposed in this invention. Detailed Implementation
[0020] The technical inventions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0021] It should be noted that the terms “front,” “back,” “left,” “right,” “up,” and “down” used in the following description refer to the directions shown in the attached diagram, while the terms “inside” and “outside” refer to the directions toward or away from the geometric center of a specific component, respectively.
[0022] Example 1 like Figures 1-2 As shown, the technical invention adopted in this invention is as follows: a risk assessment method for unsignaled intersections based on semantic analysis, the core of which lies in converting kinematic values into objective semantic text and utilizing the topic clustering capability of a large language model for risk mining, specifically including the following detailed implementation steps: S1. Extract numerical data of hazardous scenarios at unsignalized intersections based on the post-collision time (PET) index; Using open-source traffic datasets from roadside cameras, millimeter-wave radar, or drones, identify any two vehicles with spatially interacting trajectories within an unsignalized intersection. and Extract its time point Micro kinematic state vector ,in These are the length and width of the vehicle, respectively.
[0023] Based on the vehicle's geometric center coordinates and heading angle Construct a dynamic rectangular projection surface of the vehicle in a two-dimensional plane. The two vehicles were within the preset observation time window. The intersection of the sweeping trajectories within the area constitutes the conflict zone. Its expression is:
[0024] make and Representing vehicles Distance from conflict zone during driving The distance to the boundary and the distance to the boundary.
[0025] Extract the vehicles that have passed through first The absolute moment when the tail completely crossed the boundary of the conflict zone :
[0026] After extraction, through the vehicle The absolute moment when the leading end reaches the boundary of the conflict zone. :
[0027] According to traffic flow conflict theory, the post-collision time PET is calculated using the following formula:
[0028] Set high-risk interaction time threshold Filter out those that meet the requirements The interaction sequence, and extract the intervals before and after the extreme value moment. The trajectory characteristics are used as numerical data for dangerous scenarios.
[0029] S2. Pure kinematic objective natural language transformation based on expert role guidance; Traditional numerical matrices cannot directly trigger the deep reasoning capabilities of Large Language Models (LLMs). This step constructs a cross-modal "numerical-text" conversion mechanism. By setting strict system roles and mapping rules, the extracted numerical data of dangerous scenarios are transformed into objective technical descriptions that are easy for Large Language Models to parse. The specific implementation mechanism is as follows: (1) Semantic discretization mapping of kinematic features: Extract trajectory slices in S1, downsample their key frames (such as braking start point and PET minimum point), and construct a "kinematic-natural language" discretization mapping dictionary.
[0030] (2) Spatial topology mapping: Based on the high-precision map of the intersection, the two-dimensional coordinate sequence is mapped.
[0031] Transform into lane-level topology description (e.g., "vehicle") Vehicles enter along the inner lane of the south entrance. "Entering the conflict zone along the eastern entrance road"
[0032] (3) Longitudinal dynamic mapping: Calculate the first derivative of velocity (acceleration) ), set the threshold dictionary. If Mapped to "significant deceleration"; if Mapped to "emergency braking triggered".
[0033] (4) Lateral avoidance mapping: Calculate the rate of change of vehicle heading angle If within a single frame This is mapped to "perform a sharp turn to avoid an obstacle".
[0034] (5) Interactive limit mapping: Extract the minimum Euclidean distance between the bounding boxes of the two vehicle trajectories and the corresponding time. Generate an interactive convergence description.
[0035] The discretized mapped text fragments are concatenated and input into a large language model for sentence smoothing. To ensure patent-level data objectivity, the system constructs a three-layer Prompt template with "cognitive isolation" functionality: a system instruction layer; subjective inference prohibition rules; and a format assembly layer.
[0036] S3. Incremental generation of risk topics based on iterative prompts and contextual learning; This step employs an improved TopicGPT topic generation architecture that does not rely on the full dataset. Instead, it leverages the in-context learning capabilities of a large language model for incremental topic mining at the sample level, completely eliminating the need for pre-setting the number of topics required by traditional LDA topic models. Furthermore, it faces the challenge of being unable to output natural language descriptions. The specific implementation logic is as follows: (1) Dynamic sample sampling and topic initialization: Considering that large-scale corpus input to a large language model will lead to token truncation and high computing power costs, the system starts from the text set Uniform random sampling is used to construct a system containing... A sample set of generated scenario texts is created. Simultaneously, a dynamic, expert-constructed library of few-shot example topics is initialized in system memory. Each topic includes a "hierarchical number, a short label, and a one-sentence description".
[0037] (2) Incremental topic mining based on iterative prompts: Traversing each scene text in the sample set The system will display the current text. With dynamically updated theme libraries The data is simultaneously input into a large language model, and structured discrimination and generation instructions are issued.
[0038] (3) Dynamic update and shutdown criteria for the topic pool: If the language model targets text New risk themes were output. The system automatically parses the output and appends it to the current topic library, i.e., updates it. The updated theme library will be the next text. Prior contextual knowledge during analysis. As the number of documents traversed increases, the discovery rate of new topics decreases exponentially. The system monitors the generation rate of new topics in real time, and when processing continuous... Scene text (such as) When all large language models output None (i.e., no new topics were found), the topic space is considered saturated, the traversal and generation loop is terminated early, and the current initial set of risk topics is output. This approach significantly reduces computational overhead while ensuring the completeness of the data mining.
[0039] S4. Risk topic semantic merging and order reduction optimization based on vector feature embedding; Because the initial topics generated by the large language model in stage S3 may have highly overlapping semantics (for example, "non-motorized vehicle intrusion into the path" and "two-wheeled vehicle peeking out" describe similar risk causes), noise reduction optimization must be performed through high-dimensional vector computation and secondary inference. This requires using a text vectorization model (such as Sentence-Transformer) to transform the initial risk topic set... Mapped to A high-dimensional dense set of semantic vectors Each topic tag and its descriptive text are transformed into a coordinate point in a vector space, and its coordinate position reflects the semantic connotation of the risk scenario.
[0040] For vector sets Any two topic vectors in and Calculate the cosine similarity between the two. :
[0041] The system automatically filters out (e.g., threshold) Topic pairs are identified as semantic redundancies and require a secondary merging process. The filtered redundant topic pairs, along with their detailed descriptions, are input again into the large language model, accompanied by a dedicated "merging prompt." The frequency of each topic in the sample set generation process is then statistically analyzed. Introduce a rejection threshold. ,like If a topic is deemed an isolated, marginal scenario and lacks systematic evaluation value, it is removed from the topic pool. This ultimately results in a highly cohesive and mutually exclusive set of standard risk topics. .
[0042] S5. Automatic allocation and source tracing of large language models based on semantic alignment; After determining the standard topic pool, all hazardous scenario texts need to be included. Accurately map to the corresponding risk categories and ensure the classification process is "auditable". As a candidate tag list, the large language model is guided to assign tags in a single scenario: Due to the risk of probabilistic fluctuations in generative models, the system has a built-in "format validator" and "tag validator". If the large language model outputs a tag that is not... If a false label is detected (i.e., a "hallucination" is generated), the system will automatically detect the anomaly and send an error message back to the model for retry.
[0043] S6. This step transforms qualitative text clustering into quantitative risk level indicators by quantifying the semantic allocation results, thereby achieving an accurate assessment of the safety situation at unsignalized intersections. (1) Statistical analysis of risk themes for each standard Total number of scene allocations And extract the corresponding PET mean. The reliability of the assessment results was verified by comparing the consistency between the "semantic severity description" and the "numerical conflict intensity (PET)".
[0044] (2) Calculation of comprehensive risk coefficient: Calculate the first The probability of occurrence of each risk theme Severity Index : Frequency percentage:
[0045] Severity Index:
[0046] Overall risk score: ( (For custom weight parameters) Risk assessment matrix With report output: Constructing a multi-dimensional risk assessment matrix for unsignalized intersections:
[0047] (3) Based on the matrix The scores prioritize the potential hazards at the intersection. For example, if the score for "[3] Blind spot left turn conflict" is the highest, the system will automatically extract the corresponding "evidence text" from step S5 as a case and generate an in-depth risk assessment report that includes explanations of causes, typical evidence, numerical strength, and improvement suggestions.
[0048] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, material, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, material, or apparatus.
[0049] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A risk assessment method for unsignalized intersections based on semantic analysis, characterized in that, Includes the following steps: S1. Use roadside detection equipment or drone aerial photography to obtain high-precision datasets of vehicle interaction trajectory data at unsignalized intersections. Calculate and filter the interaction trajectories based on the post-collision time (PET) threshold to extract numerical data of dangerous scenarios at unsignalized intersections. S2. Set the system role of the large language model as a trajectory semantic analysis expert, design structured prompt word templates, and transform the numerical data of dangerous scenes containing coordinates, speed and relative distance extracted in S1 into natural language text with pure kinematic objective description, and construct a dangerous scene text set. S3. Extract a preset number of sample texts from the dangerous scene text set. Based on the context learning capability of the large language model, input a small number of sample example topics and prompt the large language model to generate potential risk topics for the sample texts to obtain an initial risk topic set. S4. Call the text vectorization model to embed features into the initial risk topic set, calculate the cosine similarity between topics, use the large language model to semantically merge topic pairs with similarity greater than a preset threshold, and count the generation frequency of the merged topics. Remove topics with a frequency lower than the threshold to obtain the optimized standard risk topic set. S5. Input the optimized standard risk topic set as candidate labels into the big language model, prompting the big language model to evaluate all scenario texts in the dangerous scenario text set, assign the most relevant risk topic to each scenario text, and extract the original text fragments that support the assignment result from the corresponding scenario text as a basis. S6. Calculate the frequency and proportion of each standard risk subject at the specified unsignaled intersection, construct the unsignaled intersection risk assessment matrix, and generate a safety risk assessment report.
2. The method for risk assessment of unsignaled intersections based on semantic analysis according to claim 1, characterized in that, In step S1, the extraction of numerical data for hazardous scenarios at unsignalized intersections, based on a preset trajectory interaction conflict detection algorithm, specifically includes: (1) Extract any two vehicles with spatial trajectory intersections within an unsignalized intersection. and The kinematic data matrix, including time steps Center coordinates Vehicle dimensions and instantaneous speed ; (2) Based on the geometric dimensions and driving trajectories of the two vehicles, calculate the intersecting polygons on the two-dimensional plane and mark them as the conflict area. ; (3) Based on the kinematic state, calculate the positions of the two vehicles in the conflict zone. The spatiotemporal boundary conditions are used to extract the vehicles that have passed through first. The rear of the vehicle completely drove out of the conflict zone. The moment of the boundary and the vehicles that passed through afterwards The front end reaches the conflict zone The moment of the boundary ; (4) Calculate the post-collision time PET using the following formula: ; (5) Set a high-risk interaction time threshold Filter the interaction trajectory segments; when an interactive vehicle is detected that meets the requirements... When the interaction is deemed a dangerous scenario, preset time intervals (e.g., before and after the specified time) are extracted. The pure kinematic trajectory parameter sequence of ) is used as numerical data for dangerous scenarios.
3. The method for risk assessment of unsignaled intersections based on semantic analysis according to claim 1, characterized in that, In step S2, the transformation into natural language text that is a purely kinematic objective description must follow the following refined instruction constraint system: (1) Cognitive isolation and the rule prohibiting subjective inference When constructing the system instructions for the large language model, a "subjective filter" module is preset, which mandates that the large language model, when analyzing the trajectory, strictly prohibits the use of words that characterize the driver's psychological state, anticipated intentions, or non-observable motivations; the prohibited words include at least: "intention," "want," "attempt," "distracted," and "hesitant"; all action descriptions must correspond to specific changes in physical quantities, such as replacing "attempt to overtake" with "increased longitudinal acceleration accompanied by left lateral displacement"; (2) Semantic mapping matrix of kinematic parameters The text description must contain structured kinematic information in the following five dimensions: (2.1) Spatial entry topology: describes the absolute entry position of a vehicle relative to the center point of the intersection (e.g., entering from the south entrance, entering from the northwest entrance). (2.2) Instantaneous dynamic state: describes the instantaneous velocity of the vehicle before and after the critical conflict point. Tangential acceleration and centripetal acceleration ; (2.3) Trajectory geometric evolution: describing the sequence of vehicle center point coordinates The change in the radius of curvature formed by it is used to characterize the turning, straight-going or avoidance path; (2.4) Spatiotemporal occupancy derivative: describes the rate of change of vehicle speed, used to characterize braking intensity or acceleration urgency; (2.5) Relative kinematic interaction sequence: describes the dynamic Euclidean distance between the interacting vehicle pairs. Relative velocity And the rate of change of azimuth angle.
4. The method for risk assessment of unsignaled intersections based on semantic analysis according to claim 1, characterized in that, In step S4, the feature embedding and merging optimization of the initial risk topic set specifically includes: (1) Use the Sentence-Transformer model to convert the initial risk subject set Mapped to high-dimensional semantic vectors ; (2) Calculate any two topic vectors and Cosine similarity between : (3) When At that time, the corresponding topic will be... Input a large language model and prompt it to merge two topics into a new topic while preserving the core semantics. ; (4) Calculate the merged theme Frequency of occurrence Set a rejection threshold ,when When that happens, remove the topic from the collection.
5. The method for risk assessment of unsignaled intersections based on semantic analysis according to claim 1, characterized in that, In step S5, the most relevant risk topics are assigned and fragments are extracted. The output format of the large language model must strictly follow the triple structure: [topic tag]: [topic description]. If the output format of the large language model is incorrect or a tag that does not exist in the standard risk topic set is assigned, a self-correction mechanism is triggered, and the error type and the original text are re-input into the model for secondary assignment.
6. The method for risk assessment of unsignaled intersections based on semantic analysis according to claim 1, characterized in that, In step S6, the statistical analysis of the frequency and proportion of each standard risk subject at a specified unsignalized intersection, the construction of an unsignalized intersection risk assessment matrix, and the generation of a safety risk assessment report include: (1) Suppose that the final set of standard risk topics determined after optimization and elimination includes One valid risk theme, denoted as ; (2) Statistics on current unsignalized intersections In the dangerous scenario text, it was assigned to the first by the large language model. One risk theme Total number of scenes Calculate the first The probability of occurrence of each risk theme : ; (3) Extracting topics assigned to the topic All scene clips Extreme value set , (4) Introduce an average severity index based on exponential decay. The formula for quantifying the intensity of conflict under this semantic topic is as follows: ; (5) Combining the probability of occurrence Severity Index Calculate the first The overall risk coefficient of each risk subject in, and Assign weights to the preset probability weights and severity, and (6) Based on the calculated indicators, construct a multi-dimensional risk assessment matrix for the unsignalized intersection. : Extracting the matrix The main risk characteristics are based on the comprehensive risk coefficient. Arrange the data in descending order to generate a comprehensive safety risk assessment report for the intersection.