Intelligent analysis and responsibility identification method for traffic accidents based on multi-modal language model

CN122821501APending Publication Date: 2026-09-25周品希
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611136101.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

车辆、驾驶人、地点、时间等关键要素之间的潜在关联无法被有效挖掘,历史事故数据的价值难以充分发挥

Benefits of technology

(1)在本发明中,采用自适应帧采样策略从行车记录仪视频中提取关键帧序列,通过Base64编码将视频帧转换为适合的数据格式。利用视觉语言模型,理解交通事故现场视频中的语义信息,自动生成包含车辆行驶轨迹、道路环境、事故发生过程等关键信息的结构化描述文本;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821501A_ABST
    Figure CN122821501A_ABST
Patent Text Reader

Abstract

The application discloses a kind of traffic accident intelligent analysis and responsibility identification method based on multi-modal language model, belong to the technical field of traffic video image processing;Adaptive frame sampling strategy is used to extract key frame sequence from traffic accident video, and video frame is converted into suitable data format by Base64 encoding. Utilize visual language model, understand the semantic information in traffic accident scene video, automatically generate accident structured description text. Secondly, the application constructs the traffic regulations knowledge base and reasoning engine based on deep learning. The system carries out structured processing to legal texts such as "Road Traffic Safety Law Implementation Regulations", and establishes the mapping relationship between regulations clauses and accident situations. Through the deep semantic understanding ability of large language model, the system can accurately identify illegal behavior in accident description, match the corresponding legal clauses, and conduct responsibility reasoning based on legal logic, to ensure the explainability and accuracy of responsibility identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of traffic video image processing, and in particular to a method for intelligent analysis and liability determination of traffic accidents based on a multimodal language model. Background Technology

[0002] The primary technical bottleneck currently facing traffic accident handling lies in the insufficient intelligent analysis capabilities of video evidence. Although dashcams can objectively record the entire accident process, the video data itself is unstructured information, making it difficult for computers to automatically and accurately understand the dynamic interactions in complex traffic scenarios. As a result, a large amount of evidence, including dashcam recordings and surveillance videos, can only be analyzed through subjective human judgment, lacking automated methods for analyzing traffic accident processes.

[0003] Secondly, the accident liability determination process lacks intelligent support. Current liability determinations primarily rely on the professional experience and subjective judgment of traffic police, involving manual comparison and analysis based on the Road Traffic Safety Law and its implementing regulations. This model is not only time-consuming and labor-intensive but also prone to inconsistencies in determinations due to differing interpretations among multiple parties. Although the relevant legal provisions are clear, accurately mapping specific accident scenarios to abstract legal clauses remains a major challenge in practical application. This is especially true when multiple parties are involved or when road conditions are complex, as manual analysis is prone to omissions or errors.

[0004] Furthermore, a structured storage and knowledge association mechanism for accident information has not yet been established. Currently, accident data exists mostly as independent records, lacking systematic correlation analysis capabilities. Potential correlations between key elements such as vehicles, drivers, locations, and times cannot be effectively mined, and the value of historical accident data cannot be fully realized. This information silo phenomenon severely restricts the development of advanced applications such as accident prevention and risk assessment. Summary of the Invention

[0005] The purpose of this invention is to provide a method for intelligent analysis and liability determination of traffic accidents based on a multimodal language model. By deeply integrating video understanding technology, natural language processing, and knowledge graph construction technology, it achieves automatic parsing of traffic accident videos, structured description of the accident process, and intelligent determination of liability. The core innovation of this method lies in constructing an end-to-end accident analysis process that can automatically extract key information from raw video data, perform reasoning analysis in conjunction with a traffic regulations knowledge base, and generate a standardized liability determination report.

[0006] To achieve the above objectives, this invention provides a method for intelligent analysis and liability determination of traffic accidents based on a multimodal language model, comprising the following steps: Step 1: Preprocess and encode the traffic accident video based on adaptive frame sampling to obtain the corresponding frame sequence of the traffic accident video; Step 2: Analyze the frame sequence obtained in Step 1 using the Visual Language Model (VLM) to form a structured description of the accident process; Step 3: Store the structured accident description process and traffic accident video together, and build the index based on the relational database SQLite; Step 4: Establish a searchable traffic law knowledge base to generate corresponding legal knowledge based on the accident description process; Step 5: Based on the existing large language model, integrate the accident description and corresponding legal knowledge to generate a structured liability determination report.

[0007] Preferably, the process of step one is as follows: For the input video We use the OpenCV library to extract basic metadata from the video. The video reading function is defined as follows: ; In the above formula, This indicates the video reading object created for the input video. This refers to a function in the OpenCV library used to open an input video and return a video read object, obtaining the raw frame rate of the video. as follows: ; In the above formula, This represents the number of frames per second in the video; the target sampling rate is set according to the needs of accident analysis. And calculate the frame interval accordingly. : ; In the above formula, This represents the floor function. The interval between frames for extracting one image frame was determined. Perform keyframe extraction and initialize the frame counter. and empty frame sequence set For the input video Perform a frame-by-frame traversal, and for frames that meet the sampling conditions... Perform JPEG compression encoding; the encoding function is as follows: ; The encoding function returns the encoding status. and byte buffer Then, convert the binary stream to a binary byte stream, and then to Base64 text format. The Base64 encoding function is defined as follows: ; in, This represents the UTF-8 text string obtained after JPEG compression and Base64 encoding of the j-th sampled frame. This indicates that the byte stream is encoded as a Base64 byte sequence. Convert it to a UTF-8 string, which corresponds to the frame sequence representation of the traffic accident video: ; in The total number of frames extracted is calculated using the following formula: ; in The total video duration is in seconds. Enter the video... With frame sequence Store them together to complete the preprocessing stage, and you will get: ; In the above formula This is used to output the frame sequence of a traffic accident video.

[0008] Preferably, the process of step two is as follows: Multimodal input message construction is formed based on the extracted frame sequence. To construct an image message sequence that meets the input requirements of the Visual Language Model (VLM), the corresponding image message generation function is as follows: ; The complete image message sequence is constructed from m image messages as follows: ; Design a special prompt template Generate a structured accident description. The requirements for the prompt word template include: Task description: Please describe in detail the process of the traffic accident recorded in the video; Output requirements: List the key events in chronological order as a process sequence; Format specifications: Each process is labeled with a sequence number, describing vehicle behavior, road conditions and collision details. Corresponding text message definition: ; The complete user message content is constructed by combining image messages and text messages. This user message content refers to the set of multimodal input data submitted by the user role in a single visual language model inference request. The multimodal input data set includes a sequence of image messages arranged chronologically according to the video's timeline. And the TextMsg text message used to define the traffic accident analysis task and output format: ; The resulting message structure is as follows: ; Call the visual language model for reasoning, and set... The parameters specify the VLM model to use, the input is a message structure, streaming is used, and incremental output is achieved. Each All are limited to a maximum output length of 2048 tokens; based on the streaming method used, incremental blocks are used. The response is returned in a specific format, and the streaming output aggregation function is defined as follows: ; in This represents a string concatenation operation. Total number of blocks; final generated incident description It contains the following structured information: ; Each of them It includes the accident process number, description of the vehicle involved, driving behavior, road location, and description of abnormal events.

[0009] Preferably, the process of step three is as follows: First, define the table structure in the database, including the accident information table. as follows: ; The fields are defined as follows: Primary Key Each accident record is uniquely identified using an integer auto-incrementing method; license plate number The text type is text, with a non-null constraint; the time of the incident is "time", the text type is in ISO 8601 format, and has a non-null constraint; the location of the incident is... Text type, non-null constraint; detailed accident description Text type, storing the data generated in step two. ; Enter the video address A text type used to associate with the original video; records the creation timestamp `created_at`, which defaults to the current time. To insert a new accident record, define an insert function: ; in The data tuple to be inserted: ; The above formula represents a newly inserted accident. A corresponding query function based on the license plate number should be designed: ; In the above formula, This function retrieves license plate numbers, and the results are sorted in descending order of creation time, prioritizing the most recent records.

[0010] Preferably, step four is as follows: Let the original regulatory text be... The total number of characters is To segment a long text into several semantically complete blocks, first define a chapter boundary detection function. Matched using the following regular expression: ; The following is a set of extracted chapter start positions: ; Where c is the number of chapters, and the text is divided into blocks according to the starting position of each chapter: ; Each block Corresponding to a complete chapter; vectorizing legal knowledge, for each text block The vector representation is generated using a pre-trained text embedding model, as shown in the following formula: ; in For the embedding dimension, a corresponding vector index of the regulatory knowledge base is constructed: ; For accident inquiry Generate query vectors: ; Calculate cosine similarity: ; The following are the relevant Top-K regulations: ; Context-enhanced regulatory text return for the retrieved Top-K regulatory blocks The combination methods are as follows: ; in This indicates a text concatenation operation, which uses the combined text as legal knowledge for reasoning in a large language model.

[0011] Preferably, step five proceeds as follows: Integrating accident description, legal knowledge, and large language model reasoning capabilities to generate a structured liability determination report, first defining the liability determination function: ; In the above formula, This refers to the accident description in step two; This refers to the legal knowledge in step four; This section provides task guidance prompts and includes output format requirements: analysis of the parties responsible for the accident; primary responsible party and reasons; secondary responsible party and reasons; proportion of responsibility; relevant legal basis; and handling recommendations. Set constraints on the proportion of responsibility allocation: ; This ultimately forms the complete LLM input message sequence: ; Final model output It contains the following structured information: ; Legal basis for model output Cross-validate with the legal knowledge base built in step 4: mark clauses that fail verification as "[to be confirmed]" to ensure the accuracy of legal citations.

[0012] Therefore, the present invention employs the above-mentioned intelligent analysis and liability determination method for traffic accidents based on a multimodal language model, which has the following advantages: (1) In this invention, an adaptive frame sampling strategy is used to extract key frame sequences from dashcam videos, and the video frames are converted into a suitable data format through Base64 encoding. Using a visual language model, the semantic information in the traffic accident scene video is understood, and a structured descriptive text containing key information such as vehicle trajectory, road environment, and accident occurrence process is automatically generated; (2) In this invention, a traffic law knowledge base and reasoning engine based on deep learning are used. The system structures legal texts such as the "Regulations for the Implementation of the Road Traffic Safety Law" and establishes a mapping relationship between legal provisions and accident scenarios. Through the deep semantic understanding capabilities of the large language model, the system can accurately identify illegal behaviors in accident descriptions, match corresponding legal provisions, and perform liability reasoning based on legal logic to ensure the interpretability and accuracy of liability determination; (3) In this invention, a relational database is set up to realize efficient management and related query of accident information. The system stores key information such as license plate number, accident time, location, video path, and accident description in a structured form, and supports fast retrieval based on license plate number and historical accident tracing.

[0013] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating the implementation of the intelligent traffic accident analysis and liability determination method based on a multimodal language model, as described in this invention. Figure 2This is a video analysis setup diagram in an embodiment of the intelligent traffic accident analysis and liability determination method based on a multimodal language model according to the present invention. Figure 3 This is a collision video analysis diagram in an embodiment of the intelligent traffic accident analysis and liability determination method based on a multimodal language model according to the present invention; Figure 4 This is a schematic diagram of the text portion one in an embodiment of the intelligent analysis and liability determination method for traffic accidents based on a multimodal language model of the present invention; Figure 5 This is a schematic diagram of the second text portion in an embodiment of the intelligent analysis and liability determination method for traffic accidents based on a multimodal language model of the present invention; Figure 6 This is a schematic diagram of the text portion three in an embodiment of the intelligent analysis and liability determination method for traffic accidents based on a multimodal language model of the present invention. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Specific model specifications need to be selected and determined according to the actual specifications of the device, etc. The specific selection calculation method adopts existing technology in the art, and therefore will not be described in detail.

[0016] Example like Figure 1 As shown, this invention provides a method for intelligent analysis and liability determination of traffic accidents based on a multimodal language model, including the following steps: Step 1: Preprocess and encode the traffic accident video based on adaptive frame sampling to obtain the corresponding frame sequence of the traffic accident video. The process is as follows: For the input video We use the OpenCV library to extract basic metadata from the video. The video reading function is defined as follows: ; In the above formula, This indicates the video reading object created for the input video. This refers to a function in the OpenCV library used to open an input video and return a video read object, obtaining the raw frame rate of the video. as follows: ; In the above formula, This represents the number of frames per second in the video; the target sampling rate is set according to the needs of accident analysis. And calculate the frame interval accordingly. : ; In the above formula, This represents the floor function. The interval between frames for extracting one image frame was determined. Perform keyframe extraction and initialize the frame counter. and empty frame sequence set For the input video Perform a frame-by-frame traversal, and for frames that meet the sampling conditions... Perform JPEG compression encoding; the encoding function is as follows: ; The encoding function returns the encoding status. and byte buffer Then, convert the binary stream to a binary byte stream, and then to Base64 text format. The Base64 encoding function is defined as follows: ; in, This represents the UTF-8 text string obtained after JPEG compression and Base64 encoding of the j-th sampled frame. This indicates that the byte stream is encoded as a Base64 byte sequence. Convert it to a UTF-8 string, which corresponds to the frame sequence representation of the traffic accident video: ; in The total number of frames extracted is calculated using the following formula: ; in The total video duration (in seconds) is the input video. With frame sequence Store them together to complete the preprocessing stage, and you will get: ; The above formula represents the frame sequence of the output traffic accident video.

[0017] Step Two: Analyze the frame sequence obtained in Step One using a visual language model to form a structured description of the accident process; the process is as follows: construct a multimodal input message based on the extracted frame sequence. To construct an image message sequence that meets the input requirements of the Visual Language Model (VLM), the corresponding image message generation function is as follows: ; The complete image message sequence is constructed from m image messages as follows: ; Design a special prompt template Generate a structured accident description. The requirements for the prompt word template include: Task description: Please describe in detail the process of the traffic accident recorded in the video; Output requirements: List the key events in chronological order as a process sequence; Format specifications: Each process is labeled with a sequence number, describing vehicle behavior, road conditions and collision details. Corresponding text message definition: ; The complete user message content is constructed by combining image messages and text messages. The user message content refers to the set of multimodal input data submitted by the user role in a single visual language model inference request. This multimodal input data set includes a sequence of image messages arranged chronologically according to the video's timeline. And the TextMsg text message used to define the traffic accident analysis task and output format: ; The message structure of the call is as follows: ; Call the visual language model for reasoning, and set... The parameters specify the VLM model to use, which is incorporated into the message structure and streamed for incremental output. Each All are limited to a maximum output length of 2048 tokens; based on the streaming method used, incremental blocks are used. The response is returned in a specific format, and the streaming output aggregation function is defined as follows: ; in This represents a string concatenation operation. Total number of blocks; final generated incident description It contains the following structured information: ; Each of them It includes the accident process number, description of the vehicle involved, driving behavior, road location, and description of abnormal events.

[0018] Step 3: Store the structured accident description process and video data together, using the relational database SQLite for storage and indexing. The process is as follows: First, define the table structure in the database, including the accident information table. as follows: ; The fields are defined as follows: Primary Key Each accident record is uniquely identified using an integer auto-incrementing method; license plate number The text type is text, with a non-null constraint; the time of the incident is "time", the text type is in ISO 8601 format, and has a non-null constraint; the location of the incident is... Text type, non-null constraint; detailed accident description Text type, storing the generated ; Enter the video address A text type used to associate with the original video; records the creation timestamp `created_at`, which defaults to the current time. To insert a new accident record, define an insert function: ; in The data tuple to be inserted: ; The above formula represents a newly inserted accident. Design a query function based on the license plate number: ; In the above formula, This function retrieves license plate numbers, and the results are sorted in descending order of creation time, prioritizing the most recent records.

[0019] Step 4: Establish a searchable traffic regulation knowledge base to generate corresponding regulatory knowledge based on the accident description process; the process is as follows: Assume the original regulatory text is... The total number of characters is To segment a long text into several semantically complete blocks, first define a chapter boundary detection function. Matched using the following regular expression: ; The following is a set of extracted chapter start positions: ; Where c is the number of chapters, and the text is divided into blocks according to the starting position of each chapter: ; Each block Corresponding to a complete chapter; vectorizing legal knowledge, for each text block The vector representation is generated using a pre-trained text embedding model, as shown in the following formula: ; in For the embedding dimension, a corresponding vector index of the regulatory knowledge base is constructed: ; For accident inquiry Generate query vectors: ; Calculate cosine similarity: ; The following are the relevant Top-K regulations: ; Context-enhanced regulatory text return for the retrieved Top-K regulatory blocks The combination methods are as follows: ; in This indicates a text concatenation operation, which uses the combined text as legal knowledge for reasoning in a large language model.

[0020] Step 5: Based on the existing large language model, integrate the accident description and corresponding legal knowledge to generate a structured liability determination report. The process is as follows: Integrate the accident description, legal knowledge, and the reasoning ability of the large language model to generate a structured liability determination report. First, define the liability determination function: ; In the above formula, This refers to the accident description in step two; This refers to the legal knowledge in step four; This section provides task guidance prompts and includes output format requirements: analysis of the parties responsible for the accident; primary responsible party and reasons; secondary responsible party and reasons; proportion of responsibility; relevant legal basis; and handling recommendations. Set constraints on the proportion of responsibility allocation: ; This ultimately forms the complete LLM input message sequence: ; Final model output It contains the following structured information: ; Legal basis for model output Cross-validate with the legal knowledge base built in step 4: mark clauses that fail verification as "[to be confirmed]" to ensure the accuracy of legal citations.

[0021] Here is a specific analysis process: Select a traffic accident video for processing. The video is as follows. Figure 2 and Figure 3 As shown, Figure 2 The analysis frame rate is set to 5 frames per second. Figure 3 That is the final result of the car crash, corresponding to the formation Figures 4 to 6 The text description generates the following structured text: 1. Analysis of the parties responsible for the accident, including the driver of the yellow truck and the driver of the white construction truck; 2. Primary Responsibility and Reasons: The yellow truck bears primary responsibility because, according to the video description, the yellow truck made contact with the white construction truck while adjusting its position. According to Article 47 of the "Regulations for the Implementation of the Road Traffic Safety Law," "When a motor vehicle overtakes, it shall turn on its left turn signal in advance, switch between high and low beam headlights, or sound its horn. The following vehicle shall overtake from the left side of the vehicle in front only after confirming that there is sufficient safe distance." The yellow truck failed to maintain a sufficient safe distance on the slippery road surface and made its position adjustment without ensuring safety.

[0022] 3. Secondary Liability Party and Reasons: The white construction truck bears secondary responsibility; Reason: According to Article 35 of the "Regulations for the Implementation of the Road Traffic Safety Law," "Road maintenance and construction vehicles and machinery shall be equipped with warning lights and painted with conspicuous markings, and shall turn on warning lights and hazard warning flashers when operating." Construction vehicles should take more obvious warning measures when operating on roads, especially in adverse weather conditions.

[0023] 4. The responsibility division ratio is as follows: 70% primary responsibility (yellow truck) and 30% secondary responsibility (white construction truck). 5. The relevant legal basis is as follows: Article 46 of the "Regulations for the Implementation of the Road Traffic Safety Law": When a motor vehicle encounters low visibility weather conditions such as rain, snow, sandstorm, or hail, the maximum driving speed shall not exceed 30 kilometers per hour; Article 47 of the Regulations for the Implementation of the Road Traffic Safety Law: A following vehicle shall overtake a vehicle in front only after confirming that there is sufficient safe distance. Article 35 of the Regulations for the Implementation of the Road Traffic Safety Law: Operating vehicles shall turn on their warning lights and hazard warning flashers; Article 62 of the Regulations for the Implementation of the Road Traffic Safety Law: Drivers of motor vehicles shall not engage in any behavior that impedes safe driving.

[0024] Therefore, the above analysis results prove that the technical solution of the present invention can generate a standardized liability determination report by analyzing video.

[0025] Therefore, this invention employs a traffic accident intelligent analysis and liability determination method based on a multimodal language model. By deeply integrating video understanding technology and natural language processing, it achieves automatic parsing of traffic accident videos, structured description of the accident process, and intelligent liability determination. The core innovation of this method lies in constructing an end-to-end accident analysis process that can automatically extract key information from raw video data, perform reasoning analysis in conjunction with a traffic regulations knowledge base, and generate a standardized liability determination report.

[0026] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for intelligent analysis and liability determination of traffic accidents based on a multimodal language model, characterized in that: Includes the following steps: Step 1: Preprocess and encode the traffic accident video based on adaptive frame sampling to obtain the corresponding frame sequence of the traffic accident video; Step 2: Analyze the frame sequence obtained in Step 1 using the Visual Language Model (VLM) to form a structured description of the accident process; Step 3: Store the structured accident description process and traffic accident video together, and build the index based on the relational database SQLite; Step 4: Establish a searchable traffic law knowledge base to generate corresponding legal knowledge based on the accident description process; Step 5: Based on the existing large language model, integrate the accident description and corresponding legal knowledge to generate a structured liability determination report.

2. The method for intelligent analysis and liability determination of traffic accidents based on a multimodal language model according to claim 1, characterized in that: The process of step one is as follows: For the input video We use the OpenCV library to extract basic metadata from the video. The video reading function is defined as follows: ; In the above formula, This indicates the video reading object created for the input video. This refers to a function in the OpenCV library used to open an input video and return a video read object, obtaining the raw frame rate of the video. as follows: ; In the above formula, This represents the number of frames per second in the video; the target sampling rate is set according to the needs of accident analysis. And calculate the frame interval accordingly. : ; In the above formula, This represents the floor function. The interval between frames for extracting one image frame was determined. Perform keyframe extraction and initialize the frame counter. and empty frame sequence set For the input video Perform a frame-by-frame traversal, and for frames that meet the sampling conditions... Perform JPEG compression encoding; the encoding function is as follows: ; The encoding function returns the encoding status. and byte buffer Then, convert the binary stream to a binary byte stream, and then to Base64 text format. The Base64 encoding function is defined as follows: ; in, This represents the UTF-8 text string obtained after JPEG compression and Base64 encoding of the j-th sampled frame. This indicates that the byte stream is encoded as a Base64 byte sequence. Convert it to a UTF-8 string, which corresponds to the frame sequence representation of the traffic accident video: ; in The total number of frames extracted is calculated using the following formula: ; in The total video duration is in seconds. Enter the video... With frame sequence Store them together to complete the preprocessing stage, and you will get: ; In the above formula This is used to output the frame sequence of a traffic accident video.

3. The intelligent analysis and liability determination method for traffic accidents based on a multimodal language model according to claim 2, characterized in that: The process of step two is as follows: Multimodal input message construction is formed based on the extracted frame sequence. To construct an image message sequence that meets the input requirements of the Visual Language Model (VLM), the corresponding image message generation function is as follows: ; The complete image message sequence is constructed from m image messages as follows: ; Design a special prompt template Generate a structured accident description. The requirements for the prompt word template include: Task description: Please describe in detail the process of the traffic accident recorded in the video. Output requirements: List key events in chronological order as a process sequence; Format specifications: Each process should be numbered and describe vehicle behavior, road conditions, and collision details; Corresponding text message definition: ; The complete user message content is constructed by combining image messages and text messages. This user message content refers to the set of multimodal input data submitted by the user role in a single visual language model inference request. The multimodal input data set includes a sequence of image messages arranged chronologically according to the video's timeline. And the TextMsg text message used to define the traffic accident analysis task and output format: ; The resulting message structure is as follows: ; Call the visual language model for reasoning, and set... The parameters specify the VLM model to use, the input is a message structure, streaming is used, and incremental output is achieved. Each All are limited to a maximum output length of 2048 tokens; based on the streaming method used, incremental blocks are used. The response is returned in a specific format, and the streaming output aggregation function is defined as follows: ; in This represents a string concatenation operation. Total number of blocks; final generated incident description It contains the following structured information: ; Each of them It includes the accident process number, description of the vehicle involved, driving behavior, road location, and description of abnormal events.

4. The intelligent analysis and liability determination method for traffic accidents based on a multimodal language model according to claim 3, characterized in that: The process of step three is as follows: First, define the table structure in the database, including the accident information table. as follows: ; The fields are defined as follows: Primary Key Each accident record is uniquely identified using an integer auto-incrementing method; license plate number The text type is text, with a non-null constraint; the time of the incident is "time", the text type is in ISO 8601 format, and has a non-null constraint; the location of the incident is... Text type, non-null constraint; detailed accident description Text type, storing the data generated in step two. ; Enter the video address A text type used to associate with the original video; records the creation timestamp `created_at`, which defaults to the current time. To insert a new accident record, define an insert function: ; in The data tuple to be inserted: ; The above formula represents the newly inserted accident. A corresponding query function based on the license plate number should be designed: ; In the above formula, This function retrieves license plate numbers, and the results are sorted in descending order of creation time, prioritizing the most recent records.

5. The method for intelligent analysis and liability determination of traffic accidents based on a multimodal language model according to claim 4, characterized in that: The process of step four is as follows: Let the original regulatory text be... The total number of characters is To segment a long text into several semantically complete blocks, first define a chapter boundary detection function. Matched using the following regular expression: ; The following is a set of extracted chapter start positions: ; Where c is the number of chapters, and the text is divided into blocks according to the starting position of each chapter: ; Each block Corresponding to a complete chapter; vectorizing legal knowledge, for each text block The vector representation is generated using a pre-trained text embedding model, as shown in the following formula: ; in For the embedding dimension, a corresponding vector index of the regulatory knowledge base is constructed: ; For accident inquiry Generate query vectors: ; Calculate cosine similarity: ; The following are the relevant regulatory blocks for Top-K: ; Context-enhanced regulatory text return for the retrieved Top-K regulatory blocks The combination methods are as follows: ; in This indicates a text concatenation operation, which uses the combined text as legal knowledge for reasoning in a large language model.

6. The intelligent analysis and liability determination method for traffic accidents based on a multimodal language model according to claim 5, characterized in that: The process of step five is as follows: Integrating accident description, legal knowledge, and large language model reasoning capabilities, a structured liability determination report is generated. First, the liability determination function is defined: ; In the above formula, This refers to the accident description in step two; This refers to the legal knowledge in step four; This section provides task guidance prompts and includes the following output format requirements: analysis of the parties responsible for the accident; primary responsible party and reasons; secondary responsible party and reasons; proportion of responsibility; relevant legal basis; and recommended handling. Set constraints on the proportion of responsibility allocation: ; This ultimately forms the complete LLM input message sequence: ; Final model output It contains the following structured information: ; Legal basis for model output Cross-validate with the legal knowledge base built in step 4: mark clauses that fail verification as "[to be confirmed]" to ensure the accuracy of legal citations.