A traffic accident scene derivation method and system based on unmanned aerial vehicle cooperation and a multi-modal large model, a terminal, and a storage medium

CN122368934BActive Publication Date: 2026-09-08SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610787912.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-09-08
Estimated Expiration
2046-06-03

AI Technical Summary

Technical Problem

[0007]本发明的主要目的在于提供一种基于无人机协同与多模态大模型的交通事故现场推导方法、系统、终端及计算机可读存储介质,旨在解决针对现有交通事故处理技术中存在的三维重建速度慢、针对形变车辆的分割精度低、以及自动责任判定缺乏可解释逻辑链的问题

Benefits of technology

[0018] This invention constructs an air-ground collaborative sensing network. Roadside sensors within this network monitor accidents. When an accident is detected, a drone within the network is automatically dispatched to receive real-time multi-angle image data of the accident scene. A 3D Gaussian sputtering algorithm is used to perform SFM sparse reconstruction, Gaussian ellipsoid initialization, and rasterization rendering on the multi-angle image data to construct a high-fidelity 3D scene of the accident. Multi-granularity segmentation masks and soft-scale gating mechanisms are introduced, combined with 3D cue-based segmentation technology, to extract accident vehicles, road traces, and key objects from the high-fidelity 3D scene. Based on these elements, a traffic accident scene map is constructed. This map is then processed into a text serialization module to construct multimodal input, driving a visual language model to deduce the accident logic chain and output an accident analysis report. This invention improves the real-time performance and high fidelity of 3D accident scene reconstruction, significantly shortens the accident processing cycle, and greatly enhances the objectivity, accuracy, and credibility of liability determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122368934B_ABST
    Figure CN122368934B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of computer vision, and discloses a traffic accident scene derivation method and system based on cooperation of unmanned aerial vehicles and a multimodal large model, a terminal and a storage medium, the method comprising the following steps: when an accident is monitored by a roadside sensor, triggering an unmanned aerial vehicle to be dispatched to collect multi-angle image data of the accident scene in real time; performing SFM sparse reconstruction, Gaussian ellipsoid initialization and rasterization rendering on the multi-angle image data to construct a high-fidelity three-dimensional scene of the accident scene; stripping the accident vehicle, road traces and key objects from the high-fidelity three-dimensional scene; constructing a traffic accident scene graph according to the accident vehicle, the road traces and the key objects, performing text serialization processing on the traffic accident scene graph, constructing a multimodal input, driving a visual language model to derive an accident logic chain according to the multimodal input, and outputting an accident analysis report. The application improves the real-time performance and high fidelity of three-dimensional reconstruction of the accident scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a method, system, terminal, and computer-readable storage medium for deducing traffic accident scenes based on UAV collaboration and multimodal large models. Background Technology

[0002] With the rapid development of intelligent connected transportation systems and drone inspection technology, there is a growing demand for minute-level or even second-level response times for the "pre-accident warning-in-accident handling-post-accident evidence collection" process for road traffic accidents. Traditional handling procedures rely on manual alarms, the arrival of on-duty police officers, and handheld cameras or laser surveying, which suffers from drawbacks such as delayed information acquisition, long on-site closure times, high risk of secondary accidents, and limited evidence dimensions. In recent years, the industry has attempted to introduce new technologies to alleviate these issues, but they still fall short of meeting the comprehensive requirements of high timeliness, comprehensiveness, and high accuracy.

[0003] While traditional MVS (Multi-View Stereo) technology can generate point clouds or mesh models from images captured by drones, its long modeling cycle and high computational cost make it difficult to meet the timeliness requirements of "instant results" at traffic accident scenes. Furthermore, for objects like accident vehicles with highly reflective surfaces or missing textures, the MVS algorithm is prone to producing holes or deformations, leading to insufficient reconstruction accuracy. With the development of computer 3D vision technology, the emergence of NeRF (Neural Radiance Fields) has significantly improved reconstruction accuracy. Although it has improved rendering quality, its training and inference speeds are extremely slow, often requiring several hours to complete the reconstruction of a single scene. Moreover, it struggles to directly support real-time scene editing and semantic interaction, failing to meet the demands of rapid accident scene cleanup.

[0004] In terms of scene segmentation and semantic understanding, traditional segmentation algorithms are often based on geometric features or predefined categories, making it difficult to handle severely deformed vehicles or atypical scattered objects in traffic accidents, resulting in blurred segmentation edges and objects sticking together. In addition, existing algorithms have weak semantic interaction capabilities, mainly passively "identifying" objects and lacking the ability to actively "retrieve" using text prompts or contextual information.

[0005] In the core aspects of accident handling—accident analysis and liability determination—current methods primarily rely on manual investigation or automated algorithms based on simple rules. This approach is not only inefficient and highly subjective, but also prone to overlooking crucial evidence in complex accidents. Traditional automated algorithms are mostly based on object detection in two-dimensional images, lacking a deep understanding of three-dimensional spatial relationships. Even when deep learning models are applied, existing end-to-end solutions are often "black boxes," directly outputting results without an interpretable logical reasoning process. Models cannot generate a complete accident logic chain by combining traffic regulations, vehicle trajectories, and environmental factors, unlike human traffic police, resulting in liability determination reports that lack persuasiveness and reference value.

[0006] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0007] The main objective of this invention is to provide a method, system, terminal, and computer-readable storage medium for deducing traffic accident scenes based on UAV collaboration and multimodal large models. This aims to solve the problems of slow 3D reconstruction speed, low segmentation accuracy for deformed vehicles, and lack of interpretable logical chains for automatic liability determination in existing traffic accident processing technologies.

[0008] To achieve the above objectives, this invention provides a method for deriving traffic accident scene data based on UAV collaboration and a multimodal large model. The method includes the following steps: An air-ground collaborative sensing network is constructed, and accidents are monitored through roadside sensors in the sensing network. When the roadside sensors detect an accident, the drones in the sensing network are automatically dispatched to receive multi-angle image data of the accident scene collected in real time by the drones. Using the 3D Gaussian sputtering algorithm, SFM sparse reconstruction, Gaussian ellipsoid initialization and rasterization rendering are performed on the multi-angle image data to construct a high-fidelity three-dimensional scene of the accident scene. By introducing multi-granularity segmentation masks and soft-scale gating mechanisms, combined with 3D cue-based segmentation technology, accident vehicles, road traces, and key objects are extracted from the high-fidelity 3D scene. Based on the accident vehicle, road traces, and key objects, a traffic accident scene diagram is constructed. The traffic accident scene diagram is then processed into text serialization to construct a multimodal input. This drives a visual language model to deduce the accident logic chain based on the multimodal input and outputs an accident analysis report.

[0009] Optionally, the method for deriving traffic accident scenes based on UAV collaboration and multimodal large models, wherein the construction of an air-ground collaborative perception network, the monitoring of accidents through roadside sensors of the perception network, and the automatic triggering of UAVs in the perception network to schedule and receive multi-angle image data of the accident scene collected in real time by the UAVs when the roadside sensors detect an accident, specifically includes: Construct an air-ground collaborative sensing network, which includes roadside sensors and drones; Accidents are monitored by the roadside sensors. When an accident is detected by the roadside sensors, an accident alarm signal is sent to the intelligent decision-making platform for accident handling. The intelligent decision-making platform for accident handling dispatches the UAV to the accident site to perform data collection tasks, controls the UAV to adopt a multi-altitude layer circling flight strategy, collects multi-angle image data of the core area of ​​the accident, and transmits the multi-angle image data back to the intelligent decision-making platform for accident handling in real time through a high-bandwidth communication link.

[0010] Optionally, in the method for deducing traffic accident scenes based on UAV collaboration and multimodal large models, the multi-angle image data includes multi-angle high-definition image sequences and synchronously recorded POS data containing GPS and IMU attitude information.

[0011] Optionally, the method for deriving traffic accident scenes based on UAV collaboration and multimodal large models, wherein the step of using a 3D Gaussian sputtering algorithm to perform SFM sparse reconstruction, Gaussian ellipsoid initialization, and rasterization rendering on the multi-angle image data to construct a high-fidelity 3D scene of the accident scene specifically includes: The intelligent decision-making platform for accident handling performs SFM sparse reconstruction on the multi-angle image data, estimates the camera pose through feature extraction and matching, generates sparse point clouds, and automatically removes incorrect matches. The sparse point cloud is initialized with sparse point cloud or randomly initialized with Gaussian to obtain the target point cloud. The target point cloud is then densified and pruned with Gaussian ellipsoids to obtain Gaussian ellipsoids. The Gaussian ellipsoid is projected onto a two-dimensional plane and rasterized to obtain a high-fidelity three-dimensional scene of the accident site.

[0012] Optionally, the method for deriving traffic accident scenes based on UAV collaboration and multimodal large models, wherein the Gaussian ellipsoid densification and pruning of the target point cloud specifically includes: During the point densification stage, the density of Gaussians is adaptively increased. For regions with missing geometric features or scattered Gaussians, densification is performed after a certain number of iterations. For cloning small Gaussians, a copy of the Gaussian is created and moved towards the position gradient. For split large Gaussians, two small Gaussians are used to replace one large Gaussian, and the scale is reduced according to a specific factor. During the point pruning phase, redundant Gaussians are removed, Gaussians with transparency below a specified threshold and large Gaussians in world space or view space are eliminated, and the transparency of the Gaussian ellipsoid is set to a preset value after a fixed number of iterations of the Gaussians.

[0013] Optionally, the method for deducing traffic accident scenes based on UAV collaboration and multimodal large models, wherein the introduction of multi-granularity segmentation masks and soft-scale gating mechanisms, combined with 3D cue-based segmentation technology, to extract accident vehicles, road traces, and key objects from the high-fidelity 3D scene, specifically includes: Obtain the multi-view two-dimensional mask extracted by SAM, use the multi-view two-dimensional mask to learn Gaussian affinity features, and by attaching an affinity feature to each three-dimensional Gaussian, the three-dimensional Gaussian has a new attribute for segmentation, thus obtaining the target three-dimensional Gaussian features. A soft-scale gating mechanism is used to project the three-dimensional Gaussian features of the target onto gating feature subspaces of different scales; Given a specific viewpoint, 3D visual cues are converted into corresponding 3D scale-gated query features, and the feature similarity between the 3D scale-gated query features and 3D affinity features is evaluated to segment 3D targets and extract accident vehicles, road traces, and key objects from the high-fidelity 3D scene.

[0014] Optionally, the method for deducing traffic accident scenes based on UAV collaboration and a multimodal large model, wherein the step of constructing a traffic accident scene diagram based on the accident vehicle, the road traces, and the key objects, performing text serialization processing on the traffic accident scene diagram, constructing multimodal input, driving a visual language model to deduce the accident logic chain based on the multimodal input, and outputting an accident analysis report, specifically includes: The accident vehicle, the road traces, and the key objects are transformed into a traffic accident scene graph. The extracted key entities are defined as nodes of the graph, each node is attached with attribute information, and the connection edges between nodes are established according to the relative position of the entities in three-dimensional space to represent spatial or semantic relationships. The traffic accident scene diagram and its corresponding numerical features are converted into a natural language text sequence, and the traffic accident scene diagram is traversed to generate descriptive text; The segmented and extracted 3D model of the accident vehicle is rendered into a high-definition 2D snapshot from multiple evidence collection angles and used as a visual input stream. The natural language text sequence, the descriptive text, and the target text are combined to form a text input stream; The visual input stream and the text input stream are encapsulated into a set of structured multimodal cue words. The multimodal cue words are input into a visual language model. The visual language model deduces the incident logic chain based on the multimodal cue words and outputs an incident analysis report.

[0015] Furthermore, to achieve the above objectives, the present invention also provides a traffic accident scene derivation system based on UAV collaboration and multimodal large model, wherein the traffic accident scene derivation system based on UAV collaboration and multimodal large model includes: The data acquisition module is used to construct an air-ground collaborative sensing network. It monitors accidents through roadside sensors in the sensing network. When the roadside sensors detect an accident, it automatically triggers the scheduling of drones in the sensing network and receives multi-angle image data of the accident scene collected in real time by the drones. The scene construction module is used to perform SFM sparse reconstruction, Gaussian ellipsoid initialization and rasterization rendering on the multi-angle image data using the 3D Gaussian sputtering algorithm to construct a high-fidelity three-dimensional scene of the accident scene. The 3D segmentation module is used to introduce multi-granularity segmentation masks and soft-scale gating mechanisms, combined with 3D cue-based segmentation technology, to extract accident vehicles, road traces and key objects from the high-fidelity 3D scene. The accident deduction module is used to construct a traffic accident scene diagram based on the accident vehicle, the road traces, and the key objects, perform text serialization processing on the traffic accident scene diagram, construct multimodal input, drive the visual language model to deduce the accident logic chain based on the multimodal input, and output an accident analysis report.

[0016] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a traffic accident scene derivation program based on UAV collaboration and multimodal large model stored in the memory and executable on the processor. When the traffic accident scene derivation program based on UAV collaboration and multimodal large model is executed by the processor, it implements the steps of the traffic accident scene derivation method based on UAV collaboration and multimodal large model as described above.

[0017] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a traffic accident scene derivation program based on UAV collaboration and multimodal large model, and when the traffic accident scene derivation program based on UAV collaboration and multimodal large model is executed by a processor, it implements the steps of the traffic accident scene derivation method based on UAV collaboration and multimodal large model as described above.

[0018] This invention constructs an air-ground collaborative sensing network. Roadside sensors within this network monitor accidents. When an accident is detected, a drone within the network is automatically dispatched to receive real-time multi-angle image data of the accident scene. A 3D Gaussian sputtering algorithm is used to perform SFM sparse reconstruction, Gaussian ellipsoid initialization, and rasterization rendering on the multi-angle image data to construct a high-fidelity 3D scene of the accident. Multi-granularity segmentation masks and soft-scale gating mechanisms are introduced, combined with 3D cue-based segmentation technology, to extract accident vehicles, road traces, and key objects from the high-fidelity 3D scene. Based on these elements, a traffic accident scene map is constructed. This map is then processed into a text serialization module to construct multimodal input, driving a visual language model to deduce the accident logic chain and output an accident analysis report. This invention improves the real-time performance and high fidelity of 3D accident scene reconstruction, significantly shortens the accident processing cycle, and greatly enhances the objectivity, accuracy, and credibility of liability determination. Attached Figure Description

[0019] Figure 1 This is a flowchart of a preferred embodiment of the traffic accident scene derivation method based on UAV collaboration and multimodal large model of the present invention; Figure 2 This is a schematic diagram of the entire accident deduction method implementation process in a preferred embodiment of the traffic accident scene deduction method based on UAV collaboration and multimodal large model of the present invention; Figure 3 This is a structural diagram of a preferred embodiment of the traffic accident scene derivation system based on UAV collaboration and multimodal large model of the present invention; Figure 4 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0021] The preferred embodiment of the present invention describes a method for deriving traffic accident scene information based on UAV collaboration and a multimodal large model, such as... Figure 1 and Figure 2 As shown, the method for deriving traffic accident scene data based on UAV collaboration and multimodal large model includes the following steps: Step S10: Construct an air-ground collaborative sensing network. The roadside sensors of the sensing network monitor accidents. When the roadside sensors detect an accident, the drones of the sensing network are automatically dispatched to receive multi-angle image data of the accident scene collected in real time by the drones.

[0022] Specifically, an air-ground collaborative sensing network is constructed, which includes roadside sensors and drones, with drones representing the air and roadside sensors representing the ground, thus forming an air-ground collaborative sensing network.

[0023] Accidents are monitored by the roadside sensors. When an accident is detected, an accident alarm signal is sent to the intelligent decision-making platform for accident handling (i.e., an emergency notification is sent to the intelligent decision-making platform for accident handling). The intelligent decision-making platform for accident handling dispatches the UAV to the accident site to perform data acquisition tasks. The UAV is controlled to adopt a multi-altitude layer circling flight strategy to collect multi-angle image data of the core area of ​​the accident, that is, to collect multi-angle high-definition image sequences with high overlap rate, and simultaneously record POS data (Position and Orientation System data) containing GPS and IMU attitude information. Then, the multi-angle image data (i.e., on-site image data with precise pose labels) is transmitted back to the intelligent decision-making platform for accident handling in real time through a high-bandwidth communication link, providing an immediate and complete data foundation for subsequent high-fidelity 3D reconstruction.

[0024] Step S20: Using the 3D Gaussian sputtering algorithm, perform SFM sparse reconstruction, Gaussian ellipsoid initialization and rasterization rendering on the multi-angle image data to construct a high-fidelity 3D scene of the accident site.

[0025] Specifically, the accident handling intelligent decision-making platform performs SFM (Structure-from-Motion) sparse reconstruction on the multi-angle image data, estimates the camera pose through feature extraction and matching, generates a sparse point cloud, and automatically removes incorrect matches (incorrect matches of feature points).

[0026] The sparse point cloud is initialized with sparse point cloud or randomly initialized with Gaussian to obtain the target point cloud. Then, the target point cloud is densified and pruned with Gaussian ellipsoids to control the density of 3D Gaussian ellipsoids to obtain Gaussian ellipsoids.

[0027] In the point densification stage, the density of Gaussians is adaptively increased to better capture scene details. This process pays special attention to regions with missing geometric features or overly dispersed Gaussians. After a certain number of iterations, densification is performed, focusing on regions with missing geometric features or overly dispersed Gaussians. This includes cloning small Gaussians in insufficiently reconstructed regions or splitting large Gaussians in over-reconstructed regions. For cloning small Gaussians, a copy of the Gaussian is created and moved towards the position gradient. For splitting large Gaussians, two small Gaussians are used to replace one large Gaussian, and the scale is reduced by a specific factor. This step aims to find the optimal distribution and representation of Gaussians in 3D space to enhance the overall quality of the reconstruction.

[0028] During the point pruning phase, removing redundant or less influential Gaussians can be viewed as a form of regularization. This typically involves eliminating nearly transparent Gaussians (α below a specified threshold) and excessively large Gaussians in world or view space. Specifically, it eliminates Gaussians with transparency below a specified threshold and large Gaussians in world or view space. Furthermore, to prevent an unreasonable increase in Gaussian density near the input camera, the transparency of the Gaussian ellipsoid is set to a preset value (e.g., setting Gaussian α to a value close to 0) after a fixed number of iterations. This step saves computational resources while maintaining the accuracy and effectiveness of the Gaussians.

[0029] Finally, the Gaussian ellipsoid is projected onto a two-dimensional plane and rasterized to obtain a high-fidelity three-dimensional scene of the accident site (i.e., a high-fidelity scene reconstruction result).

[0030] Step S30: Introduce multi-granularity segmentation mask and soft-scale gating mechanism, and combine with 3D cue-based segmentation technology to extract accident vehicles, road traces and key objects from the high-fidelity 3D scene.

[0031] Specifically, the multi-view two-dimensional mask extracted by SAM is obtained, and Gaussian affinity features are learned using the multi-view two-dimensional mask. By attaching an affinity feature to each three-dimensional Gaussian, the three-dimensional Gaussian has a new attribute for segmentation, and the target three-dimensional Gaussian features are obtained. The similarity between two affinity features indicates whether the corresponding three-dimensional Gaussian belongs to the same three-dimensional target.

[0032] Meanwhile, in order to handle the inherent multi-granularity ambiguity of 3D suggestible segmentation, a soft-scale gating mechanism is adopted to project the target 3D Gaussian features onto gating feature subspaces of different scales.

[0033] During the inference phase, given a specific viewpoint, 3D visual cues (points with scale) are converted into corresponding 3D scale-gated query features. The feature similarity between the 3D scale-gated query features and 3D affinity features is evaluated to segment the 3D target. Accident vehicles, road traces, and key objects are extracted from the high-fidelity 3D scene. With well-trained affinity features, 3D scene decomposition can be achieved through simple clustering.

[0034] Furthermore, by combining it with CLIP (Contrastive Language-Image Pre-training), open vocabulary segmentation can be performed without requiring a language field.

[0035] After the above processing, a Gaussian point cloud set containing only the accident vehicle, road traces, and key objects is output. This segmentation result preserves the vehicle's three-dimensional geometric structure and texture details, and is transmitted as an independent semantic object to the subsequent VLM (Visual Language Model) for processing.

[0036] Step S40: Based on the accident vehicle, the road traces, and the key objects, construct a traffic accident scene diagram, perform text serialization processing on the traffic accident scene diagram, construct multimodal input, drive the visual language model to deduce the accident logic chain based on the multimodal input, and output an accident analysis report.

[0037] Specifically, the accident vehicles, road traces, and key objects are transformed into a traffic accident scene graph. The extracted key entities are defined as nodes in the graph, including accident vehicles (e.g., vehicle A, vehicle B), road facilities (e.g., traffic lights, solid / dashed lines), and environmental traces (e.g., brake mark length, collision debris distribution). Each node is accompanied by attribute information (e.g., damaged parts of the vehicle, degree of damage, and orientation angle). Based on the relative positions of the entities in three-dimensional space, connecting edges are established between nodes to represent spatial or semantic relationships, such as "vehicle A - impact - vehicle B", "vehicle A - crosses - double yellow lines", and "brake marks - point of contact - collision point". The traffic accident scene graph abstracts unstructured three-dimensional point cloud data into a machine-readable topological structure, clarifying the interaction relationships between the accident subjects.

[0038] The traffic accident scene image and its corresponding numerical features are converted into a natural language text sequence. The traffic accident scene image is traversed to generate descriptive text. The segmented and extracted 3D model of the accident vehicle is rendered into a high-definition 2D snapshot from multiple evidence collection angles as a visual input stream. The natural language text sequence, the descriptive text, and the target text are combined as a text input stream. The visual input stream and the text input stream are encapsulated into a set of structured multimodal cue words. The multimodal cue words are input into a visual language model. The visual language model deduces the accident logic chain based on the multimodal cue words and outputs an accident analysis report.

[0039] To facilitate input and processing by the visual language model, the constructed traffic accident scene diagram and related numerical features need to be converted into a natural language text sequence. This involves traversing the traffic accident scene diagram to generate descriptive text. For example, "Node: Vehicle A, Attribute: Left Front Side Dent" is serialized into the text: "Vehicle A's left front bumper has severe dent deformation." Simultaneously, edge information is converted into spatial relationship descriptions, such as: "Vehicle A is located to the left rear of Vehicle B, and the contact point between the two is located east of the road centerline." The segmented and extracted 3D model of the accident vehicle is rendered into a high-resolution 2D snapshot from key evidence-gathering angles such as frontal view, top view, and driver's viewpoint, serving as the visual input stream. Simultaneously, the generated scene diagram serialization description is combined with relevant traffic regulations text, serving as the text input stream. Finally, the visual image and textual context are encapsulated into a set of structured multimodal prompts, thereby driving the visual language model to deeply understand and align the visual information of the accident scene within a context with prior legal knowledge. The constructed multimodal data is input into the pre-trained VLM to generate a complete accident evidence logic chain covering "phenomenon confirmation - causal backtracking - regulatory determination - conclusion generation".

[0040] This invention addresses the problems of slow 3D reconstruction speed, low segmentation accuracy for deformed vehicles, and lack of interpretable logical chains in automatic liability determination in existing traffic accident processing technologies. It proposes a method for 3D reconstruction and intelligent reasoning of traffic accident scenes based on UAV collaboration and a multimodal large model. First, an air-ground collaborative perception network is constructed, utilizing roadside sensors to monitor accidents and automatically trigger UAV scheduling to collect multi-angle image data of the accident scene in real time. Second, based on SFM sparse reconstruction and Gaussian ellipsoid initialization techniques, a 3D Gaussian Splatting algorithm is used to quickly construct a high-fidelity 3D scene of the accident scene. Next, multi-granularity segmentation masks and soft-scale gating mechanisms are introduced, combined with 3D cue-based segmentation technology, to accurately extract accident vehicles, road traces, and key objects from the high-fidelity 3D scene. Finally, a traffic accident scene graph is constructed and processed into text serialization. By constructing multimodal inputs, a Visual Language Model (VLM) is driven to deduce the accident logical chain, ultimately assisting traffic police in accident analysis and liability determination.

[0041] This invention utilizes 3D Gaussian sputtering technology to significantly improve the real-time performance and high fidelity of 3D reconstruction of accident scenes, effectively solving the problem of long modeling time in traditional methods. Through soft-scale gating mechanism and 3D prompting segmentation, it overcomes the technical bottleneck of accurately separating severely deformed vehicles from complex backgrounds. Furthermore, it innovatively combines a visual language model (VLM) to construct an interpretable accident logic chain, which significantly shortens the accident handling cycle while greatly improving the objectivity, accuracy, and legal credibility of liability determination.

[0042] The technical effects that this invention can bring are as follows: (1) Unmanned Aerial Vehicle (UAV) Automated Evidence Collection Solution: The UAV scheduling and optimal path planning are automatically triggered by the abnormal detection of roadside sensors. The multi-altitude layer circling flight strategy is used to collect accident scene images with high overlap rate and accurate POS attitude information. The high bandwidth and low latency link is used to realize the real-time data transmission, thus completing the automated closed loop from passive accident discovery to active, standardized and high-quality on-site data acquisition. (2) Key evidence extraction framework for 3D reconstruction and segmentation: The Gaussian ellipsoid is initialized using SFM sparse point cloud to achieve millisecond-level rendering and high-fidelity reconstruction of the accident scene. The soft scale gating function based on density and opacity is innovatively designed and combined with 3D cue-based segmentation technology. While effectively suppressing background noise, it accurately preserves the edge geometric features and fine fragment details of severely deformed vehicles, solving the technical problem that traditional segmentation algorithms cannot handle irregular damaged objects.

[0043] (3) Accident logic chain derivation process based on VLM multimodal model: The structured three-dimensional scene diagram is converted into a serialized text description, and a multimodal prompt word is constructed together with the two-dimensional snapshot of the key perspective and the traffic law library. This drives the visual language model (VLM) to perform thought chain reasoning and automatically derives a complete causal logic chain containing "phenomenon confirmation - causal backtracking - legal judgment - conclusion generation". This overcomes the shortcomings of traditional end-to-end models that lack reasoning transparency and legal basis.

[0044] Furthermore, such as Figure 3 As shown, based on the above-mentioned method for deriving traffic accident scenes using UAV collaboration and a multimodal large model, this invention also provides a traffic accident scene derivation system based on UAV collaboration and a multimodal large model, wherein the traffic accident scene derivation system based on UAV collaboration and a multimodal large model includes: The data acquisition module 51 is used to construct an air-ground collaborative sensing network, monitor accidents through roadside sensors of the sensing network, and automatically trigger the scheduling of drones in the sensing network when the roadside sensors detect an accident, and receive multi-angle image data of the accident scene collected by the drones in real time. Scene construction module 52 is used to use the 3D Gaussian sputtering algorithm to perform SFM sparse reconstruction, Gaussian ellipsoid initialization and rasterization rendering on the multi-angle image data to construct a high-fidelity three-dimensional scene of the accident scene. The three-dimensional segmentation module 53 is used to introduce multi-granularity segmentation masks and soft-scale gating mechanisms, combined with 3D cue-based segmentation technology, to extract accident vehicles, road traces and key objects from the high-fidelity three-dimensional scene. The accident deduction module 54 is used to construct a traffic accident scene diagram based on the accident vehicle, the road traces and the key objects, perform text serialization processing on the traffic accident scene diagram, construct multimodal input, drive the visual language model to deduce the accident logic chain based on the multimodal input, and output an accident analysis report.

[0045] Furthermore, such as Figure 4 As shown, based on the above-mentioned method and system for deducing traffic accident scenes based on UAV collaboration and multimodal large models, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 4 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0046] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a traffic accident scene derivation program 40 based on UAV collaboration and multimodal large model, which can be executed by the processor 10 to realize the traffic accident scene derivation method based on UAV collaboration and multimodal large model in this application.

[0047] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the traffic accident scene derivation method based on UAV collaboration and multimodal large model.

[0048] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The terminal's processor 10, memory 20, and display 30 communicate with each other via a system bus.

[0049] In one embodiment, when the processor 10 executes the traffic accident scene derivation program 40 based on UAV collaboration and multimodal large model stored in the memory 20, the following steps are performed: An air-ground collaborative sensing network is constructed, and accidents are monitored through roadside sensors in the sensing network. When the roadside sensors detect an accident, the drones in the sensing network are automatically dispatched to receive multi-angle image data of the accident scene collected in real time by the drones. Using the 3D Gaussian sputtering algorithm, SFM sparse reconstruction, Gaussian ellipsoid initialization and rasterization rendering are performed on the multi-angle image data to construct a high-fidelity three-dimensional scene of the accident scene. By introducing multi-granularity segmentation masks and soft-scale gating mechanisms, combined with 3D cue-based segmentation technology, accident vehicles, road traces, and key objects are extracted from the high-fidelity 3D scene. Based on the accident vehicle, road traces, and key objects, a traffic accident scene diagram is constructed. The traffic accident scene diagram is then processed into text serialization to construct a multimodal input. This drives a visual language model to deduce the accident logic chain based on the multimodal input and outputs an accident analysis report.

[0050] The aforementioned construction of an air-ground collaborative sensing network involves monitoring accidents through roadside sensors within the network. When an accident is detected by the roadside sensors, the network automatically triggers the dispatch of drones to receive real-time multi-angle image data of the accident scene collected by the drones. Specifically, this includes: Construct an air-ground collaborative sensing network, which includes roadside sensors and drones; Accidents are monitored by the roadside sensors. When an accident is detected by the roadside sensors, an accident alarm signal is sent to the intelligent decision-making platform for accident handling. The intelligent decision-making platform for accident handling dispatches the UAV to the accident site to perform data collection tasks, controls the UAV to adopt a multi-altitude layer circling flight strategy, collects multi-angle image data of the core area of ​​the accident, and transmits the multi-angle image data back to the intelligent decision-making platform for accident handling in real time through a high-bandwidth communication link.

[0051] The multi-angle image data includes multi-angle high-definition image sequences and synchronously recorded POS data containing GPS and IMU attitude information.

[0052] Specifically, the method of using a 3D Gaussian sputtering algorithm to perform SFM sparse reconstruction, Gaussian ellipsoid initialization, and rasterization rendering on the multi-angle image data to construct a high-fidelity 3D scene of the accident site includes: The intelligent decision-making platform for accident handling performs SFM sparse reconstruction on the multi-angle image data, estimates the camera pose through feature extraction and matching, generates sparse point clouds, and automatically removes incorrect matches. The sparse point cloud is initialized with sparse point cloud or randomly initialized with Gaussian to obtain the target point cloud. The target point cloud is then densified and pruned with Gaussian ellipsoids to obtain Gaussian ellipsoids. The Gaussian ellipsoid is projected onto a two-dimensional plane and rasterized to obtain a high-fidelity three-dimensional scene of the accident site.

[0053] Specifically, the process of Gaussian ellipsoid densification and pruning of the target point cloud includes: During the point densification stage, the density of Gaussians is adaptively increased. For regions with missing geometric features or scattered Gaussians, densification is performed after a certain number of iterations. For cloning small Gaussians, a copy of the Gaussian is created and moved towards the position gradient. For split large Gaussians, two small Gaussians are used to replace one large Gaussian, and the scale is reduced according to a specific factor. During the point pruning phase, redundant Gaussians are removed, Gaussians with transparency below a specified threshold and large Gaussians in world space or view space are eliminated, and the transparency of the Gaussian ellipsoid is set to a preset value after a fixed number of iterations of the Gaussians.

[0054] Specifically, the introduction of multi-granularity segmentation masks and soft-scale gating mechanisms, combined with 3D cue-based segmentation technology, to extract accident vehicles, road traces, and key objects from the high-fidelity 3D scene includes: Obtain the multi-view two-dimensional mask extracted by SAM, use the multi-view two-dimensional mask to learn Gaussian affinity features, and by attaching an affinity feature to each three-dimensional Gaussian, the three-dimensional Gaussian has a new attribute for segmentation, thus obtaining the target three-dimensional Gaussian features. A soft-scale gating mechanism is used to project the three-dimensional Gaussian features of the target onto gating feature subspaces of different scales; Given a specific viewpoint, 3D visual cues are converted into corresponding 3D scale-gated query features, and the feature similarity between the 3D scale-gated query features and 3D affinity features is evaluated to segment 3D targets and extract accident vehicles, road traces, and key objects from the high-fidelity 3D scene.

[0055] Specifically, the process of constructing a traffic accident scene diagram based on the accident vehicle, road traces, and key objects; performing text serialization processing on the traffic accident scene diagram; constructing a multimodal input; driving a visual language model to deduce the accident logic chain based on the multimodal input; and outputting an accident analysis report includes: The accident vehicle, the road traces, and the key objects are transformed into a traffic accident scene graph. The extracted key entities are defined as nodes of the graph, each node is attached with attribute information, and the connection edges between nodes are established according to the relative position of the entities in three-dimensional space to represent spatial or semantic relationships. The traffic accident scene diagram and its corresponding numerical features are converted into a natural language text sequence, and the traffic accident scene diagram is traversed to generate descriptive text; The segmented and extracted 3D model of the accident vehicle is rendered into a high-definition 2D snapshot from multiple evidence collection angles and used as a visual input stream. The natural language text sequence, the descriptive text, and the target text are combined to form a text input stream; The visual input stream and the text input stream are encapsulated into a set of structured multimodal cue words. The multimodal cue words are input into a visual language model. The visual language model deduces the incident logic chain based on the multimodal cue words and outputs an incident analysis report.

[0056] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a traffic accident scene derivation program based on UAV collaboration and multimodal large model, and the traffic accident scene derivation program based on UAV collaboration and multimodal large model, when executed by a processor, implements the steps of the traffic accident scene derivation method based on UAV collaboration and multimodal large model as described above.

[0057] In summary, this invention provides a method, system, terminal, and computer-readable storage medium for deducing traffic accident scenes based on UAV collaboration and a multimodal large model. The method includes: constructing an air-ground collaborative perception network; monitoring accidents through roadside sensors in the perception network; automatically triggering the scheduling of UAVs in the perception network when an accident is detected by the roadside sensors; receiving multi-angle image data of the accident scene collected in real time by the UAVs; using a 3D Gaussian sputtering algorithm to perform SFM sparse reconstruction, Gaussian ellipsoid initialization, and rasterization rendering on the multi-angle image data to construct a high-fidelity 3D scene of the accident scene; introducing multi-granularity segmentation masks and soft-scale gating mechanisms, combined with 3D cue-based segmentation technology, to extract accident vehicles, road traces, and key objects from the high-fidelity 3D scene; constructing a traffic accident scene graph based on the accident vehicles, road traces, and key objects; performing text serialization processing on the traffic accident scene graph to construct multimodal input; driving a visual language model to deduce the accident logic chain based on the multimodal input; and outputting an accident analysis report. This invention improves the real-time performance and high fidelity of three-dimensional reconstruction of accident scenes, significantly shortens the accident handling cycle, and greatly enhances the objectivity, accuracy, and credibility of liability determination.

[0058] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.

[0059] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.

[0060] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A method for deriving traffic accident scene information based on UAV collaboration and multimodal large model, characterized in that, The method for deriving traffic accident scene data based on UAV collaboration and multimodal large model includes: An air-ground collaborative sensing network is constructed, and accidents are monitored through roadside sensors in the sensing network. When the roadside sensors detect an accident, the drones in the sensing network are automatically dispatched to receive multi-angle image data of the accident scene collected in real time by the drones. Using the 3D Gaussian sputtering algorithm, SFM sparse reconstruction, Gaussian ellipsoid initialization and rasterization rendering are performed on the multi-angle image data to construct a high-fidelity three-dimensional scene of the accident scene. The process of using a 3D Gaussian sputtering algorithm to perform SFM sparse reconstruction, Gaussian ellipsoid initialization, and rasterization rendering on the multi-angle image data to construct a high-fidelity 3D scene of the accident site specifically includes: The accident handling intelligent decision-making platform performs SFM sparse reconstruction on the multi-angle image data, estimates the camera pose through feature extraction and matching, generates sparse point clouds, and automatically removes incorrect matches. The sparse point cloud is initialized with sparse point cloud or randomly initialized with Gaussians to obtain the target point cloud. In the point densification stage, the density of Gaussians is adaptively increased to capture scene details. For regions with missing geometric features or scattered Gaussians, densification is performed after a certain number of iterations. For cloning small Gaussians, a copy of the Gaussian is created and moved towards the position gradient. For split large Gaussians, two small Gaussians are used to replace one large Gaussian. The scale is reduced according to a specific factor to seek the optimal distribution and representation of Gaussians. In the point pruning stage, redundant Gaussians are removed, Gaussians with transparency below a specified threshold and large Gaussians in world space or view space are eliminated. After a fixed number of iterations of Gaussians, the transparency of the Gaussian ellipsoid is set to a preset value to obtain the Gaussian ellipsoid, saving computational resources. The Gaussian ellipsoid is projected onto a two-dimensional plane and rasterized to obtain a high-fidelity three-dimensional scene of the accident site. A multi-view 2D mask extracted by SAM is obtained, and Gaussian affinity features are learned using the multi-view 2D mask. By attaching an affinity feature to each 3D Gaussian, the 3D Gaussian has a new attribute for segmentation, resulting in target 3D Gaussian features. A soft-scale gating mechanism is used to project the target 3D Gaussian features onto different scale-gated feature subspaces. Given a specific viewpoint, 3D visual cues are converted into corresponding 3D scale-gated query features, and the feature similarity between the 3D scale-gated query features and 3D affinity features is evaluated. The 3D target is segmented, and accident vehicles, road traces, and key objects are extracted from the high-fidelity 3D scene. Based on the accident vehicle, the road traces, and the key objects, a traffic accident scene diagram is constructed. The traffic accident scene diagram is then processed into text serialization to construct a multimodal input. This drives a visual language model to deduce the accident logic chain based on the multimodal input and outputs an accident analysis report. The process involves constructing a traffic accident scene diagram based on the accident vehicle, road traces, and key objects; performing text serialization processing on the traffic accident scene diagram to construct a multimodal input; driving a visual language model to deduce the accident logic chain based on the multimodal input; and outputting an accident analysis report. Specifically, this includes: The accident vehicle, the road traces, and the key objects are transformed into a traffic accident scene graph. The extracted key entities are defined as nodes of the graph, each node is attached with attribute information, and the connection edges between nodes are established according to the relative position of the entities in three-dimensional space to represent spatial or semantic relationships. The traffic accident scene diagram and its corresponding numerical features are converted into a natural language text sequence, and the traffic accident scene diagram is traversed to generate descriptive text; The segmented and extracted 3D model of the accident vehicle is rendered into a high-definition 2D snapshot from multiple evidence collection angles and used as a visual input stream. The natural language text sequence, the descriptive text, and the target text are combined to form a text input stream; The visual input stream and the text input stream are encapsulated into a set of structured multimodal cue words. The multimodal cue words are input into a visual language model. The visual language model deduces the incident logic chain based on the multimodal cue words and outputs an incident analysis report.

2. The method for deriving traffic accident scene information based on UAV collaboration and multimodal large model as described in claim 1, characterized in that, The aforementioned air-ground collaborative sensing network monitors accidents through roadside sensors. When an accident is detected by the roadside sensors, the network automatically triggers the dispatch of drones to receive multi-angle image data of the accident scene collected in real time by the drones. Specifically, this includes: Construct an air-ground collaborative sensing network, which includes roadside sensors and drones; Accidents are monitored by the roadside sensors. When an accident is detected by the roadside sensors, an accident alarm signal is sent to the intelligent decision-making platform for accident handling. The intelligent decision-making platform for accident handling dispatches the UAV to the accident site to perform data collection tasks, controls the UAV to adopt a multi-altitude layer circling flight strategy, collects multi-angle image data of the core area of ​​the accident, and transmits the multi-angle image data back to the intelligent decision-making platform for accident handling in real time through a high-bandwidth communication link.

3. The method for deriving traffic accident scene information based on UAV collaboration and multimodal large model as described in claim 2, characterized in that, in, The multi-angle image data includes multi-angle high-definition image sequences and synchronously recorded POS data containing GPS and IMU attitude information.

4. The method for deriving traffic accident scene information based on UAV collaboration and multimodal large model as described in claim 1, characterized in that, The multi-angle image data includes: a multi-angle high-definition image sequence with a high overlap rate, and POS data containing GPS and IMU attitude information.

5. A traffic accident scene derivation system based on UAV collaboration and multimodal large model, characterized in that, The traffic accident scene derivation system based on UAV collaboration and multimodal large model is used to implement the traffic accident scene derivation method based on UAV collaboration and multimodal large model as described in any one of claims 1-4. The traffic accident scene derivation system based on UAV collaboration and multimodal large model includes: The data acquisition module is used to construct an air-ground collaborative sensing network. It monitors accidents through roadside sensors in the sensing network. When the roadside sensors detect an accident, it automatically triggers the scheduling of drones in the sensing network and receives multi-angle image data of the accident scene collected in real time by the drones. The scene construction module is used to perform SFM sparse reconstruction, Gaussian ellipsoid initialization and rasterization rendering on the multi-angle image data using the 3D Gaussian sputtering algorithm to construct a high-fidelity three-dimensional scene of the accident scene. The 3D segmentation module is used to introduce multi-granularity segmentation masks and soft-scale gating mechanisms, combined with 3D cue-based segmentation technology, to extract accident vehicles, road traces and key objects from the high-fidelity 3D scene. The accident deduction module is used to construct a traffic accident scene diagram based on the accident vehicle, the road traces, and the key objects, perform text serialization processing on the traffic accident scene diagram, construct multimodal input, drive the visual language model to deduce the accident logic chain based on the multimodal input, and output an accident analysis report.

6. A terminal, characterized in that, The terminal includes: a memory, a processor, and a traffic accident scene derivation program based on UAV collaboration and multimodal large model stored in the memory and executable on the processor. When the traffic accident scene derivation program based on UAV collaboration and multimodal large model is executed by the processor, it implements the steps of the traffic accident scene derivation method based on UAV collaboration and multimodal large model as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a traffic accident scene derivation program based on UAV collaboration and multimodal large model. When the traffic accident scene derivation program based on UAV collaboration and multimodal large model is executed by the processor, it implements the steps of the traffic accident scene derivation method based on UAV collaboration and multimodal large model as described in any one of claims 1-4.