Intelligent accident liability judgment method and system fused with video tracking
By forming an evidence chain through frame hashing and timestamps, and combining it with a shared backbone multi-branch network and motion model, the problems of evidence credibility and liability determination in traffic accident liability identification are solved, achieving efficient and interpretable liability determination.
Patent Information
- Application Number
- CN202511394575.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-09-28
AI Technical Summary
Existing technologies for determining liability in traffic accidents suffer from problems such as insufficient credibility of evidence, unstable cross-frame/cross-view tracking, insufficient road semantic fusion, and inconsistent legal reasoning, resulting in low efficiency and lack of interpretability in liability determination.
The evidence chain is formed by using frame hashing and timestamps, and target detection and segmentation are performed by combining a shared backbone multi-branch network. Cross-frame association is achieved through appearance re-identification and motion model, which is mapped to the BEV scene graph. Events are extracted and the responsibility ratio is determined according to regulations and rules to generate a credible report.
It has enabled efficient and interpretable determination of liability in traffic accidents, improved the traceability of the evidence chain and the consistency of liability determination, and enhanced processing efficiency and transparency.
Smart Images

Figure CN120894754A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to an accident responsibility intelligent determination method and system fusing video tracking. BACKGROUND
[0002] Traffic accident responsibility identification has long relied on manual retrieval of monitoring and vehicle recorders video, on-site investigation and witness statements for comprehensive judgment. With the increasing complexity of urban road traffic environment and the diversification of video sources (roadside monitoring, vehicle recorders, law enforcement recorders, etc.), the screening, comparison and review of massive evidence is costly and time-consuming, and different law enforcement or claims personnel have different concerns and standards for the same fact details, which may introduce subjectivity. Especially in situations such as multi-vehicle merging, signal-controlled intersections and mixed pedestrian and vehicle traffic, it is difficult to restore the space-time relationship and yielding relationship of each subject by manual frame extraction and visual inspection, affecting the objectivity and consistency of responsibility division.
[0003] To improve the objectivity, existing technologies introduce target detection, semantic / instance segmentation and multi-target tracking methods based on computer vision, and attempt to use appearance re-identification to alleviate the problems of occlusion and identity switching; at the same time, some works also segment road elements (lane lines, stop lines, zebra crossings, traffic lights) from videos to assist in interpretation. However, these perception algorithms mostly stay at the level of "seeing" in single frame or short time sequence, and still have deficiencies in cross-frame consistency, long-term occlusion recovery, cross-camera splicing and unified geometric scale; the recognition stability of small targets (such as distant traffic lights) is not high, and camera shaking and light changes cause trajectory jitter, which is difficult to directly use for event extraction; more importantly, there is a lack of structured mapping between perception results and traffic semantics such as road priority and signal phase, making it difficult to directly support rule-based responsibility reasoning.
[0004] In terms of evidence credibility, video secondary compression, frame stretching / insertion, and uploading of intercepted segments are common in practice, and existing solutions mostly only save file timestamps or simple checksums, lacking a "chain hash + signature + trusted timestamp" forensics mechanism throughout the acquisition-transmission-storage chain, making it difficult to trace and identify single frame / single segment-level modifications. Multi-source data (IMU, GPS, OBD) can assist in restoring vehicle attitude, speed and braking behavior, but there is often a deviation and drift between the time axis of the video and the multi-source data, and when alignment and consistency checking are insufficient, it may introduce new disputes. In addition, personal privacy areas (faces, license plates) lack unified desensitization and minimum necessity principle control in evidence flow, which also restricts evidence sharing and compliance use.
[0005] In the aspect of responsibility judgment and quantification, the existing methods mostly rely on artificial induction based on regulations and experience, or configure a number of trigger conditions (such as crossing the line, running the red light) in the rule engine, which is difficult to deal with the problems such as clause competition, sequence, scope of action and avoidability in complex scenarios; there is a lack of stable and unified calculation base for key quantitative factors such as "conflict point arrival time", "remaining distance" and "reaction time", which leads to the difficulty in reproducibility and review of responsibility proportion. Even if some systems can output conclusions such as "whether to run the red light" and "whether to be impolite", they generally lack traceable links and key frame guidance from evidence to rule triggering, which is not conducive to the transparency and persuasiveness of law enforcement, claims and judicial procedures.
[0006] In summary, the existing technology has not formed an end-to-end consistent framework in terms of evidence credible evidence collection, stable cross-frame / cross-perspective tracking and unified geometric calibration, road semantic fusion and event extraction, deterministic reasoning based on regulations and priority, and responsibility quantification and interpretable report generation, which is difficult to meet the requirements of efficient, strongly interpretable and verifiable accident responsibility determination. Therefore, there is an urgent need for a technical solution that can organically combine video tracking, evidence chain construction, time-space reconstruction, event extraction and rule-based reasoning, and output quantifiable and reviewable conclusions. SUMMARY
[0007] In view of the above problems in the prior art, the present application provides an intelligent accident responsibility determination method and system integrating video tracking, S1 aligns the video and IMU / GPS / OBD evidence collection, and forms an evidence chain by using frame-based hashing, chain-based hashing and timestamp; S2 completes detection and segmentation on a shared backbone multi-branch network; S3 realizes cross-frame association combined with appearance re-identification and motion model, and outputs trajectory, speed and occlusion recovery; S4 maps the trajectory to BEV according to depth and calibration, and constructs a time-space scene graph by fusing lanes, signals, speed limits, etc.; S5 extracts events such as lane changing, merging and signal running, and locates conflict points and yielding relationships; S6 generates fault factors according to the rule base and priority; S7 quantifies the responsibility proportion according to the weight, collision participation and reaction time; S8 generates a report for storing evidence of key frames and rule lists. The system is composed of acquisition and evidence collection, video understanding and tracking, time-space reconstruction, event extraction, rule reasoning, responsibility quantification and report storage modules.
[0008] The present application provides an intelligent accident responsibility determination method integrating video tracking, which comprises the following steps: S1: acquiring a sequence of accident scene video frames, and synchronously acquiring IMU, GPS and OBD multi-source data; performing frame-based content hashing on the multi-source data and concatenating with time information to form chain-based hashing, and adding trusted timestamp and digital signature to form an evidence chain; S2: adopt a shared backbone and multi-branch video understanding network to detect and segment the video, obtain the mask and semantic graph of the vehicle, pedestrian, non-motor vehicle, lane line and signal lamp targets, and perform multi-target cross-frame data association based on the appearance re-identification vector and the motion model, output the trajectory, speed and acceleration of each target in the image coordinate system, and perform occlusion recovery and ID maintenance; S3: based on the appearance re-identification vector and the motion model, perform multi-target cross-frame data association, output the trajectory, speed and acceleration of each target in the image coordinate system, and perform occlusion recovery and ID maintenance; S4: map the trajectory to the bird's eye view BEV according to the depth estimation and camera calibration, fuse the lane geometry, priority, speed limit and signal phase road elements, and construct the space-time scene graph; S5: extract lane changing, merging, U-turn, signal light violation, sudden acceleration and deceleration, and reverse driving on the trajectory and scene graph, and locate potential conflict points and yielding relationships; S6: map the events and subject participation relationships to the regulation rule library and road priority graph, and use causal graph and temporal logic reasoning to obtain a set of fault factors; S7: quantize the responsibility of each subject according to the fault factor weight, collision energy participation and reaction time, and output the responsibility proportion; S8: generate a report containing a time axis, key frames, trajectory superposition and rule trigger list, and store the report and evidence chain in a trusted manner.
[0009] Preferably, the step S1 comprises: establishing a unified time axis based on a monotonically increasing clock, aligning the video frames and IMU, GPS and OBD data according to a fixed time window, and assigning a unique frame number and modal identification to each frame; generating a content digest hash for the frame content of each modal; concatenating the terminal identification, modal identification, frame number, unified timestamp, current content digest hash and chain value of the previous record in a fixed field order, calculating the current chain value by digest, and taking the seed value written by the device at the beginning of the segment or the segment initialization value as the previous chain value; constructing an evidence item with the record containing the field, the previous chain value and the current chain value, digitally signing the evidence item by the terminal private key, and applying for a timestamp from a trusted timestamp service to bind it; appending the signed and timestamped evidence item to the log in chronological order, and aggregating the items to generate a segment digest according to the preset number or length, periodically performing remote storage of the segment digest, including on-chain storage, notarization storage; in the case of link interruption or packet loss, the previous chain value is retained and an placeholder and an abnormal marker are inserted, and after recovery, the extension is continued based on the previous chain value to ensure chain continuity; in the playback and verification, the signature validity, timestamp continuity, chain value matching relationship and content digest consistency are checked in sequence, and any field that is changed will cause the subsequent chain value to be unmatched and be identified, thereby forming a traceable and non-repudiable evidence chain.
[0010] Preferably, the shared backbone and multi-branch video understanding network comprises: after size and brightness normalization of the input video frame, the normalized video frame is input into the shared backbone composed of convolution residual units and temporal context modules, and the shared backbone outputs multi-scale features through a feature pyramid; a multi-branch head is arranged on the shared backbone, comprising: a target detection branch for vehicles, pedestrians and non-motor vehicles, which performs candidate region generation, class determination and position regression; a mask branch for instance segmentation, which samples the detected candidate regions for region alignment and generates pixel-level masks of instances with dynamic parameters; a semantic segmentation branch for road elements, which outputs drivable area and lane line semantic maps on high-resolution features, and generates continuous lane strips through thin line target enhancement and connectivity decoding; a small target detection and state recognition branch for signal lights, which performs upsampling and fine-grained classification on high-level and middle-level features to output the spatial position and light color state of the signal lights; in the inference stage, the results of each branch are smoothed for temporal consistency, non-maximum suppression and boundary refinement, and are fused at the pixel level: the instance mask is used to constrain the conflicting areas in the semantic map, the lane line is used to guide the vehicle instance boundary to fit the road structure, and the signal light state is used to provide temporal markers for event extraction, finally obtaining vehicle, pedestrian, non-motor vehicle instance masks and lane line and drivable area semantic maps aligned with the original resolution.
[0011] Preferably, the step S3 comprises: extracting a fixed-length appearance re-identification vector in the target instance region to represent RGB color and shape; predicting the current position using a constant-speed motion model combined with lane direction and speed limit scene constraints; performing global matching to complete cross-frame association by first performing geometric gating with a prediction window, and then performing global matching by comprehensive scoring of position overlap, motion consistency and appearance similarity; outputting trajectory points, bounding boxes, tracking IDs and timestamps in chronological order, and calculating velocity and acceleration based on adjacent position changes; maintaining the original ID for a temporary target when the appearance and motion thresholds are met.
[0012] Preferably, the step S4 comprises: calibrating the scale and jitter of the target trajectory in the image coordinates by depth estimation combined with IMU; projecting the target trajectory in the image coordinates to the metric BEV grid according to a preset mapping relationship, and performing temporal smoothing and closed-loop correction with lane line and stop line anchor points; time aligning and spatial splicing in overlapping areas in multi-phase scenes, constructing a space-time scene graph based on lane geometry, priority, speed limit and signal phase road elements on the BEV, and combining projected trajectories: nodes include main targets, lane segments, conflict areas and signal phases, edges include traffic connections, yielding and time connections; adding timestamps and confidence to nodes and edges, incrementally updating according to a sliding window, and outputting lane occupancy, conflict line arrival time and remaining distance.
[0013] Preferably, the step S5 comprises: extracting events on the BEV trajectory and the scene graph according to rules-thresholds: lateral displacement crossing lane boundary and keeping stable direction as lane changing; two vehicles longitudinally overlapping and converging to the same lane as merging; heading direction reversing and crossing opposite lane as U-turn; signal phase being red and crossing stop line and entering conflict zone as red light running; speed change rate exceeding threshold as sudden acceleration or deceleration; heading direction opposite to the lane direction as reverse driving, and potential conflict points being determined by the intersection of the trajectories of the two subjects and the conflict line and conflict zone and the arrival time; yielding relationship being determined according to the priority graph and the signal phase and the pedestrian priority rules, and the responsible subject, the triggering frame and the evidence segment being output.
[0014] Preferably, the step S6 comprises: generating a record for each event, the fields containing the subject ID, the event type code, the occurrence time and the road element number; locating the road element number in the space-time scene graph and reading the signal phase and speed limit at the time; determining the subject role as priority traffic, should yield, signal controlled or pedestrian priority according to the road priority graph; searching the rule template in the regulation rule library with the event type code, the road type, the subject role and the phase state as the key; determining in a fixed order: the applicable premise is established, the exemption condition is not established, the subject is located in the controlled area and the time is in the effective interval, that is, the rule triggering is determined; generating a fault factor record for the triggered rule, including the rule number, the fault code, the subject ID and the evidence reference; when there are multiple factor conflicts in the same window, the rule priority is retained in ascending order, and if the same level, the higher one is retained according to the severity level.
[0015] Preferably, the steps S7 and S8 further comprise: summarizing the triggered fault factors for each subject, scoring according to the rule library preset weight, and taking the highest for the same type; determining the collision energy participation level as a weighted coefficient according to the relative speed before contact, the damage site and the braking trace; calculating the difference between the available reaction time and the necessary reaction time, and adding weight to the negative value and reducing the positive value; normalizing the weighted score among the subjects to obtain the responsibility proportion, generating a report in time sequence: containing the time axis, the key frame, the BEV trajectory superposition, the rule triggering list and the evidence reference; and appending the report abstract and the evidence chain entry to the increasing only log, attaching a digital signature and a trusted timestamp, and regularly storing the evidence remotely.
[0016] The application also provides an accident responsibility intelligent judgment system fusing video tracking, comprising: A collection module collects a sequence of video frames of an accident scene, and synchronously collects multi-source data of IMU, GPS and OBD; performs frame content hashing on the multi-source data and concatenates time information to form a chain type hash, and adds a trusted timestamp and a digital signature to form an evidence chain; A target detection and instance segmentation module performs target detection and instance segmentation on the video by using a video understanding network with a shared backbone and multiple branches to obtain masks and semantic graphs of vehicle, pedestrian, non-motor vehicle, lane line and signal lamp targets; The association module performs multi-target cross-frame data association based on the appearance re-identification vector and the motion model, outputs the trajectory, speed and acceleration of each target in the image coordinate system, and performs occlusion recovery and ID maintenance; The space-time scene graph construction module maps the trajectory to the bird's eye view (BEV) according to the depth estimation and camera calibration, fuses road elements such as lane geometry, priority, speed limit and signal phase, and constructs a space-time scene graph; The extraction module extracts lane changing, merging, U-turn, signal light violation, sudden acceleration and deceleration, and reverse driving on the trajectory and scene graph, and locates potential conflict points and yielding relationships; The reasoning module maps the events and subject participation relationships to the regulation rule library and the road priority graph, and obtains a set of fault factors by using causal graph and temporal logic reasoning; The responsibility division proportion determination module quantifies the responsibility of each subject according to the fault factor weight, collision energy participation and reaction time, and outputs the responsibility division proportion; The generation module generates a report containing a time axis, key frames, trajectory superposition and rule triggering list, and performs credible notarization of the report and evidence chain.
[0017] Preferably, the acquisition module comprises: establishing a unified time axis based on a monotonically increasing clock, aligning video frames and IMU, GPS and OBD data according to a fixed time window, and assigning a unique frame number and modal identifier to each frame; generating a content digest hash for the frame content of each modal; concatenating the terminal identifier, modal identifier, frame number, unified timestamp, current content digest hash and chain value of the previous record in a fixed field order, obtaining the current chain value through digest calculation, and taking the seed value written by the device at the factory or the segment head initialization value as the previous chain value of the segment head record; the record comprising the fields, the previous chain value and the current chain value constitutes an evidence item, the evidence item is digitally signed by the terminal private key, and a timestamp is applied to the timestamp service to bind it; the signed and timestamped evidence item is appended to the log in chronological order, and the segment digest is generated by aggregating the items according to the preset number or time length, the segment digest is periodically stored remotely, including on-chain, notarization storage; in the case of link interruption or packet loss, the previous chain value is retained and an placeholder and an abnormal marker are inserted, and after recovery, the extension is continued based on the previous chain value to ensure chain continuity; when playing back and verifying, the signature validity, timestamp continuity, chain value matching relationship and content digest consistency are sequentially checked, and any field that is changed will cause the subsequent chain value to be mismatched and be identified, thereby forming a traceable and irrefutable evidence chain.
[0018] The application provides a kind of fusion video tracking accident responsibility intelligent determination method and system, and the beneficial technical effects that can be realized are as follows: 1. The application generates content hash by frame collection end, and forms chain hash by concatenating previous chain value; terminal signature is affixed to each record and trusted timestamp is bound, evidence entries are written in time sequence to only-increase log, segment abstracts are regularly gathered and chained or notarized; placeholder is inserted in chain interruption, and chain value is continued after recovery; combined with pull frame / two compression detection and video-IMU / GPS / OBD consistency comparison, tampering or loss can be quickly found in playback verification, forming a reviewable evidence closed loop.
[0019] 2. The application completes detection and segmentation of vehicles, pedestrians, lane lines and signal lights through shared backbone multi-branch network, and guarantees cross-frame ID stability through re-identification and motion model; trajectories are unified to BEV through calibration and depth calibration, and are fused into space-time scene graph with lane geometry, priority, speed limit and phase; lane changing, merging and signal violation events are extracted on this base and conflict points are located; error factors are output with clause number and evidence reference according to fixed order of "applicable premise-controlled area-effective period" trigger rule.
[0020] 3. The application weights and normalizes each subject based on error factor weight, collision energy participation and reaction time difference, to obtain reproducible responsibility proportion; report automatically aggregates timeline, key frame and BEV overlay, lists rule trigger list and evidence guide, facilitating review and review; edge cloud collaboration and desensitization rendering balance timeliness and privacy, which can be landed in claims settlement, law enforcement and judicial identification, significantly improving processing efficiency and conclusion consistency. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, a brief introduction will be given below to the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.
[0022] Figure 1 is a step flow chart of an accident responsibility intelligent determination method of the present application fusing video tracking; Figure 2 is a schematic diagram of an accident responsibility intelligent determination system of the present application fusing video tracking. DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0024] Embodiment 1 In view of the above problems mentioned in the prior art, to solve the above technical problems, as shown in the accompanying drawings Figure 1 The application provides an accident responsibility intelligent determination method fusing video tracking, comprising the following steps: S1: Collecting an accident scene video frame sequence, and synchronously collecting IMU, GPS, and OBD multi-source data; performing frame content hashing on the multi-source data and concatenating with time information to form a chain hashing, and appending a trusted timestamp and a digital signature to form an evidence chain; in some embodiments, an integrated driving recorder is installed on the accident vehicle, which is internally provided with a camera, an IMU, a GNSS module (outputting GPS time), an OBD interface, a trusted security chip (storing a terminal private key and a certificate), and an only-increasing log partition (eMMC / SSD). After power-on, a terminal identifier and a certificate are generated or imported; a "monotonically increasing clock" (from a local high-stability clock) is established, and a GNSS time is used for external time service, which is only used for display and checking, and does not change the monotonicity; a fixed time window (such as 33 ms corresponding to 30 fps) is set, and a modal identifier table is determined: VID (video frame), IMU, GPS, and OBD. The time alignment and numbering collection thread rotates by window: a frame of video is taken from the camera and the IMU / GPS / OBD data packet in the same window; a unified timestamp is generated for the window, and a frame serial number (global self-increment) and a modal identifier are assigned; if there is no data for a certain modal (such as no new solution for GPS), a placeholder record is created and a "missing reason code" is marked. A content digest and chain concatenation is performed: a content digest hash is generated for each "frame / packet" (for example, the original frame or the YUV main body region before compression, the IMU original three-axis data, the GPS position and speed field, and the OBD key parameters are normalized and then digested). The fixed field order is spliced: terminal identifier→modal identifier→frame serial number→unified timestamp→current content digest→chain value of the previous record (the first segment uses "initial value").
[0025] The current chain value is calculated, and the evidence entry (containing the above-mentioned fields, the previous chain value and the current chain value, the missing / abnormal marker, etc.) is formed. The signature and timestamp send the evidence entry into the secure chip signed by the terminal private key; when the network is available, a request is initiated to the trusted timestamp service at the same time, and the timestamp token is bound; if offline, the signed entry is cached first and marked as "to be supplemented with timestamp", and after networking, it is supplemented in batches. Only the logged and segmented evidence entries signed and timestamped are appended to the only-increase log file (rolling file, which cannot be modified after writing) in chronological order; every 500 entries are accumulated as a segment, and the segment digest of the segment is calculated and a segment record is generated. The segment record is stored remotely according to the strategy: one is uploaded to the business server (written into the audit table); two is archived to the third-party notary agency regularly; three is optionally written into the alliance chain as an anchor (only the segment digest and index information are stored). When the link is interrupted or packet loss, the previous chain value is retained and a placeholder entry is inserted, recording the abnormal type, duration and local ring buffer index; after recovery, the chain is continued from the previous chain value before interruption, ensuring chain continuity; the transmission layer uses "file block number + check + retransmission window" for breakpoint resume, and records the evidence entry number and time of each retransmission. The local tamper-proof and playback verification WORM (write-once-read-multiple) strategy is enabled in the only-increase log partition; the key directory is protected by secure boot and access control. During playback verification, four steps are performed in order: a. Verify that the terminal signature is valid and the certificate is within the valid period; b. Verify that the trusted timestamp corresponds to the entry one by one and the time is monotonic; c. Verify that the matching relationship between the chain values is continuous without interruption (the placeholder entry must be consistent with the abnormal record); d. Recalculate the digest of the original content, and check whether it is consistent with the content digest in the entry. If any step is inconsistent, it is determined that the entry or the chain after it has been tampered with, and an alarm report is output (listing the first inconsistent entry, corresponding key frame / data packet index and possible cause). Multi-source consistency and privacy minimization: consistency comparison is performed on the video speed estimate and OBD / GPS speed, and when the difference exceeds the threshold, "credibility weight reduction" is marked; desensitization rendering is performed on the face, license plate and other regions at the collection end (while retaining the non-desensitized version in the controlled area, the correspondence between the desensitized and original is maintained by entry number and chain value index). Retrieval and evidence delivery: generate a retrieval index for each segment of the log (according to time, location, event marker); when issuing an evidence package to the outside, export: original video slices, multi-source data, corresponding evidence entries, segment digest and third-party evidence storage credentials within the selected time window, and attach a one-key verification tool (only read, cannot be written), ensuring that the third party can complete the whole process of verification in an offline environment.
[0026] S2: A shared backbone and multi-branch video understanding network is used for target detection and instance segmentation of the video, to obtain masks and semantic maps of vehicle, pedestrian, non-motor vehicle, lane line and signal light targets. In an embodiment, the input and pre-processed dashcam is collected at 1920x1080, 30fps. Before inference, each frame is scaled to 1280 in the long direction, and the short direction is padded to be divisible by 32 (without changing the aspect ratio), and brightness, contrast and white balance normalization is performed; adaptive histogram equalization is enabled for night and strong light scenes; random cropping, color jittering and slight affine enhancement are added in the training stage. The shared backbone (convolution residual + temporal context) uses a lightweight residual network stacking network model, outputting 1 / 4, 1 / 8, 1 / 16 level features; at the tail of each level, a temporal context module is inserted to make light interaction of the corresponding position features of adjacent 5 frames to enhance the discrimination of occlusion and blur scenes. A feature pyramid is built thereon to fuse multi-scale information from top to bottom, obtaining shared features at P3, P4 and P5 scales. The multi-branch head target detection branch (vehicle / pedestrian / non-motor vehicle): decoupled detection heads are set at each scale of P3-P5, respectively outputting class confidence and box position; candidate boxes are merged and redundant suppressed across scales, with high confidence and consistent across scales being prioritized. The instance mask branch: region alignment sampling is performed on the detected candidate boxes to obtain fixed-length features; a set of dynamic parameters is generated for each candidate box to act on the shared feature clipping region, generating a pixel-level mask for the instance; after two upsampling and edge-guided refinement, the mask is pasted back to the original size. The road semantic segmentation branch (drivable area / lane line): the semantic map is output on the high-resolution branch at 1 / 4 resolution; a thin line target enhancement module (direction-sensitive linear response aggregation) and connectivity decoding (connecting and repairing broken thin lines) are introduced for lane lines, to obtain continuous lane strips and drivable areas. The signal light small target detection and state recognition branch: the candidate region is enlarged on P4 and P5 and classified in fine granularity, outputting the signal light center, bounding box and state (red / yellow / green and arrow direction); the state of the same position within 5 frames is smoothed by voting to suppress flickering and misjudgment. Inference stage fusion and post-processing temporal consistency: the class and position of the same target in adjacent frames are smoothed by exponential smoothing, and the instance mask is fine-tuned by morphology to reduce jitter. Non-maximum suppression: candidates of the same class are suppressed at the same scale first, and then suppressed again across scales to avoid repeated detection. The improved output layer Sigmoid function used in the lightweight residual network stacking network model is O(x):
[0027] where e is a natural index, x is an input, K is a variance of the output 1 / 4, 1 / 8, 1 / 16 three-level features; K is positive and bounded, taking K Kmin-Kmax, the parameter K is introduced, which is adaptive to feature statistics, so that the output layer has dynamic sensitivity and dynamic amplitude adjustment capability: in the area with clear texture and low noise, larger K makes the curve steeper and converges faster, which is beneficial to make a decisive judgment on the confidence target; in the weak light, occlusion or motion blur area, smaller K reduces the amplification effect, avoids overfitting and overconfidence, thereby reducing gradient saturation and false positives. The adaptive gate replaces the fixed Sigmoid, which can be used with the normalization / recognition branch to significantly improve the cross-scene stability and calibration, improve the detection and tracking continuity of small targets (such as signal lights), edge targets and occluded targets, while keeping the simple implementation of end-to-end training and inference. Before the output layer, the variance or root mean square is calculated on the third layer feature map according to the channel and local window as the stability index; the candidate value of K is obtained after the statistical quantity is linearly mapped and positively activated (such as softplus); and the candidate value is limited in the safe range of Kmin-Kmax; the sliding average / momentum is used to smooth the inter-batch fluctuations during the training stage, and the learned mapping parameters and cumulative mean are used to directly generate K during the inference stage. To prevent parameter drift, a light regular and gradient clipping are added to K; when the statistics is unavailable or of low quality, it is rolled back to the default Kdefault. The K thus obtained is adaptively matched with the scene complexity, texture and noise level, ensuring numerical stability and interpretability.
[0028] Pixel layer fusion rules: a) instance mask priority covers the conflict area in the semantic map to ensure the clear dynamic target boundary; b) the mask boundary of the lane line guiding the vehicle and the non-motor vehicle is fitted to the road structure to eliminate the "floating" phenomenon; c) the signal light state is written into the time tag of the same frame to provide reliable phase basis for subsequent event extraction. Output: vehicle / pedestrian / non-motor vehicle instance mask aligned with the original resolution, lane line and drivable area semantic map, signal light spatial position and state list; at the same time, the confidence and mask quality score of each instance are output. Training and data: public city traffic data and self-built driving data are mixed for training: day / night, sunny / rainy / backlight, urban expressway / signal intersection are balanced sampling; lane line labeling includes solid and dashed lines, diversion area and stop line; signal light samples cover different heights, small targets at a distance and multi-color arrow combinations. During training, four branches are trained in a multi-task joint manner, and the lane line and signal light samples are resampled according to the scarcity to reduce the class imbalance. Deployment and performance: inference in FP16 on the vehicle-mounted edge computing unit, single-frame end-to-end delay about 35-45 ms; the delay of congested intersections containing multiple targets and multiple light groups does not increase by more than 20%. Under the conditions of night rain, backlight and partial occlusion, the recall of vehicle and pedestrian detection remains stable; after 5-frame voting, the misjudgment rate of signal light state is significantly reduced; the connectivity of lane line in curved and dashed line segments is maintained well, meeting the needs of subsequent BEV mapping and event extraction. Abnormal and degradation strategy: when the image is too dark or raindrops block the confidence, the instance mask branch automatically reduces the threshold and increases the timing smoothing strength; when the detection branch and the semantic segmentation in the same area conflict for a long time, the one with the highest timing consistency is used as the standard and the conflict mark is recorded for upstream diagnosis and retraining. This embodiment completes the cooperative inference of "target detection-instance mask-road semantics-signal light state" on the unified backbone with four special branches, and outputs stable results aligned with the original resolution, which can be directly used for subsequent tracking, BEV mapping and rule judgment through timing and pixel layer fusion.
[0029] S3: Perform multi-target cross-frame data association based on appearance re-identification vector and motion model, output each target's trajectory, speed and acceleration in image coordinate system, and perform occlusion recovery and ID maintenance; in an embodiment, the scene and input urban intersection monocular video, 1080p, 30fps. S2 has output the vehicle / pedestrian / non-motor vehicle bounding box or instance mask, class and confidence of each frame. The appearance re-identification vector extracts the region alignment cropping and normalization for each detection instance, and inputs the lightweight ReID subnetwork to obtain a fixed-length appearance vector (such as 128 dimensions). An appearance template queue (length such as 10) is maintained for each trajectory, which is updated according to the "latest priority, quality decay" strategy, and low clarity or strong light samples are not queued. The motion model and scene constraints use the constant speed assumption to predict the next frame position and scale for each existing trajectory; the angular velocity / acceleration of the IMU is used for camera jitter compensation; the lane direction and speed limit constraints are combined to limit the turning angle and speed upper limit to avoid unreasonable jumps. A search window is generated for the predicted position (for example, with the prediction box as the center, and the long and short axes are enlarged adaptively according to the speed and historical stability). Geometric gating only retains candidate detections that fall within the search window, and requires an overlap with the prediction box to reach a minimum threshold (such as 0.2); candidates with scale changes exceeding a certain proportion (such as ±30%) are excluded.
[0030] The comprehensive matching and one-time association calculate a comprehensive score for each "on-track trajectory-candidate detection", taking into account position overlap, speed direction consistency and appearance similarity; a global matching algorithm is used for one-to-one assignment, which preferentially locks pairs with high scores and few conflicts; new detections that are not assigned create new trajectories, and existing trajectories that are not assigned enter a "lost" state. Occlusion processing and recovery "lost" trajectories continue to extrapolate positions according to the motion model within a limited time window and retain the number; when a new detection appears while meeting the appearance and motion double thresholds, it directly returns to the original number; if the same detection is competed for by two adjacent trajectories, the one with a longer history and more consistent appearance is preferentially retained, and identity switching suppression is performed to prevent ID exchange. Sudden jumps across lanes and instantaneous large turns opposite to the traffic direction are automatically denied. Trajectory maintenance and termination Each trajectory maintains attributes such as "age, hit count, lost count, quality score"; long-term loss or quality score below threshold is terminated and archived segment; new trajectory needs to be hit for a minimum number of frames before being marked as "stable".
[0031] Velocity and acceleration calculation outputs center point or bounding box, instance mask, tracking ID, timestamp for each associated track in time sequence; velocity is calculated based on displacement and frame interval of adjacent frames, and acceleration is obtained from velocity change; smoothing is done within a three to five frame window, and "low confidence" is marked when occlusion or interpolation segment is encountered. Shortage frame and interpolation when only a small number of frames are missing, linear interpolation is done on the time axis to complete the track points, and the "interpolation" mark is marked for this section, which does not participate in the judgment of high sensitivity events. Quality monitoring and alarm continuously statistics each frame association success rate, average appearance similarity, motion residual and other indicators; when the overall quality decreases (such as rainy night or strong backlight), increase the appearance weight, relax the geometric gate and enhance the timing smoothing; if there is large-scale ID jitter, trigger the reset strategy (only clean unstable short tracks, keep long tracks). The output interface provides standardized results to the upper module by frame: track: tracking ID, category, bounding box or mask, center point; kinematics: velocity, acceleration and its confidence; state: normal / temporarily lost / restored, occlusion mark, quality score; event hook: create, merge, terminate, suspected exchange logs, etc. for audit and trace back.
[0032] S4: Map the track to the bird's eye view (BEV) according to the depth estimation and camera calibration, fuse the lane geometry, priority, speed limit and signal phase road elements, and construct the space-time scene graph; in one embodiment, a roadside camera (fixed high view) is deployed at the city intersection and the accident vehicle event data recorder is used as a supplementary view. Before going online, the internal and external parameters and installation posture of the two types of cameras are calibrated respectively; permanent anchor points (stop line, lane boundary point, distance measuring scale) are sprayed or placed on the road, and the calibration results and anchor point coordinates are written into the configuration file and solidified with the device. Depth and image stabilization calibration The edge runs monocular depth estimation to obtain the relative depth of each frame; combined with the pitch / roll angle of the IMU, the camera jitter is compensated; the relative depth is aligned to the metric scale with the measured lane width, stop line distance, etc. as the scale reference; when the light suddenly changes or it is raining at night, the image stabilization strength is automatically increased and the "low confidence" mark is triggered. The track is projected to the BEV from the target center point or contour output from S3, projected to the ground plane according to the preset mapping relationship, and the metric position and orientation are obtained; the projected track is compared with the anchor points (lane line, stop line, zebra crossing) to smooth the sliding time window; when the track and anchor points have small drift, closed-loop correction is performed to make the track fit the road geometry.
[0033] Multi-camera time alignment and spatial stitching Each video stream is synchronized using a unified time source; in the overlapping field of view, fine-grained time alignment is performed using the time when the vehicle passes through the same anchor point; spatially, the roadside camera is taken as the reference, and the vehicle-mounted perspective is projected and spliced into a unified BEV coordinate system; when the same target is observed by two routes at the same time, the side with a closer distance to the anchor point and a clearer mask boundary is preferred, and the fusion source is recorded. Road element fusion Road elements come from two parts: one is the lane center line, boundary, speed limit, and intersection topology of the offline electronic map; the other is the lane line, drivable area, stop line, and flow guide area obtained by online semantic segmentation. The system checks the consistency of the two, and when there is a contradiction, the online result is covered in the short term, and the offline data is corrected in the long term. The position and phase table of the signal machine is provided by the intersection controller or external platform, and is written into the space-time index. The construction of the space-time scene graph maintains the scene in the structure of “node-edge”: node: main target, lane segment, conflict area, stop line, signal phase, speed limit sign; edge: lane connection relationship, yielding relationship, phase control relationship, and time connection of the same main body. Time stamp, effective interval, source, and confidence are attached to the node and edge; a sliding window is used for incremental updating, and only the affected subgraph is recalculated; grid or hash index is used to speed up the positioning of “target→lane segment” and “target→conflict area”.
[0034] Key quantity output Lane occupancy: the proportion and queue length of each lane segment occupied by the main target are counted, and the occupancy percentage and tail coordinates are output; conflict line arrival time: based on the current BEV speed and lane geometry, the expected arrival time of the main body to the intersection conflict line is calculated, and the stability evaluation within the time window is attached; remaining distance: the arc length distance to the conflict line or stop line along the lane center line is accumulated; the quality score of each trajectory and abnormal markers (such as “low-texture area” and “rainy night reflection”) are output synchronously. Abnormality and degradation strategy When the calibration drift is detected (the trajectory deviates from the anchor point for a long time), the system automatically switches to short-term local self-calibration and issues a maintenance alert; when the signal phase is missing, the BEV trajectory and lane fusion are still maintained, but the phase-related event generation is suspended; when the vehicle-mounted perspective and roadside perspective conflict, the one with higher confidence is preferred, and the conflict segment is marked for offline review. Interface to the upper module For S5 / S6, the query capability is provided: given the main body ID and time, the lane segment, the remaining distance to the stop line / conflict line, the expected arrival time, the phase state, and the priority role are returned; the event callback of “subscribe to lane occupancy changes within a time window” and “subscribe to target entering conflict area” is supported.
[0035] S5: Lane change, merge, U-turn, red light running, hard acceleration / deceleration, and reverse driving are extracted from the trajectory and scene graph, and potential conflict points and yielding relationships are located. The environment and input come from the BEV trajectory (m position, orientation, speed, acceleration, confidence) from S4, and the scene graph (lane geometry, stop line, conflict zone / line, signal and phase, priority relationship). Time step 33 ms. Rule parameters (configurable) stable time window: 1.0 seconds; minimum event duration: 0.5 seconds. Lane crossing determination buffer: 0.3 meters. Merge longitudinal overlap: at least 1.5 meters and lasts ≥0.5 seconds. Hard acceleration / hard deceleration threshold: speed increase / decrease ≥5 km / h in 0.5 seconds; extreme cases can be increased to 10 km / h. Reverse driving determination: the angle between the heading and the direction of the lane is ≥150° and lasts ≥0.5 seconds. Red light running: crossing the stop line during the red light effective period and entering the conflict zone within 3 seconds.
[0036] Event extraction process (a) Lane change: detect the trajectory crossing the lane boundary in the lateral direction, and the heading change remains continuous without turning back within the stable time window; if the line is only pressed for less than 0.3 meters or the duration is less than 0.5 seconds, it is marked as "invalid crossing". Output: start frame, crossing frame, completion frame, original / target lane ID. Evidence: three key frames (line in, crossing, lane in) and BEV overlay.(b) Merge: two vehicles have a continuous overlap in the longitudinal direction, and the lateral distance converges to the same lane, and the target vehicle has a lateral trend to the lane within 1 second; if the overlap is insufficient or there is a stop and wait, it is not judged as merging. Output: main / supplementary vehicle ID, merging lane, overlap period. Evidence: overlap area highlight, two vehicle trajectory overlay.(c) U-turn: detect the heading change from forward to reverse, and cross the opposite lane boundary; if it is a U-turn-only lane and there is a release phase, it is marked as "legal U-turn", otherwise "illegal U-turn". Output: U-turn center position, use lane, phase state. Evidence: crossing center line key frame and U-turn completion frame.(d) Red light running: the subject crosses the stop line during the red light phase; then enters any conflict zone (straight / left / right conflict line set) within 3 seconds. If it crosses the stop line and then stops without entering the conflict zone, it is marked as "crossing the line without entering the conflict zone". Output: stop line crossing frame, conflict zone entry frame, phase number. Evidence: two key frames of stop line and conflict zone, phase time axis segment.(e) Hard acceleration / deceleration: speed change exceeds threshold within 0.5 second window; if wet or sudden obstacle in front is detected simultaneously, it can be reduced to "defensive braking". Output: start and end frames, peak change, road condition label. Evidence: speed time piece, forward key frame.(f) Reverse driving: the subject is in the lane, the heading is opposite to the lane direction and lasts more than the threshold; if it is in the diversion area or the U-turn lane and is in the release phase, it is excluded. Output: reverse driving start and end position, lane segment. Evidence: reverse driving segment BEV overlay and heading indication.
[0037] Potential conflict point positioning: For any two subjects, extrapolate 3 seconds along their respective BEV trajectories, find the earliest meeting point with the intersection conflict line / conflict zone and the predicted arrival time; if the interval between the two predicted arrival times is less than 1.2 seconds, mark it as a "high-risk conflict point". Output: conflict point coordinates, two subject predicted arrival times, time difference level (high / medium / low). Yield relationship determination: based on priority diagram: straight ahead priority over left turn, main road priority over branch road, pedestrian priority over motor vehicle, green light release priority over red light control. Combined with the predicted arrival order of the two subjects at the conflict point, determine the "should yield" and "enjoy priority" parties. If there is a special phase or traffic sign (yield / stop), override the sign priority. Output records and evidence fragments: generate standard records for each event: event type code, subject / relative subject ID, occurrence time, spatial location (lane segment or conflict zone ID), phase state, yield role (if applicable), credibility. Evidence fragments include: three to five key frames (start / trigger / complete) with BEV overlay and element highlighting; time axis summary (event duration, phase switching point); evidence chain entry number (corresponding to the log index of S1), for independent verification.
[0038] Conflict and priority processing: If "lane change" and "merge" meet at the same time window, output "merge" first; when "signal light violation" and "crossing the line without entering the conflict zone" conflict, decide by whether entering the conflict zone; merge short repeated events of the same type into one continuous event to avoid fragmentation. Abnormal and degradation: When trajectory confidence is low or lane lines are missing, suspend events strongly bound to this element (such as lane change, merge), and only keep kinematic-based sudden acceleration / deceleration; when phase data is missing, signal light violation event does not trigger, but records a "phase missing" marker for subsequent calculation.
[0039] S6: Map the events and subject engagement relations to the regulation rule base and road priority graph, and obtain a set of fault factors by using causal graph and temporal logic reasoning; in some embodiments, rule mapping and reasoning - intersection "red light running + not courteous to pedestrians" scenario, the scenario and input are derived from S5: Event E1: Vehicle A crosses the stop line and enters the straight conflict zone in the red light phase (type code: crossing into conflict zone). Event E2: Vehicle A and pedestrian P have a high-risk conflict point at the zebra crossing, and the priority relationship is "pedestrian priority, motor vehicle should yield to pedestrians". Scene graph elements: stop line ID = S1, straight conflict zone ID = C1, zebra crossing ID = Z1, signal ID = TL-3 (including phase record corresponding to time); road type = signal-controlled intersection. Time reference: t0 is the crossing frame of E1, t1 is the frame of entering the conflict zone, and t2 is the expected arrival time of the conflict point near Z1. The event record generation system generates a record for each event: E1: subject ID = A, event type code = RL_CROSS, occurrence time = t0 / t1, road element number = S1 / C1. E2: subject ID = A, relative subject ID = P, event type code = YIELD_CONFLICT, occurrence time = t2, road element number = Z1. Element positioning and time sequence information reading positions S1, C1, and Z1 in the scene graph; query the phase and speed limit information of signal TL-3 according to t0, t1, t2: t0, t1 are red light, phase number = R; the speed limit at Z1 is identified as 30 km / h (used for exemption judgment, such as low-speed through buffer stop, etc.). Subject role determination is based on the priority graph: vehicle A's role at S1 / C1 = "signal-controlled". At the zebra crossing Z1, the role of pedestrian P = "pedestrian priority", and the role of vehicle A = "should yield".
[0040] Rule template retrieval with (event type code, road type, subject role, phase status) as retrieval keys: Hit rule R101 (Red light prohibited entry into conflict zone): Applicable premise = signalized intersection; Trigger condition = crossing stop line and entering any conflict zone during red light validity; Exemption condition = police gesture release or special vehicle performing task; Controlled region = stop line and corresponding conflict zone; Priority = 1; Severity level = high. Hit rule R205 (Zebra crossing yield to pedestrian): Applicable premise = pedestrian occupying zebra crossing or passing through; Trigger condition = motor vehicle not yielding and approaching / entering conflict point during yield obligation validity; Exemption condition = pedestrian jaywalking and motor vehicle has taken sufficient avoidance; Controlled region = Z1; Priority = 2; Severity level = high. Rule decision order (deterministic process) for R101: Applicable premise is true (signalized intersection); Exemption condition is not true (no police release, non-special vehicle); Controlled region and time match (S1 / C1 at t0 / t1 during red light validity); => Decision "trigger". For R205: Applicable premise is true (P passing through Z1); Exemption condition is not true (no jaywalking P detected and A has not yielded sufficiently); Controlled region and time match (Z1 at t2 during yield obligation validity); => Decision "trigger". Fault factor generation generates two fault factor records for vehicle A: F101: Rule number = R101, fault code = RUN_RED_IN_CONFLICT, subject ID = A, evidence references = stop line crossing key frame KF_E1a, conflict zone entry key frame KF_E1b, trajectory segment TRJ_A[t0-t1], evidence chain entries IDX_S1 / IDX_C1. F205: Rule number = R205, fault code = FAIL_TO_YIELD_PEDESTRIAN_AT_ZEBRA, subject ID = A, evidence references = zebra crossing conflict point key frame KF_E2, pedestrian trajectory TRJ_P[t2±Δ], evidence chain entries IDX_Z1. No fault factor output for pedestrian P.
[0041] Same window conflict resolution If the same time window also hits the general rule R100 "Red light prohibited crossing stop line" (not including entering conflict zone), which has priority > 1, it is overridden by R101, and only F101 is retained. If two yield-type factors of the same level occur, corresponding to the same location and time window, the one with higher severity level is retained (e.g. "Failure to yield to pedestrian - causing forced stop" takes precedence over "Failure to yield to pedestrian - warning").
[0042] The output is the set of fault factors of A in the current time window {F101, F205}, accompanied by: trigger time and location (S1 / C1, Z1); role description (controlled by signal, should yield); evidence guide list (key frame ID, track segment ID, evidence chain index); judgment log (premise / exemption / zone / valid period verification result of each rule). The above information is provided to S7 for responsibility quantification, and supports independent playback verification: third parties can locate the original entry according to the evidence chain index, check the signature, timestamp, and chain value continuity. Abnormal and degraded processing If the phase data is missing, only R205 is judged and the record is marked "phase missing, not calculating red light running"; if the zebra crossing recognition confidence is insufficient, R205 is suspended and the "element low credibility" label is output; if there is a traffic police on-site command event at the same time, the exemption conditions of R101 / R205 take effect first, and the exemption reason is automatically recorded without triggering.
[0043] S7: According to the fault factor weight, collision energy participation, and reaction time, the responsibility of each subject is quantified, and the responsibility proportion is output; S8: Generate a report containing timeline, key frames, trajectory overlays, and rule-triggered checklist, and credibly notarize the report with the evidence chain. In one embodiment, in the liability quantification and interpretable report generation, a collision occurs between a straight-going car B and a left-turning car A in the conflict zone. From S6: A triggers "red-light entry into conflict zone" (R101, severe), "failure to yield to pedestrian / priority party" (R205, severe); B has no violation, only a "braking insufficient" prompt (prompt level, not counted as fault). From S4 / S5: The time interval between the expected arrival times of the two vehicles at the conflict line is less than 1 second; from S1: On-site extraction of B's approximately 12-meter continuous braking mark and A's approximately 3.5-meter slight braking mark; the relative speed before the collision and the damage location are A's left front and B's front. Fault factor scoring (highest in the same category) The system aggregates the fault factors triggered by each subject within the same window and scores them according to the rule library's pre-set levels: severe > general > prompt. A's highest category is R101 and R205; B's prompt level is not counted and is only used as context description. The "original score" is obtained. The collision energy participation level is determined according to the relative speed before the collision, the impact angle, the damage location, and the braking mark length: A is "high" (the main impact party, short braking mark, large angle), and B is "medium" (the impacted party, significant braking energy reduction). The system applies this level as a "weighting coefficient" to each "original score". Reaction time difference weighting / reduction The system automatically traces back the interval between the "first perceivable time" (the time when the target first enters the visible and predicted conflict zone) and the "actual braking start time" to obtain the "available reaction time"; the "necessary reaction time" is estimated based on the default perception-reaction time and the braking demand under local road conditions. A's difference is negative (braking later than the necessary point), which is recorded as "weighting"; B's difference is positive (earlier than the necessary point), which is recorded as "reduction". The system increases A's overall weight by "weighting" and decreases B's overall weight by "reduction". Normalization The percentage is obtained by normalizing the scores adjusted by "energy participation" and "reaction time difference" between the subjects. The case output: A assumes about 80%, and B assumes about 20%. At the same time, the credibility interval is output, and the key basis for adjustment (phase record, trajectory stability, braking mark recognition quality) is marked. Report timeline and key frames Automatic generation of timeline: t0: A crosses the stop line; t1: A enters the conflict zone; t2: B starts braking; t3: contact occurs. Key frames select t0, t1, t2, and t3, and overlay the trajectories of the two vehicles, the conflict line, the remaining distance, and the speed labels of each frame in the BEV.
[0044] The rule trigger list and explanation items are listed in the "rule list": R101 (entering the conflict area during the red light), R205 (not yielding at the zebra crossing / priority relationship, if applicable); the applicable premise of each rule, trigger evidence reference (stop line crossing frame, conflict area entering frame, phase segment), exemption check result (no police release, non-special vehicle). At the same time, "unaccounted error prompt items" (such as B brake deficiency) and their reasons are listed. Each key frame and trajectory segment is attached with "evidence chain index": including acquisition terminal ID, item number, previous / current chain value summary and time stamp token. The "one-key verification list" is provided in the report, and the third party can verify the signature, time stamp and chain value continuity one by one in the offline tool. The report product and structured report include: cover summary (responsibility ratio, main basis), timing page (time axis + key frame), BEV page (trajectory superposition and distance marking), rule page (trigger / exemption list and clause number), quantification page (itemized score, energy level, reaction time difference explanation), evidence page (evidence chain index table). At the same time, machine-readable attachments (such as JSON) are exported to record the same fields, which is convenient for the case system to dock. The report abstract (hash and index) is generated by trusted notarization, and the evidence chain item number involved in the case is written into the increasing only log; use terminal / server private key signature and obtain trusted time stamp; according to the strategy, the segment abstract is uploaded to the remote notarization (notarization or on-chain anchoring) regularly. The report is transferred to the outside version using desensitization pictures (face, license plate blur), the original piece is left in the controlled domain, and can be traced back through the evidence chain index. If new evidence of phase or brake trace is supplemented subsequently, the system recalculates according to the same process, generates a "revised report", and concatenates the new and old report abstracts in the log, retains the change reason and impact item, and ensures the process traceability.
[0045] Preferably, the step S1 comprises: establishing a unified time axis based on a monotonically increasing clock, aligning the video frames and the IMU, GPS and OBD data in a fixed time window, and assigning a unique frame number and a modal identifier to each frame; generating a content digest hash for the frame content of each modal; concatenating the terminal identifier, the modal identifier, the frame number, the unified timestamp, the current content digest hash and the chain value of the previous record in a fixed field order, obtaining the current chain value through digest calculation, and taking the seed value written by the device at the factory or the segment head initialization value as the previous chain value for the first record in the segment; taking the record containing the fields, the previous chain value and the current chain value as an evidence item, digitally signing the evidence item by the terminal private key, and applying for a timestamp from a trusted timestamp service to bind it; appending the signed and timestamped evidence item to the log in chronological order, and aggregating the items to generate a segment digest according to a preset number or time length, which is periodically stored remotely as evidence, including on-chain storage, notarization storage; retaining the previous chain value and inserting a placeholder and an abnormality marker when the link is interrupted or packets are lost, and continuing to extend based on the previous chain value after recovery to ensure chain continuity; sequentially checking the signature validity, the timestamp continuity, the matching relationship between the chain values and the content digest consistency when playing back and verifying, and any field being tampered with will cause the subsequent chain value to be mismatched and be identified, thereby forming a traceable and non-repudiable evidence chain.
[0046] Preferably, the shared backbone and multi-branch video understanding network comprises: after size and brightness normalization of the input video frame, the normalized frame is sent to a shared backbone composed of convolution residual units and a temporal context module, and the shared backbone outputs multi-scale features through a feature pyramid; a multi-branch head is arranged on the shared backbone, including: a target detection branch for vehicles, pedestrians and non-motor vehicles, which performs candidate region generation, class determination and position regression; a mask branch for instance segmentation, which samples the detected candidate regions for region alignment and generates pixel-level masks of instances with dynamic parameters; a semantic segmentation branch for road elements, which outputs drivable area and lane line semantic maps on high-resolution features, and generates continuous lane strips through thin line target enhancement and connectivity decoding; a small target detection and state recognition branch for traffic lights, which performs upsampling and fine-grained classification on high-level and middle-level features to output the spatial position and light color state of the traffic light; in the inference stage, the branch results are smoothed for temporal consistency, non-maximum suppression and boundary refinement, and are fused at the pixel layer: the instance mask constrains the conflicting areas in the semantic map, the lane line guides the vehicle instance boundary to fit the road structure, and the traffic light state provides temporal markers for event extraction, finally obtaining vehicle, pedestrian and non-motor vehicle instance masks and lane line and drivable area semantic maps aligned with the original resolution.
[0047] Preferably, the step S3 comprises: extracting fixed-length appearance re-identification vectors in the target instance region, representing RGB color, shape; predicting the current position using a constant-speed-based motion model combined with lane direction, speed limit scene constraints; performing global matching to complete cross-frame association by first using a prediction window for geometric gating, and then using a comprehensive score of position overlap, motion consistency and appearance similarity; outputting trajectory points, bounding boxes, tracking IDs and timestamps in chronological order, and calculating speed and acceleration based on adjacent position changes; and maintaining the original ID for a temporarily lost target by motion extrapolation and restoring the original ID when the appearance and motion thresholds are met.
[0048] Preferably, the step S4 comprises: completing internal and external participation installation pose calibration for the acquisition camera, using depth estimation combined with IMU for scale and jitter calibration; projecting the target trajectory in image coordinates to the metric BEV grid according to a preset mapping relationship, and performing temporal smoothing and closed-loop correction with lane line and stop line anchor points; completing time alignment and spatial splicing in overlapping areas for multi-phase scenes, and constructing a space-time scene graph based on lane geometry, priority, speed limit and signal phase road elements on the BEV, combined with projected trajectories: nodes include main targets, lane segments, conflict areas and signal phases, edges include passing connections, yielding and time connections; adding timestamps and confidence to nodes and edges, updating in a sliding window increment, and outputting lane occupancy, conflict line arrival time and remaining distance.
[0049] Preferably, the step S5 comprises: extracting events on the BEV trajectory and scene graph according to rules-thresholds: lateral displacement crossing lane boundaries and maintaining a stable direction as lane changing; two vehicles longitudinally overlapping and converging to the same lane as merging; heading and road direction reversal and crossing opposite lanes as U-turns; signal phase is red and crosses the stop line and enters the conflict area as running a red light; speed change rate exceeds threshold as sudden acceleration and sudden deceleration; heading and lane direction are opposite as reverse driving; potential conflict points are determined by the intersection of two main trajectories, conflict lines and conflict areas, and arrival time; yielding relationship is determined according to priority graph, signal phase and pedestrian priority rules, and outputs responsible subjects, trigger frames and evidence segments.
[0050] Preferably, step S6 includes: generating a record for each event, with fields including subject ID, event type code, occurrence time, and road element number; locating the road element number in the spatiotemporal scene map and reading the signal phase and speed limit according to the time; determining the subject role as priority passage, yielding, signal-controlled, or pedestrian priority based on the road priority map; retrieving rule templates from the regulatory rule base using the event type code, road type, subject role, and phase status as keys; determining in a fixed order: if the applicable premise is met, the exemption condition is not met, the subject is located in a controlled area, and the time is within the valid range, i.e., the rule is considered triggered; generating an error factor record for the triggered rule, including rule number, error code, subject ID, and evidence citation; when multiple factors conflict, retaining them in ascending order of rule priority, and if they are of the same level, retaining the one with the higher severity level.
[0051] Preferably, steps S7 and S8 further include: summarizing the triggered fault factors for each subject, scoring them according to the pre-set weights in the rule base, and taking the highest score for each category; determining the collision energy participation level based on the relative velocity before contact, the damaged location, and the braking traces, as a weighting coefficient; calculating the difference between the available reaction time and the necessary reaction time, with negative values increasing the severity and positive values decreasing the severity; normalizing the weighted scores among the subjects to obtain the responsibility ratio, and generating a report in chronological order: including a timeline, keyframes, BEV trajectory overlay, rule trigger list, and evidence citations; and appending the report summary and evidence chain entries to the incremental log, attaching a digital signature and a trusted timestamp, and periodically storing the evidence remotely.
[0052] This invention also provides an intelligent accident liability determination system that integrates video tracking, such as... Figure 2 As shown, the hardware components include: A. Vehicle Subsystem (any vehicle involved / law enforcement vehicle): Multi-view driving cameras: front / rear / side (1080p / 4K, HDR, night vision illumination). Vehicle IMU: 6 / 9 axes, acceleration / angular velocity. GNSS module: GPS / BeiDou, supporting PPS (pulse) timing. OBD-II data collector: CAN / CAN-FD for reading vehicle speed, braking, turn signals, etc. Vehicle edge computing unit: Industrial-grade SoC (CPU+GPU / NPU), gigabit network / USB3 / MIPI interface. Security chip / TPM or TEE: stores private keys / certificates and completes signing. Local storage: eMMC / SSD, including WORM (increment-only) log partition. Wireless communication: 5G / 4G router (dual SIM backup) and Wi-Fi. Power management: DC-DC (9–36V) + vehicle mini UPS.
[0053] B. Roadside subsystem (fixed point) Roadside camera: high gun camera / dome camera (PoE, IP66), optional panoramic / fisheye. Roadside edge node: industrial computer (GPU / NPU), gigabit PoE switch (supports PTP). RSU roadside unit: C-V2X / DSRC (can receive vehicle broadcast status).
[0054] Signal access device: interface with signal control machine, export phase / scheme (TCP / IP or RS485). Time service device: GNSS time service receiver + PTP grandmaster; NTP as redundancy. Environmental sensor: light / rainfall (can be used for threshold self-adaptation).
[0055] C. Central platform (machine room / cloud), inference and rule cluster: application server (container orchestration). Data and evidence storage: object storage (supports version / WORM), relational database, time series database. Log and audit: message queue / log service (Kafka / Elastic, etc.). Trusted service: timestamp service (TSA) and third-party notarization / on-chain anchor node (optional). Unified device management: OTA, certificate issuance, health monitoring.
[0056] Connection mode and interface, 1) In-vehicle internal connection camera → in-vehicle edge: MIPICSI / USB3; local video stream encoding (H.264 / H.265). IMU → in-vehicle edge: I 2 C / SPI; GNSS → in-vehicle edge: UART (PPS time service into interrupt). OBD-II → in-vehicle edge: CAN / CAN-FD (USB-CAN gateway or direct connection). Security chip / TPM ↔ in-vehicle edge: SPI / I 2C; for frame / segment level signature. Onboard edge ↔ wireless router: Ethernet; external VPN / TLS out. Power supply: vehicle ACC → DC-DC → camera / edge / router; UPS for power outage endurance and safe shutdown. 2) Roadside internal connection, roadside camera ↔ PoE switch ↔ roadside edge: Gigabit Ethernet (RTSP / ONVIF). Signal machine / phase acquisition ↔ roadside edge: TCP / IP; RS485 / Modbus to Ethernet for old devices. RSU ↔ roadside edge: Ethernet (C-V2X / DSRC messages). GNSS ↔ PTP Grandmaster ↔ PoE switch (PTP) ↔ camera / edge time synchronization; NTP as backup. Roadside edge ↔ center: leased line / 5G industrial gateway (IPsec / OpenVPN / TLS). 3) Vehicle-to-roadside-to-center data channel, video and features: RTSP / GB28181 upload or only upload key frames / trajectories (with chain hash index). Metadata and events: MQTT / HTTPS (JSON / Protobuf). Evidence chain reporting: HTTPS / gRPC, TSA timestamp returned to write index back. Device management: HTTPS + certificate two-way authentication; OTA fragmented download, breakpoint resume.
[0057] Data flow and division of responsibilities, S1 forensics and signature (on-board / roadside edge), camera / sensor raw data into edge computing; generate content digest by frame / packet and splice with "last chain value"; call security chip to complete signature, apply for timestamp from TSA offline / online; evidence entries are written into local WORM log and periodically segmented to report to the center and notarized / anchored on the chain. S2 perception and segmentation (edge first), edge nodes run shared backbone + multi-branch network, output detection frame, instance mask, lane semantics, signal light state; only upload results and necessary key frames to save bandwidth. S3 tracking and association (edge), use ReID + motion model to complete cross-frame association; output trajectory, speed, acceleration, occlusion label and trajectory quality; abnormal paragraphs are buffered locally with evidence index. S4 BEV / spatial-temporal graph (edge→center) edge completes camera calibration mapping and BEV projection; multi-camera stitching and large-scale consistency can be checked again in the center; generate "lane occupancy / conflict line arrival / residual distance" and other key quantities. S5 event extraction (edge / center collaboration) lane changing, merging, signal running, reverse driving, etc. are triggered in real time based on rules and thresholds at the edge; complex competition / cross-road scenarios are checked in the center. S6 rule reasoning (center) the center maps events and principal roles to the rule library and priority graph, outputs a set of fault factors according to the deterministic process, and attaches evidence references. Responsibility quantification (center) aggregates fault weight, collision energy participation and reaction time difference, and normalizes the output responsibility proportion; records parameters and basis to support re-calculation. Report and evidence storage (center) generates a report with timeline, key frames, BEV overlay, rule list and evidence chain index; report abstract and associated evidence entries are written again into the center WORM and stored remotely.
[0058] The collection module collects a sequence of video frames of the accident scene, and synchronously collects multi-source data of IMU, GPS and OBD; the multi-source data are subjected to frame-by-frame content hashing and concatenated with time information to form a chain hash, and a trusted timestamp and a digital signature are added to form an evidence chain; The target detection and instance segmentation module uses a shared backbone and a multi-branch video understanding network to perform target detection and instance segmentation on the video, and obtains masks and semantic graphs of vehicle, pedestrian, non-motor vehicle, lane line and signal lamp targets; The association module performs multi-target cross-frame data association based on an appearance re-identification vector and a motion model, outputs a trajectory, a speed and an acceleration of each target in an image coordinate system, and performs occlusion recovery and ID maintenance; The time-space scene graph construction module maps the trajectory to a bird's-eye view (BEV) according to depth estimation and camera calibration, fuses road elements such as lane geometry, priority, speed limit and signal phase, and constructs a time-space scene graph; The extraction module extracts lane changing, merging, U-turn, signal running, sudden acceleration and deceleration, and reverse driving on the trajectory and scene graph, and locates potential conflict points and yielding relationships; The reasoning module maps the events and the relationships between the participants to a legal rule base and a road priority map, and uses causal graphs and temporal logic reasoning to obtain a set of fault factors. The responsibility allocation ratio determination module quantifies the responsibility of each entity based on the fault factor weight, collision energy participation, and reaction time, and outputs the responsibility allocation ratio. The generation module generates a report containing a timeline, keyframes, trajectory overlays, and a list of rule triggers, and then reliably stores the report and the evidence chain.
[0059] Preferably, the acquisition module includes: establishing a unified timeline based on a monotonically increasing clock; aligning video frames and IMU, GPS, and OBD data according to a fixed time window; assigning a unique frame sequence number and modality identifier to each frame; generating content digest hashes for the frame content of each modality; concatenating the terminal identifier, modality identifier, frame sequence number, unified timestamp, current content digest hash, and the chain value of the previous record in a fixed field order; calculating the current chain value through digest calculation; using the seed value or initial value of the segment header as the previous chain value; and constructing an evidence entry with a record containing the field, the previous chain value, and the current chain value, and using the terminal's private key to verify the evidence entry. Digital signatures are generated and timestamps are applied for from a trusted timestamp service to bind them. The evidence entries with signatures and timestamps are appended to a log that only adds and does not modify, in chronological order. The entries are aggregated to generate segment digests according to a preset number of entries or duration. The segment digests are periodically stored remotely, including on-chain storage and notarized storage. When the link is interrupted or packets are lost, the previous chain value is retained and placeholders and anomaly markers are inserted. After recovery, the chain continues to extend based on the previous chain value to ensure chain continuity. During playback and verification, the validity of the signature, the continuity of the timestamps, the matching relationship between the chain values before and after, and the consistency of the content digest are verified in sequence. If any field is modified, the subsequent chain values will be identified as mismatched, thus forming a traceable and non-repudiable chain of evidence.
[0060] This invention provides a method and system for intelligent accident liability determination that integrates video tracking, and the beneficial technical effects it can achieve are as follows: 1. This application generates content hashes frame by frame at the acquisition end and concatenates the previous chain values to form a chain hash; each record is stamped with a terminal signature and bound with a trusted timestamp; evidence items are written to an incremental log in chronological order; segment summaries are periodically aggregated and uploaded to the chain or notarized; placeholders are inserted when the link is interrupted, and the chain value is continued after recovery; combined with frame pulling / double compression detection and video-IMU / GPS / OBD consistency comparison, tampering or missing data can be quickly detected during playback verification, forming a verifiable evidence loop.
[0061] 2. The application completes the detection and segmentation of vehicles, pedestrians, lane lines and signal lights through a shared backbone multi-branch network, and ensures cross-frame ID stability through re-identification and motion models; the trajectories are unified to BEV through calibration and depth calibration, and are fused with lane geometry, priority, speed limit and phase into a space-time scene graph; events such as lane changing, merging and signal violation are extracted and conflict points are located on this base; the error factors are output with clause numbers and evidence citations according to the fixed order of "applicable premise - controlled area - effective period" triggering rules.
[0062] 3. The application weights and scores each subject based on the weight of the error factor, the collision energy participation and the reaction time difference, and normalizes it to obtain a reproducible responsibility ratio; the report automatically summarizes the timeline, key frames and BEV superposition, lists the rule triggering list and evidence guide, and is convenient for review; edge cloud collaboration and desensitization rendering take into account timeliness and privacy, and can be landed in claims settlement, law enforcement and judicial identification, significantly improving processing efficiency and conclusion consistency.
[0063] The above describes in detail an accident responsibility intelligent judgment method and system fusing video tracking; specific examples are applied in this paper to explain the principles and implementation modes of the application; the above examples are only used to help understand the core idea of the application; for those skilled in the art, the specific implementation mode and application range will change according to the idea and method of the application; in summary, the content of this specification should not be understood as a limitation of the application.
Claims
1. A method for intelligent accident liability determination integrating video tracking, characterized in that, Including the following steps: S1: Collect video frame sequences from the accident scene and simultaneously collect multi-source data from IMU, GPS, and OBD; perform frame-by-frame content hashing on the multi-source data and chain it with time information to form a chain hash, and attach a trusted timestamp and digital signature to form a chain of evidence; S2: A video understanding network with a shared backbone and multiple branches is used to perform target detection and instance segmentation on the video to obtain masks and semantic maps of vehicles, pedestrians, non-motorized vehicles, lane lines and traffic lights. S3: Based on the appearance re-identification vector and motion model, perform multi-target cross-frame data association, output the trajectory, velocity and acceleration of each target in the image coordinate system, and perform occlusion recovery and ID preservation; S4: Based on depth estimation and camera calibration, the trajectory is mapped to the bird's-eye view coordinate system BEV, and road elements such as lane geometry, priority, speed limit and signal phase are integrated to construct a spatiotemporal scene map; S5: Extract lane changes, lane merges, U-turns, running red lights, sudden acceleration / deceleration, and driving against traffic from the trajectory and scene map, and locate potential conflict points and yield relationships; S6: Map the events and subject participation relationships to the legal rule base and road priority map, and use causal graphs and temporal logic reasoning to obtain the set of fault factors; S7: Quantify the responsibility of each subject based on the fault factor weight, collision energy participation, and reaction time, and output the responsibility ratio; S8: Generate a report containing a timeline, keyframes, trajectory overlays, and a list of rule triggers, and store the report and the evidence chain in a reliable manner.
2. The intelligent accident liability determination method integrating video tracking as described in claim 1, characterized in that, Step S1 includes: establishing a unified timeline based on a monotonically increasing clock; aligning video frames and IMU, GPS, and OBD data according to a fixed time window; assigning a unique frame sequence number and modality identifier to each frame; generating content digest hashes for the frame content of each modality; concatenating the terminal identifier, modality identifier, frame sequence number, unified timestamp, current content digest hash, and the chain value of the previous record in a fixed field order; calculating the current chain value through digest calculation; using the seed value or initial value written at the device factory as the previous chain value for the segment header record; and constructing an evidence entry with a record containing the field, the previous chain value, and the current chain value, and then using the terminal's private key to perform operations on the evidence entry. Digital signatures are generated and timestamps are applied for from a trusted timestamp service to bind the signatures. The evidence entries following the signatures and timestamps are appended to a log that only adds and never modifies the entries in chronological order. The entries are aggregated to generate segment digests according to a preset number of entries or duration. The segment digests are periodically stored remotely, including on-chain storage and notarized storage. When the link is interrupted or packets are lost, the previous chain value is retained and placeholders and anomaly markers are inserted. After recovery, the chain continues to extend based on the previous chain value to ensure chain continuity. During playback and verification, the validity of the signature, the continuity of the timestamps, the matching relationship between the chain values before and after, and the consistency of the content digest are verified in sequence. If any field is modified, the subsequent chain values will be identified as mismatched, thus forming a traceable and non-repudiable chain of evidence.
3. The method for intelligent accident liability determination based on video tracking as described in claim 1, characterized in that, The shared backbone and multi-branch video understanding network includes: after normalizing the size and brightness of the input video frames, they are fed into a shared backbone composed of convolutional residual units and a temporal context module. The shared backbone outputs multi-scale features through a feature pyramid. Multiple branch heads are set on the shared backbone, including: a target detection branch for vehicles, pedestrians, and non-motorized vehicles, which performs candidate region generation, category determination, and location regression; a mask branch for instance segmentation, which performs region alignment sampling on detected candidate regions and generates pixel-level masks for instances using dynamic parameters; and a semantic segmentation branch for road elements, which outputs drivable areas and lane markings on high-resolution features. The semantic map is generated by enhancing the fine-line targets and decoding connectivity to create continuous lane strips. A branch for small target detection and state recognition of traffic lights is used to amplify and classify high- and mid-level features, outputting the spatial location and color status of the traffic lights. During the inference stage, the results of each branch are subjected to temporal consistency smoothing, non-maximum suppression, and boundary refinement, and then fused at the pixel level: the semantic map is constrained by the instance mask to avoid conflicting regions, the vehicle instance boundaries are guided to fit the road structure by the lane lines, and the traffic light status is used to provide temporal markers for event extraction. Finally, the semantic map of vehicle, pedestrian, and non-motorized vehicle instances, as well as lane lines and drivable areas, is obtained with the original resolution as the final resolution.
4. The method for intelligent accident liability determination based on video tracking as described in claim 1, characterized in that, Step S3 includes: extracting a fixed-length appearance re-identification vector in the target instance region to represent RGB color and shape; predicting the current position using a motion model based on constant speed and combined with lane direction and speed limit scene constraints; first performing geometric gating with the prediction window, and then performing global matching to complete cross-frame association using a comprehensive score of position overlap, motion consistency and appearance similarity; outputting trajectory points, bounding boxes, tracking IDs and timestamps arranged in chronological order, and calculating speed and acceleration based on changes in adjacent positions; maintaining temporarily lost targets by motion extrapolation and restoring the original ID when appearance and motion thresholds are met.
5. The intelligent accident liability determination method integrating video tracking as described in claim 1, characterized in that, Step S4 includes: calibrating the intrinsic and extrinsic parameters and installation attitude of the acquisition camera; performing scale and jitter calibration using depth estimation and IMU; projecting the target trajectory in the image coordinates onto the metric BEV grid according to a preset mapping relationship, and performing temporal smoothing and closed-loop correction using lane lines and stop line anchor points; performing time alignment and spatial stitching of the multi-camera scene in the overlapping area; constructing a spatiotemporal scene graph based on the road elements of lane geometry, priority, speed limit, and signal phase on the BEV, in conjunction with the projected trajectory: nodes include the main target, lane segment, conflict zone, and signal phase, and edges include passing connections, yielding connections, and time connections; adding timestamps and confidence scores to nodes and edges, updating incrementally according to the sliding window, and outputting lane occupancy, conflict line arrival time, and remaining distance.
6. The intelligent accident liability determination method integrating video tracking as described in claim 1, characterized in that, Step S5 includes: extracting events based on rules and thresholds on the BEV trajectory and scene map: lateral displacement crossing the lane boundary and maintaining a stable direction constitutes a lane change; longitudinal overlap of two vehicles converging towards the same lane constitutes a lane merge; reversal of heading and road direction and crossing the oncoming lane constitutes a U-turn; red signal phase and crossing the stop line and entering the conflict zone constitutes running a red light; speed change rate exceeding the threshold constitutes rapid acceleration or deceleration; heading opposite to the direction of its lane constitutes driving in the wrong direction. Potential conflict points are determined by the intersection of the two main trajectories, the conflict line, the conflict zone, and the arrival time; yielding relationship is determined based on the priority map, signal phase, and pedestrian priority rules, and the responsible party, trigger frame, and evidence fragment are output.
7. The intelligent accident liability determination method integrating video tracking as described in claim 1, characterized in that, Step S6 includes: generating a record for each event, with fields including subject ID, event type code, occurrence time, and road element number; locating the road element number in the spatiotemporal scene map and reading the signal phase and speed limit according to the time; determining the subject role as priority passage, yielding, signal-controlled, or pedestrian priority based on the road priority map; retrieving rule templates from the regulatory rule base using the event type code, road type, subject role, and phase status as keys; determining the rule trigger in a fixed order: if the applicable premise is met, the exemption condition is not met, the subject is located in a controlled area, and the time is within the valid range; generating an error factor record for the triggered rule, including rule number, error code, subject ID, and evidence citation; when multiple factors conflict, retaining them in ascending order of rule priority, and if they are of the same level, retaining the one with the higher severity level.
8. The intelligent accident liability determination method integrating video tracking as described in claim 1, characterized in that, Steps S7 and S8 also include: summarizing the triggered fault factors for each subject, scoring them according to the pre-set weights in the rule base, and taking the highest score for each category; determining the collision energy participation level based on the relative velocity before contact, the damaged location, and the braking traces, using it as a weighting coefficient; calculating the difference between the available reaction time and the necessary reaction time, with negative values increasing the severity and positive values decreasing it; normalizing the weighted scores among the subjects to obtain the responsibility ratio, and generating a report in chronological order: including a timeline, keyframes, BEV trajectory overlay, rule trigger list, and evidence citations; and appending the report summary and evidence chain entries to the incremental log, attaching a digital signature and a trusted timestamp, and periodically storing the evidence remotely.
9. An intelligent accident liability determination system integrating video tracking, characterized in that, include: The acquisition module collects video frame sequences from the accident scene and simultaneously acquires multi-source data from IMU, GPS, and OBD. Multi-source data is segmented into frames and hashed, then linked with time information to form a chain hash. A trusted timestamp and digital signature are added to form a chain of evidence. The target detection and instance segmentation module uses a video understanding network with a shared backbone and multiple branches to perform target detection and instance segmentation on the video, and obtains masks and semantic maps of targets such as vehicles, pedestrians, non-motorized vehicles, lane lines and traffic lights; The association module performs multi-target cross-frame data association based on appearance re-identification vectors and motion models, outputs the trajectory, velocity and acceleration of each target in the image coordinate system, and performs occlusion recovery and ID preservation. The spatiotemporal scene map module is constructed by mapping the trajectory to the bird's-eye view coordinate system BEV based on depth estimation and camera calibration, and integrating road elements such as lane geometry, priority, speed limit and signal phase to construct the spatiotemporal scene map; The extraction module extracts lane changes, lane merges, U-turns, running red lights, sudden acceleration / deceleration, and driving against traffic from the trajectory and scene map, and locates potential conflict points and yield relationships. The reasoning module maps the events and the relationships between the main participants to a legal rule base and a road priority map, and uses causal graphs and temporal logic reasoning to obtain a set of fault factors. The responsibility allocation ratio determination module quantifies the responsibility of each entity based on the fault factor weight, collision energy participation, and reaction time, and outputs the responsibility allocation ratio. The generation module generates a report containing a timeline, keyframes, trajectory overlays, and a list of rule triggers, and then reliably stores the report and the evidence chain.
10. The intelligent accident liability determination system integrating video tracking as described in claim 9, characterized in that, The acquisition module includes: establishing a unified timeline based on a monotonically increasing clock; aligning video frames and IMU, GPS, and OBD data according to a fixed time window; assigning a unique frame sequence number and modality identifier to each frame; generating content digest hashes for the frame content of each modality; concatenating the terminal identifier, modality identifier, frame sequence number, unified timestamp, current content digest hash, and the chain value of the previous record in a fixed field order; calculating the current chain value through digest calculation; using the seed value or initial value written at the device factory as the previous chain value for the segment header record; and constructing an evidence entry with a record containing the field, the previous chain value, and the current chain value, which is then processed by the terminal's private key. Digital signatures are generated and timestamps are applied for from a trusted timestamp service to bind the signatures. The evidence entries following the signatures and timestamps are appended to a log that only adds and never modifies the entries in chronological order. The entries are aggregated to generate segment digests according to a preset number of entries or duration. The segment digests are periodically stored remotely, including on-chain storage and notarized storage. When the link is interrupted or packets are lost, the previous chain value is retained and placeholders and anomaly markers are inserted. After recovery, the chain continues to extend based on the previous chain value to ensure chain continuity. During playback and verification, the validity of the signature, the continuity of the timestamps, the matching relationship between the chain values before and after, and the consistency of the content digest are verified in sequence. If any field is modified, the subsequent chain values will be identified as mismatched, thus forming a traceable and non-repudiable chain of evidence.
Citation Information
Patent Citations
Intelligent analysis method for vehicle and pedestrian collision accident liability
CN120217868A
Cited By
Multi-terminal detection score offset correction method and system
CN121073624A
Logistics abnormal event intelligent diagnosis method and system based on knowledge graph
CN121258362A
Automatic driving / driving assistance accident analysis method and device and medium
CN121459447A
An autonomous / drive assist accident analysis method, device and medium
CN121459447B
Semantic hierarchical multi-path video return method and system for vehicle insurance evidence collection and claim settlement
CN121585825A