Traffic accident risk prediction method and system based on generative isovariant causal learning

Through generative and other variable causal learning methods, video generation models and uncertain-perceived video content encoding networks are used to generate video clips in combination with text prompts, which solves the problems of limited generalization ability and causal confusion in the existing technology, and significantly improves the accuracy of traffic accident warning.

CN120014565APending Publication Date: 2025-05-16XI AN JIAOTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510077533.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When existing traffic anomaly detection technologies deal with scenarios with limited labeled data, they have limited generalization capabilities, making it difficult to effectively identify the real causal relationships of the accident, and the background confusion and data set deviation have a great impact.

Method used

Using generative and other variable causal learning methods, video clips containing accidents and not containing accidents are generated through video generation models and uncertain-perceived video content encoding networks, and video clips containing accidents are generated, combined with text prompts for training, time video content representations are obtained, and converted into probability scores for traffic accidents.

Benefits of technology

It significantly improves the accuracy of traffic accident warning, reduces causal confusion and background confusion, enhances the practical application value of the warning system, and can improve the robustness of accident prediction in complex traffic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014565A_ABST
    Figure CN120014565A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic accident risk prediction method and system based on generative isovariant causal learning, and the method comprises the steps: collecting driving video data, constructing a video generation model based on the video frame of the driving video data, training the video generation model through combining with text description prompt fusion before an accident, and obtaining a trained video generation model; generating a video clip containing an accident and a video clip not containing the accident based on the trained video generation model; acquiring time video content representation from the generated video clip containing the accident and the video clip not containing the accident by using a video content coding network perceived by uncertainty; and converting the time video content into a probability score of traffic accident occurrence at a set moment. By effectively solving the problem of causal confusion, the accident prediction robustness of the model in a long-tail, complex and uncertain traffic scene is improved, so that the driving safety is remarkably enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of traffic safety, and in particular to a traffic accident risk prediction method and system based on generative equivariant causal learning. Background Art

[0002] Traffic accidents have unique rarity characteristics in terms of temporal and spatial distribution. Typically, accidents occur in a short period of time, involving approximately 10 to 30 frames, during which road users (including pedestrians, cyclists, vehicles, etc.) account for a relatively small proportion of the frame. Therefore, given the large proportion of normal background in video data, the causal analysis of accidents has the problem of causal confusion. Current traffic anomaly detection (TAA) instances construct a supervised learning framework that includes accident time windows and annotated events, but this design has limited generalization capabilities when dealing with scenarios with limited annotated data. Many traffic anomaly detection tasks focus mainly on object-based processes to ensure consistency in the temporal domain, or extract features by analyzing spatial interactions. However, complex traffic scenarios make it difficult for these processes to achieve satisfactory traffic anomaly detection performance. Summary of the invention

[0003] The purpose of the present invention is to provide a traffic accident risk prediction method and system based on generative equivariant causal learning to solve the above problems.

[0004] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a traffic accident risk prediction method based on generative equivariant causal learning, comprising: Collect driving video data, build a video generation model based on the video frames of the driving video data, and train the video generation model by integrating the text description prompts before the accident to obtain a trained video generation model; Based on the trained video generation model, generate video clips containing accidents and video clips not containing accidents; The generated video clips containing accidents and the generated video clips not containing accidents are used to obtain temporal video content representation using an uncertainty-aware video content coding network; Convert temporal video content into a probability score of a traffic accident occurring at a set moment.

[0005] Furthermore, constructing a video generation model based on video frames includes: Automatically encode video clips of length N frames into latent representations; The latent vector is operated on to introduce k-step noise in the representation by applying a denoising diffusion probability model, generating a noisy latent representation. The video feature vector is obtained through a structure consisting of six 3D criss-cross attention blocks.

[0006] Furthermore, the video generation model is trained by combining the text prompt fusion to obtain a trained video generation model, including: Leveraging multimodal vision and text learning, the CLIP text model encodes text cues; A low-rank adaptive fine-tuning strategy is adopted by integrating the entire 3D cross-attention block and the CLIP text model into the low-rank adaptation model (LoRA) framework, and freezing the pre-trained latent diffusion model and CLIP text model encoder.

[0007] Furthermore, the generating of the video clips containing the accident and the video clips not containing the accident based on the trained video generation model includes: Map a video clip with N frames into a latent representation through an autoencoder; The denoising diffusion implicit model is used to perform back diffusion to remove the noise in the latent representation and obtain a pure latent representation. Randomly extract text prompts from the text prompt pool and encode them using the same CLIP model that was trained in the training phase; A video segment containing the accident and a video segment not containing the accident are generated.

[0008] Furthermore, the generated video clips containing accidents and the generated video clips not containing accidents are obtained by using an uncertainty-aware video content coding network to obtain temporal video content representation, including: The generated video segment triplets are processed by the spatial transformer block, and each frame of video is encoded into a corresponding spatial vector; Adaptive token sampling blocks are used to transform spatial vectors into semantic representations of key object regions to select tokens with important meanings. Cross-frame token embedding and time domain transformer are used to build a temporal model and obtain time-related video content representation.

[0009] Furthermore, the step of converting the temporal video content into a probability score of a traffic accident at a set time includes: Using the accident score decoder, the temporal video content is converted into a probability score of a traffic accident at a specific moment.

[0010] In a second aspect, the present invention provides a traffic accident risk prediction system based on generative equivariant causal learning, comprising: A model building module is used to build a video generation model based on video frames, and train the video generation model by integrating the text description prompts before the accident to obtain a trained video generation model; A video generation module, used to generate video clips containing accidents and video clips not containing accidents based on the trained video generation model; A content representation module is used to obtain temporal video content representation by using an uncertainty-aware video content coding network to generate video clips containing accidents and video clips not containing accidents; The evaluation module is used to convert the temporal video content into a probability score of a traffic accident at a set time.

[0011] Furthermore, constructing a video generation model based on video frames includes: Automatically encode video clips of length N frames into latent representations; The latent vector is operated on to introduce k-step noise in the representation by applying a denoising diffusion probability model, generating a noisy latent representation. The video feature vector is obtained through a structure consisting of six 3D criss-cross attention blocks.

[0012] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of a traffic accident risk prediction method based on generative equivariant causal learning when executing the computer program.

[0013] In a fourth aspect, the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the traffic accident risk prediction method based on generative equivariant causal learning.

[0014] Compared with the prior art, the present invention has the following technical effects: The present invention aims to solve the problems of existing technologies in identifying the true causal relationship of accidents, eliminating the influence of data set bias, reducing background confusion, and the difficulty of collecting and annotating accident videos in actual driving situations. A new equivariant video diffusion model is introduced, which can generate additional accident video clips by converting parts of normal dashcam videos into accident scenes, and making targeted modifications to causal video frames based on text guidance of accidents or no accidents, while ensuring that the time dependency of the generated video sequence is maintained, so as to achieve clear construction of causal relationships without additional annotations, and directly use widely collected driving scene data sets for training, which significantly improves the accuracy of traffic accident warnings and greatly enhances the practical application value of the warning system.

[0015] The present invention first uses well-labeled training samples to convert video clips containing N frames into representations in latent space through an autoencoder. Then, with the help of a denoising diffusion probability model, k-step noise is added to the latent representation to generate a noisy latent representation. Subsequently, the text prompt is encoded through the CLIP text model and processed together with the noisy latent representation using a 3D cross-attention mechanism to estimate the noise and align the spatiotemporal text-video content correlation; then, in the inference stage, the noise in the latent representation is removed by using the trained diffusion model and the denoising diffusion implicit model is applied to generate video clips without and with accidents, and then the video content representation in the time dimension is obtained through an uncertainty-aware video content coding network. Finally, the probability score of a traffic accident is calculated using an accident scoring decoder. By effectively solving the causal confusion problem, the model's robustness in accident prediction in long-tail, complex, and uncertain traffic scenarios is improved, thereby significantly enhancing driving safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flow chart of the traffic accident risk prediction method of generative equivariant causal learning in the present invention; Figure 2 Schematic diagram of the structure of the traffic accident risk prediction model of generative equivariant causal learning in the present invention. DETAILED DESCRIPTION

[0017] The present invention is further described below in conjunction with the accompanying drawings: Example 1, please refer to Figure 1 The present invention provides a traffic accident risk prediction method based on generative equivariant causal learning, comprising: A video generation model is constructed based on video frames, and the video generation model is trained by integrating the text description prompts before the accident to obtain a trained video generation model; Based on the trained video generation model, generate video clips containing accidents and video clips not containing accidents; The generated video clips containing accidents and the generated video clips not containing accidents are used to obtain temporal video content representation using an uncertainty-aware video content coding network; Convert temporal video content into a probability score of a traffic accident occurring at a set moment.

[0018] The present invention transforms part of the normal dashcam video into an accident scene, and based on text guidance of whether there is an accident or not, makes targeted modifications to the causal video frames, while ensuring that the time dependency of the generated video sequence is maintained, thereby achieving clear construction of the causal relationship without the need for additional annotation, and directly using a widely collected driving scene dataset for training, which significantly improves the accuracy of traffic accident warnings and greatly enhances the practical application value of the warning system.

[0019] Embodiment 2, the present invention provides a traffic accident risk prediction method based on generative equivariant causal learning, comprising: Step S1, training the video frames through event detection and video analysis to generate accident videos; Step S2, training the text prompt fused with the video frame representation together; Step S3, generating a specific traffic accident video clip for the trained generation model; Step S4, using an uncertainty-aware video content coding network to obtain a temporal video content representation; Step S5, using an accident score decoder, converting the temporal video content at a specific moment into a probability score of a traffic accident occurring.

[0020] Step 1 includes the following steps: Step S11, automatically encoding a video clip with a length of N frames into a latent representation; Step S12, operating the latent vector by applying a denoising diffusion probability model to introduce k-step noise into the representation, thereby generating a latent representation containing noise; Step S13, obtaining a video feature vector through a structure including six 3D cross attention blocks.

[0021] Step 2 includes the following steps: Step S21, encoding the text prompt using the CLIP model; In step S22, a low-rank adaptive fine-tuning strategy is adopted to integrate the entire 3D cross-attention block and the CLIP text model into the LoRA framework, and the pre-trained latent diffusion model and CLIP text model encoder are frozen, which significantly improves the training efficiency of fine-tuning the diffusion model on a specific dataset.

[0022] Step 3 includes the following steps: Step S31, mapping a segment video with N frames into a potential representation through an automatic encoder; Step S32, performing back diffusion using a denoising diffusion implicit model, aiming to eliminate noise in the latent representation, thereby obtaining a pure latent representation; Step S33, randomly extracting text prompts from the text prompt pool, and encoding them using the same CLIP model that has been trained in the training phase; Step S34, generating a large number of video clips containing accidents and a large number of video clips not containing accidents.

[0023] Step 4 includes the following steps: Step S41, the generated video segment triples are processed by a spatial transformer block, and each frame of video is encoded into a corresponding spatial vector; Step S42, using an adaptive token sampling block to convert the spatial vector into a semantic representation of the key object area, so as to select tokens with important significance; Step S43 uses cross-frame token embedding and a temporal transformer to construct a temporal model, thereby obtaining a temporally relevant video content representation.

[0024] In step 5, an accident score decoder is used to convert the temporal video content at a specific moment into a probability score of a traffic accident occurring.

[0025] Example 3 The 22 video clips from the labeled training samples are converted into latent representations by an autoencoder. These latent vectors are then fed into a denoising diffusion probabilistic model, which adds noise to the latent vectors and produces noisy latent representations in an 8000-step generation process. The text is encoded by the CLIP text model and then fed into a 3D criss-cross attention block along with the latent representation of the video. The block consists of a residual network (ResNet) module, a criss-cross attention layer that fuses text and video information, a spatial attention layer, and a temporal attention layer, with the ResNet layer as the core and a layer dedicated to encoding spatiotemporal features. By adopting the LoRA fine-tuning strategy, we integrate the 3D criss-cross attention block and the CLIP text model into the LoRA framework and freeze the pre-trained LDM and CLIP text model encoders, which significantly improves the efficiency of fine-tuning the diffusion model on a specific dataset while ensuring that the knowledge features in the latent representation are effectively preserved when learning the alignment between video and text.

[0026] In the inference phase, we use the denoising diffusion implicit model to replace the traditional denoising diffusion probability model. By performing 8,000 steps of denoising operations, the noisy latent representation is denoised into a clean latent representation. Using the trained model, a large number of video clips containing accidents and those not containing accidents are generated.

[0027] By building an uncertainty-aware video content coding network, we are able to generate a video content representation in the temporal dimension, which consists of a spatial transformer network, an adaptive token sampling network, and a temporal transformer for cross-frame token embedding. The input video frame is set to 5 frames, and a linear layer with an output dimension of 192 and a position embedding module are integrated in the spatial transformer. The adaptive token sampling block uses an 8-head attention mechanism and linear normalization to identify and select important tokens with semantic features of core object regions. The temporal information is modeled through multi-head attention and maximum pooling layers of the temporal transformer structure, and then the accident probability score is calculated at time point t through two fully connected layers and a normalized exponential (softmax) function, realizing the conversion process from 192 dimensions to 64 dimensions and finally to 2 dimensions.

[0028] In the accident video generation model of the present invention, we use a learning rate parameter of 5e−6 and a batch size of 2. FaceBook's open source Transformers acceleration model Xformers is used to improve hardware memory efficiency. In the inference stage, we use consistent settings to generate videos. The number of inference steps K is set to 50, and the bootstrap ratio is 8.0. In the accident risk prediction stage, the number of layers of the spatial transformer block is set to 2, while the number of layers of the temporal transformer for cross-frame token embedding is set to 3. The AdamW optimizer is used, the learning rate is set to 1e−4, the parameters beta1 are adjusted to 0.9, beta2 to 0.999, and the weight decay is configured to 0.01. In view of the sufficient sample size, all experiments in the present invention use a batch size of 1 in single-cycle training. The experiments were performed on a high-performance platform equipped with three Nvidia RTX 3090 Ti GPUs and 128GB RAM.

[0029] The present invention optimizes the training process of accident video generation by training on a large-scale multimodal accident dataset CAP-DATA, which is derived from driving scenes and contains 5,842 pre-accident clips and 7,419 accident frame clips, generating a total of 13,271 text-video clips for training. Each clip contains 22 frames and corresponding text descriptions. For traffic accident prediction, the method we adopted was evaluated on the DADA-2000, A3D and CCD datasets. DADA-2000 contains 2000 accident videos, of which 1000 are used for training and testing, and all test videos are extracted as 150-frame samples. A3D integrates 1500 dashcam videos. In view of the differences in time and location of traffic accidents in different videos, we crop each A3D video to ensure that a continuous 150-frame, i.e. 5-second video clip, is obtained. We selected 30% of the A3D dataset as the test set. CCD collected 1500 dashcam accident videos. Each video is processed to contain 50 frames, corresponding to a frame rate of 10 frames per second. Each accident is arranged in the last two seconds of its respective accident video. Following the setting of CCD, a total of 900 accident videos were selected in our project for performance evaluation. In the video generation experiment, the present invention shows significant advantages over the existing advanced models in terms of generation quality, text and video alignment accuracy, and computational efficiency. In the traffic accident prediction experiment, the model used showed superior performance to other diffusion models on the A3D and CCD datasets, and achieved a higher AUC value on the DADA-2000 dataset.

[0030] In yet another embodiment of the present invention, a traffic accident risk prediction system based on generative equivariant causal learning is provided, which can be used to implement the above-mentioned traffic accident risk prediction method based on generative equivariant causal learning. Specifically, the system includes: A model building module is used to build a video generation model based on video frames, and train the video generation model by integrating the text description prompts before the accident to obtain a trained video generation model; A video generation module, used to generate video clips containing accidents and video clips not containing accidents based on the trained video generation model; A content representation module is used to obtain temporal video content representation by using an uncertainty-aware video content coding network to generate video clips containing accidents and video clips not containing accidents; The evaluation module is used to convert the temporal video content into a probability score of a traffic accident at a set time.

[0031] The division of modules in the embodiments of the present invention is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional module in each embodiment of the present invention may be integrated into one processor, or may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.

[0032] In another embodiment of the present invention, a computer device is provided, the computer device including a processor and a memory, the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, which are suitable for implementing one or more instructions, and are specifically suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of a traffic accident risk prediction method of generative equivariant causal learning.

[0033] In another embodiment of the present invention, the present invention also provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in a computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by a processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the traffic accident risk prediction method of a generative equivariant causal learning in the above embodiment.

[0034] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0035] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0036] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0037] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0038] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A traffic accident risk prediction method based on generative equivariant causal learning, characterized in that: include: Collect driving video data, build a video generation model based on the video frames of the driving video data, and train the video generation model by integrating the text description prompts before the accident to obtain a trained video generation model; Based on the trained video generation model, generate video clips containing accidents and video clips not containing accidents; The generated video clips containing accidents and the generated video clips not containing accidents are used to obtain temporal video content representation using an uncertainty-aware video content coding network; Convert temporal video content into a probability score of a traffic accident occurring at a set moment.

2. The traffic accident risk prediction method based on generative equivariant causal learning according to claim 1 is characterized in that: The constructing of a video generation model based on video frames comprises: Automatically encode video clips of length N frames into latent representations; The latent vector is operated on to introduce k-step noise in the representation by applying a denoising diffusion probability model, generating a noisy latent representation. The video feature vector is obtained through a structure consisting of six 3D criss-cross attention blocks.

3. The traffic accident risk prediction method based on generative equivariant causal learning according to claim 1 is characterized in that: The video generation model is trained by combining text prompt fusion to obtain a trained video generation model, including: Leveraging multimodal vision and text learning, the CLIP text model encodes text cues; A low-rank adaptive fine-tuning strategy is adopted by integrating the entire 3D cross-attention block and the CLIP text model into the low-rank adaptation model LoRA framework, and freezing the pre-trained latent diffusion model and CLIP text model encoder.

4. The traffic accident risk prediction method based on generative equivariant causal learning according to claim 3 is characterized in that: The generating of the video clips containing the accident and the video clips not containing the accident based on the trained video generation model includes: Map a video clip with N frames into a latent representation through an autoencoder; The denoising diffusion implicit model is used to perform back diffusion to remove the noise in the latent representation and obtain a pure latent representation. Randomly extract text prompts from the text prompt pool and encode them using the same CLIP model that was trained in the training phase; A video segment containing the accident and a video segment not containing the accident are generated.

5. The traffic accident risk prediction method based on generative equivariant causal learning according to claim 1 is characterized in that: The generated video clips containing the accident and the generated video clips not containing the accident are obtained by using an uncertainty-aware video content coding network to obtain a temporal video content representation, including: The generated video segment triplets are processed by the spatial transformer block, and each frame of video is encoded into a corresponding spatial vector; Adopting adaptive token sampling block to transform spatial vector into semantic representation of key object regions, so as to select tokens with important meanings; Cross-frame token embedding and time domain transformer are used to build a temporal model and obtain time-related video content representation.

6. The traffic accident risk prediction method based on generative equivariant causal learning according to claim 5 is characterized in that: The step of converting the temporal video content into a probability score of a traffic accident at a set time includes: Using the accident score decoder, the temporal video content is converted into a probability score of a traffic accident at a specific moment.

7. A traffic accident risk prediction system based on generative equivariant causal learning, characterized in that: include: A model building module is used to build a video generation model based on video frames, and train the video generation model by integrating the text description prompts before the accident to obtain a trained video generation model; A video generation module, used to generate video clips containing accidents and video clips not containing accidents based on the trained video generation model; A content representation module is used to obtain temporal video content representation by using an uncertainty-aware video content coding network to generate video clips containing accidents and video clips not containing accidents; The evaluation module is used to convert the temporal video content into a probability score of a traffic accident at a set time.

8. The traffic accident risk prediction system based on generative equivariant causal learning according to claim 7 is characterized in that: The constructing of a video generation model based on video frames comprises: Automatically encode video clips of length N frames into latent representations; The latent vector is operated on to introduce k-step noise in the representation by applying a denoising diffusion probability model, generating a noisy latent representation. The video feature vector is obtained through a structure consisting of six 3D criss-cross attention blocks.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of a traffic accident risk prediction method based on generative equivariant causal learning as described in any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of a traffic accident risk prediction method based on generative equivariant causal learning as described in any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Vehicle accident prediction method and device

    CN121168725A

  • Traffic accident detection method and system based on combined scene video generation and decoupling representation learning

    CN122116244A