Vehicle accident real-time identification method based on multi-scale feature fusion and fuzzy time slicing
By employing multi-scale feature fusion and fuzzy time slicing, the problem of slow response and insufficient accuracy in vehicle accident detection in existing technologies is solved, achieving high-precision, low-latency real-time accident identification and improving the intelligence and reliability of traffic management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTHEAST UNIV
- Filing Date
- 2025-12-15
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies for vehicle accident detection in complex scenarios suffer from slow response, insufficient accuracy, and poor robustness, making it difficult to achieve real-time and accurate vehicle accident identification.
By employing a method based on multi-scale feature fusion and fuzzy time slicing, and through video data acquisition and preprocessing, moving target detection and dynamic tracking, temporal motion template construction, feature extraction and fusion, fuzzy time slice feature selection, and accident prediction and detection based on deep neural networks, high-precision and low-latency identification of vehicle accidents is achieved.
It achieves high-precision, low-latency real-time accident identification and alarm in complex traffic scenarios, providing reliable technical support for intelligent traffic management.
Smart Images

Figure CN122067201A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent transportation systems, and in particular to a method for real-time vehicle accident identification based on multi-scale feature fusion and fuzzy time slicing. Background Technology
[0002] In recent years, with the rapid growth of urban traffic and the increasing complexity of road conditions, the frequency of vehicle accidents has been rising year by year. This has not only caused serious casualties and property losses, but also posed a great threat to traffic efficiency and public safety. Therefore, the frequent occurrence of vehicle accidents has become a major problem that needs to be solved.
[0003] Traditional accident detection methods primarily rely on video data collected by fixed monitoring equipment, using conventional image processing, target detection, and trajectory tracking techniques for accident assessment. However, these methods often suffer from slow response times, insufficient accuracy, and poor robustness in complex scenarios. While some methods focus on target detection within images, they are frequently limited by factors such as lighting, occlusion, and low resolution. Meanwhile, vehicle trajectory analysis-based techniques depend on high-precision trajectory data and stable tracking algorithms, making them susceptible to interference from multiple targets and data noise in practical applications, leading to misjudgments or missed detections.
[0004] To overcome the above-mentioned shortcomings, a new method is needed that can make full use of the multi-dimensional spatiotemporal information in the video to achieve real-time and accurate identification of vehicle accidents, thereby improving the accuracy and real-time performance of detection and providing more intelligent and reliable technical support for traffic management. Summary of the Invention
[0005] To address the shortcomings of the prior art, the purpose of this invention is to provide a real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing, so as to solve one or more problems in the prior art.
[0006] To achieve the above objectives, the technical solution of the present invention is as follows:
[0007] A real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing includes the following steps:
[0008] S1. Video data acquisition and preprocessing: This involves sequentially processing real-time vehicle monitoring videos to provide stable and accurate data support, including:
[0009] S11, Frame segmentation processing, converting the acquired video data into a discrete image sequence;
[0010] S12. Noise reduction processing, based on Gaussian filter processing to eliminate the influence of factors;
[0011] S13. Image enhancement processing, based on histogram equalization processing, improves image contrast;
[0012] S2. Moving target detection and dynamic tracking to achieve stable tracking of multiple targets in complex scenes, including:
[0013] S21. Moving target detection: background modeling and foreground segmentation based on Gaussian mixture model (GMM).
[0014] S22. Target region extraction: Extract the binary image obtained from the segmentation in S21. Perform row connectivity analysis and group adjacent non-zero pixels into the same target region;
[0015] S23. Dynamic target tracking: Based on the dynamic tracking algorithm and combined with the Kalman filter, the state of the detected target is estimated and updated.
[0016] S3. Construction of temporal motion templates to capture continuous motion features of the target and enhance sensitivity to anomalous time periods, including:
[0017] S31. Motion trajectory data acquisition and preprocessing: Based on dynamic tracking, obtain the position and velocity data of each target in consecutive frames;
[0018] S32, Temporal window partitioning and fuzzy weight assignment: Based on the sliding window technique, several fixed-duration windows are used to divide the continuous motion trajectory.
[0019] S33. Multi-level motion template construction to fully consider the motion indication of the current frame, as well as the template information and target motion speed and direction factors of the previous moment;
[0020] S34. Cubic spline interpolation smoothing and multi-scale fusion are used to obtain the final template;
[0021] S4. Feature extraction and fusion: Utilizing spatial and temporal information from the video to improve feature discriminative power, including:
[0022] S41. Multi-scale spatial feature extraction, based on multi-scale complex Gabor filtering and scale integration nesting for preprocessing each frame of the image. Extract local texture and structural information;
[0023] S42. Extraction of high-order temporal dynamic features: Establishing a nested formula that includes local time weighted integrals and high-order derivatives;
[0024] S43. Deep fusion of multimodal features: Based on multi-layer nested integration, summation and nonlinear mapping, the spatial and temporal features extracted in steps S41 and S42 are deeply fused.
[0025] S44. Feature normalization and dimensionality reduction mapping: The fusion results in S43 are normalized and dimensionality reduced based on nested optimization mapping.
[0026] S5. Fuzzy time slice feature selection to finely segment key moments and highlight important information before and after the accident, including:
[0027] S51. Time slice division and fuzzy membership function construction: The overall time interval is divided into several sub-intervals based on the adaptive sliding window method, and a fuzzy membership function is constructed in each time slice to assign weights to each moment.
[0028] S52. Multi-layer time-weighted feature aggregation: In each time slice, the original time-series features are weighted and integrated, and first-order and second-order time change information are fused at the same time.
[0029] S53. Fuzzy logic filtering for redundancy removal and noise suppression, performing nonlinear mapping processing on the aggregation features of each slice;
[0030] S54. Determination of key time windows and feature output: Define key confidence and select slices as key windows;
[0031] S6. Accident prediction and detection based on deep neural networks: By integrating multimodal features, accurate prediction of accident probability is achieved, including:
[0032] S61. Feature processing and branch fusion: Spatial features, temporal features, and fuzzy time slice features are fused through nonlinear transformation.
[0033] S62. Convolutional layer feature extraction: Local features are extracted based on multi-layer convolution, and convolutional expressions are constructed through nested integrals.
[0034] S63. Dynamic modeling of recurrent layers: Time dynamic modeling based on recurrent neural networks combined with weighted time integrals and nonlinear mapping.
[0035] S64. Weighted fusion of attention mechanisms, combined with self-attention modules to focus on key dynamic information, to construct attention scores and context vectors;
[0036] S65, Fully Connected Layer Prediction Mapping: Maps the context vector in S65 to the accident probability prediction output through the fully connected layer;
[0037] S66. Design and optimization of composite loss function: Construct a composite loss function including cross-entropy loss, regularization term and auxiliary consistency loss.
[0038] S67. Online learning and feedback updates, based on an incremental update strategy, are adjusted online through real-time feedback;
[0039] S7. Post-processing and Alarm: By constructing a multi-level post-processing and decision-making mechanism, accident risk classification identification and alarm are achieved, including:
[0040] S71. Temporal smoothing and continuity correction of accident prediction results: a smoothing function is constructed based on nested integrals and recursive filtering.
[0041] S72. Based on the prediction probability and threshold, an alarm triggering mechanism is constructed to generate an alarm decision signal.
[0042] S73. Multi-level alarm strategy construction: Define a hierarchical decision model, and determine the alarm level based on the integral response of the accumulated alarm decision signal over different time periods. ;
[0043] S74. Interface Design and Real-time Alarm Information Transmission: Design an alarm information encoding and transmission model, specifying alarm levels. Real-time transmission combining latency compensation and data fusion;
[0044] S75. Alarm feedback mechanism and performance evaluation: Construct a composite feedback model to process and update alarm results.
[0045] Furthermore, in S21, the Gaussian Mixture Model (GMM) background modeling is performed by analyzing a pixel at time... Observations Modeling as The weighted sum of Gaussian distributions is shown in the following formula:
[0046]
[0047] In the formula: The number of mixture components represents the number of Gaussian distributions used to model each pixel; Indicates the first The Gaussian component at time... The weights below; Indicates the first The Gaussian component at time... The mean of the following; Indicates the first The Gaussian component at time... The standard deviation below.
[0048] in Let represent the Gaussian probability density function, and its formula is shown below:
[0049]
[0050] This function is used to calculate observed values. In the The probability of occurrence under a given distribution.
[0051] Furthermore, the multi-level motion template construction in S33 is established using the following formula:
[0052]
[0053] In the formula: This represents the direct contribution weight of the motion information in the current frame; The retention weight for template information from the previous moment; To control the decay rate of information from the previous moment; The proportion of motion speed and direction information incorporated into the dynamic weighting factor during template updates; Let be the motion indicator function, if at pixel (x, y) at time... If motion is detected, the value is 1; otherwise, the value is 0. This represents the motion template value from the previous moment, reflecting historical motion information. This refers to the inter-frame time interval.
[0054] in The dynamic weighting factor, calculated based on the target velocity v and direction, is defined as follows:
[0055]
[0056] In the formula: , indicating the current direction of the target's movement; This is a preset reference direction used to correct for differences in motion direction.
[0057] Furthermore, S34 cubic spline interpolation smoothing and multi-scale fusion are performed to obtain the final template, and the calculation formula is shown below:
[0058]
[0059] In the formula: Indicates the first Smooth motion templates at various scales; To integrate weights, satisfy ; The total number of time scales.
[0060] Furthermore, the Gabor filter in S41 is defined as follows:
[0061]
[0062] In the formula: Wavelength; In the direction of the filter; Standard deviation; Aspect ratio; For phase shift; This is the scaling weight function.
[0063] Furthermore, the formula for the fuzzy membership function constructed in S51 is shown below:
[0064]
[0065] In the formula: For the first The center moment of each slice; Scale parameters for controlling the temporal spread range; As a shape control factor, The time interval is defined as follows.
[0066] Furthermore, in S52, the following feature aggregation formula is constructed to achieve multi-level time-weighted feature aggregation:
[0067]
[0068] In the formula: The motion template features at time t; and These represent the first and second order changes of the feature, respectively; , and These are the layer weighting coefficients, which can be adaptively optimized through subsequent training.
[0069] Furthermore, the formula for the composite loss function constructed in S66 is shown below:
[0070]
[0071] In the formula: , and To adjust the weights of different loss contributions, For cross-entropy loss, For weight decay, To mitigate consistency loss.
[0072] Furthermore, the smoothing function in S71 is constructed as shown in the following equation:
[0073]
[0074] In the formula: Normalization factor; Control the width of the smoothing kernel; Define the range of the integration window; This represents the feedback weight of the smoothing result from the previous time step; For feedback delay; It is a non-linear activation function.
[0075] Furthermore, S73 defines a hierarchical decision-making model by constructing a nested integral and exponential mapping model, as shown in the following equation:
[0076]
[0077]
[0078] In the formula: The decay rate over time; Define the integration interval; and These are the threshold values for determining primary and emergency alarms, respectively. The values represent different alarm levels, corresponding to primary alert, emergency alarm, and automatic response, respectively.
[0079] Compared with the prior art, the beneficial technical effects of the present invention are as follows:
[0080] This invention first provides a stable and clear foundation of moving target data for subsequent analysis through video data acquisition and preprocessing, along with moving target detection and dynamic tracking. Based on this, temporal motion template construction, feature extraction, and fusion enable the system to finely characterize the local texture and continuous motion patterns of vehicles. Meanwhile, fuzzy time slice feature selection adaptively weights and focuses on key time periods before and after an accident, thereby extracting highly discriminative dynamic feature representations. These features are efficiently processed by accident prediction and detection based on deep neural networks, achieving accurate prediction of accident probabilities. Finally, post-processing and alarm systems denoise the prediction results and implement intelligent decision-making, effectively filtering instantaneous interference and triggering appropriate alarm levels based on the cumulative risk effect. This achieves high-precision, low-latency, and environmentally adaptive real-time accident recognition and alarming in complex real-world traffic scenarios, providing reliable technical support for intelligent traffic management. Attached Figure Description
[0081] Figure 1 The diagram shows a flowchart of a real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing according to an embodiment of the present invention. Detailed Implementation
[0082] To make the objectives, technical solutions, and advantages of this invention clearer, the following detailed description of a real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing, in conjunction with the accompanying drawings and specific embodiments, will further illustrate the present invention. The advantages and features of this invention will become clearer from the following description. It should be noted that the accompanying drawings are in a very simplified form and use non-precise proportions, used only to facilitate and clearly illustrate the purpose of the embodiments of this invention. Please refer to the accompanying drawings to make the objectives, features, and advantages of this invention more apparent and understandable. It should be understood that the structures, proportions, sizes, etc., depicted in the accompanying drawings are only for illustrative purposes to aid those skilled in the art and are not intended to limit the implementation conditions of this invention. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in proportions, or adjustments to the size, without affecting the effects and objectives achieved by this invention, should still fall within the scope of the technical content disclosed in this invention.
[0083] Please refer to the following: Figure 1 A real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing includes the following steps:
[0084] S1. Video data acquisition and preprocessing.
[0085] Intelligent preprocessing of raw video data is achieved through real-time frame segmentation, noise filtering, and image enhancement of vehicle monitoring videos. An edge computing platform is used to split the continuous video stream into discrete image sequences, and the frame sampling rate is dynamically adjusted based on image quality. Adaptive denoising algorithms and brightness / contrast optimization techniques are employed to eliminate low-quality and redundant frames, ensuring that each frame is clear and rich in detail. The resulting high-quality image data, generated through weighted fusion, accurately and comprehensively reflects the scene, providing stable and accurate data support for subsequent moving target detection and deep learning prediction.
[0086] Specifically, the steps include the following:
[0087] S11, Frame splitting.
[0088] Use the input vehicle monitoring video This represents the spatial coordinates, where (x, y) represent the coordinates. Indicates time. The video is processed according to a certain sampling interval. Extract as a single image. If the video's sampling rate is... Frames / second, then every If one frame is extracted per second, the expression is as shown in Equation 1 below:
[0089]
[0090] In the formula: For the first Frame image. This is the start time. This represents the total number of frames extracted.
[0091] The video is converted into a discrete image sequence by frame-by-frame processing, so that each frame can be processed separately.
[0092] S12, noise reduction processing.
[0093] By employing a Gaussian filter to smooth the image, the effects of sensor noise, compression noise, or changes in illumination are eliminated. The filtering formula is shown in Equation 2 below:
[0094]
[0095] In the formula: This represents the pixel value at position (xi, yj) in the original image. Represents the radius of the filter kernel (window). This represents the pixel value at position (x, y) after filtering. The weight of the two-dimensional Gaussian kernel at offset (i, j) is shown in Equation 3 below:
[0096]
[0097] In the formula: The normalization factor, the exponential part This determines that pixels farther from the center receive smaller weights, resulting in a typical bell-shaped distribution. The standard deviation is denoted as .
[0098] The entire Gaussian filtering process is achieved through... Within the window, a weighted average is applied to each pixel, with a focus on pixels near the center that have a higher weight, while pixels further away from the center have a lower weight. This achieves image smoothing and noise removal while preserving the main structure and details of the image as much as possible.
[0099] S13, Image enhancement processing.
[0100] The image after S12 denoising may suffer from low contrast or insufficient brightness, affecting the accuracy of subsequent object detection. Therefore, image enhancement processing is necessary. Histogram equalization is used to even out the distribution of pixel values in the image, improving image contrast. Assuming the original image's pixel grayscale values are... The range of values is Typically, L=256, and the mapping function for histogram equalization is shown in Equation 4 below:
[0101]
[0102] In the formula: grayscale value The number of pixels. This represents the total number of pixels in the image.
[0103] S2, Moving Target Detection and Dynamic Tracking
[0104] A Gaussian mixture model is used to model the background and segment the foreground for each pixel in the video, thus distinguishing moving targets from the background. Then, connected component analysis is applied to the binary image obtained from the foreground segmentation to extract target regions, and the bounding boxes and centroids of each region are calculated. Finally, a dynamic tracking algorithm is used to predict the state of detected targets, correlate data, and update the data, achieving stable tracking of multiple targets in consecutive frames and ensuring that the motion trajectory of each target is continuously recorded.
[0105] Specifically, the steps include the following:
[0106] S21. Moving target detection.
[0107] Gaussian Mixture Model (GMM) is used for background modeling and foreground segmentation, treating the grayscale or color value of each pixel as a mixture of multiple Gaussian distributions. For a given pixel at time... Observations Modeling as The weighted sum of Gaussian distributions is shown in Equation 5 below:
[0108]
[0109] In the formula: The number of mixture components represents the number of Gaussian distributions used to model each pixel. It is commonly set to 3 to 5. More components can better adapt to complex background changes, but at the same time, it will increase the amount of computation. Indicates the first The Gaussian component at time... The weights are the contributions of the distribution to the probability of the current pixel value. Components with larger weights are usually considered to represent "common" or background information. Indicates the first The Gaussian component at time... The mean of the distribution is the expected value of the pixel value. Indicates the first The Gaussian component at time... The standard deviation below describes the range of pixel values fluctuating around the mean; a smaller value indicates a higher standard deviation. The value indicates that the distribution is sensitive to changes in pixel values, and larger values indicate that the distribution is sensitive to changes in pixel values. It is suitable for slight fluctuations in lighting or background. This represents the Gaussian probability density function, which is used to calculate the observed values. In the The probability of occurrence under a given distribution is given by the following formula, Equation 6:
[0110]
[0111] Subsequently, for the current pixel observation Check all in turn We have a Gaussian component, and we check if there exists a component that satisfies the following equation 7:
[0112]
[0113] In the formula: The threshold, typically set to 2.5 or 3, is used to determine whether the current observation "matches" a Gaussian distribution. If it matches, the component is considered to likely describe the background. If a component meets the matching condition, its parameters are updated using the current observation. Let the learning rate be... The update formulas are shown in equations 8 to 11 below:
[0114] Mean update:
[0115]
[0116] Variance update:
[0117]
[0118] Weight update:
[0119]
[0120] For unmatched components, their weights are updated as follows:
[0121]
[0122] If no component is with If a match is found, it is considered that a new background change or foreground target has appeared. In this case, the component with the smallest weight is usually replaced with an initial higher variance and a lower weight, and its mean is set to... .
[0123] Based on the weight and variance of each component, it is usually arranged according to... Value sorting is used, and the top few items whose cumulative weight reaches a certain threshold are used as the background model. If the current pixel... If a pixel falls within the background component, it is marked as background; otherwise, it is marked as foreground. This method continuously updates the multi-Gaussian model for each pixel, enabling the system to adaptively capture background changes in the scene. Simultaneously, when an observation deviates significantly from the background model, it is classified as a foreground target, providing accurate segmentation results for subsequent moving target detection and dynamic tracking.
[0124] S22, Target Region Extraction.
[0125] Binary image obtained after foreground segmentation Connectivity analysis is performed to group adjacent non-zero pixels into the same target region. Each connected region yields a bounding box and centroid coordinates, providing basic features for subsequent tracking. Noisy regions that are too small or do not conform to the preset shape are filtered out based on features such as region area and shape, retaining regions that may represent valid targets such as vehicles.
[0126] S23, Dynamic Target Tracking.
[0127] To maintain target consistency across consecutive frames, a dynamic tracking algorithm is employed, using a Kalman filter to estimate and update the state of each detected target. (State vector) To describe the motion state of a target in an image, it is usually defined as shown in Equation 12 below:
[0128]
[0129] In the formula: This represents the target's coordinates in the horizontal direction (x-axis of the image). This indicates the target's coordinates in the vertical direction (y-axis of the image). This represents the target's velocity in the horizontal direction, i.e., its horizontal displacement in each frame. This indicates the target's velocity in the vertical direction.
[0130] State transition matrix It is used to predict the current state based on the state of the previous time step. Under the uniform motion model, it is defined as shown in Equation 13:
[0131]
[0132] In the formula: This represents the inter-frame time interval, used to convert velocity information into displacement, thereby updating the target position.
[0133] Predicted state Indicates based on the state at the previous time step Through the state transition matrix The predicted state at the current moment is calculated using the following formula (Equation 14):
[0134]
[0135] In the current frame, the detection information obtained by extracting the target region, such as the target centroid, constitutes the measurement vector. Using measurement matrix To obtain the predicted measurement value As shown in Equation 15 below:
[0136]
[0137] Subsequently, through Kalman gain The updated state is calculated as shown in Equation 16 below:
[0138]
[0139] Kalman gain The calculation will be dynamically adjusted based on information such as prediction covariance and measurement noise.
[0140] In practical applications, multiple targets often exist in a video. In this case, it is necessary to associate the detection results of the current frame with each tracked target. First, the similarity between the detected bounding box and the predicted bounding box is calculated, and then the Hungarian algorithm is used for global optimal matching, as shown in Equation 17 below:
[0141]
[0142] In the formula: The target bounding box detected in the current frame. The bounding box of the target obtained by the Kalman filter prediction.
[0143] By setting appropriate A threshold or distance threshold is used to determine whether the detection result matches an existing target. For unmatched detection results, a new tracking target is initialized with an initial state and a high degree of uncertainty. If a tracking target fails to match for several consecutive frames, it is considered a tracking loss and the target is removed from the tracking list. This dynamic tracking method enables the system to continuously track moving targets in complex scenes, maintain target consistency across consecutive frames, and provide continuous and reliable target position information for subsequent steps.
[0144] S3, Temporal Motion Template Construction
[0145] Employing multi-scale spatiotemporal information fusion technology, this module divides continuous motion information into multiple fixed-duration temporal windows based on target motion trajectory data obtained through dynamic tracking. A fuzzy weighting strategy is introduced to dynamically assign importance to different time frames. Building upon traditional motion template construction, this module further integrates target motion velocity, direction, and spatiotemporal attenuation characteristics, achieving multi-level weighted fusion of current frame motion information and historical motion data. Simultaneously, cubic spline interpolation smoothing and multi-scale fusion techniques are introduced to meticulously reconstruct and uniformly fuse motion templates at each scale. Ultimately, the constructed comprehensive motion template accurately and robustly reflects the target's motion intensity, trend, and spatiotemporal distribution over continuous time, providing high-quality and stable input data for subsequent motion feature extraction and anomaly detection.
[0146] Specifically, the steps include the following:
[0147] S31. Motion trajectory data acquisition and preprocessing.
[0148] The dynamic tracking module acquires the position and velocity data of each target in consecutive frames to form a motion trajectory sequence, such as the target at time t. The position and velocity are respectively and This data provides timing information for the subsequent construction of motion templates.
[0149] S32. Temporal Window Division and Fuzzy Weight Assignment. The continuous motion trajectory is divided into several fixed-duration time windows using a sliding window technique. Each window provides local temporal data for constructing the motion template. To emphasize the importance of different time frames within the window, a fuzzy weight function is introduced as shown in Equation 18:
[0150]
[0151] In the formula: Indicates the time at the center of the window. Controlling the extent of weight decay, Used to adjust the shape of the decay curve. This function makes frames closer to the center have higher weights and frames farther from the center have lower weights, thus highlighting information at key moments in template construction.
[0152] S33, Multi-level motion template construction.
[0153] Within each time window, an improved motion template update formula is proposed, which not only considers the motion indication of the current frame, but also incorporates the template information from the previous moment and the target's motion speed and direction factors. The formula is shown in Equation 19 below:
[0154]
[0155] In the formula: This represents the direct contribution weight of the motion information in the current frame. The retention weight for template information from the previous moment. To control the decay rate of information from the previous moment. The proportion of motion speed and direction information incorporated into the dynamic weighting factor during template updates. Let be the motion indicator function, if at pixel (x, y) at time... If motion is detected, the value is 1; otherwise, it is 0. This represents the motion template value at the previous moment, reflecting historical motion information. This represents the inter-frame time interval. Based on target speed The dynamic weighting factor for the direction calculation is defined as shown in Equation 20 below:
[0156]
[0157] In the formula: , indicating the current direction of movement of the target. This is a preset reference direction used to correct for differences in motion direction.
[0158] S34, Cubic spline interpolation smoothing and multi-scale fusion.
[0159] To make the constructed motion template smoother and able to capture motion features at different time scales, cubic spline interpolation is used to reconstruct the template data within each time window. The specific interpolation formula used is shown in Equation 21 below:
[0160]
[0161] In the formula: These are cubic B-spline basis functions used for smooth interpolation in the time domain. These are the interpolation coefficients calculated within a local window. Let be the order of the spline function.
[0162] Meanwhile, to adapt to the multi-scale characteristics of the target motion, the motion templates constructed in different time windows are weighted and fused to obtain the final template, as shown in Equation 22 below:
[0163]
[0164] In the formula: Indicates the first Smooth motion templates at various scales. To integrate weights, satisfy . The total number of time scales.
[0165] The construction of temporal motion templates can not only capture the motion intensity of a target at a single moment, but also comprehensively consider the temporal continuity, speed and direction changes of the motion. At the same time, the robustness and subtlety of the template expression are enhanced by using cubic spline reconstruction and smoothing techniques, providing richer and more accurate spatiotemporal dynamic features for subsequent motion feature extraction and accident detection.
[0166] S4. Feature Extraction and Fusion
[0167] This paper describes the joint extraction and fusion of multi-scale spatial and temporal dynamic information from preprocessed vehicle monitoring video sequences. Multi-scale complex Gabor filtering and scale integration mechanisms are used to capture local texture and structural information of the images. Simultaneously, local temporal weighted integration and higher-order derivative operations are employed to reveal the dynamic characteristics of moving targets. Deep fusion of spatial and temporal information is achieved through nested integration, layer-by-layer summation, and nonlinear mapping strategies. After normalization and dimensionality reduction mapping, the generated comprehensive features possess high-dimensional discriminative power, providing stable and accurate input data for accident prediction and detection based on deep neural networks.
[0168] Specifically, the steps include the following:
[0169] S41, Multi-scale spatial feature extraction.
[0170] Preprocessing images for each frame ( For spatial coordinates, For time (i.e., when), multi-scale complex Gabor filtering and scale integration are nested to extract local texture and structural information, as defined in Equation 23 below:
[0171]
[0172] In the formula: This indicates taking the real part of the complex number. This represents the number of filters. The weight is a dimensionless weight. For the first The Gabor filter is defined as shown in Equation 24 below:
[0173]
[0174] In the formula: λ is the wavelength. This indicates the filter direction. The standard deviation is denoted as . The aspect ratio is 1. This is the phase shift. This is the scaling weight function.
[0175] S42, High-order temporal dynamic feature extraction.
[0176] To capture the temporal dynamics of a moving target, a nested formula incorporating local time-weighted integrals and higher-order derivatives is designed, as shown in Equation 25 below:
[0177]
[0178] In the formula: Let be the highest order of the derivative under consideration. For the corresponding Dimensionless weights of time features. It reflects the instantaneous changes in motion. The weight function is a normal distribution. Offset to the center time. To expand the scope. This indicates the size of the time domain for extracting high-order dynamic features.
[0179] S43, Deep fusion of multimodal features.
[0180] The spatial and temporal dynamic features extracted by S41 and S42 are deeply fused. Multi-layer nested integrals, summations, and nonlinear mappings are used to construct a comprehensive feature representation. The fusion formula is shown in Equation 26 below:
[0181]
[0182] In the formula: This is the activation function. This represents the number of local region divisions. For the first Fusion weights for local regions. Indicates the first The integration domain of a local region. This is the normalization factor. This is the local neighborhood offset, used to represent the local sampling offset relative to the reference position (x, y). These are spatial, temporal, and interactive features in the region. The weights in the algorithm can be adaptively learned through subsequent backpropagation. Let be the value of the spatial feature from time t at the local offset position. The time feature from time t to t takes the value at the same offset position. This refers to the interaction between spatial and temporal characteristics at this location.
[0183] S44, Feature Normalization and Dimensionality Reduction Mapping.
[0184] To improve the robustness and computational efficiency of fused features, Normalization and dimensionality reduction are performed using nested optimal mapping, as shown in Equation 27 below:
[0185]
[0186] In the formula: Definition . and ε and φ are the mean and standard deviation of the fused features on the training samples, respectively. ϵ is a small constant to prevent the denominator from being zero. is the dimension reduction mapping matrix. b is the bias vector. The target is a low-dimensional feature representation.
[0187] S5. Selection of features for fuzzy time slices.
[0188] Fine-grained time segmentation is performed on the temporal motion template data, and fuzzy membership functions are constructed to evaluate the importance of each time segment. A multi-layer time weighting strategy is adopted to weight and aggregate the original features and their first- and second-order transformations within each time slice, and noise and redundant information are filtered out through nonlinear mapping and fuzzy logic. After key time windows are selected based on fuzzy confidence, the features of each slice are fused to form a highly discriminative dynamic feature representation, providing input data for accident prediction and detection based on deep neural networks.
[0189] Specifically, the steps include the following:
[0190] S51. Time slice partitioning and fuzzy membership function construction.
[0191] Using an adaptive sliding window method to divide the overall time interval Divided into several sub-intervals Within each time slice, a fuzzy membership function is constructed, and a weight is assigned to each time point, as shown in Equation 28 below:
[0192]
[0193] In the formula: For the first The center moment of each slice To control the scale parameters of the time-varying range, As a shape control factor, The time interval is defined as follows.
[0194] S52, Multi-layer time-weighted feature aggregation.
[0195] Within each slice, the original temporal features are weighted and integrated, while first-order and second-order temporal variation information are fused to construct the feature aggregation formula as shown in Equation 29 below:
[0196]
[0197] In the formula: The motion template features at time t. and These represent the first and second order changes of the feature, respectively. , and These are the layer weighting coefficients, which can be adaptively optimized through subsequent training.
[0198] S53. Fuzzy logic filtering for redundancy removal and noise suppression.
[0199] To remove redundant and noisy information, the aggregation features of each slice are... After performing nonlinear mapping processing, the composite filter formula is designed as shown in Equation 30 below:
[0200]
[0201] In the formula: It serves as a balance factor between linear and nonlinear processing. and Adjusting the shape of the nonlinear mapping. Used for noise reduction and smoothing. This indicates a fuzzy logic normalization operation, ensuring that the output is within a preset range.
[0202] S54. Determination of key time windows and feature output.
[0203] Based on the integral value of the fuzzy membership function within each slice, the key confidence level is defined as shown in Equation 31 below:
[0204]
[0205] In the formula: For the first Time slice The integral value of the fuzzy membership function. Let be the integration variable. Choose one that satisfies... ( A slice with a preset threshold is used as the key window. Finally, the filtered feature vectors within each key window are output as a dynamic feature representation through feature concatenation or fusion operations, as shown in Equation 32 below:
[0206]
[0207] In the formula: , The feature vector within the selected key window. The filter condition indicates that only the key confidence level is selected. Greater than the preset threshold The design of the fuzzy time slice feature module enables fine-grained fuzzy time slice division, dynamic weighted feature aggregation, redundancy filtering, and key window determination of the temporal motion template, forming a dynamic feature representation with high-dimensional discriminative power. This provides innovative and complex input data for subsequent accident prediction and detection based on deep neural networks.
[0208] S6. Accident Prediction and Detection Based on Deep Neural Networks
[0209] By integrating spatial, temporal, and fuzzy time-slice features using a hybrid model, real-time prediction of accident probabilities is achieved. This step constructs a multi-layer convolutional and recurrent unit network for deep feature extraction and dynamic modeling, enhances the response to key information through a self-attention mechanism, completes feature mapping and output through fully connected layers, and designs a composite loss function and online feedback update strategy to adaptively optimize the model, ensuring high accuracy and robustness of the system.
[0210] S61: Feature preprocessing and branch fusion.
[0211] Spatial features, temporal features, and fuzzy time slice features are fused into a unified representation through nonlinear transformation, and the calculation formula is shown in Equation 33 below:
[0212]
[0213] In the formula: , and For adaptive weights, This represents a non-linear activation function.
[0214] S62, Convolutional layer feature extraction.
[0215] Local features are extracted using multi-layer convolution operations, and a complex convolution expression is constructed through nested integrals, as shown in Equation 34 below:
[0216]
[0217] In the formula: The feature map of the current convolutional layer output. For convolution kernel function, For activation function, For the convolutional layer bias term, the integration region Covering a local spatiotemporal range, Current position Relative to the current position The local offset.
[0218] S63, Dynamic modeling of loop layers.
[0219] A recurrent neural network is used for time dynamic modeling. The dynamic state is constructed by weighted time integral and nonlinear mapping, as shown in Equation 35 below:
[0220]
[0221] In the formula: For time-weighted functions, This indicates the state update operation of the loop unit. For a moment Input features, For a moment The hidden layer state, Define the integration interval.
[0222] S64, Attention Mechanism Weighted Fusion.
[0223] By focusing on key dynamic information through a self-attention module, an attention score and context vector are constructed, as shown in Equation 36 below:
[0224]
[0225] In the formula: , and For attention parameters, It is the hidden layer state at all times. The weighted sum.
[0226] S65, Fully Connected Layer Prediction Mapping.
[0227] The context vector is mapped to the accident probability prediction output through a fully connected layer, as shown in Equation 37 below:
[0228]
[0229] In the formula: For a suitable activation function, For bias terms of fully connected layers, This is the weight matrix of the output layer.
[0230] S66. Design and optimization of composite loss function.
[0231] The composite loss function, which includes cross-entropy loss, regularization term, and auxiliary consistency loss, is constructed as shown in Equation 38 below:
[0232]
[0233] In the formula: , and To adjust the weights of different loss contributions, an adaptive optimization algorithm is used to update the model parameters. For cross-entropy loss, For weight decay, To mitigate consistency loss.
[0234] S67. Online learning and feedback updates.
[0235] The model is adjusted online based on real-time feedback, using an incremental update strategy, as shown in Equation 39 below:
[0236]
[0237] In the formula: For learning rate, Reflecting the amount of feedback correction, For the updated model parameters, For the current model parameters, Represents the loss function For the current model parameters The gradient.
[0238] By designing a multi-layered hybrid neural network structure, a deep dynamic prediction model is constructed using nested integrals, attention weighting, and composite loss, thereby enabling real-time and accurate determination of the probability of accident occurrence.
[0239] S7. Post-processing and alarms.
[0240] To address the accident prediction results output by deep neural networks, a multi-level post-processing and decision-making mechanism is constructed. The prediction results undergo temporal smoothing, noise suppression, and continuity correction. An alarm triggering model is constructed by incorporating time accumulation and decay effects. Through complex nonlinear mapping, a graded identification of accident risks is achieved, thereby determining different response levels. Simultaneously, interfaces with monitoring and vehicle control systems are designed to enable real-time transmission of alarm information. A feedback model is established to statistically analyze and optimize alarm records, continuously improving system robustness and prediction accuracy.
[0241] Specifically, the steps include the following:
[0242] S71. Temporal smoothing and continuity correction of accident prediction results.
[0243] A smoothing function is constructed using nested integrals and recursive filtering to smooth the original accident prediction signal output by the deep neural network. To perform nonlinear smoothing and delayed feedback correction, the following composite smoothing formula is designed, as shown in Equation 40:
[0244]
[0245] In the formula: This is the normalization factor. Control the width of the smoothing kernel. Define the range of the integration window. This represents the feedback weight of the smoothing result from the previous time step. For feedback delay. It is a non-linear activation function used to suppress abnormal fluctuations.
[0246] S72, Alarm triggering mechanism based on predicted probability and threshold.
[0247] Construct a trigger function to predict the smoothed signal. By integrating the time-cumulative effect and the exponential decay model, an alarm decision signal is formed. As shown in Equation 41 below:
[0248]
[0249] In the formula: This represents the length of the time window. This is the attenuation factor. This is the trigger threshold. This indicates an activation or step function that converts continuous values into alarm indication signals. This represents the local neighborhood offset.
[0250] S73, Multi-level alarm strategy construction.
[0251] Design a hierarchical decision-making model to determine the alarm level based on the integral response of the cumulative alarm signal over different time periods. The nested integral-exponential mapping model is constructed as shown in equations 42 and 43 below:
[0252]
[0253]
[0254] In the formula: This represents the time decay rate. Define the integration interval. and These are the threshold values for primary and emergency alarms, respectively. The values represent different alarm levels, corresponding to primary alert, emergency alarm, and automatic response, respectively.
[0255] S74, Interface Design and Real-time Transmission of Alarm Information.
[0256] Design an alarm information encoding and transmission model, including alarm levels. Real-time transmission is achieved by combining delay compensation and data fusion. The transmission coding formula is constructed as shown in Equation 44 below:
[0257]
[0258] In the formula: A numerical code indicating the alarm level. Indicates transmission delay. This is the time delay compensation factor. The weights are adjusted based on feedback. This indicates real-time diagnostic data or correction signals.
[0259] S75, Alarm Feedback Mechanism and Performance Evaluation.
[0260] A composite feedback model is constructed to record, statistically analyze, and update subsequent decision parameters based on alarm results. The composite loss function is designed as shown in Equation 45 below:
[0261]
[0262] In the formula: and The losses are false alarms and missed alarms, respectively. The third constraint is continuity. It measures the smoothness of changes in the alarm signal. , and These are the corresponding weighting coefficients.
[0263] By using an adaptive optimization strategy to update the overall system parameters online, the alarm mechanism can be continuously improved.
[0264] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0265] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing, characterized in that, The steps include the following: S1. Video data acquisition and preprocessing: This involves sequentially processing real-time vehicle monitoring videos to provide stable and accurate data support, including: S11, Frame segmentation processing, converting the acquired video data into a discrete image sequence; S12. Noise reduction processing, based on Gaussian filter processing to eliminate the influence of factors; S13. Image enhancement processing, based on histogram equalization processing, improves image contrast; S2. Moving target detection and dynamic tracking to achieve stable tracking of multiple targets in complex scenes, including: S21. Moving target detection: background modeling and foreground segmentation based on Gaussian mixture model (GMM). S22. Target region extraction: Extract the binary image obtained from the segmentation in S21. Perform row connectivity analysis and group adjacent non-zero pixels into the same target region; S23. Dynamic target tracking: Based on the dynamic tracking algorithm and combined with the Kalman filter, the state of the detected target is estimated and updated. S3. Construction of temporal motion templates to capture continuous motion features of the target and enhance sensitivity to anomalous time periods, including: S31. Motion trajectory data acquisition and preprocessing: Based on dynamic tracking, obtain the position and velocity data of each target in consecutive frames; S32, Temporal window partitioning and fuzzy weight assignment: Based on the sliding window technique, several fixed-duration windows are used to divide the continuous motion trajectory. S33. Multi-level motion template construction to fully consider the motion indication of the current frame, as well as the template information and target motion speed and direction factors of the previous moment; S34. Cubic spline interpolation smoothing and multi-scale fusion are used to obtain the final template; S4. Feature extraction and fusion: Utilizing spatial and temporal information from the video to improve feature discriminative power, including: S41. Multi-scale spatial feature extraction, based on multi-scale complex Gabor filtering and scale integration nesting for preprocessing each frame of the image. Extract local texture and structural information; S42. Extraction of high-order temporal dynamic features: Establishing a nested formula that includes local time weighted integrals and high-order derivatives; S43. Deep fusion of multimodal features: Based on multi-layer nested integration, summation and nonlinear mapping, the spatial and temporal features extracted in steps S41 and S42 are deeply fused. S44. Feature normalization and dimensionality reduction mapping: The fusion results in S43 are normalized and dimensionality reduced based on nested optimization mapping. S5. Fuzzy time slice feature selection to finely segment key moments and highlight important information before and after the accident, including: S51. Time slice division and fuzzy membership function construction: The overall time interval is divided into several sub-intervals based on the adaptive sliding window method, and a fuzzy membership function is constructed in each time slice to assign weights to each moment. S52. Multi-layer time-weighted feature aggregation: In each time slice, the original time-series features are weighted and integrated, and first-order and second-order time change information are fused at the same time. S53. Fuzzy logic filtering for redundancy removal and noise suppression, performing nonlinear mapping processing on the aggregation features of each slice; S54. Determination of key time windows and feature output: Define key confidence and select slices as key windows; S6. Accident prediction and detection based on deep neural networks: By integrating multimodal features, accurate prediction of accident probability is achieved, including: S61. Feature processing and branch fusion: Spatial features, temporal features, and fuzzy time slice features are fused through nonlinear transformation. S62. Convolutional layer feature extraction: Local features are extracted based on multi-layer convolution, and convolutional expressions are constructed through nested integrals. S63. Dynamic modeling of recurrent layers: Time dynamic modeling based on recurrent neural networks combined with weighted time integrals and nonlinear mapping. S64. Weighted fusion of attention mechanisms, combined with self-attention modules to focus on key dynamic information, to construct attention scores and context vectors; S65, Fully Connected Layer Prediction Mapping: Maps the context vector in S65 to the accident probability prediction output through the fully connected layer; S66. Design and optimization of composite loss function: Construct a composite loss function including cross-entropy loss, regularization term and auxiliary consistency loss. S67. Online learning and feedback updates, based on an incremental update strategy, are adjusted online through real-time feedback; S7. Post-processing and Alarm: By constructing a multi-level post-processing and decision-making mechanism, accident risk classification identification and alarm are achieved, including: S71. Temporal smoothing and continuity correction of accident prediction results: a smoothing function is constructed based on nested integrals and recursive filtering. S72. Based on the prediction probability and threshold, an alarm triggering mechanism is constructed to generate an alarm decision signal. S73. Multi-level alarm strategy construction: Define a hierarchical decision model, and determine the alarm level based on the integral response of the accumulated alarm decision signal over different time periods. ; S74. Interface Design and Real-time Alarm Information Transmission: Design an alarm information encoding and transmission model, specifying alarm levels. Real-time transmission combining latency compensation and data fusion; S75. Alarm feedback mechanism and performance evaluation: Construct a composite feedback model to process and update alarm results.
2. The real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing as described in claim 1, characterized in that: In S21, Gaussian Mixture Model (GMM) background modeling is performed by analyzing a pixel at time... Observations Modeling as The weighted sum of Gaussian distributions is shown in the following formula: ; In the formula: The number of mixture components represents the number of Gaussian distributions used to model each pixel; Indicates the first The Gaussian component at time... The weights below; Indicates the first The Gaussian components at time... The mean of the following; Indicates the first The Gaussian component at time... The standard deviation below; in Let represent the Gaussian probability density function, and its formula is shown below: ; This function is used to calculate observed values. In the The probability of occurrence under a given distribution.
3. The real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing as described in claim 1, characterized in that: The multi-level motion template construction in S33 is established using the following formula: ; In the formula: This represents the direct contribution weight of the motion information in the current frame; The retention weight for template information from the previous moment; To control the decay rate of information from the previous moment; The proportion of motion speed and direction information incorporated into the dynamic weighting factor during template updates; Let be the motion indicator function, if at pixel (x, y) at time... If motion is detected, the value is 1; otherwise, the value is 0. This represents the motion template value from the previous moment, reflecting historical motion information. This refers to the inter-frame time interval. in The dynamic weighting factor, calculated based on the target velocity v and direction, is defined as follows: ; In the formula: , indicating the current direction of the target's movement; This is a preset reference direction used to correct for differences in motion direction.
4. The real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing as described in claim 1, characterized in that: S34 cubic spline interpolation smoothing and multi-scale fusion are used to obtain the final template, and the calculation formula is shown below: ; In the formula: Indicates the first Smooth motion templates at various scales; To integrate weights, satisfy ; The total number of time scales.
5. The real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing as described in claim 1, characterized in that: The Gabor filter in S41 is defined as follows: ; In the formula: Wavelength; In the direction of the filter; Standard deviation; Aspect ratio; For phase shift; This is the scaling weight function.
6. The real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing as described in claim 1, characterized in that: The formula for the fuzzy membership function constructed in S51 is shown below: ; In the formula: For the first The center moment of each slice; Scale parameters for controlling the temporal spread range; As a shape control factor, The time interval is defined as follows.
7. The real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing as described in claim 1, characterized in that: In S52, the following feature aggregation formula is constructed to achieve multi-level time-weighted feature aggregation: ; In the formula: The motion template features at time t; and These represent the first and second order changes of the feature, respectively; , and These are the layer weighting coefficients, which can be adaptively optimized through subsequent training.
8. The real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing as described in claim 1, characterized in that: The formula for the composite loss function constructed in S66 is shown below: ; In the formula: , and To adjust the weights of different loss contributions, For cross-entropy loss, For weight decay, To mitigate consistency loss.
9. The real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing as described in claim 1, characterized in that: The smoothing function in S71 is constructed as follows: ; In the formula: Normalization factor; Control the width of the smoothing kernel; Define the range of the integration window; This represents the feedback weight of the smoothing result from the previous time step; For feedback delay; It is a non-linear activation function.
10. The real-time vehicle accident identification method based on multi-scale feature fusion and fuzzy time slicing as described in claim 1, characterized in that: S73 defines a hierarchical decision-making model, which is constructed by a nested integral and exponential mapping model, as shown in the following equation: ; ; In the formula: The decay rate over time; Define the integration interval; and These are the threshold values for determining primary and emergency alarms, respectively. The values represent different alarm levels, corresponding to primary alert, emergency alarm, and automatic response, respectively.