A method to improve event detection accuracy based on multi-model fusion
Through multi-model fusion and multi-modal data preprocessing technology, the problems of event detection accuracy and robustness in complex monitoring scenarios are solved, and the detection effects of high accuracy, low false alarm rate and good real-time performance are achieved.
Patent Information
- Application Number
- CN202510280132.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The prior art is difficult to achieve high accuracy and robust event detection in complex and variable monitoring scenarios, especially when environmental factors change, detection performance is significantly reduced and it is difficult to balance real-time and false alarm rates.
Multi-model fusion technology is adopted to generate a standardized feature matrix through multimodal data preprocessing, and feature-level fusion is performed by combining YOLOv5 improved model, ResNet-Transformer hybrid model and LSTM spatiotemporal analysis model, and the detection results are verified through the three-level logic judgment mechanism and the false positive filtering module.
It improves the accuracy and robustness of event detection, effectively eliminates shadow, reflection and noise interference, reduces false alarm rates, ensures the reliability of output results, and maintains good real-time performance in complex scenarios.
Smart Images

Figure CN119785162B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and video surveillance, and more specifically, to a method for improving event detection accuracy based on multi-model fusion. Background Art
[0002] In the field of video surveillance, event detection is one of the key technologies for realizing intelligent monitoring and is widely used in security, traffic management, smart cities and other fields. Existing event detection methods mainly rely on a single model or simple feature extraction technology, such as traditional background subtraction and optical flow methods. These methods can achieve certain results in simple scenarios, but in complex environments, such as illumination changes, shadow interference, low illumination, etc., the detection accuracy will drop significantly. In recent years, the development of deep learning technology has brought new opportunities for event detection, such as target detection models based on convolutional neural networks (CNNs) and spatiotemporal analysis models based on recurrent neural networks (RNNs). However, a single deep learning model still has limitations when facing complex and changing monitoring scenarios, such as insufficient detection capabilities for targets of different scales and inaccurate recognition of abnormal trajectories and aggregation behaviors.
[0003] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the prior art: a single model is difficult to cope with complex and changeable monitoring scenarios, resulting in insufficient accuracy and robustness of event detection; the existing methods have poor adaptability to environmental factors (such as light, visibility, etc.), especially at night or in severe weather conditions, the detection performance will be significantly reduced; in addition, the prior art is difficult to achieve a balance between real-time and false alarm rate, and it is difficult to meet the dual requirements of actual monitoring systems for high real-time and low false alarm rate. Summary of the invention
[0004] The present invention provides a method for improving the accuracy of event detection based on multi-model fusion, comprising the following steps: S1. performing multi-modal data preprocessing on a monitoring video stream to generate a standardized feature matrix.
[0005] S2. The outputs of the YOLOv5 improved model, the ResNet-Transformer hybrid model, and the LSTM spatiotemporal analysis model are fused at the feature level through the adaptive weighted fusion module.
[0006] S3. Perform multi-dimensional event detection based on the fused feature matrix and generate a preliminary set of detection results.
[0007] S4. Establish a three-level logical judgment mechanism to verify the confidence of the test results.
[0008] S5. Output the verified final event detection results to the monitoring system.
[0009] Furthermore, the S1 includes: S11. Using an improved Retinex algorithm to perform image enhancement processing to eliminate shadow and reflection interference.
[0010] S12. Apply an adaptive bilateral filter to perform noise suppression, and the filter parameters are dynamically adjusted according to the environmental visibility.
[0011] S13. Construct a spatiotemporal feature pyramid and extract multi-scale spatiotemporal features.
[0012] Among them, the illumination component estimation formula of the improved Retinex algorithm is: ; In the formula, is the illumination component, For color channels The pixel value of is the Gaussian kernel function, , It is the real-time visibility parameter.
[0013] Furthermore, the adaptive weighted fusion module in S2 adopts a dynamic attention mechanism, and the fusion formula is: ; In the formula, To fusion features, is the dynamic weight coefficient, is the sigmoid function, is the learnable parameter matrix, is the environmental feature vector, including light intensity, visibility, and rainfall parameters.
[0014] Further, the S3 includes: S31. performing target recognition based on the improved YOLOv5 model, using the CIoU loss function and the adaptive anchor box mechanism; S32. applying the dynamic time warping algorithm (DTW) to perform abnormal trajectory analysis; S33. detecting abnormal aggregation behavior through a density-aware network; wherein the loss function of the improved YOLOv5 model is: ; In the formula, , is the target size parameter, .
[0015] Furthermore, the three-level logic judgment mechanism of S4 includes: S41. First-level judgment: based on the confidence threshold Conduct a preliminary screening.
[0016] S42. Second level judgment: trajectory continuity analysis is performed through the spatiotemporal consistency verification module.
[0017] S43. Third level judgment: Applying Bayesian network to calculate posterior probability ,when It is considered as a valid event.
[0018] Furthermore, it also includes a night mode processing module, specifically including: S61. Using a dark channel prior defogging algorithm to process low-light images.
[0019] S62. Apply improved Retinex-Net for illumination component decomposition.
[0020] S63. Amplifying effective motion signals through motion feature enhancement network.
[0021] The improved Retinex-Net reflection component estimation formula is: ; In the formula, is the adaptive mask, , is the local contrast parameter.
[0022] Furthermore, it also includes a real-time optimization module, which is specifically implemented as follows: S71. Establish a priority queue management detection task, and the queue is sorted according to ; In the formula, , is the event urgency parameter, The duration of the event.
[0023] S72. Use model sharding loading technology to dynamically allocate computing resources.
[0024] S73. Apply feature cache reuse mechanism to reduce repeated calculations.
[0025] Further, the adaptive bilateral filter parameters in S12 are set to: ; In the formula, is the estimated noise level, is the image contrast parameter.
[0026] Furthermore, the spatiotemporal consistency verification of S42 adopts a trajectory prediction model: ; In the formula, To predict the location, is the current position, is the motion control quantity, is the adaptive weight coefficient, is the historical trajectory variance.
[0027] Furthermore, it also includes a false alarm filtering module, which is specifically implemented as follows: S101. Establish a multi-dimensional feature database to store historical false alarm samples.
[0028] S102. Apply contrastive learning network to calculate the similarity between the current detection result and the false positive sample: .
[0029] S103. When The false alarm filtering mechanism is activated and the detection result is not output to the monitoring system.
[0030] According to the above-mentioned embodiments of the present invention, at least the following beneficial effects are achieved: the present invention can improve the accuracy and robustness of event detection through multimodal data preprocessing and multi-model fusion technology. The improved Retinex algorithm and adaptive bilateral filter can effectively eliminate shadows, reflections and noise interference, enhance image quality, and provide clearer input data for subsequent detection. At the same time, the feature-level output of the fusion YOLOv5 improved model, ResNet Transformer hybrid model and LSTM spatiotemporal analysis model can make full use of the advantages of each model to achieve accurate target identification, abnormal trajectory analysis and abnormal aggregation behavior detection, thereby improving event detection performance in complex scenarios. In addition, the three-level logic judgment mechanism and false alarm filtering module can further verify the confidence of the detection results, effectively reduce the false alarm rate, and ensure the reliability of the output results.
[0031] The present invention also has good real-time optimization capabilities and can meet the needs of actual monitoring systems. By establishing a priority queue to manage detection tasks, adopting model sharding loading technology and feature cache reuse mechanism, it can dynamically allocate computing resources, reduce repeated calculations, and improve detection efficiency. At the same time, the night mode processing module can optimize processing for low-light environments, further improving the detection performance at night or under harsh lighting conditions, so that the present invention can be widely used in monitoring event detection in various complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood by reading the following detailed description with reference to the accompanying drawings, in which several embodiments of the present invention are shown by way of example and not limitation.
[0033] in: Figure 1 A flowchart of a method for improving event detection accuracy based on multi-model fusion provided in one embodiment of the present invention. DETAILED DESCRIPTION
[0034] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0035] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, apparatus, method or computer program product. Therefore, the present invention can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0036] It should be noted that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0037] Reference below Figure 1 , Figure 1 A flow chart of a method for improving event detection accuracy based on multi-model fusion provided by an embodiment of the present invention. Figure 1 As shown, a method 100 for improving the accuracy of event detection based on multi-model fusion includes the following steps: S1. performing multi-modal data preprocessing on the monitoring video stream to generate a standardized feature matrix.
[0038] S2. The outputs of the YOLOv5 improved model, the ResNet-Transformer hybrid model, and the LSTM spatiotemporal analysis model are fused at the feature level through the adaptive weighted fusion module.
[0039] S3. Perform multi-dimensional event detection based on the fused feature matrix and generate a preliminary set of detection results.
[0040] S4. Establish a three-level logical judgment mechanism to verify the confidence of the test results.
[0041] S5. Output the verified final event detection results to the monitoring system.
[0042] It should be noted that the present invention proposes a method for improving the accuracy of event detection based on multi-model fusion, and the method realizes efficient detection of events in surveillance video streams through a series of steps. In the multimodal data preprocessing stage, the generation of a standardized feature matrix is completed by performing operations such as image enhancement, noise suppression and feature extraction on the surveillance video stream. Among them, the improved Retinex algorithm is a technology for image enhancement, which eliminates shadow and reflection interference by estimating the illumination component, thereby improving the clarity and contrast of the image. The adaptive bilateral filter is a noise suppression technology, and its parameters are dynamically adjusted according to the environmental visibility to adapt to the noise level in different environments. The spatiotemporal feature pyramid is a multi-scale feature extraction method used to capture the spatiotemporal information in the video stream and provide a rich feature basis for subsequent event detection.
[0043] Specifically, the improved Retinex algorithm achieves image enhancement through a specific illumination component estimation formula. The Gaussian kernel function in the formula is used to smooth the color channels of the image, while the real-time visibility parameter is used to dynamically adjust the parameters of the algorithm to adapt to different lighting and environmental conditions. The parameter setting of the adaptive bilateral filter is based on the noise level estimate and the image contrast parameter. These parameters are calculated through a specific formula and can automatically adjust the performance of the filter according to changes in the environment. The construction of the spatiotemporal feature pyramid involves the extraction of spatiotemporal features of different scales in the video stream. These features can include the motion trajectory of the target, the shape change of the target, etc., providing multi-dimensional feature support for subsequent event detection.
[0044] Preferably, in the multimodal data preprocessing stage, the parameter settings of the Retinex algorithm can be further optimized and improved. For example, the real-time visibility parameters can be obtained by analyzing the image brightness distribution in the video stream to more accurately reflect the lighting conditions of the current environment. At the same time, the parameters of the adaptive bilateral filter can be adjusted according to the actual application scenario, such as increasing the filtering intensity in a high-contrast environment to better suppress noise. In addition, the construction of the spatiotemporal feature pyramid can be combined with deep learning technology to automatically extract more representative spatiotemporal features by training convolutional neural networks, thereby further improving the performance of event detection.
[0045] In some embodiments, the S1 includes: S11. Using an improved Retinex algorithm to perform image enhancement processing to eliminate shadow and reflection interference.
[0046] S12. Apply an adaptive bilateral filter to perform noise suppression, and the filter parameters are dynamically adjusted according to the environmental visibility.
[0047] S13. Construct a spatiotemporal feature pyramid and extract multi-scale spatiotemporal features.
[0048] Among them, the illumination component estimation formula of the improved Retinex algorithm is: ; In the formula, is the illumination component, For color channels The pixel value of is the Gaussian kernel function, , It is the real-time visibility parameter.
[0049] It should be noted that the present invention adopts an improved Retinex algorithm to perform image enhancement processing in the multimodal data preprocessing stage to eliminate shadow and reflection interference. The improved Retinex algorithm is an image enhancement technology based on illumination component estimation. It enhances the details and contrast of the image by separating the illumination component and the reflection component of the image. At the same time, an adaptive bilateral filter is applied for noise suppression, and the filter parameters are dynamically adjusted according to the environmental visibility to adapt to different environmental conditions. In addition, a spatiotemporal feature pyramid is constructed to extract multi-scale spatiotemporal features, which helps to capture dynamic information in the video stream. Among them, the illumination component estimation formula of the improved Retinex algorithm is: ; In the formula, is the illumination component, For color channels The pixel value of is the Gaussian kernel function, , It is the real-time visibility parameter.
[0050] Specifically, the improved Retinex algorithm smoothes the pixel values of each color channel using a Gaussian kernel function to estimate the illumination component. Dynamically adjusted according to real-time visibility parameters, which can be obtained by analyzing the image brightness distribution in the video stream, for example, by calculating the image contrast or brightness variance. The parameter setting of the adaptive bilateral filter is based on the noise level estimate and the image contrast parameter, and the specific formula is: ;in, is the estimated noise level, is the image contrast parameter. The dynamic adjustment of these parameters enables the filter to automatically optimize performance according to different environmental conditions. The construction of the spatiotemporal feature pyramid involves the extraction of spatiotemporal features of different scales in the video stream. These features can include the motion trajectory of the target, the shape change of the target, etc., providing multi-dimensional feature support for subsequent event detection.
[0051] Preferably, in the multimodal data preprocessing stage, the parameter settings of the Retinex algorithm can be further optimized and improved. For example, the real-time visibility parameter can be obtained by analyzing the image brightness distribution in the video stream, and specifically by calculating the brightness histogram or brightness variance of the image to reflect the lighting conditions of the current environment. At the same time, the parameters of the adaptive bilateral filter can be adjusted according to the actual application scenario, such as increasing the filtering intensity in a high-contrast environment to better suppress noise.
[0052] Furthermore, the construction of spatiotemporal feature pyramids can be combined with deep learning technology to automatically extract more representative spatiotemporal features by training convolutional neural networks, thereby further improving the performance of event detection. For example, pre-trained convolutional neural networks (such as ResNet or VGG) can be used to extract features, and the feature extraction process can be further optimized through transfer learning.
[0053] In some embodiments, the adaptive weighted fusion module in S2 adopts a dynamic attention mechanism, and the fusion formula is: ; In the formula, To fusion features, is the dynamic weight coefficient, is the sigmoid function, is the learnable parameter matrix, is the environmental feature vector, including light intensity, visibility, and rainfall parameters.
[0054] It should be noted that the present invention adopts an adaptive weighted fusion module in the feature fusion stage, which fuses the output features of different models through a dynamic attention mechanism. The core of the adaptive weighted fusion module lies in the calculation of dynamic weight coefficients, which can be dynamically adjusted according to the characteristics of the input data, so as to better combine the advantages of each model. Specifically, the fusion formula is: ;in, To fusion features, is the dynamic weight coefficient, is the sigmoid function, is a learnable parameter matrix, and E is an environmental feature vector, including parameters such as light intensity, visibility, and rainfall. This dynamic weighted fusion method can effectively improve the flexibility and adaptability of feature fusion.
[0055] Specifically, the adaptive weighted fusion module calculates the weight coefficient of each model through the dynamic attention mechanism. It includes parameters such as light intensity, visibility, and rainfall, which can be collected by sensors or estimated from video streams. For example, light intensity can be calculated from the brightness histogram of the image, visibility can be obtained from the contrast analysis of the image, and rainfall can be estimated from the texture features of raindrops in the image. Learned through the training process, used to map the environmental feature vector to the weight coefficient. It is used to normalize the weight coefficient to the range of [0,1] to ensure the rationality and stability of the weight coefficient. This dynamic weighting mechanism can automatically adjust the weights of each model according to different environmental conditions and input data characteristics, thereby achieving a better feature fusion effect.
[0056] Preferably, in the adaptive weighted fusion module, the parameter extraction method of the environmental feature vector can be further optimized. For example, the light intensity can be estimated by calculating the average brightness value of the image, the visibility can be estimated by analyzing the contrast distribution of the image, and the rainfall can be estimated by detecting the raindrop texture features in the image.
[0057] Furthermore, the parameter matrix can be learned It can be trained through a deep learning framework (such as TensorFlow or PyTorch). During the training process, the labeled data set can be used to optimize the parameter matrix to improve the accuracy of the weight coefficient. As an alternative, in addition to using the sigmoid function for normalization, the softmax function can also be used to calculate the weight coefficient. The softmax function can normalize the weight coefficient to a probability distribution, further improving the rationality and stability of the weight coefficient.
[0058] In some embodiments, S3 includes: S31. Target recognition is performed based on an improved YOLOv5 model, using a CIoU loss function and an adaptive anchor box mechanism.
[0059] S32. Apply the dynamic time warping (DTW) algorithm to perform abnormal trajectory analysis.
[0060] S33. Detect abnormal aggregation behavior through density-aware network; the loss function of the improved YOLOv5 model is: ; In the formula, , is the target size parameter, .
[0061] It should be noted that the present invention adopts an improved The model performs target recognition and combines the dynamic time warping algorithm (DTW) and density-aware network to perform abnormal trajectory analysis and abnormal aggregation behavior detection. The model improves the accuracy and robustness of target detection by optimizing the loss function and adaptive anchor box mechanism. The loss function of the model is: ;in, , is the target size parameter, The CIoU loss function is an improved IoU (intersection over union) loss function that can better handle the regression problem of the target box. The adaptive anchor box mechanism dynamically adjusts the size of the anchor box according to the size of the target, thereby improving the detection accuracy.
[0062] Specifically, the improved YOLOv5 model optimizes the regression process of the target box through the CIoU loss function. The CIoU loss function not only considers the overlapping area of the target box, but also introduces the penalty terms of shape and scale, so as to more accurately adjust the position and size of the target box. The adaptive anchor box mechanism dynamically adjusts the size of the anchor box according to the size of the target. The specific formula is: ;in, is the target size parameter, which indicates the relative size of the target in the image. In this way, the model can better adapt to targets of different sizes and improve detection accuracy. The dynamic time warping algorithm (DTW) is used to analyze the motion trajectory of the target and detect abnormal trajectories by calculating the similarity between trajectories. The density-aware network detects abnormal aggregation behaviors by analyzing the distribution density of the target, such as abnormal aggregation of people or congestion of vehicles.
[0063] Preferably, in the multi-dimensional event detection stage, the parameter settings of the improved YOLOv5 model can be further optimized. For example, the target size parameter It can be obtained by analyzing the aspect ratio of the target box, so as to more accurately reflect the size change of the target. In addition, the shape and scale penalty terms in the CIoU loss function can be adjusted experimentally to further improve the accuracy of target box regression. For the dynamic time warping algorithm (DTW), more trajectory features such as speed, acceleration, etc. can be introduced to more comprehensively analyze the motion trajectory of the target. For density-aware networks, deep learning techniques can be combined, such as using convolutional neural networks (CNNs) to extract the density features of the target, so as to more accurately detect abnormal aggregation behavior. As an alternative, other advanced target detection models (such as YOLOv7 or EfficientDet) can be used instead of the improved YOLOv5 model to further improve detection performance.
[0064] In some embodiments, the three-level logic judgment mechanism of S4 includes: S41. First-level judgment: based on the confidence threshold Conduct a preliminary screening.
[0065] S42. Second level judgment: trajectory continuity analysis is performed through the spatiotemporal consistency verification module.
[0066] S43. Third level judgment: Applying Bayesian network to calculate posterior probability ,when It is considered as a valid event.
[0067] It should be noted that the present invention adopts a three-level logic judgment mechanism in the event detection result verification stage to ensure the accuracy and reliability of the detection results. This mechanism verifies the confidence of the detection results step by step, from preliminary screening to final confirmation, to ensure that only high-confidence events are output. Specifically, the first-level judgment is based on the confidence threshold for preliminary screening, the second-level judgment analyzes the trajectory continuity through the spatiotemporal consistency verification module, and the third-level judgment uses the Bayesian network to calculate the posterior probability and finally determine the validity of the event. Among them, the Bayesian network is a model based on probabilistic reasoning that can combine prior knowledge and observation data to calculate the posterior probability of an event.
[0068] Specifically, each level of the three-level logic judgment mechanism has its own specific parameters and functions. In the first level judgment, the confidence threshold is set as: ;in, is the target size parameter, which indicates the relative size of the target in the image. This parameter can be dynamically adjusted according to the size of the target to meet the detection requirements in different scenarios. The second level judgment performs trajectory continuity analysis through the spatiotemporal consistency verification module, which uses the trajectory prediction model: ;in, To predict the location, is the current position, is the motion control quantity, is the adaptive weight coefficient, is the historical trajectory variance. In the third level judgment, the Bayesian network calculates the posterior probability: ;when When , the event is determined to be a valid event. Indicates that in the observed data Events under the conditions The posterior probability of occurrence, Indicates in the event Data observed under the conditions The probability of Indicates an event The prior probability of Represents observation data The total probability of .
[0069] Preferably, in the three-level logic judgment mechanism, the parameter settings and verification methods of each stage can be further optimized. For example, in the first-level judgment, the confidence threshold can be adjusted according to the actual application scenario, such as appropriately lowering the threshold in a high-noise environment to reduce the missed detection rate. In the second-level judgment, the parameters of the trajectory prediction model can be optimized through experiments, such as adjusting the adaptive weight coefficient by analyzing historical trajectory data. Calculation formula. In the third level judgment, the prior probability of the Bayesian network can be estimated by a large amount of historical data to improve the accuracy of the posterior probability calculation. As an alternative, a deep learning model (such as a recurrent neural network RNN or a long short-term memory network LSTM) can be introduced to replace the spatiotemporal consistency verification module to more accurately analyze the continuity and dynamic changes of the trajectory. In addition, a variety of probability models (such as Markov models or hidden Markov models) can be combined to replace the Bayesian network to further improve the accuracy and reliability of event verification.
[0070] In some embodiments, a night mode processing module is also included, specifically including: S61. Using a dark channel prior defogging algorithm to process low-light images.
[0071] S62. Apply improved Retinex-Net for illumination component decomposition.
[0072] S63. Amplify the effective motion signal through the motion feature enhancement network; the reflection component estimation formula of the improved Retinex-Net is: ; In the formula, is the adaptive mask, , is the local contrast parameter.
[0073] It should be noted that the present invention introduces a night mode processing module to address the problem of surveillance video event detection in low-light environments. The night mode processing module uses a series of image processing technologies, including a dark channel prior defogging algorithm, an improved Retinex-Net illumination component decomposition, and a motion feature enhancement network to improve the quality of low-light images and the accuracy of event detection. The dark channel prior defogging algorithm is a defogging technology based on image statistical characteristics, which can effectively remove fog from the image and enhance the clarity of the image. The improved Retinex-Net is an illumination component decomposition method based on deep learning, which can decompose the image into illumination components and reflection components, thereby enhancing the details of the image. The motion feature enhancement network improves the detection capability of moving targets by amplifying the motion signals in the image.
[0074] Specifically, each step of the night mode processing module has its own specific parameters and functions. First, the dark channel prior defogging algorithm estimates the concentration of fog by calculating the dark channel of the image and using the defogging formula: ;in, is the input image, is the image after dehazing, is the transmittance, is atmospheric light. Transmittance It can be estimated by dark channel a priori, atmospheric light It can be estimated by the brightness distribution of the image. Secondly, the improved Retinex-Net estimates the reflection component through the formula: ;in, is the reflection component, is the pixel value of the input image, is the illumination component, is the adaptive mask coefficient, is the local contrast parameter. Adaptive mask coefficient and the local contrast parameter It can be dynamically adjusted according to the brightness and contrast of the image. Finally, the motion feature enhancement network improves the detection ability of moving targets by amplifying the motion signal in the image, and its parameters can be optimized through the training data set.
[0075] Preferably, in the night mode processing module, the parameter settings and processing methods of each step can be further optimized. For example, in the dark channel prior defogging algorithm, the transmittance The estimation of can improve the dehazing effect by introducing deep learning models (such as convolutional neural networks). In the improved Retinex-Net, the adaptive mask coefficient It can be dynamically adjusted according to the brightness distribution of the image to better adapt to different lighting conditions.
[0076] Furthermore, the motion feature enhancement network can further improve the detection ability of moving targets by introducing an attention mechanism. As an alternative, other advanced defogging algorithms (such as a defogging network based on deep learning) can be used to replace the dark channel prior defogging algorithm to further improve the defogging effect. At the same time, a variety of illumination component decomposition methods (such as a multi-scale Retinex algorithm) can be combined to replace the improved Retinex-Net to further enhance the details and contrast of the image.
[0077] In some embodiments, a real-time optimization module is also included, which is specifically implemented as follows: S71. Establish a priority queue management detection task, and sort the queues according to ; In the formula, , is the event urgency parameter, The duration of the event.
[0078] S72. Use model sharding loading technology to dynamically allocate computing resources.
[0079] S73. Apply feature cache reuse mechanism to reduce repeated calculations.
[0080] It should be noted that the present invention introduces a real-time optimization module to ensure that the event detection system can still maintain efficient real-time performance under high load and complex scenarios. The real-time optimization module dynamically allocates computing resources and reduces repeated calculations by establishing a priority queue management detection task, adopting model fragmentation loading technology and feature cache reuse mechanism, thereby improving the response speed and processing efficiency of the system. The priority queue management detection task is to sort the tasks by comprehensively considering the urgency and duration of the event to ensure that high-priority tasks can be processed first. The model fragmentation loading technology decomposes the complex model into multiple small fragments and loads them on demand, thereby saving memory and computing resources. The feature cache reuse mechanism avoids repeated calculations and improves computing efficiency by caching the calculated features.
[0081] Specifically, each part of the real-time optimization module has its own specific parameters and functions. The priority queue management detection task is sorted according to the formula: ;in, is the task priority, and is the weight coefficient, Urgency is the event urgency parameter, and Duration is the event duration. The event urgency parameter can be set according to the event type and scenario. For example, for events involving security threats, the urgency parameter can be set to a higher value. The event duration is measured by the length of time from the detection of the event to the present. The model fragment loading technology breaks down complex models into multiple small fragments, each of which is dynamically loaded into the memory when needed, thereby saving memory and computing resources. The feature cache reuse mechanism improves computing efficiency by caching the calculated features to avoid repeated calculations. For example, for frequently occurring features, they can be cached, and the cached results can be used directly when they are needed again.
[0082] Preferably, in the real-time optimization module, the parameter settings and processing methods of each part can be further optimized. For example, in the priority queue management detection task, the weight coefficient and It can be adjusted according to the actual application scenario to better reflect the priority of different events. For example, in some scenarios, the duration of an event may be more important than the urgency. In the model sharding loading technology, more efficient model compression technologies such as quantization and pruning can be introduced to further reduce the model's memory usage and computational complexity. In the feature cache reuse mechanism, intelligent cache strategies can be introduced, such as dynamically adjusting the cache size and content according to the frequency and importance of feature usage.
[0083] Furthermore, as an alternative, other task scheduling algorithms (such as deep learning-based task scheduling networks) can be used to replace the priority queue management detection task to further improve the intelligence level of task scheduling. At the same time, a variety of model optimization technologies (such as model distillation) can be combined to replace the model sharding loading technology to further improve the operation efficiency of the model.
[0084] In some embodiments, the adaptive bilateral filter parameters in S12 are set to: ; In the formula, is the estimated noise level, is the image contrast parameter.
[0085] It should be noted that the present invention has made detailed provisions for the parameter setting of the adaptive bilateral filter to ensure that it can be dynamically adjusted under different environmental conditions, so as to better suppress noise and retain image details. The adaptive bilateral filter is a filtering technology that combines spatial distance and pixel similarity, which can maintain edge clarity while removing noise. Its parameter setting depends on the noise level estimate and image contrast parameters, which can reflect the noise characteristics and contrast information of the image, thereby realizing adaptive adjustment of the filter.
[0086] Specifically, the parameter setting formula of the adaptive bilateral filter is: ;in, is the standard deviation in the spatial domain, which is used to control the spatial weight of the filter; is the standard deviation of the color domain, used to control the color weight of the filter. Noise level estimate It can be obtained by analyzing the local variance of the image, which reflects the intensity of the noise in the image. Image contrast parameter It can be obtained by calculating the local contrast of the image, which reflects the difference in brightness and darkness of the image. Through the dynamic adjustment of these parameters, the adaptive bilateral filter can effectively remove noise and retain image details under different noise levels and contrast conditions.
[0087] Preferably, in the parameter setting of the adaptive bilateral filter, the calculation method of the noise level estimation value and the image contrast parameter can be further optimized. For example, the noise level estimation value can be obtained by analyzing the high-frequency components of the image, and specifically, wavelet transform or Fourier transform can be used to extract high-frequency information. The image contrast parameter can be obtained by calculating the histogram distribution of the image, so as to more accurately reflect the brightness difference of the image. In addition, more image quality assessment indicators, such as the structural similarity index (SSIM), can be introduced to further optimize the parameter setting of the filter. As an alternative, other advanced noise suppression algorithms, such as non-local means filtering (Non-Local Means) or a denoising network based on deep learning, can be used to replace the adaptive bilateral filter to further enhance the denoising effect of the image. At the same time, a variety of filtering techniques can be combined, such as the combination of bilateral filtering and Gaussian filtering, to achieve better noise suppression and image detail retention effects.
[0088] In some embodiments, the spatiotemporal consistency verification in S42 adopts a trajectory prediction model: ; In the formula, To predict the location, is the current position, is the motion control quantity, is the adaptive weight coefficient, is the historical trajectory variance.
[0089] It should be noted that the present invention adopts a specific trajectory prediction model in the spatiotemporal consistency verification module to ensure that the detected target trajectory is continuous and reasonable. Spatiotemporal consistency verification is an important link in event detection, which verifies the reliability of the detection result by analyzing the motion trajectory of the target. The trajectory prediction model uses the current position, motion control amount and historical trajectory information to predict the position of the target at the next moment, thereby determining whether the trajectory is continuous. In this way, abnormal trajectories caused by noise or false detection can be effectively filtered out, thereby improving the accuracy of event detection.
[0090] Specifically, the trajectory prediction model formula in the spatiotemporal consistency verification module is: ;in, To predict the location, is the current position, is the motion control quantity, is the adaptive weight coefficient, is the influence coefficient of the historical trajectory. Adaptive weight coefficient According to the variance of historical trajectory Dynamic adjustment, the specific formula is: ;in, is the maximum value of the historical trajectory variance, used to normalize the variance. Motion control amount It can be calculated from the target's speed and acceleration, reflecting the target's motion trend. Used to adjust the influence of historical trajectory on the predicted position, usually set to a small positive value. By setting these parameters, the trajectory prediction model can accurately predict the next moment's position based on the target's historical motion information, thereby verifying the continuity of the trajectory.
[0091] Preferably, in the spatiotemporal consistency verification module, the parameter setting and verification method of the trajectory prediction model can be further optimized. For example, the adaptive weight coefficient The calculation formula can be adjusted according to the actual application scenario. For example, in a scene where the target motion is relatively stable, the value can be appropriately increased. The value of can be reduced to rely more on the current position information. In the case of complex target motion, 's value to refer more to historical trajectory information.
[0092] Furthermore, the motion control The calculation of can introduce more motion features, such as acceleration, angular velocity, etc., to more comprehensively reflect the motion state of the target. As an alternative, other advanced trajectory prediction models, such as recurrent neural networks (RNN) or long short-term memory networks (LSTM) based on deep learning, can be used to replace the current trajectory prediction model to further improve the accuracy and robustness of trajectory prediction. At the same time, multiple verification methods can be combined, such as trajectory verification based on geometric constraints or trajectory verification based on statistical models, to further improve the effect of spatiotemporal consistency verification.
[0093] In some embodiments, a false alarm filtering module is also included, and the specific implementation method is: S101. Establish a multi-dimensional feature database to store historical false alarm samples.
[0094] S102. Apply contrastive learning network to calculate the similarity between the current detection result and the false positive sample: .
[0095] S103. When The false alarm filtering mechanism is activated and the detection result is not output to the monitoring system.
[0096] It should be noted that the present invention introduces a false alarm filtering module to reduce the false alarm rate in the event detection system and improve the reliability of the detection results. The false alarm filtering module stores historical false alarm samples by establishing a multidimensional feature database, and uses a contrastive learning network to calculate the similarity between the current detection result and the false alarm sample, thereby determining whether the current detection result is a false alarm. The multidimensional feature database is a collection of historical false alarm sample features stored to provide a benchmark for contrastive learning. The contrastive learning network is a model based on deep learning that can learn the similarity between samples to determine the degree of similarity between the current detection result and the known false alarm sample. When the similarity exceeds the set threshold, the false alarm filtering mechanism is activated to prevent false alarm results from being output to the monitoring system.
[0097] Specifically, the core of the false alarm filtering module lies in the construction of a multidimensional feature database and the implementation of a contrastive learning network. The multidimensional feature database stores the feature vectors of historical false alarm samples, which can be obtained by extracting multidimensional information such as images, trajectories, and behaviors of false alarm samples. The contrastive learning network determines whether it is a false alarm by calculating the similarity between the current detection result and the false alarm sample. The similarity calculation formula is: ;in, is the similarity, is the feature vector of the current detection result, is the feature vector of the false positive sample. When the false alarm filtering mechanism is activated, the current detection result is not output to the monitoring system. The parameter setting of the false alarm filtering module includes the determination of the similarity threshold, which can be adjusted according to the actual application scenario to balance the false alarm rate and the missed alarm rate.
[0098] Preferably, in the false alarm filtering module, the construction of the multidimensional feature database and the implementation of the contrastive learning network can be further optimized. For example, the multidimensional feature database can be regularly updated to include the latest false alarm sample features, thereby improving the accuracy and timeliness of false alarm filtering. The training of the contrastive learning network can introduce more false alarm samples and normal samples to enhance the generalization ability of the network.
[0099] Furthermore, the similarity threshold can be dynamically adjusted according to the needs of different scenarios. For example, in scenarios with high security requirements, the threshold can be appropriately increased to reduce false alarms. As an alternative, other advanced similarity calculation methods, such as those based on cosine similarity or Jaccard similarity, can be used to replace the current Euclidean distance similarity calculation formula to further improve the accuracy and efficiency of similarity calculation. At the same time, multiple false alarm filtering technologies can be combined, such as rule-based false alarm filtering or cluster analysis-based false alarm filtering, to further improve the effect of false alarm filtering.
[0100] The above-mentioned embodiments of the present invention have the following beneficial effects: First, through multimodal data preprocessing and multi-model fusion technology, the accuracy and robustness of event detection can be effectively improved. The improved Retinex algorithm and adaptive bilateral filter can eliminate shadow, reflection and noise interference, enhance image quality, and provide clearer input data for subsequent detection. At the same time, the adaptive weighted fusion module uses a dynamic attention mechanism to flexibly adjust the weights of each model, give full play to the advantages of the YOLOv5 improved model, the ResNet Transformer hybrid model and the LSTM spatiotemporal analysis model, and realize accurate identification of targets, abnormal trajectory analysis and abnormal aggregation behavior detection. The three-level logic judgment mechanism and the false alarm filtering module can further verify the confidence of the detection results, effectively reduce the false alarm rate, and ensure the reliability of the output results. In addition, the night mode processing module optimizes the processing for low-light environments, and through the dark channel prior defogging algorithm and the improved Retinex-Net, the detection performance at night or under harsh lighting conditions can be improved, further expanding the application scenarios of the present invention.
[0101] Secondly, the present invention performs well in real-time optimization. By establishing a priority queue to manage detection tasks and dynamically sorting events based on their urgency and duration, computing resources can be reasonably allocated and important events can be prioritized. Model sharding loading technology and feature cache reuse mechanism can reduce repeated calculations, improve detection efficiency, and ensure that the system can still maintain good real-time performance under high load. These optimization measures enable the present invention to not only meet the high accuracy requirements in complex scenarios, but also achieve fast and efficient event detection in actual monitoring systems, providing a more reliable and efficient technical solution for the field of intelligent monitoring.
[0102] Furthermore, the storage medium of the embodiment of the present application stores program instructions that can implement all the above methods, wherein the program instructions can be stored in the above storage medium in the form of a software product, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or terminal devices such as a computer, a server, a mobile phone, and a tablet.
[0103] The above descriptions are only some preferred embodiments of the present invention and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the above features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present invention.
Claims
1. A method for improving event detection accuracy based on multi-model fusion, characterized in that: The following steps are involved: S1. Perform multimodal data preprocessing on the surveillance video stream to generate a standardized feature matrix; S2. The outputs of the YOLOv5 improved model, the ResNet-Transformer hybrid model and the LSTM spatiotemporal analysis model are fused at the feature level through an adaptive weighted fusion module; S3. Multi-dimensional event detection is performed based on the fused feature matrix to generate a preliminary detection result set; S4. Establish a three-level logic judgment mechanism to verify the confidence of the test results; S5. Output the verified final event detection results to the monitoring system; The adaptive weighted fusion module in S2 adopts a dynamic attention mechanism, and the fusion formula is: In the formula, To fusion features, is the dynamic weight coefficient, is the sigmoid function, is the learnable parameter matrix, is the environmental feature vector, including light intensity, visibility, and rainfall parameters; The three-level logic judgment mechanism of S4 includes: S41. First-level judgment: based on confidence threshold Conduct preliminary screening; S42. Second level judgment: perform trajectory continuity analysis through the spatiotemporal consistency verification module; S43. Third level judgment: apply Bayesian network to calculate posterior probability ,when It is considered as a valid event.
2. The method according to claim 1, characterized in that The S1 includes: S11. Using the improved Retinex algorithm to perform image enhancement processing to eliminate shadow and reflection interference; S12. Applying an adaptive bilateral filter to suppress noise, and the filter parameters are dynamically adjusted according to the environmental visibility; S13. Constructing a spatiotemporal feature pyramid to extract multi-scale spatiotemporal features; wherein the illumination component estimation formula of the improved Retinex algorithm is: In the formula, is the illumination component, For color channels The pixel value of is the Gaussian kernel function, , It is the real-time visibility parameter.
3. The method according to claim 1, characterized in that: The S3 includes: S31. Target recognition is performed based on the improved YOLOv5 model, using the CIoU loss function and the adaptive anchor box mechanism; S32. Abnormal trajectory analysis is performed using the dynamic time warping algorithm (DTW); S33. Abnormal aggregation behavior is detected through a density-aware network; wherein the loss function of the improved YOLOv5 model is: In the formula, , is the target size parameter, .
4. The method according to claim 1, characterized in that It also includes a night mode processing module, specifically including: S61. Using a dark channel prior defogging algorithm to process low-light images; S62. Applying an improved Retinex-Net to decompose illumination components; S63. Amplifying effective motion signals through a motion feature enhancement network; wherein the reflection component estimation formula of the improved Retinex-Net is: In the formula, is the adaptive mask, , , is the local contrast parameter.
5. The method according to claim 1, characterized in that It also includes a real-time optimization module, which is specifically implemented as follows: S71. Establish a priority queue management detection task, and sort the queue based on ; In the formula, , is the event urgency parameter, is the duration of the event; S72. Use model sharding loading technology to dynamically allocate computing resources; S73. Apply feature cache reuse mechanism to reduce repeated calculations.
6. The method according to claim 2, characterized in that The adaptive bilateral filter parameters in S12 are set to: In the formula, is the estimated noise level, is the image contrast parameter.
7. The method according to claim 1, characterized in that The spatiotemporal consistency verification of S42 adopts the trajectory prediction model: In the formula, To predict the location, is the current position, is the motion control quantity, is the adaptive weight coefficient, is the historical trajectory variance.
8. The method according to claim 1, characterized in that It also includes a false alarm filtering module, which is specifically implemented as follows: S101. Establish a multidimensional feature database to store historical false alarm samples; S102. Apply a contrastive learning network to calculate the similarity between the current detection result and the false alarm sample: S103. When The false alarm filtering mechanism is activated and the detection result is not output to the monitoring system.
Citation Information
Patent Citations
Vehicle adaptive fusion detection method and system
CN114332655A
Military field annotation data correction and event detection method
CN117217222A