Video data preprocessing method fusing long and short distance event streams and target detection method

By fusing the instantaneous and long-distance information of the video event stream, the problem of long-term motion characteristics of the target in the prior art is solved, and the effect of improving the accuracy of target detection is achieved.

CN120198828APending Publication Date: 2025-06-24ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311774968.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the prior art, it is difficult to effectively learn the long-term motion characteristics of the target to be detected in object detection, resulting in large amounts of model parameters, long calculation time and low detection accuracy.

Method used

By calculating the event stream of the video, the long-distance event stream image is calculated using an adaptive quadratic time integration algorithm, and the instantaneous event stream information image is fused with the original image through the channel, increasing the information entropy of the image to enhance the target characteristics.

Benefits of technology

With smaller memory overhead and calculation amount, this method can effectively enhance the motion characteristics of the target to be detected and improve the accuracy of object detection, which is particularly suitable for motion object detection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004622166470000021
    Figure BDA0004622166470000021
  • Figure BDA0004622166470000023
    Figure BDA0004622166470000023
  • Figure BDA0004622166470000031
    Figure BDA0004622166470000031
Patent Text Reader

Abstract

The invention discloses a video data preprocessing method fusing long and short distance event streams, which comprises the following steps of: (1) sampling images according to a set interval for an infrared online video obtained in real time to obtain a current original image, and carrying out difference calculation on the original image and an original image obtained by sampling at the previous moment to obtain an instantaneous event stream image; (2) obtaining a quadratic integration result image representing long-distance variable weight event flow information through an adaptive quadratic time integration algorithm; and (3) fusing the obtained instantaneous event flow image, the secondary integration result image and the original image to obtain a fused image corresponding to the current moment. According to the method, the information entropy of the image can be greatly improved, and the effects of visual tasks such as a target detection algorithm and a training deep learning model are improved; the method is especially suitable for a moving target detection task, can enhance the features of a to-be-detected target, and serves as a training set to train a target detection model so as to improve the final detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent detection, and specifically relates to a video data preprocessing method and an object detection method that fuse long- and short-distance event streams. Background Art

[0002] Deep learning technology refers to the technology of training a deep learning model using training data (such as images, audio, text, etc.) so that the model can learn certain feature associations in the data and generate results. The quality of the training data is crucial for the deep learning model.

[0003] For some fields of object detection, the motion characteristics of the object to be detected are very important. In order to enable the model to obtain this feature, Bhatt et al. used eight consecutive frames of the video image as input (Bhatt R, Uzunbas M G, Hoang T, et al. Segmentation of low-level temporal plume patterns from IR video[C]. 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition Workshops. Long Beach, 2020: 847-854.), and trained a U-net-based network so that the network can learn the motion characteristics of the object to be detected. However, the model trained by this method has a large number of parameters, long training and inference calculation time, and cannot learn the long-term motion characteristics of the object to be detected. Summary of the Invention

[0004] The preprocessing method of the present invention calculates the event stream of the video, and calculates the long-distance event stream image through the adaptive quadratic time integration algorithm. Combining the instantaneous event stream information image, the two images are fused into the original image through the channel fusion method to improve the information entropy of the original image. This method is particularly suitable for the motion object detection task, can enhance the features of the object to be detected, and is used as a training set to train the object detection model to improve the accuracy of the final detection.

[0005] A video data preprocessing method that fuses long- and short-distance event streams includes the following steps:

[0006] (1) For the real-time acquired infrared online video, sample images at a set interval to obtain the current original image, and perform differential calculation on the original image and the original image sampled at the previous moment to obtain the instantaneous event stream image;

[0007] (2) Obtain the quadratic integral result image representing the variable-weight event stream information of the long distance through the adaptive quadratic time integration algorithm;

[0008] (3) Fuse the obtained instantaneous event stream image, the result image of the second integral, and the original image to obtain the fused image corresponding to the current moment.

[0009] Preferably, the differential calculation result is the absolute value of the difference between the original image at the current moment and the original image corresponding to the previous moment.

[0010] The instantaneous event stream image is calculated using the following formula:

[0011] diff F (i, j) = |I F (i, j) - I F-gap (i, j)|

[0012] where diff F (i, j) is the result image of the differential algorithm, that is, the instantaneous event stream image corresponding to the current moment; I is the sampled image as the input of the algorithm: I F (i, j) is the original image collected at the current moment, that is, the F-th frame of the original image collected since the start of image collection; I F-gap (i, j) is the original image collected at the previous moment, that is, the (F - gap)-th frame of the image collected, where gap is the sampling interval.

[0013] In step (2), the formula of the second time integration algorithm is as follows:

[0014]

[0015] where represents the result image of the first time integration, f represents the total number of frames currently collected; ω1 is the degradation factor of the first time integration, b1 is the weak bias of the first time integration, ω2 is the degradation factor of the second time integration, and b2 is the weak bias of the second time integration.

[0016] Specifically, the long-distance variable-weight event stream information is obtained through the adaptive second time integration algorithm, while avoiding excessive memory consumption. The algorithm first performs a first time integration calculation, and the calculation formula is as follows:

[0017]

[0018] where FI f (i, j) represents the result image of the first time integration, f represents the total number of samples currently taken, ω1 is the degradation factor of the first time integration, which determines the attenuation speed of the weight of each frame, and b1 is the weak bias of the first time integration, which can avoid the over-accumulation of the response of the time integration.

[0019] The second - order time integration is to perform another time integration on the basis of the first - order time integration. The second - order time integration can further accumulate the event - stream information, enabling the result image to obtain event - stream information over a longer distance. At the same time, the second - order time integration can also avoid the interference caused by camera jitter or the instantaneous movement of interfering objects in the video to the result. The expression of the second - order time integration is as follows:

[0020]

[0021] Among them, SI f (i, j) represents the result image of the second - order time integration, ω2 is the degradation factor of the second - order time integration, and b2 is the weak bias of the second - order time integration.

[0022] In the entire algorithm of the second - order time integration, the three values of the degradation factor, the weak bias, and the distance between sampling frames have a great impact on the final result. Since videos in different scenarios have different factors such as contrast and the motion intensity of the target to be detected, and the overall change intensity of the frames at different moments in the same video is also different, it may cause the result of the second - order time integration to accumulate to an overly large value or an overly small value. Therefore, the second - order time integration algorithm can be made applicable to different scenarios by changing the degradation factor, the weak bias, and the distance between sampling frames. Thus, an adaptive second - order time integration algorithm is designed.

[0023] Preferably, in step (2), the calculation of the integral change trend is carried out simultaneously, and based on the obtained integral change trend, it is determined whether to adjust the degradation factor and the weak bias in the pre - processing of the next original image; the degradation factor and the weak bias can be the degradation factor and the weak bias of the first - order time integration, or the degradation factor and the weak bias of the second - order time integration.

[0024] Preferably, the calculation of the integral change trend adopts the following formula:

[0025]

[0026] Among them, μ represents the integral change trend, w represents the width of the image, h represents the height of the image, and t represents time.

[0027] Preferably, after obtaining the integral change trend, if the integral change trend of consecutive specific frames is greater than the set threshold, then the first time integral degradation factor and the first time integral weak deviation are increased accordingly; otherwise, the first time integral degradation factor and the first time integral weak deviation are reduced. Alternatively, as another option, after obtaining the integral change trend, if the integral change trend of consecutive specific frames is greater than the set threshold, then the degradation factor of the second time integral and the weak deviation of the second time integral are increased accordingly; otherwise, the second time integral degradation factor and the weak deviation of the second time integral are reduced.

[0028] Furthermore, when the movement of objects in a scene is too violent, or the camera shakes greatly, causing the pixels in the picture to change violently, the time integral response will also increase. The result of time integral has continuity. If the time integral result of the previous frame is too large, it will affect the time integral result of the next frame, and eventually the time integral result will gradually increase, making the detection effect worse. Therefore, a formula for calculating the integral change trend is designed to evaluate whether the integral response is too large at this time by the change of the pixel value of the overall integral image over time. The expression is as follows:

[0029]

[0030] In the formula, μ represents the trend of integral change, w represents the width of the image, h represents the height of the image, and t represents time. For example, if μ is greater than a threshold for 30 consecutive frames (or other set frames, which can be adjusted according to actual application scenarios), the degradation factor ω2 and the weak deviation b2 are increased accordingly, the weight of the previous frame difference result's contribution to the second time integral is reduced, and the weak deviation is increased to reduce the cumulative effect of the integral result. On the contrary, if μ is less than a value for 30 consecutive frames, the degradation factor ω2 and the weak deviation b2 are reduced. The adjustment of the degradation factor ω2 and the weak deviation b2 is limited within a certain range, that is, the value range of the degradation factor ω2 and the weak deviation b2 is set. When the current degradation factor ω2 and the weak deviation b2 reach the set upper or lower limit, they will no longer increase or decrease, and are always within the set range, further increasing the stability of the processing.

[0031] In step (3), image fusion is performed by channel merging. The instantaneous motion information of the current image can be expressed by the event stream image of the video, and the long-term historical motion information can be obtained by the result of the adaptive quadratic time integration algorithm. The results of these two algorithms are combined with the original image by channel merging to perform image fusion. Preferably, a three-channel mode can be used for fusion, for example, the original image is used as the blue channel, the event stream image is used as the red channel, and the quadratic integration result image is used as the green channel.

[0032] A target detection method based on motion features includes training a model and performing real-time leakage detection with the fused image obtained by processing using the preprocessing method described in any of the above technical solutions as the input.

[0033] The preprocessing method of the present invention fuses time-domain information into the image and is applicable to models with image data as the input, such as YOLO series models, RCNN series models, DETR models, etc. It is suitable for application fields that rely on the time-domain characteristics of data, such as target detection.

[0034] Preferably, the target detection is gas detection based on a video stream, and the detection target is whether leakage occurs.

[0035] Preferably, at the initial stage of detection, the sampling interval in the preprocessing process is set to a relatively large value, and then the specific value of the sampling interval in the preprocessing process is adjusted according to the detection result of the target to be detected.

[0036] Taking the preprocessing algorithm parameters of the gas detection model as an example, the initial value of gap is 21 (for other systems, it can be other set values of gap max ), and it takes values between 1 and 21 in subsequent detections (for other systems, it can be the corresponding gap min ~gap max ). When gap is greater than or equal to 3 (gap min +2), when the model continuously detects the target to be detected for 3 frames (or any n frames less than gap max ), in order to increase the sampling frequency, the gap value is decreased by 2 (or △gap determined according to the actual application scenario) as the input parameter of the preprocessing algorithm for the next frame; conversely, when gap is less than or equal to 19 (or when it is less than or equal to gap max -△gap), when the model does not detect the target to be detected for 2 consecutive frames (△gap), the gap value is increased by 2 (△gap) as the input parameter of the preprocessing algorithm for the next frame, and it will not be decreased within 5 minutes (or other set time t), so as to enhance the preprocessing effect and ensure the stability of subsequent detections.

[0037] Through the preprocessing algorithm, the present invention can summarize historical long-term and instantaneous motion information into two pictures with relatively small memory overhead and computational complexity, and then fuse these two pictures into the single-channel original image through channel fusion, enhancing the motion features of the image. Using the preprocessed image as the input of the model can obtain motion information while effectively solving the technical problems existing in the prior art.

[0038] A video data preprocessing method that fuses long- and short-distance event streams proposed by the present invention can significantly increase the information entropy of images and improve the effects of visual tasks such as object detection algorithms and training deep learning models; this method is particularly suitable for moving object detection tasks, can enhance the features of the objects to be detected, and be used as a training set to train an object detection model to improve the accuracy of the final detection. Description of the Drawings

[0039] Figure 1 It is a frame in the infrared video containing gas leakage in the embodiment;

[0040] Figure 2 It is the event stream result diagram of the gas leakage infrared video in the embodiment;

[0041] Figure 3 It is the first integral result diagram of the gas leakage infrared video in the embodiment;

[0042] Figure 4 It is the second integral result diagram of the gas leakage in the embodiment;

[0043] Figure 5 It is the fused image obtained by using the original image as the blue channel, the event stream image as the red channel, and the second integral result image as the green channel in the embodiment;

[0044] Figures 6 - 7 It is the gas leakage detection image result when verifying the existence of personnel by using the method of the present invention ( Figure 6 It is the original image at a certain moment, Figure 7 It is the processed fused image);

[0045] Figure 8 and Figure 9 It is the gas leakage detection image result when verifying the existence of moving objects by using the method of the present invention ( Figure 8 It is the original image at a certain moment, Figure 9 It is the processed fused image);

[0046] Figure 10 It is the fused images preprocessed by using the method of the present invention in 8 different scenarios;

[0047] Figure 11 It is the confusion matrix in the appendix. Detailed Embodiments

[0048] Taking the gas leakage detection task as an example, this article elaborates on the principle of the video data preprocessing method that fuses long- and short-distance event streams. The input of this object detection task is an infrared video that may contain leaked gas, and the detection of leaked gas is achieved by training a convolutional neural network. Figure 1It is a frame in an infrared video containing a gas leak. The characteristics of the gas leak in the infrared are not obvious, making it difficult to confirm the location of the gas leak. As data for training a convolutional neural network, it is likely to cause the network to overfit and the network detection effect is poor.

[0049] Implementation method:

[0050] 1. Obtain instantaneous event stream images

[0051] Sample images at a certain interval in the video to obtain the current original image. Perform differential calculation on the current sampled original image and the previous sampled original image to obtain the instantaneous event stream image diff F (i, j), and its calculation formula is as follows

[0052] diff F (i, j) = |I F (i, j) - I F-gap (i, j) |

[0053] Among them, diff F (i, j) is the instantaneous event stream image obtained through the differential algorithm. I is the original image obtained by the current sampling, which is used as the input of the algorithm. (i, j) represents the pixel position. The superscript F represents the current frame number, and gap represents the number of frame distances between sampled frames. Figure 2 It is the event stream result image of the gas leak infrared video, that is, the instantaneous event stream image of the current image.

[0054] 2. Obtain the second integral result image

[0055] Obtain long-distance variable-weight event stream information through the adaptive second-order time integration algorithm, while avoiding excessive memory consumption. The algorithm first performs a first-order time integration calculation, and the calculation formula is as follows

[0056]

[0057] Among them, FI f (i, j) represents the result image of the first-order time integration. f represents the number of all frames from the start of the video to the current time. ω1 is the degradation factor of the first-order time integration, which determines the attenuation speed of the weight of each frame. b1 is the weak deviation of the first-order time integration, which can avoid the over-accumulation of the response of the time integration. The first-order integration result of the gas leak infrared video is as Figure 3 shown.

[0058] The second time integration is to perform another time integration on the basis of the first time integration. The second time integration can further accumulate the event flow information, so that the event flow information of a longer distance can be obtained in the result image. At the same time, the second time integration can also avoid the interference caused by the instantaneous motion of the camera shake or the interfering objects in the video. The expression of the second time integration is as follows:

[0059]

[0060] Among them, SI f (i, j) represents the result image of the second time integration, ω2 is the degradation factor of the second time integration, and b2 is the weak deviation of the second time integration.

[0061] In the entire quadratic time integration algorithm, the three values ​​of degradation factor, weak deviation and distance between sampling frames have a great influence on the final result. Videos in different scenes are different due to factors such as contrast and motion intensity of the target to be detected, and the overall change intensity of the picture at different times of the same video is also different, which may cause the result of the quadratic time integration to accumulate to a value that is too large or too small. Therefore, it is necessary to change the degradation factor, weak deviation and distance between sampling frames to make the quadratic time integration algorithm suitable for different scenes. Therefore, the present invention designs an adaptive quadratic time integration algorithm.

[0062] When an object in a scene moves too violently, or the camera shakes too much, causing the pixels in the picture to change violently, the time integral response will also increase. The result of the time integral is continuous. If the time integral result of the previous frame is too large, it will affect the time integral result of the next frame, and eventually the time integral result will gradually increase, making the detection effect worse. Therefore, a formula for calculating the integral change trend is designed to evaluate whether the integral response is too large at this time by comparing the change of the pixel value of the integral image over time. The expression is as follows

[0063]

[0064] In the formula, μ represents the trend of integral change, w represents the width of the image, h represents the height of the image, and t represents time. If μ is greater than a threshold for 30 consecutive frames, the degradation factor ω1 and weak deviation b1 are increased accordingly, the weight of the previous frame difference result's contribution to the first time integral is reduced, and the weak deviation is increased to reduce the cumulative effect of the integral result. Conversely, if μ is less than a value for 30 consecutive frames, the degradation factor ω1 and weak deviation b1 are reduced. The adjustment of the degradation factor ω1 and weak deviation b1 is within a certain range.

[0065] Using the method of the present invention, the secondary integral result of gas leakage is as follows: Figure 4 shown.

[0066] In addition, a larger sampling frame interval gap means that the time interval between the two frames participating in the inter-frame difference algorithm is farther, so the displacement distance of the moving object is more, and it is easier to generate a larger response in the frame difference result. But at the same time, it also means a reduction in the sampling points of the algorithm, making the response speed of the algorithm slower. Therefore, a suitable sampling frame interval is particularly important. Therefore, a relatively large initial sampling interval value gap = 5 is set to ensure the effect of the quadratic time integration algorithm. If a suspected target to be detected is found in subsequent detections (various existing detection algorithms based on image processing (such as image filtering, operators, and psychological transformations) or a gas detection model based on deep learning can be used to obtain the target to be detected. The data obtained through the preprocessing method is better, so that the detection algorithm can be better designed and the detection accuracy is higher), then the sampling interval is reduced until it affects the detection effect (unstable detection). Specifically, taking the preprocessing algorithm parameters of the gas detection model as an example, the initial value of gap is 21, and it takes values between 1 and 21 in subsequent detections. When gap is not 1, when the model continuously detects the target to be detected for 3 frames, in order to increase the sampling frequency, the gap value is reduced by 2 as the input parameter of the next frame of the preprocessing algorithm; conversely, when gap is not 20, when the model does not detect the target to be detected for 2 consecutive frames, the gap value is increased by 2 as the input parameter of the next frame of the preprocessing algorithm, and it will not be reduced within 5 minutes, so as to enhance the preprocessing effect and ensure the stability of subsequent detections.

[0067] 3. Image Fusion

[0068] Through the event stream image of the video, the instantaneous motion information of the current image can be expressed, and through the result of the adaptive quadratic time integration algorithm, the long-term historical motion information can be obtained. The results of these two algorithms and the original image are merged through channels for image fusion. Among them, the original image is used as the blue channel, the event stream image is used as the red channel, and the quadratic integration result image is used as the green channel. The result of the image fusion is as Figure 5 shown.

[0069] To verify the effectiveness of the present invention, algorithm verification is carried out:

[0070] The algorithm of the present invention has a fast calculation speed. When testing the algorithm on the intel i5 CPU platform, the algorithm can process a single-channel image with a size of 640*512 within 10 ms (the infrared image is taken by a Hepro-gas300 cooled infrared camera), fully meeting the real-time requirements. This preprocessing can greatly improve the features of moving objects in the image, Figures 6 - 9 and there are more gas detection task examples (the red and green colored areas are gas leakage areas).

[0071] Figure 6 and Figure 7is a set of data, Figure 8 and Figure 9 is a set of data. The purpose of the two sets of images is to compare the infrared images ( Figure 6 and Figure 8 ) and the corresponding preprocessed images ( Figure 7 and Figure 9 ). By observing the images of the gas leakage area with the naked eye, it can be found that the gas gray level in the infrared image is close to the background, and there is no specific shape. The contour edge is blurred. Before preprocessing, it is difficult to determine whether there is gas leakage and the gas leakage location by human judgment. Therefore, the gas characteristics in the infrared image are weak, and it is difficult to train an effective deep learning model for gas detection with this data. For the preprocessed image, there are obvious red-green blocks in the gas leakage area, and the red and green blocks have a unique relative shape different from other moving objects, which is caused by the special floating movement of the gas. The gas characteristics in the preprocessed image are rich (rectangular frame), and an effective deep learning model can be trained.

[0072] The preprocessed data is used as the training set to train a gas leakage convolutional neural network (YOLOv8 convolutional neural network). Since the preprocessed data enhances the information entropy and the characteristics of the target to be detected, the detection effect of the gas leakage model is significantly enhanced compared with the network model trained with the data without preprocessing. The following are the detection results ( Figure 10 ) of the convolutional neural network model trained with the preprocessed data and the detection result table 1 (testing the model on the Tesla t4 GPU platform).

[0073] Table 1

[0074]

[0075]

[0076] Figure 10 Among them, eight images correspond to different scenarios for gas leakage simulation detection. The eight different scenarios are as follows;

[0077] video1: Micro gas leakage at a long distance in a road scene, with interference such as shaking leaves and passing electric vehicles in the scene.

[0078] Video2: Micro gas leakage at a long distance in a road scene, with interference such as shaking leaves and passing cars in the scene.

[0079] Video3: Medium gas leakage at a long distance in a park scene, with interference such as passing vehicles in the scene.

[0080] Video4: Medium gas leakage at a long distance in a road scene, with interference such as passing pedestrians in the scene.

[0081] Video5: Medium gas leakage at a long distance in a road scene, with interference such as leaf shaking in the scene.

[0082] Video6: Medium gas leakage at a long distance in a park scene, with interference such as worker operation in the scene.

[0083] Video7: Medium gas leakage at a long distance in a park scene, with interference such as worker operation in the scene.

[0084] Video8: Medium gas leakage at a long distance in a park scene, with interference such as a cyclist passing by in the scene.

[0085] In the table, "trace leakage" and "medium leakage" indicate the size of the pixel area occupied by the leaked gas in the image. "(New)" indicates that this scene does not appear in the training dataset. Other scenes appear in the training set (there are overlapping scenes in the training set and the test set, but they are from different gas leakage videos, only with the same background). "FPS" is the number of frames the model infers per second. The meanings of Precision, Recall, and F1-score are shown in the appendix (Ahmed Iqbal, Muhammad Usman, Zohair Ahmed. Tuberculosis chest X-ray detection using CNN-based hybrid segmentation and classification approach[J]. Biomedical Signal Processing and Control, 2023(84):104667.).

[0086] As can be seen from Table 1, when modeling with the images processed by the preprocessing method of the present invention, the target detection rate is high and the robustness is better.

[0087] Appendix:

[0088] Evaluation Metrics for Gas Detection Algorithms

[0089] When conducting hypothesis testing, there are two types of errors, namely the first type of error and the second type of error. For target detection, the first type of error is incorrectly predicting a detection target in an area without a target, which can also be called False Positive (FP). The second type of error is incorrectly not predicting a detection target in an area with a target, which can also be called False Negative (FN). Additionally, in target detection, successfully detecting a target is called True Positive (TP), and successfully not detecting a target in an area without a target is called True Negative (TN). Integrate the above concepts as shown in Figure 11In this case, this figure is also known as the confusion matrix.

[0090] Based on the concept of the confusion matrix, several metrics for evaluating the performance of object detection algorithms are derived, namely precision, recall, and F1-score.

[0091] Precision

[0092] Precision is the probability of correct detection among all detected objects, and its expression is as follows

[0093]

[0094] Precision measures the credibility of the algorithm's detection results. The corresponding one is the false positive rate. The higher the precision, the lower the false positive rate of the algorithm. However, using precision to measure the algorithm's performance also has certain limitations, that is, precision cannot evaluate the object detection ability of the algorithm. If the number of objects successfully detected by the algorithm is very small, that is, TP is small, and if FP is also small, the algorithm can still obtain a high precision.

[0095] Recall

[0096] Recall is the probability of correct recognition among all positive samples, and its expression is as follows

[0097]

[0098] Recall measures the object detection ability of the algorithm. The corresponding one is the miss rate. The higher the recall, the fewer the undetected positive samples FN, and the lower the miss rate. Using recall to measure the algorithm's performance also has certain limitations, that is, it cannot evaluate the credibility of the algorithm's detection results. When the number of the algorithm's detection results TP + FP is large and FP accounts for a large proportion, it means that the algorithm has a large false positive rate, but FN is small, and the algorithm can still obtain a high recall.

[0099] F1-score

[0100] Neither precision nor recall can comprehensively and accurately measure the performance of an algorithm, nor can they be used as a basis for parameter tuning. When increasing the confidence threshold of the detection algorithm, FP is more likely to be excluded by the algorithm. Therefore, FP decreases more than TP, resulting in an increase in precision. However, since TP also decreases correspondingly, the corresponding FN increases, leading to a decrease in recall. Conversely, if the confidence threshold of the detection algorithm is decreased, the algorithm is more likely to produce more FPs, resulting in a decrease in precision, while the increase in TP also leads to an increase in recall. Therefore, when adjusting the confidence threshold of the detection algorithm, the changes in precision and recall are often relative. To obtain an indicator that can comprehensively measure the performance of the algorithm, F1score can be used as an indicator to measure the performance of the algorithm. The formula is as follows

[0101]

[0102] As can be seen from the formula, F1score is essentially the harmonic mean of precision and recall, which is obtained by combining the scores of precision and recall. It can also represent the performance of the detection algorithm well for algorithms with unbalanced precision and recall scores.

Claims

1. A video data preprocessing method for fusing long- and short-distance event streams, characterized in that, It includes the following steps: (1) For the real-time acquired infrared online video, sample images at a set interval to obtain the current original image, and perform differential calculation on this original image and the original image sampled at the previous moment to obtain the instantaneous event stream image; (2) Obtain the second integral result image representing the variable-weight event stream information of long distance through the adaptive second-order time integration algorithm; (3) Fuse the obtained instantaneous event stream image, the second integral result image and the original image to obtain the fused image corresponding to the current moment.

2. The video data preprocessing method for fusing long- and short-distance event streams according to claim 1, wherein The result of the differential calculation is the absolute value of the difference between the original image at the current moment and the corresponding original image at the previous moment.

3. The video data preprocessing method for fusing long- and short-distance event streams according to claim 1, characterized in that The formula of the second-order time integration algorithm is as follows: SI f (i,j) = ∫ f 0 ((FI f (i,j) - b2)ω2 F ) dF Among them, FI f (i,j) = -∫ f 0 ((diff F (i,j) - b1)ω1 F )dF represents the result image of the first temporal integration, f represents the total number of frames currently acquired; ω1 is the degradation factor of the first temporal integration, b1 is the weak bias of the first temporal integration, ω2 is the degradation factor of the second temporal integration, and b2 is the weak bias of the second temporal integration.

4. The video data preprocessing method for fusing long- and short-distance event streams according to claim 3, characterized in that In step (2), the calculation of the integral change trend is carried out at the same time, and according to the obtained integral change trend, it is determined whether to adjust the degradation factor and the weak deviation in the preprocessing of the next original image.

5. The video data preprocessing method for fusing long- and short-distance event streams according to claim 4, wherein, The calculation of the integral change trend adopts the following formula: Wherein, μ represents the integral change trend, w represents the width of the image, h represents the height of the image, and t represents time.

6. The video data preprocessing method for fusing long- and short-distance event streams according to claim 4, wherein After obtaining the integral change trend, if the integral change trend of consecutive specific frames is greater than the set threshold, the degradation factor and the weak deviation are correspondingly increased; otherwise, the degradation factor and the weak deviation are decreased; the degradation factor and the weak deviation are the degradation factor and the weak deviation of the first-order time integration or the degradation factor and the weak deviation of the second-order time integration.

7. The video data preprocessing method for fusing long- and short-distance event streams according to claim 1, wherein Image fusion is carried out by means of channel merging.

8. A target detection method based on motion features, characterized in that, Using the fused image processed by the preprocessing method described in any one of claims 1 to 7 as the input for the training of the model and the real-time detection of air leakage.

9. The method for target detection based on motion features according to claim 8, wherein The model is selected from the YOLO series models, the RCNN series models, and the DETR model.

10. The object detection method based on motion features according to claim 8, characterized in that, In the initial stage of preprocessing, the sampling interval is set to a relatively large value, and when air leakage is detected by this detection method, the sampling interval is reduced.