Multi-stage filtering road thrown object detection method based on dynamic difference analysis

By combining dynamic difference analysis and multi-level filtering methods with the YOLO model and multi-dimensional feature analysis, the problems of high false detection rate and poor adaptability of debris detection on highways have been solved, achieving efficient and stable debris detection.

CN120932185APending Publication Date: 2025-11-11CCCC HUAKONG (TIANJIN) CONSTR GRP CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202511032846.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies for detecting debris on highways are susceptible to changes in lighting and shadows, resulting in a high false alarm rate. Furthermore, deep neural network methods require a large number of samples for training and have weak generalization ability, making them difficult to effectively detect debris.

Method used

A multi-level filtering method based on dynamic difference analysis is adopted, including target localization, interference elimination, multi-dimensional filtering and spatiotemporal consistency verification. Interference is eliminated by using the YOLO model, and stable targets are selected by combining color histogram similarity, structural similarity and brightness contrast.

Benefits of technology

It significantly reduces the false detection rate caused by factors such as shadows and water stains, improves the adaptability and stability of detection, is suitable for various lighting and weather conditions, and has good interpretability and ease of deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932185A_ABST
    Figure CN120932185A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-stage filtering road spilled object detection method based on dynamic difference analysis, which is suitable for automatic identification of unstructured foreign matters in video monitoring. The method comprises the following steps: firstly, extracting a reference image road mask, eliminating vehicle and pedestrian interference by using YOLOv8 detection, and extracting a motion candidate area through a frame difference method and background modeling; and then context expansion and super-resolution reconstruction are carried out on the candidate region, the candidate region is converted into an HSV space, multi-dimensional features such as color similarity, structural similarity and shadow determination are synthesized for screening, false detection is further removed in combination with inter-frame time sequence consistency, and finally a stable detection result is output. The method provided by the invention has the advantages of strong anti-interference capability, high adaptability, high detection precision and the like, and is suitable for the intelligent recognition task of the expressway thrown objects in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical fields of intelligent traffic monitoring, image processing and target detection, and in particular to a multi-level filtering method for detecting road debris based on dynamic difference analysis. Background Technology

[0002] With the increasing frequency of road transport, incidents of cargo falling and vehicle parts detaching from highways are on the rise, posing a serious threat to traffic safety. Current mainstream detection methods include:

[0003] 1. Traditional frame difference method is easily affected by changes in lighting and shadows, and is sensitive to complex backgrounds, shadows and changes in lighting, resulting in a high false alarm rate;

[0004] 2. Although end-to-end neural network methods have high accuracy, they rely on deep neural network training, require a large number of sample data, have weak generalization ability, and lack logical interpretability and stability verification mechanisms.

[0005] 3. Most models misclassify static areas such as medians, curbs, reflective surfaces, and water stains.

[0006] Therefore, there is an urgent need for a lightweight method that integrates dynamic image analysis, target detection, and image feature comparison, with good adaptability and real-time performance, to overcome training dependence and improve the accuracy of projectile detection. Summary of the Invention

[0007] This invention aims to address the shortcomings of existing technologies by providing a multi-level filtering method for detecting road debris based on dynamic difference analysis. Addressing the challenge of direct detection due to the diverse morphologies and indistinct features of road debris, this invention achieves efficient and robust debris detection through a combination of dynamic target localization, interference elimination, multi-dimensional filtering, and spatiotemporal consistency verification.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] A multi-level filtering method for detecting road debris based on dynamic difference analysis includes the following steps:

[0010] Step S1: Obtain the reference image as I ref And generate the reference mask M smoothed ;

[0011] Step S2: Extract video frame I from the video to be detected. k Calculate the video frame mask M frame With reference mask M smoothed Similarity S mask To determine whether to conduct spill detection;

[0012] Step S3: Train the YOLO object detection model and use the trained YOLOv8 model to detect the reference image I. ref and video frame I k Target detection is performed to obtain vehicles, pedestrians, and guardrails in the image. Vehicles, pedestrians, guardrails, and off-road areas are then filled with a solid color to obtain a reference image I' after eliminating dynamic interference targets. ref and video frame I' k ;

[0013] Step S4: Apply Gaussian mixture background modeling algorithm to I' ref and I' k Frame difference processing is performed to obtain the initial dynamic target candidate region M. diff ;

[0014] Step S5, for M diff For each candidate target in the dataset, the candidate bounding box is dynamically expanded to obtain the expanded candidate region R. extend ;

[0015] Step S6, for R extend Super-resolution enhancement is performed to obtain a high-resolution image R. high-res Then R high-res Candidate region H is obtained by converting from RGB color space to HSV color space. cand ;

[0016] Step S7: For candidate region H cand Perform color histogram similarity analysis, structural similarity analysis, shadow determination, and brightness comparison to comprehensively filter out non-sprayed targets;

[0017] Step S8: Calculate the frequency of candidate regions in consecutive frames, combine with the IOU threshold to filter stable targets, and output the coordinates of the projectiles and the visualization results.

[0018] In step S1, low-saturation, medium-brightness road regions are extracted using a preset threshold as a reference mask M. smoothed , where the color thresholds are H∈[0,180], S∈[0,50], and V∈[50,200].

[0019] Furthermore, the specific process of step S1 is as follows:

[0020] First, obtain a clear, unobstructed reference image. Then, it is converted to the HSV color space. Based on the common low saturation and moderate brightness color characteristics of road areas, the color threshold ranges H∈[0,180], S∈[0,50], V∈[50,200] are set to extract the road area and generate the initial mask M. road .

[0021] To further improve the robustness of the mask, morphological erosion is used to remove isolated noise points, and then dilation is used to repair edge gaps, resulting in a smooth and continuous road area mask M. smoothed This provides a stable spatial constraint for subsequent differential detection and significantly reduces the error caused by background variations.

[0022] In step S2, the road mask M of the video frame is first calculated. frame With reference mask M smoothed Similarity S mask When the similarity S mask Detection is performed when the value is higher than 0.5.

[0023] Furthermore, the process of step S2 is as follows:

[0024] Extract video frames from the video at preset time intervals. Each frame of the image is processed using the same color thresholding and morphological processing methods as in step S1 to obtain the road mask M. frame Compare it with the reference mask M smoothed The overlap ratio is compared using the following formula: if the mask similarity S... mask If the value is below a threshold (e.g., 0.5), the current frame is considered to have a significant environmental change, and the analysis is skipped. Otherwise, the subsequent steps of debris detection are performed.

[0025]

[0026] In step S3, the YOLO object detection algorithm is used to identify the reference image I. ref and video frame I k The system detects vehicles, pedestrians, and guardrails, and eliminates their interference with debris detection by masking.

[0027] Furthermore, the process of step S3 is as follows:

[0028] Data from highway scenarios was collected to construct a dataset containing various categories such as vehicles, pedestrians, guardrails, and safety cones. A YOLO object detection model was trained on this dataset. The trained YOLOv8 model was then used to detect common interfering targets in the images, such as pedestrians, vehicles, and guardrails. The bounding boxes of these targets were expanded and filled with solid colors to obtain the processed reference image I'. ref and video frame I' k This effectively eliminates the problem of interference from other targets in the frame difference method, and improves the purity and stability of subsequent detection.

[0029] In step S4, the frame difference processing steps are as follows: Reference image I ref and video frame I kThe image is converted to grayscale and blurred using Gaussian. The difference between the two frames is calculated and then binarized to obtain the initial dynamic target candidate region M. diff .

[0030] Furthermore, the process of step S4 is as follows:

[0031] Will I' ref and I' k After converting to grayscale, Gaussian blurring is performed to reduce noise interference. The frame difference between the two images is calculated and thresholded to obtain a binary difference map M. diff At the same time, for video frame I' k Using the MOG2 background modeling method, the dynamic foreground region M is extracted. fg The two are bitwise ANDed to obtain the candidate dynamic mask M. candidate Median filtering and smoothing closing operations are applied to the mask to enhance connectivity, and connected components with sizes within a reasonable range (e.g., width and height between 20 and 100 pixels) are extracted as a preliminary set of candidate bounding boxes. This step significantly improves the sensitivity to moving projectiles by fusing frame difference and background modeling as two dynamic extraction methods, and uses size constraints to initially filter false alarms.

[0032] In step S5, the expansion ratio is dynamically calculated based on the target size, and the detection box is expanded spatially using the quadratic boundary expansion method. The expanded area is then masked.

[0033] Furthermore, the process of step S5 is as follows:

[0034] The coordinates of each initial candidate box are proportionally expanded in context to obtain the expanded region of interest R. extend,p This dynamic expansion mechanism ensures that the target and its surrounding background information are fully included in the analysis. It can effectively alleviate the problem of missing information caused by blurred object edges and overly tight cropping, and helps to improve the accuracy and robustness of subsequent feature judgment.

[0035] In step S6, the candidate region is iteratively reconstructed using a projection-to-convex-set super-resolution algorithm, combined with Lanczos interpolation and Gaussian blurring to improve the texture details of the low-resolution region.

[0036] Furthermore, the process of step S6 is as follows:

[0037] Each region R extend,p The image was magnified by a factor of two using bilinear interpolation and then subjected to 50 iterations of the PoCS algorithm for image reconstruction, resulting in a high-resolution image R with greater detail. high-resThe image is then converted to HSV color space, and the luminance (V) and saturation (S) channels are extracted for subsequent analysis. The image obtained through this step has higher texture clarity and can effectively support feature comparison analysis at the color, structure, and lighting levels. It also has good performance when processing small or blurred targets.

[0038] In step S7, the color histogram similarity analysis includes calculating candidate regions R. extend With reference image I ref The color histogram similarity ρ of the corresponding regions is used. If the similarity ρ > 0.5, the region is considered a static background and is not retained. Brightness contrast includes comparing candidate regions R. extend Average gray value inside and outside; shadow determination includes analyzing candidate regions R extend The brightness and saturation characteristics are analyzed. If the saturation is below 20 and the brightness is less than 25% of the maximum brightness, it is considered to be caused by lighting shadows. The structural similarity calculation includes the analysis of candidate regions R. extend and in reference image I ref SSIM calculation is performed on the corresponding region. If the SSIM value is greater than 0.5, it is considered that the change is insufficient and is excluded.

[0039] Furthermore, the process of step S7 is as follows:

[0040] For each high-resolution image region H after conversion cand Perform the following feature analysis and judgment respectively:

[0041] 1. Region H, same as the reference diagram ref Perform color histogram matching, and calculate the formula as follows. If the correlation ρ>0.5, it is considered a static background and will not be retained.

[0042]

[0043] In the formula, The standard deviation of the candidate region histogram. This represents the standard deviation of the histogram for the reference region.

[0044] 2. Use the SSIM structural similarity index to evaluate the structural changes in the target region. The calculation formula is as follows. If the SSIM value is greater than 0.5, the changes are considered insufficient and the region can be excluded.

[0045]

[0046] In the formula, For candidate region grayscale image G cand Average brightness value, For the reference region grayscale image G ref Average brightness value, Let V be the variance of the grayscale image of the candidate region. The variance of the grayscale image in the reference area. C1 and C2 are small constants used to avoid zero denominators, representing the covariance between the candidate region and the reference region grayscale images.

[0047] 3. Statistically measure the average saturation and brightness of the area. If the saturation is below 20 and the brightness is less than 25% of the maximum brightness, it is considered to be caused by lighting and shadows.

[0048] 4. Perform local brightness enhancement contrast analysis on the grayscale image. If the brightness difference between the central area and the edge exceeds 20% after the contrast enhancement, it is identified as a potential target.

[0049] After filtering through the four dimensions mentioned above, only regions that simultaneously meet the conditions of significant change, large color difference, and no shadow occlusion are retained as valid candidate targets. This multi-dimensional feature fusion judgment strategy significantly improves the filtering accuracy and avoids misjudgment caused by relying on a single feature, especially showing higher adaptability in complex environments such as highly reflective road surfaces and uneven nighttime lighting.

[0050] In step S8, the number of consecutive frames that meet the IOU threshold is counted by calculating the IOU value of the candidate region in consecutive frames, and only the detection targets that appear stably in multiple consecutive frames are retained.

[0051] Furthermore, the process of step S8 is as follows:

[0052] During multi-frame detection, the Interchange of Units (IOU) metric is used to analyze the temporal continuity between candidate targets. If two adjacent frames contain candidate targets with an IOU ≥ 0.5, they are considered to be the same target entity. The frequency of target occurrence in consecutive frames is counted, and short-term false detection targets appearing only in isolated frames are removed. Finally, the candidate set that passes consistency verification is output as the projectile detection result. This temporal consistency strategy effectively filters random noise or occasional interference, enhancing the continuity and reliability of the detection results.

[0053] The beneficial effects of this invention are:

[0054] 1. The method of the present invention introduces multi-dimensional feature fusion methods such as color and structural similarity calculation and HSV brightness and saturation analysis during the candidate region screening process, and constructs a multi-level filtering mechanism, which can significantly reduce the false detection rate caused by factors such as shadows, water stains, and road surface pollution.

[0055] 2. The method of this invention integrates image super-resolution reconstruction and brightness enhancement technology, which improves the discriminability of images in low light, night and backlight environments, and makes the detection process adaptable to various lighting and weather conditions.

[0056] 3. The method of the present invention introduces a time consistency verification mechanism before the detection result is output. By comparing the IOU continuity of candidate targets between frames, it effectively removes the occasional false alarm area and ensures that the output result has stability and temporal continuity.

[0057] 4. The method of this invention is mainly based on image frame difference analysis and traditional image processing methods to realize the detection logic. It only relies on the pre-trained YOLO model in the target exclusion stage, without the need for large-scale retraining. It has good interpretability and easy deployment, and is suitable for the rapid integration and expansion of existing video surveillance systems. Attached Figure Description

[0058] Figure 1 This is a flowchart of the method of the present invention;

[0059] Figure 2 This is a schematic diagram of the projectile in a specific embodiment of the present invention;

[0060] The following will describe in detail, with reference to the accompanying drawings, embodiments of the present invention. Detailed Implementation

[0061] The principles and features of the present invention are described below with reference to the accompanying drawings. The embodiments given are for illustrative purposes only and are not intended to limit the scope of the invention. The invention is described more specifically in the following paragraphs by way of example with reference to the accompanying drawings. The advantages and features of the invention will become clearer from the following description. It should be noted that the drawings are in a very simplified form and use non-precise proportions, and are only used to facilitate and clarify the illustration of the embodiments of the invention.

[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0063] The present invention will be further described below with reference to the accompanying drawings and embodiments:

[0064] like Figure 1 As shown, the multi-level filtering method for detecting road debris based on dynamic difference analysis includes the following steps:

[0065] Step S1: First, extract the road surface region from the input reference image. Let the input image be... (RGB color space) is converted to HSV color space, and the grayscale thresholds are defined as: H∈[0,180], S∈[0,50], V∈[50,200], to generate a binary mask M for the road surface. road∈{0,1} H×W As shown in formula (1):

[0066]

[0067] Then, the mask is subjected to morphological processing, mainly including two operations: erosion and dilation. First, isolated noise points are eliminated by erosion, and then the mask edges are smoothed by dilation to ensure the continuity and integrity of the road area. The operation process is shown in formula (2) and formula (3) respectively:

[0068]

[0069] In the formula Indicates morphological corrosion, Indicates morphological expansion, B erode and B dilate These represent the structural elements for corrosion and expansion operations, respectively.

[0070] The processed mask M smoothed It serves as a benchmark for determining the scene consistency between subsequent video frames and the reference image.

[0071] Step S2: Extract a set of video frames from the video stream to be detected at a preset frequency (e.g., 1 frame per second). Ensure the video coverage changes dynamically. For each keyframe in this set, process it using the method in step S1 to obtain the video frame mask M. frame Calculate the video frame mask M frame With reference image mask M smoothed Similarity S mask The similarity calculation formula is shown in formula (4). If it is lower than the preset threshold (such as 0.5), it is considered that the scene difference is too large and the current video stream processing is skipped; otherwise, subsequent operations are performed.

[0072]

[0073] Step S3: Construct a target detection dataset for highway scenes, labeling categories including vehicles, pedestrians, guardrails, etc. Train the YOLO model on this dataset to obtain the YOLOv8 model with better detection performance in highway scenes.

[0074] The trained YOLOv8 model is used to perform object detection on reference images and video keyframes, identifying vehicles, pedestrians, and guardrails in the images, and outputting a set of object bounding boxes. Among them B i =(x i ,y i ,w i ,h iThe horizontal and vertical axes are respectively proportional to the specified width and height α of the target. x and α y Expand the target box to obtain B' i =(x i -α x w i ,y i -α y h i ,x i +(1+α x )w i ,y i +(1+α y )h i This ensures complete coverage of the target and surrounding interference areas.

[0075] Using the road mask of the reference image as a reference, the expanded interference target area and non-road area are filled with pure black, as shown in formula (5), to obtain the processed reference image I'. ref and video frame I' k This allows for the elimination of these dynamic targets and external disturbances during subsequent processing.

[0076]

[0077] In the formula I original (x,y) represents the pixel at (x,y) in the original reference image or video frame.

[0078] Step S4: Place the reference image I' ref and video frame I' k Convert to grayscale image G' ref , Two images are smoothed using a Gaussian blur algorithm to reduce noise. The absolute difference between the two grayscale images is calculated, and a threshold (T) is set. diff =30) Perform binarization processing to generate the initial dynamic target mask M. diff As shown in formulas (6) and (7) respectively:

[0079] D(x,y)=|G′ ref (x,y)-G′k(x,y)|; (6)

[0080]

[0081] The Gaussian Mixture Background Modeling (MOG2) algorithm is used to extract the foreground from video frames and generate a foreground target mask M. fg .

[0082] The dynamic target mask M obtained by the frame difference method diffForeground mask M for background modeling fg Perform a logical AND operation to fuse the detection results from both tests and generate a fused mask M. candidate .

[0083] The fused mask is first subjected to median filtering (kernel size 5×5) to remove salt-and-pepper noise. Then, a closing operation (kernel size 3×3, iterations 2) is used to connect broken regions and smooth the target edges. The contours of connected regions are extracted, and candidate boxes with width and height within a preset range are selected, i.e., 20 ≤ w. p ,h p ≤100, where w p and h p Let be the width and height of the p-th candidate box.

[0084] Step S5: For each candidate target, given the expansion ratio β, dynamically expand the candidate box according to its size to ensure that the target edge and surrounding environment information are fully captured. The expanded range is shown in formulas (8) and (9):

[0085] x′ min =x p -βw p ,x′ max =x p +(1+β)w p (8)

[0086] y′ min =y p -βh p ,y′ max =y p +(1+β)h p (9)

[0087] Based on the above extended coordinates, generate mask M. ROI,p ∈{0,1} H×W Extract the extended ROI region R extend,p For subsequent feature analysis, as shown in formula (10):

[0088] R extend,p =I k ⊙N ROI,p (10)

[0089] In the formula, I k These are keyframes extracted from the video, and ⊙ indicates element-wise multiplication.

[0090] Step S6: For the expanded low-resolution region Magnify using bilinear interpolation. As shown in formula (11):

[0091]

[0092] In the formula, (u,v) represents the target high-resolution image R. init The coordinates (i,j) in the image are the original low-resolution image R. extend,p The coordinates in the diagram.

[0093] The magnified region is iteratively optimized 50 times, as shown in formula (12). By combining Gaussian blur and image reconstruction constraints, the image detail clarity is gradually improved and the jagged effect is reduced, resulting in the final high-resolution image R. high-res .

[0094]

[0095] In the formula, For convex set projection constraints, Let be a Gaussian kernel function, and σ g =1.0.

[0096] The enhanced high-resolution image R high-res Convert from RGB color space to HSV color space, and extract brightness (V) and saturation (S) for subsequent shadow determination.

[0097] Step S7: Perform multi-dimensional false alarm filtering on each candidate box in the candidate region, and retain only the candidate boxes that meet the conditions.

[0098] S701, Color Histogram Similarity Analysis:

[0099] Candidate regions H after super-resolution reconstruction cand Corresponding region H in the reference figure ref Divide the data into 8×8×8 histograms and calculate the histograms for the three RGB channels respectively. Normalize the histograms and compare the similarity ρ between the two using the correlation coefficient method, as shown in formula (13). If ρ>0.5, it is determined to be a false background detection and does not meet the basic conditions for spilled material.

[0100]

[0101] In the formula, The standard deviation of the candidate region histogram. This represents the standard deviation of the histogram for the reference region.

[0102] S702, Structural similarity verification:

[0103] Candidate regions H after super-resolution reconstruction cand Corresponding region H in the reference figure ref Convert to grayscale image G cand and G refCalculate the similarity index (SSIM) between the two, as shown in formula (14). If SSIM>0.5, it is determined to be a false background detection, which does not meet the basic conditions of spilled material.

[0104]

[0105] In the formula, For candidate region grayscale image G cand Average brightness value, For the reference region grayscale image G ref Average brightness value, Let V be the variance of the grayscale image of the candidate region. The variance of the grayscale image in the reference area. C1 and C2 are small constants used to avoid zero denominators, representing the covariance between the candidate region and the reference region grayscale images.

[0106] S703, Shadow Detection:

[0107] Calculate candidate region H cand Average saturation in HSV space and average brightness like and It is then identified as a shadow and excluded.

[0108] S704, Brightness Contrast:

[0109] Calculate candidate region H cand The internal average gray value μ in and the average gray value μ of the surrounding area out .like It is then identified as a potential spill.

[0110] After filtering through multiple dimensions, the target boxes that meet all the filtering conditions are retained as the final candidate boxes.

[0111] Step S8: Calculate the Intersection over Union (IOU) of the candidate boxes in consecutive frames, as shown in Formula (15). If IOU ≥ 0.5, the candidate box is determined to be a unified target.

[0112]

[0113] In the formula, B t and B t+1 These are the candidate boxes in frame t and frame t+1, respectively.

[0114] Count the number of times the target appears in frame T, N count ,like Then retain the target and eliminate transient interference (such as birds or short-lived objects).

[0115] The final retained detection bounding boxes are then visually expanded, and the locations of the spilled material are marked on the image, such as... Figure 2 As shown, the labeled image is saved to the specified path, and a JSON file containing coordinate information is output.

[0116] The present invention has been described above by way of example with reference to the accompanying drawings. Obviously, the specific implementation of the present invention is not limited to the above-described manner. Any improvements made using the inventive concept and technical solution of the present invention, or direct application to other occasions without modification, are all within the protection scope of the present invention.

Claims

1. A multi-level filtering method for detecting road spills based on dynamic difference analysis, characterized in that, Includes the following steps: Step S1: Obtain the reference image as I ref And generate the reference mask M smoothed ; Step S2: Extract video frame I from the video to be detected. k Calculate the video frame mask M frame With reference mask M smoothed Similarity S mask To determine whether to conduct spill detection; Step S3: Train the YOLO object detection model and use the trained YOLOv8 model to detect the reference image I. ref and video frame I k Target detection is performed to obtain vehicles, pedestrians, and guardrails in the image. Vehicles, pedestrians, guardrails, and off-road areas are then filled with a solid color to obtain a reference image I' after eliminating dynamic interference targets. ref and video frame I' k ; Step S4: Apply Gaussian mixture background modeling algorithm to I' ref and I' k Frame difference processing is performed to obtain the initial dynamic target candidate region M. diff ; Step S5, for M diff For each candidate target in the dataset, the candidate bounding box is dynamically expanded to obtain the expanded candidate region R. extend ; Step S6, for R extend Super-resolution enhancement is performed to obtain a high-resolution image R. high-res Then R high-res Candidate region H is obtained by converting from RGB color space to HSV color space. cand ; Step S7: For candidate region H cand Perform color histogram similarity analysis, structural similarity analysis, shadow determination, and brightness comparison to comprehensively filter out non-sprayed targets; Step S8: Calculate the frequency of candidate regions in consecutive frames, combine with the IOU threshold to filter stable targets, and output the coordinates of the projectiles and the visualization results.

2. The multi-stage filtering method for detecting road spills based on dynamic difference analysis according to claim 1, characterized in that, In step S1, low-saturation, medium-brightness road regions are extracted using a preset threshold as a reference mask M. smoothed , where the color thresholds are H∈[0,180], S∈[0,50], and V∈[50,200].

3. The multi-stage filtering method for detecting road spills based on dynamic difference analysis according to claim 1, characterized in that, In step S2, the road mask M of the video frame is first calculated. frame With reference mask M smoothed Similarity S mask : When the similarity S mask Perform the test when the value is higher than 0.

5.

4. The multi-stage filtering method for detecting road spills based on dynamic difference analysis according to claim 1, characterized in that, In step S3, the YOLO object detection algorithm is used to identify the reference image I. ref and video frame I k The system detects vehicles, pedestrians, and guardrails, and eliminates their interference with debris detection by masking.

5. The multi-stage filtering method for detecting road spills based on dynamic difference analysis according to claim 1, characterized in that, In step S4, the frame difference processing steps are as follows: Reference image I ref and video frame I k The image is converted to grayscale and blurred using Gaussian. The difference between the two frames is calculated and then binarized to obtain the initial dynamic target candidate region M. diff .

6. The multi-stage filtering method for detecting road spills based on dynamic difference analysis according to claim 1, characterized in that, In step S5, the expansion ratio is dynamically calculated based on the target size, and the detection box is expanded spatially using the quadratic boundary expansion method. The expanded area is then masked.

7. The multi-stage filtering method for detecting road spills based on dynamic difference analysis according to claim 1, characterized in that, In step S6, the candidate region is iteratively reconstructed using a projection-to-convex-set super-resolution algorithm, combined with Lanczos interpolation and Gaussian blurring to improve the texture details of the low-resolution region.

8. The multi-stage filtering method for detecting road spills based on dynamic difference analysis according to claim 1, characterized in that, In step S7, the color histogram similarity analysis includes calculating candidate regions R. extend With reference image I ref The color histogram similarity ρ of the corresponding regions is used. If the similarity ρ > 0.5, it is considered a static background and not retained. Brightness contrast includes comparing candidate regions R. extend Average gray value inside and outside; shadow determination includes analyzing candidate regions R extend The brightness and saturation characteristics are considered. If the saturation is below 20 and the brightness is less than 25% of the maximum brightness, it is considered to be caused by lighting and shadows. Structural similarity calculation includes the calculation of candidate regions R extend and in reference image I ref SSIM calculation is performed on the corresponding region. If the SSIM value is greater than 0.5, it is considered that the change is insufficient and is excluded.

9. The multi-stage filtering method for detecting road spills based on dynamic difference analysis according to claim 1, characterized in that, In step S8, the number of consecutive frames that meet the IOU threshold is counted by calculating the IOU value of the candidate region in consecutive frames, and only the detection targets that appear stably in multiple consecutive frames are retained.

Citation Information

Cited By

  • Crack detection method and device based on bounding box guide constraint and electronic equipment

    CN121504915A

  • Crack detection method and device based on bounding box guidance constraint and electronic equipment

    CN121504915B

  • Industrial scene anomaly detection method based on multi-dimensional feature decoupling and double-track state machine

    CN121504931A

  • Road cleanliness detection method and device, electronic equipment and storage medium

    CN121661052A

  • Road cleanliness testing methods, devices, electronic equipment and storage media

    CN121661052B