Foreign matter segmentation method based on dynamic background suppression and weak supervised learning
By using the methods of dynamic background suppression and weakly supervised learning, and utilizing fully connected autoencoders and U-Net networks, the problem of foreign object segmentation in complex dynamic backgrounds is solved, thereby improving the accuracy and adaptability of foreign object detection.
Patent Information
- Application Number
- CN202511169916.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-10-10
AI Technical Summary
Existing foreign object segmentation technology has insufficient performance when dealing with complex dynamic backgrounds, especially in multimodal dynamic scenarios where the false detection rate is high, making it difficult to deploy on a large scale in railway foreign object detection.
A method based on dynamic background suppression and weakly supervised learning is adopted. The static background is reconstructed by a fully connected autoencoder. The dynamic background is learned by combining a Gaussian model and a U-Net network. A sliding window and neighborhood consistency constraints are used to generate a foreign object probability map and perform morphological post-processing.
It significantly improves the detection accuracy of foreign body segmentation, reduces the interference of dynamic background on foreground detection, and improves the adaptability of foreign body monitoring in complex environments.
Smart Images

Figure CN120766104A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and intelligent monitoring technology, specifically a foreign object segmentation method based on dynamic background suppression and weakly supervised learning. This method is particularly suitable for outdoor monitoring scenarios with complex dynamic interference (such as swaying vegetation, water surface ripples, and sudden changes in illumination). It can be widely used in transportation hub perimeter protection, railway safety protection, and railway foreign object detection. Background Art
[0002] The current mainstream foreign body segmentation technology has the following bottlenecks:
[0003] 1. Performance needs improvement: Traditional Gaussian mixture model (GMM)-based methods assume that background pixels follow a static distribution, making them difficult to handle multimodal dynamic scenes (such as the simultaneous presence of swaying trees and flowing water). For example, in the "badWeather" category of the CDnet 2014 dataset, the GMM method has a false positive rate of up to 41.7%.
[0004] 2. Insufficient modeling capabilities for dynamic backgrounds: Existing methods are insufficiently capable of processing the periodic motion of dynamic backgrounds, such as leaves and flowing water. This is especially true when faced with scenes such as sudden changes in lighting and shadows at night, where the false detection rate increases significantly.
[0005] Summary of core issues: Existing methods have significant deficiencies in dynamic background modeling and their performance needs to be improved, which limits their large-scale deployment in railway foreign object detection industrial scenarios. Summary of the Invention
[0006] In order to solve the above problems in the prior art, the present invention proposes a foreign body segmentation method based on dynamic background suppression and weakly supervised learning.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] A foreign body segmentation method based on dynamic background suppression and weakly supervised learning includes the following steps:
[0009] Step 1: Train a static background reconstruction network. Filter video frames that do not contain foreign objects from the input video sequence and input them into a fully connected autoencoder. The fully connected autoencoder consists of two parts: an encoder and a decoder. The encoder learns and outputs a normalized static background image by minimizing the L1 reconstruction loss function.
[0010] Step 2: Based on the static background image reconstructed in step 1, a Gaussian model is used for adaptive background modeling. A sliding window background queue of length K frames is maintained. The static background corresponding to each pixel is updated based on the sliding average. The foreground pixels are determined using the 3σ criterion, and a neighborhood consistency constraint is introduced to filter dynamic noise to obtain a binary image of the foreground.
[0011] Step 3: Automatically or manually identify the area containing the dynamic background using the binary image of the foreground. Filter the sequences containing only the dynamic background and no foreign objects as training data, generate binary labels for the dynamic background, and train the U-Net network to predict the probability distribution of the dynamic background.
[0012] Step 4: Perform pixel-by-pixel operations on the foreground binary image from step 2 and the dynamic background probability predicted by the U-Net network in step 3 to suppress dynamic background noise and generate a foreign object probability map;
[0013] Step 5: Binarize the foreign body probability map obtained in step 4, and then perform 3×3 corrosion operation and 5×5 expansion operation in sequence for morphological post-processing to remove isolated noise and restore the integrity of the foreign body contour to obtain the final foreign body segmentation result.
[0014] The structure of the fully connected autoencoder described in step 1 is:
[0015] The encoder contains four fully connected layers with 512, 256, 128, and 64 neurons, respectively, and uses the SELU activation function; the decoder contains four fully connected layers with 128, 256, 512, and 1024 neurons, respectively. The first three layers use the SELU activation function, and the last layer uses the Sigmoid activation function to normalize the output to the range [0, 1].
[0016] The L1 reconstruction loss function is defined as:
[0017]
[0018] in: is the reconstruction loss function, B′ is the reconstructed background output by the autoencoder, B is the input background frame without foreign objects, (x, y) is the pixel coordinate, B′(x, y) is the pixel value of the reconstructed background at the pixel point (x, y), B(x, y) is the pixel value of the input background frame without foreign objects at the pixel point (x, y), N is the number of pixels in the x direction, and M is the number of pixels in the y direction; this loss function guides the network training by constraining the L1 loss between the reconstructed image and the original image, that is, the absolute value of the pixel-by-pixel difference.
[0019] The strategy for updating the static background corresponding to each pixel based on sliding average in step 2 is:
[0020] Input current video frame I t , using the static background reconstruction network trained in the previous step, obtain the corresponding static background B t ; and update the background using a sliding average, as shown below:
[0021] BGt =α·BG t-1 +(1-α)·B t ,
[0022] Where α∈[0.9, 0.99] is the forgetting factor, B t The current frame corresponds to the static background frame estimated by the neural network in step 2, BG t is the background model obtained by sliding average at time t, BG t-1 is the background model obtained by sliding average at time t-1; the criterion for determining foreground foreign objects is:
[0023]
[0024] in It is an indicator function, which outputs 1 if the condition is met and 0 if the condition is not met; F t is the binary image at time t; (x, y) is the coordinate of the pixel point, F t (x, y) is the value of the binary image at the coordinate (x, y) at time t; I t (x, y) is the video frame I at time t t The value at the pixel point (x, y), BG t (x, y) is the value of the background model at the pixel point (x, y) obtained by sliding average, σ t (x, y) is the standard deviation of the pixel (x, y) at time t, and its calculation process is as follows:
[0025]
[0026] Among them, σ t (x, y) is the standard deviation of the pixel (x, y) at time t, s is the time index symbol, tK is the starting time, t-1 is the ending time, I s (x, y) is the video frame I at time s s The value at the pixel point (x, y), BG s (x, y) is the value of the background model obtained by sliding average at the pixel point (x, y) at time s, and K is the total number of video frames in the sequence window;
[0027] At the same time, a neighborhood consistency constraint is imposed: when the proportion of foreground pixels in the 3×3 neighborhood of a pixel is lower than the threshold θ=0.1, the pixel is reset to the background, thereby obtaining the binary image F at the updated time t t .
[0028] The U-Net network structure described in step 3 includes:
[0029] Encoder part: 4 downsampling modules, each module contains two 3×3 convolutional layers, ReLU activation and 2×2 maximum pooling;
[0030] Decoder part: 4 upsampling modules, each module contains 2×2 deconvolution, feature concatenation and two 3×3 convolution layers;
[0031] Skip connection: concatenates the features of each encoder layer with the corresponding decoder layer to preserve multi-scale information;
[0032] Output layer: 1×1 convolution followed by Sigmoid activation to output dynamic background probability map.
[0033] The U-Net network learns the dynamic background in the video frame and inputs the video frame I at the current time t. t , output the probability map D of the dynamic background of the image t ;
[0034] The binary graph F at time t is obtained in step 2 t As a label, automatically filter out the pixels of the input dynamic background to train the U-Net network; automatically filter the state change frequency V of the pixel point (x, y) in the time window t (x, y) is given by:
[0035]
[0036] Among them, s is the time index symbol, t-k+1 is the starting time, t-1 is the ending time, F s is the binary image corresponding to the initial foreground residual at time s, F s (x, y) is the value of the binary image corresponding to the initial foreground residual at time s at (x, y), XOR is the exclusive OR operation, K is the total number of video frames in the sequence window; if (x, y) is a dynamic background, then V t The value of (x, y) is larger, otherwise it is smaller; V is automatically filtered out through a threshold t Pixels with larger (x, y) values are used to train the U-Net network; or dynamic backgrounds are manually filtered to train the U-Net network.
[0037] The calculation formula for the foreign matter probability map in step 4 is:
[0038] P t (x,y)=F t (x,y)·(1-D t (x,y))·C t (x,y)
[0039] Among them, (x, y) is the coordinate of the pixel point, P tis the output probability map, F t is the initial foreground residual, D t is the probability of dynamic background, C t is the time consistency weight; P t (x, y) is the value of the output probability map at the pixel point (x, y), F t (x, y) is the value of the initial foreground residual at the pixel point (x, y), D t (x, y) is the probability of dynamic background at the pixel point (x, y), C t (x, y) is the value of the temporal consistency weight at the pixel point (x, y);
[0040] The calculation formula of the value of the temporal consistency weight at the pixel point (x, y) is as follows:
[0041] C t (x,y)=exp(-λ·V t (x,y))
[0042] Where λ is the adjustment parameter, V t (x, y) is the frequency of state change of pixel (x, y) within the time window, which is given by the following formula:
[0043]
[0044] Among them, s is the time index symbol, t-K+1 is the starting time, t-1 is the ending time, F s is the binary image corresponding to the initial foreground residual at time s, F s (x, y) is the value of the binary image corresponding to the initial foreground residual at time s at (x, y), XOR is the exclusive OR operation, and k is the total number of video frames in the sequence window.
[0045] The adaptive threshold for binarization in step 5 is automatically determined based on the OTSU algorithm.
[0046] The specific method of morphological post-processing in step 5 is:
[0047] Use 3×3 corrosion operation to eliminate isolated noise points; purpose: remove small isolated noise points, smooth the boundaries of objects, and eliminate small protrusions;
[0048] A 5×5 dilation operation is used to restore the complete outline of the foreign object; the purpose is to restore the complete outline of the target foreign object that was reduced by the corrosion operation and fill the void or broken parts.
[0049] Perform connected region analysis on the binary image after morphological operation, calculate the area of each connected region, that is, the number of pixels in each connected region, and filter the connected regions whose area is less than the minimum threshold Amin, where Amin is the defined minimum area threshold.
[0050] In step 2, the value range of K frame is 50-200 frames;
[0051] In step 3, training data is obtained through manual annotation or automatic screening based on motion consistency.
[0052] Beneficial effects
[0053] Quantitative comparisons of the proposed method with other state-of-the-art foreground detection methods on CDnet 2014 and railway surveillance video datasets revealed that the proposed method achieved superior performance in terms of detection accuracy (F1 score), a comprehensive indicator of responsiveness. For the "camera shake" feature in CDnet 2014, the proposed method achieved a score of 0.945, a 5.4% improvement over 0.904 (BSUV-2.0); for the "dynamic background" feature, the proposed method achieved a score of 0.898, a 5.0% improvement over 0.848 (SemanticBGS); and for the "bad weather" feature, the proposed method achieved a score of 0.922, a 3.2% improvement over 0.889 (STPNet).
[0054] The constructed railway surveillance video consists of 1,000 short railway surveillance videos, including scenes such as rainy and snowy days, mud-covered railways, pedestrians passing by, falling rocks, mudslides, etc., and the railway foreign object foreground is marked as a label. The final detection accuracy (F1 score) was 0.909, which is better than all other methods. It is proved that the method of the present invention has good performance under various complex conditions, including dynamic noise, low-light noise at night, and railway surveillance videos. Since the probability map of the dynamic background is learned and the weight of the dynamic background area is suppressed, the impact on foreground detection is reduced, the interference of the dynamic background on the foreground object detection is suppressed, the versatility of moving object detection under various lighting and various environments is improved, and the adaptability of foreign object monitoring is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a flow chart of the method of the present invention.
[0056] Figure 2 It is a network structure used for static background recovery.
[0057] Figure 3 It is a U-Net network structure for dynamic background detection.
[0058] Figure 4 This is the optimized result diagram for processing dynamic background.
[0059] Figure 5 This is a qualitative result diagram for processing the cdnet2014 dataset and the railway surveillance video dataset. DETAILED DESCRIPTION
[0060] The present invention is further described in detail below with reference to the accompanying drawings and specific examples.
[0061] Figure 1 The flowchart of the method of the present invention is as follows. Specifically, in order to realize the detection of moving objects in the video, the present invention is divided into five parts:
[0062] 1) Train a static background reconstruction network, select foreign object-free frames from the video sequence as the input of the fully connected autoencoder, and learn the pure static background of the scene by minimizing the L1 reconstruction loss; 2) Based on the reconstructed static background, use a Gaussian model for adaptive background modeling, maintain a sliding window background queue, and improve the model robustness through the 3σ criterion and neighborhood consistency constraint; 3) Use the background subtraction method to identify areas containing foreign objects and dynamic background, select sequences containing only dynamic background to train the U-Net network, so that it learns to predict periodic dynamic background patterns; 4) Perform pixel-by-pixel operations on the foreground binary image in step 2 and the dynamic background probability predicted by the U-Net network in step 3 to suppress dynamic background noise and generate a foreign object probability map; 5) Perform 3×3 corrosion and 5×5 dilation morphological post-processing on the foreign object probability map to optimize the integrity of the foreign object contour and obtain the final foreign object segmentation result.
[0063] Figure 2 This network structure is used for static background restoration, using a fully connected autoencoder to generate static background images. It consists of two components: an encoder and a decoder. The encoder maps the input to a compressed code, and the decoder reconstructs the input based on the code, aiming to make the output as close to the input as possible. The encoder contains four fully connected layers with 512, 256, 128, and 64 neurons, respectively, using the SELU activation function. The decoder contains four fully connected layers with 128, 256, 512, and 1024 neurons, respectively. The first three layers use the SELU activation function, and the final layer uses the Sigmoid activation function to normalize the output to the [0, 1] range.
[0064] Figure 3 It is a U-Net network structure for dynamic background detection. The U-Net network learns to predict dynamic background and uses the U-Net network to learn dynamic background pixels.
[0065] Construct a dynamic background detection network. Use the binary image obtained in step 2 as a label and manually or automatically filter out the dynamic background based on it to train the dynamic detection network U-Net, which mainly learns the probability map of the dynamic background.
[0066] Encoder part: 4 downsampling modules, each module contains two 3×3 convolutional layers, ReLU activation and 2×2 maximum pooling;
[0067] Decoder part: 4 upsampling modules, each module contains 2×2 deconvolution, feature concatenation and two 3×3 convolution layers;
[0068] Skip connection: concatenates the features of each encoder layer with the corresponding decoder layer to preserve multi-scale information;
[0069] Output layer: 1×1 convolution followed by Sigmoid activation to output dynamic background probability map.
[0070] The binary image obtained in step 2 is used as a label. Since it shows a dynamic background, the U-Net network is used to learn the dynamic background in the video frame and the output is a probability map D of the dynamic background. t ; Using D t Correct the binary image and remove the dynamic background to improve the monitoring effect.
[0071] Evaluation indicators: Figure 4 and Figure 5 The qualitative results in Figure 1 and the quantitative results in Tables 1 and 2 are shown. The input is a surveillance video with a dynamic background. After processing by the present invention, the output video is a binary image of the location of moving objects without the dynamic background noise. The evaluation metric is detection accuracy (F1 score), which is defined by the following metric:
[0072] Number of pixels correctly identified as foreground (TP)
[0073] Number of pixels correctly identified as background (TN)
[0074] Number of background pixels incorrectly identified as foreground (FP)
[0075] Number of foreground pixels incorrectly identified as background (FN)
[0076] Accuracy = TP / (TP+FP);
[0077] Recall = TP / (TP+FN)
[0078] The detection index is defined as:
[0079] Detection accuracy (F1 score) = 2*(precision*recall) / (precision+recall)
[0080] The quantitative and qualitative results of the proposed method are compared with other state-of-the-art foreground detection methods on the CDNet2014 and railway surveillance video datasets. The comparison results are shown in Tables 1 and 2.
[0081] Table 1: Comparative results of the proposed method on CDNet-2014
[0082] method Benchmark Camera shake Dynamic Background Bad weather Low resolution F1 score STPNet 0.958 0.7721 0.805 0.889 0.729 0.799 BUSV-2.0 0.962 0.904 0.814 0.884 0.790 0.838 Semantic 0.950 0.815 0.848 0.861 0.559 0.751 SWCD 0.921 0.741 0.764 0.823 0.737 0.758 RTSS 0.959 0.836 0.735 0.866 0.671 0.792 RT-SBS 0.965 0.823 0.821 0.829 0.734 0.804 The present invention 0.963 0.945 0.898 0.922 0.714 0.832
[0083] Table 2: Comparative effect of the present invention on a railway monitoring video set constructed from 1000 videos
[0084] method Accuracy Recall F1 score STPNet 0.872 0.815 0.842 BUSV-2.0 0.891 0.843 0.866 Semantic 0.855 0.802 0.827 SWCD 0.832 0.784 0.807 RTSS 0.868 0.812 0.839 RT-SBS 0.883 0.834 0.858 The present invention 0.923 0.896 0.909
[0085] Among them, in terms of the comprehensive indicator of detection accuracy (F1 score), for "camera shake" in CDnet 2014, the present invention obtained a score of 0.945, which is 5.4% higher than 0.904 (BSUV-2.0); for "dynamic background", the present invention obtained a score of 0.898, which is 5.0% higher than 0.848 (SemanticBGS); for "bad weather", the present invention obtained a score of 0.922, which is 3.2% higher than 0.889 (STPNet).
[0086] The constructed railway surveillance video consists of 1,000 short railway surveillance videos, including scenes such as rainy and snowy days, mud-covered railways, pedestrians passing by, falling rocks, mudslides, etc., and the railway foreign object foreground is marked as a label. The final detection accuracy (F1 score) was 0.909, which is better than all other methods. It is proved that the method of the present invention has good performance under various complex conditions, including dynamic noise, low-light noise at night, and railway surveillance videos. Since the probability map of the dynamic background is learned and the weight of the dynamic background area is suppressed, the impact on foreground detection is reduced, the interference of the dynamic background on the foreground object detection is suppressed, the versatility of moving object detection under various lighting and various environments is improved, and the adaptability of foreign object monitoring is improved.
[0087] Figure 4 This is the optimization result diagram for the dynamic background. The first column is the original video frame, the second column is the reconstructed background image, and the third column is the foreign object foreground segmented using the method of the present invention. The first row is a picture of an ordinary traffic road. One frame is taken out for analysis. The foreign object detection method of the dynamic background detection network of the present invention can optimize the noise caused by the wind blowing the leaves nearby. The second row is a scene where people are obscured indoors. Due to the dynamically updated background, an unobstructed foreground can be obtained, and the accuracy is optimized. The third row is the dynamic background optimization of the water surface ripples, which removes the dynamic noise of the water surface ripples. The fourth row is the result of removing the noise of large-scale leaf shaking.
[0088] Figure 5The following are qualitative results for processing the Railway Monitoring Dataset and the Cdnet 2014 dataset. In the qualitative results, a, b, c, d, and e represent the original video frame, the groundtruth label, the results of our method, the results of BSUV-2.0, and the results of SemanticBGS, respectively. In Cdnet 2014, the first row shows that our method performs better than the other two methods for optimizing dynamic wind-blown leaf noise. The second row shows that our method is superior in removing dynamic background noise from glass in the video. The third row primarily compares the removal of dynamic human figures, where our method performs better. The fourth row shows that our method performs better for vehicle segmentation in rainy and snowy weather, removing dynamic snowflake noise.
[0089] On the railway monitoring dataset, compared with the other two methods, the method of the present invention is more effective in foreign body segmentation and noise removal in night scenes, low-resolution monitoring videos, and dynamic noise such as leaf shaking.
Claims
1. A foreign body segmentation method based on dynamic background suppression and weakly supervised learning, characterized by: The following steps are involved: Step 1: Train a static background reconstruction network. Filter video frames that do not contain foreign objects from the input video sequence and input them into a fully connected autoencoder. The fully connected autoencoder consists of two parts: an encoder and a decoder. The encoder learns and outputs a normalized static background image by minimizing the L1 reconstruction loss function. Step 2: Based on the static background image reconstructed in step 1, a Gaussian model is used for adaptive background modeling. A sliding window background queue of length K frames is maintained. The static background corresponding to each pixel is updated based on the sliding average. The foreground pixels are determined using the 3σ criterion, and a neighborhood consistency constraint is introduced to filter dynamic noise to obtain a binary image of the foreground. Step 3: Automatically or manually identify the area containing the dynamic background using the binary image of the foreground. Filter the sequences containing only the dynamic background and no foreign objects as training data, generate binary labels for the dynamic background, and train the U-Net network to predict the probability distribution of the dynamic background. Step 4: Perform pixel-by-pixel operations on the foreground binary image from step 2 and the dynamic background probability predicted by the U-Net network in step 3 to suppress dynamic background noise and generate a foreign object probability map; Step 5: Binarize the foreign body probability map obtained in step 4, and then perform 3×3 corrosion operation and 5×5 expansion operation in sequence for morphological post-processing to remove isolated noise and restore the integrity of the foreign body contour to obtain the final foreign body segmentation result.
2. The foreign body segmentation method according to claim 1, characterized in that: The structure of the fully connected autoencoder described in step 1 is: The encoder contains four fully connected layers with 512, 256, 128, and 64 neurons, respectively, and uses the SELU activation function; the decoder contains four fully connected layers with 128, 256, 512, and 1024 neurons, respectively. The first three layers use the SELU activation function, and the last layer uses the Sigmoid activation function to normalize the output to the range [0, 1]. The L1 reconstruction loss function is defined as: in: is the reconstruction loss function, B′ is the reconstructed background output by the autoencoder, B is the input background frame without foreign objects, (x, y) is the pixel coordinate, B′(x, y) is the pixel value of the reconstructed background at the pixel point (x, y), B(x, y) is the pixel value of the input background frame without foreign objects at the pixel point (x, y), N is the number of pixels in the x direction, and M is the number of pixels in the y direction; this loss function guides the network training by constraining the L1 loss between the reconstructed image and the original image, that is, the absolute value of the pixel-by-pixel difference.
3. The foreign body segmentation method according to claim 1, characterized in that: The strategy for updating the static background corresponding to each pixel based on sliding average in step 2 is: Input current video frame I t , using the static background reconstruction network trained in the previous step, obtain the corresponding static background B t ; and update the background using a sliding average, as shown below: BG t =a·BG t-1 +(1-a).B t , Where α∈[0.9, 0.99] is the forgetting factor, B t The current frame corresponds to the static background frame estimated by the neural network in step 2, BG t is the background model obtained by sliding average at time t, BG t-1 is the background model obtained by sliding average at time t-1; the criterion for determining foreground foreign objects is: in It is an indicator function, which outputs 1 if the condition is met and 0 if the condition is not met; F t is the binary image at time t; (x, y) is the coordinate of the pixel point, F t (x, y) is the value of the binary image at the coordinate (x, y) at time t; I t (x, y) is the video frame I at time t t The value at the pixel point (x, y), BG t (x, y) is the value of the background model at the pixel point (x, y) obtained by sliding average, σ t (x, y) is the standard deviation of the pixel (x, y) at time t, and its calculation process is as follows: Among them, σ t (x, y) is the standard deviation of the pixel (x, y) at time t, s is the time index symbol, tK is the starting time, t-1 is the ending time, I s (x, y) is the video frame I at time s s The value at the pixel point (x, y), BG s (x, y) is the value of the background model obtained by sliding average at the pixel point (x, y) at time s, and K is the total number of video frames in the sequence window; At the same time, a neighborhood consistency constraint is imposed: when the proportion of foreground pixels in the 3×3 neighborhood of a pixel is lower than the threshold θ=0.1, the pixel is reset to the background, thereby obtaining the binary image F at the updated time t t .
4. The foreign body segmentation method according to claim 1, characterized in that: The U-Net network structure described in step 3 includes: Encoder part: 4 downsampling modules, each module contains two 3×3 convolutional layers, ReLU activation and 2×2 maximum pooling; Decoder part: 4 upsampling modules, each module contains 2×2 deconvolution, feature concatenation and two 3×3 convolution layers; Skip connection: concatenates the features of each encoder layer with the corresponding decoder layer to preserve multi-scale information; Output layer: 1×1 convolution followed by Sigmoid activation to output dynamic background probability map. The U-Net network learns the dynamic background in the video frame and inputs the video frame I at the current time t. t , output the probability map D of the dynamic background of the image t ; The binary graph F at time t is obtained in step 2 t As a label, automatically filter out the pixels of the input dynamic background to train the U-Net network; automatically filter the state change frequency V of the pixel point (x, y) in the time window t (x, y) is given by: Among them, s is the time index symbol, t-K+1 is the starting time, t-1 is the ending time, F s is the binary image corresponding to the initial foreground residual at time s, F s (x, y) is the value of the binary image corresponding to the initial foreground residual at time s at (x, y), XOR is the exclusive OR operation, K is the total number of video frames in the sequence window; if (x, y) is a dynamic background, then V t The value of (x, y) is larger, otherwise it is smaller; V is automatically filtered out through a threshold t Pixels with larger (x, y) values are used to train the U-Net network; or dynamic backgrounds are manually filtered to train the U-Net network.
5. The foreign body segmentation method according to claim 1, characterized in that: The calculation formula for the foreign matter probability map in step 4 is: P t (x,y)=F t (x,y)·(1-D t (x,y))·C t (x,y) Among them, (x, y) is the coordinate of the pixel point, P t is the output probability map, F t is the initial foreground residual, D t is the probability of dynamic background, C t is the time consistency weight; P t (x, y) is the value of the output probability map at the pixel point (x, y), F t (x, y) is the value of the initial foreground residual at the pixel point (x, y), D t (x, y) is the probability of dynamic background at the pixel point (x, y), C t (x,y) is the value of the temporal consistency weight at the pixel point (x,y); The calculation formula of the value of the temporal consistency weight at the pixel point (x, y) is as follows: C t (x,y)=exp(-λ·V t (x,y)) Where λ is the adjustment parameter, V t (x, y) is the frequency of state change of pixel (x, y) within the time window, which is given by the following formula: Among them, s is the time index symbol, t-K+1 is the starting time, t-1 is the ending time, F s is the binary image corresponding to the initial foreground residual at time s, F s (x, y) is the value of the binary image corresponding to the initial foreground residual at time s at (x, y), XOR is the exclusive OR operation, and K is the total number of video frames in the sequence window.
6. The foreign body segmentation method according to claim 1, characterized in that: The adaptive threshold for binarization in step 5 is automatically determined based on the OTSU algorithm.
7. The foreign body segmentation method according to claim 1, characterized in that: The specific method of morphological post-processing in step 5 is: Use 3×3 corrosion operation to eliminate isolated noise points; Purpose: Remove small isolated noise points, smooth the boundaries of objects, and eliminate small protrusions; Use 5×5 dilation operation to restore the complete outline of the foreign body; Purpose: To restore the complete outline of the target foreign body that has been reduced by the corrosion operation and fill the void or broken part. Perform connected region analysis on the binary image after morphological operation, calculate the area of each connected region, that is, the number of pixels in each connected region, and filter the connected regions whose area is less than the minimum threshold Amin, where Amin is the defined minimum area threshold.
8. The foreign body segmentation method according to claim 1, characterized in that: In step 2, the value range of K frame is 50-200 frames; In step 3, training data is obtained through manual annotation or automatic screening based on motion consistency.