Dictionary learning and incoherent item fused deep expansion foreground detection method
By integrating dictionary learning with deep unfolding of incoherent terms for foreground detection, this method addresses the shortcomings of interpretability and robustness in existing technologies, achieving efficient foreground object detection and enhancing adaptability to complex scenarios.
Patent Information
- Application Number
- CN202511084065.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-18
AI Technical Summary
Existing foreground target detection methods suffer from insufficient interpretability and robustness, as well as low computational efficiency.
A deep unfolded foreground detection method that integrates dictionary learning and incoherent terms is proposed. By constructing a deep unfolded foreground detection network, dictionary learning is used to model low-rank background and sparse foreground respectively. Mask generation and low-rank background modeling are combined, an incoherent term loss function is introduced, and the ADMM algorithm is used for optimization.
It improves the interpretability and robustness of the model, enhances its adaptability to complex scenarios, reduces computational costs, and achieves efficient foreground object detection.
Smart Images

Figure CN120976257A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision and image processing technology, specifically to a deep unfolding foreground detection method that integrates dictionary learning and incoherent terms. Background Technology
[0002] Foreground object detection, as one of the core tasks of computer vision, has significant application value in fields such as object recognition, image segmentation, and video surveillance. However, in real-world scenarios, due to factors such as environmental complexity and changes in object size, achieving effective separation of foreground and background remains a major challenge. Currently, foreground object detection methods mainly include traditional low-rank sparse decomposition (RPCA), deep learning methods, and hybrid methods. Traditional low-rank sparse decomposition methods rely on manual parameter tuning, have poor generalization ability, and struggle to handle dynamic backgrounds and camera shake. The black-box nature of deep learning methods leads to insufficient interpretability and requires a large amount of labeled data for training. Hybrid methods, due to the lack of sufficient integration of foreground mask generation and low-rank background modeling in existing deep unfolded networks, result in blurred object edges or artifact residue. Summary of the Invention
[0003] The present invention addresses the problems of insufficient interpretability and robustness, as well as low computational efficiency, in existing foreground object detection methods, and provides a deep unfolding foreground detection method that integrates dictionary learning and incoherent terms.
[0004] To solve the above problems, the present invention is achieved through the following technical solution:
[0005] A method for foreground detection that integrates dictionary learning and deep unfolding of irrelevant terms includes the following steps:
[0006] Step 1: Construct a deep unfolded foreground detection network;
[0007] The deep unfolded foreground detection network consists of an input layer, at least one hidden layer, and an output layer. All hidden layers are connected sequentially. The input of the input layer serves as the input of the deep unfolded foreground detection network. The output of the input layer is connected to the input of the first hidden layer. The output of the last hidden layer is connected to the input of the output layer. The output of the output layer serves as the output of the deep unfolded foreground detection network.
[0008] Each hidden layer consists of 8 sub-blocks: L-block, S-block, X-block, B-block, M-block, Y1-block, Y2-block, and Y3-block. Input is sent to the L-block, and output is sent out from the L-block. Y3 i-1 Inputs are sent to the S-block, and outputs are sent out from the S-block. Input to X-block, output of X-block Y1 i-1 , Y2 i-1 , Input to B-block, output of B-block Y1 i-1 , Input to M-block, output of M-block Input to Y1 -block, output of Y1 -block Y1 i ; Input to Y2-block, output of Y2-block Y2 i ; Input to Y3-block, output of Y3-block Y3 i ;
[0009] where, and represent the auxiliary background variables of the i-th hidden layer and the (i-1)-th hidden layer output, respectively, and represent the auxiliary foreground variables of the i-th hidden layer and the (i-1)-th hidden layer output, respectively, and represent the reconstruction variables of the i-th hidden layer and the (i-1)-th hidden layer output, respectively, and represent the low-rank background variables of the i-th hidden layer and the (i-1)-th hidden layer output, respectively, and represent the sparse foreground variables of the i-th hidden layer and the (i-1)-th hidden layer output, respectively, Y1 i and Y1 i-1 represent the first augmented Lagrange multipliers of the i-th hidden layer and the (i-1)-th hidden layer output, respectively, and represent the second augmented Lagrange multipliers of the i-th hidden layer and the (i-1)-th hidden layer output, respectively, and represent the third augmented Lagrange multipliers of the i-th hidden layer and the (i-1)-th hidden layer output, respectively, represents the background dictionary of the (i-1)-th hidden layer, represents the foreground dictionary of the (i-1)-th hidden layer;
[0010] Step 2, obtain a historical image dataset, and obtain a training sample set after pre-processing each historical image in the historical image dataset; train the deep unfolding foreground detection network constructed in step 1 using the training sample set to obtain a deep unfolding foreground detection model;
[0011] Step 3, collect a to-be-detected image, pre-process the to-be-detected image, and input the pre-processed to-be-detected image into the deep unfolding foreground detection model obtained in step 2 to obtain a foreground target of the to-be-detected image.
[0012] The L-block consists of 3 3D convolution layers, 2 batch normalization layers, 2 linear rectifier function layers, and 1 addition layer; the input of the first 3D convolution layer and one input of the addition layer are collectively used as the input of the L-block, the output of the first 3D convolution layer is connected to the input of the first batch normalization layer, the output of the first batch normalization layer is connected to the input of the first linear rectifier function layer, the output of the first linear rectifier function layer is connected to the input of the second 3D convolution layer, the output of the second 3D convolution layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second linear rectifier function layer, the output of the second linear rectifier function layer is connected to the input of the third 3D convolution layer, the output of the third 3D convolution layer is connected to the other input of the addition layer, and the output of the addition layer is used as the output of the L-block.
[0013] The S-block consists of 3 3D convolution layers, 2 batch normalization layers, 2 linear rectifier function layers, 1 soft threshold function layer, and 1 subtraction layer; the input of the first 3D convolution layer and one input of the subtraction layer are collectively used as the input of the S-block, the output of the first 3D convolution layer is connected to the input of the first batch normalization layer, the output of the first batch normalization layer is connected to the input of the first linear rectifier function layer, the output of the first linear rectifier function layer is connected to the input of the second 3D convolution layer, the output of the second 3D convolution layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second linear rectifier function layer, the output of the second linear rectifier function layer is connected to the input of the soft threshold function layer, the output of the soft threshold function layer is connected to the input of the third 3D convolution layer, the output of the third 3D convolution layer is connected to the other input of the subtraction layer, and the output of the subtraction layer is used as the output of the S-block.
[0014] The X-block consists of 3 3D convolution layers, 2 batch normalization layers and 2 linear rectifier function layers; the input of the first 3D convolution layer is taken as the input of the X-block, the output of the first 3D convolution layer is connected to the input of the first batch normalization layer, the output of the first batch normalization layer is connected to the input of the first linear rectifier function layer, the output of the first linear rectifier function layer is connected to the input of the second 3D convolution layer, the output of the second 3D convolution layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second linear rectifier function layer, the output of the second linear rectifier function layer is connected to the input of the third 3D convolution layer, and the output of the third 3D convolution layer is taken as the output of the X-block.
[0015] The B-block consists of 3 3D convolution layers, 2 batch normalization layers and 2 linear rectifier function layers; the input of the first 3D convolution layer is taken as the input of the B-block, the output of the first 3D convolution layer is connected to the input of the first batch normalization layer, the output of the first batch normalization layer is connected to the input of the first linear rectifier function layer, the output of the first linear rectifier function layer is connected to the input of the second 3D convolution layer, the output of the second 3D convolution layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second linear rectifier function layer, the output of the second linear rectifier function layer is connected to the input of the third 3D convolution layer, and the output of the third 3D convolution layer is taken as the output of the B-block.
[0016] The M-block consists of 3 3D convolution layers, 2 batch normalization layers and 2 linear rectifier function layers; the input of the first 3D convolution layer is taken as the input of the M-block, the output of the first 3D convolution layer is connected to the input of the first batch normalization layer, the output of the first batch normalization layer is connected to the input of the first linear rectifier function layer, the output of the first linear rectifier function layer is connected to the input of the second 3D convolution layer, the output of the second 3D convolution layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second linear rectifier function layer, the output of the second linear rectifier function layer is connected to the input of the third 3D convolution layer, and the output of the third 3D convolution layer is taken as the output of the M-block.
[0017] The Y1-block consists of 3 3D convolution layers; the input of the first 3D convolution layer is taken as the input of the Y1-block, the output of the first 3D convolution layer is connected to the input of the second 3D convolution layer, the output of the second 3D convolution layer is connected to the input of the third 3D convolution layer, and the output of the third 3D convolution layer is taken as the output of the Y1-block.
[0018] The Y2-block consists of 3 3D convolution layers; the input of the first 3D convolution layer is taken as the input of the Y2-block, the output of the first 3D convolution layer is connected to the input of the second 3D convolution layer, the output of the second 3D convolution layer is connected to the input of the third 3D convolution layer, and the output of the third 3D convolution layer is taken as the output of the Y2-block.
[0019] The Y3-block consists of 3 3D convolution layers; the input of the first 3D convolution layer is taken as the input of the Y3-block, the output of the first 3D convolution layer is connected to the input of the second 3D convolution layer, the output of the second 3D convolution layer is connected to the input of the third 3D convolution layer, and the output of the third 3D convolution layer is taken as the output of the Y3-block.
[0020] In the above step 2, the loss function L of the deep unfolding foreground detection network is trained total is:
[0021]
[0022] wherein, and represent the actual sparse foreground and the predicted sparse foreground respectively, and represent the actual auxiliary foreground and the predicted auxiliary foreground respectively, represents the predicted auxiliary background; BCE(·) is a binary cross-entropy loss function, Tversky(·) is a segmentation loss function, MSE(·) is a mean square error loss function, and Frobenius(·) is a non-coherent item loss function; and in are weight parameters corresponding to the loss functions; represents Hadamard product.
[0023] Compared with the prior art, the present application has the following characteristics:
[0024] 1. Model interpretability: mapping mathematical optimization problems to neural network layers to preserve theoretical interpretability;
[0025] 2. Dynamic scene adaptability: modeling low-rank background and sparse foreground through double dictionary learning to enhance robustness in complex scenes;
[0026] 3. Computational efficiency: using a deep unfolding network to convert iterative optimization to a single forward propagation, reducing computational cost. BRIEF DESCRIPTION OF DRAWINGS
[0027] Figure 1 is a whole architecture diagram of the deep unfolding foreground detection network.
[0028] Figure 2are sub-module structure diagrams, (a) L-block, (b) S-block, (c) X-block, (d) B-block, (e) M-block, (f) Y1-block, (g) Y2-block, (h) Y3-block.
[0029] Figure 3 are result diagrams of experiments of the algorithm of the present application and existing algorithms on a CDnet2014 dataset.
[0030] Figure 4 are IOU evaluation index diagrams of experiments of the algorithm of the present application and existing algorithms on a CDnet2014 dataset.
[0031] Figure 5 is an ablation experiment result diagram. DETAILED DESCRIPTION
[0032] In order to make the objectives, technical solutions and advantages of the present application clearer and more comprehensible, the present application is further described in detail below with reference to specific examples and the accompanying drawings.
[0033] The present application fuses the advantages of model driving and data driving, and proposes a deep unfolding foreground detection method fusing dictionary learning and incoherent terms, which includes the following steps:
[0034] Step 1, constructing a deep unfolding foreground detection network, as shown in Figure 1
[0035] The deep unfolding foreground detection network (DUM-Net) is composed of an input layer, K serial hidden layers and an output layer; the input of the input layer is taken as the input of the deep unfolding foreground detection network, the output of the input layer is connected to the input of the first hidden layer, the output of the last hidden layer is connected to the input of the output layer, and the output of the output layer is taken as the output of the deep unfolding foreground detection network. In the implementation of low-rank sparse decomposition by the traditional ADMM, 8 types of variables must be maintained and updated at the same time, while the present application adopts a deep unfolding network structure containing K hidden layers, each hidden layer corresponds to one ADMM iteration, and the K hidden layers correspond to K iterations.
[0036] In order to express the foreground detection objective function in a separable form, the final mathematical model of the present application is:
[0037]
[0038] wherein, and represent background, and represent foreground, and represent foreground dictionary and background dictionary, respectively, represents Hadamard product, ||·||1 * represents tensor nuclear norm, ||·||2 2,1 represents l 2,1 norm, ||·||1 represents l1 norm, ||·||2 F represents Frobenius norm.
[0039] In order to solve the problem of lack of spatiotemporal continuity of the traditional model, the third term In order to further distinguish the structural features of the foreground and the background, a non-coherent term is introduced The above mathematical model is finally changed into:
[0040]
[0041] Wherein, Y1, Y2, Y3 are augmented Lagrange multipliers, and μ1, μ2, μ3 are penalty operators.
[0042] The above mathematical model can be decomposed into 8 sub-problems by using ADMM algorithm. Each sub-problem corresponds to a sub-module, that is, L-block, S-block, X-block, B-block, M-block, Y1-block, Y2-block and Y3-block, and the structures of the sub-modules are as shown in Figure 2
[0043] The L-block is composed of 3 3D convolution layers, 2 batch normalization layers, 2 linear rectifier function layers and 1 addition layer. The input of the first 3D convolution layer and one input of the addition layer are collectively taken as the input of the L-block, and the input The output of the first 3D convolution layer is connected to the input of the first batch normalization layer, the output of the first batch normalization layer is connected to the input of the first linear rectifier function layer, the output of the first linear rectifier function layer is connected to the input of the second 3D convolution layer, the output of the second 3D convolution layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second linear rectifier function layer, the output of the second linear rectifier function layer is connected to the input of the third 3D convolution layer, and the output of the third 3D convolution layer is connected to the other input of the addition layer; the output of the addition layer is taken as the output of the L-block, and the output Input data Firstly, feature fusion is performed by 3x3 convolution kernel (stride 1, channel number 1->32), and then batch normalization processing is performed on the 32-channel features and a ReLU activation function is applied; then 3x3 convolution is performed again in the 32-channel feature space (keeping the channel number 32), and batch normalization and ReLU activation processing are also performed; finally, the feature dimension is restored to the original dimension by 3x3 convolution kernel (stride 1, channel number 32->1), and the updated
[0044] The S-block is composed of 3 three-dimensional convolution layers, 2 batch normalization layers, 2 linear rectifier function layers, 1 soft threshold function layer and 1 subtraction layer. The input of the first three-dimensional convolution layer and one input of the subtraction layer are jointly taken as the input of the S-block, and the input of the first three-dimensional convolution layer is The output of the first three-dimensional convolution layer is connected to the input of the first batch normalization layer, the output of the first batch normalization layer is connected to the input of the first linear rectifier function layer, the output of the first linear rectifier function layer is connected to the input of the second three-dimensional convolution layer, the output of the second three-dimensional convolution layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second linear rectifier function layer, the output of the second linear rectifier function layer is connected to the input of the soft threshold function layer, the output of the soft threshold function layer is connected to the input of the third three-dimensional convolution layer, and the output of the third three-dimensional convolution layer is connected to the other input of the subtraction layer; the output of the subtraction layer is taken as the output of the S-block, and the output Input data Firstly, feature extraction is performed by 3x3 convolution kernel (stride 1, channel number 1->32), and then batch normalization is performed on the 32-channel features and a ReLU activation is applied; then 3x3 convolution is performed again in the 32-channel feature space (keeping the channel number 32), and 3x3 grouped convolution (group number=2, channel number kept 32->32) is used to reduce the parameter quantity and enhance the local feature independence, and batch normalization and ReLU activation processing are also performed; then feature-level adaptive sparsification is realized by a learnable soft thresholding operation, in which the threshold parameter is automatically optimized by back propagation; finally, the feature map is restored to the original dimension by 3x3 convolution kernel (stride 1, channel number 32->1), and the updated
[0045] The X-block is composed of 3 three-dimensional convolution layers, 2 batch normalization layers and 2 linear rectifier function layers. The input of the first three-dimensional convolution layer is taken as the input of the X-block, the output of the first three-dimensional convolution layer is connected to the input of the first batch normalization layer, and the input The output of the first batch normalization layer is connected to the input of the first linear rectifier function layer, the output of the first linear rectifier function layer is connected to the input of the second 3D convolution layer, the output of the second 3D convolution layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second linear rectifier function layer, and the output of the second linear rectifier function layer is connected to the input of the third 3D convolution layer; the output of the third 3D convolution layer is taken as the output of the X-block, and the output The input data is first Feature fusion is performed through a 3x3 convolution kernel (step 1, channel number 1->32), and then batch normalization processing is performed on the 32-channel features and ReLU activation is applied; then 3x3 convolution is performed again in the 32-channel feature space (keeping the channel number 32), and batch normalization and ReLU activation processing are also performed; finally, the feature dimension is restored through a 3x3 convolution kernel (step 1, channel number 32->1), and the updated
[0046] The B-block is composed of 3 3D convolution layers, 2 batch normalization layers and 2 linear rectifier function layers. The input of the first 3D convolution layer is taken as the input of the B-block, and the input The output of the first 3D convolution layer is connected to the input of the first batch normalization layer, the output of the first batch normalization layer is connected to the input of the first linear rectifier function layer, the output of the first linear rectifier function layer is connected to the input of the second 3D convolution layer, the output of the second 3D convolution layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second linear rectifier function layer, and the output of the second linear rectifier function layer is connected to the input of the third 3D convolution layer; the output of the third 3D convolution layer is taken as the output of the B-block, and the output The input data First, a 3x3 convolution kernel (step 1) is used to expand the channel number from 1 to 32, and then batch normalization processing and ReLU activation are performed; then, 3x3 convolution (step 1) is applied again on the 32-channel feature map, and batch normalization and ReLU activation are also performed; finally, the channel number is reduced from 32 to 1 through a third 3x3 convolution (step 1), and the original data dimension is restored, and the updated
[0047] The M-block is composed of 3 3D convolution layers, 2 batch normalization layers and 2 linear rectifier function layers. The input of the first 3D convolution layer is taken as the input of the M-block, and the input Y1 i-1 、 The output of the first 3D convolutional layer is connected to the input of the first batch normalization layer, the output of the first batch normalization layer is connected to the input of the first linear rectifier function layer, the output of the first linear rectifier function layer is connected to the input of the second 3D convolutional layer, the output of the second 3D convolutional layer is connected to the input of the second batch normalization layer, the output of the second batch normalization layer is connected to the input of the second linear rectifier function layer, and the output of the second linear rectifier function layer is connected to the input of the third 3D convolutional layer; the output of the third 3D convolutional layer is taken as the output of the M-block, and the output The input data is first Y1 i-1 , The features are extracted through a 3x3 convolution kernel (step 1, channel number 1→32), and then the 32-channel features are standardized through batch normalization and ReLU activation; then the second 3x3 convolution is performed in the 32-channel feature space (the channel number is kept at 32), and batch normalization and ReLU activation are also performed to enhance the feature expression ability; finally, the feature map is restored to the original dimension through a 3x3 convolution kernel (step 1, channel number 32→1), and the updated
[0048] The Y1-block is composed of three 3D convolutional layers. The input of the first 3D convolutional layer is taken as the input of the Y1-block, and the input The output of the first 3D convolutional layer is connected to the input of the second 3D convolutional layer, and the output of the second 3D convolutional layer is connected to the input of the third 3D convolutional layer; the output of the third 3D convolutional layer is taken as the output of the Y1-block, and the output Y1 is output i . First, the input data is subjected to a 3x3 convolution kernel with a step of 1 and a channel number of 1→32; second, the convolution channel number is kept at 32, the step is 1, and the convolution kernel is 3x3; finally, a 3x3 convolution kernel is performed with a step of 1 and a channel number of 32→1 to restore the original data dimension, and the updated Y1 is output i .
[0049] The Y2-block is composed of three 3D convolutional layers. The input of the first 3D convolutional layer is taken as the input of the Y2-block, and the input The output of the first 3D convolutional layer is connected to the input of the second 3D convolutional layer, and the output of the second 3D convolutional layer is connected to the input of the third 3D convolutional layer; the output of the third 3D convolutional layer is taken as the output of the Y2-block, and the output First, the input data convolution with 3x3 kernel, stride 1, channel number 1→32; secondly, keep the convolution channel number 32 unchanged, convolution with 3x3 kernel, stride 1; finally, convolution with 3x3 kernel, stride 1, channel number 32→1, restore to the original data dimension, output the updated
[0050] The Y3-block is composed of 3 3D convolution layers. The input of the first 3D convolution layer is taken as the input of the Y3-block, and the input of the first 3D convolution layer is The output of the first 3D convolution layer is connected to the input of the second 3D convolution layer, the output of the second 3D convolution layer is connected to the input of the third 3D convolution layer; the output of the third 3D convolution layer is taken as the output of the Y3-block, and the output of the third 3D convolution layer is First, the input data convolution with 3x3 kernel, stride 1, channel number 1→32; secondly, keep the convolution channel number 32 unchanged, convolution with 3x3 kernel, stride 1; finally, convolution with 3x3 kernel, stride 1, channel number 32→1, restore to the original data dimension, output the updated
[0051] wherein, and represent the auxiliary background variables of the output of the i-th hidden layer and the (i-1)-th hidden layer, respectively, and represent the auxiliary foreground variables of the output of the i-th hidden layer and the (i-1)-th hidden layer, respectively, and represent the reconstruction variables of the output of the i-th hidden layer and the (i-1)-th hidden layer, respectively, and represent the low-rank background variables of the output of the i-th hidden layer and the (i-1)-th hidden layer, respectively, and represent the sparse foreground variables of the output of the i-th hidden layer and the (i-1)-th hidden layer, respectively, Y1 i and Y1 i-1 represent the first augmented Lagrange multipliers of the output of the i-th hidden layer and the (i-1)-th hidden layer, respectively, and represent the second augmented Lagrange multipliers of the output of the i-th hidden layer and the (i-1)-th hidden layer, respectively, and represent the third augmented Lagrange multipliers of the output of the i-th hidden layer and the (i-1)-th hidden layer, respectively, represents the background dictionary of the (i-1)-th hidden layer, represents the foreground dictionary of the (i-1)-th hidden layer.
[0052] Step 2: Obtain the historical image dataset and preprocess each historical image in the historical image dataset to obtain the training sample set; use the training sample set to train the depth unfolded foreground detection network constructed in Step 1 to obtain the depth unfolded foreground detection model.
[0053] In this embodiment, the historical image dataset is obtained by segmenting historical videos into preset frames (e.g., 50 frames). When preprocessing the historical images, they are first grayscaled and then normalized to [0,1]. Then, bilinear interpolation is used to limit the frame width to a preset number of pixels (e.g., 128 pixels) to balance computational efficiency and resolution.
[0054] The process of training the depth unfolded foreground detection network using the training sample set is as follows:
[0055] 1) Construct a learnable dual dictionary system ( and Generate dictionary atoms through random initialization, where the dictionary is a set of trainable parameters. Initialize the penalty operator μ (e.g., μ1 = μ2 = μ3 = 0.01) and set the total number of training iterations (e.g., 90 iterations).
[0056] 2) Each training sample in the training sample set is fed into the deep unfolded foreground detection network. Each sub-module of the deep unfolded foreground detection network adopts a modular decomposition method and performs calculations on different sub-problems for the training samples.
[0057] 3) To optimize the foreground detection and background modeling performance of the model, the following semi-supervised composite loss function is used to train the parameters of the deep unfolded foreground detection network;
[0058]
[0059] in, and These represent the actual sparse prospect and the predicted sparse prospect, respectively. and These represent the actual auxiliary prospects and the predicted auxiliary prospects, respectively. The auxiliary background represents the prediction; BCE(·) is the binary cross-entropy loss function used to optimize... Tversky(·) is a segmentation loss function used to better control precision and recall in segmentation tasks, especially in cases of class imbalance, for optimization. MSE(·) is the mean squared error loss function, used for optimization. Through the mask Error is calculated only in the background region. Frobenius(·) is the incoherent loss function used to measure the difference between the foreground mask and the low-rank background. α1, α2, and λ inare weight parameters for balancing the effects of different loss terms (α1, α2 are set to 0.5 in the experiment, and λ in is set to 0.01).
[0060] 4) When the deep unfolding foreground detection network reaches the set total number of training or converges, the iteration is terminated, and a deep unfolding foreground detection model is obtained, which can be used for foreground target detection on images.
[0061] Step 3, collect the image to be detected, and after pre-processing the image to be detected, send it to the deep unfolding foreground detection model obtained in step 2 to obtain the foreground target of the image to be detected, and realize foreground target detection on the image to be detected.
[0062] The application synchronously extracts the low-rank feature of the background and the sparse structure feature of the foreground by constructing a double-channel dictionary learning mechanism, combines a mask generation module, converts an additive decomposition model of a traditional RPCA into a mask product form, realizes collaborative optimization of a foreground mask M and a low-rank background S through end-to-end learning, and effectively promotes foreground-background separation in a deep unfolding network by introducing a constraint layer of a non-coherent term.
[0063] The effect of the application is verified through a simulation experiment below.
[0064] The experiment is performed on a Windows 10 operating system through a Pycharm framework, and is composed of a 64-bit operating system, 8GB memory and an Intel Core i5-8279U processor.
[0065] Simulation content:
[0066] (1) The application algorithm and the existing algorithm are experimented on a CDnet2014 dataset, and the results are shown in Table 1. Figure 3 The first row to the tenth column respectively represent the original framework, the real foreground, Frmc, Horpca, Ialm, L112, L1-2, Refrpca, Roman rand the present application. The first to tenth rows represent "Highway", "Pedestrians", "PETS2006", "Traffic", "Boulevard", "Sofa", "StreetLight", "WinterDriveway", "Corridor" and "Library" monitoring videos, respectively. The first three rows correspond to three groups of video samples of the baseline scene, the fourth to fifth rows present the contrast of camera jitter scenes, the sixth to eighth rows focus on three groups of test sequences of intermittent target motion scenes, and the ninth to tenth rows select two groups of typical data of thermal imaging scenes.
[0067] ①In the baseline scene experiment comparison, FrmC, Horpca, Ialm, L112, L1-2 and Refrpca all have local hole phenomenon in the target area in the three test video sequences, among which the L112 algorithm is particularly obvious in the detection of "Pedestrians", and the output result not only has a large area of internal hole, but also has persistent ghost residue. It is worth noting that although the Roman-r algorithm has improved in target contour extraction, it still has motion target ghost phenomenon in dynamic scenes. In comparison, the model in this paper can still maintain the complete structure and clear boundary features of the foreground target under complex motion patterns and light changes.
[0068] ②In the comparison test of camera jitter scenes, FrmC, Horpca, Ialm, L112 and L1-2 algorithms generally have the problems of artifact residue and target internal hole, among which the Ialm algorithm has a large area of structural hole phenomenon in the detection of "Traffic". Although Refrpca and Roman_r have improved in target positioning, they still have foreground boundary blurring and local area fracture phenomenon. In comparison, the model in this paper effectively overcomes the interference caused by camera jitter, and at the same time maintains the clear and sharp edge features of the target contour, significantly improving the integrity of the foreground region segmentation.
[0069] ③In the comparison experiment of intermittent target motion scenes, FrmC, HoRPCA, Ialm, L112, L1-2 and Refrpca all have hole phenomenon in the three video sequences. It is particularly worth noting that the L112 algorithm has obvious motion ghost interference in the foreground detection of "WinterDriveway", in addition to the hole phenomenon, which leads to a significant decrease in detection accuracy. Although the Roman-r algorithm is improved compared to other methods, it still has partial area missing phenomenon in the detection result. Through the above comparison, the method in this paper realizes more complete contour capture in the target non-continuous motion scene, and the foreground detection result is more accurate.
[0070] (4) In the contrast experiment of thermal imaging scenes, the foreground region holes exist in the "Library" of FrmC, HorPca, IAlm, L112 and L1-2 algorithms, and the dense ghost interference is caused by the thermal radiation feature misjudgment of FrmC and L112 algorithms; in the detection of "Corridor", RefRPca, Roman_r and the algorithm in the paper partially capture the target subject, but there are still local hole problems, but in the comparison of objective evaluation indexes, we have a higher F-measure value. In comparison, the method of the paper realizes accurate detection in the thermal imaging scene, and the generated foreground mask completely covers the target area, which shows that the model has strong adaptability to thermal imaging data.
[0071] In summary, the foreground detection of the application is complete, the overall effect is good, and the advantages in target integrity and edge sharpness are obvious.
[0072] (2) The algorithm of the application and the existing algorithm are experimented on the CDnet2014 dataset, and the IOU evaluation index is as shown in Figure 4 . The advantages of the application are better shown by the objective indexes, and it is proved that the application is better than the existing algorithm in most scenes.
[0073] (3) In order to further verify the effectiveness of the application in the foreground detection task, a series of ablation experiments are designed to in-depth analyze the contribution of the key components in the model. First, the spatiotemporal continuity constraint module in the model is removed, and the remaining part is retained for experiment, so as to evaluate the influence of spatiotemporal continuity modeling on the precision of foreground detection. Secondly, the non-coherent term optimization part in the model is further removed to study its role in enhancing the separation of foreground and background. Finally, the complete model and the above two ablation variants are experimented on the same test set, and the results are as shown in Figure 5 . The first row to the tenth column respectively represent "Highway", "Pedestrians", "PETS2006", "Traffic", "Boulevard", "Sofa", "StreetLight", "WinterDriveway", "Corridor" and "Library" monitoring videos. The first row to the fifth row respectively represent the original video frame, the true foreground, the foreground detection graph obtained by the application, the foreground detection graph obtained by removing the spatiotemporal continuity constraint term in the step 1 mathematical model, and the foreground detection graph obtained by removing the non-coherent term . .
[0074] The experimental results show that when the spatiotemporal continuity constraint is removed, the foreground segmentation accuracy of the model in the dynamic scene decreases significantly, especially in the object edge area, obvious artifacts appear, and on the "Boulevard" data set, the foreground target is not detected, resulting in a decrease in the detection accuracy of the model; and after removing the incoherent term, the suppression ability of the model to the complex background is obviously weakened, resulting in an increase in the false detection rate. The ablation experiment results fully prove that each component in the application has an irreplaceable important role, the spatiotemporal continuity constraint effectively improves the temporal stability of the foreground mask, and the incoherent term significantly enhances the foreground and background separation function of the model. Through ablation experiment analysis, we verify the rationality of the design of the application.
[0075] It should be noted that although the above embodiments of the application are illustrative, this is not a limitation of the application, therefore the application is not limited to the above specific embodiments. Any other embodiments obtained by those skilled in the art under the inspiration of the application without departing from the principles of the application are considered to be within the protection of the application.
Claims
1. A method for deep unfolding foreground detection that integrates dictionary learning and irrelevant terms, characterized by: The steps include the following: Step 1: Construct a deep unfolded foreground detection network; The deep unfolded foreground detection network consists of an input layer, at least one hidden layer, and an output layer. All hidden layers are connected sequentially. The input of the input layer serves as the input of the deep unfolded foreground detection network. The output of the input layer is connected to the input of the first hidden layer. The output of the last hidden layer is connected to the input of the output layer. The output of the output layer serves as the output of the deep unfolded foreground detection network. Each hidden layer consists of 8 sub-blocks: L-block, S-block, X-block, B-block, M-block, Y1-block, Y2-block, and Y3-block. Input is sent to the L-block, and output is sent out from the L-block. Input is sent to the S-block, and output is sent out from the S-block. Input is sent to the X-block, and output is sent out from the X-block. Input is sent to the B-block, and output is sent out from the B-block. Input is sent to the M-block, and output is sent out from the M-block. The input is fed into Y1-block, and the output of Y1-block is sent out to Y1. i ; Input is sent to the Y2-block, and output is sent out from the Y2-block. Input is sent to the Y3-block, and output is sent out from the Y3-block. in, and These represent the auxiliary background variables output by the i-th hidden layer and the (i-1)-th hidden layer, respectively. and Let represent the auxiliary foreground variables output by the i-th hidden layer and the (i-1)-th hidden layer, respectively. and These represent the reconstruction variables output by the i-th hidden layer and the (i-1)-th hidden layer, respectively. and Let represent the low-rank background variables output by the i-th hidden layer and the (i-1)-th hidden layer, respectively. and Y1 represents the sparse foreground variable output by the i-th hidden layer and the (i-1)-th hidden layer, respectively. i and Y1 i-1 These represent the first augmented Lagrange multipliers of the outputs of the i-th and (i-1)-th hidden layers, respectively. and These represent the second augmented Lagrange multipliers of the outputs of the i-th and (i-1)-th hidden layers, respectively. and The third augmented Lagrange multipliers represent the outputs of the i-th and (i-1)-th hidden layers, respectively. This represents the background dictionary of the (i-1)th hidden layer. Represents the foreground dictionary of the (i-1)th hidden layer; Step 2: Obtain the historical image dataset and preprocess each historical image in the historical image dataset to obtain the training sample set; use the training sample set to train the depth unfolded foreground detection network constructed in Step 1 to obtain the depth unfolded foreground detection model. Step 3: Acquire the image to be detected, and after preprocessing the image, send it into the depth unfolding foreground detection model obtained in Step 2 to obtain the foreground target of the image to be detected.
2. The method for deep unfolding foreground detection that integrates dictionary learning and incoherent terms as described in claim 1, is characterized in that, The L-block consists of three 3D convolutional layers, two batch normalization layers, two linear rectified function layers, and one additive layer. The input of the first 3D convolutional layer and one input of the addition layer are used as the input of the L-block. The output of the first 3D convolutional layer is connected to the input of the first batch normalization layer. The output of the first batch normalization layer is connected to the input of the first linear rectified function layer. The output of the first linear rectified function layer is connected to the input of the second 3D convolutional layer. The output of the second 3D convolutional layer is connected to the input of the second batch normalization layer. The output of the second batch normalization layer is connected to the input of the second linear rectified function layer. The output of the second linear rectified function layer is connected to the input of the third 3D convolutional layer. The output of the third 3D convolutional layer is connected to another input of the addition layer. The output of the addition layer is used as the output of the L-block.
3. The method for deep unfolding foreground detection that integrates dictionary learning and incoherent terms as described in claim 1, is characterized in that, The S-block consists of three 3D convolutional layers, two batch normalization layers, two linear rectified function layers, one soft thresholding function layer, and one subtraction layer. The input of the first 3D convolutional layer and one input of the subtraction layer are used as the input of the S-block. The output of the first 3D convolutional layer is connected to the input of the first batch normalization layer. The output of the first batch normalization layer is connected to the input of the first linear rectified function layer. The output of the first linear rectified function layer is connected to the input of the second 3D convolutional layer. The output of the second 3D convolutional layer is connected to the input of the second batch normalization layer. The output of the second batch normalization layer is connected to the input of the second linear rectified function layer. The output of the second linear rectified function layer is connected to the input of the soft thresholding function layer. The output of the soft thresholding function layer is connected to the input of the third 3D convolutional layer. The output of the third 3D convolutional layer is connected to another input of the subtraction layer. The output of the subtraction layer is used as the output of the S-block.
4. The method for deep unfolding foreground detection that integrates dictionary learning and incoherent terms as described in claim 1, characterized in that, X-block consists of 3 3D convolutional layers, 2 batch normalization layers, and 2 linear rectified function layers; The input of the first 3D convolutional layer is used as the input of the X-block. The output of the first 3D convolutional layer is connected to the input of the first batch normalization layer. The output of the first batch normalization layer is connected to the input of the first linear rectified function layer. The output of the first linear rectified function layer is connected to the input of the second 3D convolutional layer. The output of the second 3D convolutional layer is connected to the input of the second batch normalization layer. The output of the second batch normalization layer is connected to the input of the second linear rectified function layer. The output of the second linear rectified function layer is connected to the input of the third 3D convolutional layer. The output of the third 3D convolutional layer is used as the output of the X-block.
5. The method for deep unfolding foreground detection that integrates dictionary learning and incoherent terms as described in claim 1, characterized in that, The B-block consists of three 3D convolutional layers, two batch normalization layers, and two linear rectified function layers. The input of the first 3D convolutional layer is used as the input of the B-block. The output of the first 3D convolutional layer is connected to the input of the first batch normalization layer. The output of the first batch normalization layer is connected to the input of the first linear rectified function layer. The output of the first linear rectified function layer is connected to the input of the second 3D convolutional layer. The output of the second 3D convolutional layer is connected to the input of the second batch normalization layer. The output of the second batch normalization layer is connected to the input of the second linear rectified function layer. The output of the second linear rectified function layer is connected to the input of the third 3D convolutional layer. The output of the third 3D convolutional layer is used as the output of the B-block.
6. The method for deep unfolding foreground detection that integrates dictionary learning and incoherent terms as described in claim 1, characterized in that, The M-block consists of three 3D convolutional layers, two batch normalization layers, and two linear rectified function layers. The input of the first 3D convolutional layer is used as the input of the M-block. The output of the first 3D convolutional layer is connected to the input of the first batch normalization layer. The output of the first batch normalization layer is connected to the input of the first linear rectified function layer. The output of the first linear rectified function layer is connected to the input of the second 3D convolutional layer. The output of the second 3D convolutional layer is connected to the input of the second batch normalization layer. The output of the second batch normalization layer is connected to the input of the second linear rectified function layer. The output of the second linear rectified function layer is connected to the input of the third 3D convolutional layer. The output of the third 3D convolutional layer is used as the output of the M-block.
7. The method for deep unfolding foreground detection that integrates dictionary learning and incoherent terms as described in claim 1, characterized in that, Y1-block consists of three 3D convolutional layers; The input of the first 3D convolutional layer is used as the input of the Y1-block. The output of the first 3D convolutional layer is connected to the input of the second 3D convolutional layer. The output of the second 3D convolutional layer is connected to the input of the third 3D convolutional layer. The output of the third 3D convolutional layer is used as the output of the Y1-block.
8. The method for deep unfolding foreground detection that integrates dictionary learning and incoherent terms as described in claim 1, characterized in that, Y2-block consists of three 3D convolutional layers; The input of the first 3D convolutional layer is used as the input of the Y2-block. The output of the first 3D convolutional layer is connected to the input of the second 3D convolutional layer. The output of the second 3D convolutional layer is connected to the input of the third 3D convolutional layer. The output of the third 3D convolutional layer is used as the output of the Y2-block.
9. The method for deep unfolding foreground detection that integrates dictionary learning and incoherent terms as described in claim 1, characterized in that, The Y3-block consists of three 3D convolutional layers; The input of the first 3D convolutional layer is used as the input of the Y3-block. The output of the first 3D convolutional layer is connected to the input of the second 3D convolutional layer. The output of the second 3D convolutional layer is connected to the input of the third 3D convolutional layer. The output of the third 3D convolutional layer is used as the output of the Y3-block.
10. The method for deep unfolding foreground detection that integrates dictionary learning and incoherent terms as described in claim 1, characterized in that, In step 2, the loss function L is trained for the deep unfolded foreground detection network. total for: in, and These represent the actual sparse prospect and the predicted sparse prospect, respectively. and These represent the actual auxiliary prospects and the predicted auxiliary prospects, respectively. The auxiliary background for prediction is represented by BCE(·), which is the binary cross-entropy loss function, Tversky(·), which is the segmentation loss function, MSE(·), which is the mean squared error loss function, and Frobenius(·), which is the incoherence term loss function; α1, α2, and λ in These are the weight parameters of the corresponding loss function; Represents Hadema.