Diffusion process-based multivariate target detection method
By introducing a detection method based on diffusion process in the multivariate variable prediction model, combining a lightweight detection network and a diffusion probability network, the problem of difficulty in predicting the movement position of multiple targets in the existing technology is solved, and more flexible and efficient multivariate target detection is achieved.
Patent Information
- Application Number
- CN202411989771.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-23
AI Technical Summary
The existing multivariate prediction model is difficult to effectively predict the movement positions of different targets, and the model generalization ability is limited, making it difficult to adapt to different scenarios.
The multi-object detection method based on the diffusion process is adopted, and by building a lightweight detection network and diffusion probability network, combining the timing fusion module, the timing information and feature differences are retained, the model's situation capture ability of different frames is improved, and the multi-object position data is learned through the diffusion probability network to sample and restore the position information of the real data.
It improves the flexibility of multivariate probability distribution prediction, and can continue to detect targets in the absence of discontinuous or detection loss, which improves the probability and target detection rate of weak target detection.
Smart Images

Figure CN120032242A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multivariate prediction, and in particular relates to a multivariate target detection method based on diffusion process. Background Art
[0002] Multi-objective prediction refers to predicting the development trends, behaviors or states of multiple different objectives or variables in a complex system at the same time. Its research results have been widely used in weather forecasting, finance, traffic supervision and other fields. In order to better and more comprehensively understand and predict the relationship between multiple factors in a complex system, the multivariate probability distribution is modeled to approximate the real prediction model. A common method is to preset an easy-to-handle distribution or a low-rank approximation of the real data distribution. This type of method often involves subjective judgment, which limits the generalization ability of the model and makes it difficult to effectively adapt to different scenarios. Another way is to use autoregressive or normalized generation methods based on deep learning. Although the likelihood function of this type of method cannot be solved, it can use the variational lower bound to fit the optimal solution to predict multivariate target variables.
[0003] Univariate forecasting focuses on the prediction of a single target, usually with low data requirements and relatively simple models. Multivariate forecasting, on the other hand, involves the joint prediction of multiple related targets, requiring more data and complex models, and is suitable for situations where complex decisions need to be made based on multiple factors.
[0004] There are two main types of multivariate prediction models at present. One is to establish a multivariate model based on point prediction. This method is used in situations where deterministic results are required and aims to provide a most likely specific numerical prediction for each target variable. The other is to establish a multivariate model based on probability, providing a probability distribution for each target variable, which is suitable for scenarios related to uncertainty and risk.
[0005] It can be seen that the existing multi-target prediction mainly predicts multiple variables by building a single model. This method is difficult to effectively predict the movement positions of different targets. Summary of the invention
[0006] The technology of the present invention solves the problem: overcomes the shortcomings of the prior art, provides a multi-target detection method based on a diffusion process, designs a lightweight detection network, is helpful for the detection of weak multi-targets, and combines the multi-target positions and feature information through a time series fusion module, retains the time series information and feature differences, and improves the model's ability to capture the situation of different frames. Finally, by constructing a diffusion probability network based on learning the distribution obeyed by the multi-target position data, and by sampling the distribution, the position information of the real data is restored.
[0007] In order to solve the above technical problems, the present invention discloses a multi-target detection method based on a diffusion process, comprising:
[0008] A multi-objective prediction model is constructed;
[0009] Obtain an image to be recognized;
[0010] The image to be identified is used as the input of the multi-target prediction model to obtain the predicted target.
[0011] In the above-mentioned multi-target detection method based on diffusion process, a multi-target prediction model is constructed, including:
[0012] Obtain sample data and preprocess the sample data to obtain training sets and test sets;
[0013] Determine the model architecture of the multi-objective prediction model;
[0014] Based on the training set and the test set, as well as the model architecture of the multi-target prediction model, iterative optimization training of the model is performed to obtain the multi-target prediction model.
[0015] In the above-mentioned multi-target detection method based on diffusion process, sample data is obtained and preprocessed to obtain a training set and a test set, including:
[0016] Obtain annotated aerial detection images as sample data; wherein the aerial detection images include: visible light images and infrared images;
[0017] Use image data enhancement technology to preprocess sample data and expand sample data;
[0018] After normalization and scaling, the expanded sample data is divided into a training set and a test set.
[0019] In the above-mentioned multi-target detection method based on diffusion process, the model architecture of the multi-target prediction model includes: a lightweight detection network, a diffusion probability network and a post-processing module;
[0020] The lightweight detection network uses the U2Net model or the U2Netp model as the backbone network, which uses the training data in the training set as input to perform target feature detection and output a feature matrix;
[0021] The diffusion probability network uses the DDPM model, which uses the feature matrix output by the lightweight detection network as input to perform diffusion probability analysis, obtain the probability of the predicted position corresponding to each pixel, and compare it with the set threshold. The predicted position with a probability greater than the set threshold is output as the predicted target corresponding to the training sample data;
[0022] The post-processing module is used to verify the output results of the diffusion probability network based on the test data in the test set, so as to realize iterative optimization training of the model and obtain the final multi-objective prediction model.
[0023] In the above-mentioned multi-target detection method based on the diffusion process, the training sample data is used as the input of the lightweight detection network; the lightweight detection network uses the U2Net model or the U2Netp model as the backbone network, and designs multi-layer branches to extract multi-level feature information; at the same time, a temporal fusion module is introduced into the backbone network encoding and decoding structure, and the input of the temporal fusion module fuses the feature information of the previous levels. The output of the temporal fusion module and the output of the residual network extract the significant information of the small target through the cross-attention mechanism; after multiple rounds of iterative training, the optimal weight parameters are obtained, and the significant information is processed by contour detection to obtain a fusion feature matrix as the input of the diffusion probability network.
[0024] In the above-mentioned multi-target detection method based on diffusion process, the lightweight detection network includes: an encoding module, a decoding module and a temporal fusion module;
[0025] The encoding module includes: residual block A, residual block B, residual block C, residual block D, residual block E and residual block F; wherein, residual block A includes: unit A1, unit A2 and unit A3; unit A1 includes: convolution layer, BN layer and ReLu layer; unit A2 is a pooling layer; unit A3 is an upsampling layer;
[0026] Decoding module, including: cross attention unit, residual block G, residual block H, residual block I, residual block J and residual block K;
[0027] The temporal fusion module includes: a feature extraction module, an up-down module, a down-up module and a transfer module; wherein the feature extraction module includes: a convolution layer, a BN layer and a ReLu layer, the convolution window size is 1, and the step size is 1; the up-down module includes: an adaptive pooling layer, a convolution layer, a BN layer, a ReLu layer, a convolution layer, a BN layer and a sigmoid layer, the output size of the adaptive pooling layer is (1,1), the convolution window size is 1, and the step size is 1; the down-up module includes: a convolution layer, a BN layer, a ReLu layer, a cross attention mechanism unit and a sigmoid layer, the convolution window size is 1, the step size is 1, and the cross attention mechanism unit input size is 3; the transfer module includes: a convolution layer, a BN layer and a ReLu layer, the convolution window size is 3, the step size is 1, and the padding is 1.
[0028] The present invention has the following advantages:
[0029] (1) The present invention discloses a multivariate target detection method based on a diffusion process. The prediction of multivariate target variables obeys different distributions. The prediction of multivariate target variables is achieved through a diffusion probability model. The model is achieved by estimating its gradient and sampling from the data distribution of the current frame, aiming to improve the flexibility of multivariate probability distribution estimation.
[0030] (2) The present invention discloses a multivariate target detection method based on a diffusion process. The deduction prediction of the diffusion process helps to achieve the continued detection of the target in the case of discontinuous target motion or detection loss. The diffusion model is usually accompanied by uncertainty modeling. The target position prediction is not a definite point, but has a certain uncertainty, wherein the target position is modeled as a probability distribution, representing the uncertainty around the possible position.
[0031] (3) The present invention discloses a multi-target detection method based on a diffusion process, designs a more lightweight detection network, and improves the probability of detecting small and weak targets. The cross-attention mechanism is used to help transmit the complex relationship between the feature information of small and weak targets and improve the target detection rate. The present invention is further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a flow chart of a multi-target detection method based on a diffusion process in an embodiment of the present invention;
[0033] Figure 2 It is a schematic diagram of the relevant contents of temporal feature extraction in a temporal fusion module in an embodiment of the present invention;
[0034] Figure 3 It is a schematic diagram of extracting related contents from differences between adjacent frames in a time series fusion module in an embodiment of the present invention;
[0035] Figure 4 It is a schematic diagram of a lightweight detection network in an embodiment of the present invention. DETAILED DESCRIPTION
[0036] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments disclosed in the present invention will be further described in detail below with reference to the accompanying drawings.
[0037] The present invention discloses a multivariate target detection method based on a diffusion process. It does not need to preset the distribution probability of the model in advance. First, an autoregressive model is established for multi-target variables to perform position prediction, which obeys different distributions, that is, sampling from the multi-target data distribution by using a diffusion probability model. At the same time, by designing a lightweight detection network, a relatively small multivariate target position is given, and the feature information and position information of the multivariate are passed as input to the autoregressive model to predict the multivariate position trajectory.
[0038] Due to the irreversibility of neural networks, the present invention introduces a diffusion probability process to learn the mapping relationship from latent space to sample space, that is, by using the multivariate target position information, feature information and covariate in the previous frame image as the condition for generating the multivariate target position and feature information of the current frame. The position and feature information of the multivariate target are generated by a lightweight weak target detection network, which uses the U2net structure as the backbone network, introduces a cross-attention mechanism to fuse multi-level features and outputs position information and feature information. The above information and covariate are used as input to the time series fusion module to obtain the hidden variables of the time series fusion module. Then the above hidden variables are input into the diffusion probability network, which uses the correlation in the time data and adopts the form of self-supervised training to optimize the diffusion model. Since the previous frame position, feature and covariate information are already included in this multivariate target prediction framework, when modeling the multivariate distribution based on probability, this information is passed to the diffusion probability network, and the distribution fitted by the diffusion probability network is the conditional probability distribution of the position and other information of the current frame based on the position, feature and covariate of the previous frame. The present invention utilizes a diffusion probability network to simulate the probability distribution of multiple targets, rather than assuming that they obey a certain specific distribution. At the same time, a lightweight detection network is designed to output the position and feature information of the target, and the above information and covariance are used as the input of a time series fusion module to enhance the above features. Finally, the diffusion probability network is used to fit the conditional probability of the current frame.
[0039] like Figures 1 to 4 In this embodiment, the multi-target detection method based on diffusion process includes:
[0040] Step 1: Construct a multi-objective prediction model.
[0041] In this embodiment, the construction of the multi-objective prediction model mainly includes three aspects:
[0042] (11) Obtain sample data and preprocess the sample data to obtain a training set and a test set.
[0043] Obtain labeled aerial inspection images as sample data; aerial inspection images include visible light images and infrared images. Use image data enhancement technology (such as Mosaic and other methods) to preprocess and expand the sample data to enhance the generalization performance of the model. After normalizing and scaling (scaling to 640×640) the expanded sample data, divide it into a training set and a test set.
[0044] (12) Determine the model architecture of the multi-target prediction model. The model architecture of the multi-target prediction model includes: lightweight detection network, diffusion probability network and post-processing module. Specifically:
[0045] The lightweight detection network uses the U2Net (Ultra-Deep Supervised Salient Object Detection Network, U2Net, ultra-deep supervised saliency detection network) model or the more lightweight U2Netp model as the backbone network, which is used to detect target features with the training data in the training set as input and output the feature matrix. The training sample data is used as the input of the lightweight detection network; the lightweight detection network uses the U2Net model or the U2Netp model as the backbone network, and designs multi-layer branches to extract multi-level feature information, so as to extract the feature information of targets with extremely small size but obvious; at the same time, the temporal fusion module (Information Combined Assigned) is introduced into the backbone network codec structure. The input of the temporal fusion module integrates the feature information of the previous levels. The output of the temporal fusion module and the output of the residual network extract the salient information of small targets through the cross-attention mechanism; after multiple rounds of iterative training, the optimal weight parameters are obtained, and the salient information is processed by contour detection to obtain the fused feature matrix as the input of the diffusion probability network.
[0046] The lightweight detection network can specifically include: encoding module, decoding module and time series fusion module. The encoding module consists of six residual blocks, including: residual block A, residual block B, residual block C, residual block D, residual block E and residual block F; the residual block is mainly used to extract target information of multiple levels of different feature scales, including target location information and high-level semantic information; each residual block corresponds to multiple operation modules, such as Figure 1In the figure, the residual block A includes: unit A1, unit A2 and unit A3; unit A1 includes: convolution layer, BN layer and ReLu layer; unit A2 is a pooling layer; unit A3 is an upsampling layer. Specifically, residual block A first passes through a layer of unit A1, passes through unit A1 and unit A2 five times, then passes through unit A1 twice, then passes through unit A1 and unit A3 five times, and finally outputs through unit A1. The convolution layer window size is 3, the pooling layer window size is 2, and the step size is 2. The input in unit A1 requires multi-level jump links for the last five times. Similarly, residual block B first passes through a layer of unit B1, passes through unit B1 and unit B2 four times, then passes through unit B1 twice, then passes through unit B1 and unit B3 four times, and finally outputs through unit B1. By analogy, residual block C and residual block D first pass through a layer of unit C / D1, passes through unit C / D1 and unit C / D2 three times and twice respectively, then passes through unit C / D1 twice, then passes through unit C / D1 and unit C / D3 three times and twice respectively, and finally outputs through unit C / D1. Finally, residual block E and residual block F pass through 8 layers of convolution layers, BN layers, and unit E / F1 composed of ReLu, respectively. The decoding module includes: cross-attention unit (the cross-attention unit mainly selects the spatiotemporal cross-attention mechanism, considering the spatial and temporal dimension information, which helps the model to better extract the complex structure of the data), residual block G, residual block H, residual block I, residual block J and residual block K; the function of the residual block is to extract multi-level feature information at a deeper level; residual block G is similar to residual block E, and passes through 8 convolutional layers, BN layers and unit 1 composed of ReLu respectively; the structures of residual block H, residual block I, residual block J and residual block K are similar to residual block D, residual block C, residual block B and residual block A respectively. The temporal fusion module includes: feature extraction module, up-down module, down-up module and transfer module; the feature extraction module includes: convolution layer, BN layer and ReLu layer, the convolution window size is 1, and the step size is 1; the up-down module includes: adaptive pooling layer, convolution layer, BN layer, ReLu layer, convolution layer, BN layer and sigmoid layer, the output size of the adaptive pooling layer is (1,1), the convolution window size is 1, and the step size is 1; the down-up module includes: convolution layer, BN layer, ReLu layer, cross attention mechanism unit and sigmoid layer, the convolution window size is 1, the step size is 1, and the cross attention mechanism unit input size is 3; the transfer module includes: convolution layer, BN layer and ReLu layer, the convolution window size is 3, the step size is 1, and the padding is 1.
[0047] The diffusion probability network uses the DDPM (Diffusion Probabilistic Model, DDPM) model, which uses the feature matrix output by the lightweight detection network as input, performs diffusion probability analysis, obtains the probability of the predicted position corresponding to each pixel, and compares it with the set threshold. The predicted position with a probability greater than the set threshold is output as the predicted target corresponding to the training sample data.
[0048] The post-processing module is used to verify the output results of the diffusion probability network based on the test data in the test set, so as to realize iterative optimization training of the model and obtain the final multi-objective prediction model.
[0049] (13) Based on the training set and the test set, as well as the model architecture of the multi-objective prediction model, iterative optimization training of the model is performed to obtain the multi-objective prediction model.
[0050] Preferably, the specific construction process of the multi-objective prediction model can be as follows:
[0051] The first step is to preprocess the training set and the test set and construct a tensor matrix of the current frame image and the previous frame image divided by 255.
[0052] In the second step, the constructed tensor matrix is used as the input matrix of the lightweight detection network, and the expanded feature map is obtained through the encoding module and enters the decoding module.
[0053] In the third step, the data output by the decoding module passes through the cross attention unit, the residual block G, and the upsampling layer, and finally outputs the target feature information.
[0054] In the fourth step, after the encoding-decoding structure, the target output of the detection network is obtained, but further processing is still required at this time. During processing, the minimum enclosing rectangle is drawn by calling the contour detection method to calculate the position information of the corresponding target.
[0055] In the fifth step, the relevant position information and the covariance of the current frame calculated by unbiased estimation are used as the input of the time series fusion module. The time series fusion module consists of a feature extraction module and a feature difference module, such as Figure 2 and Figure 3 The feature extraction module consists of a pooling layer, a resizing layer, a 3D convolution layer, a resizing layer, and an activation layer. The final output is a time series feature. The representative input includes the covariance matrix and the input image data features, and the feature difference module consists of a spatial pooling layer, a two-dimensional convolution, a resizing layer, a one-dimensional convolution, a resizing layer, a two-dimensional convolution, and an activation layer. The final output is the difference feature After weights α, α∈[0,1] are assigned, the fused feature h t-1 for:
[0056]
[0057] h t-1 Contains the state characteristics of the current target state, where the parameters α, α∈[0,1] are obtained from experience.
[0058] The sixth step is to convert h containing the target state features into t-1 And the image data of the training set obeys As input into the diffusion probability network for training, the following process is repeated until convergence:
[0059] Initialize the noise level n ~ uniform(1,…,N) and the diffusion network parameter ε: N(0,1)
[0060] Then the gradient of the objective function constraint formula is derived as follows:
[0061]
[0062] The output distribution parameter ε of the multivariate target is obtained.
[0063] Step 7: Update the position information. A multi-target prediction model is constructed based on the multi-target prediction of the diffusion process. During training, the feature state h is obtained through the sixth step. t-1 , according to the annealed Langevin dynamics sampling process, the next time step t of the current frame can be obtained, which can be autoregressed to the time series fusion module c together with the covariate t Get the next hidden state h t+1 , repeat multiple times (depending on the actual data training situation, the number of times in the present invention is about 100) until the expected prediction range (between 0 and 1) has been reached. Finally, the multivariate variable prediction position information is obtained.
[0064] In the eighth step, feature extraction is performed on the candidate area to realize multi-target positioning, and the target position is predicted by filtering through Euclidean distance and IOU intersection-over-union operations, thereby realizing multi-target prediction.
[0065] Step 2: Obtain the image to be identified.
[0066] Step 3: Use the image to be identified as the input of the multi-target prediction model to obtain the predicted target.
[0067] Although the present invention has been disclosed as above in the form of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art may make possible changes and modifications to the technical solution of the present invention by using the methods and technical contents disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the protection scope of the technical solution of the present invention.
[0068] The contents not described in detail in the specification of the present invention belong to the common knowledge of the professionals in this field.
Claims
1. A multi-target detection method based on diffusion process, characterized in that: include: A multi-objective prediction model is constructed; Obtain an image to be recognized; The image to be identified is used as the input of the multi-target prediction model to obtain the predicted target.
2. The multi-target detection method based on diffusion process according to claim 1, characterized in that: The multi-objective prediction model is constructed, including: Obtain sample data and preprocess the sample data to obtain training sets and test sets; Determine the model architecture of the multi-objective prediction model; Based on the training set and the test set, as well as the model architecture of the multi-target prediction model, iterative optimization training of the model is performed to obtain the multi-target prediction model.
3. The multi-target detection method based on diffusion process according to claim 2, characterized in that: Obtain sample data and preprocess the sample data to obtain training sets and test sets, including: Obtain annotated aerial detection images as sample data; wherein the aerial detection images include: visible light images and infrared images; Use image data enhancement technology to preprocess sample data and expand sample data; After normalization and scaling, the expanded sample data is divided into a training set and a test set.
4. The multi-target detection method based on diffusion process according to claim 3, characterized in that: The model architecture of the multi-target prediction model includes: lightweight detection network, diffusion probability network and post-processing module; The lightweight detection network uses the U2Net model or the U2Netp model as the backbone network, which uses the training data in the training set as input to perform target feature detection and output a feature matrix; The diffusion probability network uses the DDPM model, which uses the feature matrix output by the lightweight detection network as input to perform diffusion probability analysis, obtain the probability of the predicted position corresponding to each pixel, and compare it with the set threshold. The predicted position with a probability greater than the set threshold is output as the predicted target corresponding to the training sample data; The post-processing module is used to verify the output results of the diffusion probability network based on the test data in the test set, so as to realize iterative optimization training of the model and obtain the final multi-objective prediction model.
5. The multi-target detection method based on diffusion process according to claim 4, characterized in that: The training sample data is used as the input of the lightweight detection network; the lightweight detection network uses the U2Net model or the U2Netp model as the backbone network, and designs multi-layer branches to extract multi-level feature information; at the same time, a temporal fusion module is introduced into the backbone network encoding and decoding structure, and the input of the temporal fusion module fuses the feature information of the previous levels. The output of the temporal fusion module and the output of the residual network extract the salient information of small targets through the cross-attention mechanism; after multiple rounds of iterative training, the optimal weight parameters are obtained, and the salient information is processed through contour detection to obtain a fused feature matrix as the input of the diffusion probability network.
6. The multi-target detection method based on diffusion process according to claim 5, characterized in that: Lightweight detection network, including: encoding module, decoding module and time series fusion module; The encoding module includes: residual block A, residual block B, residual block C, residual block D, residual block E and residual block F; wherein, residual block A includes: unit A1, unit A2 and unit A3; unit A1 includes: convolution layer, BN layer and ReLu layer; unit A2 is a pooling layer; unit A3 is an upsampling layer; Decoding module, including: cross attention unit, residual block G, residual block H, residual block I, residual block J and residual block K; The temporal fusion module includes: a feature extraction module, an up-down module, a down-up module and a transfer module; wherein the feature extraction module includes: a convolution layer, a BN layer and a ReLu layer, the convolution window size is 1, and the step size is 1; the up-down module includes: an adaptive pooling layer, a convolution layer, a BN layer, a ReLu layer, a convolution layer, a BN layer and a sigmoid layer, the output size of the adaptive pooling layer is (1,1), the convolution window size is 1, and the step size is 1; the down-up module includes: a convolution layer, a BN layer, a ReLu layer, a cross attention mechanism unit and a sigmoid layer, the convolution window size is 1, the step size is 1, and the cross attention mechanism unit input size is 3; the transfer module includes: a convolution layer, a BN layer and a ReLu layer, the convolution window size is 3, the step size is 1, and the padding is 1.
Citation Information
Cited By
Multi-temporal remote sensing crop extraction method in combination with MoCo self-supervised learning
CN121545065A