Low-rank correlation multi-source target detection method based on adversarial data enhancement
By adopting a low-rank association method that counter data enhancement in multi-source heterogeneous data fusion, virtual data sets are generated, which solves the problem of insufficient data association relationship utilization and improves the generalization ability and performance of the target detection model.
Patent Information
- Application Number
- CN202510035079.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art fails to fully utilize the correlation relationship within heterogeneous data in multi-source heterogeneous data fusion, and GAN-based image enhancement methods have problems of training instability and hyperparameter sensitivity. Low rank representations may not be able to fully capture all details of the image, resulting in loss of information.
A low-rank correlation multi-source object detection method based on adversarial data enhancement is adopted. By collecting image sets from different domains, dimensionality reduction processing is performed, a unified clustering structure is learned, and a pre-trained generative adversarial network is used to generate virtual data sets to enhance the training data of the object detection model.
It effectively solves the problem of insufficient utilization of data association relationships in multi-source heterogeneous data fusion. By generating high-quality virtual data, the generalization ability and performance of the model are enhanced and the effect of target detection is improved.
Smart Images

Figure CN120107633A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-source target detection method, in particular to a low-rank associated multi-source target detection method based on adversarial data enhancement. Background Art
[0002] This section merely provides background information related to the present disclosure and is not necessarily prior art.
[0003] Typical scenarios for multi-source heterogeneous data fusion include situations such as emotion recognition of personnel based on multimodal perception means and target identity recognition based on multimodal battlefield perception data. Usually, the multimodal data is used for feature extraction, and then the correlation calculation at the heterogeneous modal feature level is performed, and finally the result is obtained through decision-level fusion. The fusion of multiple modal data is only performed at the decision-making level, and the internal correlation of heterogeneous data is not well utilized. Therefore, using artificial intelligence algorithms such as generative adversarial networks (GAN) and robust principal component analysis (RPCA) to extract features between heterogeneous models and perform feature alignment and comparison analysis has become a research hotspot.
[0004] Generative adversarial networks (GANs) are used to generate higher quality images in the field of image enhancement. For example, the semantic adversarial data enhancement algorithm (ASDA) proposed by Tencent Youtu Lab has made a breakthrough in the task of human 2D pose estimation. It simulates challenging cases that are difficult for the network to handle by replacing and transforming images at a semantic granularity, effectively improving the prediction difficulty and accuracy of the pose estimation network. Adversarial training enhances the robustness of the model by introducing adversarial samples during the training process. For example, the fast gradient sign method (FGSM) proposed by Ian Goodfellow et al. is used to generate adversarial samples. This method quickly generates adversarial samples by adding perturbations in the gradient direction of the loss function, thereby improving the model's defense against adversarial attacks.
[0005] Robust principal component analysis (RPCA) is a low-rank model that handles outliers or noise in data by decomposing the data matrix into the sum of a low-rank matrix and a sparse matrix. This method is widely used in image processing, video surveillance, and text analysis. For example, it can separate the background and foreground in video surveillance to detect abnormal events in the video. LRR is a method for subspace clustering that discovers the intrinsic structure of data by representing the data as a linear combination of multiple subspaces. This method has been successfully applied in fields such as face recognition, image segmentation, and recommendation systems. For example, in face recognition, LRR can help identify facial features in an image and distinguish them from other features.
[0006] However, the above existing technologies all have some defects, as follows: Although GAN has achieved remarkable results in image enhancement, the training process may be unstable and prone to mode collapse, that is, the generator starts to generate very similar or identical samples instead of diverse data. In addition, GAN is very sensitive to the choice of hyperparameters and needs to be carefully tuned to achieve optimal performance. Low-rank-based methods may not work well when the image content is complex or the salient area is not much different from the background. In addition, the low-rank representation may not fully capture all the details of the image, resulting in the loss of some important information.
[0007] The target detection algorithm based on traditional machine learning needs to manually extract discriminative features, which requires professional knowledge and experience; the target detection algorithm based on neural networks may require different features to be extracted from different images, so this method is susceptible to noise interference, has certain limitations, and is difficult to adapt to the detection of objects of different forms.
[0008] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0009] Purpose of the invention: The technical problem to be solved by the present invention is to provide a low-rank associated multi-source target detection method based on adversarial data enhancement in response to the deficiencies of the prior art.
[0010] In order to solve the above technical problems, the present invention discloses a low-rank associated multi-source target detection method based on adversarial data enhancement, comprising the following steps:
[0011] Step 1: collect image sets from different domains as two original data sets X and Y;
[0012] Step 2: Input the data in the two original data sets into the encoder for dimensionality reduction processing to obtain two new data sets X′ and Y′;
[0013] Step 3, by applying the association matrices U and V, a unified clustering structure S is introduced and the unified clustering structure S is learned from the shared features of the datasets X′ and Y′;
[0014] Step 4: Generate virtual datasets for two new datasets X′ and Y′ using the pre-trained generative adversarial network and the clustering structure S. and
[0015] Step 5: Use the generated virtual data together with the real data to train the existing target detection model, and use the trained target detection model to perform target detection.
[0016] Furthermore, the two image sets collected in different domains described in step 1 include:
[0017] The image sets from different domains are collected as the original data sets, which are expressed as follows:
[0018] {X}={x 1 ,x 2 ,…,x n}
[0019] {Y}={y 1 ,y 2 ,…,y n′}
[0020] Among them, n and n′ represent the number of data samples in different domain image sets respectively.
[0021] Furthermore, the dimensionality reduction process described in step 2 includes:
[0022] All data X and Y in the original data set are input into the encoder for dimensionality reduction processing, and new data X' and Y' are obtained to form two new data sets, which are expressed as follows:
[0023] X′∈R r×n′
[0024] Y′∈R r×n′
[0025] Among them, r represents the dimension of the new data after the encoder reduces the dimension.
[0026] Furthermore, the unified clustering structure S is learned from the shared features of the datasets X′ and Y′ as described in step 3, including:
[0027] Step 3-1: Introduce a unified clustering structure S and construct the objective function
[0028] Step 3-2: Based on the Lagrangian function, the objective function Solve the variables in to obtain the public low-rank matrix S and the noise data E 1 and E 2 .
[0029] Furthermore, the objective function constructed in step 3-1 The details are as follows:
[0030]
[0031] stX′=X′S+E 1 ,Y′=Y′S+E 2,S1=1,diag(S)=0,S≥0,rank(L S ) = nc
[0032] Among them, U and V are incidence matrices, S∈R n×n represents the common low-rank matrix of X′ and Y′, which is also the shared subspace affinity matrix in the hidden layer, λ 2 and λ 3 is the regularization parameter, ‖·‖ * is the kernel function used to solve the rank minimization problem of the public low-rank matrix S, ‖E 1 ‖ 2,1 and ‖E 2 ‖ 2,1 represents the regularization strategy of the error, where E 1 and E 2 represents the noisy data, and r 1 and r 2 represents the sample dimension, η 1 and η 2 is the regularization parameter, 1∈R n×1 is a unit vector, L S is the Laplacian matrix, D is a diagonal matrix, and the diagonal elements are the sum of the weights of the nodes.
[0033] Furthermore, the incidence matrices U and V in step 3-1 are expressed as follows:
[0034] X′ T U=Y′ T V+E
[0035] in, and is the correlation matrix of the cross-domain datasets X′ and Y′, E∈R n×n Indicates the effect of noisy data.
[0036] Furthermore, the Laplace matrix described in step 3-1 is a symmetric positive definite matrix, which is expressed as follows:
[0037] L S =D s -(S+S T ) / 2
[0038] Among them, D s ∈R n×n is the degree matrix.
[0039] Furthermore, the generative adversarial network described in step 4 is a generative adversarial network based on a fully convolutional layer.
[0040] Furthermore, the pre-training described in step 4 transforms the decoder in the generative adversarial network into a generator, and adds the input of the original decoder to the random vector input generator, identifies the generated result through the discriminator, iterates the above process until the Nash equilibrium is reached, and realizes the pre-training.
[0041] Furthermore, the virtual data sets described in step 4 are generated for the two new data sets X′ and Y′ respectively. and include:
[0042] Step 4-1, in the generative adversarial network, the generator not only receives the same random noise vector z, but also receives data from two new data sets as additional inputs, and maps them to the data space to generate virtual data and Among them, G 1 and G 2 Respectively represent two generators in the generative adversarial network;
[0043] Step 4-2: The discriminator evaluates the input data X′ or Y′ and outputs a scalar indicating the probability D that X′ or Y′ is a true sample. 1 (X′) or D 2 (Y′).
[0044] Beneficial effects:
[0045] The method provided by the present invention solves the problem of multi-source heterogeneous target detection with a small number of samples, uses generative adversarial networks (GANs) for data enhancement, and solves the problem of sample sparsity by generating new data with a distribution similar to that of real data. High-quality virtual data is generated through the adversarial game between the generator and the discriminator, which increases the diversity and quantity of data and improves the generalization ability and performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.
[0047] Figure 1 It is a schematic diagram of the overall process of the present invention.
[0048] Figure 2 Schematic diagram of the data set after noise addition.
[0049] Figure 3 Schematic diagram of experimental results in the examples.
[0050] Figure 4 Diagram of the network structure for the pre-trained model. DETAILED DESCRIPTION
[0051] The present invention provides a low-rank correlation multi-source target detection method based on adversarial data enhancement, such as Figure 1 As shown, the specific technical solution is as follows:
[0052] Step 1: Collect images from different domains {x} = {x 1 ,x 2 ,…,x n} and image {y}={y 1 ,y 2 ,…,y n′} as the required data set, where n and n′ represent the number of different samples respectively;
[0053] Step 2: Input the data X and data Y in the datasets {x} and {y} into the encoder for dimensionality reduction processing, and obtain X′∈R r×n′ and Y′∈R r×n′ r represents the sample dimension of the reduced-dimensional data obtained by the encoder;
[0054] Among them, data X and data Y are respectively input into the encoder for dimensionality reduction processing to obtain X' and Y', and the obtained linear data X' and Y' are used to extract the most critical information. At the same time, noise elements are added to the association representation of cross-domain data to make it more robust to low-quality images. Then the correlation between X' and Y' can be expressed by the following formula:
[0055] X′ T U=Y′ T V+E
[0056] in and is the correlation matrix of cross-domain data X′ and Y′, E∈R n×n Indicates the effect of noisy data.
[0057] Step 3: To strengthen the correlation between X′ and Y′, a unified clustering structure is introduced by applying the correlation matrices U and V. A unified clustering structure S is learned from the shared features of data X′ and Y′, and S is also an affinity matrix;
[0058] Among them, the application of the correlation matrix U and V introduces a unified clustering structure S, and then obtains the objective function It is expressed as:
[0059]
[0060] stX′=X′S+E 1 ,Y′=Y′S+E 2 ,S1=1,diag(S)=0,S≥0,rank(L S) = nc
[0061] where S∈R n×n represents the common low-rank matrix of X′ and Y′, λ 2 and λ 3 is the regularization parameter. ‖·‖ * is the kernel function used to solve the rank minimization problem of S, which is also the shared subspace affinity matrix in the hidden layer. 1 ‖ 2,1 and ‖E 2 ‖ 2,1 Represents the regularization strategy of the error, which is used to reduce the impact of noisy data, where E 1 and E 2 represents the noisy data, and r 1 and r 2 Represents the sample dimension. η 1 and η 2 are regularization parameters, which are multiplied by In order to optimize and update the parameters. n×1 is a unit vector. The Laplace matrix L S =D S -(S+S T ) / 2 is a symmetric positive definite matrix, D S ∈R n×n is the degree matrix. D is a diagonal matrix whose diagonal elements are the sum of the weights of the nodes.
[0062] Based on the Lagrangian function, the process of updating and solving the variables S, U, V, E1, and E2 is as follows:
[0063] Step 1: According to the Lagrangian function, the public low-rank matrix S can be obtained by optimizing and solving the following formula:
[0064]
[0065] in, represents the loss function for S, x i ′ represents the i-th image sample, x j ′ represents the jth image sample, S ij represents the weight of the distance from the i-th sample to the j-th sample, y i ′ represents the i-th image sample, y j ′ represents the jth image sample, represents the square of the two-norm, represents the square of the F norm, E1, E2 represent noise data, M 5 and M 6is the Lagrange multiplier and μ is the penalty parameter.
[0066] Step 2: By taking the partial derivative of the common low-rank matrix S of the Lagrangian function and setting it to 0, we can calculate:
[0067]
[0068] Where η represents the Lagrange multiplier, M 3 Represented as a vector, where the jth element is M 3j , n i and e ij Also set it up in the same way;
[0069] Step 3: By taking the partial derivatives of the Lagrangian function U and V and setting them to 0, we can get:
[0070] U=(X′L S X′ T +μX′X′ T +μI) -1 (μX′(Y′ T V+E)-X′M 4 -M 1 )
[0071] V=(Y′L S Y′ T +μY′Y′ T +μI) -1 (μY′(X′ T UE)+Y′M 4 -M 2 )
[0072] Among them, M 1 、M 2 and M 4 is the Lagrange multiplier and I represents the identity matrix.
[0073] Step 4: Based on the S, U and V calculated in the above steps, the residual term E 1 and E 2 It can be obtained by optimizing the following formula:
[0074]
[0075] The function E1 * τ (T) is represented as follows:
[0076]
[0077] Among them, (:,i) represents the i-th column, t i represents the i-th row of the matrix, and τ represents is the Lagrange multiplier;
[0078] Function E2 * τ (T) is represented as follows:
[0079]
[0080] Among them, (:,i) represents the i-th column, t i represents the i-th row of the matrix, and τ represents is the Lagrange multiplier;
[0081] Step 4: In order to solve the problem that the generalization performance of the model is reduced due to the small number of samples in the target domain, the present invention applies a generative adversarial network to generate a large number of virtual samples similar to the real data distribution for the source domain and the target domain respectively, thereby bridging the quantitative gap between the source domain and the target domain samples.
[0082] Among them, two generative adversarial networks are designed. The generator not only receives the same random noise vector, but also receives the source domain sample as an additional input and maps it to the data space to generate fake samples. The discriminator evaluates the input sample x (or y) and outputs a scalar representing the probability that x (or y) is a real sample.
[0083] Step 5: Add a pre-training stage to the generative adversarial network in step 4 and transform the original decoder into a generator to further improve the structural similarity between the generated virtual data and the real data, and improve the performance of the generative adversarial network in the case of few samples.
[0084] The network structure of the model pre-training is as follows Figure 4 As shown, specifically including:
[0085] The original data X and Y are input into the encoder to obtain X′ and Y′, and the linear correlation between the two is explored to learn the projection matrix. Then, the unified clustering structure S of cross-domain data is learned through projection and smoothing. Then, the input of the original decoder and the random vector are fed into the generator respectively. Finally, the generated result and the source domain data are fed into the discriminator. The optimal generator is trained repeatedly until the Nash equilibrium is reached.
[0086] Step 6: Use the generated virtual data together with the real data to train the target detection model, and use the trained target detection model to perform target detection.
[0087] Example:
[0088] In this implementation example, the ImageCLEF-DA image dataset is selected as the cross-domain dataset {X, Y}. This dataset is a benchmark dataset consisting of three domains with a total of 12 classes: Caltech-256 (C), ImageNet ILSVRC 2012 (I), and Pascal VOC 2012 (P). Each class has 50 images and each domain has 600 images. The dataset {X, Y} is processed by adding noise. Figure 2 shown.
[0089] The present invention measures the correlation between cross-domain data through typical association analysis, and introduces manifold learning so that cross-domain data can share a unified Laplace clustering structure to extract the correlation information between cross-domain data; at the same time, in the case of high noise density, the influence of noise on the correlation information of cross-domain data can be well suppressed by applying low-rank constraints, thereby obtaining a lower-rank matrix to achieve better noise reduction effect; finally, a generative adversarial network is applied to generate a large number of virtual samples similar to the real data distribution for the source domain and the target domain, thereby bridging the gap in the number of samples between the source domain and the target domain. In order to more clearly illustrate the effect of the present invention, as shown in FIG. Figure 2 Different degrees of noise are applied to the cross-domain dataset images, and the cross-domain dataset images with different degrees of noise are input into the scheme of the present invention, the general domain adaptation method (DCC), the semi-supervised domain adaptation method (CDAC), the unsupervised domain adaptation method (SRDC), the self-attention interactive generation clustering method (H-SRDC), and the active domain adaptation method (CLUE) for processing, and the clustering accuracy of the images is calculated respectively. Figure 3 Under different noise levels, the processing accuracy of the CLUE method is second only to the method proposed in the present invention, the processing accuracy of the SRDC method is the lowest, and the processing accuracy of the method proposed in the present invention is always higher than that of other comparison algorithms.
[0090] In a specific implementation, the present application provides a computer storage medium and a corresponding data processing unit, wherein the computer storage medium can store a computer program, and the computer program can run the invention content of a low-rank correlation multi-source target detection method based on adversarial data enhancement provided by the present invention and some or all of the steps in each embodiment when executed by the data processing unit. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0091] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of computer programs and their corresponding general hardware platforms. Based on such an understanding, the technical solutions in the embodiments of the present invention can be essentially or partly contributed to the prior art in the form of computer programs, i.e., software products, which can be stored in a storage medium and include several instructions for enabling a device including a data processing unit (which can be a personal computer, a server, a single-chip microcomputer, an MCU or a network device, etc.) to execute the methods described in various embodiments of the present invention or certain parts of the embodiments.
[0092] The present invention provides a method and idea for low-rank correlation multi-source target detection based on adversarial data enhancement. There are many methods and approaches to implement the technical solution. The above is only a preferred implementation of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented by existing technologies.
Claims
1. A low-rank correlation multi-source target detection method based on adversarial data enhancement, characterized in that: The following steps are involved: Step 1: collect image sets from different domains as two original data sets X and Y; Step 2: Input the data in the two original data sets into the encoder for dimensionality reduction processing to obtain two new data sets X′ and Y′; Step 3, by applying the association matrices U and V, a unified clustering structure S is introduced and the unified clustering structure S is learned from the shared features of the datasets X′ and Y′; Step 4: Generate virtual datasets for two new datasets X′ and Y′ using the pre-trained generative adversarial network and the clustering structure S. and Step 5: Use the generated virtual data together with the real data to train the existing target detection model, and use the trained target detection model to perform target detection.
2. According to claim 1, a low-rank correlation multi-source target detection method based on adversarial data enhancement is characterized in that: The two image sets of different domains described in step 1 are collected, including: The image sets from different domains are collected as the original data sets, which are expressed as follows: {X}={x1,x2,...,x n } {Y}={y1,y2,...,y n ,} Among them, n and n′ represent the number of data samples in different domain image sets respectively.
3. According to claim 2, a low-rank correlation multi-source target detection method based on adversarial data enhancement is characterized in that: The dimensionality reduction process described in step 2 includes: All data X and Y in the original data set are input into the encoder for dimensionality reduction processing, and new data X' and Y' are obtained to form two new data sets, which are expressed as follows: X′∈R r×n′ Y′∈R r×n′ Among them, r represents the dimension of the new data after the encoder reduces the dimension.
4. The low-rank correlation multi-source target detection method based on adversarial data enhancement according to claim 3 is characterized in that: The unified clustering structure S is learned from the shared features of the datasets X′ and Y′ as described in step 3, including: Step 3-1: Introduce a unified clustering structure S and construct the objective function Step 3-2: Based on the Lagrangian function, the objective function Solve the variables in to obtain the public low-rank matrix S and the noise data E1 and E2.
5. The low-rank correlation multi-source target detection method based on adversarial data enhancement according to claim 4 is characterized in that: Construct the objective function as described in step 3-1 The details are as follows: s.t.X′=X'S+E1,Y′=Y'S+E2,S1=1,diag(S)=0,S≥0,rank(L S )=n-c Among them, U and V are incidence matrices, S∈R n×n represents the common low-rank matrix of X′ and Y′, which is also the shared subspace affinity matrix in the hidden layer. λ2 and λ3 are regularization parameters. ||·|| * is the kernel function used to solve the rank minimization problem of the public low-rank matrix S, ||E1|| 2,1 and ||E2|| 2,1 represents the regularization strategy of the error, where E1 and E2 represent the noisy data, and r1 and r2 represent the sample dimensions, η1 and η2 are regularization parameters, 1∈R n×1 is a unit vector, L S is the Laplacian matrix, D is a diagonal matrix, and the diagonal elements are the sum of the weights of the nodes.
6. The low-rank correlation multi-source target detection method based on adversarial data enhancement according to claim 5 is characterized in that: The incidence matrices U and V in step 3-1 are expressed as follows: X′ T U=Y′ T V+E in, and is the correlation matrix of the cross-domain datasets X′ and Y′, E∈R n×n Indicates the effect of noisy data.
7. The low-rank correlation multi-source target detection method based on adversarial data enhancement according to claim 6 is characterized in that: The Laplacian matrix described in step 3-1 is a symmetric positive definite matrix and is represented as follows: L S =D S -(S+S T ) / 2 Among them, D S ∈R n×n is the degree matrix.
8. The low-rank correlation multi-source target detection method based on adversarial data enhancement according to claim 7 is characterized in that: The generative adversarial network described in step 4 is a generative adversarial network based on a fully convolutional layer.
9. The low-rank correlation multi-source target detection method based on adversarial data enhancement according to claim 8, characterized in that: The pre-training described in step 4 transforms the decoder in the generative adversarial network into a generator, adds the input of the original decoder to the random vector input generator, identifies the generated result through the discriminator, iterates the above process until the Nash equilibrium is reached, and realizes the pre-training.
10. The low-rank correlation multi-source target detection method based on adversarial data enhancement according to claim 9, characterized in that: Generate virtual data sets for the two new data sets X' and Y' as described in step 4 and include: Step 4-1, in the generative adversarial network, the generator not only receives the same random noise vector z, but also receives data from two new data sets as additional inputs, and maps them to the data space to generate virtual data and Wherein, G1 and G2 represent two generators in the generative adversarial network respectively; In step 4-2, the discriminator evaluates the input data X′ or Y′ and outputs a scalar representing the probability D1(X′) or D2(Y′) that X′ or Y′ is a true sample.