A method for identifying unknown type targets under open conditions
By generating unknown type targets in the input layer or intermediate layer of the classification neural network, combining data enhancement and improvement of network structure, the large amount of computation and data dependence of unknown type target recognition under open conditions is solved, and efficient and accurate recognition effect is achieved.
Patent Information
- Application Number
- CN202210917882.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-01
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-08-01
AI Technical Summary
The prior art has a large amount of calculation and relies on a large amount of data in the recognition of unknown type targets under open conditions, making it difficult to achieve fast and accurate identification.
By randomly cropping and splicing or random linear interpolation at the input layer or intermediate layer of the classification neural network, unknown type targets are generated, combined with data enhancement and improved classification neural network structure, K+1 class classification model is trained to judge unknown targets using threshold values.
It realizes that without relying on a large number of training samples, simplifies the calculation amount, improves the accuracy and efficiency of target recognition of unknown types, effectively avoids overfitting, suppresses background noise interference, and improves the classification accuracy of SAR image targets.
Smart Images

Figure CN115393628B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image information processing, and particularly relates to a method for identifying unknown type targets under open conditions. Background Art
[0002] Traditional target recognition methods assume that prior information about all target types is known. The target recognition process is essentially a process of finding the best match for information, which can be regarded as target recognition under closed conditions. However, from the perspective of practical applications, this assumption is unrealistic. On the one hand, the target types themselves are dynamically changing, and new target types are constantly increasing; on the other hand, it is difficult to obtain data of non-cooperating parties or enemy targets. Therefore, conducting synthetic aperture radar target recognition under dynamic open conditions, making full use of the existing target type data, and accurately rejecting unknown type targets has very important theoretical research significance.
[0003] In recent years, in order to solve the problem of unknown target discrimination, scholars at home and abroad have carried out a large number of studies. Existing algorithms can be classified into two categories: one is the recognition method based on discriminant models. This family of methods believes that known type targets have a higher output confidence after passing through the classification model under closed conditions, while unknown type targets have a lower output confidence. Therefore, by setting a threshold for the output confidence, the model can correctly discriminate between known type targets and unknown type targets. Although this method is simple and intuitive, they are mainly used as post-processing methods after CNN network feature extraction. The setting of the decision boundary depends heavily on the information obtained after the end of network training. When the extracted features are not rich enough, the recognition effect will be greatly affected; the other is the recognition method based on generative models. This family of methods can be further divided into methods based on the reconstruction of known type targets and methods based on the generation of unknown type targets. The basic idea of the method based on the reconstruction of known type targets is to reconstruct the distribution information of known type targets by means of GAN, Auto-Encoder, Flow-based Model, etc. Compared with known class targets, unknown class targets have a larger reconstruction error due to the lack of corresponding fitting models. Therefore, the unknown targets can be judged based on the reconstruction error. The method based on the generation of unknown type targets is to simulate the distribution information of unknown type targets by means of similar means, thereby transforming the target recognition problem under open conditions into a traditional recognition problem to achieve the classification of known targets and the discrimination of unknown targets.
[0004] However, the disadvantage of these two traditional generative models based on means such as GAN, Auto-Encoder, Flow-based Model, etc. is that in order to more accurately simulate the type distribution, a very sufficient amount of data and complex computational amount are often required. The generation process is time-consuming and it is difficult to achieve an ideal speed. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides a method for identifying unknown type targets under open conditions. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0006] A method for identifying unknown type targets under open conditions, the method for identifying unknown type targets includes:
[0007] Step 1, obtain a first training target set, where the first training target set includes K types of known type targets;
[0008] Step 2, perform data augmentation preprocessing on the first training target set to obtain a second training target set;
[0009] Step 3, select n types of target data with a size of B from different types from the second training target set, and adjust the order of the n types of target data n times to obtain n input data with different permutation ways, so as to obtain an unknown target source array based on the splicing result of the n input data with different permutation ways, where 2 ≤ n ≤ K;
[0010] Step 4, input the unknown target source array into a classification neural network, and use random cropping and splicing or random linear interpolation to process the unknown target source array to obtain an unknown type target;
[0011] Step 5, input the unknown type target and the second training target set into the classification neural network to be trained to train the classification neural network, and obtain a K + 1 class classification model after the loss function converges;
[0012] Step 6, iteratively input the validation data set into the K + 1 class classification model, and finally perform Softmax normalization processing on the output activation vector after passing through the fully connected layer to obtain an output probability score. The maximum value in the output probability scores corresponding to each data in the validation data set is used as the maximum probability output value of the data; sort the maximum probability output values corresponding to all data in the validation data set in ascending order, and select a maximum probability output value as the threshold ε1 based on the first preset position from the sorted maximum probability output values. The first preset position is the position of the maximum probability output value at the 10% ratio position close to the minimum value on the left side within the sorting interval, and this maximum probability output value is used as the threshold ε1; sort the output probability scores under the K + 1 class corresponding to all data in the validation data set in ascending order, and select one as the threshold ε2 based on the second preset position from the sorted output probability scores. The second preset position is the position of the output probability score at the 90% ratio position close to the minimum value on the left side within the sorting interval, and this output probability score is used as the threshold ε2;
[0013] Step 7: Input the target to be recognized into the K + 1 class classification model to obtain the maximum probability output value of the target to be recognized, and determine the recognition result of the target to be recognized according to the relationship between the maximum probability output value of the target to be recognized and the threshold ε1 and the relationship between the output probability score of the target in the K + 1 category and the threshold ε2.
[0014] In an embodiment of the present invention, the classification neural network is obtained by adding a combination layer between two original combination layers in the middle position of the classifier32 classification neural network and adding a combination layer between the last two original combination layers, and the combination layer includes a convolutional layer and an activation function layer.
[0015] In an embodiment of the present invention, obtaining the unknown target source array based on the splicing result of the input data in the n different permutation ways includes:
[0016] Splice the input data in the n different permutation ways by column to obtain an array of B×n;
[0017] Check each row of the B×n array to determine whether there is data of the same type in each row. If so, delete the data in that row; if not, retain the data in that row to obtain the unknown target source array.
[0018] In an embodiment of the present invention, step 4 includes:
[0019] Input the unknown target source array into the classification neural network to randomly crop and splice the unknown target source array at the input layer of the classification neural network to obtain an unknown type target, denoted as the first technical route; or,
[0020] Input the unknown target source array into the classification neural network to perform random linear interpolation on the unknown target source array at the middle layer of the classification neural network to obtain an unknown type target, denoted as the second technical route, where the middle layer is the combination layer in the middle position of the classification neural network.
[0021] In an embodiment of the present invention, when n is an even number, inputting the unknown target source array into the classification neural network to randomly crop and splice the unknown target source array at the input layer of the classification neural network to obtain an unknown type target includes:
[0022] S1.1: First, randomly select 1 h and (n / 2 - 1) w according to the beta distribution to form (n / 2 - 1) coordinates, denoted as (w1, h), (w2, h) to (w n / 2-1 , h);
[0023] S1.2. Use the (n / 2 - 1) coordinates obtained in S1.1 as the center points to draw horizontal and vertical lines along the x-axis and y-axis, respectively, to divide the image into n cropping regions, and denote their shapes as v0(a0, b0), v1(a1, b1) to v n-1 (a n-1 , b n-1 ), where v0(a0, b0) to v n / 2-1 (a n / 2-1 , b n / 2-1 ) are above the horizontal line, and v n / 2 (a n / 2 , b n / 2 ) to v n-1 (a n-1 , b n-1 ) are below the horizontal line, where a i and b i represent width and length respectively, and 0 ≤ i ≤ n - 1;
[0024] S1.3. Generate random numbers x i and y i within the ranges of (0, a - a i ) and (0, b - b i ) respectively. Let (x i , y i ) represent the starting position of the cropping region of the i-th image. Taking (x i , y i ) as the upper left corner of the cropping region, crop the i-th image according to the shape of v i (a i , b i ) to obtain the cropped image v i ′(a i , b i );
[0025] S1.4. Horizontally splice the images v′0(a0, b0) to the image v′ n / 2-1 (a n / 2-1 , b n / 2-1 ) column by column along the x-axis to obtain the image Horizontally splice the images v′ n / 2 (a n / 2 , b n / 2 ) to the image v′ n-1 (a n-1 , b n-1 ) column by column along the x-axis to obtain the image
[0026] S1.5. Vertically splice the image and the image row by row along the y-axis to obtain an object of unknown type.
[0027] In one embodiment of the present invention, when n is odd, the unknown target source array is input into the classification neural network to randomly crop and splice the unknown target source array at the input layer of the classification neural network, and an unknown type target is obtained, including:
[0028] S2.1. First, randomly select 1 h and (n - 1) / 2 w according to the beta distribution to form (n - 1) / 2 coordinates, which are respectively denoted as (w1, h), (w2, h) to (w (n-1) / 2 , h);
[0029] S2.2. Make vertical lines along the vertical axis with the (n - 1) / 2 coordinates obtained in S2.1 as the center points, and make horizontal lines within the range of (0, w (n-1) / 2 ) along the horizontal axis to divide the image into n cropping regions, and their shapes are respectively denoted as v0(a0, b0), v1(a1, b1) to v n-1 (a n-1 , b n-1 ), where v0(a0, b0) to v (n-3) / 2 (a (n-3) / 2 , b (n-3) / 2 ) are above the horizontal line, v (n+1) / 2 (a (n+1) / 2 , b (n+1) / 2 ) to v n-1 (a n-1 , b n-1 ) are below the horizontal line, v (n-1) / 2 (a (n-1) / 2 , b (n-1) / 2 ) is the region not divided by the horizontal line, a i , b i respectively represent the length and width, 0 ≤ i ≤ n - 1;
[0030] S2.3. For the random numbers x i and y i generated respectively within the ranges of (0, a - a i ) and (0, b - b i ), let (x i , y i ) represent the starting position of the cropping region of the i-th image, with (x i , y i ) as the upper left corner of the cropping region, and crop the i-th image according to the shape of v i (a i , b i ) to obtain the cropped image v″ i (a i , b i );
[0031] S2.4. Concatenate the images v″0(a0, b0) to v″ (n-3) / 2 (a (n-3) / 2 , b (n-3) / 2 ) column by column along the horizontal axis to obtain an image Concatenate the images v″ (n+1) / 2 (a (n+1) / 2 , b (n+1) / 2 ) to v″ n-1 (a n-1 , b n-1 ) column by column along the horizontal axis to obtain an image
[0032] S2.5. Concatenate the image and the image row by row along the vertical axis to obtain an image
[0033] S2.6. Concatenate the image and the image v′ (n-1) / 2 (a (n-1) / 2 , b (n-1) / 2 ) column by column along the horizontal axis to obtain an unknown type target.
[0034] In an embodiment of the present invention, the unknown type target obtained by random linear interpolation is represented as:
[0035]
[0036] Wherein, is the unknown type target, C1 to C n are weight coefficients, C1 + C2 + … + C n = 1, C1 to C n are random numbers that satisfy the beta distribution, is all the layers before the combination layer at the middle position of the classification neural network.
[0037] In an embodiment of the present invention, for the first technical route, step 5 includes:
[0038] Step 5.1. Concatenate the unknown type target and the second training target set, and input them together into the classification neural network to be trained to obtain the output feature distribution of the unknown type target and the output feature distribution of the known type target;
[0039] Step 5.2. Continuously train the neural network to optimize the output feature distribution of the classification neural network to be trained until the first loss function converges, and obtain a K + 1 class classification model after the first loss function converges, where the first loss function is composed of the decision losses of all K + 1 class training targets;
[0040] For the second technical route, step 5 includes:
[0041] Step 5.1: Input the target of unknown type into the classification neural network in the second half to obtain the output feature distribution of the target of unknown type. At the same time, the second training target set is input into the complete classification neural network to be trained to obtain the output feature distribution of the target of known type, where the classification neural network in the second half is the combination layer to the fully connected layer in the middle position;
[0042] Step 5.2: By continuously training the neural network, optimize the output feature distribution of the classification neural network to be trained until the second loss function converges, and obtain the K + 1 class classification model after the second loss function converges, where the second loss function is composed of the decision loss of the unknown class and the classification loss of the known class.
[0043] In an embodiment of the present invention, the first loss function is:
[0044]
[0045] where l t1 is the first loss function, y k is the true label of the input sample, is the preset loss function, is the output feature distribution of all K + 1 class training targets;
[0046] The second loss function is:
[0047] l t2 = l k + α u · l u
[0048]
[0049]
[0050] where l t2 is the second loss function, l k is the classification loss of the known type, y k is the true label of the input sample, is the preset loss function, is the output feature distribution of the known type samples, is the output feature distribution of the unknown type samples, l u is the decision loss of the unknown type, α u is the hyperparameter.
[0051] In one embodiment of the present invention, the prediction result of the target to be recognized is determined according to the relationship between the maximum probability output value of the target to be recognized and the threshold ε1 and the relationship between the output probability score of the target in the (K + 1)-th category and the threshold ε2, including:
[0052] Judge the relationship between the maximum probability output value of the target to be recognized and the threshold ε1. If the maximum probability output value of the target to be recognized is less than or equal to the threshold ε1, the recognition result is an unknown target. If the maximum probability output value of the target to be recognized is greater than the threshold ε1, then judge the relationship between the output probability score of the (K + 1)-th category and the threshold ε2. If the output probability score is greater than or equal to the threshold ε2, the recognition result is an unknown target. Otherwise, the recognition result of the target is the type to which the maximum probability output value among the first K known categories belongs.
[0053] Advantages of the present invention:
[0054] First, the amount of calculation required to generate unknown class targets in the method proposed by the present invention is simple, and an ideal recognition effect can be achieved without relying on a large number of training samples;
[0055] Second, the method proposed by the present invention can extract more effective high-dimensional feature information by means of newly generated unknown class samples, push the decision boundary towards its corresponding clustering, and compact the within-class spacing;
[0056] Third, the method proposed by the present invention can construct diverse new pixel-level features on the basis of repeatedly learning local features of known classes, effectively avoiding the overfitting phenomenon of the network.
[0057] Fourth, the method proposed by the present invention can effectively suppress the interference of background noise in SAR images on classification by means of random cropping, and improve the classification accuracy of SAR image targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 is a schematic flow chart of a method for recognizing unknown type targets under open conditions provided by an embodiment of the present invention;
[0059] Figure 2 is a structural diagram of a classification network improved based on classifier32 provided by an embodiment of the present invention;
[0060] Figure 3 is a schematic diagram of a segmentation region when n is an even number provided by an embodiment of the present invention;
[0061] Figure 4 is a schematic diagram of a segmentation region when n is an odd number provided by an embodiment of the present invention;
[0062] Figure 5It is a schematic diagram of an unknown target generated by randomly splicing four types of samples provided by an embodiment of the present invention;
[0063] Figure 6 It is a schematic diagram of an unknown target generated by randomly interpolating (weighted combination) two types of samples in the middle layer provided by an embodiment of the present invention;
[0064] Figure 7 It is a schematic diagram of an unknown target generated according to the first technical route provided by an embodiment of the present invention;
[0065] Figure 8 It is a schematic diagram of an unknown target generated according to the second technical route provided by an embodiment of the present invention. Detailed implementation manners
[0066] The following further describes the present invention in detail with specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0067] Embodiment 1
[0068] Aiming at the problems existing in the discriminant model and the traditional generation model, the present invention proposes a simple and effective generation model method for simulating unknown type targets, and realizes the efficient and accurate classification of synthetic aperture radar targets under dynamic open conditions on the basis of combining the characteristics of synthetic aperture radar images (SAR).
[0069] The essence of the target recognition process is to seek the best matching of feature information. In the classification network based on deep learning, with the help of the activation vector of the last fully connected layer, the network can effectively identify the positive half space of each class. When the sample is closer to the center in the positive half space, the corresponding activation score value is higher. On the contrary, as the sample gradually moves away from the center of the positive half space, the activation score value will gradually decrease. Considering that the unknown targets under real open conditions can be divided into two types: the first type includes targets that have no connection with all known class samples, and the second type includes targets that have similarity with a certain known class in a certain aspect, the present invention specifically proposes a solution to identify these two types of unknown type targets simultaneously.
[0070] First, for the former, it is considered that the formation of these samples is due to the lack of any common features with the known class targets. In this case, the output activation scores of all types will be very low. Therefore, these types of targets can be directly rejected by setting a threshold for the highest activation score. Second, for the latter, it cannot be simply rejected by means of a threshold, so a simple and effective generation model method is introduced to construct the prior information of this type of unknown target for the identification of unknown targets.
[0071] For all samples of the same type of target It is generally considered that they can span the linear subspace of this type:
[0072]
[0073] wherein is a coefficient vector, and each set of imaging data of this type can be regarded as a specific element on a linear subspace. However, for the linear combination of two or more sets of imaging data from different target types, it is theoretically not spanned in the corresponding known types:
[0074]
[0075] Therefore, it can be largely ensured that it does not belong to any known type, and by virtue of the powerful learning ability and feature extraction ability of deep learning, the target obtained through this combination method often has high-dimensional features similar to some original imaging data, and can effectively simulate the second type of unknown target.
[0076] In addition, compared with general optical images, SAR images are granular speckle images with extremely strong noise. By linearly combining multiple SAR images, while retaining the original local area information, the speckle noise of the mixed image is further enhanced, and its distribution in the feature space is more likely to be near the decision boundary, effectively restricting the known types.
[0077] Therefore, the present invention proposes a method for randomly combining different types of samples in the input layer or intermediate hidden layer of a classification network to generate samples of unknown types, and specifically realizes accurate identification of unknown targets in the problem of synthetic aperture radar target recognition, including two technical routes of random cropping and stitching and random linear interpolation.
[0078] The method for identifying unknown type targets under open conditions of the present invention:
[0079] Please refer to Figure 1 , Figure 1 which is a schematic flow chart of a method for identifying unknown type targets under open conditions provided by an embodiment of the present invention. The method for identifying unknown type targets under open conditions provided by an embodiment of the present invention includes steps 1-step 7, wherein:
[0080] Step 1, obtain a first training target set, where the first training target set includes K types of known type targets, and the targets are SAR image data.
[0081] Assume that the original training target set (i.e., the first training target set or the SAR image training set) is represented as X = {x1, x2,..., x K}, where K represents the number of types of known targets.
[0082] Step 2, perform data augmentation preprocessing on the first training target set to obtain a second training target set.
[0083] Specifically, the data volume of the general SAR image training set is too small, which will cause problems of overfitting and falling into local optimal solutions in the deep learning network. Therefore, in this embodiment, data augmentation is first performed on the original training target set, which specifically includes five operations: horizontal flipping, random erasing, random cropping, random rotation, and color transformation, so as to increase the data volume of training and improve the generalization ability of the model. The training set after data augmentation is denoted as
[0084] Step 3: Select n types of target data with a size of B from different types from the second training target set, and adjust the order of the n types of target data n times to obtain n input data with different permutation methods, so as to obtain an unknown target source array based on the splicing result of the n input data with different permutation methods, where 2 ≤ n ≤ K.
[0085] Specifically, assuming that n types of target data of known types from different types are randomly selected to generate an unknown target, first, the input sample data with a batch size of B (i.e., n types of target data of known types randomly selected from different types) needs to be shuffled n times to obtain n input data with different permutation methods data. Thus, the final unknown target source array can be obtained according to the splicing result of the n input data with different permutation methods data. The unknown target source array is denoted as novel score, which represents an array of known class samples used to simulate the unknown target.
[0086] In a specific embodiment, obtaining an unknown target source array based on the splicing result of n input data with different permutation methods includes:
[0087] S1: Concatenate the n input data with different permutation methods column by column to obtain a B×n array.
[0088] S2: Check each row of the B×n array to determine whether there is data of the same type in each row. If so, delete the data in that row; if not, keep the data in that row to obtain the unknown target source array.
[0089] That is to say, when there are at least two data of the same type in any row of the B×n array, delete the data in that row; otherwise, keep the data in that row.
[0090] Step 4: Input the unknown target source array into the classification neural network to process the unknown target source array by using random cropping splicing or random linear interpolation to obtain an unknown type target.
[0091] Specifically, a classification neural network model is constructed. The network structure selected in this embodiment is the classifier32 classification network structure in the ARPL (Adversarial Reciprocal Points Learning for Open Set Recognition) algorithm, and some improvements are made on this basis, such as Figure 2 as shown. Considering that compared with visible light images, SAR images have stronger speckle noise and are very sensitive to the imaging azimuth angle. The imaging results of the same target at different azimuth angles vary greatly. Therefore, optimizing the design of the deep convolutional neural network and learning higher-dimensional features are the key factors to improve the recognition accuracy of SAR images. In this embodiment, a deeper network needs to be used to extract higher-dimensional abstract features for SAR images. Specifically, two combined layers containing a convolutional layer and an activation function layer are added to the middle hidden layer (the middle hidden layer is all the layers between the input layer and the output layer (i.e., the fully connected layer)) of the original classifier32 network structure to increase the network depth. At the same time, with the help of the regional merging and channel merging effects of the convolutional kernel on information, the merged information will have higher abstraction. Therefore, this new network structure is conducive to obtaining high-dimensional feature information for SAR images specifically.
[0092] Furthermore, the classification neural network is obtained by adding a combined layer between two original combined layers in the middle of the classifier32 classification neural network and adding a combined layer between the last two original combined layers. The combined layer contains a convolutional layer and an activation function layer. Among them, the middle hidden layer includes multiple stacked original combined layers, and each original combined layer includes a convolutional layer (conv2d), a batch normalization layer (Batchnorm2d), and a (activation function layer) LeakyReLU layer. For example, if there are a total of 9 original combined layers, then a combined layer is added between the 4th and 5th original combined layers, and a combined layer is added between the 8th and 9th original combined layers.
[0093] Assume that the model of the classification neural network is represented by f(x): x → f, that is, an abstract mapping function from imaging data to the output activation score of the last fully connected layer. The output activation score characterizes the output feature distribution of the imaging data passing through the classification network. Further, f(x) can be regarded as a combination of a feature extraction function and a linear closed-set classifier, expressed as:
[0094]
[0095] where W represents the weight matrix of linear classification, represents the abstract embedding function of feature extraction, All the layers before the intermediate hidden layer of the neural network, which map the input data to the intermediate hidden layer to obtain abstract features; correspondingly, Map the features of the intermediate hidden layer to the features of the final output layer.
[0096] In this embodiment, step 4 may specifically include:
[0097] Input the unknown target source array into the classification neural network to randomly crop and splice the unknown target source array at the input layer of the classification neural network to obtain an unknown type target, denoted as the first technical route; or, input the unknown target source array into the classification neural network to perform random linear interpolation (weighted combination) on the unknown target source array at the intermediate layer of the classification neural network to obtain an unknown type target, denoted as the second technical route, where the intermediate layer is the combined layer in the middle position of the classification neural network.
[0098] That is to say, in this embodiment, an unknown type target can be obtained by randomly cropping and splicing the unknown target source array, or an unknown type target can be obtained by performing random linear interpolation on the unknown target source array.
[0099] Randomly crop the samples from different classes at the input layer, and then splice the cropping results to generate unknown class samples. In this way, the classification neural network can repeatedly learn the local features contained in the original dataset from these unknown samples, which is beneficial to extracting deeper, abstract, and comprehensive feature information and effectively constraining the known class boundaries. In addition, the newly generated targets form new global features due to their own combination, effectively avoiding the overfitting phenomenon of these samples during the deep learning process.
[0100] In a specific embodiment, when n is an even number, input the unknown target source array into the classification neural network to randomly crop and splice the unknown target source array at the input layer of the classification neural network to obtain an unknown type target, including:
[0101] S1.1. First, randomly select 1 h and (n / 2 - 1) w according to the beta distribution to form (n / 2 - 1) coordinates, denoted as (w1, h), (w2, h) to (w n / 2-1 , h).
[0102] Specifically, assume that the image size of the known type target is uniformly a×b. Generate 1 random number h within the range of (0, b) and generate (n / 2 - 1) random numbers w within the range of (0, a).
[0103] S1.2. As Figure 3, Using the (n / 2 - 1) coordinates obtained in S1.1 as the center points, draw horizontal and vertical lines along the x-axis and y-axis directions to divide the image into n cropping regions, and denote their shapes as v0(a0, b0), v1(a1, b1) to v n-1 (a n-1 , b n-1 ), where v0(a0, b0) to v n / 2-1 (a n / 2-1 , b n / 2-1 ) are above the horizontal line, and v n / 2 (a n / 2 , b n / 2 ) to v n-1 (a n-1 , b n-1 ) are below the horizontal line. a i and b i represent width and length respectively, where 0 ≤ i ≤ n - 1;
[0104] S1.3. For randomly generated numbers x i and y i within the ranges (0, a - a i ) and (0, b - b i ) respectively, let (x i , y i ) represent the starting position of the cropping region of the i-th image. Taking (x i , y i ) as the upper left corner of the cropping region, crop the i-th image according to the shape of v i (a i , b i ) to obtain the cropped image v i ′(a i , b i ). The n images correspond to form an array novelscore of B×n structure. Therefore, there are a total of n images v′ i (a i , b i );
[0105] S1.4. Concatenate the images v′0(a0, b0) to the image v′ n / 2-1 (a n / 2-1 , b n / 2-1 ) column by column along the x-axis to obtain the image Concatenate the images v′ n / 2 (a n / 2 , b n / 2 ) to the image v′ n-1 (a n-1 , b n-1 ) column by column along the x-axis to obtain the image
[0106] S1.5. Concatenate the image and the image vertically by rows to obtain a target of unknown type.
[0107] In a specific embodiment, when n is odd, input the unknown target source array into the classification neural network to randomly crop and splice the unknown target source array at the input layer of the classification neural network to obtain a target of unknown type, including:
[0108] S2.1. First, randomly select 1 h and (n - 1) / 2 w according to the beta distribution to form (n - 1) / 2 coordinates, denoted as (w1, h), (w2, h) to (w (n-1) / 2 , h).
[0109] Specifically, assume that the image size of the known type target is uniformly a×b. Generate 1 random number h within the range of (0, b) and (n - 1) / 2 random numbers w within the range of (0, a).
[0110] S2.2. As Figure 4 , draw vertical lines along the vertical axis with the (n - 1) / 2 coordinates obtained in S2.1 as the center points, and draw horizontal lines within the range of (0, w (n-1) / 2 ) along the horizontal axis to divide the image into n cropping regions, and denote their shapes as v0(a0, b0), v1(a1, b1) to v n-1 (a n-1 , b n-1 ), where v0(a0, b0) to v (n-3) / 2 (a (n-3) / 2 , b (n-3) / 2 ) are above the horizontal line, v (n+1) / 2 (a (n+1) / 2 , b (n+1) / 2 ) to v n-1 (a n-1 , b n-1 ) are below the horizontal line, and v (n-1) / 2 (a (n-1) / 2 , b (n-1) / 2 ) is the region not divided by the horizontal line. a i and b i represent the length and width respectively, 0 ≤ i ≤ n - 1;
[0111] S2.3. For the random numbers x i and y i generated within the ranges of (0, a - a i ) and (0, b - b i ) respectively, let (x i , y i) represents the starting position of the cropping region of the i-th image, with (x i , y i ) as the upper left corner of the cropping region, and the i-th image is cropped according to the shape of v i (a i , b i ) to obtain the cropped image v″ i (a i , b i ). The i-th image corresponds to the i-th column obtained from the unknown target source array.
[0112] S2.4. Stitch the images v″0(a0, b0) to v″ (n-3) / 2 (a (n-3) / 2 , b (n-3) / 2 ) column by column along the horizontal axis to obtain the image Stitch the images v″ (n+1) / 2 (a (n+1) / 2 , b (n+1) / 2 ) to v″ n-1 (a n-1 , b n-1 ) column by column along the horizontal axis to obtain the image
[0113] S2.5. Stitch the image and the image row by row along the vertical axis to obtain the image
[0114] S2.6. Stitch the image and the image v′ (n-1) / 2 (a (n-1) / 2 , b (n-1) / 2 ) column by column along the horizontal axis to obtain the unknown type target.
[0115] For example, in this embodiment, taking the cropping and stitching of four different category samples as an example, assume that the data obtained by shuffling the order four times are data1, data2, data3, and data4. Then, a B×4 structured array is obtained by stitching column by column. After that, through the judgment of whether they belong to the same category, the novel score is obtained by screening row by row, as shown in the appendix Figure 5 .
[0116] Assume that the image sizes of the known class targets are uniformly a×b. Randomly generate two numbers within the ranges of (0, a) and (0, b), denoted as w and h. Then, (w, h) constitutes a coordinate point on the two-dimensional image plane. Since the selection of a certain number within the given range is random, and compared with the normal distribution, the beta distribution can generate different distribution shapes more flexibly. Therefore, random numbers are generated according to the beta distribution, as shown in the following formula:
[0117] w = round(w'a), w' ~ B(α, α)
[0118] h = round(h'a), h' ~ B(α, α)
[0119] where α ∈ (0, ∞), representing a hyperparameter.
[0120] First, taking (w, h) as the boundary positions, determine the shapes of the four cropping regions, and then crop the four columns of sample images of the novel score according to the set shapes respectively. The specific process is as follows:
[0121] For the image set corresponding to novel score[:, 0], generate random numbers x0 and y0 in the ranges (0, w + 1) and (0, b - h + 1) respectively. Let (x0, y0) represent the starting position of the cropping region of each cropped image, and the cropping region is expressed as:
[0122] v0 = v<(x0, y0), (x0 + w, y0 + h)>
[0123] where v<·> represents cropping a rectangle with the line segment determined by two points as the diagonal.
[0124] For the image set corresponding to novel score[:, 1], generate random numbers x1 and y1 in the ranges (0, a - w + 1) and (0, b - h + 1) respectively. Let (x1, y1) represent the starting position of the cropping region of each cropped image, and the cropping region is expressed as:
[0125] v1 = v<(x1, y1), (x1 + a - w, y1 + h)>
[0126] For the image set corresponding to novel score[:, 2], generate random numbers x2 and y2 in the ranges (0, a - w + 1) and (0, h + 1) respectively. Let (x2, y2) represent the starting position of the cropping region of each cropped image, and the cropping region is expressed as:
[0127] v2 = v<(x2, y2), (x2 + w, y2 + b - h)>
[0128] For the image set corresponding to novel score[:, 3], generate random numbers x3 and y3 in the ranges (0, w + 1) and (0, h + 1) respectively. Let (x3, y3) represent the starting position of the cropping region of each cropped image, and the cropping region is expressed as:
[0129] v3 = v<(x3, y3), (x3 + a - w, y3 + b - h)>
[0130] Finally, to ensure the The size of the target image is the same as that of the known-class target image. The image obtained by the above cropping is patched according to the boundary positions (w, h), that is, v0 and v1, v2 and v3 are concatenated column by column to obtain v' and v'', and then v' and v'' are concatenated row by row to obtain the unknown-class sample, which is expressed as:
[0131]
[0132]
[0133]
[0134] where represents concatenation along the horizontal axis, represents concatenation along the vertical axis. The concatenated unknown-type target is expanded to the original training set and input into the classification network together. Finally, the output feature distribution of all training targets is expressed as
[0135] For the second technical route, if linear interpolation is directly performed on the target at the input layer, that is the resulting mixed sample may appear in another class y i between y j and y k nearby, which will seriously affect the feature information learned by the original data set and damage the distinguishability between known classes. On the contrary, interpolation in the middle layer can make the decision boundary smoother and move away from the known-class data in all directions, thus forming a compact embedding space. In addition, since the features in the input space are fixed, the mixed samples obtained by linear interpolation at the input layer are deterministic data and cannot be further optimized by the network. In contrast, the mixed samples obtained in the middle layer can be continuously optimized through the feature extraction function between the middle layer and the output layer to better simulate the feature distribution of the second type of unknown target. Therefore, this embodiment proposes to perform random linear interpolation on the samples from different classes in the middle hidden layer of the classification network, where the weights are selected as random numbers that satisfy the beta distribution. Under the action of the random weighted combination, the generated mixed targets can effectively constrain the known-class boundaries, as shown in the appendix Figure 6 shown.
[0136] In a specific embodiment, the unknown-type target obtained by random linear interpolation is expressed as:
[0137]
[0138] where, is the unknown-type target, C1 to Cn is the weight coefficient, C1 + C2 + … + C n = 1, C1 to C n are random numbers that satisfy the beta distribution, are all the layers before the combination layer at the middle position of the classification neural network.
[0139] For example, taking the combination of two different types of samples as an example, assuming the data obtained by shuffling the order twice are data1 and data2, a B×2 structured array is obtained by concatenating them column by column. Then, through the judgment of whether they belong to the same category, row-by-row screening is performed to obtain the novel score. The novel score is input into the middle layer to generate samples of unknown categories, which is expressed as:
[0140]
[0141] where λ ~ [0, 1]. In order to generate different distribution shapes more flexibly in the [0, 1] interval and reflect the weighted trend suitable for each event, λ is also sampled according to the beta distribution. represents the unknown type target generated in the middle layer.
[0142] Step 5: Input the unknown type target and the second training target set into the classification neural network to be trained to train the classification neural network, and obtain a K + 1 class classification model after the loss function converges, that is, the trained classification neural network is the K + 1 class classification model.
[0143] In a specific embodiment, for the first technical route, Step 5 may specifically include:
[0144] Step 5.1: Concatenate the unknown type target and the second training target set and input them together into the classification neural network to be trained to obtain the output feature distribution of the unknown type target and the output feature distribution of the known type target.
[0145] Specifically, the target of unknown type is expanded to the training set (that is, the second training target set), and then directly input into the complete classification network to obtain the output feature distribution of all targets
[0146] Step 5.2: Please refer to Figure 7 , by continuously training the neural network, optimize the output feature distribution of the classification neural network to be trained until the first loss function converges, and obtain a K + 1 class classification model after the first loss function converges, where the first loss function is composed of the decision losses of all K + 1 class training targets. The first loss function is:
[0147]
[0148] where l t1 is the first loss function, y k is the true label of the input sample, is the preset loss function, is the output feature distribution of all K + 1 types of training targets.
[0149] In another specific embodiment, for the second technical route, step 5 may specifically include:
[0150] Step 5.1: Input the target of unknown type into the classification neural network in the second half to obtain the output feature distribution of the target of unknown type. At the same time, the second training target set is input into the complete classification neural network to be trained to obtain the output feature distribution of the target of known type. Among them, the classification neural network in the second half is the combination layer to the fully connected layer in the middle position.
[0151] Specifically, the target of unknown type passes through the network layer and the linear classifier W to obtain the output feature distribution of the target of unknown type while the training set (i.e., the second training target set) is directly input into the complete classification network to obtain the output feature distribution of the target of known type
[0152] Step 5.2: By continuously training the neural network, optimize the output feature distribution of the classification neural network to be trained until the second loss function converges, and obtain the K + 1 class classification model after the second loss function converges. Among them, the second loss function is composed of the decision loss of the unknown class and the classification loss of the known class.
[0153] Specifically, use the output feature distribution result to construct the loss function of the classification model. Please refer to Figure 8 . The overall loss function includes the classification loss of the known type and the decision loss of the unknown type. First, for the classification loss of the known type, it is desired to make the output feature distribution of the known type sample match its true label's one-hot encoding distribution. Then, the decision loss expression of the known type is defined as:
[0154]
[0155] where y k represents the true label of the input sample, represents the preset loss function, generally selected as the cross-entropy loss function. Of course, in practical applications, It is also possible to select various loss functions such as Hinge Loss, Huber Loss, and BCELoss.
[0156] For the decision loss of unknown types, it is necessary to first classify all unknown targets into K + 1 classes, and then optimize the output feature distribution of unknown class samples to match the one-hot encoding distribution form of K + 1 label values. Then the corresponding decision loss expression for unknown types is defined as:
[0157]
[0158] Finally, a weighted sum of these two loss functions is taken to obtain the overall loss function, and the expression is:
[0159] l t2 = l k + α u · l u
[0160] where α u is a hyperparameter. Through the stochastic gradient descent method, continuously train and optimize the overall loss function of the neural network until convergence.
[0161] After training the network to convergence through the above two methods, a classification network model under the closed-set condition with a total of K + 1 categories is obtained. Among them, the first K categories are still known category targets, and the (K + 1)-th category represents all unknown category targets. The output activation vector after the last fully connected layer at the output end is subjected to Softmax normalization processing to normalize the activation score range to the range interval of [0, 1], and then the output probability scores for K + 1 categories can be obtained, denoted as
[0162] Step 6: Iteratively input the validation dataset into the (K + 1)-class classification model. Finally, perform Softmax normalization on the output activation vector after the fully connected layer to obtain output probability scores. The maximum value among the output probability scores corresponding to each data in the validation dataset is used as the maximum probability output value for that data. Sort the maximum probability output values corresponding to all data in the validation dataset in ascending order, and select a maximum probability output value as the threshold ε1 based on the first preset position. The first preset position is the position of the maximum probability output value at the 10% ratio position close to the minimum value on the left side within the sorted range. This maximum probability output value is used as the threshold ε1. Sort the output probability scores of the (K + 1)-th class data corresponding to all data in the validation dataset in ascending order, and select an output probability score as the threshold ε2 based on the second preset position. The second preset position is the position of the output probability score at the 90% ratio position close to the minimum value on the left side within the sorted range. This output probability score is used as the threshold ε2.
[0163] Specifically, the data in the validation dataset includes data of known types and data of unknown types. Put all the data in the validation set into the obtained (K + 1)-class classification model for 200 iterative tests. Among all the iterations, select the one with the best result and observe the output values of all data under the (K + 1) categories. For example, if there are 2149 data in the validation set, there will be 2149×(k + 1) arrays. Each data corresponds to k + 1 output probability scores. Take the maximum value among these k + 1 output probability scores as the maximum probability output value. Finally, there will be a 2149×1 array. Sort these 2149 maximum probability output values in ascending order, and take the value at the 10% ratio position close to the minimum value on the left side within this range as the threshold ε1. Similarly, the threshold ε2 can also be obtained.
[0164] Step 7: Input the target to be recognized into the (K + 1)-class classification model to obtain the maximum probability output value of the target to be recognized, and determine the recognition result of the target to be recognized according to the relationship between the maximum probability output value of the target to be recognized and the threshold ε1 and the relationship between the output probability score of the target in the (K + 1)-th category and the threshold ε2.
[0165] Specifically, judge the relationship between the maximum probability output value of the target to be recognized and the threshold ε1. If the maximum probability output value of the target to be recognized is less than or equal to the threshold ε1, the recognition result is an unknown target. If the maximum probability output value of the target to be recognized is greater than the threshold ε1, then judge the relationship between the output probability score of the (K + 1)-th category and the threshold ε2. If the output probability score is greater than or equal to the threshold ε2, the recognition result is an unknown target. Otherwise, the recognition result of the target is the type to which the maximum probability output value among the first K known classes belongs.
[0166] That is to say, in the prediction stage, assume that a certain input is First, compare the maximum probability output score with the threshold ε1. If it satisfies:
[0167]
[0168] Then this instance is directly judged as an unknown target, that is, y = K + 1;
[0169] If it satisfies:
[0170]
[0171] Then further judge the size of the output probability value of the (K + 1)-th class and the threshold ε2. If it satisfies:
[0172]
[0173] Then this instance also needs to be judged as an unknown target, that is, y = K + 1. Otherwise, it is judged as the class to which the maximum probability output value among the first K known classes belongs:
[0174]
[0175] Thus, the accurate identification of instance data is finally achieved.
[0176] Next, the actual effect of the present invention is verified.
[0177] 1. Experimental conditions:
[0178] Use the publicly available MSTAR dataset to verify the algorithm of the invention. This dataset is a classic SAR image dataset, with a total of 5173 images, corresponding to 10 types of targets, namely 2S1, BMP2, BRDM2, BTR60, BTR70, D7, T62, T72, ZIL131, ZSU23 / 4. The image resolution is 0.3m×0.3m, and the pixel size is 128×128. The experiment selects the traditional Softmax algorithm as the benchmark, and uses AUC (Area under the ROC Curve for Open Set Detection) and Closed-set Accuracy as evaluation indicators to conduct a comparative verification on the proposed method. The experimental running system is Intel(R) Core(TM) i7-11700K@3.60GHz and NVIDIA GeForce RTX 3060 GPU, 64-bit Windows 10 operating system, and the simulation software uses Python 3.6.
[0179] 2. Experimental content and result analysis:
[0180] The present invention focuses on solving the problem of achieving efficient and accurate classification of SAR targets in an open environment. The proposed method is verified using the MSTAR dataset. Regarding the division of the training set and the test set, it is directly set according to the division provided by the MSTAR dataset.
[0181] The first group of experiments: First, randomly select eight types of targets from all ten types of targets in the MSTAR dataset as known classes, and the other two types of targets serve as unknown targets during the testing process. Use the technical route of randomly cropping and splicing samples of different classes to generate unknown class samples. According to experience values, here four different types of samples are randomly selected for splicing. Then, the generated unknown targets and the original K known targets are jointly input into the Softmax classification network for training. The converged network is used to perform target recognition on all images in the test set. The AUC metric and the Closed-set Accuracy metric are used to evaluate the discrimination ability of unknown class targets in an open environment and the classification ability of known class targets in a closed environment respectively. The experimental results are shown in Table 1.
[0182] Table 1 Comparison of recognition result evaluations for the first group of experiments
[0183] AUC ACC Softmax recognition result 83.0 95.559 Recognition result of the present invention 93.6 94.09 Improvement rate 10.6 -1.469
[0184] It can be seen from the experimental results that the recognition results of the method proposed in the present invention are significantly better than the original recognition results in terms of the evaluation metric AUC, with an improvement of 10.6 magnitudes compared to the original method. This indicates that the method proposed in the present invention has a stronger discrimination ability for unknown type targets in an open environment. On the other hand, the value of the Closed-set Accuracy metric does not decrease significantly, which indicates that the SAR-ATR method based on cropping and splicing proposed in the present invention can fully guarantee the classification effect of known type targets in a closed environment, thus achieving the correct identification of unknown class targets in the open world without sacrificing the closed-set recognition performance.
[0185] The second group of experiments: First, randomly select eight types of targets from all ten types of targets in the MSTAR dataset as known classes, and then use two types of MSTAR-related clutter data generated by GAN as unknown targets during the testing process. Use the technical route of randomly cropping and splicing samples of different classes to generate unknown class samples. According to experience values, here four different types of samples are still randomly selected for splicing. Then, the generated unknown targets and the original K known targets are jointly input into the Softmax classification network for training. The converged network is used to perform target recognition on all images in the test set. The AUC metric and the Closed-set Accuracy metric are used to evaluate the discrimination ability of unknown class targets and the closed-set recognition ability of known class targets respectively. The experimental results are shown in Table 2.
[0186] Evaluation and Comparison of Recognition Results of the Second Group of Experiments in Table 2
[0187] AUC ACC Softmax recognition result 73.1 95.095 Recognition result of the present invention 85.2 94.48 Improvement rate 12.1 -0.615
[0188] It can be seen from the experimental results that the recognition results of the method proposed in the present invention are significantly better than the original recognition results in terms of the evaluation index AUC, with an improvement of 12.1 magnitudes compared to the original method. This indicates that the method proposed in the present invention has a stronger discrimination ability for unknown type targets. On the other hand, the value of the Closed-set Accuracy index slightly decreases by 0.615 magnitudes. This shows that the SAR-ATR method based on cropping and splicing proposed in the present invention will not overly sacrifice the classification ability of the model for known class targets under closed-set conditions, and overall can better solve the problem of achieving efficient and accurate classification of targets under dynamic open conditions.
[0189] The third group of experiments: First, randomly select eight types of targets from all ten types of targets in the MSTAR dataset as known classes, and the other two types of targets will act as unknown targets during the testing process. Use the technical route of randomly linearly interpolating different class samples in the middle layer to generate unknown class samples. According to empirical values, randomly select two different class samples for interpolation here, and use the Softmax classification network for training. Finally, use the converged network to perform target recognition on all images in the test set, and use the AUC index and the Closed-set Accuracy index to evaluate the discrimination ability of unknown class targets and the closed-set recognition ability of known class targets respectively. The experimental results are shown in Table 3.
[0190] Evaluation and Comparison of Recognition Results of the Third Group of Experiments in Table 3
[0191] AUC ACC Softmax recognition result 83.0 95.559 Recognition result of the present invention 92.6 95.08 Improvement rate 9.6 -0.479
[0192] It can be seen from the experimental results that the recognition results of the method proposed in the present invention are significantly better than the original recognition results in terms of the evaluation index AUC, with an improvement of 9.6 magnitudes compared to the original method. This indicates that the method proposed in the present invention has achieved a very high accuracy in the discrimination of unknown type targets. On the other hand, the value of the Closed-set Accuracy index decreases slightly by 0.479 magnitudes. This shows that the SAR-ATR method based on linear interpolation in the middle layer proposed in the present invention can achieve accurate discrimination of unknown class targets in an open environment while basically ensuring the closed-set classification performance.
[0193] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0194] Although the present application has been described herein in connection with various embodiments, however, in the process of implementing the claimed present application, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results. The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for identifying an unknown type of target under open conditions, characterized in that, The target recognition method for the unknown type includes: Step 1: Obtain a first training target set, where the first training target set includes K types of known-type targets, and the first training target set is a SAR image training set; Step 2: Perform data augmentation preprocessing on the first training target set to obtain a second training target set; Step 3: Select n types of target data of size B from different types in the second training target set, and adjust the order of the n types of target data n times to obtain n input data with different permutation methods, so as to obtain an unknown target source array based on the splicing result of the n input data with different permutation methods, where 2 ≤ n ≤ K; Step 4: Input the unknown target source array into a classification neural network to process the unknown target source array by using random cropping and splicing or random linear interpolation to obtain an unknown-type target; Step 5: Input the unknown-type target and the second training target set into the classification neural network to be trained to train the classification neural network, and obtain a K + 1-class classification model after the loss function converges; Step 6: Iteratively input the validation data set into the K + 1-class classification model, and finally perform Softmax normalization processing on the output activation vector after passing through the fully connected layer to obtain an output probability score. The maximum value in the output probability scores corresponding to each data in the validation data set is used as the maximum probability output value of the data; sort the maximum probability output values corresponding to all data in the validation data set in ascending order, and select a maximum probability output value as the threshold ε1 from the sorted maximum probability output values based on the first preset position; sort the output probability scores under the K + 1th class corresponding to all data in the validation data set in ascending order, and select one as the threshold ε2 from the sorted output probability scores based on the second preset position; Step 7: Input the target to be recognized into the K + 1-class classification model to obtain the maximum probability output value of the target to be recognized, and determine the recognition result of the target to be recognized according to the relationship between the maximum probability output value of the target to be recognized and the threshold ε1 and the relationship between the output probability score of the target in the K + 1th category and the threshold ε2.
2. The method for identifying an unknown type of target under open conditions according to claim 1, wherein The classification neural network is obtained by adding a combination layer between two original combination layers in the middle position of the classifier32 classification neural network and adding a combination layer between the last two original combination layers. The combination layer includes a convolutional layer and an activation function layer.
3. The method for identifying an unknown type of target under open conditions according to claim 1, wherein Obtaining an unknown target source array based on the splicing result of the n input data with different permutation methods includes: Splicing the n input data with different permutation methods column by column to obtain a B×n array; Check each row of the B×n array to determine whether there is data of the same type in each row. If so, delete the data in that row. If not, retain the data in that row to obtain the unknown target source array.
4. The method for identifying an unknown type of target under open conditions according to claim 1, wherein The said Step 4 includes: Input the unknown target source array into the classification neural network to randomly crop and splice the unknown target source array at the input layer of the classification neural network to obtain an unknown type target, denoted as the first technical route; or, Input the unknown target source array into the classification neural network to perform random linear interpolation on the unknown target source array at the intermediate layer of the classification neural network to obtain an unknown type target, denoted as the second technical route, where the intermediate layer is the combined layer in the middle position of the classification neural network.
5. The method for identifying an unknown type of target under open conditions according to claim 4, wherein When n is even, input the unknown target source array into the classification neural network to randomly crop and splice the unknown target source array at the input layer of the classification neural network to obtain an unknown type target, including: S1.
1. First, randomly select 1 h and (n / 2 - 1) w according to the beta distribution to form (n / 2 - 1) coordinates, which are respectively denoted as (w1, h), (w2, h) to (w n / 2-1 , h); S1.
2. Taking the (n / 2 - 1) coordinates obtained in S1.1 as the center points, draw horizontal and vertical lines along the horizontal and vertical axes to divide the image into n cropping regions, and denote their shapes as v0(a0, b0), v1(a1, b1) to v n-1 (a n-1 , b n-1 ). v0(a0, b0) to v n / 2-1 (a n / 2-1 , b n / 2-1 ) are above the horizontal line, and v n / 2 (a n / 2 , b n / 2 ) to v n-1 (a n-1 , b n-1 ) are below the horizontal line, where a i and b i represent width and length respectively, and 0 ≤ i ≤ n - 1; S1.
3. For generating random numbers x i within the ranges of (0, a - a i ), (0, b - b i ), respectively, let (x i , y i ) represent the starting position of the cropping area of the i-th image. Taking (x i , y i ) as the upper left corner of the cropping area, crop the i-th image according to the shape of v i (a i , b i ) to obtain the cropped image v′ i (a i , b i , b i ); S1.
4. Concatenate the images v′0(a0,b0) to the image v′ n / 2-1 (a n / 2-1 ,b n / 2-1 ) by columns in the horizontal axis direction to obtain the image Concatenate the image v′ n / 2 (a n / 2 ,b n / 2 ) to the image v′ n-1 (a n-1 ,b n-1 ) by columns in the horizontal axis direction to obtain the image S1.
5. Concatenate Image and Image by rows along the vertical axis to obtain a target of unknown type.
6. The method for identifying an unknown type of target under open conditions according to claim 4, wherein When n is odd, input the unknown target source array into the classification neural network to randomly crop and splice the unknown target source array at the input layer of the classification neural network to obtain an unknown type target, including: S2.
1. First, randomly select 1 h and (n - 1) / 2 ws according to the beta distribution to form (n - 1) / 2 coordinates, which are respectively denoted as (w1, h), (w2, h) to (w (n-1) / 2 , h); S2.
2. Taking the (n - 1) / 2 coordinates obtained in S2.1 as the center points, draw vertical lines along the vertical axis, and draw horizontal lines along the horizontal axis within the range of (0, w (n-1) / 2 ), so as to divide the image into n cropping regions, and record their shapes as v0(a0, b0), v1(a1, b1) to v n-1 (a n-1 , b n-1 ), where v0(a0, b0) to v (n-3) / 2 (a (n-3) / 2 , b (n-3) / 2 ) are above the horizontal line, ν (n+1) / 2 (a (n+1) / 2 , b (n+1) / 2 ) to ν n-1 (a n-1 , b n-1 ) are below the horizontal line, v (n-1) / 2 (a (n-1) / 2 , b (n-1) / 2 ) is the region not divided by the horizontal line, a i and b i represent the length and width respectively, 0 ≤ i ≤ n - 1; S2.
3. For generating random numbers x i within the ranges of (0, a - a i ), (0, b - b i ), respectively, let (x i , y i ) represent the starting position of the cropping region of the i-th image. Taking (x i , y i ) as the upper left corner of the cropping region, crop the i-th image according to the shape of v i (a i , b i ) to obtain the cropped image v″ i (a i , b i , b i ); S2.
4. Concatenate the images v″0(a0, b0) to the image v″ (n-3) / 2 (a (n-3) / 2 , b (n-3) / 2 ) by columns along the horizontal axis to obtain the image Concatenate the image v″ (n+1) / 2 (a (n+1) / 2 , b (n+1) / 2 ) to the image v″ n-1 (a n-1 , b n-1 ) by columns along the horizontal axis to obtain the image S2.
5. Concatenate Image and Image row by row along the vertical axis to obtain Image S2.
6. Concatenate the image and the image v′ (n-1) / 2 (a (n-1) / 2 , b (n-1) / 2 ) column by column along the horizontal axis to obtain an unknown type target.
7. The method for identifying a target of an unknown type under open conditions according to claim 4, wherein The unknown type target obtained by random linear interpolation is expressed as: Among them, is an unknown type target, C1 to C n are weight coefficients, C1 + C2 + … + C n = 1, C1 to C n are random numbers that satisfy the beta distribution, are all the layers before the combination layer at the middle position of the classification neural network.
8. The method for identifying an unknown type of target under open conditions according to claim 1, characterized in that For the first technical route, step 5 includes: Step 5.1: Splice the unknown type target and the second training target set together and input them into the classification neural network to be trained to obtain the output feature distribution of the unknown type target and the output feature distribution of the known type target; Step 5.2: Continuously train the neural network to optimize the output feature distribution of the classification neural network to be trained until the first loss function converges, and obtain a K + 1 class classification model after the first loss function converges, where the first loss function is composed of the decision losses of all K + 1 class training targets; For the second technical route, step 5 includes: Step 5.1: Input the unknown type target into the second half of the classification neural network to obtain the output feature distribution of the unknown type target, and at the same time, input the second training target set into the complete classification neural network to be trained to obtain the output feature distribution of the known type target, where the second half of the classification neural network is the combined layer in the middle position to the fully connected layer; Step 5.2: Continuously train the neural network to optimize the output feature distribution of the classification neural network to be trained until the second loss function converges, and obtain a K + 1 class classification model after the second loss function converges, where the second loss function is composed of two parts: the decision loss of generating the unknown class and the classification loss of the known class.
9. The method for identifying an unknown type of target under open conditions according to claim 8, wherein The first loss function is: where, l t1 is the first loss function, y k is the true label of the input sample, l is the preset loss function, is the output feature distribution of all K + 1 types of training targets; The second loss function is: l t2 =l k +α u ·l u where l t2 is the second loss function, l k is the classification loss of the known type, y k is the true label of the input sample, l is the preset loss function, is the output feature distribution of the known type samples, is the output feature distribution of the unknown type samples, l u is the decision loss of the unknown type, α u is a hyperparameter.
10. The method for identifying an unknown type of target under open conditions according to claim 1, characterized in that Determine the prediction result of the target to be recognized according to the relationship between the maximum probability output value of the target to be recognized and the threshold ε1 and the relationship between the output probability score of the target on the K + 1st category and the threshold ε2, including: Determine the relationship between the maximum probability output value of the target to be recognized and the threshold ε1. If the maximum probability output value of the target to be recognized is less than or equal to the threshold ε1, the recognition result is an unknown target. If the maximum probability output value of the target to be recognized is greater than the threshold, then determine the relationship between the output probability score of the (K + 1)-th category and the threshold ε2. If the output probability score is greater than or equal to the threshold ε2, the recognition result is an unknown target; otherwise, the recognition result of the target is the type to which the maximum probability output value among the first K known categories belongs.
Citation Information
Patent Citations
Synthetic aperture radar target identification method based on auxiliary decision update learning
CN107886123A
Radar interference semi-supervised open set identification system based on generative adversarial network
CN114241263A