A Target Recognition Method and Apparatus Based on Multi-View SAR Image Feature Coding and Fusion
By constructing a recognition model that encodes and fuses multi-view SAR image features and using viewpoints and semantic labels to train a sample set, the problem of insufficient SAR image features under a single viewpoint is solved, thereby improving the accuracy of target recognition.
Patent Information
- Application Number
- CN202510219607.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-02-26
AI Technical Summary
In existing technologies, SAR image features based on a single viewpoint have low accuracy in target classification and recognition, and the SAR image features under different observation views vary greatly, resulting in insufficient recognition accuracy.
By constructing a recognition model and training sample set, and utilizing a multi-view SAR image feature encoding and fusion method, the recognition model is trained using a sample set with pre-labeled viewpoints and semantic labels. Prior viewpoint information is used to guide the network to learn multi-view fusion features, thereby improving recognition accuracy.
It improves the accuracy of target recognition, can output accurate recognition and classification results from different perspectives, and enhances the flexibility and accuracy of the recognition model.
Smart Images

Figure CN120047674B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of target recognition, and in particular relates to a target recognition method and device based on multi-view SAR image feature coding and fusion. BACKGROUND
[0002] In recent years, the synthetic aperture radar (SAR) image target classification recognition technology based on deep learning has developed rapidly, and good results have been achieved in the classification tasks of SAR images of vehicles, aircraft and ships. However, since the SAR system essentially measures the backscattering characteristics of the target under the current observation configuration, the SAR image features of the target under a single observation view may not be sufficient, and the SAR image features under different observation views may be quite different. The accuracy of target classification and recognition using only single-view SAR images is low. In view of the above problems, how to jointly observe the target from multiple views, obtain the multi-view SAR images of the target and perform fusion processing is the key to improving the accuracy of target classification and recognition.
[0003] Therefore, there is an urgent need for a target recognition method based on multi-view SAR image feature coding and fusion to solve the above problems. SUMMARY
[0004] The present application provides a target recognition method and device based on multi-view SAR image feature coding and fusion, which can improve the accuracy of target recognition. The technical solution is as follows:
[0005] In a first aspect, the present application provides a target recognition method based on multi-view SAR image feature coding and fusion, which comprises:
[0006] obtaining SAR images of a target to be recognized under different views;
[0007] inputting the SAR images of the target to be recognized under different views into a pre-trained recognition model to obtain a recognition result of the target to be recognized; the recognition model is obtained by training a pre-constructed sample set, the sample set is composed of SAR sample images of multiple known category targets under different views, and each sample image is labeled with a view label and a semantic label, so that the recognition model performs target recognition based on the features of different view images.
[0008] In a second aspect, the present application further provides a target recognition device based on multi-view SAR image feature coding and fusion, which comprises:
[0009] an acquisition unit configured to acquire SAR images of a target to be recognized under different views;
[0010] The recognition unit is configured to input the SAR images of the target to be recognized at different viewing angles into a pre-trained recognition model to obtain a recognition result of the target to be recognized. The recognition model is trained by a pre-constructed sample set, and the sample set is composed of SAR sample images of targets of known categories at different viewing angles. Each sample image is labeled with a viewing angle label and a semantic label, so that the recognition model performs target recognition based on the features of images at different viewing angles.
[0011] In a third aspect, an electronic device is provided, including a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the method described in any of the embodiments of the present specification.
[0012] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program, when executed in a computer, causes the computer to perform the method described in any of the embodiments of the present specification.
[0013] In a fifth aspect, a computer program product is provided, which includes a computer program, and the computer program, when executed by a processor, implements the steps of the method described above.
[0014] The embodiments of the present application provide a target recognition method based on multi-view SAR image feature coding and fusion. First, an identification model and a training sample set are constructed, the sample set is composed of SAR images of targets of different categories and different viewing angles, and each image is labeled with a viewing angle label and a semantic label. Therefore, training the identification model using the above sample set can effectively utilize the viewing angle prior information to guide the identification model to learn the multi-view fusion features of the SAR target, and improve the accuracy of target classification of the identification model. Finally, when the target to be recognized needs to be classified, only the SAR images of the target to be recognized at different angles need to be input into the trained identification model, and the accurate recognition classification result can be output. Therefore, the present application can improve the accuracy of target recognition. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings described below are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0016] Figure 1 is a flowchart of a target recognition method based on multi-view SAR image feature coding and fusion provided by an embodiment of the present application;
[0017] Figure 2 is a target recognition device structure diagram based on multi-view SAR image feature coding and fusion provided by an embodiment of the present application;
[0018] Figure 3 is a hardware architecture diagram of a computer device provided by an embodiment of the present application;
[0019] Figure 4 is a network architecture diagram of a recognition model provided by an embodiment of the present application;
[0020] Figure 5 is a multi-view SAR image schematic diagram provided by an embodiment of the present application;
[0021] Figure 6 is a loss function change curve diagram in a multi-view SAR image feature coding and fusion network model training process provided by an embodiment of the present application;
[0022] Figure 7 is a prediction accuracy change curve diagram in a multi-view SAR image feature coding and fusion network model training process provided by an embodiment of the present application;
[0023] Figure 8 is a multi-view SAR image recognition result schematic diagram provided by an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0025] The specific implementation of the above concept will be described below.
[0026] Referring to Figure 1 The embodiment of the present application provides a target recognition method based on multi-view SAR image feature coding and fusion, which comprises the following steps.
[0027] In step 100, SAR images of a target to be recognized under different views are acquired.
[0028] In step 102, the SAR images of the target to be identified at different angles are input into the pre-trained identification model to obtain the identification result of the target to be identified; the identification model is trained by a pre-constructed sample set, and the sample set is composed of SAR sample images of targets of different categories at different angles, and each sample image is labeled with an angle label and a semantic label, so that the identification model performs target identification based on the features of images at different angles.
[0029] In this embodiment, the identification model and the training sample set are first constructed, the sample set is composed of SAR images of different target categories and different angles, and each image is labeled with an angle label and a semantic label. Therefore, training the identification model using the above sample set can effectively utilize the angle prior information to guide the network model to learn the multi-angle fusion features of the SAR target, thereby improving the accuracy of target classification of the identification model. Finally, when it is necessary to classify the target to be identified, only the image tensor of the target to be identified at different angles needs to be input into the trained identification model, and an accurate identification classification result can be output. Therefore, the application can improve the accuracy of target identification.
[0030] The following describes Figure 1 the execution manner of each step.
[0031] First, for step 100:
[0032] The target to be identified can be a car, an airplane, etc., and the different angles can be a front view, a side view, an oblique view, etc. The application does not make specific limitations on the types of targets and the directions of angles. Users can determine them independently according to their needs. Of course, the type of the target to be identified is one of the types of known targets in the sample set.
[0033] Second, for step 102:
[0034] First, the construction process of the sample set is introduced:
[0035] In this step, a distributed SAR system or a circumferential SAR system can be used to collect SAR images of any target at multiple angles, and semantic labels and angle labels are labeled. Assuming that the i-th angle image is i = 1, 2,..., N v , C, H and W represent the channel number, height and width of the SAR image respectively, N v represents the number of angles; the semantic (i.e., category) label corresponding to the i-th angle image is denoted as c i , the angle label is denoted as v i , c i ∈0,1,...,N c -1, N c represents the number of semantic categories; v i ∈0,1,...,Nv -1.
[0036] Secondly, in some embodiments, as shown in the figure, the network structure of the recognition model includes: a feature extraction network, a semantic-perspective encoding network, a multi-perspective feature fusion network, a semantic classification head, a perspective classification head and a fusion classification head. Figure 4
[0037] Next, each sub-network will be introduced in detail.
[0038] (1) The feature extraction network includes a semantic feature extraction module and a perspective feature extraction module that are calculated in parallel; the feature extraction network takes SAR images under different perspectives of a target as input, and for each SAR image, the semantic feature extraction module is used to extract semantic features and output a semantic feature map, and the perspective feature extraction module is used to extract perspective features and output a perspective feature map.
[0039] In some embodiments, the semantic feature extraction module and the perspective feature extraction module are both constructed based on a convolutional neural network, and when extracting the corresponding feature maps, the following formula is used for calculation:
[0040] x SFi =f SF (I i )
[0041] x VFi =f VF (I i )
[0042] In the formula, x SFi and x VFi are the semantic feature map and the perspective feature map of the i-th perspective image of the target, C FM , H FM and W FM represent the channel number, height and width of the output feature map respectively; f SF (·) represents the semantic feature extraction module, and f VF (·) represents the perspective feature extraction module.
[0043] (2) The semantic-perspective encoding network takes the semantic feature map and the perspective feature map of each perspective SAR image output by the feature extraction network as input, and is used to calculate the semantic-perspective embedding of the target multi-perspective SAR image.
[0044] In some embodiments, the semantic-perspective encoding network includes a semantic linear mapping module, a perspective linear mapping module and a fusion module, and each linear mapping module is composed of two layers of convolutional neural networks;
[0045] The semantic linear mapping module is configured to perform channel compression and vectorization on each semantic feature map of the input to obtain a corresponding semantic embedding.
[0046] The view linear mapping module is configured to perform channel compression and vectorization on each view feature map of the input to obtain a corresponding view embedding.
[0047] The fusion module is configured to add the semantic embedding and the view embedding of each view SAR image respectively to obtain a semantic-view embedding of each view SAR image, and to concatenate each semantic-view embedding with a pre-constructed learnable class token to obtain a semantic-view embedding of the target multi-view SAR image.
[0048] The theoretical calculation formulas of the convolution layers of the two linear mapping modules are described in detail as follows:
[0049] (1) The expression of the first convolution layer of each linear mapping module is as follows:
[0050] x o1 =f conv1 (x i ;C FM ,C LP ,H LP ,W LP )
[0051] In the formula, x o1 represents the output of the first convolution layer, x i represents the input of the first convolution layer, f conv1 (·) represents the first convolution layer, and the convolution layer parameters are input channel number C FM , output channel number C LP , convolution kernel height H LP , and convolution kernel width W LP , respectively.
[0052] (2) The expression of the second convolution layer of each linear mapping module is as follows:
[0053] x o2 =f conv2 (x o1 ;C LP ,D,H FM ,W FM )
[0054] In the formula, x o2 represents the output of the second convolution layer, x o1 represents the input of the second convolution layer (i.e., the output of the first convolution layer), f conv2 (·) represents the second convolution layer, and the convolution layer parameters are input channel number C LP , output channel number D, convolution kernel height HFM Convolution kernel width W FM .
[0055] Based on the above calculation formula, the semantic linear mapping module calculates the semantic embedding of each semantic feature map using the following formula:
[0056] x SEi =f conv2 ((f conv1 (x SFi )))
[0057] The viewpoint linear mapping module calculates the viewpoint embedding for each viewpoint feature map using the following formula:
[0058] x VEi =f conv2 ((f conv1 (x VFi )))
[0059] In the formula, x SEi and x VEi Let represent the semantic embedding and viewpoint embedding of the i-th viewpoint image of the target, respectively.
[0060] In addition, the semantic-view embedding of each viewpoint SAR image is calculated using the following formula:
[0061] x SVEi =x SEi +x VEi
[0062] In the formula, x represents SVEi Semantic-viewpoint embedding of the target image at the i-th viewpoint.
[0063] The semantic-view embedding of the target multi-view SAR image is calculated using the following formula:
[0064]
[0065] In the formula, x cls Indicates a learnable category token. x FE Semantic-view embeddings representing multi-view SAR images of targets. N p =N v +1.
[0066] (III) Multi-view feature fusion network, which takes the semantic-view feature embedding output by the semantic-view coding network as input to calculate the multi-view fused semantic feature vector of the target.
[0067] In some embodiments, the multi-view feature fusion network is constructed based on a preset Transformer network, the Transformer network comprising a plurality of layers of encoders connected in sequence; wherein the first layer of encoder takes the semantic-view feature embedding output by the semantic-view encoding network as input, and the other layers of encoder take the feature embedding output by the previous layer of encoder as input; each layer of encoder is configured to perform the following work: performing layer normalization on the input feature embedding, performing multi-head self-attention calculation on the layer-normalized feature embedding, and adding the result of the multi-head self-attention calculation to the input feature embedding to obtain an intermediate feature embedding; performing layer normalization on the intermediate feature embedding, performing multi-layer perceptron calculation on the layer-normalized intermediate feature embedding, and adding the result of the multi-layer perceptron calculation to the intermediate feature embedding to obtain the feature embedding output by the layer of encoder;
[0068] The multi-view fusion semantic feature vector of the target is calculated in the following manner:
[0069] The semantic-view feature embedding output by the semantic-view encoding network is input into the Transformer network and sequentially calculated by each layer of encoder until the last layer of encoder to obtain the multi-view fusion feature; the feature vector corresponding to the learnable class token in the multi-view fusion feature is taken to obtain the multi-view fusion semantic feature vector of the target.
[0070] The calculation process of the l-th layer of encoder is described below, and the formula is as follows:
[0071]
[0072] In the formula, represents the intermediate feature embedding output by the l-th layer of encoder; represents the feature embedding output by the (l-1)-th layer of encoder; LN (·) represents layer normalization, MSA (·) represents multi-head self-attention calculation, MLP (·) represents multi-layer perceptron calculation, l = 1, 2,..., L, and L is the number of layers of the encoder; represents the feature embedding output by the l-th layer of encoder.
[0073] It should be noted that, for the last layer of encoder, the output is the multi-view fusion feature. After obtaining the multi-view fusion feature, the multi-view fusion semantic feature vector x Ecls of the target is calculated using the following formula:
[0074]
[0075] Since the learnable class token is located at position 0, the feature vector corresponding to the orientation vector index 0.
[0076] The semantic classification head, the perspective classification head and the fusion classification head are all constructed based on a multi-layer perceptron layer, and the classification heads are described in detail below.
[0077] (IV) Semantic classification head.
[0078] The semantic classification head takes the semantic feature maps of the perspective SAR images output by the semantic feature extraction module as input, and is used to calculate the semantic classification prediction probability of each perspective image. The calculation process is as follows:
[0079] Based on the semantic classification head, the multi-layer perceptron is calculated on the semantic feature maps of the perspective SAR images to obtain the single-perspective semantic classification feature vector of each perspective image. The single-perspective semantic classification feature vectors are normalized based on the softmax normalization exponential function to obtain the semantic classification prediction probability of each perspective image predicted as each class.
[0080] In some embodiments, the semantic classification prediction probability of each perspective image predicted as each class is calculated using the following formula:
[0081] z sci =f MLP (x SFi )
[0082]
[0083] In the formula, z sci represents the single-perspective semantic classification feature vector of the i-th perspective image, i = 1, 2, …, N v ; f MLP (·) represents multi-layer perceptron calculation; softmax(·) represents a normalization exponential function; Pj represents the semantic prediction probability of the i-th perspective image predicted as the j-th class; yj represents the j-th class label predicted, j = 1, 2, …, N c .
[0084] (V) Perspective classification head.
[0085] The perspective classification head takes the perspective feature maps of the perspective SAR images output by the perspective feature extraction module as input, and is used to calculate the perspective classification prediction probability of each perspective image. The calculation process is as follows:
[0086] The view classification head is used for performing multi-layer perceptron calculation on the view feature maps of each view SAR image to obtain a view classification feature vector of each view image, and performing normalization processing on the view classification feature vector based on a softmax normalization exponential function to obtain a view classification prediction probability of each view image being predicted as each view.
[0087] In some embodiments, the view classification prediction probability of each view image being predicted as each view is calculated using the following formula:
[0088] z svi =f MLP (x VFi )
[0089]
[0090] In the formula, z svi represents the view classification feature vector of the i-th view image; represents the prediction probability of the i-th view image being predicted as the h-th view; represents the predicted h-th view label, h = 1, 2, …, N v .
[0091] (VI) Fusion classification head.
[0092] The fusion classification head takes the multi-view fusion feature vector output by the multi-view feature fusion network as input, and is used for calculating a multi-view fusion classification prediction probability of the target. The calculation process is as follows:
[0093] The fusion classification head is used for performing multi-layer perceptron calculation on the multi-view fusion semantic feature vector of the target to obtain a multi-view fusion classification feature vector with a length equal to the number of semantic categories, and performing normalization processing on the multi-view fusion classification feature vector based on a softmax normalization exponential function to obtain a fusion classification prediction probability of the multi-view fusion classification feature vector being predicted as each category.
[0094] In some embodiments, the fusion classification prediction probability is calculated using the following formula:
[0095] z mc =f MLP (x Ecls )
[0096]
[0097] In the formula, z mc represents the multi-view fusion classification feature vector; represents the probability of being predicted as the j-th category based on the multi-view fusion classification feature vector; represents the predicted j-th category label, j = 1, 2, …, Nc .
[0098] In addition, the semantic classification prediction probability, the perspective classification prediction probability, and the multi-perspective fusion classification prediction probability are used to determine a hybrid loss function of the recognition model.
[0099] In some embodiments, the hybrid loss function is calculated by the following ways:
[0100] (1) Based on the semantic classification prediction probability of each perspective image of the target and the real class label, a semantic classification loss of the target is calculated, and the calculation formula is as follows:
[0101]
[0102] In the formula, L sc represents the semantic classification loss of the target; y cj represents the real jthclass label; when the real class of the ithperspective image is the jthclass, y cj takes 1, otherwise, y cj takes 0.
[0103] (2) Based on the perspective classification prediction probability of each perspective image of the target and the real perspective label, a perspective classification loss of the target is calculated, and the calculation formula is as follows:
[0104]
[0105] In the formula, L sv represents the perspective classification loss of the target; y vh represents the real hthperspective label, when the real perspective of the ithperspective image is the hthperspective, y vh takes 1, otherwise, y vh takes 0. represents the prediction probability that the hthperspective image is predicted as the hthperspective, z svh represents the perspective classification feature vector of the hthperspective image; represents the prediction probability that the ithperspective image is predicted as the hthperspective; m is a positive hyperparameter; i = 1, 2, …, N v , h = 1, 2, …, N v , and h≠i.
[0106] (3) Based on the multi-perspective fusion classification prediction probability of the target and the real class label, a fusion classification loss of the target is calculated, and the calculation formula is as follows:
[0107]
[0108] In the formula, L mc represents the fusion classification loss of the target; y cjdenotes the true j-th class label; y cj = 1, otherwise, y cj = 0.
[0109] (4) The semantic classification loss, the view classification loss and the fusion classification loss are weighted and summed to obtain a hybrid loss function Loss of the recognition model, and the calculation formula is as follows:
[0110] Loss = a sc L sc + a sv L sv + a mc L mc
[0111] In the formula, a sc , a sv and a mc are super parameter weights of the semantic classification loss, the view classification loss and the fusion classification loss respectively. The super parameter weights are determined according to user needs, and the present application does not make specific limitations.
[0112] In some embodiments, the pre-constructed recognition model is trained in the following manner:
[0113] The pre-constructed SAR sample images are divided into a training set and a test set;
[0114] The SAR sample images in the training set and the test set are used to train and test the recognition model until a preset convergence condition is reached, so as to obtain a trained recognition model; the convergence condition is that the number of training times reaches a preset number or the loss value of the hybrid loss function is less than a preset value.
[0115] After the recognition model is trained, it can be used for target recognition.
[0116] In order to prove the effectiveness of the method of the present application, the inventors verified the method of the present application with the following data. Among them, the number of multi-view angles is N v = 4, the number of SAR image channels C = 1, the image height H = 224, and the image width W = 224; then the i-th view image is i = 1, 2, 3, 4; the class classification label corresponding to the i-th view image is c i , c i ∈ 0, 1,..., 9, the number of semantic classes N c = 10, the view classification label corresponding to the i-th view image is v i , v i ∈ 0, 1, 2, 3. In addition, the semantic feature map x SFi extracted by the semantic feature extraction module and the view feature map xVFi , wherein the output feature map channel number is C FM = 2048, the height is H FM = 7, and the width is W FM = 7, and thus
[0117] For each linear mapping module, the parameters of the first convolutional layer are in turn the input channel number C FM = 2048, the output channel number C LP = 128, the convolution kernel height H LP = 1, the convolution kernel width W LP = 1. The parameters of the second convolutional layer are in turn the input channel number C LP = 128, the output channel number D = 768, the convolution kernel height H FM = 7, and the convolution kernel width W FM = 7. Thus, N p = 5, The Transformer network contains L = 6 layers of encoders. The hybrid loss function hyperparameters are respectively a sc = 0.25, a sv = 0.25, and a mc = 1.
[0118] On the basis of the above parameters, the training data shown in Table 1 is used, and the format of the sample image is as shown in Figure 5 :
[0119] Table 1: Semantic categories and target quantities of multi-view SAR images
[0120]
[0121]
[0122] The recognition model is trained using Figure 5 and the sample set shown in Table 1. The hybrid loss function change curve and the prediction accuracy change curve during the training process are shown in Figure 6 and Figure 7 respectively. As can be seen from the figures, the training loss gradually decreases with the training round number, and the target classification recognition accuracy gradually improves with the training round number.
[0123] The trained model is used to identify the target, and the identification result is shown in Figure 8 . As can be seen from the figure, under any order input, the method of the present application can obtain accurate target classification recognition results.
[0124] Therefore, the application improves the recognition accuracy by constructing the view feature extraction branch and the view classification loss and using the view prior information; the semantic-view feature embedding makes the test multi-view image be able to be input in any order, which is more flexible; the semantic-view feature embedding is fused by the Transformer model, which can effectively capture the correlation and complementarity of the target features of the multi-view image and improve the final recognition accuracy.
[0125] As shown in Figure 2 、 Figure 3 , the embodiment of the application provides a target recognition device based on multi-view SAR image feature coding and fusion. The device embodiment can be realized by software, or realized by hardware or a combination of software and hardware. From the hardware layer, as shown in Figure 2 , a hardware architecture diagram of a computing device where the target recognition device based on multi-view SAR image feature coding and fusion is provided by the embodiment of the application, in addition to the processor, memory, network interface, and non-volatile memory shown in Figure 2 , the computing device where the device in the embodiment usually also includes other hardware, such as a forwarding chip responsible for processing packets, etc. Taking the software implementation as an example, as shown in Figure 3 , as a logically meaningful device, it is formed by the CPU of the computing device where it is located reading the corresponding computer program in the non-volatile memory into the memory for running.
[0126] Please refer to Figure 3 , the embodiment of the application provides a target recognition device based on multi-view SAR image feature coding and fusion, and the device includes:
[0127] The acquisition unit 300 is configured to acquire SAR images of a target to be recognized under different views.
[0128] The recognition unit 302 is configured to input the SAR images of the target to be recognized under different views into a pre-trained recognition model to obtain a recognition result of the target to be recognized. The recognition model is obtained by training a pre-constructed sample set, and the sample set is composed of SAR sample images of targets of known categories under different views, and each sample image is labeled with a view label and a semantic label, so that the recognition model performs target recognition based on the features of different view images.
[0129] In some embodiments, the network structure of the recognition model includes a feature extraction network, a semantic-view coding network, a multi-view feature fusion network, a semantic classification head, a view classification head, and a fusion classification head.
[0130] The feature extraction network comprises a semantic feature extraction module and a view feature extraction module for parallel calculation; the feature extraction network takes SAR images of different views of a target as input, and for each input SAR image, the semantic feature extraction module is used to extract semantic features and output a semantic feature map, and the view feature extraction module is used to extract view features and output a view feature map;
[0131] The semantic-view encoding network takes the semantic feature map and the view feature map of each view SAR image output by the feature extraction network as input, and is used to calculate semantic-view embedding of the target multi-view SAR image.
[0132] The multi-view feature fusion network takes the semantic-view embedding output by the semantic-view encoding network as input, and is used to calculate the multi-view fusion semantic feature vector of the target.
[0133] The semantic classification head takes the semantic feature map of each view SAR image output by the semantic feature extraction module as input, and is used to calculate the semantic classification prediction probability of each view image; the view classification head takes the view feature map of each view SAR image output by the view feature extraction module as input, and is used to calculate the view classification prediction probability of each view image; the fusion classification head takes the multi-view fusion feature vector output by the multi-view feature fusion network as input, and is used to calculate the multi-view fusion classification prediction probability of the target; the semantic classification prediction probability, the view classification prediction probability and the multi-view fusion classification prediction probability are used to determine the hybrid loss function of the recognition model.
[0134] In some embodiments, the semantic-view encoding network comprises a semantic linear mapping module, a view linear mapping module and a fusion module, and each linear mapping module is composed of two layers of convolutional neural networks.
[0135] The semantic linear mapping module is used to compress the channel and vectorize each input semantic feature map to obtain a corresponding semantic embedding.
[0136] The view linear mapping module is used to compress the channel and vectorize each input view feature map to obtain a corresponding view embedding.
[0137] The fusion module is used to add the semantic embedding and the view embedding of each view SAR image respectively to obtain the semantic-view embedding of each view SAR image, and to splice each semantic-view embedding with a pre-constructed learnable class token to obtain the semantic-view embedding of the target multi-view SAR image.
[0138] In some embodiments, the multi-view feature fusion network is constructed based on a preset Transformer network, the Transformer network comprising a plurality of layers of encoders connected in sequence; wherein the first layer of encoder takes the semantic-view feature embedding output by the semantic-view encoding network as input, and the other layers of encoder take the feature embedding output by the previous layer of encoder as input; each layer of encoder is configured to perform the following work: performing layer normalization on the input feature embedding, performing multi-head self-attention calculation on the layer-normalized feature embedding, and adding the result of the multi-head self-attention calculation to the input feature embedding to obtain an intermediate feature embedding; performing layer normalization on the intermediate feature embedding, performing multi-layer perceptron calculation on the layer-normalized intermediate feature embedding, and adding the result of the multi-layer perceptron calculation to the intermediate feature embedding to obtain the feature embedding output by the layer of encoder;
[0139] The multi-view fusion semantic feature vector of the target is calculated in the following manner:
[0140] The semantic-view feature embedding output by the semantic-view encoding network is input into the Transformer network, and is sequentially calculated by each layer of encoder until the last layer of encoder to obtain a multi-view fusion feature; the feature vector corresponding to the learnable class token in the multi-view fusion feature is taken to obtain the multi-view fusion semantic feature vector of the target.
[0141] In some embodiments, the semantic classification head, the view classification head and the fusion classification head are all constructed based on a multi-layer perceptron layer;
[0142] The calculation process of the semantic classification prediction probability is: performing multi-layer perceptron calculation on the semantic feature map of each view SAR image based on the semantic classification head to obtain a single-view semantic classification feature vector of each view image, and performing normalization processing on each single-view semantic classification feature vector based on a softmax normalization exponential function to obtain the semantic classification prediction probability of each view image being predicted as each class;
[0143] The calculation process of the view classification prediction probability is: performing multi-layer perceptron calculation on the view feature map of each view SAR image based on the view classification head to obtain a view classification feature vector of each view image, and performing normalization processing on each view classification feature vector based on a softmax normalization exponential function to obtain the view classification prediction probability of each view image being predicted as each view;
[0144] The calculation process of the multi-view fusion classification prediction probability is: based on the multi-layer perceptron calculation of the multi-view fusion semantic feature vector of the target by the fusion classification head, a multi-view fusion classification feature vector with a length equal to the number of semantic categories is obtained, and the multi-view fusion classification feature vector is normalized by a softmax normalization exponential function to obtain the fusion classification prediction probability of the multi-view fusion classification feature vector being predicted as each category.
[0145] In some embodiments, the hybrid loss function is calculated by the following way:
[0146] Based on the semantic classification prediction probability of each view image of the target and the real category label, the semantic classification loss of the target is calculated.
[0147] Based on the view classification prediction probability of each view image of the target and the real view label, the view classification loss of the target is calculated.
[0148] Based on the multi-view fusion classification prediction probability of the target and the real category label, the fusion classification loss of the target is calculated.
[0149] The semantic classification loss, the view classification loss and the fusion classification loss are weighted and summed to obtain the hybrid loss function of the recognition model.
[0150] In some embodiments, the pre-constructed recognition model is trained by the following way:
[0151] The pre-constructed SAR sample images are divided into a training set and a test set;
[0152] The SAR sample images in the training set and the test set are used to train and test the recognition model until a preset convergence condition is reached, and a trained recognition model is obtained; the convergence condition is that the number of training times reaches a preset number or the loss value of the hybrid loss function is less than a preset value.
[0153] It should be noted that: the target recognition device based on multi-view SAR image feature coding and fusion provided in the above embodiments is only exemplified by the division of the above functional modules, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the target recognition device based on multi-view SAR image feature coding and fusion provided in the above embodiments and the target recognition method based on multi-view SAR image feature coding and fusion belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0154] Embodiments of the present application also provide a computer device, please refer to Figure 3The computer device includes a processor and a memory, and the memory stores at least one instruction, at least one program, a code set or an instruction set, which are loaded and executed by the processor to implement the target recognition method based on multi-view SAR image feature coding and fusion provided in the above method embodiments.
[0155] Embodiments of the present application also provide a computer readable storage medium, which stores at least one instruction, at least one program, a code set or an instruction set, which are loaded and executed by the processor to implement the target recognition method based on multi-view SAR image feature coding and fusion provided in the above method embodiments.
[0156] Embodiments of the present application also provide a computer program product, which includes a computer program, and the processor of the computer device reads the computer program from the computer readable storage medium, and the processor executes the computer program to enable the computer device to execute the target recognition method based on multi-view SAR image feature coding and fusion provided in any of the above embodiments.
[0157] For the convenience of description, the above system or device is described in various modules or units in terms of functions. Of course, the functions of each unit can be implemented in the same or multiple software and / or hardware in the implementation of the present application.
[0158] From the above description of the embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and necessary general hardware platforms. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0159] Finally, it needs to be pointed out that, in this document, relational terms such as first, second, third, and fourth and the like can only be used to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual relationship or order between or among such entities or actions. Moreover, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0160] The above description is merely preferred embodiments of the present application, and it is obvious to those skilled in the art that, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should also be considered as falling within the scope of the present application.
Claims
1. A target recognition method based on multi-view SAR image feature coding and fusion, characterized in that, The method comprises: acquiring SAR images of a target to be identified at different viewing angles; inputting the SAR images of the target to be identified at different viewing angles into a pre-trained identification model to obtain an identification result of the target to be identified; the identification model is trained by a pre-constructed sample set, the sample set is composed of SAR sample images of multiple targets of known categories at different viewing angles, and each sample image is labeled with a viewing angle label and a semantic label, so that the identification model performs target identification based on features of images at different viewing angles; the network structure of the identification model comprises a feature extraction network, a semantic-viewing angle encoding network, a multi-view feature fusion network, a semantic classification head, a viewing angle classification head, and a fusion classification head; the feature extraction network comprises a semantic feature extraction module and a viewing angle feature extraction module that are calculated in parallel; the feature extraction network takes SAR images of a target at different viewing angles as input, and for each input SAR image, semantic features of the SAR image are extracted based on the semantic feature extraction module and semantic feature maps are output, and viewing angle features of the SAR image are extracted based on the viewing angle feature extraction module and viewing angle feature maps are output; the semantic-viewing angle encoding network takes the semantic feature maps and the viewing angle feature maps of each viewing angle SAR image output by the feature extraction network as input, and is used to calculate semantic-viewing angle embeddings of the target multi-view SAR image; the multi-view feature fusion network takes the semantic-viewing angle feature embeddings output by the semantic-viewing angle encoding network as input, and is used to calculate a multi-view fusion semantic feature vector of the target; the semantic classification head takes the semantic feature maps of each viewing angle SAR image output by the semantic feature extraction module as input, and is used to calculate semantic classification prediction probabilities of each viewing angle image; the viewing angle classification head takes the viewing angle feature maps of each viewing angle SAR image output by the viewing angle feature extraction module as input, and is used to calculate viewing angle classification prediction probabilities of each viewing angle image; the fusion classification head takes the multi-view fusion feature vector output by the multi-view feature fusion network as input, and is used to calculate a multi-view fusion classification prediction probability of the target; and the semantic classification prediction probability, the viewing angle classification prediction probability, and the multi-view fusion classification prediction probability are used to determine a hybrid loss function of the identification model.
2. The method of claim 1, wherein, the semantic-viewing angle encoding network comprises a semantic linear mapping module, a viewing angle linear mapping module, and a fusion module, and each linear mapping module is composed of two layers of convolutional neural networks; the semantic linear mapping module is used to perform channel compression and vectorization on each input semantic feature map to obtain a corresponding semantic embedding; the viewing angle linear mapping module is used to perform channel compression and vectorization on each input viewing angle feature map to obtain a corresponding viewing angle embedding; the fusion module is used to add the semantic embedding and the viewing angle embedding of each viewing angle SAR image respectively to obtain a semantic-viewing angle embedding of each viewing angle SAR image; and each semantic-viewing angle embedding is spliced with a pre-constructed learnable category token respectively to obtain a semantic-viewing angle embedding of the target multi-view SAR image.
3. The method of claim 2, wherein, The multi-view feature fusion network is constructed based on a preset Transformer network, the Transformer network comprising a plurality of layers of encoders connected in sequence; wherein the first layer of encoder takes the semantic-view feature embedding output by the semantic-view encoding network as input, and the other layers of encoder take the feature embedding output by the previous layer of encoder as input; each layer of encoder is configured to perform the following work: performing layer normalization on the input feature embedding, performing multi-head self-attention calculation on the layer-normalized feature embedding, and adding the result of the multi-head self-attention calculation to the input feature embedding to obtain an intermediate feature embedding; performing layer normalization on the intermediate feature embedding, performing multi-layer perceptron calculation on the layer-normalized intermediate feature embedding, and adding the result of the multi-layer perceptron calculation to the intermediate feature embedding to obtain the feature embedding output by the layer of encoder; The multi-view fusion semantic feature vector of the target is calculated in the following manner: The semantic-view feature embedding output by the semantic-view encoding network is input into the Transformer network, and is sequentially calculated by each layer of encoder until the last layer of encoder to obtain a multi-view fusion feature; the feature vector corresponding to the learnable class token in the multi-view fusion feature is taken to obtain the multi-view fusion semantic feature vector of the target.
4. The method of claim 1, wherein, The semantic classification head, the view classification head and the fusion classification head are all constructed based on a plurality of layers of perceptron; The calculation process of the semantic classification prediction probability is: performing multi-layer perceptron calculation on the semantic feature map of each view SAR image based on the semantic classification head to obtain a single-view semantic classification feature vector of each view image, and performing normalization processing on each single-view semantic classification feature vector based on a softmax normalization exponential function to obtain a semantic classification prediction probability of each view image predicted as each class; The calculation process of the view classification prediction probability is: performing multi-layer perceptron calculation on the view feature map of each view SAR image based on the view classification head to obtain a view classification feature vector of each view image, and performing normalization processing on each view classification feature vector based on a softmax normalization exponential function to obtain a view classification prediction probability of each view image predicted as each view; The calculation process of the multi-view fusion classification prediction probability is: performing multi-layer perceptron calculation on the multi-view fusion semantic feature vector of the target based on the fusion classification head to obtain a multi-view fusion classification feature vector with a length equal to the number of semantic classes, and performing normalization processing on the multi-view fusion classification feature vector based on a softmax normalization exponential function to obtain a fusion classification prediction probability of the multi-view fusion classification feature vector predicted as each class.
5. The method of claim 4, wherein, The hybrid loss function is calculated in the following manner: calculating a semantic classification loss of the target based on the semantic classification prediction probability of each view image of the target and the real class label; calculating a view classification loss of the target based on the view classification prediction probability of each view image of the target and the real view label; and calculating a fusion classification loss of the target based on the multi-view fusion classification prediction probability of the target and the real class label. The target-based multi-view fusion classification prediction probability and the real class label are fused to calculate a fusion classification loss of the target; The semantic classification loss, the view classification loss and the fusion classification loss are weighted and summed to obtain a hybrid loss function of the recognition model.
6. The method of claim 1, wherein, The pre-constructed recognition model is trained in the following manner: The pre-constructed SAR sample images are divided into a training set and a test set; The SAR sample images in the training set and the test set are used to train and test the recognition model until a preset convergence condition is reached, thereby obtaining a trained recognition model; the convergence condition is that the number of training reaches a preset number or the loss value of the hybrid loss function is less than a preset value.
7. A target recognition device based on multi-view SAR image feature coding and fusion, characterized in that, The device comprises: An acquisition unit configured to acquire SAR images of a target to be identified under different views; An identification unit configured to input the SAR images of the target to be identified under different views into a pre-trained recognition model to obtain an identification result of the target to be identified; the recognition model is trained by a pre-constructed sample set, the sample set is composed of SAR sample images of targets of known categories under different views, and each sample image is labeled with a view label and a semantic label, so that the recognition model performs target identification based on features of different view images; The network structure of the recognition model comprises a feature extraction network, a semantic-view encoding network, a multi-view feature fusion network, a semantic classification head, a view classification head and a fusion classification head; The feature extraction network comprises a semantic feature extraction module and a view feature extraction module that are calculated in parallel; the feature extraction network takes SAR images of a target under different views as input, and for each input SAR image, semantic features thereof are extracted based on the semantic feature extraction module and semantic feature maps are output, and view features thereof are extracted based on the view feature extraction module and view feature maps are output; The semantic-view encoding network takes the semantic feature maps and the view feature maps of each view SAR image output by the feature extraction network as input, and is configured to calculate semantic-view embeddings of the target multi-view SAR images; The multi-view feature fusion network takes the semantic-view feature embeddings output by the semantic-view encoding network as input, and is configured to calculate a multi-view fusion semantic feature vector of the target; The semantic classification head takes the semantic feature maps of each view SAR image output by the semantic feature extraction module as input, and is configured to calculate semantic classification prediction probabilities of the view images; the view classification head takes the view feature maps of each view SAR image output by the view feature extraction module as input, and is configured to calculate view classification prediction probabilities of the view images; the fusion classification head takes the multi-view fusion feature vector output by the multi-view feature fusion network as input, and is configured to calculate a multi-view fusion classification prediction probability of the target; the semantic classification prediction probability, the view classification prediction probability and the multi-view fusion classification prediction probability are used to determine a hybrid loss function of the recognition model.
8. A computing device comprising a memory having stored therein a computer program and a processor, wherein the processor, when executing the computer program, implements the method of any one of claims 1-6.
9. A computer readable storage medium having stored thereon a computer program, which, when executed in a computer, causes the computer to perform the method of any one of claims 1-6.