Target identification method and device based on multi-view SAR image feature coding and fusion

By acquiring and fusing SAR image features from different perspectives, a recognition model is constructed for target recognition, which solves the problem of insufficient features of SAR images in a single-view angle and large differences in perspectives, and improves the accuracy of target recognition.

CN120047674AActive Publication Date: 2025-05-27BEIHANG UNIV

Patent Information

Application Number
CN202510219607.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-27
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The feature expression of single-view SAR images is insufficient and the SAR images have large differences in characteristics at different perspectives, resulting in a low accuracy of target classification recognition.

Method used

By obtaining the SAR images of the target to be identified at different perspectives, and building a pre-trained recognition model, the target recognition is performed using multi-view SAR image feature encoding and fusion methods. The method includes steps such as feature extraction, semantic-view coding, multi-view feature fusion, etc., and uses the Transformer network to achieve the fusion of multi-view features.

Benefits of technology

It improves the accuracy of target recognition, can effectively utilize the prior information of perspective, improves the accuracy of target classification of the recognition model, and supports the input of multi-view images in any order.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047674A_ABST
    Figure CN120047674A_ABST
Patent Text Reader

Abstract

The invention discloses a target identification method and device based on multi-view SAR image feature coding and fusion. The method comprises the following steps: acquiring SAR images of a to-be-recognized target at different visual angles; inputting the SAR images of the to-be-recognized target under different visual angles into a pre-trained recognition model to obtain a recognition result of the to-be-recognized target; the recognition model is obtained by training a pre-constructed sample set, the sample set is composed of SAR sample images of a plurality of targets of known categories under different visual angles, and each sample image is labeled with a visual angle label and a semantic label, so that the recognition model performs target recognition based on features of images of different visual angles. According to the invention, the accuracy of target identification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target recognition, and particularly to a target recognition method and device based on multi-view SAR image feature encoding and fusion. Background Art

[0002] In recent years, the target classification and recognition technology for synthetic aperture radar (SAR) images based on deep learning has developed rapidly, and good results have been achieved in target classification tasks such as vehicles, airplanes, and ships in SAR images. However, since the SAR system essentially measures the backscattering characteristics of the target under the current observation configuration, the SAR image features of the target under a single observation view may not be fully expressed, and there may be significant differences in the SAR image features under different observation views. The accuracy of target classification and recognition using only single-view SAR images is relatively low. To address the above problems, how to perform multi-view joint observation on the target, obtain multi-view SAR images of the target, and perform fusion processing is the key to improving the accuracy of target classification and recognition.

[0003] Based on this, there is an urgent need for a target recognition method based on multi-view SAR image feature encoding and fusion to solve the above problems. Summary of the Invention

[0004] The present invention provides a target recognition method and device based on multi-view SAR image feature encoding and fusion, which can improve the accuracy of target recognition. The technical solutions are as follows:

[0005] In a first aspect, an embodiment of the present invention provides a target recognition method based on multi-view SAR image feature encoding and fusion, the method comprising:

[0006] Obtaining SAR images of the target to be recognized under different views;

[0007] Inputting the SAR images of the target to be recognized under different views into a pre-trained recognition model to obtain the recognition result of the target to be recognized; the recognition model is obtained by training with a pre-constructed sample set, and the sample set consists of SAR sample images of multiple known-class targets under different views, and each sample image is labeled with a view label and a semantic label, so that the recognition model performs target recognition based on the features of different-view images.

[0008] In a second aspect, an embodiment of the present invention further provides a target recognition device based on multi-view SAR image feature encoding and fusion, the device comprising:

[0009] An obtaining unit, configured to obtain SAR images of the target to be recognized under different views;

[0010] An identification unit for inputting SAR images of the target to be identified from different perspectives into a pre-trained identification model to obtain an identification result of the target to be identified; the identification model is obtained by training with a pre-constructed sample set, and the sample set consists of SAR sample images of multiple known-class targets from different perspectives, and each sample image is labeled with a perspective label and a semantic label, so that the identification model can perform target identification based on the features of images from different perspectives.

[0011] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the method described in any embodiment of this specification is implemented.

[0012] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in any embodiment of this specification.

[0013] In a fifth aspect, an embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method described above are implemented.

[0014] An embodiment of the present invention provides a target recognition method based on multi-perspective SAR image feature coding and fusion. First, an identification model and a training sample set are constructed. The sample set consists of SAR images of different-class targets from different perspectives, and each image is labeled with a perspective label and a semantic label. Therefore, training the identification model with the above sample set can effectively utilize the perspective prior information to guide the identification model to learn the multi-perspective fusion features of SAR targets and improve the accuracy of target classification of the identification model. Finally, when it is necessary to classify the target to be identified, only the SAR images of the target to be identified from different angles need to be input into the trained identification model, and the accurate identification classification result can be output. Thus, it can be seen that this application can improve the accuracy of target recognition. Description of the Drawings

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a flowchart of a target recognition method based on multi-perspective SAR image feature coding and fusion provided by an embodiment of the present invention;

[0017] Figure 2 It is a structural diagram of an object recognition device provided by an embodiment of the present invention based on multi - perspective SAR image feature encoding and fusion;

[0018] Figure 3 It is a hardware architecture diagram of a computer device provided by an embodiment of the present invention;

[0019] Figure 4 It is a network architecture diagram of an identification model provided by an embodiment of the present invention;

[0020] Figure 5 It is a schematic diagram of the target multi - perspective SAR image provided by an embodiment of the present invention;

[0021] Figure 6 It is a curve graph of the change of the loss function during the training process of the multi - perspective SAR image feature encoding and fusion network model provided by an embodiment of the present invention;

[0022] Figure 7 It is a curve graph of the change of the prediction accuracy during the training process of the multi - perspective SAR image feature encoding and fusion network model provided by an embodiment of the present invention;

[0023] Figure 8 It is a schematic diagram of the multi - perspective SAR image recognition result provided by an embodiment of the present invention. Detailed implementation manners

[0024] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0025] The following describes the specific implementation manners of the above concepts.

[0026] Please refer to Figure 1 , a target recognition method based on multi - perspective SAR image feature encoding and fusion provided by an embodiment of the present invention, the method includes:

[0027] Step 100, obtaining SAR images of the target to be recognized from different perspectives;

[0028] Step 102: Input the SAR images of the target to be recognized from different perspectives into a pre-trained recognition model to obtain the recognition result of the target to be recognized; the recognition model is obtained by training with a pre-constructed sample set, and the sample set consists of SAR sample images of multiple known-class targets from different perspectives. Each sample image is labeled with a perspective label and a semantic label, so that the recognition model can perform target recognition based on the features of images from different perspectives.

[0029] In this embodiment, first, a recognition model and a training sample set are constructed. The sample set consists of SAR images of different target categories and different perspectives, and each image is labeled with a perspective label and a semantic label. Therefore, training the recognition model with the above sample set can effectively utilize the perspective prior information to guide the network model to learn the multi-perspective fusion features of SAR targets and improve the accuracy of target classification of the recognition model. Finally, when it is necessary to classify the target to be recognized, only the image tensors of the target to be recognized from different angles need to be input into the trained recognition model, and the accurate recognition classification result can be output. Thus, it can be seen that this application can improve the accuracy of target recognition.

[0030] The following describes Figure 1 the execution manners of the steps shown.

[0031] First, for step 100:

[0032] The target to be recognized can be a car, an airplane, etc., and different perspectives can be a front view, a side view, an oblique view, etc. This application does not make specific limitations on the types of targets and the perspective directions. Users can determine them independently according to needs. Of course, the type of the target to be recognized is one of the types of known targets in the sample set.

[0033] Second, for step 102:

[0034] First, the construction process of the sample set is introduced:

[0035] In this step, a distributed SAR system or a circumferential SAR system can be used to collect SAR images of any target from multiple perspectives and label the semantic labels and perspective labels. Assume that the i-th perspective image is C, H, and W respectively represent the number of channels, height, and width of the SAR image, and N v represents the number of perspectives; the semantic (i.e., category) label corresponding to the i-th perspective image is denoted as c i , and the perspective label is denoted as v i , c i ∈ {0, 1,..., N c -1}, N c represents the number of semantic categories; v i ∈ {0, 1,…, N v -1}.

[0036] Secondly, in some embodiments, such as Figure 4 shown, the network structure of the recognition model includes: a feature extraction network, a semantic-perspective encoding network, a multi-perspective feature fusion network, a semantic classification head, a perspective classification head, and a fusion classification head.

[0037] Next, each sub-network will be introduced in detail.

[0038] (1) Feature extraction network, which includes a semantic feature extraction module and a perspective feature extraction module that perform parallel computing; the feature extraction network takes SAR images of the target from different perspectives as input. For each input SAR image, its semantic features are extracted based on the semantic feature extraction module and a semantic feature map is output, and its perspective features are extracted based on the perspective feature extraction module and a perspective feature map is output.

[0039] In some embodiments, both the semantic feature extraction module and the perspective feature extraction module are constructed based on a convolutional neural network. When extracting the corresponding feature maps, both modules use the following formula for calculation:

[0040] x SFi = f SF (I i )

[0041] x VFi = f VF (I i )

[0042] In the formula, x SFi and x VFi are respectively the semantic feature map and the perspective feature map of the target's i-th perspective image, C FM , H FM and W FM respectively represent the number of channels, height, and width of the output feature map; f SF (·) represents the semantic feature extraction module, and f VF (·) represents the perspective feature extraction module.

[0043] (2) Semantic-perspective encoding network, which takes the semantic feature maps and perspective feature maps of each perspective SAR image output by the feature extraction network as input, and is used to calculate the semantic-perspective embedding of the target multi-perspective SAR image.

[0044] In some embodiments, the semantic-perspective encoding network includes a semantic linear mapping module, a perspective linear mapping module, and a fusion module, and each linear mapping module consists of two layers of convolutional neural networks;

[0045] The semantic linear mapping module is used to perform channel compression and vectorization on each input semantic feature map to obtain the corresponding semantic embedding;

[0046] The perspective linear mapping module is used to perform channel compression and vectorization on each input perspective feature map to obtain the corresponding perspective embedding;

[0047] The fusion module is used to add the semantic embedding and the perspective embedding of each perspective SAR image respectively to obtain the semantic-perspective embedding of each perspective SAR image; and to concatenate each semantic-perspective embedding with the pre-constructed learnable class token respectively to obtain the semantic-perspective embedding of the target multi-perspective SAR image.

[0048] The following details the theoretical calculation formulas of each convolutional layer of the two linear mapping modules:

[0049] (1) The expression of the first convolutional layer of each linear mapping module is:

[0050] x o1 =f conv2 (x i ;C LP ,D,H FM ,W FM )

[0051] In the formula, x o1 represents the output of the first convolutional layer, x i represents the input of the first convolutional layer, f conv1 (·) represents the first convolutional layer, and the convolutional layer parameters are the number of input channels C FM , the number of output channels C LP , the height H of the convolutional kernel LP , and the width W of the convolutional kernel LP .

[0052] (2) The expression of the second convolutional layer of each linear mapping module is:

[0053] x o2 =f conv2 (x o1 ;C LP ,D,H FM ,W FM )

[0054] In the formula, x o2 represents the output of the second convolutional layer, x o1 represents the input of the second convolutional layer (i.e., the output of the first convolutional layer), f conv2 (·) represents the second convolutional layer, and the convolutional layer parameters are the number of input channels C LP , the number of output channels D, the height H of the convolutional kernel FM, the width W of the convolutional kernel FM .

[0055] Based on the above calculation formula, the semantic linear mapping module calculates the semantic embedding of each semantic feature map through the following formula:

[0056] x SEi = f conv2 ((f conv1 (x SFi )))

[0057] The perspective linear mapping module calculates the perspective embedding of each perspective feature map through the following formula:

[0058] x VEi = f conv2 ((f conv1 (x VFi )))

[0059] In the formula, x SEi and x VEi respectively represent the semantic embedding and perspective embedding of the i-th perspective image of the target.

[0060] In addition, the semantic-perspective embedding of each perspective SAR image is calculated using the following formula:

[0061] x SVEi = x SEi + x VEi

[0062] In the formula, it represents x SVEi the semantic-perspective embedding of the i-th perspective image of the target.

[0063] The semantic-perspective embedding of the target multi-perspective SAR image is calculated using the following formula:

[0064]

[0065] In the formula, x cls represents the learnable class token. x FE represents the semantic-perspective embedding of the target multi-perspective SAR image. N p = N v + 1.

[0066] (III) Multi-perspective feature fusion network, which takes the semantic-perspective feature embedding output by the semantic-perspective encoding network as input and is used to calculate the multi-perspective fusion semantic feature vector of the target.

[0067] In some embodiments, the multi-view feature fusion network is constructed based on a preset Transformer network, and the Transformer network includes multiple layers of encoders connected in sequence; among them, the first layer of encoder takes the semantic-view feature embedding output by the semantic-view encoding network as input, and other layers of encoders take the feature embedding output by the previous layer of encoder as input; each layer of encoder is used to perform the following operations: perform layer normalization on the input feature embedding, perform multi-head self-attention calculation on the layer-normalized feature embedding, and add the result of the multi-head self-attention calculation to the input feature embedding to obtain an intermediate feature embedding; perform layer normalization on the intermediate feature embedding, perform multi-layer perceptron calculation on the layer-normalized intermediate feature embedding, and add the result of the multi-layer perceptron calculation to the intermediate feature embedding to obtain the feature embedding output by this layer of encoder;

[0068] The multi-view fusion semantic feature vector of the target is calculated through the following method:

[0069] Input the semantic-view feature embedding output by the semantic-view encoding network into the Transformer network, and perform calculations through each layer of encoder in sequence until the last layer of encoder to obtain the multi-view fusion feature; take the feature vector corresponding to the learnable class token in the multi-view fusion feature to obtain the multi-view fusion semantic feature vector of the target.

[0070] The following takes the l-th layer of encoder as an example to illustrate its calculation process, and the formula is as follows:

[0071]

[0072] In the formula, represents the intermediate feature embedding output by the l-th layer of encoder; represents the feature embedding output by the (l - 1)-th layer of encoder; f LN (·) represents layer normalization, f MSA (·) represents multi-head self-attention calculation, f MLP (·) represents multi-layer perceptron calculation, l = 1, 2,..., L, where L is the number of layers of the encoder; represents the feature embedding output by the l-th layer of encoder.

[0073] It should be noted that for the last layer of encoder, its output is the multi-view fusion feature. After obtaining , the following formula is used to calculate the multi-view fusion semantic feature vector of the target

[0074]

[0075] Since the position of the learnable class token is 0, the feature vector corresponding to vector index 0 is taken.

[0076] For the semantic classification head, the perspective classification head, and the fusion classification head, they are all constructed based on multi-layer perceptron layers. The following is a detailed introduction to each classification head.

[0077] (IV) Semantic classification head.

[0078] The semantic classification head takes the semantic feature maps of SAR images of each perspective output by the semantic feature extraction module as input, and is used to calculate the semantic classification prediction probabilities of images of each perspective. The calculation process is as follows:

[0079] Based on the semantic classification head, multi-layer perceptron calculations are performed on the semantic feature maps of SAR images of each perspective to obtain the single-perspective semantic classification feature vectors of images of each perspective, and the softmax normalization exponential function is used to normalize each single-perspective semantic classification feature vector to obtain the semantic classification prediction probabilities of images of each perspective predicted as various categories.

[0080] In some embodiments, the semantic classification prediction probabilities of images of each perspective predicted as various categories are calculated using the following formula:

[0081] z sci = f MLP (x SFi )

[0082]

[0083] In the formula, z sci represents the single-perspective semantic classification feature vector of the i-th perspective image, i = 1, 2,..., N v ; f MLP (·) represents multi-layer perceptron calculation; softmax(·) represents the normalization exponential function; represents the semantic prediction probability that the i-th perspective image is predicted as the j-th category; represents the predicted j-th category label, j = 1, 2,..., N c .

[0084] (V) Perspective classification head.

[0085] The perspective classification head takes the perspective feature maps of SAR images of each perspective output by the perspective feature extraction module as input, and is used to calculate the perspective classification prediction probabilities of images of each perspective. The calculation process is as follows:

[0086] Based on the perspective classification head, multi-layer perceptron calculations are performed on the perspective feature maps of SAR images of each perspective to obtain the perspective classification feature vectors of images of each perspective, and the softmax normalization exponential function is used to normalize each perspective classification feature vector to obtain the perspective classification prediction probabilities of images of each perspective predicted as each perspective.

[0087] In some embodiments, the perspective classification prediction probability of each perspective image is calculated using the following formula:

[0088] z svi = f MLP (x VFi )

[0089]

[0090] In the formula, z svi represents the perspective classification feature vector of the i-th perspective image; represents the prediction probability that the i-th perspective image is predicted as the h-th perspective; represents the predicted h-th perspective label, h = 1, 2,..., N v .

[0091] (VI) Fusion classification head.

[0092] The fusion classification head takes the multi-perspective fusion feature vector output by the multi-perspective feature fusion network as input and is used to calculate the multi-perspective fusion classification prediction probability of the target. The calculation process is as follows:

[0093] Based on the fusion classification head, a multi-layer perceptron calculation is performed on the multi-perspective fusion semantic feature vector of the target to obtain a multi-perspective fusion classification feature vector with a length equal to the number of semantic categories, and the softmax normalization exponential function is used to normalize this multi-perspective fusion classification feature vector to obtain the fusion classification prediction probability that this multi-perspective fusion classification feature vector is predicted as each category.

[0094] In some embodiments, the fusion classification prediction probability is calculated using the following formula:

[0095] z mc = f MLP (x Ecls )

[0096]

[0097] In the formula, z mc represents the multi-perspective fusion classification feature vector; represents the probability that the multi-perspective fusion classification feature vector is predicted as the j-th category; represents the predicted j-th category label, j = 1, 2,..., N c .

[0098] In addition, the semantic classification prediction probability, the perspective classification prediction probability, and the multi-perspective fusion classification prediction probability are used to determine the hybrid loss function of the recognition model.

[0099] In some embodiments, the hybrid loss function is calculated as follows:

[0100] (1) Based on the semantic classification prediction probability and the true class label of the target's multi-view images, calculate the semantic classification loss of the target. The calculation formula is as follows:

[0101]

[0102] In the formula, L sc represents the semantic classification loss of the target; y cj represents the true j-th class label; when the true class of the i-th view image is the j-th class, y cj takes 1, otherwise, y cj takes 0.

[0103] (2) Based on the view classification prediction probability and the true view label of the target's multi-view images, calculate the view classification loss of the target. The calculation formula is as follows:

[0104]

[0105] In the formula, L sv represents the view classification loss of the target; y vh represents the true h-th view label. When the true view of the i-th view image is the h-th view, y vh takes 1, otherwise, y vh takes 0; represents the prediction probability that the h-th view image is predicted as the h-th view, z svh represents the view classification feature vector of the h-th view image; represents the prediction probability that the i-th view image is predicted as the h-th view; m is a positive hyperparameter; i = 1, 2,..., N v , h = 1, 2,..., N v , and h ≠ i.

[0106] (3) Based on the multi-view fusion classification prediction probability and the true class label of the target, calculate the fusion classification loss of the target. The calculation formula is as follows:

[0107]

[0108] In the formula, L mc represents the fusion classification loss of the target; y cj represents the true j-th class label; when the true class of the i-th view image is the j-th class, y cj takes 1, otherwise, y cj takes 0.

[0109] (4) Perform a weighted sum of the semantic classification loss, the perspective classification loss, and the fusion classification loss to obtain the hybrid loss function Loss of the recognition model. The calculation formula is as follows:

[0110] Loss = α sc L sc + α sv L sv + α mc L mc

[0111] In the formula, α sc , α sv , and α mc are the hyperparameter weights of the semantic classification loss, the perspective classification loss, and the fusion classification loss, respectively. Each hyperparameter weight is determined according to user needs and is not specifically limited in this application.

[0112] In some embodiments, the pre-constructed recognition model is trained in the following manner:

[0113] Divide the pre-constructed SAR sample images into a training set and a test set;

[0114] Use the SAR sample images in the training set and the test set to train and test the recognition model until the preset convergence condition is reached, and obtain the trained recognition model; the convergence condition is that the number of training times reaches the preset number of times or the loss value of the hybrid loss function is less than the preset value.

[0115] After the recognition model is trained, it can be used for target recognition.

[0116] To prove the effectiveness of the method of this application, the inventors verified the method of this application with the following data. Among them, the number of multi-perspectives is N v = 4, the number of channels of the SAR image C = 1, the image height H = 224, and the image width W = 224; then the i-th perspective image is The category classification label corresponding to the i-th perspective image is c i , c i ∈ {0, 1,..., 9}, the number of semantic categories N c = 10, the perspective classification label corresponding to the i-th perspective image is v i , v i ∈ {0, 1, 2, 3}. In addition, for the semantic feature map x SFi extracted by the semantic feature extraction module, and the perspective feature map extracted by the perspective feature extraction module where the number of channels of the output feature map is C FM = 2048, the height is H FM = 7, and the width is W FM = 7. Therefore

[0117] For each linear mapping module, the parameters of the first convolutional layer are, in sequence, the number of input channels C FM = 2048, the number of output channels C LP = 128, the height H of the convolutional kernel LP = 1, and the width W of the convolutional kernel LP = 1. The parameters of the second convolutional layer are, in sequence, the number of input channels C LP = 128, the number of output channels D = 768, the height H of the convolutional kernel FM = 7, and the width W of the convolutional kernel FM = 7. Therefore, N p = 5, the Transformer network contains L = 6 encoder layers. The hyperparameters of the hybrid loss function are α sc = 0.25, α sv = 0.25, and α mc = 1.

[0118] Based on the above parameters, using the training data shown in Table 1, the format of the sample images is as Figure 5 shown:

[0119] Table 1 Semantic Categories and Target Quantities of Multi-view SAR Images

[0120]

[0121]

[0122] Using Figure 5 and the sample set shown in Table 1 to train the recognition model, the change curve of the hybrid loss function and the change curve of the prediction accuracy rate during the training process are respectively as Figure 6 and Figure 7 shown. It can be seen from the figure that: the training loss gradually decreases with the number of training rounds, and the target classification recognition accuracy gradually increases with the number of training rounds.

[0123] Using the trained model to recognize the target, the recognition results are as Figure 8 shown. It can be seen from the figure that, under any order of input, the method of this application can obtain accurate target classification recognition results.

[0124] It can be seen that the present application improves the recognition accuracy by constructing a perspective feature extraction branch and a perspective classification loss and utilizing perspective prior information; through semantic-perspective feature embedding, test multi-perspective images can be input in any order, which is more flexible; through the Transformer model to fuse the semantic-perspective feature embedding, the correlation and complementarity of the target features of multi-perspective images can be effectively captured, and the final recognition accuracy can be improved.

[0125] As Figure 2 , Figure 3 shown, an embodiment of the present invention provides an object recognition device based on multi-perspective SAR image feature coding and fusion. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. In terms of the hardware level, as Figure 2 shown, it is a hardware architecture diagram of a computing device where an object recognition device based on multi-perspective SAR image feature coding and fusion provided by an embodiment of the present invention is located. In addition to Figure 2 the shown processor, memory, network interface, and non-volatile memory, the computing device where the device is located in the embodiment usually may further include other hardware, such as a forwarding chip responsible for processing packets, etc. Taking software implementation as an example, as Figure 3 shown, as a logically meaningful device, it is formed by the CPU of its computing device reading the corresponding computer program in the non-volatile memory into the memory and running it.

[0126] Please refer to Figure 3 , an embodiment of the present invention provides an object recognition device based on multi-perspective SAR image feature coding and fusion. The device includes:

[0127] An acquisition unit 300, configured to acquire SAR images of a target to be recognized from different perspectives;

[0128] A recognition unit 302, configured to input the SAR images of the target to be recognized from different perspectives into a pre-trained recognition model to obtain a recognition result of the target to be recognized; the recognition model is obtained by training with a pre-constructed sample set, and the sample set is composed of SAR sample images of multiple known-class targets from different perspectives, and each sample image is labeled with a perspective label and a semantic label, so that the recognition model performs object recognition based on the features of different perspective images.

[0129] In some embodiments, the network structure of the recognition model includes: a feature extraction network, a semantic-perspective coding network, a multi-perspective feature fusion network, a semantic classification head, a perspective classification head, and a fusion classification head;

[0130] The feature extraction network includes a semantic feature extraction module and a perspective feature extraction module that perform parallel computing; the feature extraction network takes SAR images of the target from different perspectives as input. For each input SAR image, its semantic features are extracted based on the semantic feature extraction module and a semantic feature map is output, and its perspective features are extracted based on the perspective feature extraction module and a perspective feature map is output;

[0131] The semantic-perspective encoding network takes the semantic feature maps and perspective feature maps of each perspective SAR image output by the feature extraction network as input, and is used to calculate the semantic-perspective embedding of the target multi-perspective SAR image;

[0132] The multi-perspective feature fusion network takes the semantic-perspective feature embedding output by the semantic-perspective encoding network as input, and is used to calculate the multi-perspective fusion semantic feature vector of the target;

[0133] The semantic classification head takes the semantic feature maps of each perspective SAR image output by the semantic feature extraction module as input, and is used to calculate the semantic classification prediction probability of each perspective image; the perspective classification head takes the perspective feature maps of each perspective SAR image output by the perspective feature extraction module as input, and is used to calculate the perspective classification prediction probability of each perspective image; the fusion classification head takes the multi-perspective fusion feature vector output by the multi-perspective feature fusion network as input, and is used to calculate the multi-perspective fusion classification prediction probability of the target; the semantic classification prediction probability, the perspective classification prediction probability, and the multi-perspective fusion classification prediction probability are used to determine the hybrid loss function of the recognition model.

[0134] In some embodiments, the semantic-perspective encoding network includes a semantic linear mapping module, a perspective linear mapping module, and a fusion module, and each linear mapping module consists of two layers of convolutional neural networks;

[0135] The semantic linear mapping module is used to compress the channels and vectorize each input semantic feature map to obtain the corresponding semantic embedding;

[0136] The perspective linear mapping module is used to compress the channels and vectorize each input perspective feature map to obtain the corresponding perspective embedding;

[0137] The fusion module is used to add the semantic embedding and the perspective embedding of each perspective SAR image respectively to obtain the semantic-perspective embedding of each perspective SAR image; and to splice each semantic-perspective embedding with a pre-constructed learnable class token respectively to obtain the semantic-perspective embedding of the target multi-perspective SAR image.

[0138] In some embodiments, the multi-view feature fusion network is constructed based on a preset Transformer network, and the Transformer network includes multiple layers of encoders connected in sequence; wherein, the first layer of encoder takes the semantic-view feature embedding output by the semantic-view encoding network as input, and the other layers of encoders take the feature embedding output by the previous layer of encoder as input; each layer of encoder is used to perform the following operations: perform layer normalization on the input feature embedding, perform multi-head self-attention calculation on the layer-normalized feature embedding, and add the result of the multi-head self-attention calculation to the input feature embedding to obtain an intermediate feature embedding; perform layer normalization on the intermediate feature embedding, perform multi-layer perceptron calculation on the layer-normalized intermediate feature embedding, and add the result of the multi-layer perceptron calculation to the intermediate feature embedding to obtain the feature embedding output by this layer of encoder;

[0139] The multi-view fusion semantic feature vector of the target is calculated through the following method:

[0140] Input the semantic-view feature embedding output by the semantic-view encoding network into the Transformer network, and perform calculations through each layer of encoder in sequence until the last layer of encoder to obtain the multi-view fusion feature; take the feature vector corresponding to the learnable class token in the multi-view fusion feature to obtain the multi-view fusion semantic feature vector of the target.

[0141] In some embodiments, the semantic classification head, the view classification head, and the fusion classification head are all constructed based on multi-layer perceptron layers;

[0142] The calculation process of the semantic classification prediction probability is as follows: perform multi-layer perceptron calculation on the semantic feature maps of each-view SAR images based on the semantic classification head to obtain the single-view semantic classification feature vectors of each-view images, and perform normalization processing on each single-view semantic classification feature vector based on the softmax normalization exponential function to obtain the semantic classification prediction probabilities of each-view images predicted as various categories;

[0143] The calculation process of the view classification prediction probability is as follows: perform multi-layer perceptron calculation on the view feature maps of each-view SAR images based on the view classification head to obtain the view classification feature vectors of each-view images, and perform normalization processing on each view classification feature vector based on the softmax normalization exponential function to obtain the view classification prediction probabilities of each-view images predicted as each view;

[0144] The calculation process of the multi-view fusion classification prediction probability is as follows: Based on the fusion classification head, perform a multi-layer perceptron calculation on the multi-view fusion semantic feature vector of the target to obtain a multi-view fusion classification feature vector with a length equal to the number of semantic categories, and perform normalization processing on this multi-view fusion classification feature vector based on the softmax normalization exponential function to obtain the fusion classification prediction probability that this multi-view fusion classification feature vector is predicted as each category.

[0145] In some embodiments, the hybrid loss function is calculated in the following manner:

[0146] Based on the semantic classification prediction probability and the true category label of each view image of the target, calculate the semantic classification loss of the target;

[0147] Based on the view classification prediction probability and the true view label of each view image of the target, calculate the view classification loss of the target;

[0148] Based on the multi-view fusion classification prediction probability and the true category label of the target, calculate the fusion classification loss of the target;

[0149] Perform a weighted sum of the semantic classification loss, the view classification loss, and the fusion classification loss to obtain the hybrid loss function of the recognition model.

[0150] In some embodiments, the pre-constructed recognition model is trained in the following manner:

[0151] Divide the pre-constructed SAR sample images into a training set and a test set;

[0152] Use the SAR sample images in the training set and the test set to train and test the recognition model until the preset convergence condition is reached to obtain a trained recognition model; the convergence condition is that the number of training times reaches the preset number of times or the loss value of the hybrid loss function is less than the preset value.

[0153] It should be noted that: The target recognition device based on multi-view SAR image feature encoding and fusion provided in the above embodiments is only illustrated by dividing the above functional modules. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the target recognition device based on multi-view SAR image feature encoding and fusion provided in the above embodiments and the embodiments of the target recognition method based on multi-view SAR image feature encoding and fusion belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be elaborated here.

[0154] The embodiments of the present application also provide a computer device, please refer to Figure 3, the computer device includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the target recognition method based on multi-view SAR image feature encoding and fusion provided in the above method embodiments.

[0155] An embodiment of the present application further provides a computer-readable storage medium. At least one instruction, at least one program, a code set, or an instruction set is stored on the computer-readable storage medium. The at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the target recognition method based on multi-view SAR image feature encoding and fusion provided in the above method embodiments.

[0156] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the target recognition method based on multi-view SAR image feature encoding and fusion described in any one of the above embodiments.

[0157] For convenience of description, when describing the above system or device, it is divided into various modules or units according to functions for description. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0158] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present application.

[0159] Finally, it should also be noted that in this text, relational terms such as first, second, third, and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0160] The above are only the preferred embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A target recognition method based on multi-view SAR image feature coding and fusion, characterized in that: The method comprises: Acquire SAR images of the target to be identified at different viewing angles; The SAR images of the target to be identified at different viewing angles are input into a pre-trained recognition model to obtain the recognition result of the target to be identified; the recognition model is trained by a pre-constructed sample set, and the sample set consists of SAR sample images of multiple known category targets at different viewing angles, and each of the sample images is annotated with a viewing angle label and a semantic label, so that the recognition model can perform target recognition based on the features of images with different viewing angles.

2. The method according to claim 1, characterized in that The network structure of the recognition model includes: a feature extraction network, a semantic-view encoding network, a multi-view feature fusion network, a semantic classification head, a view classification head and a fusion classification head; The feature extraction network includes a semantic feature extraction module and a view feature extraction module for parallel computing; the feature extraction network takes SAR images of the target at different view angles as input, and for each input SAR image, extracts its semantic features based on the semantic feature extraction module and outputs a semantic feature map, and extracts its view features based on the view feature extraction module and outputs a view feature map; The semantic-view encoding network uses the semantic feature map and the view feature map of each view SAR image output by the feature extraction network as input, and is used to calculate the semantic-view embedding of the target multi-view SAR image; The multi-view feature fusion network uses the semantic-view feature embedding output by the semantic-view encoding network as input to calculate the multi-view fusion semantic feature vector of the target; The semantic classification head takes the semantic feature map of each viewing angle SAR image output by the semantic feature extraction module as input, and is used to calculate the semantic classification prediction probability of each viewing angle image; the viewing angle classification head takes the viewing angle feature map of each viewing angle SAR image output by the viewing angle feature extraction module as input, and is used to calculate the viewing angle classification prediction probability of each viewing angle image; the fusion classification head takes the multi-view fusion feature vector output by the multi-view feature fusion network as input, and is used to calculate the multi-view fusion classification prediction probability of the target; the semantic classification prediction probability, the viewing angle classification prediction probability and the multi-view fusion classification prediction probability are used to determine the mixed loss function of the recognition model.

3. The method according to claim 2, characterized in that The semantic-view encoding network includes a semantic linear mapping module, a view linear mapping module and a fusion module, each linear mapping module is composed of a two-layer convolutional neural network; The semantic linear mapping module is used to perform channel compression and vectorization on each input semantic feature map to obtain corresponding semantic embedding; The perspective linear mapping module is used to perform channel compression and vectorization on each input perspective feature map to obtain a corresponding perspective embedding; The fusion module is used to add the semantic embedding and the perspective embedding of each perspective SAR image respectively to obtain the semantic-perspective embedding of each perspective SAR image; And each semantic-view embedding is concatenated with the pre-built learnable category tokens to obtain the semantic-view embedding of the target multi-view SAR image.

4. The method according to claim 3, characterized in that The multi-view feature fusion network is constructed based on a preset Transformer network, and the Transformer network includes multiple layers of encoders connected in sequence; wherein the first layer encoder takes the semantic-view feature embedding output by the semantic-view encoding network as input, and the other layers of encoders take the feature embedding output by the previous layer of encoder as input; each layer of encoder is used to perform the following tasks: perform layer normalization on the input feature embedding, perform multi-head self-attention calculation on the feature embedding after layer normalization, and add the result of the multi-head self-attention calculation to the input feature embedding to obtain an intermediate feature embedding; perform layer normalization on the intermediate feature embedding, perform multi-layer perceptron calculation on the intermediate feature embedding after layer normalization, and add the result of the multi-layer perceptron calculation to the intermediate feature embedding to obtain the feature embedding output by the encoder of this layer; The multi-view fusion semantic feature vector of the target is calculated in the following way: The semantic-perspective features output by the semantic-perspective encoding network are embedded into the input Transformer network, and are calculated in turn through each layer of encoder until the last layer of encoder to obtain a multi-perspective fusion feature; the feature vector corresponding to the learnable category token in the multi-perspective fusion feature is taken to obtain a multi-perspective fusion semantic feature vector of the target.

5. The method according to claim 2, characterized in that: The semantic classification head, the view classification head and the fusion classification head are all constructed based on a multi-layer perceptron layer; The calculation process of the semantic classification prediction probability is as follows: based on the semantic classification head, a multi-layer perceptron is used to calculate the semantic feature map of the SAR image of each viewing angle to obtain a single-view semantic classification feature vector of each viewing angle image, and each single-view semantic classification feature vector is normalized based on a softmax normalized exponential function to obtain a semantic classification prediction probability of each viewing angle image being predicted as each category; The calculation process of the view classification prediction probability is as follows: based on the view classification head, a multi-layer perceptron is used to calculate the view feature map of each view SAR image to obtain a view classification feature vector of each view image, and each view classification feature vector is normalized based on a softmax normalized exponential function to obtain a view classification prediction probability of each view image being predicted as each view; The calculation process of the multi-view fusion classification prediction probability is as follows: based on the fusion classification head, a multi-layer perceptron calculation is performed on the multi-view fusion semantic feature vector of the target to obtain a multi-view fusion classification feature vector whose length is equal to the number of semantic categories, and the multi-view fusion classification feature vector is normalized based on a softmax normalized exponential function to obtain the multi-view fusion classification feature vector predicted as a fusion classification prediction probability for each category.

6. The method according to claim 5, characterized in that The hybrid loss function is calculated as follows: Calculate the semantic classification loss of the target based on the semantic classification prediction probability and true category label of the target images at each viewpoint; Calculate the target’s view classification loss based on the predicted view classification probability and true view label of each view image of the target; Calculate the fusion classification loss of the target based on the multi-view fusion classification prediction probability and the true category label of the target; The semantic classification loss, the view classification loss and the fusion classification loss are weightedly summed to obtain a hybrid loss function of the recognition model.

7. The method according to claim 2, characterized in that The pre-built recognition model is trained as follows: The pre-constructed SAR sample images are divided into training set and test set; The recognition model is trained and tested using the SAR sample images in the training set and the test set until a preset convergence condition is reached to obtain a trained recognition model; the convergence condition is that the number of training times reaches a preset number of times or the loss value of the mixed loss function is less than a preset value.

8. A target recognition device based on multi-view SAR image feature coding and fusion, characterized in that: The device comprises: An acquisition unit, used for acquiring SAR images of a target to be identified at different viewing angles; The recognition unit is used to input the SAR images of the target to be identified at different viewing angles into a pre-trained recognition model to obtain the recognition result of the target to be identified; the recognition model is obtained by training a pre-constructed sample set, and the sample set is composed of SAR sample images of multiple known category targets at different viewing angles, and each of the sample images is annotated with a viewing angle label and a semantic label, so that the recognition model can perform target recognition based on the features of images with different viewing angles.

9. A computing device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • SAR automatic target recognition method based on multi-view deep learning framework

    CN108038445A

  • Multi-modal information fusion target identification method

    CN117351315A

  • Multi-label image classification method and apparatus, and multi-label image classification model training method and apparatus

    WO2024124770A1

Cited By

  • Multi-view target detection method and device, multi-view target detection model training method and device, equipment, storage medium and program product

    CN120411485A

  • Image classification method based on test time training backbone network

    CN121170449A

  • Multi-view fruit quality detection method, device and system, medium and product

    CN121616581A