ISAR image quality rating method based on double-flow network

The dual-stream network architecture for ISAR images effectively addresses the low accuracy of existing methods by combining global and local features, enhancing precision and robustness in ISAR image evaluation.

CN120318556APending Publication Date: 2025-07-15XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510335266.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing ISAR image quality evaluation methods rely on a single feature, making it difficult to accurately capture image details, resulting in low evaluation accuracy and difficulty in adapting to inverse synthetic aperture radars.

Method used

The ISAR image quality rating method based on dual-stream network is adopted, and the global structural information and local distortion characteristics of the image are extracted in parallel, combined with multi-type distortion data set training, dynamic feature fusion and hierarchical correction, to improve the evaluation accuracy and robustness.

Benefits of technology

It significantly improves the accuracy and robustness of ISAR image quality evaluation, with a classification accuracy of 97.07%, providing efficient and reliable technical support for ISAR image screening and subsequent interpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318556A_ABST
    Figure CN120318556A_ABST
Patent Text Reader

Abstract

The invention relates to an ISAR image quality rating method based on a double-flow network. The method comprises the following steps: acquiring an ISAR image to be evaluated; inputting the ISAR image to be evaluated into a trained ISAR image quality rating network to extract global feature information and local feature information, and combining the global feature information and the local feature information to obtain a corresponding image rating result; wherein the trained ISAR image quality rating network is obtained by performing iterative training on the ISAR image quality rating network by using a pre-generated data set containing different types of disturbance distortion images. According to the method, the quality rating precision of the ISAR image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of ISAR image processing, and particularly relates to an ISAR image quality rating method based on a dual-stream network. Background Art

[0002] Inverse Synthetic Aperture Radar (ISAR) imaging is a means of two-dimensional high-resolution imaging of an observed target by using a large time-bandwidth transmitted signal and the relative rotation between the observed target and the radar observation line of sight. It has the characteristics of all-weather, all-day and long-distance observation. ISAR images can reflect the shape, structure and other characteristics of the observed target, and have been widely used in the field of space target surveillance. Limited by the diversity of ISAR imaging scenarios and the robustness of processing algorithms, the quality of ISAR images varies. How to efficiently and accurately evaluate the quality of ISAR images, and then quickly and effectively screen high-quality ISAR images to provide support for subsequent target structure analysis and information interpretation based on ISAR images has become a research hotspot and difficulty in the field of ISAR image processing.

[0003] Currently, the existing image quality evaluation methods are mainly divided into two categories: subjective evaluation and objective evaluation. The subjective evaluation method relies on the direct perception of the observer on the image quality for scoring, and the evaluation result is relatively reliable. However, it is also affected by the subjectivity of the observer and is extremely time-consuming and laborious. Objective evaluation (for example, using a neural network for quality evaluation) mostly relies on manually specified evaluation indicators, has the problem of a single evaluation standard, and the existing objective evaluation focuses more on global feature extraction and is difficult to capture specific details of the picture, further limiting the evaluation accuracy. In short, the existing image quality evaluation methods have low evaluation accuracy and are difficult to be applied to inverse synthetic aperture radar. Summary of the Invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides an ISAR image quality rating method based on a dual-stream network. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0005] The present invention provides an ISAR image quality rating method based on a dual-stream network, including: obtaining an ISAR image to be evaluated; inputting the ISAR image to be evaluated into a trained ISAR image quality rating network to extract global feature information and local feature information, and combining the global feature information and the local feature information to obtain a corresponding image rating result; wherein, the trained ISAR image quality rating network is obtained by iteratively training the ISAR image quality rating network using a pre-generated data set containing different types of perturbed and distorted images.

[0006] Compared with the prior art, the beneficial effects of the present invention are as follows: By integrating global features and local detail features through a dual-stream network architecture and combining a pre-generated diverse perturbation distortion dataset training mechanism, the accuracy and robustness of ISAR image quality assessment are significantly improved. Among them, the dual-stream network overcomes the limitation of traditional methods relying on a single feature by parallelly extracting the global structural information of the image (such as the overall shape of the target and the low-frequency background) and local distortion features (such as noise distribution and edge defocus); while the training strategy based on multi-type distortion data including signal-to-noise ratio fluctuations, motion compensation errors, etc. enhances the generalization ability of the model to complex scenarios, enabling it to accurately identify the impact of different distortion patterns on image quality. Finally, a high-accuracy rating (reaching 97.07% in experiments) is achieved through a dynamic feature fusion and hierarchical-correction joint optimization mechanism, providing an efficient and reliable technical support for ISAR image screening and subsequent interpretation. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 FIG. is a schematic flowchart of a method for ISAR image quality rating based on a dual-stream network proposed in an embodiment of the present invention;

[0008] Figure 2 FIG. is a structural block diagram of an ISAR image quality rating network provided in an embodiment of the present invention;

[0009] Figure 3 FIG. is a structural block diagram of a high-low frequency image feature extraction unit provided in an embodiment of the present invention;

[0010] Figure 4 FIG. is a structural block diagram of a multi-scale feature extraction unit provided in an embodiment of the present invention;

[0011] Figure 5 FIG. is an example diagram of the process of obtaining a trained ISAR image quality rating network provided in an embodiment of the present invention;

[0012] Figure 6 FIG. is an example diagram of ISAR image quality rating based on a dual-stream network provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] The following further describes the present invention in detail with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0014] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0015] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0016] Although the present invention has been described in connection with various embodiments herein, however, in the process of implementing the claimed invention, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0017] Now, in conjunction with the accompanying drawings, a method for ISAR image quality rating based on a two-stream network proposed in an embodiment of the present invention will be described in detail. Figure 1 It is a schematic flowchart of a method for ISAR image quality rating based on a two-stream network proposed in an embodiment of the present invention. As Figure 1 shown, the method includes:

[0018] Step 110: Obtain the ISAR image to be evaluated.

[0019] Step 120: Input the ISAR image to be evaluated into the trained ISAR image quality rating network to extract global feature information and local feature information, and combine the global feature information and the local feature information to obtain the corresponding image rating result; wherein, the trained ISAR image quality rating network is obtained by iteratively training the ISAR image quality rating network using a pre-generated dataset containing different types of perturbed and distorted images.

[0020] Here, the pre-generated dataset containing different types of perturbed and distorted images is obtained through the following method:

[0021] (1) Generate images;

[0022] Here, in order to increase the data richness as much as possible and effectively improve the comprehensive recognition ability of the model, the PFA imaging technology is used to generate images in the ideal state, images with different signal-to-noise ratios, images with different sparse aperture observations, images with different motion compensation accuracies (translation compensation, rotation compensation), and images with different defocusing situations. Among them, in the defocus scene, two situations under different translation compensation accuracies and different rotation compensation accuracies are respectively simulated. Here, in order to take into account the quality levels of targets of different sizes, the images generated by PFA imaging include: satellite images of large sizes, and scattering model images of small sizes (such as images generated by dihedral angle and cuboid scattering models).

[0023] (2) Label division;

[0024] Here, the image labels need to be divided by two methods: objective evaluation and subjective evaluation. Specifically, the generated images are divided into 3 grades according to the quality. Images with good quality and clear target recognition are regarded as images with the quality level of "excellent"; images with medium quality, with certain noise and a small amount of defocusing phenomenon, are regarded as images with the quality level of "general"; images with poor quality, low signal-to-noise ratio and a large amount of defocusing phenomenon, are regarded as images with the quality level of "poor".

[0025] When objectively evaluating the images, the images in the ideal state are regarded as images with the quality level of "excellent". For images with different signal-to-noise ratios, the images with the signal-to-noise ratio range of 15dB - 25dB are regarded as images with the quality level of "general", and the images with the signal-to-noise ratio range of 0dB - 10dB are regarded as images with the quality level of "poor". For the case of sparse aperture observation, the images with the defect rate range of 10% - 30% are regarded as images with the quality level of "general", and the images with the error range of 40% - 60% are regarded as images with the quality level of "poor". For different translation compensation accuracy cases, the images with the error range of 1 / 8 distance unit to 1 / 4 distance unit are regarded as images with the quality level of "general", and the images with the error range of 5 / 16 distance unit to 7 / 16 distance unit are regarded as images with the quality level of "poor". For different rotation compensation accuracy cases, the images with the error range of 10% - 30% are regarded as images with the quality level of "general", and the images with the error range of 40% - 60% are regarded as images with the quality level of "poor".

[0026] When subjectively evaluating the images, experienced senior ISAR imaging researchers are used to rate them, which are divided into three grades: "excellent", "general" and "poor".

[0027] It should be noted that the number of images corresponding to each case is a fixed value (e.g., 1600 images). When dividing the labels of the images, it is inevitable that the requirements for division are not met, resulting in fluctuations in the number of images corresponding to each case. When the number of images corresponding to a certain case is less than the fixed value, PFA imaging is used to continue generating images of the corresponding type for supplementation. When the number of images corresponding to a certain case is greater than the fixed value, a part of the images are screened out.

[0028] (3) Image screening;

[0029] Select the overlapping part of the images divided by subjective evaluation and the images of objective evaluation to form a final image dataset including three levels of "excellent", "general" and "poor", and use 70% of them as the training set, 20% as the test set, and 10% as the validation set.

[0030] After obtaining the pre-generated dataset containing different types of perturbed and distorted images, Figure 2 is the structural block diagram of the ISAR image quality rating network provided by the embodiment of the present invention. As Figure 2 shown, the ISAR image quality rating network includes: a first feature extraction unit, a second feature extraction unit, a high-low frequency image feature extraction unit, a multi-scale feature extraction unit, and a classification unit; wherein, the input end of the first feature extraction unit receives the ISAR image to be evaluated, and the output end is connected to the input end of the high-low frequency image feature extraction unit. The input end of the second feature extraction unit receives the saliency map, and the saliency map is obtained by processing the ISAR image to be evaluated using the residual spectrum algorithm. The output end of the second feature extraction unit is connected to the input end of the multi-scale feature extraction unit. The output ends of the high-low frequency image feature extraction unit and the multi-scale feature extraction unit are both connected to the input end of the classification unit, and the output end of the classification unit is used to output the image rating result of the ISAR image to be evaluated.

[0031] Here, the first feature extraction unit is used to extract the features of the ISAR image to be evaluated; the second feature extraction unit is used to extract the features of the corresponding saliency map; the high-low frequency image feature extraction unit is used to calculate the global feature information using the features of the ISAR image to be evaluated; the multi-scale feature extraction unit is used to calculate the local feature information using the features of the saliency map; the classification unit is used to output the image rating result based on the global feature information and the local feature information.

[0032] Here, the first feature extraction unit and the second feature extraction unit are both used for shallow feature extraction, and both are feature extraction modules of Vgg16. The feature extraction module of Vgg16 includes a number of ninth convolution modules and a second max pooling module connected in series. Specifically, each ninth convolution module is a 3*3 convolution module. A number of 3*3 convolution modules are connected in series in sequence, and the last ninth convolution module is connected in series with the second max pooling module.

[0033] Figure 3 is the structural block diagram of the high and low frequency image feature extraction unit provided by the embodiment of the present invention. As Figure 3 shown, the high and low frequency image feature extraction unit includes: a first convolution sub-unit (3*3 convolution layer), a high frequency feature extraction sub-unit (HA), a low frequency feature extraction sub-unit (LA), and a first feature fusion sub-unit; the input end of the first convolution sub-unit is connected to the output end of the first feature extraction unit, and the output end is respectively connected to the input ends of the high frequency feature extraction sub-unit HA and the low frequency feature extraction sub-unit LA. The output ends of the high frequency feature extraction sub-unit HA and the low frequency feature extraction sub-unit LA are both connected to the input end of the first feature fusion sub-unit, and the output end of the first feature fusion sub-unit is connected to the input end of the classification unit; the high frequency feature extraction sub-unit HA is used for performing high frequency feature extraction on the ISAR image features to be evaluated to obtain high frequency feature information; the low frequency feature extraction sub-unit LA is used for performing low frequency feature extraction on the ISAR image features to be evaluated to obtain low frequency feature information; the first feature fusion sub-unit is used for performing feature fusion processing on the high frequency feature information and the low frequency feature information to obtain global feature information.

[0034] Further, the high-frequency feature extraction subunit HA includes: a first max pooling module (MP) and a first convolutional module (1*1 convolutional layer); the input end of the first max pooling module is connected to the input end of the first convolutional subunit, the output end is connected to the input end of the first convolutional module, and the output end of the first convolutional module is connected to the input end of the first feature fusion subunit. Moreover, the low-frequency feature extraction subunit LA includes: a second convolutional module (1*1 convolutional layer), a multi-head self-attention module (Self-Attn), a fully connected layer (MatMul-Linear), a multi-module perceptron module (MLP), a channel attention module (CA), and a second feature fusion subunit; the input ends of the second convolutional module and the multi-head self-attention module are both connected to the input end of the first convolutional subunit, the output end of the second convolutional module is connected to the first input end of the fully connected layer, the output end of the multi-head self-attention module is connected to the second input end of the fully connected layer, the output end of the fully connected layer is respectively connected to the input ends of the multi-module perceptron module and the channel attention module, the output ends of the multi-module perceptron module and the channel attention module are both connected to the input end of the second feature fusion subunit, and the output end of the second feature fusion subunit is connected to the input end of the first feature fusion subunit.

[0035] Figure 4 is the structural block diagram of the multi-scale feature extraction unit provided by the embodiment of the present invention. As Figure 4 shown, the multi-scale feature extraction unit includes: a third convolutional module (3*3 convolutional layer), a fourth convolutional module (3*3 convolutional layer), a spatial branch module (SA), a multi-channel branch module (MCA), and a third feature fusion module; the input end of the third convolutional module is connected to the output end of the second feature extraction unit, the output end is connected to the input end of the fourth convolutional module, the output end of the fourth convolutional module is respectively connected to the input ends of the spatial branch module, the multi-channel branch module, and the third feature fusion module, the output ends of the spatial branch module and the multi-channel branch module are both connected to the input end of the third feature fusion module, and the output end of the third feature fusion module is connected to the input end of the classification unit.

[0036] Further, the spatial branch module includes: an average pooling module (AP), a max pooling module (MP), a concatenation module (concat), a fifth convolutional module (7*7 convolutional layer), and a first activation function module; the input ends of the average pooling module and the max pooling module are both connected to the output end of the fourth convolutional module, the output ends of the average pooling module and the max pooling module are both connected to the input end of the concatenation module, the output end of the concatenation module is connected to the input end of the fifth convolutional module, the output end of the fifth convolutional module is connected to the input end of the first activation function module, and the output end of the first activation function module is connected to the input end of the third feature fusion module.

[0037] Moreover, the multi-channel branch module includes: an X-direction pooling module (XAP), a Y-direction pooling module (YAP), a sixth convolutional module (1*1 convolutional layer), a seventh convolutional module (1*1 convolutional layer), an eighth convolutional module (1*1 convolutional layer), a second activation function module, and a third activation function module; the input end of the X-direction pooling module and the input end of the Y-direction pooling module are both connected to the output end of the fourth convolutional module, the output end of the X-direction pooling module and the output end of the Y-direction pooling module are both connected to the input end of the sixth convolutional module, the first output end of the sixth convolutional module is connected to the input end of the seventh convolutional module, the second output end is connected to the input end of the eighth convolutional module, the output end of the seventh convolutional module is connected to the input end of the second activation function module, the output end of the eighth convolutional module is connected to the input end of the third activation function module, and the output end of the second activation function module and the output end of the third activation function module are both connected to the input end of the third feature fusion module.

[0038] Here, the first activation function module, the second activation function module, and the third activation function module are all sigmoid activation functions.

[0039] Here, the classification unit includes: a first classification module and a second classification module; the first input end of the first classification module is connected to the output end of the high-low frequency image feature extraction unit, the input end of the second classification module is connected to the output end of the multi-scale feature extraction unit, the first output end of the first classification module is connected to the output end of the second classification module, and the second output end of the first classification module is used to output the image rating result.

[0040] Exemplarily, both the first classification module and the second classification module are composed of a dropout layer, a non-linear activation function ReLU, a batch normalization (BatchNorm) layer, and a fully connected layer.

[0041] The above is the structural description of the ISAR image quality rating network. Now, the process of iteratively training the ISAR image quality rating network using a pre-generated image dataset to obtain a trained ISAR image quality rating network will be described. Figure 5 It is an example diagram of the process of obtaining a trained ISAR image quality rating network provided by an embodiment of the present invention.

[0042] As Figure 5 shown, in the i-th training process (i is a positive integer), 70% of the picture data in the dataset is input into the i-th ISAR image quality rating network, and the pictures in this image set are processed from two perspectives: the image stream and the saliency stream, with the image size being 256*256*3.

[0043] (a) In the image stream branch;

[0044] First, feature extraction is performed by the first feature extraction unit, and then it enters the high-frequency and low-frequency image feature extraction unit. In the high-frequency feature extraction sub-unit HA, max-pooling operation and 1*1 convolution operation are used to cover more local information within the receptive field through local convolution, thereby effectively extracting high-frequency feature information. The high-frequency feature information Y h can be expressed as: Y h = conv 1×1 (MP(X in )); where MP represents the max-pooling operation. And, in the low-frequency feature extraction sub-unit LA, the multi-head self-attention module is used to effectively capture the low-frequency components in the visual data, such as the global shape and structure information of the scene or object in the image. Subsequently, through the MLP and CA modules respectively, the non-linear feature extraction ability of the model and the information aggregation ability of the module are strengthened. Finally, the outputs of the two branches are multiplied element-wise to obtain the low-frequency feature. Specifically, when implementing, first use self-attention to calculate the attention matrix for the linearly transformed input and aggregate the token features, and further perform a 1×1 convolution operation on the attention matrix to strengthen the feature expression ability to obtain the low-frequency feature information X mix . The low-frequency feature information X mix can be expressed as: where, assuming the embedding tensor of the input token is using linear transformation to obtain the query key value

[0045] Here, the multi-module perceptron module (MLP) contains two linear layers and one GELU layer. Since different heads focus on different information of the object, information aggregation is crucial for the information extraction ability of the module. Therefore, LA considers enhancing the information aggregation ability of the module by introducing channel re-weighting on the basis of MLP, considering the importance of each channel more comprehensively and making a more reasonable aggregation decision. Specifically, when implementing, the channel attention allocation is performed by multiplying the MLP module with the channel attention matrix to perform feature transformation. The channel attention matrix is calculated by the channel attention module (CA) and is defined as follows:

[0046] It should be noted that the difference from the self-attention mechanism is that when CA calculates the channel attention matrix, it calculates along the channel dimension and uses feature covariance for feature transformation. This design makes the relevant feature channels with larger correlations be preferentially aggregated, while the relevant feature channels with smaller correlations are outlierized, which helps the model extract highly relevant information and further improves the information aggregation ability. Finally, the low-frequency feature information Y l is aggregated as follows: Y l = CA(X mix )MLP(Xmix )。

[0047] Furthermore, the global feature information output by the high-low frequency image feature extraction unit can be expressed as: Y hl = Y h + Y l 。

[0048] (b) In the saliency flow branch;

[0049] First, use the Spectral Residual (SR) module to process the dataset images to effectively suppress the interference components in the background area of the images to obtain the saliency region of the images. Subsequently, use the second feature extraction unit for feature extraction. The features of the obtained saliency map are input into the multi-scale feature extraction unit. To better perceive the local quality, the multi-scale feature extraction unit is designed as a dual-branch structure. The spatial attention (SA) branch module aims to capture complex spatial relationships, while the multi-channel attention (MCA) branch module aims to capture the uneven information contained between different channels, enabling the network to more efficiently focus on important information. SA can assign different weights to each pixel of the spatial features at each layer.

[0050] In a possible implementation, in the spatial attention branch module, after being processed by the third convolutional module and the fourth convolutional module, the obtained feature map is The spatial attention branch module aggregates X f in the channel dimension through max pooling and average pooling respectively to generate two single-channel feature maps X m and X a , where, X m = MaxPool(X f ), X a = AvgPool(X f ), The two single-channel feature maps X m and X a can respectively extract the most significant local features and the overall background information. Further, the two single-channel feature maps X m and X a are connected through a concatenation operation, and the fifth convolutional module is used to perform feature fusion on them, and then immediately followed by using the sigmoid activation function to obtain the spatial weight value W S . This spatial weight value W S can be expressed as: W S = σ(conv 7×7 (concat(X m , X a ))), where, σ represents the sigmoid function, which is used to normalize the weight to [0, 1].

[0051] In the multi-channel branch module, first, the channel attention is decomposed into 1D feature encodings in two directions. Each direction aggregates one-dimensional features, so that long-range dependencies can be captured in one direction while precise position information is retained in the other direction. Applying the attention along the horizontal and vertical directions to the input tensor simultaneously can obtain two attention maps. Each element in the two attention maps reflects whether the object of interest exists in the corresponding row and column. This encoding process allows us to more accurately locate the exact position of the object of interest. In specific implementation, the channel branch performs average pooling on all rows in each channel along the horizontal direction to obtain the horizontal-direction feature X x , and performs average pooling on all columns in each channel along the vertical direction to obtain the vertical-direction feature X y . Further, the features in the two directions are fused by concat, and channel compression is performed using a non-linear transformation. Then, two independent 1×1 convolutions are used to obtain the horizontal-direction attention W Cx and the vertical-direction attention W Cy . Finally, the attention weight values in different directions are applied to the original feature map using pixel-wise multiplication.

[0052] Accumulate the spatial attention feature map obtained by SA and the multi-channel attention feature map output by MCA to obtain the finally optimized saliency feature (local feature information) X S as: X S = W S X f + W Cx W Cy X f .

[0053] Input the global feature information Y hl into the first classification module. Based on the hierarchical loss function corresponding to the first classification module, output the i-th hierarchical loss value, and input the local feature information into the second classification module. Based on the calibration loss function corresponding to the second classification module, output the i-th calibration loss value. In order to comprehensively consider the errors existing when using the global feature information for classification and the errors existing when using the local feature information for classification, the error finally corresponding to the i-th training is defined as the sum of the two; among them, the i-th hierarchical loss value can be expressed as: where N is the number of samples, y ic represents 1 if the true category of sample i is consistent with c, otherwise 0; p ic represents the predicted probability that sample i belongs to category c; the i-th calibration loss value can be expressed as: where N is the number of samples, x i is the true value of the i-th sample, y iis the predicted value of the model for the i-th sample; the sum of the two can be expressed as: Loss = loss1 + ω1·loss2; where ω1 controls the influence degree of the calibration task, and the general empirical value is set to 0.01 - 0.1; determine whether the sum of the i-th classification loss value and the i-th calibration loss value is less than or equal to a preset value. If it is greater than the preset value, use the Adam optimizer to update the model parameters in the i-th ISAR image quality rating network to obtain the (i + 1)-th ISAR image quality rating network, and continue to input the data set into the (i + 1)-th ISAR image quality rating network for the (i + 1)-th iterative training until the sum of the two loss values is less than or equal to the preset value or reaches the preset number of iterations; if it is less than or equal to the preset value or reaches the preset number of iterations, stop training, and use the image stream branch in the i-th ISAR image quality rating network as the trained ISAR image quality rating network.

[0054] To verify the accuracy of the trained ISAR image quality rating network, Figure 6 is an example diagram of ISAR image quality rating based on a dual-stream network provided by an embodiment of the present invention.

[0055] As Figure 6 shown, first use the above image generation principle to generate a data set containing images with different types of perturbation distortions, specifically including 1,600 ideal images, 9,600 images with different signal-to-noise ratios, 19,200 images corresponding to different motion compensation accuracies, and 9,600 images corresponding to different sparse observation situations. Use 70% of the data in the data set for model training, then use 20% of the data in the finally formed data set for testing, and 10% of the data for verification. Among them, the accuracy Acc is used as the evaluation index for the accuracy of the TSAIQA model, which is defined as follows: Among them, TP represents the number of samples where the predicted image quality is consistent with the image quality label, and SUM represents the total number of samples. The final test results are shown in Table 1 below. Based on Table 1, it can be seen that the final prediction accuracy can reach 97.07%.

[0056] Table 1

[0057] Level Good Medium Poor Good 2869 10 1 Medium 36 2718 126 Poor 45 35 2800

[0058] Aiming at the problem that the evaluation accuracy of existing image quality evaluation methods is low and it is difficult to be applied to inverse synthetic aperture radar, the present invention provides an ISAR image quality rating method based on a dual-stream network, and this method has the following technical effects:

[0059] (1) Dual-stream feature fusion and multi-dimensional feature extraction

[0060] Design a dual-stream network architecture (image stream + saliency stream), and capture global structure and local detail features synchronously through high- and low-frequency feature extraction units and multi-scale branches respectively. The image stream generates global information by fusing high-frequency features (such as edge noise) and low-frequency features (overall target shape); the saliency stream combines the residual spectrum algorithm to extract the salient regions, and focuses on local distortions (such as defocus, motion compensation error) through spatial / multi-channel attention modules. The two are dynamically fused through the classification unit, solving the problem of detail loss caused by traditional methods relying on single features. Experiments show that the classification accuracy reaches 97.07%, which is significantly improved compared with traditional methods.

[0061] (2) Adaptive attention mechanism and multi-scale perception

[0062] The network integrates multi-head self-attention, channel / space attention and direction-sensitive pooling techniques to enhance the feature expression ability. The low-frequency branch correlates the global context through self-attention, and combines channel reweighting to screen out highly relevant features; the multi-scale branch uses X / Y direction pooling to capture long-range dependencies and position information, and the spatial attention dynamically assigns regional weights, enabling the model to accurately locate the distorted regions (such as defects caused by sparse apertures, signal-to-noise ratio fluctuations).

[0063] (3) Robust training strategy and data augmentation

[0064] Based on a diverse synthetic dataset (covering 6 types of perturbations such as signal-to-noise ratio, motion compensation error, etc., with a total of 44,800 images) and a subjective and objective joint annotation mechanism, ensure the authenticity of the data and the reliability of the labels. During training, a joint optimization strategy of hierarchical loss (classification error) and calibration loss (local feature error) is adopted, and the global-local feature contributions are balanced through hyperparameters to enhance the model's generalization ability to noise and annotation deviations, achieving high robustness across scenarios on the test set.

[0065] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. An ISAR image quality rating method based on a two-stream network, characterized in that Including: Obtain the ISAR image to be evaluated; Input the ISAR image to be evaluated into the trained ISAR image quality rating network to extract global feature information and local feature information, and combine the global feature information and the local feature information to obtain the corresponding image rating result; wherein, the trained ISAR image quality rating network is obtained by iteratively training the ISAR image quality rating network using a pre-generated dataset containing different types of perturbed and distorted images.

2. The method for ISAR image quality rating based on a two-stream network according to claim 1, wherein The ISAR image quality rating network includes: a first feature extraction unit, a second feature extraction unit, a high-low frequency image feature extraction unit, a multi-scale feature extraction unit, and a classification unit; Wherein, the input end of the first feature extraction unit receives the ISAR image to be evaluated, and the output end is connected to the input end of the high-low frequency image feature extraction unit. The input end of the second feature extraction unit receives the saliency map, which is obtained by processing the ISAR image to be evaluated using the residual spectrum algorithm. The output end of the second feature extraction unit is connected to the input end of the multi-scale feature extraction unit. The output ends of the high-low frequency image feature extraction unit and the multi-scale feature extraction unit are both connected to the input end of the classification unit, and the output end of the classification unit is used to output the image rating result of the ISAR image to be evaluated; The first feature extraction unit is used to extract the features of the ISAR image to be evaluated; The second feature extraction unit is used to extract the features of the saliency map; The high-low frequency image feature extraction unit is used to calculate and obtain the global feature information using the features of the ISAR image to be evaluated; The multi-scale feature extraction unit is used to calculate and obtain the local feature information using the features of the saliency map; The classification unit is used to output the image rating result based on the global feature information and the local feature information.

3. The method for ISAR image quality rating based on a two-stream network according to claim 2, wherein The high-low frequency image feature extraction unit includes: a first convolution sub-unit, a high-frequency feature extraction sub-unit, a low-frequency feature extraction sub-unit, and a first feature fusion sub-unit; The input end of the first convolution sub-unit is connected to the output end of the first feature extraction unit, and the output end is respectively connected to the input ends of the high-frequency feature extraction sub-unit and the low-frequency feature extraction sub-unit. The output ends of the high-frequency feature extraction sub-unit and the low-frequency feature extraction sub-unit are both connected to the input end of the first feature fusion sub-unit, and the output end of the first feature fusion sub-unit is connected to the input end of the classification unit; The high-frequency feature extraction sub-unit is used to perform high-frequency feature extraction on the features of the ISAR image to be evaluated to obtain high-frequency feature information; The low-frequency feature extraction sub-unit is used to perform low-frequency feature extraction on the features of the ISAR image to be evaluated to obtain low-frequency feature information; The first feature fusion sub-unit is used to perform feature fusion processing on the high-frequency feature information and the low-frequency feature information to obtain the global feature information.

4. The ISAR image quality rating method based on a two-stream network according to claim 3, characterized in that, The high-frequency feature extraction subunit includes: a first max-pooling module and a first convolutional module; The input end of the first max-pooling module is connected to the input end of the first convolutional subunit, the output end is connected to the input end of the first convolutional module, and the output end of the first convolutional module is connected to the input end of the first feature fusion subunit.

5. The method for ISAR image quality rating based on a two-stream network according to claim 3, wherein The low-frequency feature extraction subunit includes: a second convolutional module, a multi-head self-attention module, a fully connected layer, a multi-module perceptron module, a channel attention module, and a second feature fusion subunit; The input ends of the second convolutional module and the multi-head self-attention module are both connected to the input end of the first convolutional subunit. The output end of the second convolutional module is connected to the first input end of the fully connected layer, and the output end of the multi-head self-attention module is connected to the second input end of the fully connected layer. The output end of the fully connected layer is respectively connected to the input ends of the multi-module perceptron module and the channel attention module. The output ends of the multi-module perceptron module and the channel attention module are both connected to the input end of the second feature fusion subunit. The output end of the second feature fusion subunit is connected to the input end of the first feature fusion subunit.

6. The method for ISAR image quality rating based on a two-stream network according to claim 2, wherein The multi-scale feature extraction unit includes: a third convolutional module, a fourth convolutional module, a spatial branch module, a multi-channel branch module, and a third feature fusion module; The input end of the third convolutional module is connected to the output end of the second feature extraction unit, and the output end is connected to the input end of the fourth convolutional module. The output end of the fourth convolutional module is respectively connected to the input ends of the spatial branch module, the multi-channel branch module, and the third feature fusion module. The output ends of the spatial branch module and the multi-channel branch module are both connected to the input end of the third feature fusion module. The output end of the third feature fusion module is connected to the input end of the classification unit.

7. The ISAR image quality rating method based on a two-stream network according to claim 6, wherein The spatial branch module includes: an average pooling module, a max-pooling module, a splicing module, a fifth convolutional module, and a first activation function module; The input ends of the average pooling module and the max-pooling module are both connected to the output end of the fourth convolutional module. The output ends of the average pooling module and the max-pooling module are both connected to the input end of the splicing module. The output end of the splicing module is connected to the input end of the fifth convolutional module. The output end of the fifth convolutional module is connected to the input end of the first activation function module. The output end of the first activation function module is connected to the input end of the third feature fusion module.

8. The ISAR image quality rating method based on a two-stream network according to claim 6, characterized in that, The multi-channel branch module includes: an X-direction pooling module, a Y-direction pooling module, a sixth convolutional module, a seventh convolutional module, an eighth convolutional module, a second activation function module, and a third activation function module; The input ends of the X-direction pooling module and the Y-direction pooling module are both connected to the output end of the fourth convolution module. The output end of the X-direction pooling module and the output end of the Y-direction pooling module are both connected to the input end of the sixth convolution module. The first output end of the sixth convolution module is connected to the input end of the seventh convolution module, and the second output end is connected to the input end of the eighth convolution module. The output end of the seventh convolution module is connected to the input end of the second activation function module, and the output end of the eighth convolution module is connected to the input end of the third activation function module. The output ends of the second activation function module and the third activation function module are both connected to the input end of the third feature fusion module.

9. The ISAR image quality rating method based on a dual-stream network according to claim 2, characterized in that The classification unit includes: a first classification module and a second classification module; The first input end of the first classification module is connected to the output end of the high-low frequency image feature extraction unit, the input end of the second classification module is connected to the output end of the multi-scale feature extraction unit, the first output end of the first classification module is connected to the output end of the second classification module, and the second output end of the first classification module is used to output the image rating result.

10. The method for ISAR image quality rating based on a dual-stream network according to claim 2, wherein, Both the first feature extraction unit and the second feature extraction unit are feature extraction modules of Vgg16, and the feature extraction module of Vgg16 includes a plurality of ninth convolution modules and a second max pooling module connected in series.