Synthetic aperture radar image recognition method based on diffusion generation low-rank fine tuning
Through a low-rank fine-tuning method based on diffusion generation, the synthetic aperture radar image model is incrementally fine-tuned using Gaussian noise training and low-rank decomposition matrix bypass, which solves the sample shortage and overfitting problems of the synthetic aperture radar image recognition model and improves the scalability and recognition accuracy of the model.
Patent Information
- Application Number
- CN202510601581.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-09-19
AI Technical Summary
In the existing technology, the synthetic aperture radar image recognition model suffers from insufficient sample quantity and overfitting problems of traditional fine-tuning methods, resulting in decreased generalization performance and image generation quality, which limits the application scalability of the model.
A low-rank fine-tuning method based on diffusion generation is adopted. Gaussian noise is added to the synthetic aperture radar image samples to train the diffusion model. The diffusion model is incrementally fine-tuned using the low-rank decomposition matrix bypass to generate high-quality enhanced image samples. The model is then trained in combination with the image recognition model to improve the model's learning efficiency and adaptability to new sample types.
Without changing the original model parameters, the computing resource consumption is reduced, the learning efficiency and adaptability of the image recognition model to new sample types are improved, and the scalability and recognition accuracy of the model are enhanced.
Smart Images

Figure CN120673069A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of image data processing, and in particular to a synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning. Background Art
[0002] In practical applications, synthetic aperture radar (SAR) image recognition is used to automatically identify and classify specific objects in SAR images. Due to the difficulty of labeling SAR images, the number of SAR image samples is severely insufficient. Existing techniques typically train generative models using semi-supervised learning strategies. These models are then fine-tuned using a small number of labeled SAR images of new types, enabling the generative models to recognize unknown types and generate enhanced image samples of the corresponding types. This enhanced image sample is then used to train image recognition models, enabling SAR image recognition. However, traditional fine-tuning methods require iterative updates of all model parameters of the generative model. Given the limited number of labeled image samples of new types, these methods are prone to overfitting, resulting in a decline in the generalization performance of the generative model and the quality of image sample generation. This ultimately limits the scalability of SAR image recognition models. Summary of the Invention
[0003] The following is an overview of the subject matter described in detail in this disclosure. This overview is not intended to limit the scope of the claims.
[0004] The embodiments of the present disclosure provide a synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning, which improves the learning efficiency and adaptability of the image recognition model for synthetic aperture radar images of new sample types, thereby improving the scalability of the image recognition model.
[0005] The present disclosure provides a synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning, including:
[0006] Acquire a first synthetic aperture radar image sample, and add Gaussian noise to the first synthetic aperture radar image sample to obtain a noisy radar image sample;
[0007] Calling a diffusion model to perform noise prediction on the noisy radar image sample to obtain predicted noise, and training the diffusion model based on a difference between the predicted noise and the Gaussian noise to obtain the pre-trained diffusion model;
[0008] obtaining a second synthetic aperture radar image sample, adding a low-rank decomposition matrix bypass to the diffusion model, and incrementally fine-tuning the pre-trained diffusion model based on the second synthetic aperture radar image sample and the low-rank decomposition matrix bypass, wherein the second synthetic aperture radar image sample is used to fine-tune the diffusion model;
[0009] calling the incrementally fine-tuned diffusion model to generate a first enhanced synthetic aperture radar image sample, and performing incremental training on an image recognition model based on the first enhanced synthetic aperture radar image sample;
[0010] In response to the image recognition request, the trained image recognition model is called to perform image recognition on the synthetic aperture radar image to be processed, and a type corresponding to the synthetic aperture radar image to be processed is obtained.
[0011] The disclosed embodiments include at least the following advantageous effects: obtaining a first synthetic aperture radar image sample, adding Gaussian noise to the first synthetic aperture radar image sample to obtain a noisy radar image sample, invoking a diffusion model to perform noise prediction on the noisy radar image sample to obtain predicted noise, and training a diffusion model based on the difference between the predicted noise and the Gaussian noise to obtain a pre-trained diffusion model. This allows the pre-trained diffusion model to understand the relationship between noise and image, thereby predicting an accurate noise distribution. Furthermore, obtaining a second synthetic aperture radar image sample, adding a low-rank decomposition matrix bypass to the diffusion model, and incrementally fine-tuning the pre-trained diffusion model based on the second synthetic aperture radar image sample and the low-rank decomposition matrix bypass, enables the diffusion model to adapt to new synthetic aperture radar image types without changing the original model parameters, thereby reducing computational resource consumption during repeated training. Furthermore, the design of the low-rank decomposition matrix bypass reduces modifications to the original structure of the diffusion model, further reducing computational resource consumption during training. Invoking the incrementally fine-tuned diffusion model to generate a first enhanced synthetic aperture radar image sample, and incrementally training the image recognition model based on the first enhanced synthetic aperture radar image sample can enhance the learning efficiency and adaptability of the image recognition model to synthetic aperture radar images of new sample types. Therefore, in response to an image recognition request, invoking the trained image recognition model to perform image recognition on the synthetic aperture radar image to be processed enables the image recognition model to more accurately determine the type of the synthetic aperture radar image to be processed, thereby improving the scalability of the image recognition model.
[0012] Other features and advantages of the present disclosure will be set forth in the description which follows, and in part will be apparent from the description, or may be learned by practicing the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings are used to provide a further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation to the technical solution of the present disclosure.
[0014] Figure 1 An optional flowchart of a synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning provided in an embodiment of the present disclosure;
[0015] Figure 2 An optional schematic diagram of a forward noise addition training process provided in an embodiment of the present disclosure;
[0016] Figure 3 An optional schematic diagram of a backward denoising generation process provided in an embodiment of the present disclosure;
[0017] Figure 4 An optional schematic diagram of a diffusion model provided in an embodiment of the present disclosure;
[0018] Figure 5 An optional schematic diagram of an encoder module structure provided in an embodiment of the present disclosure;
[0019] Figure 6 An optional schematic diagram of a decoder module structure provided in an embodiment of the present disclosure;
[0020] Figure 7 An optional schematic diagram of a fine-tuning sub-model provided in an embodiment of the present disclosure;
[0021] Figure 8 An optional overall process of the synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning provided in the embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not intended to limit the present disclosure.
[0023] It should be noted that in various specific embodiments of the present disclosure, when it comes to the need to perform relevant processing based on data related to the characteristics of the target object, such as the target object attribute information or attribute information set, the permission or consent of the target object will be obtained first, and the collection, use and processing of these data will comply with relevant laws, regulations and standards. Among them, the target object can be a user. In addition, when the embodiment of the present disclosure needs to obtain the attribute information of the target object, the separate permission or separate consent of the target object will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary target object-related data for the normal operation of the embodiment of the present disclosure will be obtained.
[0024] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0025] To facilitate understanding of the technical solutions provided by the embodiments of the present disclosure, some key terms used in the embodiments of the present disclosure are explained here:
[0026] Synthetic Aperture Radar (SAR) imagery is a high-resolution radar image acquired using a synthetic aperture radar (SAR) system. SAR observes a target area by transmitting and receiving microwave signals. Using the coherence properties of radar signals, it processes and synthesizes echo signals received at different locations to produce a high-resolution image. SAR images can provide information about the target area's topography and features, regardless of lighting or weather conditions. They are widely used in fields such as Earth resource observation, environmental observation, and oceanographic observation.
[0027] Diffusion Model: It is a probabilistic generation model based on the diffusion process in physics. Its diffusion process includes a forward noise addition process and a backward denoising process. New data samples are generated by gradually adding noise to the data and learning the reverse denoising process.
[0028] Transformer: It is a deep learning architecture based on the attention mechanism. It adopts an encoder-decoder architecture. Through the self-attention mechanism and cross-attention mechanism, the model can automatically learn the weights of information at different positions when processing data, effectively capturing the dependencies in the data. It is widely used in natural language processing, image processing and other fields.
[0029] LORA (Low Rank Adaptation) fine-tuning technology: A method for efficiently fine-tuning pre-trained models. By introducing a bypass matrix based on the pre-trained model to adjust some model parameters, it achieves adaptation to specific tasks without significantly increasing the model storage and computational complexity. LORA fine-tuning technology is widely used in model fine-tuning in fields such as natural language processing and image processing. It can effectively improve the performance of the model on specific tasks while reducing the cost and difficulty of fine-tuning.
[0030] In practical applications, synthetic aperture radar (SAR) image recognition is used to automatically detect and classify specific objects in images. However, due to the unique nature of SAR imaging, it is difficult to directly identify image details with the naked eye, making image calibration a significant challenge. The high labor costs and technical barriers to calibration lead to a severe shortage of labeled image samples, making it difficult to train traditional deep learning models and limiting the application of SAR image recognition models. To alleviate the limited number of SAR images, existing technologies have adopted a semi-supervised learning strategy that combines a small number of labeled image samples with a large number of unlabeled image samples to reduce the generative model's reliance on labeled data. The pre-trained generative model is then retrained using a small number of labeled SAR images of a new type, enabling the generative model to recognize unknown types and generate enhanced image samples of the corresponding type. The enhanced image samples are then used to train the image recognition model to achieve SAR image recognition. However, existing generative models are prone to gradient vanishing or mode collapse when faced with high-noise, high-dimensional synthetic aperture radar images. Traditional fine-tuning methods require iterative updates of all model parameters of the generative model, which not only consumes a lot of computing resources but is also prone to overfitting problems due to the limited number of new types of image samples, resulting in a decline in the generalization performance of the generative model and the quality of image sample generation, ultimately restricting the application scalability of synthetic aperture radar image recognition models.
[0031] Based on this, the embodiments of the present disclosure provide a synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning, which improves the learning efficiency and adaptability of the image recognition model for synthetic aperture radar images of new sample types, thereby improving the scalability of the image recognition model.
[0032] Reference Figure 1 , Figure 1 An optional flowchart of a synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning provided in an embodiment of the present disclosure, the synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning includes but is not limited to the following steps S101 to S105.
[0033] Step S101: obtaining a first synthetic aperture radar image sample, and adding Gaussian noise to the first synthetic aperture radar image sample to obtain a noisy radar image sample.
[0034] Among them, the first synthetic aperture radar image sample is used for overall training of the diffusion model. The first synthetic aperture radar image sample is an unlabeled image. The processing process of the diffusion model includes a forward denoising process and a backward denoising process. Gaussian noise is noise pre-set according to a specific noise distribution and is used in the forward denoising process of the diffusion model.
[0035] Specifically, a first synthetic aperture radar image sample is obtained and Gaussian noise is added to the first synthetic aperture radar image sample time-step by time. The noise intensity at each time step is gradually increased according to the noise addition strategy until the image is converted into pure random noise, thereby obtaining a noisy radar image sample. The noise addition process for the first synthetic aperture radar image sample is a Markov process, where the image state of the first synthetic aperture radar image sample at each time step is generated based on the superposition of the image state of the first synthetic aperture radar image sample at the previous time step and the noise.
[0036] Step S102: calling a diffusion model to perform noise prediction on the noisy radar image sample to obtain predicted noise, and training the diffusion model based on the difference between the predicted noise and the Gaussian noise to obtain the pre-trained diffusion model.
[0037] The diffusion model is used to recover the first synthetic aperture radar image sample from the noisy radar image sample, the predicted noise is obtained by predicting the noisy radar image sample through the noise predictor of the diffusion model, and the predicted noise is the predicted noise distribution.
[0038] When training the diffusion model, forward noise addition training is a key step in its training. During the forward noise addition process, the diffusion model will gradually add noise in multiple time steps until the first synthetic aperture radar image sample is converted into pure random noise. Specifically, the forward noise addition training process can be referred to Figure 2 , Figure 2 An optional schematic diagram of a forward noise training process provided by an embodiment of the present disclosure includes adding Gaussian noise to a first synthetic aperture radar image sample to obtain a noisy radar image sample contaminated by the Gaussian noise. The noisy radar image sample is input into a noise predictor for noise prediction to obtain predicted noise. The difference between the predicted noise and the Gaussian noise is calculated to determine the noise prediction loss, and the reverse gradient of the difference is calculated to update the model parameters of the noise predictor to train the noise predictor. The noise prediction loss can be a mean squared error loss (MSE), and the noise prediction loss L MS2It can be expressed by the following formula, where n represents a total of n first synthetic aperture radar image samples, represents the Gaussian noise added to the first SAR image sample, y i represents the predicted noise of the i-th first synthetic aperture radar image sample.
[0039]
[0040] Through the forward denoising process, the diffusion model can understand the relationship between noise and the first synthetic aperture radar image sample, and learn how to effectively eliminate noise and restore important information in the first synthetic aperture radar image sample, thereby improving the accuracy of the diffusion model in predicting noise.
[0041] After the forward denoising process is completed to train the diffusion model, the trained diffusion model is used to implement the backward denoising generation process to achieve the generation of the first enhanced synthetic aperture radar image sample. Figure 3 , Figure 3 This is an optional schematic diagram of the backward denoising generation process provided by an embodiment of the present disclosure. A randomly sampled random Gaussian noise is obtained. The noise distribution of the noisy radar image sample is input into a noise predictor. The noise distribution of the noisy radar image sample is predicted based on the random Gaussian noise to obtain a predicted noise. The predicted noise is removed from the noise distribution of the noisy radar image sample to obtain a denoising result. The denoising result is used to update the noise distribution of the noisy radar image sample. The updated noise distribution is input into the noise predictor to continue noise prediction for the next time step. Each time step depends on the denoising result of the previous time step. In each backward step, the diffusion model uses the noisy radar image sample at the current time step and the denoising result output by the noise predictor to calculate a denoising update direction. The noisy radar image sample is then updated based on this update direction. Through iterative updates over multiple backward time steps, the noise of the noisy radar image sample is gradually reduced and its pixel values are adjusted to restore the noisy radar image sample. This gradually approximates the features and structure of the first synthetic aperture radar image sample to generate a second enhanced synthetic aperture radar image sample. After completing the forward denoising training and backward denoising generation process based on the first SAR image sample, a pre-trained diffusion model is obtained. This progressive denoising process maintains high-quality restoration of the noisy radar image sample at each time step, avoiding blurring and distortion. Furthermore, the diffusion model accurately restores detailed features in the noisy radar image sample through a gradual denoising process. This enables the generated second enhanced SAR image sample to not only retain rich semantic information but also capture subtle details in the noisy radar image sample, providing high-quality training data for subsequent image recognition model training.
[0042] It should be noted that since the training of the diffusion model is based on the forward denoising process, the backward denoising process can utilize the knowledge learned through the forward denoising process. Therefore, in the backward denoising process, the diffusion model can achieve efficient image restoration from the learned noise removal strategy.
[0043] In one possible implementation, the diffusion model includes a noise predictor, which utilizes a Transformer architecture and consists of multiple encoder modules and decoder modules. These modules are cascaded, with encoder and decoder modules at the same level connected via residual connections. When the diffusion model is invoked to predict noise on noisy radar image samples to obtain predicted noise, the encoder module can be invoked to encode the noisy radar image samples, and the decoder module can be invoked to decode the encoded results of the encoder module to obtain the predicted noise. The encoder module is used to convert the input noisy radar image samples into corresponding image feature vectors.
[0044] Specifically, refer to Figure 4 , Figure 4 An optional schematic diagram of a diffusion model provided in an embodiment of the present disclosure, Figure 4 The diffusion generative network in
[15] consists of two encoder modules, two decoder modules, and a latent feature module. Encoder module A and decoder module B are at the same level, and encoder module B and decoder module A are at the same level. The output dimensions of encoder modules at the same level are the same as those of decoder modules. The noisy radar image is input into encoder module A for encoding. The encoding result of encoder module A is then input into encoder module B. The encoding result of encoder module B is then input into the latent feature module for the same processing as the encoder module. The output of the latent feature module and the encoding result of encoder module B are input into decoder module A for decoding. The decoding result of decoder module A and the encoding result of encoder module A are then input into decoder module B for decoding, and the predicted noise is output. The noise predictor is implemented using the Transformer architecture and globally models noisy radar image samples. This effectively captures the long-range dependency features in noisy radar image samples, thereby improving the accuracy of the noise predictor's noise prediction.
[0045] In one possible implementation, in the process of calling the encoder module to encode the noisy radar image sample, the noisy radar image sample can be input into the first encoder module, the noisy radar image sample is sliced to obtain image slices, the image slices are converted into image slice vectors, the image slice vectors are input into the multi-head attention module for feature extraction to obtain the residual connection vector, the features extracted by the multi-head attention module are downsampled to obtain the backbone feature vector, and the backbone feature vector is input into the next-level encoder module for encoding until the backbone feature vector output by the last-level encoder module is obtained.
[0046] Specifically, refer to Figure 5 , Figure 5 This is an optional schematic diagram of the encoder module structure provided by an embodiment of the present disclosure. A noisy radar image sample is input to the first encoder module. In the first encoder module, the noisy radar image sample is convolved and sliced into image sample slices. The image sample slices are converted into image slice vectors. The image slice vectors are respectively input as query vectors Q, key vectors K, and value vectors V into the multi-head attention module for feature extraction. The feature extraction process of the multi-head attention module can be expressed by the following formula:
[0047]
[0048] Among them, d k is the dimension size of the image slice vector, the softmax function can be expressed as The features extracted by the multi-head attention module are used as residual connection vectors for use by the decoder module at the same level. Furthermore, the features are input into the downsampling convolution layer to reduce the length and width dimensions of the features. The reduced features are then input into the next-level encoder module as backbone feature vectors for encoding, until the backbone feature vector output by the final encoder module is obtained. Processing noisy radar image samples through the encoder module can reduce the dimensionality of the noisy radar image samples, thereby reducing the computing power required for subsequent processing. Furthermore, the encoder module can discard redundant information in the noisy radar image samples during processing, allowing the backbone feature vector output by the final encoder module to more centrally retain information reflecting the essential characteristics of the noisy radar image samples, which helps improve the accuracy of subsequent noise prediction.
[0049] It should be noted that the last-level encoder module can be a latent feature module. The processing process of the latent feature module is the same as that of the encoder module, but the features extracted by the multi-head attention module are not downsampled, and the features extracted by the multi-head attention module are directly input into the decoder module.
[0050] In one possible implementation, when calling the decoder module to decode the encoding result of the encoder module to obtain the predicted noise, the residual connection vector output by the encoder module at the same level and the backbone feature vector output by the subsequent encoder module are input into the first decoder module, the backbone feature vector output by the last encoder module after the upsampling operation is added to the residual connection vector output by the encoder module at the same level and then input into the multi-head attention module to obtain the decoding feature, and the decoding feature is input into the upper-level decoder module for decoding until the predicted noise output by the highest-level decoder module is obtained.
[0051] Specifically, refer to Figure 6 , Figure 6 An optional schematic diagram of a decoder module structure provided for an embodiment of the present disclosure. The backbone feature vector is differentially upsampled through an upsampling layer so that the size and dimension of the upsampled backbone feature vector are the same as the residual connection vector, and the upsampled backbone feature vector and the residual connection vector are added element by element to obtain a fused feature vector g. The fused feature vector g is input into the multi-head attention module as the query vector Q and the value vector V respectively, and the residual connection vector is input into the multi-head attention module as the key vector K to obtain the decoding features extracted by the multi-head attention module, and the decoding features are input into the upper-level decoder module for decoding until the decoding features output by the top-level decoder module are obtained, and the decoding features output by the top-level decoder module are used as prediction noise. By decoding based on the backbone feature vector and the residual connection vector, the cross-attention effect of the residual connection vector guiding the multi-head attention module to perform feature extraction and decoding can be achieved, thereby improving the feature extraction capability of the decoder module.
[0052] Step S103: obtaining a second synthetic aperture radar image sample, adding a low-rank decomposition matrix bypass to the diffusion model, and incrementally fine-tuning the pre-trained diffusion model based on the second synthetic aperture radar image sample and the low-rank decomposition matrix bypass.
[0053] The low-rank decomposition matrix bypass is a branch added to the pre-trained diffusion model. The low-rank decomposition matrix bypass includes multiple bypass matrices. The bypass matrix is a low-rank matrix obtained by decomposing the original weight matrix in the pre-trained diffusion model. The low-rank decomposition matrix bypass is added next to the decomposed original weight matrix and is parallel to the decomposed original weight matrix. The second synthetic aperture radar image sample is used to fine-tune the diffusion model. The second synthetic aperture radar image sample is an annotated image. The sample type of the second synthetic aperture radar image sample includes the sample type of the first synthetic aperture radar image sample and other new sample types. Incremental fine-tuning is used to allow the diffusion model to learn sample types that have not appeared in the pre-training stage. It should be noted that the method for fine-tuning the pre-trained diffusion model is the LoRA fine-tuning method.
[0054] In one possible implementation, when incrementally fine-tuning a pre-trained diffusion model based on a second synthetic aperture radar image sample and a low-rank decomposition matrix bypass, the second synthetic aperture radar image sample is obtained, fine-tuning parameters are determined based on the similarities and differences in sample types between the first synthetic aperture radar image sample and the second synthetic aperture radar image sample, the low-rank decomposition matrix bypass is added to the encoder module and decoder module of the noise predictor, and the pre-trained diffusion model is incrementally fine-tuned based on the low-rank decomposition matrix bypass, the second synthetic aperture radar image sample, and the fine-tuning parameters. The fine-tuning parameters are used to set the low-rank decomposition matrix bypass and control the contribution of the low-rank decomposition matrix bypass to the fine-tuning process, and include the decomposition matrix rank, bypass weight, and bypass drop probability.
[0055] Specifically, a second synthetic aperture radar image sample is obtained. The first synthetic aperture radar image sample and the second synthetic aperture radar image sample each have a corresponding sample type. The sample type includes the type of object displayed in the image, the type of scene displayed in the image, etc., which are not specifically limited in this application. The fine-tuning task type is determined based on the similarities and differences in sample types between the first synthetic aperture radar image sample and the second synthetic aperture radar image sample, and the corresponding fine-tuning parameters are determined based on the fine-tuning task type. Next, a low-rank decomposition matrix bypass is added to the encoder module and decoder module of the noise predictor to construct a fine-tuning submodel, and the fine-tuning submodel is set based on the fine-tuning parameters. The second synthetic aperture radar image samples are classified according to sample type, and the second synthetic aperture radar image samples of different sample types are respectively input into the fine-tuning submodel for incremental fine-tuning, thereby obtaining a fine-tuning submodel corresponding to each sample type.
[0056] In one possible implementation, when determining fine-tuning parameters based on the similarities and differences in sample types between a first SAR image sample and a second SAR image sample, if the sample type of the second SAR image sample is the same as that of the first SAR image sample, the diffusion model is initially fine-tuned, and the corresponding fine-tuning parameters are determined based on the initial fine-tuning. If the sample type of the second SAR image sample is different from that of the first SAR image sample, the diffusion model is incrementally fine-tuned, and the corresponding fine-tuning parameters are determined based on the incremental fine-tuning. The initial fine-tuning allows the diffusion model to learn sample types that have already appeared in the pre-training phase.
[0057] Specifically, the fine-tuning task type is determined based on the similarities and differences in sample types between the first and second SAR image samples. Fine-tuning task types include initial fine-tuning and incremental fine-tuning. If the sample type of the second SAR image sample is the same as the sample type of the first SAR image sample, the diffusion model is initially fine-tuned based on the second SAR image sample, and the corresponding fine-tuning parameters are determined based on the fine-tuning task type for the initial fine-tuning. If the sample type of the second SAR image sample is different from the sample type of the first SAR image sample, the diffusion model is incrementally fine-tuned based on the second SAR image sample, and the corresponding fine-tuning parameters are determined based on the fine-tuning task type for the incremental fine-tuning. Table 1 lists the fine-tuning parameters for the fine-tuning task types provided in embodiments of the present disclosure.
[0058] Table 1 Fine-tuning parameters for fine-tuning task types
[0059] Fine-tuning parameters Initial fine-tuning Incremental fine-tuning Decomposition matrix rank 36 8 Bypass Weight 0.5 0.3 Bypass drop probability 0.1 0.5
[0060] Table 1 is used as an example to illustrate the fine-tuning parameter settings for different fine-tuning task types. During initial fine-tuning, a relatively large number of second SAR image samples of the same sample type are available for initial fine-tuning. A larger decomposition matrix rank is set to capture the characteristics of the second SAR image samples. The bypass weights are comparable to those of the pre-trained diffusion model, and a smaller bypass dropout probability is used. During incremental fine-tuning, a relatively small number of second SAR image samples of new sample types are available for incremental fine-tuning. A smaller decomposition matrix rank is used to prevent overfitting of the diffusion model and avoid model forgetting caused by new sample types destroying the learned sample type characteristics of the diffusion model. Furthermore, a smaller bypass weight is set to allow the pre-trained diffusion model to have a greater dominant position, and a larger bypass dropout probability is set to prevent the low-rank decomposition matrix from overfitting the new sample types during training.
[0061] In one possible implementation, the encoder module and the decoder module both include a multi-head attention module, the multi-head attention module includes a first linear layer as a query weight matrix and a second linear layer as a key weight matrix, and in the process of adding a low-rank decomposition matrix bypass to the encoder module and the decoder module of the noise predictor, for the first linear layer as the query weight matrix, the query weight matrix is decomposed into a first bypass matrix and a second bypass matrix, for the second linear layer as the key weight matrix, the key weight matrix is decomposed into a third bypass matrix and a fourth bypass matrix, the first bypass matrix and the second bypass matrix are bypass-connected to the query weight matrix as low-rank decomposition matrix bypasses of the query weight matrix, and the third bypass matrix and the fourth bypass matrix are bypass-connected to the key weight matrix as low-rank decomposition matrix bypasses of the key weight matrix. Wherein, the ranks of the first bypass matrix, the second bypass matrix, the third bypass matrix, and the fourth bypass matrix are all the same.
[0062] Specifically, refer to Figure 7 , Figure 7 This is an optional schematic diagram of a fine-tuning sub-model provided in an embodiment of the present disclosure. The fine-tuning sub-model is applied to the noise predictor of the diffusion model and is obtained by adding a low-rank decomposition matrix bypass to the linear layer within the multi-head attention module of the encoder module and decoder module of the noise predictor. When fine-tuning using the fine-tuning sub-model, the weight matrix of the pre-trained diffusion model is frozen, i.e. Figure 7 The query weight matrix and key weight matrix in
[15] are used. When a second synthetic aperture radar image sample of any sample type passes through the first and second bypass matrices, the weights of the first and second bypass matrices are updated, and the first updated weight matrix for the bypass is output. The first updated weight matrix and the query weight matrix are weighted and summed based on the bypass weights to obtain a first weight matrix. When a second synthetic aperture radar image sample of any sample type passes through the third and fourth bypass matrices, the weights of the third and fourth bypass matrices are updated, and the second updated weight matrix for the bypass is output. The second updated weight matrix and the key weight matrix are weighted and summed based on the bypass weights to obtain a second weight matrix. The first weight matrix is multiplied by the second weight matrix and then passed through the value weight matrix to obtain the output corresponding to the sample type. The low-rank decomposition matrix bypass design of the fine-tuning sub-model reduces modifications to the original structure of the diffusion model, effectively reducing computational resource consumption during training.
[0063] It should also be noted that when fine-tuning using the fine-tuning sub-model, the bypass update weight matrix is randomly discarded with a certain probability according to the set bypass discard probability, and only the weight matrix of the pre-trained diffusion model (such as the query weight matrix and the key weight matrix) is retained to prevent overfitting problems during the fine-tuning process.
[0064] Step S104: calling the incrementally fine-tuned diffusion model to generate a first enhanced synthetic aperture radar image sample, and performing incremental training on the image recognition model based on the first enhanced synthetic aperture radar image sample.
[0065] The first enhanced SAR image sample is of a different sample type from the first and second SAR image samples and is generated by a diffusion model based on SAR images of other sample types. The image recognition model is used to identify and classify specific objects in SAR images of each sample type.
[0066] In one possible implementation, before incrementally training the image recognition model based on the first enhanced SAR image sample, the incrementally fine-tuned diffusion model is invoked to generate a second enhanced SAR image sample based on the sample types of the first and second SAR images. A sample type label is assigned to the second enhanced SAR image sample, and the second enhanced SAR image sample is input into the image recognition model to obtain a predicted type label. Initial training of the image recognition model is then performed based on the difference between the predicted type label and the sample type label. The image recognition model may be a Resnet-18 residual network, and the second enhanced SAR image sample is of the same sample type as the first and second SAR image samples, generated by the diffusion model based on the first and second SAR image samples.
[0067] Specifically, a fine-tuning sub-model in the incrementally fine-tuned diffusion model is invoked to generate second enhanced SAR image samples of multiple sample types based on the sample types of the first SAR image sample and the second sample SAR image. The second enhanced SAR image samples are labeled using the corresponding sample types as sample type labels. The second enhanced SAR image samples and the corresponding sample type labels are input into an image recognition model for image recognition, resulting in predicted type labels for the second enhanced SAR image samples. The predicted type labels are then compared with the sample type labels, and the image recognition model is initially trained based on the difference between the predicted type labels and the sample type labels.
[0068] It should be noted that after completing the training of the image recognition model, it only needs to be deployed to the application or device. Therefore, by training the image recognition model with the second enhanced synthetic aperture radar image samples generated by the fine-tuning sub-model, the model parameters and complexity of the image recognition model can be effectively reduced.
[0069] In one possible implementation, the image recognition model includes a linear classification layer, and a diffusion model after incremental fine-tuning is called to generate a first enhanced synthetic aperture radar image sample. When incrementally training the image recognition model based on the first enhanced synthetic aperture radar image sample, the diffusion model after incremental fine-tuning can be called to generate the first enhanced synthetic aperture radar image sample, the model parameters of the image recognition model can be frozen, the first enhanced synthetic aperture radar image sample and the second enhanced synthetic aperture radar image sample are mixed in a certain proportion and input into the image recognition model, and the classification linear layer of the image recognition model is trained and updated.
[0070] Specifically, an image of a different sample type from the first and second SAR image samples is acquired to obtain a SAR image of the new sample type. The diffusion model is incrementally fine-tuned based on the SAR image of the new sample type, and then the fine-tuning sub-model within the incrementally fine-tuned diffusion model is invoked to generate first enhanced SAR image samples for each new sample type. The model parameters of the image recognition model are frozen, leaving only the last processing layer of the image recognition model, the linear classification layer, open. The first enhanced SAR image sample and the second enhanced SAR image sample are mixed in a predetermined ratio to generate a mixed image sample set. For example, the first enhanced SAR image sample and the second enhanced SAR image sample can be mixed in a ratio of 3:1. The mixed image sample set is then input into the image recognition model, and the linear classification layer of the image recognition model is trained based on the mixed image sample set, and the parameters and weights of the linear classification layer are updated.
[0071] Step S105: In response to the image recognition request, the trained image recognition model is called to perform image recognition on the synthetic aperture radar image to be processed, and the type corresponding to the synthetic aperture radar image to be processed is obtained.
[0072] Specifically, after the image recognition model completes incremental training, it is deployed to downstream applications or recognition devices. When an image recognition request is received, the downstream application or recognition device responds to the image recognition request, calls the image recognition model to perform image recognition on the synthetic aperture radar image to be processed, and obtains the category corresponding to the synthetic aperture radar image to be processed.
[0073] In a possible implementation, the synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning provided by the embodiment of the present disclosure can be applied to environmental observation scenarios, referring to Figure 8 , Figure 8 An optional overall process of the synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning provided in the embodiment of the present disclosure.
[0074] The synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning provided in the embodiments of the present disclosure is jointly implemented based on a diffusion model and an image recognition model.
[0075] First, a diffusion model is constructed. This model includes a noise predictor, which adopts a transformer structure and consists of multiple encoder modules, multiple decoder modules, and a latent feature module. These multiple encoder modules and decoder modules are cascaded, and encoder and decoder modules at the same level are connected via residual connections. Both encoder and decoder modules include a multi-head attention module. The diffusion model is primarily used in the training phase. During this phase, it extracts features from the first and second synthetic aperture radar image samples and generates second enhanced synthetic aperture radar image samples for training the image recognition model. After training, the image recognition model is deployed in downstream applications or recognition devices for image recognition.
[0076] Next, an unsupervised learning strategy is used to train the diffusion model. An unlabeled first SAR image sample is obtained and used to train the diffusion model as a whole. The first SAR image sample is input into the diffusion model, where forward noise is added. Gaussian noise is added to the first SAR image sample time-step by time until the first SAR image sample is converted into pure random noise, resulting in a noisy radar image sample. During the forward noise training process, the noisy radar image sample is input into a noise predictor for noise prediction, obtaining predicted noise. The predicted noise is then compared with Gaussian noise, and a noise prediction loss is determined based on the difference between the predicted noise and the Gaussian noise. The forward noise training process of the diffusion model is then trained based on the noise prediction loss.
[0077] Next, after the diffusion model is trained using the forward denoising process, the trained diffusion model is used to perform the backward denoising process. The noise predictor predicts the noise distribution in the noisy radar image samples based on randomly sampled Gaussian noise, obtaining the predicted noise. The predicted noise is then removed from the noisy radar image samples to obtain the denoised results. The denoised results are then used to update the noise distribution of the noisy radar image samples. The updated noise distribution is then fed back into the noise predictor, and the noise prediction continues for the next time step until the denoised results for the final time step are obtained. The second enhanced SAR image samples are generated based on the predicted results from the final time step. After completing the forward denoising training and backward denoising process based on the unlabeled first SAR image samples, the pre-trained diffusion model is obtained.
[0078] When a noisy radar image sample is input to the noise predictor for noise prediction, the first encoder module of the noise predictor performs convolution on the noisy radar image sample, slicing it into multiple image sample slices. These image sample slices are then converted into image slice vectors. These image slice vectors are then fed into the multi-head attention mechanism as query vectors, key vectors, and value vectors for feature extraction. The features extracted by the attention mechanism are then used as residual connection vectors for use by the decoder module at the same level. Simultaneously, the features extracted by the attention mechanism are downsampled to obtain a backbone feature vector, which is then fed into the encoder module at the next level for further encoding, until the backbone feature vector of the final encoder module is obtained. The final encoder module is the latent feature module, which undergoes the same processing as the encoder module, but without the downsampling operation. The backbone feature vector output by the latent feature module is input into the decoder module, and the backbone feature vector is upsampled in the decoder module so that the backbone feature vector has the same size and dimension as the residual connection vector of the encoder module at the same level. The upsampled backbone feature vector and the residual connection vector of the encoder module at the same level are added element by element, and then input into the multi-head attention mechanism as the query vector and value vector respectively. The residual connection vector is input into the multi-head attention mechanism as the key vector, and feature extraction is performed in the multi-head attention mechanism. The extracted decoding features are used as the input of the next-level decoder module until the decoding features of the top-level decoder module are obtained, and the decoding features output by the top-level decoder module are used as prediction noise.
[0079] Next, a fine-tuning sub-model is constructed and used to fine-tune the pre-trained diffusion model. To better fine-tune the pre-trained diffusion model, the fine-tuning task is divided into two types: initial fine-tuning and incremental fine-tuning. Fine-tuning parameters are set for each of these two types. A low-rank decomposition matrix bypass is then added to the multi-head attention module of all encoder and decoder modules in the noise predictor to construct the fine-tuning sub-model. The low-rank decomposition matrix bypass consists of multiple bypass matrices, each obtained by performing a rank decomposition of the original weight matrix of the pre-trained diffusion model. A second synthetic aperture radar image sample is obtained. When fine-tuning the fine-tuning sub-model, the first step is to determine whether the sample type of the second synthetic aperture radar image sample is the same as that of the first synthetic aperture radar image sample. If so, initial fine-tuning is performed and the corresponding fine-tuning parameters are set for the fine-tuning sub-model. If not, incremental fine-tuning is performed and the corresponding fine-tuning parameters are set for the fine-tuning sub-model. The second synthetic aperture radar image sample is input into the diffusion model, and forward denoising training is performed based on the fine-tuning sub-model. Then, a backward denoising process is performed based on the trained forward denoising process to complete the fine-tuning of the diffusion model.
[0080] Next, the fine-tuned diffusion model is invoked, and the fine-tuned sub-model is used to generate second enhanced SAR image samples of the same sample type as the first and second SAR image samples. Initial training of the image recognition model is performed based on the second enhanced SAR image samples. Next, an image of a different sample type from the first and second SAR image samples is obtained, i.e., a SAR image of a new sample type. After incremental fine-tuning the diffusion model based on the SAR image of the new sample type, the fine-tuned sub-model is used to generate first enhanced SAR image samples of the new sample type. Incremental training of the image recognition model is then performed based on the first enhanced SAR image samples.
[0081] After the image recognition model completes incremental training, the image recognition model is deployed to downstream applications or recognition devices to perform image recognition on the synthetic aperture radar image to be processed and obtain the category of the synthetic aperture radar image to be processed.
[0082] The disclosed embodiments provide a synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning. By using unlabeled first synthetic aperture radar image samples to perform unsupervised training on the diffusion model, the diffusion model can effectively reduce its dependence on labeled data. This is especially true when synthetic aperture radar image annotation is difficult. This unsupervised training strategy provides an efficient data-driven method for the diffusion model. Secondly, by adding a low-rank decomposition matrix bypass and introducing the LORA fine-tuning technology to fine-tune the pre-trained diffusion model, the diffusion model can flexibly adapt to various types without large-scale parameter updates. The design of the low-rank decomposition matrix bypass does not excessively modify the original model structure of the diffusion model. Combined with the LORA fine-tuning technology, the diffusion model is subjected to low-rank adaptive fine-tuning, which can reduce the number of parameters and computational complexity during the training process, thereby improving the training efficiency of the diffusion model while maintaining high accuracy. Furthermore, by incrementally fine-tuning the diffusion model and incrementally training the image recognition model, the diffusion model and image recognition model can learn new types of SAR images when they appear, requiring only a small number of previously labeled SAR image samples for fine-tuning. This incremental learning capability enables the SAR image recognition method based on diffusion-generated low-rank fine-tuning provided in the present embodiment to adapt to dynamically changing application scenarios. It can quickly and effectively adjust the diffusion model and image recognition model in the face of constantly changing SAR images, thereby reducing the computational resource consumption of repeated training. Finally, the SAR image recognition method based on diffusion-generated low-rank fine-tuning provided in the present embodiment decouples the tasks of SAR image feature extraction, incremental fine-tuning, and image generation from the diffusion model. This simplifies the initial and incremental training processes of the image recognition model. Furthermore, the image recognition model utilizes a lightweight network with a small number of model parameters, which facilitates its deployment in downstream applications or recognition devices across multiple scenarios, improving its scalability.
[0083] The terms "first," "second," "third," "fourth," and the like (if any) in the specification of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate to describe embodiments of the present disclosure, e.g., capable of being implemented in orders other than those illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or elements is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0084] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0085] It should be understood that in the description of the embodiments of the present disclosure, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood to exclude the number itself, and above, below, within, etc. are understood to include the number itself.
[0086] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0087] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0088] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0089] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.
[0090] It should also be understood that the various implementations provided in the embodiments of the present disclosure can be arbitrarily combined to achieve different technical effects.
[0091] The above is a specific description of the preferred implementation of the present disclosure, but the present disclosure is not limited to the above implementation. Technical personnel familiar with the art can also make various equivalent modifications or substitutions under the shared conditions that do not violate the spirit of the present disclosure. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present disclosure.
Claims
1. A synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning, characterized in that: include: Acquire a first synthetic aperture radar image sample, and add Gaussian noise to the first synthetic aperture radar image sample to obtain a noisy radar image sample; Calling a diffusion model to perform noise prediction on the noisy radar image sample to obtain predicted noise, and training the diffusion model based on a difference between the predicted noise and the Gaussian noise to obtain the pre-trained diffusion model; obtaining a second synthetic aperture radar image sample, adding a low-rank decomposition matrix bypass to the diffusion model, and incrementally fine-tuning the pre-trained diffusion model based on the second synthetic aperture radar image sample and the low-rank decomposition matrix bypass, wherein the second synthetic aperture radar image sample is used to fine-tune the diffusion model; calling the incrementally fine-tuned diffusion model to generate a first enhanced synthetic aperture radar image sample, and performing incremental training on an image recognition model based on the first enhanced synthetic aperture radar image sample; In response to the image recognition request, the trained image recognition model is called to perform image recognition on the synthetic aperture radar image to be processed, and a type corresponding to the synthetic aperture radar image to be processed is obtained.
2. The synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning according to claim 1, characterized in that: The diffusion model includes a noise predictor, which adopts a Transformer architecture. The noise predictor includes multiple encoder modules and multiple decoder modules. The multiple encoder modules and the multiple decoder modules are cascaded in sequence, and the encoder modules and the decoder modules at the same level are connected via residual connections. The diffusion model is called to perform noise prediction on the noisy radar image sample to obtain predicted noise, including: The encoder module is called to encode the noisy radar image sample, and the decoder module is called to decode the encoding result of the encoder module to obtain the predicted noise.
3. The synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning according to claim 1, characterized in that: The acquiring of a second synthetic aperture radar image sample, adding a low-rank decomposition matrix bypass to the diffusion model, and incrementally fine-tuning the pre-trained diffusion model based on the second synthetic aperture radar image sample and the low-rank decomposition matrix bypass, comprises: Acquiring the second synthetic aperture radar image sample; determining a fine-tuning parameter according to similarities and differences in sample types between the first synthetic aperture radar image sample and the second synthetic aperture radar image sample; A low-rank decomposition matrix bypass is added to the encoder module and the decoder module of the noise predictor, and the pre-trained diffusion model is incrementally fine-tuned based on the low-rank decomposition matrix bypass, the second synthetic aperture radar image sample, and the fine-tuning parameters.
4. The synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning according to claim 3, characterized in that: The determining of the fine-tuning parameter according to the similarities and differences in sample types between the first synthetic aperture radar image sample and the second synthetic aperture radar image sample includes: If the sample type of the first synthetic aperture radar image sample is the same as the sample type of the second synthetic aperture radar image sample, performing initial fine-tuning on the diffusion model, and determining the corresponding fine-tuning parameter according to the initial fine-tuning; If the sample type of the first synthetic aperture radar image sample is different from the sample type of the second synthetic aperture radar image sample, incremental fine-tuning is performed on the diffusion model, and the corresponding fine-tuning parameter is determined according to the incremental fine-tuning.
5. The synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning according to claim 3, characterized in that: The encoder module and the decoder module each include a multi-head attention module, the multi-head attention module including a first linear layer as a query weight matrix and a second linear layer as a key weight matrix, and adding a low-rank decomposition matrix bypass to the encoder module and the decoder module of the noise predictor, comprising: For the first linear layer serving as the query weight matrix, decomposing the query weight matrix into a first bypass matrix and a second bypass matrix; for the second linear layer serving as the key weight matrix, decomposing the key weight matrix into a third bypass matrix and a fourth bypass matrix; The first bypass matrix and the second bypass matrix are bypass-connected to the query weight matrix as the low-rank decomposition matrices of the query weight matrix, and the third bypass matrix and the fourth bypass matrix are bypass-connected to the key weight matrix as the low-rank decomposition matrices of the key weight matrix.
6. The synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning according to claim 2, characterized in that: The calling the encoder module to encode the noisy radar image sample includes: Inputting the noisy radar image sample into the first encoder module, slicing the noisy radar image sample to obtain image slices, and converting the image slices into image slice vectors; The image slice vector is input into the multi-head attention module for feature extraction to obtain a residual connection vector, the features extracted by the multi-head attention module are downsampled to obtain a backbone feature vector, and the backbone feature vector is input into the encoder module of the next level for encoding until the backbone feature vector output by the encoder module of the last level is obtained.
7. The synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning according to claim 2, characterized in that: The calling of the decoder module to decode the encoding result of the encoder module to obtain the predicted noise includes: Inputting the residual connection vector output by the encoder module at the same level and the backbone feature vector output by the last encoder module into the first decoder module; The backbone feature vector output by the last encoder module after the upsampling operation is added to the residual connection vector output by the encoder module at the same level and then input into the multi-head attention module to obtain the decoding feature, and the decoding feature is input into the decoder module at the previous level for decoding until the predicted noise output by the decoder module at the highest level is obtained.
8. The synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning according to claim 1, characterized in that: The image recognition model includes a linear classification layer, calling the incrementally fine-tuned diffusion model to generate a first enhanced synthetic aperture radar image sample, and performing incremental training on the image recognition model based on the first enhanced synthetic aperture radar image sample, including: calling the incrementally fine-tuned diffusion model to generate a first enhanced synthetic aperture radar image sample, where the first enhanced synthetic aperture radar image sample is an image of a different sample type from the first synthetic aperture radar image sample and the second synthetic aperture radar image sample; Freeze the model parameters of the image recognition model, mix the first enhanced synthetic aperture radar image sample and the second enhanced synthetic aperture radar image sample in a certain proportion and input the mixed sample into the image recognition model, and train and update the classification linear layer of the image recognition model.
9. The synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning according to claim 8, characterized in that: Before incrementally training the image recognition model based on the first enhanced synthetic aperture radar image sample, the synthetic aperture radar image recognition method based on diffusion-generated low-rank fine-tuning further includes: calling the incrementally fine-tuned diffusion model to generate a second enhanced synthetic aperture radar image sample based on the sample types of the first synthetic aperture radar image sample and the second sample synthetic aperture radar image, and assigning a sample type label to the second enhanced synthetic aperture radar image sample; The second enhanced synthetic aperture radar image sample is input into the image recognition model to obtain a predicted type label, and the image recognition model is initially trained based on a difference between the predicted type label and the sample type label.