SAR image classification model training method and related device

By adopting iterative training and distillation techniques in the SAR image classification model, the shortcomings of the existing models in classification effect and generalization capabilities are solved, and more efficient new category identification and data distribution imbalance processing are achieved.

CN119992210AActive Publication Date: 2025-05-13NANKAI UNIV +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510124917.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-13
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

There is a gap in the classification effect of existing SAR image classification models, especially when facing data distribution imbalance and new category recognition, the model's generalization ability is insufficient.

Method used

A training method of SAR image classification model is adopted, through iterating the training data multiple times, using the distillation technology of the student model and the teacher model, combined with the shared encoder of the perceptual projection module, the total loss is calculated and the model parameters are updated to improve the classification effect of the model.

Benefits of technology

By improving the long-tail problem of training data, the generalization ability of the model is improved, the classification effect of the SAR image classification model is improved, and the new categories can be more effectively identified and the problem of unbalanced data distribution can be handled.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992210A_ABST
    Figure CN119992210A_ABST
Patent Text Reader

Abstract

The invention discloses an SAR image classification model training method and a related device, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the weighted fusion and centralization processing of a teacher projection head based on a weight factor, and obtaining a distillation balance projection head, and calculating the total loss of iteration based on the truth value label, the distillation balance projection head, the student projection head, the first projection head and the second projection head, and because the weight factor is inversely correlated with the number of images belonging to the image category of the sample SAR image in the training data, the long tail problem of the training data can be improved, and further, the training efficiency is improved. According to the method, the total loss is obtained by performing supervised and unsupervised combined training on the distillation balance projection head, the student projection head, the first projection head and the second projection head based on the truth value labels of all sample SAR images in the iterative training data, so that the training effect is improved, and the classification effect of the SAR image classification model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a training method and related devices for a SAR image classification model. Background Art

[0002] Synthetic Aperture Radar (SAR) is an active microwave remote sensing imaging radar that acquires images by emitting coherent electromagnetic waves to illuminate the surface and then receiving scattered echoes from surface targets. SAR images reflect the scattering characteristics of microwaves from objects and targets, and have the unique advantage of all-day and all-weather imaging.

[0003] SAR image classification models are usually used to detect and identify target features and models from SAR images. Currently, SAR image classification models used for SAR image category discovery have the technical problem of poor classification effect. Summary of the invention

[0004] In view of the above problems, the present application provides a training method and related device for a SAR image classification model to achieve the purpose of improving the classification effect of the SAR image classification model. The specific scheme is as follows:

[0005] In a first aspect, the present application provides a training method for a SAR image classification model, comprising multiple iterations, wherein the iterations include:

[0006] Acquire the iterative training data, wherein the training data includes a sample SAR image and a true value label, wherein the true value label identifies an image category;

[0007] For any one of the sample SAR images, obtaining a first enhanced image and a second enhanced image of a sample characteristic image of the sample SAR image;

[0008] Inputting the first enhanced image of the sample SAR image into the student model and the teacher model, obtaining the student projection head output by the student model and the teacher projection head output by the teacher model, wherein the student model is a SAR image classification model to be trained, the SAR image classification model is composed of an encoder and a first projector, and the teacher model is a mirror model of the SAR image classification model;

[0009] Performing weighted fusion and centering processing on the teacher projection heads based on a weight factor to obtain a distilled balanced projection head, wherein the weight factor is inversely correlated with the number of images belonging to the image category of the sample SAR image in the training data;

[0010] Input the first enhanced image and the second enhanced image into a perception projection module, obtain a first projection head corresponding to the first enhanced image output by the perception projection module and a second projection head corresponding to the second enhanced image, wherein the perception projection module shares an encoder with the SAR image classification model;

[0011] Calculate the total loss of the iteration based on the true value labels of all the sample SAR images in the training data of the iteration, the distillation balanced projection head, the student projection head, the first projection head, and the second projection head;

[0012] If the preset training completion conditions are met, the training ends;

[0013] If the training completion condition is not met, the model parameters are updated and the next iteration is performed.

[0014] In one possible implementation, the encoder is a frequency attention module, and the first projector is a linear layer;

[0015] The perception projection module is composed of the encoder and a second projector, and the second projector is a multi-layer perceptron.

[0016] In a possible implementation, obtaining a first enhanced image and a second enhanced image of a sample characteristic image of the sample SAR image includes:

[0017] The sample SAR image is linearly projected and position-encoded to obtain the sample feature image;

[0018] The sample feature image is resized back to its original size by random cropping, and the sample feature image is transformed by random color distortion and random Gaussian blur to generate a first enhanced image and a second enhanced image of the sample feature image.

[0019] In a possible implementation, calculating the total loss of the iteration based on the true value labels of all the sample SAR images in the training data of the iteration, the distillation balance projection head, the student projection head, the first projection head, and the second projection head includes:

[0020] The distilled balanced projection head and the student projection head are normalized based on a preset normalized softmax function, and the normalized result corresponding to the distilled balanced projection head is used as a soft pseudo label, and the normalized result corresponding to the student projection head is used as a predicted value;

[0021] Using the first projection head and the second projection head of the sample SAR image as the first contrast value and the second contrast value respectively;

[0022] Acquire a first projection head corresponding to a first enhanced image of other sample SAR images as a comparison value, wherein the other sample images are images other than the sample SAR image in the iterative training data;

[0023] Calculating classification learning loss based on the soft pseudo label, the predicted value and the true value label;

[0024] Calculating a representation of learning loss based on the control comparison value, the first comparison value, and the second comparison value;

[0025] The sum of the classification learning loss and the representation learning loss is calculated as the total loss of the iteration.

[0026] In a possible implementation, calculating the classification learning loss based on the soft pseudo label, the predicted value, and the true value label includes:

[0027] Based on the soft false label and the predicted value, using a first unsupervised contrast loss function to perform unsupervised classification learning on the iterative training data to obtain a first unsupervised loss value, wherein the first unsupervised contrast loss function is equal to a cross entropy function minus an average entropy maximization function;

[0028] Based on the predicted value and the true value label, using a first supervised contrast loss function to perform supervised classification learning on the iterative training data to obtain a first supervised loss value, wherein the first supervised contrast loss function is a cross entropy loss function;

[0029] Based on a preset first joint training weighting coefficient, the first unsupervised loss value and the first supervised loss value are weightedly added to obtain the classification learning loss.

[0030] In a possible implementation, calculating a representation learning loss based on the control comparison value, the first comparison value, and the second comparison value includes:

[0031] Obtaining a positive contrast value from the control contrast value, wherein the positive contrast value is a first projection head corresponding to a first enhanced image of a sample image of the same type, and the sample image of the same type is an image of the same image category as the sample SAR image in the iterative training data;

[0032] Based on the positive contrast value, the first contrast value and the second contrast value, using a second unsupervised contrast loss function to perform unsupervised contrast learning on the iterative training data to obtain a second unsupervised loss value;

[0033] Based on the positive contrast value, the control contrast value, the first contrast value and the second contrast value, using a second supervised contrast loss function to perform supervised contrast learning on the iterative training data to obtain a second supervised loss value;

[0034] Based on a preset second joint training weighting coefficient, the second unsupervised loss value and the second supervised loss value are weightedly added to obtain the representation learning loss.

[0035] A second aspect of the present application provides a training device for a SAR image classification model, including a joint training module, wherein the joint training module includes:

[0036] A training data acquisition unit, configured to acquire the iterative training data, wherein the training data includes a sample SAR image and a true value label, wherein the true value label identifies an image category;

[0037] An image enhancement unit, configured to obtain, for any one of the sample SAR images, a first enhanced image and a second enhanced image of a sample characteristic image of the sample SAR image;

[0038] a first projection unit, configured to input the first enhanced image of the sample SAR image into a student model and a teacher model, and obtain a student projection head output by the student model and a teacher projection head output by the teacher model, wherein the student model is a SAR image classification model to be trained, the SAR image classification model is composed of an encoder and a first projector, and the teacher model is a mirror model of the SAR image classification model;

[0039] A distillation balancing unit, configured to perform weighted fusion and centering processing on the teacher projection heads based on a weight factor to obtain a distillation balanced projection head, wherein the weight factor is inversely correlated with the number of images belonging to the image category of the sample SAR image in the training data;

[0040] a second projection unit, configured to input the first enhanced image and the second enhanced image into a perception projection module, and obtain a first projection head corresponding to the first enhanced image output by the perception projection module and a second projection head corresponding to the second enhanced image, wherein the perception projection module shares an encoder with the SAR image classification model;

[0041] A total loss calculation unit, configured to calculate the total loss of the iteration based on the true value labels of all the sample SAR images in the training data of the iteration, the distillation balance projection head, the student projection head, the first projection head, and the second projection head;

[0042] The training condition determination unit is used to terminate the training if the preset training completion condition is met; if the training completion condition is not met, the model parameters are updated and the next iteration is performed.

[0043] A third aspect of the present application provides a computer program product, comprising computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements the training method of the SAR image classification model of the above-mentioned first aspect or any implementation manner of the first aspect.

[0044] A fourth aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:

[0045] The memory is used to store computer programs;

[0046] The processor is used to execute the computer program so that the electronic device can implement the training method of the SAR image classification model of the above-mentioned first aspect or any implementation manner of the first aspect.

[0047] A fifth aspect of the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the motion recognition method of a web page object according to the first aspect or any implementation of the first aspect.

[0048] By means of the above technical scheme, the present application provides a training method and related device for a SAR image classification model, and obtains iterative training data, wherein the training data includes a sample SAR image and a true value label, and the true value label identifies the image category. For any sample SAR image, obtain the first enhanced image and the second enhanced image of the sample feature image of the sample SAR image. Input the first enhanced image of the sample SAR image into the student model and the teacher model, obtain the student projection head output by the student model and the teacher projection head output by the teacher model, the student model is the SAR image classification model to be trained, the SAR image classification model is composed of an encoder and a first projector, and the teacher model is a mirror model of the SAR image classification model. Based on the weight factor, the teacher projection head is weightedly fused and centralized to obtain a distillation balanced projection head, and the weight factor is inversely correlated with the number of images belonging to the image category of the sample SAR image in the training data. Input the first enhanced image and the second enhanced image into the perception projection module, obtain the first projection head corresponding to the first enhanced image output by the perception projection module, and obtain the second projection head corresponding to the second enhanced image, and the perception projection module and the SAR image classification model share an encoder. The total loss of the iteration is calculated based on the true value labels of all sample SAR images in the iterative training data, the distillation balanced projection head, the student projection head, the first projection head, and the second projection head. If the preset training completion condition is met, the training is terminated. If the training completion condition is not met, the model parameters are updated and the next iteration is performed. This scheme performs weighted fusion and centralization processing on the teacher projection head based on the weight factor to obtain the distillation balanced projection head, and calculates the total loss of the iteration based on the true value label, the distillation balanced projection head, the student projection head, the first projection head, and the second projection head. Since the weight factor is inversely correlated with the number of images belonging to the image category of the sample SAR image in the training data, the long tail problem of the training data can be improved. Further, the total loss is obtained by joint supervised and unsupervised training based on the true value labels of all sample SAR images in the iterative training data, the distillation balanced projection head, the student projection head, the first projection head, and the second projection head, thereby improving the training effect, and then improving the classification effect of the SAR image classification model. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.

[0050] Figure 1 A schematic diagram of a system architecture provided for this application;

[0051] Figure 2 An optional hardware structure diagram of the terminal 100 is shown;

[0052] Figure 3 A schematic diagram of the structure of a server 200 is shown;

[0053] Figure 4 A flowchart of a training method for a SAR image classification model provided in an embodiment of the present application;

[0054] Figure 5 A specific implementation flow chart of a SAR image classification method provided in an embodiment of the present application;

[0055] Figure 6a A schematic diagram of the specific structure of a SAR image classification model provided in an embodiment of the present application;

[0056] Figure 6b A schematic diagram of a specific structure of a frequency module provided in an embodiment of the present application;

[0057] Figure 6c A schematic diagram of a specific structure of an attention module provided in an embodiment of the present application;

[0058] Figure 7 A schematic diagram of the specific structure of a BKD-CL model provided in an embodiment of the present application;

[0059] Figure 8a A basic framework of SimCLR is demonstrated;

[0060] Figure 8b A schematic diagram of the specific structure of a multi-layer perceptron provided in an embodiment of the present application;

[0061] Fig. 9 A flowchart of a training method for a SAR image classification model provided in an embodiment of the present application;

[0062] Fig.10 A flowchart of a total loss calculation method provided in an embodiment of the present application;

[0063] Fig.11 A schematic diagram of the structure of a training device for a SAR image classification model provided in an embodiment of the present application;

[0064] Fig.12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0065] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation method section of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.

[0066] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0067] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, which is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0068] The present application can be applied in the field of artificial intelligence technology, and specifically in the field of synthetic aperture radar automatic recognition, for image classification tasks of detecting and identifying target features and models from SAR images.

[0069] At present, the technical difficulties are: (1) Traditional SAR ATR relies on a large amount of annotated training data, but in actual scenarios, data annotation is difficult. When the movable military targets in the SAR images undergo technical upgrades, such as adding turrets to tanks or replacing aircraft chassis and wing structures, these modifications often lead to the emergence of new target categories. (2) The distribution of real-world data samples is unbalanced and often presents a long-tail distribution, that is, a few categories occupy a large number of samples. This unbalanced data distribution can easily cause the model to overfit to the head category and ignore the tail category, thereby reducing the generalization ability of the model. Therefore, in order to promote the intelligent implementation of SAR systems, it is necessary to solve the problems of new category recognition and unknown data distribution.

[0070] In order to solve the above technical problems, the present invention proposes a long-tail SAR category discovery method based on the BKD-CL model, which performs supervised and unsupervised joint training on the SAR image classification model to overcome the problem of poor generalization ability of the SAR image classification model caused by unbalanced sample distribution. It can realize category discovery in an open environment with unknown data distribution, and improve the intelligence level of the system through autonomous learning of SAR and recognition of new targets.

[0071] This application can be applied to, but is not limited to, applications with model training functions or cloud services provided by cloud-side servers, etc. The following will introduce them separately:

[0072] See also Figure 1 , Figure 1A schematic diagram of a system architecture is shown. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1 In the example, a server is included, and the server 200 can provide the method provided in the embodiment of the present application for one or more terminals.

[0073] Among them, a training application for a SAR image classification model can be installed on the terminal 100. The above application and web page can provide an interface. The terminal 100 can receive relevant parameters input by the user on the training interface of the SAR image classification model, and send the above parameters to the server 200. The server 200 can obtain processing results based on the received parameters and return the processing results to the terminal 100.

[0074] It should be understood that in some optional implementations, the terminal 100 can also complete the action of obtaining the processing result based on the received parameters by itself without the cooperation of the server, and the embodiments of the present application are not limited to this.

[0075] Next describe Figure 1 The product form of the mid-terminal 100;

[0076] The terminal 100 in the embodiment of the present application can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiment of the present application does not impose any limitation on this.

[0077] Figure 2 An optional hardware structure diagram of the terminal 100 is shown.

[0078] refer to Figure 2 As shown, the terminal 100 may include a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160 (optional), a speaker 161 (optional), a microphone 162 (optional), an earphone jack 163 (optional), a processor 170, an external interface 180, a power supply 190 and other components. Those skilled in the art will appreciate that Figure 2 These are merely examples of terminals or multi-function devices and do not constitute limitations on the terminals or multi-function devices, which may include more or fewer components than those shown in the figures, or combinations of certain components, or different components.

[0079] The input unit 130 can be used to receive input digital or character information, and generate key signal input related to the user settings and function control of the portable multifunctional device. Specifically, the input unit 130 may include a touch screen 131 (optional) and / or other input devices 132. The touch screen 131 can collect the user's touch operations on or near it (such as the user's operation on or near the touch screen using any suitable object such as fingers, joints, stylus, etc.), and drive the corresponding connection device according to a pre-set program. The touch screen can detect the user's touch action on the touch screen, convert the touch action into a touch signal and send it to the processor 170, and can receive and execute the command sent by the processor 170; the touch signal at least includes the touch point coordinate information. The touch screen 131 can provide an input interface and an output interface between the terminal 100 and the user. In addition, the touch screen can be implemented using multiple types such as resistive, capacitive, infrared and surface acoustic wave. In addition to the touch screen 131, the input unit 130 can also include other input devices. Specifically, other input devices 132 may include, but are not limited to, one or more of a physical keyboard, function keys (such as a volume control key, a switch key, etc.), a trackball, a mouse, a joystick, and the like.

[0080] Among them, the input device 132 can receive input data and the like.

[0081] The display unit 140 may be used to display information input by the user or provided to the user, various menus of the terminal 100, interactive interfaces, file display, and / or playback of any multimedia file. In the embodiment of the present application, the display unit 140 may be used to display the training interface, processing results, etc. of the SAR image classification model.

[0082] The memory 120 can be used to store instructions and data. The memory 120 can mainly include an instruction storage area and a data storage area. The data storage area can store various data, such as multimedia files, texts, etc.; the instruction storage area can store software units such as operating systems, applications, instructions required for at least one function, or their subsets and extensions. It can also include a non-volatile random access memory; provide the processor 170 with hardware, software and data resources including management of computing and processing equipment, and support control software and applications. It is also used for the storage of multimedia files, and the storage of running programs and applications.

[0083] The processor 170 is the control center of the terminal 100. It uses various interfaces and lines to connect various parts of the entire terminal 100. By running or executing instructions stored in the memory 120 and calling data stored in the memory 120, it executes various functions of the terminal 100 and processes data, thereby controlling the terminal device as a whole. Optionally, the processor 170 may include one or more processing units; preferably, the processor 170 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface and application program, and the modem processor mainly processes wireless communication. It is understandable that the above-mentioned modem processor may not be integrated into the processor 170. In some embodiments, the processor and the memory may be implemented on a single chip, and in some embodiments, they may also be implemented separately on separate chips. The processor 170 may also be used to generate corresponding operation control signals, send them to corresponding components of the computing and processing device, read and process data in the software, especially read and process data and programs in the memory 120, so that each functional module therein performs corresponding functions, thereby controlling the corresponding components to act according to the requirements of the instructions.

[0084] Among them, the memory 120 can be used to store software codes related to the training method of the SAR image classification model, the processor 170 can execute the steps of the training method of the SAR image classification model, and can also schedule other units (such as the above-mentioned input unit 130 and the display unit 140) to implement corresponding functions.

[0085] The radio frequency unit 110 (optional) can be used for receiving and sending information or receiving and sending signals during a call, for example, after receiving the downlink information of the base station, it is sent to the processor 170 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (Low Noise Amplifier, LNA), a duplexer, etc. In addition, the radio frequency unit 110 can also communicate with network devices and other devices through wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (Global System of Mobile communication, GSM), General Packet Radio Service (General Packet Radio Service, GPRS), Code Division Multiple Access (Code Division Multiple Access, CDMA), Wideband Code Division Multiple Access (Wideband Code Division Multiple Access, WCDMA), Long Term Evolution (Long Term Evolution, LTE), email, Short Messaging Service (SMS), etc.

[0086] In this embodiment of the present application, the RF unit 110 can send data to the server 200 and receive processing results sent by the server 200.

[0087] It should be understood that the radio frequency unit 110 is optional and can be replaced by other communication interfaces, such as a network port.

[0088] The terminal 100 also includes a power supply 190 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 170 through a power management system, so that the power management system can manage functions such as charging, discharging, and power consumption.

[0089] The terminal 100 further includes an external interface 180 , which may be a standard Micro USB interface or a multi-pin connector, and may be used to connect the terminal 100 to communicate with other devices, or to connect a charger to charge the terminal 100 .

[0090] Although not shown, the terminal 100 may also include a flashlight, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which are not described in detail here. Some or all of the methods described below may be applied in the following embodiments. Figure 2 In the terminal 100 shown.

[0091] Next describe Figure 1 The product form of the server 200;

[0092] Figure 3 A schematic diagram of the structure of a server 200 is shown. Figure 3 As shown, the server 200 includes a bus 201, a processor 202, a communication interface 203, and a memory 204. The processor 202, the memory 204, and the communication interface 203 communicate with each other via the bus 201.

[0093] The bus 201 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0094] The processor 202 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0095] The memory 204 may include a volatile memory, such as a random access memory (RAM). The memory 204 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard drive (HDD), or a solid state drive (SSD).

[0096] The memory 204 may be used to store software codes related to the training method of the SAR image classification model, and the processor 202 may execute the steps of the training method of the SAR image classification model of the chip, and may also schedule other units to implement corresponding functions.

[0097] It should be understood that the above-mentioned terminal 100 and server 200 can be centralized or distributed devices, and the processors in the above-mentioned terminal 100 and server 200 (such as processor 170 and processor 202) can be hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general-purpose processor, digital signal processor (DSP), microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the processor can be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.

[0098] The present application embodiment provides a training method for a SAR image classification model. The training method for a SAR image classification model of the present application embodiment is described in detail below with reference to the accompanying drawings.

[0099] Reference Figure 4 , Figure 4 A flowchart of a training method for a SAR image classification model provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, a training method for a SAR image classification model provided in an embodiment of the present application includes multiple iterations, and one iteration may include steps S401 to S404. These steps are described in detail below.

[0100] S401: Obtain iterative training data.

[0101] In this embodiment, the training data includes sample SAR images and true value labels, and the true value labels identify image categories.

[0102] S402: For any sample SAR image, obtain a first enhanced image and a second enhanced image of a sample characteristic image of the sample SAR image.

[0103] S403: Input the first enhanced image of the sample SAR image into the student model and the teacher model, and obtain the student projection head output by the student model and the teacher projection head output by the teacher model.

[0104] In this embodiment, the student model is a SAR image classification model to be trained, the SAR image classification model is composed of an encoder and a first projector, and the teacher model is a mirror model of the SAR image classification model.

[0105] In an optional embodiment, the encoder is a frequency attention module and the first projector is a linear layer.

[0106] S404: Perform weighted fusion and centralization processing on the teacher projection heads based on the weight factors to obtain a distilled balanced projection head.

[0107] In this embodiment, the weight factor is inversely correlated with the number of images belonging to the image category of the sample SAR image in the training data.

[0108] S405: Input the first enhanced image and the second enhanced image to the perception projection module, and obtain a first projection head corresponding to the first enhanced image and a second projection head corresponding to the second enhanced image output by the perception projection module.

[0109] In this embodiment, the perception projection module is composed of an encoder and a second projector, the second projector is a multi-layer perceptron, and the perception projection module and the SAR image classification model share an encoder, that is, the perception projection module is a FAViT module.

[0110] S406 , calculating the total loss of the iteration based on the true value labels of all sample SAR images in the iterative training data, the distillation balanced projection head, the student projection head, the first projection head, and the second projection head.

[0111] S407: If the preset training completion condition is met, the training ends.

[0112] S408: If the training completion condition is not met, the model parameters are updated and the next iteration is performed.

[0113] It can be seen from the above technical solution that a training method for a SAR image classification model provided in an embodiment of the present application performs weighted fusion and centralization processing on the teacher projection head based on a weight factor to obtain a distilled balanced projection head, and calculates the iterative total loss based on the true value label, the distilled balanced projection head, the student projection head, the first projection head, and the second projection head. Since the weight factor is inversely correlated with the number of images belonging to the image category of the sample SAR image in the training data, the long tail problem of the training data can be improved. Furthermore, supervised and unsupervised joint training is performed based on the true value labels of all sample SAR images in the iterative training data, the distilled balanced projection head, the student projection head, the first projection head, and the second projection head to obtain the total loss, thereby improving the training effect, and then improving the classification effect of the SAR image classification model.

[0114] See also Figure 5 , Figure 5 A specific implementation flow chart of a SAR image classification method provided in an embodiment of the present application is as follows: Figure 5 As shown, this method specifically includes S501 to S507, as follows:

[0115] S501, constructing a SAR image classification model.

[0116] In this embodiment, the SAR image classification model is composed of a frequency attention module and a linear layer. The SAR image classification model is used to input a SAR image to be identified, extract the image features of the SAR image to be identified through N backbone networks, and output the category recognition result based on the image features of the SAR image to be identified through a linear layer, that is, a fully connected layer. The linear layer is a normalized linear layer, and the number of neurons in the output layer is equal to the number of labeled categories. The frequency attention module includes N backbone networks, and each backbone network includes multiple frequency modules and multiple attention modules.

[0117] Figure 6a A schematic diagram of the specific structure of a SAR image classification model provided in an embodiment of the present application.

[0118] like Figure 6a As shown, the frequency attention module includes N backbone networks, each backbone network includes 2 frequency modules and 4 attention modules.

[0119] Figure 6b A specific structural diagram of a frequency module provided in an embodiment of the present application, Figure 6b As shown, the frequency module includes a normalization Norm layer, a spectrum gating network, and a multi-layer perceptron MLP, wherein the spectrum gating network includes a fast Fourier transform layer FFT, a weighted gating layer, and an inverse Fourier layer IFFT. The frequency module introduces a global filter using a spectrum gating network to capture different frequency components of the SAR image to understand the local frequency.

[0120] First, for the input The frequency module obtains the 2D fast Fourier transform through FFT Spectrum , thereby converting the physical space into the spectrum space, as shown in formula (1):

[0121] (1);

[0122] in, represents the 2D fast Fourier transform, is a complex tensor representing spectrum.

[0123] Furthermore, the frequency module passes through a weighted gating layer with learnable weight parameters Determine the weight of each frequency component , in order to properly capture the lines and edges of the SAR image and obtain the modulated spectrum , the formula is as follows (2):

[0124] (2);

[0125] in, Represents element-by-element multiplication, global filter With spectrum have the same dimensions.

[0126] Furthermore, the frequency module converts the modulated spectrum into Transform back to physical space to achieve The update formula is as follows (3):

[0127] (3);

[0128] in, Represents the inverse of the 2D Fast Fourier Transform.

[0129] Furthermore, the frequency module performs nonlinear transformation and processing on the output data of the spectrum gating network through MLP to enhance the expression ability of the frequency module.

[0130] Figure 6c A schematic diagram of a specific structure of an attention module provided in an embodiment of the present application, Figure 6c As shown in the figure, the attention module contains a normalization layer, a multi-head self-attention layer, and a multi-layer perceptron MLP.

[0131] First, the attention module uses a fully connected layer to generate the initial values ​​of q, k, and v through a multi-head self-attention layer, then uses reconstruction and dimension swapping to adjust them, and finally uses slicing operations to obtain separate Q, K, and V. The attention module introduces a multi-head self-attention mechanism to enhance the stability and robustness of the network.

[0132] Furthermore, the attention module performs nonlinear transformation and processing on the output data of the multi-head self-attention layer through MLP to enhance the expression ability of the attention module.

[0133] S502: Construct a joint training model for training a SAR image classification model.

[0134] In this embodiment, the joint training model is called the BKD-CL model. Figure 7 A structural diagram of a joint training model provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, the joint training model includes a preprocessing module, a data enhancement module, a SAR image classification model to be trained as a student model, a teacher model of the SAR image classification model, a multi-layer perceptron, and a loss function module.

[0135] The preprocessing module includes a linear projection layer and a position encoding layer connected in sequence, and the input data of the data preprocessing module is training data, and the training data includes labeled data and unlabeled data. The preprocessing module is used to output a sample feature image of a sample SAR image.

[0136] The SAR image classification model to be trained includes a frequency attention module and a linear layer, which are used to output the student classification results. The structures of the teacher model and the student model are completely consistent, and can be different versions of the SAR image classification model to be trained, which are used to output the teacher classification results. For the specific structure, please refer to the above embodiment.

[0137] In this embodiment, the FAViT module includes a frequency attention module in a SAR image classification model and a multi-layer perceptron connected to the output end of the frequency attention module. The FAViT module is constructed based on SimCLR (Simple Framework for Contrastive Learning of Representations, self-supervised learning based on contrastive learning). The frequency attention module of the BKD-CL model constitutes an encoder in the basic framework of SimCLR, and the multi-layer perceptron constitutes a projection head in the basic framework of SimCLR. FIG8 illustrates a basic framework of SimCLR, and the basic framework of SimCLR includes:

[0138] The data augmentation module is used to resize the sample SAR image x back to its original size by random cropping, random color distortion, and random Gaussian blur transformation of the input sample SAR image to generate two enhanced images of the same sample SAR image x and , and Form a pair of positive samples.

[0139] Encoder , used to extract the representation vector from each enhanced image, for example, using a residual network to extract the representation vector to obtain the representation vector and ,in, and They are all representation vectors output after the average pooling layer.

[0140] Projector , which is used to map the representation vector to the latent space of contrastive learning to obtain the projection result. For example, an MLP (Multilayer Perceptron) with a hidden layer is used to extract and In practical applications, the specific structures of the encoder and the projector (also called projection head) can be selected according to requirements.

[0141] The contrast loss function is used to extract 2N enhanced images generated by a training subset including N sample SAR images (that is, the training data set and used in one iteration), where each sample SAR image generates 2 enhanced images. For a sample SAR image x in the training subset, when a positive sample pair of the sample SAR image x is determined ( , ), then enhance the image And the other 2(N-1) enhanced images form a negative sample pair, denoted as ( , ), For 2N enhanced images, except for the positive sample pairs ( , ), The number of is 2(N-1).

[0142] The contrast loss function is used to perform contrastive learning based on the projection results of the positive sample pair and the negative sample pair, so that the projection results of the positive sample pair can be as consistent as possible. The contrast loss function Defined as formula (4):

[0143] (4);

[0144] in, is an indicator function that evaluates to 1, Represents the temperature parameter.

[0145] In this embodiment, the category discovery task is a joint task of comprehensive supervised and unsupervised classification. The contrastive learning in the SimCLR model can learn the image features of unlabeled sample SAR images by training the model which data points are similar or different.

[0146] In this embodiment, the frequency attention module outputs a first characterization vector and a second characterization vector respectively based on the input first enhanced image and the second enhanced image, and the multilayer perceptron is used to output a first projection head and a second projection head respectively based on the first characterization vector and the second characterization vector.

[0147] Figure 8b A schematic diagram of a specific structure of a multi-layer perceptron provided in an embodiment of the present application is shown in FIG. Figure 8b As shown in the figure, the multilayer perceptron includes 2048 hidden layers and 256 output layers. Therefore, after passing through the multilayer perceptron, the tensor changes from (64, 384) to a projection feature of (64, 256).

[0148] The loss function module includes a distillation balance module, a Softmax module, and a total loss calculation module, wherein the distillation balance module is connected to the output end of the linear layer of the teacher model, the Softmax module is respectively connected to the output end of the distillation balance module and the output end of the SAR image classification model, and the total loss calculation module is respectively connected to the output end of the multilayer perceptron and the output end of the Softmax module. The distillation balance module includes a fusion module and a centralization module connected in sequence, which are used to receive the teacher projection head output by the teacher model and output the fused teacher projection head. The Softmax module is used to use the normalized softmax function to output a soft pseudo label based on the received teacher projection head and output a predicted value based on the received student projection head output by the student model.

[0149] The total loss calculation module calculates the total loss based on the received true value labels, soft false labels, predicted values, and contrast values, and updates the model parameters based on the total loss, wherein the contrast values ​​include a first projection head as a first contrast value, a second projection head as a second contrast value, and a control contrast value, wherein the control contrast value is determined based on the first projection head of multiple negative samples of the sample SAR image.

[0150] S503: Acquire sample SAR image data, and divide the sample SAR image data into labeled data and unlabeled data.

[0151] In this embodiment, the sample SAR image data includes sample SAR images of various categories, and the sample SAR images are acquired by a high-resolution spotlight synthetic aperture radar sensor under various scalable conditions such as target occlusion, camouflage, configuration change, etc. Optionally, the radar operates in the X-band and uses HH polarization mode.

[0152] Specifically, the number of categories of sample SAR images in the sample SAR image data is n, n1 categories are randomly selected as known categories, and the remaining n2 categories are selected as unknown categories, where n= n1+ n2.

[0153] Select a preset first ratio of sample SAR images from each known class of sample SAR images as the first class of images, use the first class of images as labeled images, and construct labeled data .in represents the u-th labeled image, Represents the true value label of the u-th labeled image. The true value label of the labeled image is the category, u∈[1, ], Indicates the number of labeled images.

[0154] The sample SAR images of known classes except the first class images are taken as the second class images, and the second class images and the sample SAR images of unknown classes are taken as unlabeled images to construct unlabeled data. ,in, represents the vth unlabeled image, Represents the label of the vth unlabeled image. The label of an unlabeled image can be a null value, v∈[1, ], Represents the number of unlabeled images.

[0155] Taking the sample SAR image data including sample SAR images of 10 categories as an example, 5 categories are selected as known categories and 5 categories are selected as unknown categories. 50% of the sample SAR images in the known categories are selected as labeled images to construct labeled data. The remaining 50% of the sample SAR images of the known class and the sample SAR images of the unknown class are taken as unlabeled images, which together constitute the unlabeled data .

[0156] S504 , using the training data, and performing supervised and unsupervised joint training on the SAR image classification model based on the BKD-CL model until a preset training completion condition is met, thereby obtaining a trained SAR image classification model.

[0157] In this embodiment, the supervised and unsupervised joint training of the SAR image classification model based on the BKD-CL model includes multiple iterative trainings. The execution method of each iterative training is as described in the above embodiment.

[0158] In this embodiment, the training data includes labeled data and unlabeled data That is, the training data includes + The training data is input into the BKD-CL model, the maximum number of iterative training cycles is set to 300, and the batch size of the training data for each iteration is 32, that is, 32 sample SAR images are input each time to fit the BKD-CL model. It should be noted that the size of the training data for a single iteration is (32, 3, 224, 224), where 32 is the batch size, 3 is the number of channels, and 224*224 is the width and height of the sample SAR image.

[0159] In this embodiment, the initial learning rate is 0.001, and the SGD (Stochastic Gradient Descent) optimization method is used to update all parameters of the SAR image classification model. The training completion conditions include that the error is less than the preset error threshold or the maximum number of iterations is met, that is, the iteration is stopped after the error is less than the preset error threshold or the maximum number of iterations is met, and a trained SAR image classification model is obtained.

[0160] Specifically, the model training platform is 2 NVIDIA GeForce RTX 2080Ti GPUs, and Table 1 lists the training parameter configurations.

[0161] Table 1

[0162]

[0163] S505: input the verification data into the trained SAR image classification model, and output the predicted categories of the sample SAR images belonging to the known class and the predicted categories of the sample SAR images belonging to the unknown class in the verification data through the SAR image classification model.

[0164] In this embodiment, the unlabeled data The input is fed into the trained SAR image classification model, and the predicted category corresponding to each sample SAR image in the unlabeled data is output.

[0165] In this embodiment, the accuracy rate is a measure of the overall correctness of the model. By counting the number of correctly predicted categories in the SAR images, the ratio of correctly identified sample SAR images to the verification data is obtained.

[0166] In this embodiment, the evaluation index includes the first classification accuracy, the second classification accuracy, and the third classification accuracy. The first classification accuracy represents the ratio of correctly identified sample SAR images to the verification data, the second classification accuracy represents the ratio of correctly identified sample SAR images of known classes to all sample SAR images of known classes, and the third classification accuracy represents the ratio of correctly identified sample SAR images of unknown classes to all sample SAR images of unknown classes.

[0167] Specifically, the model accuracy is calculated based on various evaluation indicators, as shown in formula (5):

[0168] (5);

[0169] Among them, TP represents the number of samples with the same predicted category and label and correct, and N represents the total number of samples.

[0170] Furthermore, based on the category discovery problem of unbalanced data distribution, this solution introduces an imbalance factor To change the number of sample SAR images of known classes in the unlabeled data, that is, the number of images of the second category and the number of sample SAR images of unknown classes , so as to complete the test under different long tail ratios, that is, .

[0171] Table 2 lists the classification accuracy performance of the BKD-CL model provided by this solution and the GCD (Generalized Category Discovery) model, SimGCD (SimpleGeneralized Category Discovery) model and ORCA (Open-world with uncertain based adaptive margin) model in the prior art at different long-tail ratios. From Table 2, it can be seen that the BKD-CL model has higher accuracy in the classification of known classes, unknown classes, and all classes at different long-tail ratios. As the long-tail ratio increases, the third classification accuracy of the other three models decreases significantly. In contrast, the balanced knowledge distillation of the BKD-CL model can well solve the problem of low third classification accuracy.

[0172] Table 2

[0173]

[0174] S506: If the preset verification condition is met, output the trained SAR image classification model.

[0175] In this embodiment, the verification conditions include that the first classification accuracy, the second classification accuracy, and the third classification accuracy are not less than the corresponding indicator thresholds, and the model accuracy is not less than the preset accuracy threshold.

[0176] S507 , obtaining a SAR image to be identified, inputting the SAR image to be identified into a trained SAR image classification model, and obtaining a category recognition result of the SAR image to be identified output by the SAR image classification model.

[0177] The embodiment of the present application implements a training method for a SAR image classification model through a joint training model. Fig. 9 A flowchart of a training method for a SAR image classification model provided in an embodiment of the present application is shown in FIG. Fig. 9 As shown, the training method includes multiple iterations. Optionally, the batch size of the training data of a single iteration is a preset batch size. The mth iteration includes:

[0178] S901. A preprocessing module is used to linearly project and position-code the sample SAR image in the training data of the mth iteration to obtain a sample feature image.

[0179] In this embodiment, taking the i-th sample SAR image as an example, the sample SAR image in the training data is input into the preprocessing module, and the tensor of the image block obtained after the sample SAR image is patched by the linear projection layer of the preprocessing module is (64, 196, 384). The position code is embedded into the image block and normalized by the position coding layer, and the tensor of the normalized image block obtained is (64, 384), wherein the position code expresses the position information of the image block in the sample SAR image. The sample feature image xi output by the position coding module is composed of normalized image blocks with a tensor of (64, 384).

[0180] S902: Input the sample feature image into a data enhancement module, and generate a first enhanced image and a second enhanced image of the sample feature image through the data enhancement module.

[0181] In this embodiment, the data enhancement module resizes the sample feature image back to its original size by random cropping, performs random color distortion, and performs random Gaussian blur conversion on the input sample feature image to generate two enhancements of the same sample feature image, namely a first enhanced image and a second enhanced image.

[0182] like Figure 7 As shown, the sample feature image Input to the data enhancement module, the data enhancement module generates the first enhanced image of the sample feature image and the second enhanced image .

[0183] S903, input the first enhanced image to the student model, output the representation vector of the first enhanced image to the linear layer through the frequency attention module, and output the student projection head through the linear layer.

[0184] like Figure 7 As shown, the first enhanced image Input to the student model (i.e. the classification model of the SAR image to be trained), and the frequency attention module Encode and output the first representation vector ,Will Input to the linear layer, through the linear layer Projection, output student projection head .

[0185] S904, input the first enhanced image into the teacher model, output the representation vector of the first enhanced image to the linear layer through the frequency attention module, and output the teacher projection head through the linear layer.

[0186] In this embodiment, the teacher model is a mirror image of the classification model of the SAR image. Specifically, the teacher model and the student model may be classification models of SAR images of different versions.

[0187] like Figure 7 As shown, Input to the teacher model, the teacher model performs the same operation as the student model and outputs the teacher projection head .

[0188] In another optional embodiment, the first projection head output by the student model is directly copied to obtain the teacher projection head.

[0189] S905. Input the teacher projection head into the distillation balance module in the loss function module. The distillation balance module performs weighted fusion on the teacher projection head based on the weight factor through the fusion module. The weighted fusion teacher projection head is centralized through the centralization module to output the distillation balanced projection head.

[0190] like Figure 7 As shown, the fusion module in the balanced distillation module is based on the weight factor Projection head for teachers Perform weighted fusion. Specifically, Divide by the temperature parameter of the teacher model Later introduction right Perform weighted fusion and update the teacher output after centering through the centering module to be the distillation balanced projection head .

[0191] In self-supervised learning, when the model maps multiple input data to the same feature representation, only considering part of the data representation and ignoring the features of other data samples will have a significant negative impact on the robustness of the model. Using a centralization module can prevent the model from collapsing and prevent a part of the features from dominating. The specific process is to add a bias term In the teacher model, The updating strategy adopts exponential moving average method.

[0192] S906. Input the distilled balanced projection head and the student projection head into the Softmax module. The Softmax module outputs a soft pseudo label based on the distilled balanced projection head and outputs a predicted value based on the student projection head.

[0193] In this embodiment, the student output is divided by the temperature parameter of the student model Finally, the softmax function is applied to update the student output to obtain the predicted value.

[0194] like Figure 7As shown, through the Softmax module Normalize and get the predicted value ,right Normalize and get soft pseudo labels .

[0195] S907, input the first enhanced image and the second enhanced image into the FAViT module, output the representation vectors of the first enhanced image and the second enhanced image to the multi-layer perceptron through the frequency attention module, and output the first projection head and the second projection head through the multi-layer perceptron.

[0196] In this embodiment, the first enhanced image and the second enhanced image are input into the frequency attention module to obtain the representation vector of the first enhanced image and the representation vector of the second enhanced image output by the frequency attention module as an encoder, that is, the first representation vector and the second representation vector, and the first representation vector and the second representation vector are input into the multilayer perceptron to obtain the projection head of the first representation vector and the projection head of the second representation vector output by the multilayer perceptron as a projection head, that is, the first projection head and the second projection head.

[0197] like Figure 7 As shown, and Input to the frequency attention module, the frequency attention module outputs the first representation vector and the second characterization vector ,Will and Input to the multi-layer perceptron, the multi-layer perceptron outputs the first projection head and the second projection head .

[0198] S908, input the true value label, soft false label, predicted value and comparison value into the total loss calculation module.

[0199] In this embodiment, the contrast value is a first contrast value and a second contrast value, that is, a first projection value and a second projection value, such as Figure 7 As shown, the true value label , soft fake label , predicted value And the first comparison value and the second comparison value Input to the total loss calculation module.

[0200] S909. For the training data of the mth iteration, the total loss is calculated by substituting the true value label, soft false label, predicted value, first contrast value, second contrast value and reference contrast value of each sample SAR image into the total loss function through the total loss calculation module to obtain the total loss.

[0201] In this embodiment, the first projection head and the second projection head of the sample SAR image are respectively used as the first contrast value and the second contrast value. The first projection head corresponding to the first enhanced image of other sample SAR images is obtained as the control contrast value, and the other sample images are images other than the sample SAR image in the iterative training data. Based on the soft pseudo label, the predicted value and the true value label, the classification learning loss is calculated. Based on the control contrast value, the first contrast value and the second contrast value, the representation learning loss is calculated. The sum of the classification learning loss and the representation learning loss is calculated as the total loss of the iteration.

[0202] It should be noted that a specific method for calculating the total loss can be found in Fig.10 .

[0203] S910: Determine whether a preset training completion condition is met.

[0204] In this embodiment, the training completion condition includes that the number of iterations reaches a threshold number, or the total loss is less than a preset loss threshold.

[0205] S911. If yes, end the iteration, and configure the classification model of the SAR image based on the current values ​​of the model parameters to obtain the trained classification model of the SAR image.

[0206] S912: If not, update the model parameters and return to S901 to perform the (m+1)th iteration.

[0207] Fig.10 A schematic diagram of a total loss calculation block diagram provided in an embodiment of the present application, such as Fig.10 As shown, S909 substitutes the predicted value, comparison value, soft-false label and true value label of each sample SAR image in the training data of one iteration through the total loss calculation module to calculate the total loss of the mth iteration.

[0208] In this embodiment, the total loss calculation module performs joint learning of representation learning and classification learning based on the total loss function, wherein representation learning includes supervised and unsupervised learning, and classification learning includes supervised and unsupervised learning. Specifically, the representation learning of the total loss calculation module includes the frequency attention module in the student branch for labeled data. Implement supervised learning and self-supervised learning on training data. The total loss calculation module for classification learning comes from the cross entropy loss between the predicted value and the soft pseudo label, or the cross entropy loss between the predicted value and the true value label.

[0209] In this embodiment, the total loss calculation module substitutes the true value label, soft false label, predicted value, first contrast value, second contrast value and control contrast value of each sample SAR image into the total loss function to calculate the total loss. The specific method includes:

[0210] S1. Based on the soft pseudo labels and the predicted values, unsupervised classification learning is performed on the iterative training data using the first unsupervised contrast loss function to obtain a first unsupervised loss value.

[0211] In this embodiment, the first unsupervised contrast loss function is equal to the cross entropy function minus the average entropy maximization function.

[0212] In this embodiment, self-supervised training can transfer a small amount of labeled sample SAR image data to a large amount of unlabeled sample SAR image data through soft pseudo-labels, thereby improving the quality of features. This scheme uses self-distillation technology to perform self-supervised training classification learning, thereby solving the problem of new class discovery.

[0213] The cross entropy loss function between the soft pseudo labels and the predicted values ​​is used to perform self-supervised classification learning on the training data. The loss function of self-supervised classification learning is equal to the cross entropy minus the average entropy maximization result, as shown in formula (6):

[0214] (6);

[0215] in, Indicates a soft fake label. Represents the predicted value. and are the average prediction and entropy of the batch (iteration of training data), ( ) is the average entropy maximization result, and the calculation methods are as follows:

[0216] (7);

[0217] (8).

[0218] In this embodiment, the quality of soft pseudo labels can be improved by directly connecting the linear layer to the frequency attention module, thereby achieving good new category classification performance. However, in the category discovery problem of long-tail data, the trained model is biased towards the head class. During the distillation process, the prediction information of the tail class may be overwhelmed by the prediction information of the head class. Therefore, the teacher model guided by such a biased model may perform worse.

[0219] Based on this, this scheme introduces a balanced knowledge distillation method through the distillation balance module to transfer knowledge by considering the prior information about the number of categories. Specifically, the weight factor of the sample SAR image is calculated as shown in formula (9):

[0220] (9);

[0221] in, is a hyperparameter, It is the sum of the number of corresponding categories in the entire training data to which the sample SAR image belongs.

[0222] Therefore, the weight factor of the sample SAR image is introduced to update the formula (6), and the first unsupervised contrast loss function is obtained: As shown in formula (10):

[0223] (10).

[0224] S2. Based on the predicted value and the true value label, supervised classification learning is performed on the iterative training data using the first supervised contrast loss function to obtain a first supervised loss value.

[0225] In this embodiment, the first supervised contrast loss function is a cross entropy loss function.

[0226] like Fig.10 As shown, using the predicted value and the true value label The cross entropy loss between them is used to perform supervised classification learning on the training data. The first supervised contrast loss function is as shown in formula (11):

[0227] (11);

[0228] in, Indicates labeled data. is with The corresponding true value label.

[0229] S3. Based on the preset first joint training weighting coefficient, weighted addition is performed on the first unsupervised loss value and the first supervised loss value to obtain the classification learning loss.

[0230] Therefore, by first jointly training the weighted coefficients combination and Get the classification learning loss function , as shown in formula (12):

[0231] (12).

[0232] S4. Based on the control contrast value, the positive contrast value, the first contrast value and the second contrast value, use a second unsupervised contrast loss function to perform unsupervised contrast learning on the iterative training data to obtain a second unsupervised loss value.

[0233] In this embodiment, the training data of this iteration includes N sample images as an example. , compared with the contrast value Additional sample images included The first projection head corresponding to the first enhanced image is facing the contrast value Includes similar sample images The first enhanced image corresponding to the first projection head, the positive contrast value is extracted from the control contrast value.

[0234] In this embodiment, in representation learning, for labeled data, features of sample SAR images of the same category are brought closer in feature space, and features of sample SAR images of different categories are pushed further away.

[0235] Specifically, the second unsupervised contrast loss function As shown in formula (13):

[0236] (13).

[0237] in, is the index corresponding to a batch of images (training data for the mth iteration), for example, =N. Including all the same batch except The index of the set of first enhanced images other than sample images . The first projection head corresponds to the first enhanced image of other sample images.

[0238] In the unsupervised contrastive learning task, the true value label is empty, that is, there is no annotation information. The positive sample pair is usually two augmented images of the same sample SAR image, that is, two enhanced images after image enhancement, and the negative sample comes from other samples in the same batch. Therefore, unsupervised learning solves the classification problem of multiple samples belonging to different categories. However, when there are multiple sample SAR images in the same category, only the augmented positive samples can be classified into the same category, while other samples belonging to the same category are far away from the category. Therefore, supervised learning is introduced.

[0239] S5. Based on the positive contrast value, the first contrast value, and the second contrast value, a second supervised contrast loss function is used to perform supervised contrast learning on the iterative training data to obtain a second supervised loss value.

[0240] In this embodiment, the second supervised contrast loss function As shown in formula (14):

[0241] (14);

[0242] in, For the same batch The indices of all other sample images with the same true value label (referred to as similar sample images), is the positive contrast value, which is the first projection head corresponding to the first enhanced image of the same sample image, is a pre-configured temperature parameter.

[0243] S6. Based on the preset second joint training weighting coefficient, weighted addition is performed on the second unsupervised loss value and the second supervised loss value to obtain a representation learning loss.

[0244] In this embodiment, the second joint training weighting coefficient Combining supervised contrastive loss functions and unsupervised contrast loss function Get the loss function representing learning As shown in formula (15):

[0245] (15).

[0246] In this embodiment, the FAViT module for contrastive learning contains a frequency attention module and a multi-layer perceptron for learnable features. The projection features of the image can be obtained by directly cascading the multi-layer perceptron with the frequency attention module. The traditional self-supervised classification network is based on the projection features, and such a network is more suitable for proxy tasks rather than classification tasks.

[0247] S7. Calculate the sum of the classification learning loss and the representation learning loss as the total loss of the iteration.

[0248] In this embodiment, the total loss function of the total loss calculation module is As shown in formula (16):

[0249] (16).

[0250] In summary, this step calculates the loss value of this iteration by substituting the predicted value, comparison value, soft-false label and true value label of each sample SAR image in the training data of this iteration into the above-mentioned total loss calculation module.

[0251] From the above technical solutions, it can be seen that the introduction of the FAViT module in this solution can enhance data representation by using supervised contrastive learning and unsupervised contrastive learning for labeled data and training data, and cascading a projection head with entropy regularization can improve the recognition of new classes. In view of the long-tail distribution problem in the open world, the introduction of balanced knowledge distillation can improve the recognition accuracy of the tail class.

[0252] A training method for a SAR image classification model provided in an embodiment of the present application is introduced above. A device for executing the training method for the SAR image classification model is introduced below.

[0253] See also Fig.11 , Fig.11 The structural diagram of a training device for a SAR image classification model provided in an embodiment of the present application is shown in FIG. The training device for a SAR image classification model is configured in a client, such as Fig.11 As shown, the training device 1100 of the SAR image classification model includes a joint training module for performing multiple iterations, and the joint training module includes:

[0254] A training data acquisition unit 1101 is used to acquire the iterative training data, wherein the training data includes a sample SAR image and a true value label, and the true value label identifies an image category;

[0255] An image enhancement unit 1102 is configured to obtain, for any one of the sample SAR images, a first enhanced image and a second enhanced image of a sample characteristic image of the sample SAR image;

[0256] A first projection unit 1103 is used to input the first enhanced image of the sample SAR image into a student model and a teacher model, and obtain a student projection head output by the student model and a teacher projection head output by the teacher model, wherein the student model is a SAR image classification model to be trained, the SAR image classification model is composed of an encoder and a first projector, and the teacher model is a mirror model of the SAR image classification model;

[0257] A distillation balancing unit 1104 is used to perform weighted fusion and centralization processing on the teacher projection heads based on a weight factor to obtain a distillation balanced projection head, wherein the weight factor is inversely correlated with the number of images belonging to the image category of the sample SAR image in the training data;

[0258] A second projection unit 1105 is used to input the first enhanced image and the second enhanced image into a perception projection module, obtain a first projection head corresponding to the first enhanced image output by the perception projection module and a second projection head corresponding to the second enhanced image, wherein the perception projection module shares an encoder with the SAR image classification model;

[0259] A total loss calculation unit 1106, configured to calculate the total loss of the iteration based on the true value labels of all the sample SAR images in the iterative training data, the distillation balance projection head, the student projection head, the first projection head, and the second projection head;

[0260] The training condition determination unit 1107 is used to terminate the training if a preset training completion condition is met; if the training completion condition is not met, update the model parameters and execute the next iteration.

[0261] In one possible implementation, the encoder is a frequency attention module, and the first projector is a linear layer;

[0262] The perception projection module is composed of the encoder and a second projector, and the second projector is a multi-layer perceptron.

[0263] In a possible implementation, when the image enhancement unit is used to obtain the first enhanced image and the second enhanced image of the sample characteristic image of the sample SAR image, it is specifically used to:

[0264] The sample SAR image is linearly projected and position-encoded to obtain the sample feature image;

[0265] The sample feature image is resized back to its original size by random cropping, and the sample feature image is transformed by random color distortion and random Gaussian blur to generate a first enhanced image and a second enhanced image of the sample feature image.

[0266] In a possible implementation, the total loss calculation unit is used to calculate the total loss of the iteration based on the true value labels of all the sample SAR images in the iterative training data, the distillation balance projection head, the student projection head, the first projection head, and the second projection head, specifically for:

[0267] The distilled balanced projection head and the student projection head are normalized based on a preset normalized softmax function, and the normalized result corresponding to the distilled balanced projection head is used as a soft pseudo label, and the normalized result corresponding to the student projection head is used as a predicted value;

[0268] Using the first projection head and the second projection head of the sample SAR image as the first contrast value and the second contrast value respectively;

[0269] Acquire a first projection head corresponding to a first enhanced image of other sample SAR images as a comparison value, wherein the other sample images are images other than the sample SAR image in the iterative training data;

[0270] Calculating classification learning loss based on the soft pseudo label, the predicted value and the true value label;

[0271] Calculating a representation of learning loss based on the control comparison value, the first comparison value, and the second comparison value;

[0272] The sum of the classification learning loss and the representation learning loss is calculated as the total loss of the iteration.

[0273] In a possible implementation, the total loss calculation unit is used to calculate the classification learning loss based on the soft pseudo label, the predicted value and the true value label, specifically to:

[0274] Based on the soft false label and the predicted value, using a first unsupervised contrast loss function to perform unsupervised classification learning on the iterative training data to obtain a first unsupervised loss value, wherein the first unsupervised contrast loss function is equal to a cross entropy function minus an average entropy maximization function;

[0275] Based on the predicted value and the true value label, using a first supervised contrast loss function to perform supervised classification learning on the iterative training data to obtain a first supervised loss value, wherein the first supervised contrast loss function is a cross entropy loss function;

[0276] Based on a preset first joint training weighting coefficient, the first unsupervised loss value and the first supervised loss value are weightedly added to obtain the classification learning loss.

[0277] In a possible implementation, the total loss calculation unit is used to calculate the learning loss based on the control comparison value, the first comparison value, and the second comparison value, specifically to:

[0278] Obtaining a positive contrast value from the control contrast value, wherein the positive contrast value is a first projection head corresponding to a first enhanced image of a sample image of the same type, and the sample image of the same type is an image of the same image category as the sample SAR image in the iterative training data;

[0279] Based on the positive contrast value, the first contrast value and the second contrast value, using a second unsupervised contrast loss function to perform unsupervised contrast learning on the iterative training data to obtain a second unsupervised loss value;

[0280] Based on the positive contrast value, the control contrast value, the first contrast value and the second contrast value, using a second supervised contrast loss function to perform supervised contrast learning on the iterative training data to obtain a second supervised loss value;

[0281] Based on a preset second joint training weighting coefficient, the second unsupervised loss value and the second supervised loss value are weightedly added to obtain the representation learning loss.

[0282] The present application also provides an electronic device in an embodiment. Fig.12 As shown, it shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Fig.12 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0283] like Fig. 9 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1202 or a program loaded from a storage device 1208 to a random access memory (RAM) 1203. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 1203. The processing device 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0284] Typically, the following devices may be connected to the I / O interface 1205: an input device 1206 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1207 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1208 including, for example, a memory card, a hard disk, etc.; and a communication device 1209. The communication device 1209 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Fig.12 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0285] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any training method for the SAR image classification model provided in the embodiment of the present application.

[0286] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any training method for the SAR image classification model provided in the embodiment of the present application.

[0287] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.

[0288] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0289] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0290] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.

Claims

1. A training method for a SAR image classification model, characterized in that: Comprising a plurality of iterations, the iterations comprising: Acquire the iterative training data, wherein the training data includes a sample SAR image and a true value label, wherein the true value label identifies an image category; For any one of the sample SAR images, obtaining a first enhanced image and a second enhanced image of a sample characteristic image of the sample SAR image; Inputting the first enhanced image of the sample SAR image into the student model and the teacher model, obtaining the student projection head output by the student model and the teacher projection head output by the teacher model, wherein the student model is a SAR image classification model to be trained, the SAR image classification model is composed of an encoder and a first projector, and the teacher model is a mirror model of the SAR image classification model; Performing weighted fusion and centering processing on the teacher projection heads based on a weight factor to obtain a distilled balanced projection head, wherein the weight factor is inversely correlated with the number of images belonging to the image category of the sample SAR image in the training data; Input the first enhanced image and the second enhanced image into a perception projection module, obtain a first projection head corresponding to the first enhanced image output by the perception projection module and a second projection head corresponding to the second enhanced image, wherein the perception projection module shares an encoder with the SAR image classification model; Calculate the total loss of the iteration based on the true value labels of all the sample SAR images in the training data of the iteration, the distillation balanced projection head, the student projection head, the first projection head, and the second projection head; If the preset training completion conditions are met, the training ends; If the training completion condition is not met, the model parameters are updated and the next iteration is performed.

2. The training method of the SAR image classification model according to claim 1, characterized in that: The encoder is a frequency attention module, and the first projector is a linear layer; The perception projection module is composed of the encoder and a second projector, and the second projector is a multi-layer perceptron.

3. The training method of the SAR image classification model according to claim 1, characterized in that: The step of acquiring a first enhanced image and a second enhanced image of a sample characteristic image of the sample SAR image comprises: The sample SAR image is linearly projected and position-encoded to obtain the sample feature image; The sample feature image is resized back to its original size by random cropping, and the sample feature image is transformed by random color distortion and random Gaussian blur to generate a first enhanced image and a second enhanced image of the sample feature image.

4. The training method of the SAR image classification model according to claim 3, characterized in that: The calculating the total loss of the iteration based on the true value labels of all the sample SAR images in the training data of the iteration, the distillation balance projection head, the student projection head, the first projection head, and the second projection head comprises: The distilled balanced projection head and the student projection head are normalized based on a preset normalized softmax function, and the normalized result corresponding to the distilled balanced projection head is used as a soft pseudo label, and the normalized result corresponding to the student projection head is used as a predicted value; Using the first projection head and the second projection head of the sample SAR image as the first contrast value and the second contrast value respectively; Acquire a first projection head corresponding to a first enhanced image of other sample SAR images as a comparison value, wherein the other sample images are images other than the sample SAR image in the iterative training data; Calculating classification learning loss based on the soft pseudo label, the predicted value and the true value label; Calculating a representation of learning loss based on the control comparison value, the first comparison value, and the second comparison value; The sum of the classification learning loss and the representation learning loss is calculated as the total loss of the iteration.

5. The training method of the SAR image classification model according to claim 4, characterized in that: The calculating the classification learning loss based on the soft pseudo label, the predicted value and the true value label includes: Based on the soft false label and the predicted value, using a first unsupervised contrast loss function to perform unsupervised classification learning on the iterative training data to obtain a first unsupervised loss value, wherein the first unsupervised contrast loss function is equal to a cross entropy function minus an average entropy maximization function; Based on the predicted value and the true value label, using a first supervised contrast loss function to perform supervised classification learning on the iterative training data to obtain a first supervised loss value, wherein the first supervised contrast loss function is a cross entropy loss function; Based on a preset first joint training weighting coefficient, the first unsupervised loss value and the first supervised loss value are weightedly added to obtain the classification learning loss.

6. The training method of the SAR image classification model according to claim 4, characterized in that: The calculating, based on the control comparison value, the first comparison value and the second comparison value, represents the learning loss, comprises: Obtaining a positive contrast value from the control contrast value, wherein the positive contrast value is a first projection head corresponding to a first enhanced image of a sample image of the same type, and the sample image of the same type is an image of the same image category as the sample SAR image in the iterative training data; Based on the positive contrast value, the first contrast value and the second contrast value, using a second unsupervised contrast loss function to perform unsupervised contrast learning on the iterative training data to obtain a second unsupervised loss value; Based on the positive contrast value, the control contrast value, the first contrast value and the second contrast value, using a second supervised contrast loss function to perform supervised contrast learning on the iterative training data to obtain a second supervised loss value; Based on a preset second joint training weighting coefficient, the second unsupervised loss value and the second supervised loss value are weightedly added to obtain the representation learning loss.

7. A training device for a SAR image classification model, characterized in that: Comprising a joint training module, the joint training module comprising: A training data acquisition unit, configured to acquire the iterative training data, wherein the training data includes a sample SAR image and a true value label, wherein the true value label identifies an image category; An image enhancement unit, configured to obtain, for any one of the sample SAR images, a first enhanced image and a second enhanced image of a sample characteristic image of the sample SAR image; a first projection unit, configured to input the first enhanced image of the sample SAR image into a student model and a teacher model, and obtain a student projection head output by the student model and a teacher projection head output by the teacher model, wherein the student model is a SAR image classification model to be trained, the SAR image classification model is composed of an encoder and a first projector, and the teacher model is a mirror model of the SAR image classification model; A distillation balancing unit, configured to perform weighted fusion and centering processing on the teacher projection heads based on a weight factor to obtain a distillation balanced projection head, wherein the weight factor is inversely correlated with the number of images belonging to the image category of the sample SAR image in the training data; a second projection unit, configured to input the first enhanced image and the second enhanced image into a perception projection module, and obtain a first projection head corresponding to the first enhanced image output by the perception projection module and a second projection head corresponding to the second enhanced image, wherein the perception projection module shares an encoder with the SAR image classification model; A total loss calculation unit, configured to calculate the total loss of the iteration based on the true value labels of all the sample SAR images in the training data of the iteration, the distillation balance projection head, the student projection head, the first projection head, and the second projection head; The training condition determination unit is used to terminate the training if the preset training completion condition is met; if the training completion condition is not met, the model parameters are updated and the next iteration is performed.

8. A computer program product, characterized in that The method comprises computer-readable instructions, and when the computer-readable instructions are executed on an electronic device, the electronic device implements the training method of the SAR image classification model as claimed in any one of claims 1 to 6.

9. An electronic device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the training method of the SAR image classification model as described in any one of claims 1 to 6.

10. A computer storage medium, characterized in that: The storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the training method of the SAR image classification model as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for training three-dimensional target detection model based on cross-modal knowledge distillation

    CN115690708A

  • SAR (Synthetic Aperture Radar) image recognition method based on mask self-distillation network and related device

    CN118429825A

  • Systems and methods for training a video object detection machine learning model with a teacher-student framework

    DE102023212504A1

  • Systems and methods for multi-teacher group-distillation for long-tail classification

    US20240096067A1

  • Systems and methods for training video object detection machine learning model with teacher and student framework

    US20240220848A1