Multi-view medical image classification method based on low-rank domain self-adaption

By combining low-rank domain adaptive technology, summarizing bias prior information and self-attention mechanism in the medical image classification task, and implementing a multi-view multi-level feature fusion strategy, the problems of insufficient utilization of multi-view correlation and domain offset in the existing technology are solved, and a more efficient and robust medical image classification effect is achieved.

CN120125873APending Publication Date: 2025-06-10XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510076160.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize the correlation between multiple views in medical image classification tasks, and the visual basic model has domain offset problems when migrating to the medical image domain, resulting in limited performance improvement.

Method used

By combining low-rank domain adaptive technology, summarizing biased prior information and self-attention mechanism, and implementing a multi-view multi-level feature fusion strategy, an efficient and robust medical image classification model is built.

Benefits of technology

It significantly improves the classification performance and training efficiency of the model, enhances the perception ability of local structural information, and has broader applicability and stronger robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125873A_ABST
    Figure CN120125873A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-view medical image classification method based on low-rank domain self-adaption, and the method comprises the steps: firstly, inputting a multi-view medical image; secondly, low-rank domain self-adaption and multi-scale feature extraction of inductive bias perception is carried out; then, multi-scale feature multi-level fusion is carried out; and finally, performing classification prediction to obtain a final identification result. Low-rank domain self-adaption and inductive bias prior information is fused into a multi-scale feature extraction backbone network based on a self-attention mechanism, so that the field offset problem occurring when a visual basic model migrates from a general image domain to a specific task domain is effectively solved, and the sensing ability of local structure information of the visual basic model is enhanced; the model classification performance and the training efficiency are obviously improved; meanwhile, by implementing a multi-level fusion strategy of the features in the views and between the views, the unique description capability of the features in the views and the complementary advantages of the features between the views are fully utilized, the task description performance of the model is improved, and the recognition model has wider applicability and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a multi-view medical image classification method based on low-rank domain adaptation. Background Art

[0002] The medical image classification task refers to the process of automatically identifying and classifying medical image data by using computer vision and machine learning technologies. Currently, the medical image classification methods based on deep learning mainly focus on single-view analysis. However, when radiologists examine images (such as X-ray pictures), they usually comprehensively evaluate all views and utilize the valuable correlations provided between multi-views to accurately identify the types of lesions. This highlights the importance of cross-view data analysis for anomaly recognition and diagnosis in the healthcare field, and also emphasizes the advantages of multi-view-based computer-aided diagnosis schemes compared to single-image schemes.

[0003] Currently, vision foundation models pre-trained on large-scale natural image datasets, such as Swin Transformer, have demonstrated their powerful feature extraction capabilities and have been successfully applied to tasks such as classification, detection, and segmentation, achieving remarkable results in multiple fields. However, the medical image field has its uniqueness and complexity, with fine structures and often containing local subtle lesion features, which pose many challenges for vision foundation models in medical image classification tasks.

[0004] To effectively apply vision foundation models to medical image classification tasks, domain adaptation is particularly important as it can ensure that the model adapts to the uniqueness of the medical image field and optimize its performance. Currently, the mainstream methods of domain adaptation cover local fine-tuning, full fine-tuning, and low-rank domain adaptation techniques, but each method has its limitations. Although local fine-tuning attempts to adjust according to the target domain features, its characterization ability is often insufficient, resulting in limited improvement in model performance. Although global fine-tuning is more comprehensive, it requires re-training the entire network model, which not only increases the computational complexity but also requires a large training sample set. As for low-rank domain adaptation, although it optimizes the model parameters through low-rank constraints and reduces the number of parameters, it may not be able to fully describe the key local information in medical images due to the lack of sufficient inductive bias. In medical image classification tasks, these local information are crucial for accurately identifying and locating the lesion sites. Summary of the Invention

[0005] To overcome the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to provide a multi-view medical image classification method based on low-rank domain adaptation, which combines low-rank domain adaptation, inductive bias prior information and self-attention mechanism, and implements a multi-level feature fusion strategy to construct an efficient and robust medical image classification model; while solving the domain shift problem, this model enhances its ability to perceive local structural information, significantly improves the classification performance and training efficiency, and has broader applicability and stronger robustness.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0007] A multi-view medical image classification method based on low-rank domain adaptation specifically includes the following steps:

[0008] Step S1, input multi-view image data:

[0009] Let the input multi-view image data be X (i) ∈R H×W×3 , where i = {1, 2} represents the i-th view, and H and W respectively represent the height and width of the image, and H = W = 384;

[0010] Step S2, inductive bias-aware low-rank domain adaptation and multi-scale feature extraction:

[0011] Construct an inductive bias-aware low-rank domain adaptation and multi-scale feature extraction backbone network. The specific method is as follows:

[0012] In the preprocessing stage, first perform Patch partitioning on X (i) in the same way as Swin Transformer Tiny to obtain the output: Then input it into the subsequent 4 stages of the network for inductive bias-aware low-rank domain adaptation and multi-scale feature extraction;

[0013] In the first stage, first use the same linear embedding as Swin Transformer Tiny to obtain the output: where C = 96; then input it into the first low-rank domain adaptation Transformer block for feature extraction and output:

[0014] In the second stage, first use the same Patch merging as Swin Transformer Tiny to obtain the output: Then input it into the second low-rank domain adaptation Transformer block for feature extraction and output:

[0015] In the third stage, first, Patch merging is performed to obtain the output: Subsequently, three consecutive low-rank domain adaptive Transformer blocks are input for feature extraction, and the outputs of each low-rank domain adaptive Transformer block are respectively: k = 3, 4, 5; Let respectively represent the inputs of the latter two low-rank domain adaptive Transformer blocks;

[0016] In the fourth stage, first, Patch merging is performed to obtain the output: Subsequently, the sixth low-rank domain adaptive Transformer block is input for feature extraction and outputs:

[0017] The calculation process of the low-rank domain adaptive Transformer block in step 2 is as follows:

[0018] Let the input of the j-th low-rank domain adaptive Transformer block be Its corresponding output can be calculated by formula (1), where j = 1, 2,..., 6:

[0019]

[0020] Among them, W-MLRA(·) represents window multi-head low-rank domain adaptive attention, SW-MLRA(·) represents sliding window multi-head low-rank domain adaptive attention, the window size is set to win = 12, LR-MLP(·) represents low-rank domain adaptive multi-layer perceptron, and LN(·) represents layer normalization.

[0021] The calculation processes of the low-rank domain adaptive attention and low-rank domain adaptive multi-layer perceptron in step 2 are as follows:

[0022] Let the input of the low-rank domain adaptive attention be The low-rank encoder and low-rank decoder are respectively represented by the projection transformation matrices Then, the output of the low-rank domain adaptive attention is calculated by formula (2)

[0023]

[0024] Among them, The dimension of Same, SA(·) represents the self-attention operation, reshape(p, h, w) represents reshaping the first dimension of matrix p into a matrix of h×w, flatten(p, l) represents flattening the first l dimensions of p, concat(·) represents concatenating operations in the last dimension direction, conv(p, kel, pad, in, out) represents performing a 2D convolution operation on p with a kernel size of kel, a padding size of pad, an input channel number of in, an output channel number of out, and a stride of 1. The rank r = 16, and d is the value of the second dimension of; the set of learnable parameters involved in the conv(·) operation is denoted as

[0025] Let the input of the low-rank domain adaptive multi-layer perceptron be The low-rank encoder and the low-rank decoder are represented by the projection transformation matrices respectively. Then, calculate the output of the low-rank domain adaptive multi-layer perceptron according to equation (3)

[0026]

[0027] Among them, has the same dimension as Same, MLP(·) represents the multi-layer perceptron;

[0028] Step S3, multi-scale feature multi-level fusion:

[0029] Construct a multi-scale feature multi-level fusion module. The specific method is as follows:

[0030] First, perform reshape, dimension increase, and downsampling processing on the multi-scale features output at each stage according to equation (4) to obtain the feature map where m = 1, 2, 5, 6, and n = 1, 2, 3, 4:

[0031]

[0032] Among them, down(p, ratio) represents bilinearly interpolating and shrinking p by a factor of ratio. The set of learnable parameters involved in the conv(·) operation is denoted as

[0033] Secondly, perform in-view feature fusion according to equation (5) to obtain the feature map after fusion of each view

[0034]

[0035] Among them, bn_relu(·) represents batch normalization first and then relu non-linear activation. The set of learnable parameters involved in the conv(·) operation is denoted as ρ;

[0036] Finally, perform inter-view feature fusion according to equation (6) to obtain the final feature map F ∈ R 8C :

[0037] F = ga_pooling(bn_relu(conv(concat(F 1 , F 2 ), 1, 0, 16C, 8C))) #(6)

[0038] Among them, ga_pooling(·) represents the global average pooling operation. The set of learnable parameters involved in the conv(·) operation is denoted as

[0039] Step S4, classification prediction:

[0040] The classifier consists of 1 fully connected layer FC. Its learnable parameter matrix is denoted as σ ∈ R 8C×cls , where cls is the number of classes to be predicted; the prediction probability p of the classifier is p = softmax(FC(flatten(F, 3))) ∈ R cls will be used as the final recognition result, where flatten(p, l) means flattening the first l dimensions of p;

[0041] Furthermore, in steps 2, 3, and 4, the inductive bias-aware low-rank domain adaptation and multi-scale feature extraction backbone network are initialized with the parameters of the Swin Transformer Tiny network pre-trained on the ImageNet-21K dataset, and the network parameters are frozen; the learnable parameters of the feature extraction backbone network, multi-scale feature multi-level fusion module, and classifier are trained through the cross-entropy loss and the gradient descent algorithm: ρ, σ, and the initial values of each parameter are initialized with random values that follow the standard normal distribution.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] A multi-view medical image classification method based on low-rank domain adaptation, and its advantages are prominently reflected in the following steps: In step 2, low-rank domain adaptation and inductive bias prior information are incorporated into the multi-scale feature extraction backbone network based on the self-attention mechanism, effectively solving the domain shift problem that occurs when the vision-based model migrates from the general image domain to the specific task domain, and enhancing its ability to perceive local structural information, significantly improving the model classification performance and training efficiency; In step 3, by implementing a multi-level fusion strategy for intra-view and inter-view features, the unique descriptive ability of intra-view features and the complementary advantages of inter-view features are fully utilized, further improving the task characterization performance of the model, making the recognition model have wider applicability and robustness. Brief Description of the Drawings

[0044] Figure 1 is the flowchart of a multi-view medical image classification method based on low-rank domain adaptation of the present invention.

[0045] Figure 2 is the network structure diagram of a multi-view medical image classification method based on low-rank domain adaptation of the present invention.

[0046] Figure 3 is the structure diagram of the low-rank domain adaptation Transformer block of a multi-view medical image classification method based on low-rank domain adaptation of the present invention.

[0047] Figure 4 is the structure diagram of the low-rank domain adaptation attention and low-rank domain adaptation multi-layer perceptron of a multi-view medical image classification method based on low-rank domain adaptation of the present invention. Detailed Embodiment

[0048] Now, the example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this invention will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The described features or characteristics can be combined in any suitable manner in one or more embodiments.

[0049] As Figure 1 shown, a multi-view medical image classification method based on low-rank domain adaptation includes the following steps:

[0050] Step S1, input multi-view image data:

[0051] Let the input multi-view image data be X (i) ∈R H×W×3 , where i = {1, 2} represents the i-th view, H and W respectively represent the height and width of the image, and H = W = 384;

[0052] Step S2, Inductive Bias-Aware Low-Rank Domain Adaptation and Multi-Scale Feature Extraction:

[0053] Construct an inductive bias-aware low-rank domain adaptation and multi-scale feature extraction backbone network, and the specific approach is as follows:

[0054] In the preprocessing stage, first perform Patch partitioning on X (i) in the same way as Swin Transformer Tiny to obtain the output: Then input it into the subsequent 4 stages of the network for inductive bias-aware low-rank domain adaptation and multi-scale feature extraction;

[0055] In the first stage, first adopt the same linear embedding as Swin Transformer Tiny to obtain the output: where C = 96; subsequently, input it into the first low-rank domain adaptation Transformer block for feature extraction and output:

[0056] In the second stage, first perform Patch merging in the same way as Swin Transformer Tiny to obtain the output: Subsequently, input it into the second low-rank domain adaptation Transformer block for feature extraction and output:

[0057] In the third stage, first perform Patch merging to obtain the output: Subsequently, input it into 3 consecutive low-rank domain adaptation Transformer blocks for feature extraction, and the outputs of each low-rank domain adaptation Transformer block are respectively: k = 3, 4, 5; let respectively represent the inputs of the latter 2 low-rank domain adaptation Transformer blocks;

[0058] In the fourth stage, first perform Patch merging to obtain the output: Subsequently, input it into the sixth low-rank domain adaptation Transformer block for feature extraction and output:

[0059] The calculation process of the low-rank domain adaptation Transformer block in step 2 is as follows:

[0060] Let the input of the j-th low-rank domain adaptation Transformer block be Its corresponding output can be calculated by equation (1), where j = 1, 2,..., 6:

[0061]

[0062] Among them, W-MLRA(·) represents window multi-head low-rank domain adaptive attention, SW-MLRA(·) represents sliding window multi-head low-rank domain adaptive attention, the window size is set to win = 12, LR-MLP(·) represents low-rank domain adaptive multi-layer perceptron, and LN(·) represents layer normalization.

[0063] The calculation processes of the low-rank domain adaptive attention and the low-rank domain adaptive multi-layer perceptron in step 2 are as follows:

[0064] Let the input of the low-rank domain adaptive attention be The low-rank encoder and the low-rank decoder are represented by the projection transformation matrices respectively. Then, the output of the low-rank domain adaptive attention is calculated by equation (2)

[0065]

[0066] Among them, has the same dimension as . SA(·) represents self-attention operation, reshape(p, h, w) represents reshaping the first dimension of matrix p into a matrix of h×w, flatten(p, l) represents flattening the first l dimensions of p, concat(·) represents concatenation operation in the last dimension direction, conv(p, kel, pad, in, out) represents performing a 2D convolution operation on p with a kernel size of kel, a padding size of pad, an input channel number of in, an output channel number of out, and a sliding stride number of 1. The rank r = 16, and d is the value of the second dimension of . The set of learnable parameters involved in the conv(·) operation is denoted as

[0067] Let the input of the low-rank domain adaptive multi-layer perceptron be The low-rank encoder and the low-rank decoder are represented by the projection transformation matrices respectively. Then, the output of the low-rank domain adaptive multi-layer perceptron is calculated by equation (3)

[0068]

[0069] Among them, has the same dimension as . MLP(·) represents multi-layer perceptron;

[0070] Step S3, multi-scale feature multi-level fusion:

[0071] Construct a multi-scale feature multi-level fusion module, and the specific method is as follows:

[0072] First, for the multi-scale features output at each stage Perform reshape, dimension increase, and downsampling processing according to formula (4) to obtain a feature map where m = 1, 2, 5, 6 and n = 1, 2, 3, 4:

[0073]

[0074] where down(p, ratio) represents bilinear interpolation to reduce p by ratio times, and the set of learnable parameters involved in the conv(·) operation is denoted as

[0075] Secondly, perform in-view feature fusion according to formula (5) to obtain the feature map after in-view fusion

[0076]

[0077] where bn_relu(·) represents first performing batch normalization and then relu non-linear activation, and the set of learnable parameters involved in the conv(·) operation is denoted as ρ;

[0078] Finally, perform inter-view feature fusion according to formula (6) to obtain the final feature map F ∈ R 8C :

[0079] F = ga_pooling(bn_relu(conv(concat(F 1 , F 2 ), 1, 0, 16C, 8C))) #(6)

[0080] where ga_pooling(·) represents the global average pooling operation, and the set of learnable parameters involved in the conv(·) operation is denoted as

[0081] Step S4, classification prediction:

[0082] The classifier consists of 1 fully connected layer FC, and its learnable parameter matrix is denoted as σ ∈ R 8C×cls , where cls is the number of categories to be predicted; the prediction probability p of the classifier is p = softmax(FC(flatten(F, 3))) ∈ R cls will be used as the final recognition result, where flatten(p, l) represents flattening the first l dimensions of p;

[0083] Furthermore, in steps 2, 3, and 4, the inductive bias-aware low-rank domain adaptation and multi-scale feature extraction backbone network are initialized with the network parameters of the Swin Transformer Tiny network pre-trained on the ImageNet-21K dataset, and the network parameters are frozen; the learnable parameters of the feature extraction backbone network, the multi-scale feature multi-level fusion module, and the classifier are trained through the cross-entropy loss and the gradient descent algorithm: ρ, σ, and the initial values of each parameter are initialized with random values obeying the standard normal distribution.

[0084] The effects of the present invention can be further illustrated by the following simulation experiments:

[0085] I. Simulation conditions:

[0086] The simulation experiment of the present invention is carried out in the hardware environment of 8 NVIDIA A100 GPUs and the software environment of the PyTorch deep learning framework.

[0087] II. Simulation content:

[0088] The dataset used in the simulation experiment of the present invention is the internationally public multi-view mammography image set CBIS-DDSM ("A curated mammography data set for use in computer-aided detection and diagnosis research", Scientific data, vol. 4, no. 1, pp. 1–9, 2017).

[0089] In order to better compare with the method proposed by the present invention, the following 4 benchmark methods are set in the simulation experiment:

[0090] Multi-view (without low-rank domain adaptation): Replace the low-rank domain adaptation attention module in step 2 with a general attention module;

[0091] Multi-view (without bias inductive perception): Remove the 3 parallel convolutions in the low-rank domain adaptation attention module in step 2;

[0092] Single-view (mediolateral oblique): Adjust the multi-view input in step 1 to a mediolateral oblique single-view input, and at the same time remove the inter-view feature fusion in step 3;

[0093] Single-view (craniocaudal): Adjust the multi-view input in step 1 to a craniocaudal single-view input, and at the same time remove the inter-view feature fusion in step 3.

[0094] III. Analysis of simulation effects:

[0095] Table 1 shows the comparison of the classification accuracies obtained by three methods in the simulation. As can be seen from Table 1, the present invention incorporates low-rank domain adaptation and inductive bias prior information into the multi-scale feature extraction backbone network based on the self-attention mechanism, which can significantly improve the performance of the vision-based model in medical image classification tasks. In addition, after fusing multi-view features, the model further utilizes the complementary advantages of features between views, enhancing the model's task characterization ability.

[0096] Table 1 List of classification accuracies obtained by three methods in the simulation

[0097] Simulation method Classification accuracy The method of the present invention 71.7% Multi-view (without low-rank domain adaptation) 65.3% Multi-view (without bias inductive perception) 69.2% Single-view (internal and external oblique positions) 67.9% Single-view (cephalocaudal position) 68.4%

[0098] In summary, the present invention constructs an efficient and robust medical image classification model by combining low-rank domain adaptation technology, inductive bias prior information and self-attention mechanism, and implementing a multi-view multi-level feature fusion strategy. While addressing the domain shift problem, the model effectively enhances its ability to capture local structural information, thus significantly improving the classification accuracy and training efficiency. This design endows the model with broader adaptability and stronger stability, making it have broad application prospects in the field of medical image analysis.

[0099] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed by the present invention. The specification and examples are only illustrative, and the true scope and spirit of the present invention are pointed out by the appended claims.

Claims

1. A multi-view medical image classification method based on low-rank domain adaptation, characterized in that: The specific steps include: Step S1, input multi-view image data: Assume that the input multi-view image data is X (i) ∈R H×W×3 , where i = {1, 2} represents the i-th view, H and W represent the height and width of the image respectively, H = W = 384; Step S2, inductive bias-aware low-rank domain adaptation and multi-scale feature extraction: Construct an inductive bias-aware low-rank domain adaptation and multi-scale feature extraction backbone network. The specific approach is as follows: In the preprocessing stage, X (i) Perform the same patch division as Swin Transformer Tiny and get the output: Then it is input into the subsequent four stages of the network for inductive bias-aware low-rank domain adaptation and multi-scale feature extraction; Step S3, multi-scale feature multi-level fusion: Construct a multi-scale feature multi-level fusion module. The specific steps are as follows: First, the multi-scale features output by each stage According to formula (1), reshape, dimension increase and downsampling are performed to obtain the feature map Where m = 1, 2, 5, 6, n = 1, 2, 3, 4: Among them, reshape(p,h,w) means reshaping the first dimension of matrix p into a h×w matrix, conv(p,kel,pad,in,out) means performing a 2D convolution operation on p with a kernel size of kel, a padding size of pad, an input channel number of in, an output channel number of out, and a sliding stride of 1, down(p,ratio) means reducing p by a factor of ratio by bilinear interpolation, and the set of learnable parameters involved in the conv(·) operation is recorded as Secondly, perform intra-view feature fusion according to formula (2) to obtain the fused feature maps of each view. Among them, concat(·) means concatenation operation in the last dimension direction, bn_relu(·) means batch normalization first, then relu nonlinear activation, and the set of learnable parameters involved in conv(·) operation is denoted as ρ; Finally, according to formula (3), the inter-view feature fusion is performed to obtain the final feature map F∈R 8C : F=ga_pooling(bn_relu(conv(concat(F1,F2),1,0,16C,8C)))#(3) Among them, ga_pooling(·) represents the global average pooling operation, and the set of learnable parameters involved in the conv(·) operation is recorded as Step S4, classification prediction: The classifier consists of a fully connected layer FC, and its learnable parameter matrix is ​​denoted by σ∈R 8C×cls , where cls is the number of categories to be predicted; the predicted probability of the classifier p = softmax(FC(flatten(F,3)))∈Rcls will be used as the final recognition result, where flatten(p,l) means flattening the first l dimensions of p.

2. The multi-view medical image classification method based on low-rank domain adaptation according to claim 1, characterized in that: The low-rank domain adaptive Transformer block described in step 2 is calculated as follows: Let the input of the j-th low-rank domain adaptive Transformer block be The corresponding output can be calculated by formula (4), where j = 1, 2, ..., 6: Among them, W-MLRA(·) represents windowed multi-head low-rank domain adaptive attention, SW-MLRA(·) represents sliding window multi-head low-rank domain adaptive attention, the window size is set to win=12, LR-MLP(·) represents low-rank domain adaptive multilayer perceptron, and LN(·) represents layer normalization.

3. The multi-view medical image classification method based on low-rank domain adaptation according to claim 2, characterized in that: The calculation process of low-rank domain adaptive attention and low-rank domain adaptive multi-layer perceptron is as follows: Assume that the input of low-rank domain adaptive attention is The low-rank encoder and low-rank decoder use the projection transformation matrix Represented by, then the output of low-rank domain adaptive attention is calculated by (5) in, The dimensions and Same, SA(·) represents the self-attention operation, rank r = 16, d is The second dimension value of ; the set of learnable parameters involved in the conv(·) operation is recorded as Assume that the input of the low-rank domain adaptive multilayer perceptron is The low-rank encoder and low-rank decoder use the projection transformation matrix Represented by, then the output of the low-rank domain adaptive multilayer perceptron is calculated by (6) in, The dimensions and Similarly, MLP(·) represents a multi-layer perceptron.

4. The multi-view medical image classification method based on low-rank domain adaptation according to claim 1, characterized in that: The training process of the entire network is as follows: The inductive bias-aware low-rank domain adaptation and multi-scale feature extraction backbone network is initialized using the Swin Transformer Tiny network parameters pre-trained on the ImageNet-21K dataset, and the network parameters are frozen; the learnable parameters of the feature extraction backbone network, multi-scale feature multi-level fusion module, and classifier are trained using the cross entropy loss and gradient descent algorithm: ρ, σ, and the initial values ​​of each parameter are initialized with random values ​​that obey the standard normal distribution.