Industrial anomaly detection method based on semi-supervised contrast learning
By adopting a semi-supervised contrast learning method in industrial anomaly detection, a two-branch contrast learning architecture and a shared weight encoder are used, combining label and label-free data to generate discriminant feature representations, solving the problem of scarcity of label data and high-dimensional data processing, and achieving efficient abnormality detection.
Patent Information
- Application Number
- CN202510311759.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art faces the scarcity of label data and high-dimensional data processing problems in industrial anomaly detection, making it difficult to generate effective discriminant feature representations.
Using the industrial anomaly detection method (CLAD) based on semi-supervised contrast learning, a discriminant feature representation is generated through a dual-branch contrast learning architecture and a shared weight encoder, combining label data and label-free data. Specific steps include data augmentation, iterative training and reverse adjustment of model parameters.
The accuracy and robustness of anomaly detection is significantly improved, and 98.7% AUROC can be achieved using only 10% of the label data, which is better than the existing traditional semi-supervised approach.
Smart Images

Figure CN120198397A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial image anomaly detection, and specifically relates to an industrial anomaly detection method based on semi-supervised contrastive learning. Background Art
[0002] In modern industrial manufacturing systems, anomaly detection is a key technology for ensuring production safety and quality assurance. Traditional anomaly detection methods usually rely on a large amount of labeled data and manual rules. However, in actual industrial scenarios, due to the rarity of defect occurrences and the high cost of expert annotation, obtaining a large amount of labeled data is both expensive and time-consuming. Although existing semi-supervised learning methods have alleviated the problem of scarce labeled data to a certain extent, when dealing with high-dimensional industrial data, it is often difficult to generate effective discriminative feature representations. As an unsupervised representation learning technique, contrastive learning learns discriminative features of data by contrasting positive and negative sample pairs and has achieved remarkable results in the visual field. However, existing contrastive learning methods still face two main challenges in industrial anomaly detection: 1) Ambiguity of negative samples: Randomly sampling negative samples from unlabeled data may lead to the mixing of potential anomaly samples, breaking the compactness assumption of normal samples; 2) Underutilization of partial labels: Most contrastive learning methods only use unsupervised learning in industrial anomaly detection, ignoring the small amount of labeled data available in actual production systems. Summary of the Invention
[0003] The object of the present invention is to propose an industrial anomaly detection method (CLAD) based on semi-supervised contrastive learning, namely Semi-Supervised Contrastive Learning with Augmented Negatives, aiming to solve the problem of scarce labeled data in the prior art. The core idea of the present invention is to generate discriminative feature representations by combining labeled data and unlabeled data through a two-branch contrastive learning architecture.
[0004] The technical solution adopted by the present invention to achieve the above object is as follows:
[0005] An industrial anomaly detection method based on semi-supervised contrastive learning, comprising the following steps:
[0006] Perform data augmentation on the image dataset of the set scenario to expand negative samples;
[0007] Construct a CLAD network model for semi-supervised contrastive learning through a two-branch contrastive learning architecture and a shared-weight encoder method; based on the dataset after data augmentation, perform iterative training on the CLAD network model and reverse-adjust the model parameters to obtain an ideal model, and the ideal model identifies and detects the input image and outputs discriminative feature representations;
[0008] Collect the images of the actual industrial scenario and input them into the ideal model to automatically detect whether the current industrial scenario is abnormal.
[0009] The image dataset of the set scenario uses the MVTec-AD dataset, including image data of 10 object categories and 5 texture categories.
[0010] The image dataset is marked and divided according to whether there is a label and whether it is a positive or negative sample.
[0011] The semi-supervised contrastive learning CLAD network model adopts a two-branch contrastive learning architecture:
[0012] The supervised contrastive learning branch is used to expand the labeled negative sample image data and perform explicit category-aware contrastive learning;
[0013] The unsupervised contrastive learning branch is used to automatically discover the potential relationships between the unlabeled positive sample image data through the self-supervised contrastive learning mechanism.
[0014] Automatically discover potential data relationships through similarity-aware instance discrimination.
[0015] Both the supervised or unsupervised contrastive learning branches include:
[0016] An enhancement module that uses the CutPaste data enhancement method to generate negative sample images based on positive samples;
[0017] An encoder module that adopts the network structure of a two-branch ResNet-50 encoder;
[0018] A projection head module that maps the features output by the encoder module to a low-dimensional space for contrastive learning to generate discriminative representations.
[0019] The encoder modules of the supervised or unsupervised contrastive learning branches share weights through the regularization term γ∥θ1 - θ2∥2 to make the feature consistency; where θ1 and θ2 are the weight parameters of the two encoder modules respectively.
[0020] Iteratively train the CLAD network model and inversely adjust the model parameters to obtain the ideal model, including:
[0021] 1) Initialize the ResNet-50 encoder and projection head with shared weights;
[0022] 2) Sample a batch of data from the labeled data and unlabeled data respectively and input them into the two-branch contrastive learning architecture;
[0023] 3) Perform supervised contrastive learning on the labeled data and calculate the supervised contrastive loss
[0024] 4) Perform unsupervised contrastive learning on the unlabeled data and calculate the unsupervised contrastive loss.
[0025] 5) Calculate the total loss for joint optimization. And update the model parameters by the gradient descent method.
[0026]
[0027] Among them, β and γ are hyperparameters that respectively control the influence of the unsupervised contrastive loss and the weight consistency constraint.
[0028] 6) Repeat the above steps until the total loss function of the model converges, and stop the iteration to obtain the ideal model.
[0029] During the training process, the Adam optimizer is used for parameter update.
[0030] The present invention has the following beneficial effects and advantages:
[0031] 1. Efficient utilization of labeled data: By enhancing negative sample contrastive learning through CutPaste, the present invention can generate discriminative feature representations under limited labeled data, significantly improving the accuracy of anomaly detection.
[0032] 2. Adaptive learning of unlabeled data: The present invention designs an adaptive contrastive learning mechanism that can automatically discover the potential relationships in unlabeled data without explicit annotation, further enhancing the generalization ability of the model.
[0033] 3. Shared-weight encoder architecture: Through the shared-weight encoder architecture, the present invention ensures the feature consistency between labeled data and unlabeled data, avoiding the divergence problem of the model during joint optimization.
[0034] 4. Significant performance improvement: In the MVTec AD benchmark test, the present invention can achieve 98.7% AUROC by using only 10% of the labeled data, significantly outperforming existing traditional semi-supervised methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a schematic diagram of the overall framework of the method of the present invention.
[0036] Figure 2 It is the result of effect comparison on the MVTEC AD dataset.
[0037] Figure 3 It is an ROC curve graph.
[0038] Figure 4 It is a PR curve graph.
[0039] Figure 5It is a t-SNE visualization graph. Detailed implementation manners
[0040] To make the above objects, features and advantages of the present invention more obvious and understandable, the following will describe the specific implementation methods of the present invention in detail with reference to the accompanying drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the invention. Therefore, the present invention is not limited by the specific implementations disclosed below.
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0042] I. This application mainly targets industrial anomaly detection scenarios, and the data objects include:
[0043] 1. Image data: from the MVTec-AD dataset, covering 15 categories, including 10 object categories (such as bottles, cables, capsules, etc.) and 5 texture categories (such as grids, wood, etc.).
[0044] 2. Normal samples and abnormal samples: Normal samples are standard product images in industrial production, and abnormal samples contain various defects (such as scratches, cracks, stains, etc.).
[0045] 3. Label data: Due to the scarcity of abnormal samples and the high annotation cost, there are only a small number of labeled abnormal samples in the dataset, and most of the data is unlabeled.
[0046] II. An industrial anomaly detection method (CLAD) based on semi-supervised contrastive learning proposed by the present invention, namely Semi-Supervised Contrastive Learning with Augmented Negatives, aims to solve the problem of scarce label data in the prior art. The core idea of the present invention is to generate discriminative feature representations by combining label data and unlabeled data through a two-branch contrastive learning architecture. Specifically, the present invention includes the following three main innovative points:
[0047] 1. CutPaste - based Enhanced Negative - sample Contrastive Learning: For limited labeled data, the present invention uses the CutPaste data augmentation technique to generate negative samples and generates discriminative feature representations through explicit class - aware contrastive learning. The CutPaste technique simulates defect variations in the real world by randomly cropping and pasting rectangular regions in normal samples while maintaining geometric consistency.
[0048] 2. Adaptive Contrastive Learning for Unlabeled Data: For a large amount of unlabeled data, the present invention designs an adaptive contrastive learning mechanism that automatically discovers potential data relationships through similarity - aware instance discrimination. This mechanism can effectively utilize the structural information in unlabeled data without explicit annotation.
[0049] 3. Encoder Architecture with Shared Weights: The present invention adopts a dual - branch encoder architecture with shared weights, combined with hierarchical projection heads, to ensure joint representation learning between supervised anomaly discrimination and unsupervised representation alignment. Through the weight - sharing mechanism, the present invention ensures feature consistency between labeled data and unlabeled data.
[0050] The present invention has conducted extensive experiments on the MVTec AD benchmark. The results show that only 10% of the labeled data can achieve an AUROC of 98.7%, significantly outperforming existing traditional semi - supervised methods. The success of the present invention provides an efficient and low - cost solution for industrial anomaly detection and has broad application prospects.
[0051] III. Overview of the System Architecture
[0052] The specific implementation manner of the present invention relates to an industrial anomaly detection system based on semi - supervised contrastive learning, aiming to solve the problem of scarce labeled data in industrial scenarios by combining contrastive learning and data augmentation techniques. The system mainly includes the following modules:
[0053] 1) Data Pre - processing Module: Responsible for pre - processing the input industrial images, including operations such as image enhancement and normalization.
[0054] 2) Dual - branch Contrastive Learning Module: Consisting of two branches, which separately process labeled data and unlabeled data and extract features through contrastive learning.
[0055] 3) Shared - weight Encoder Module: Used to ensure the consistency of labeled data and unlabeled data during feature extraction.
[0056] 4) Anomaly Detection Module: Based on the learned features, uses the Kernel Density Estimation (KDE) model for anomaly detection.
[0057] 5) Model Optimization and Hyperparameter Tuning Module: Further improves the performance of the system through hyperparameter tuning and model optimization.
[0058] 2. Data preprocessing module
[0059] In the data preprocessing module, the system first preprocesses the input industrial images to ensure the quality and consistency of the data. The specific steps are as follows:
[0060] 1) Image enhancement: To increase the diversity of the data, the system performs enhancement operations such as random cropping, horizontal flipping, and color jittering on the images. These operations help the model learn more robust features during training. For example, random cropping can simulate images from different perspectives, horizontal flipping can increase the symmetry of the image, and color jittering can simulate images under different lighting conditions.
[0061] 2) Normalization: The images are normalized to scale the pixel values to a fixed range (such as [0,1]) to accelerate the convergence of the model. Normalization can reduce the differences between images, making it easier for the model to learn effective features.
[0062] 3) Data splitting: The dataset is divided into two parts: labeled data and unlabeled data. The labeled data contains a small number of normal samples and abnormal samples, while the unlabeled data is mainly composed of normal samples and may contain a small number of unlabeled abnormal samples. This data splitting method simulates the situation where labeled data is scarce in actual industrial scenarios.
[0063] 3. Dual-branch contrastive learning module
[0064] As Figure 1 shown, it shows the dual-branch contrastive learning architecture, including enhanced negative sample contrastive learning based on CutPaste and adaptive contrastive learning of unlabeled data. In the figure, Supervised Contrastive Learning is the supervised contrastive learning branch, Unsupervised Contrastive Learning is the unsupervised contrastive learning branch, NegativeAugmentation is the negative sample enhancement module, Positive Augmentation is the positive sample enhancement module, ResNetencoder is the encoder, and Projection Head is the projection head. The dual-branch contrastive learning module is the core part of the present invention, aiming to extract effective feature representations from labeled data and unlabeled data through contrastive learning. That is, feature embedding: The test samples are subjected to feature extraction through the trained encoder to obtain high-dimensional feature representations. These feature representations contain the semantic information of the samples and can effectively distinguish normal samples and abnormal samples. This module contains two branches: the supervised contrastive learning branch and the unsupervised contrastive learning branch.
[0065] 3.1 Supervised contrastive learning branch
[0066] The supervised contrastive learning branch is mainly used to process labeled positive and negative sample data, and enhance the model's discriminative ability for abnormal samples through contrastive learning. The specific steps are as follows:
[0067] 1) Sample pair generation: For each labeled sample, the system generates positive sample pairs and negative sample pairs. Positive sample pairs are composed of normal samples of the same category, while negative sample pairs are composed of normal samples and abnormal samples generated by CutPaste augmentation. The CutPaste augmentation technique simulates abnormalities by randomly cropping and pasting a rectangular area in a normal image, thus generating negative sample pairs.
[0068] 2) Feature extraction: Use a ResNet-50 encoder with shared weights to extract features from the samples, obtaining high-dimensional feature representations. ResNet-50 is a deep convolutional neural network with strong feature extraction capabilities, capable of effectively capturing detailed information in images.
[0069] Contrastive loss calculation: By calculating the contrastive loss, optimize the model parameters to make the feature representations of positive sample pairs as close as possible, while the feature representations of negative sample pairs are as far apart as possible. The calculation formula for the contrastive loss is as follows:
[0070]
[0071] where P(i) represents positive sample pairs, A(i) represents negative sample pairs, τ is the temperature parameter used to control the softness and hardness of the similarity distribution, z i represents the feature representation of sample i, z p represents the feature representation of positive sample p, z a represents the feature representation of negative sample a, B represents the set of all samples in the current training batch, and exp represents the exponential.
[0072] 3.2 Unsupervised contrastive learning branch
[0073] The unsupervised contrastive learning branch is used to process unlabeled positive and negative sample data, and automatically discover potential relationships in the data through self-supervised contrastive learning. The specific steps are as follows:
[0074] 1) Sample pair generation: For each unlabeled sample, the system generates positive sample pairs through data augmentation and uses other samples in the same batch as negative sample pairs. Data augmentation operations include random cropping, horizontal flipping, and color jittering, etc., to increase the diversity of the data.
[0075] 2) Feature extraction: Use a ResNet-50 encoder with shared weights to extract features from the samples, obtaining high-dimensional feature representations.
[0076] 3) Contrastive Loss Calculation: By calculating the unsupervised contrastive loss, the model parameters are optimized to make the feature representations of positive sample pairs as close as possible, while making the feature representations of negative sample pairs as far away as possible. The formula for the unsupervised contrastive loss is as follows:
[0077]
[0078] where τ is the temperature parameter used to control the softness and hardness of the similarity distribution, k represents the sample index, z i represents the feature representation of sample i, B represents the set of all samples in the current training batch, z j represents the feature representation of positive sample j, z k represents the feature representation of negative sample k, and exp represents the exponential function.
[0079] This loss function enables the model to learn effective feature representations from unlabeled data by maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs.
[0080] 4. Shared-Weight Encoder Module
[0081] To ensure consistency between labeled and unlabeled data during feature extraction, the system adopts an encoder architecture with shared weights. The specific steps are as follows:
[0082] 1) Encoder Initialization: The system initializes two ResNet-50 encoders and sets their weights to be shared, i.e., θ1 = θ2. This weight-sharing mechanism ensures that the same feature extractor is used for both labeled and unlabeled data during feature extraction, thus avoiding inconsistencies in feature representations.
[0083] 2) Weight Sharing: After feature extraction of labeled and unlabeled data through two encoders with or without supervision respectively, high-dimensional feature representations are obtained. Due to the weight sharing of the encoders, the features extracted from the two branches are consistent, enabling the model to better utilize labeled and unlabeled data.
[0084] 3) Weight Consistency Constraint: During training, the system ensures the consistency of the weights of the two encoders through the regularization term γ∥θ1 - θ2∥2, preventing weight divergence during the joint optimization process. This regularization term enforces consistency by penalizing the difference between the weights of the two encoders.
[0085] 5. Anomaly Detection Module (Projection Head Module)
[0086] In the anomaly detection module, the system uses a Kernel Density Estimation (KDE) model for anomaly detection based on the learned feature representations. The specific steps are as follows:
[0087] 1) Kernel density estimation: Use the kernel density estimation model to calculate the anomaly score of the test sample. Kernel density estimation is a non-parametric probability density estimation method that can calculate the probability density according to the characteristics of the sample. The higher the anomaly score, the more likely the sample is an abnormal sample.
[0088] 2) Anomaly determination: Determine whether the sample is an abnormal sample according to the preset threshold. By adjusting the threshold, the sensitivity and specificity of the model can be controlled to adapt to different application scenarios.
[0089] 6. Model Optimization and Hyperparameter Tuning Module
[0090] To further improve the performance of the system, the present invention also designs a model optimization and hyperparameter tuning module. This module ensures that the system can perform well under different datasets and scenarios through hyperparameter tuning and model optimization. The specific steps are as follows:
[0091] 1) Hyperparameter tuning: The system tunes the hyperparameters through the grid search method, including the weight β of the unsupervised contrast loss, the coefficient γ of the weight consistency constraint, and the temperature parameter τ in the contrast loss. By conducting experiments under different parameter combinations, the optimal hyperparameter combination is selected.
[0092] 2) Model optimization: During the training process, the system uses the Adam optimizer to update the parameters. The Adam optimizer combines the momentum method and the adaptive learning rate adjustment strategy, which can effectively accelerate the convergence of the model and avoid falling into local optima.
[0093] 3) Early stopping strategy: To prevent the model from overfitting, the system adopts an early stopping strategy. Monitor the performance of the model on the validation set, and when the validation loss no longer decreases, terminate the training process in advance.
[0094] 7. Training Process
[0095] The training process of the system is as follows:
[0096] 1) Initialization: Initialize the ResNet-50 encoder and projection head with shared weights.
[0097] 2) Batch sampling: Sample a batch of data from the labeled data and unlabeled data respectively.
[0098] 3) Supervised contrastive learning: Perform supervised contrastive learning on the labeled data and calculate the supervised contrast loss
[0099] 4) Unsupervised contrastive learning: Perform unsupervised contrastive learning on the unlabeled data and calculate the unsupervised contrast loss
[0100] 5) Joint Optimization: Calculate the total loss and update the model parameters using gradient descent. The formula for the total loss is as follows:
[0101]
[0102] where β and γ are hyperparameters that control the influence of the unsupervised contrastive loss and the weight consistency constraint, respectively.
[0103] 6) Iterative Training: Repeat the above steps until the model converges.
[0104] 8. Experimental Results
[0105] To verify the effectiveness of the present invention, experiments were conducted on the MVTec-AD industrial anomaly detection dataset, as Figure 2 shown. The experimental results show that the present invention achieved 87.0% AU-ROC and 91.9% AU-PR with only 10% labeled data, significantly outperforming existing semi-supervised anomaly detection methods. The specific experimental results are as follows:
[0106] 1) Comparison with Baseline Methods: Compared with traditional methods such as SVDD, Convolutional Autoencoder (CAE), Deep SVDD, and Deep SAD, the present invention has achieved significant improvements in both AU-ROC and AU-PR metrics. Especially in the AU-PR metric, the present invention has improved by 11.4% compared to existing semi-supervised methods, indicating that the present invention has significant advantages in detecting rare anomaly samples.
[0107] 2) As Figure 3 (a)-(d) are ROC curves showing the anomaly detection performance of the present invention on the MVTec AD dataset. As Figure 4 (a)-(d) are PR curves showing the precision-recall performance of the present invention on the MVTec AD dataset.
[0108] By plotting the ROC curve and the PR curve, it can be intuitively observed that the present invention has excellent performance in distinguishing normal samples and anomaly samples. The ROC curve shows a high true positive rate (TPR) and a low false positive rate (FPR), while the PR curve shows high precision and high recall, indicating that the present invention has high accuracy and robustness in detecting anomaly samples.
[0109] 3) As Figure 5 (a)-(d) are t-SNE visualization graphs showing that the feature representations generated by contrastive learning can clearly distinguish normal samples and anomaly samples.
[0110] Feature visualization: Through the t-SNE visualization technique, the present invention reduces the high-dimensional feature representation to a two-dimensional space, and it is observed that normal samples and abnormal samples are significantly separable in the feature space. This indicates that the present invention can effectively extract discriminative features, thereby improving the performance of anomaly detection.
[0111] The present invention proposes an efficient semi-supervised industrial anomaly detection method by combining contrastive learning and data augmentation techniques. This method can still effectively extract discriminative features in the case of scarce labeled data, significantly improving the accuracy and robustness of anomaly detection. Experimental results show that the present invention has broad application prospects in industrial anomaly detection tasks.
[0112] The above are only the preferred embodiments of the present invention and do not impose any limitations on the present invention. Any simple modifications, changes, and equivalent structural changes made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. An industrial anomaly detection method based on semi-supervised contrastive learning, characterized in that: The following steps are involved: Perform data augmentation on the image dataset of the set scene to expand the negative samples; A CLAD network model for semi-supervised contrastive learning is constructed through a dual-branch contrastive learning architecture and a shared weight encoder method. Based on the data-enhanced dataset, the CLAD network model is iteratively trained and the model parameters are reversely adjusted to obtain an ideal model, which recognizes and detects the input image and outputs a discriminative feature representation. Collect images of actual industrial scenes and input them into the ideal model to automatically detect whether the current industrial scene is abnormal.
2. The industrial anomaly detection method based on semi-supervised contrastive learning according to claim 1 is characterized in that: The image dataset of the setting scene adopts the MVTec-AD dataset, which includes image data of 10 object categories and 5 texture categories.
3. The industrial anomaly detection method based on semi-supervised contrastive learning according to claim 1 is characterized in that: The image data set is marked and divided according to whether it has a label and whether it is a positive or negative sample.
4. The industrial anomaly detection method based on semi-supervised contrastive learning according to claim 1 is characterized in that: The semi-supervised contrastive learning CLAD network model adopts a dual-branch contrastive learning architecture: The supervised contrastive learning branch is used to expand the labeled negative sample image data and perform explicit category-aware contrastive learning; The unsupervised contrastive learning branch is used to automatically discover the potential relationship between unlabeled positive sample image data through a self-supervised contrastive learning mechanism. Automatically discover latent data relationships through similarity-aware instance discrimination.
5. The industrial anomaly detection method based on semi-supervised contrastive learning according to claim 4 is characterized in that: The supervised or unsupervised contrastive learning branch includes: The enhancement module uses the CutPaste data enhancement method to generate negative sample images based on positive samples; The encoder module adopts the network structure of the dual-branch ResNet-50 encoder; The projection head module maps the features output by the encoder module to a low-dimensional space for contrastive learning and generates discriminative representations.
6. The industrial anomaly detection method based on semi-supervised contrastive learning according to claim 5, characterized in that: The encoder modules with or without supervised contrastive learning branches share weights through the regularization term γ||θ1-θ2||2 to ensure feature consistency; wherein θ1 and θ2 are weight parameters of the two encoder modules respectively.
7. The industrial anomaly detection method based on semi-supervised contrastive learning according to claim 1 is characterized in that: Iteratively train the CLAD network model and reversely adjust the model parameters to obtain the ideal model, including: 1) Initialize the ResNet-50 encoder and projection head with shared weights; 2) Sample a batch of data from the labeled data and the unlabeled data respectively, and input them into the dual-branch contrastive learning architecture respectively; 3) Perform supervised contrastive learning on the labeled data and calculate the supervised contrast loss 4) Perform unsupervised contrastive learning on unlabeled data and calculate unsupervised contrastive loss 5) Calculate the total loss of joint optimization And update the model parameters through the gradient descent method; Among them, β and γ are hyperparameters, which control the influence of unsupervised contrast loss and weight consistency constraint respectively; 6) Repeat the above steps until the total loss function of the model converges, and stop iterating to obtain the ideal model.
8. The industrial anomaly detection method based on semi-supervised contrastive learning according to claim 5, characterized in that: During the training process, the Adam optimizer is used to update the parameters.