Multi-modal mouse experiment data trend prediction and analysis method based on deep learning

The multimodal analysis framework was constructed through deep learning methods, which solved the problem of difficult analysis of mouse experimental data, and achieved efficient evaluation and analysis of mouse CT, HE, Masson, TEM, and IHC data, improved the accuracy and robustness of the model, and evaluated the effectiveness of integrated traditional Chinese and Western medicine treatment.

CN120339203APending Publication Date: 2025-07-18CHONGQING ACAD OF CHINESE MATERIA MEDICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510376316.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In modern medical experiments, the experimental data analysis of mice CT, HE, Masson, TEM, and IHC has problems such as small sample size, unevenness, and unclear pictures, which makes it difficult for researchers to evaluate and analyze data between groups.

Method used

A multimodal analysis method based on deep learning, including data preprocessing, network model construction, self-distillation training and voting strategy, and a convolutional neural network was used to analyze the experimental data of mouse CT, HE, Masson, TEM, and IHC, and a multimodal learning framework was constructed. The CBS, FFM, and GAM modules were used for feature extraction and fusion, and the self-distillation training and Focal loss loss function were used to combine the voting strategy for result inference.

Benefits of technology

It improves the evaluation effectiveness of drug intervention animal experimental models in diagnosis and treatment, improves the accuracy and robustness of the model, can quantitatively analyze trends and indicator correlations between groups, and evaluates the effectiveness of integrated traditional Chinese and Western medicine treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a multi-modal mouse experimental data trend prediction and analysis method based on deep learning, which comprises the steps of acquiring a to-be-analyzed pathological section picture, enhancing off-line data, constructing a network model, enhancing on-line data, performing self-distillation training and performing a voting strategy, and aims at quantitatively analyzing experimental data results of CT, HE, Masson, TEM and IHC of a mouse by applying a convolutional neural network (CNN). A powerful multi-modal learning framework, namely a multi-index picture data analysis model, is established through methods such as data preprocessing and enhancement, model architecture design, model training, model evaluation, model optimization and evaluation and analysis of experimental results, trends among groups are predicted through quantitative analysis, and a multi-index picture data analysis model is established. And displaying the relevance among the indexes through pictures, namely integrally analyzing the relevance among the indexes, so as to analyze the treatment response of the lung cancer complicated with the idiopathic pulmonary fibrosis and evaluate the effectiveness of the combined treatment of the traditional Chinese medicine and the western medicine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical image processing, and particularly relates to a method for predicting and analyzing trends of multi-modal mouse experimental data based on deep learning. Background Art

[0002] In the basic research of modern medical experiments, it is very common to separately analyze the experimental data of mouse CT, HE, Masson, electron microscopy (TEM), and immunohistochemistry (IHC). However, due to the lack of proficiency in experimental operations by researchers and the difficulty in the modeling process of animal models, there are often some data problems such as a small number of experimental samples collected in animal experiments, uneven number of samples in each group, unclear pathological section pictures, and incomplete picture acquisition during picture collection, resulting in difficulties for researchers in evaluating, quantifying, and analyzing the experimental data among groups. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for predicting and analyzing trends of multi-modal mouse experimental data based on deep learning to solve the above problems of data analysis difficulties and improve the evaluation of the effectiveness of drug intervention in animal experimental models in diagnosis and treatment.

[0004] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows: A method for predicting and analyzing trends of multi-modal mouse experimental data based on deep learning, comprising the following steps: Step 1: Obtain the pathological section pictures to be analyzed; Step 2: Offline data augmentation: Preprocess the pathological section pictures in an offline state; Step 3: Construct a network model: The network model is constructed by stacking several CBS modules, FFM modules, and GAM modules, and a classification head CH is set at the end of the model for classification; Step 4: Online data augmentation: Send the pictures into the network model, perform online data augmentation on each picture during the data loading stage, randomly select pictures from the obtained pictures for mosaic mode splicing, and the obtained large pictures enter the network model for training; Step 5: Self-distillation training: Perform self-distillation training on the network model, determine the loss function, and obtain multiple models of different modalities; Step 6: Voting strategy: Use the voting strategy for the model to perform inference on the data corresponding to different modalities to obtain the results.

[0005] To further implement the present invention, the preprocessing in step two includes random vertical flipping, random rotation, random scaling, random brightness adjustment, random hue-saturation adjustment, random contrast adjustment, and random Gaussian blur.

[0006] To further implement the present invention, the online data augmentation in step four includes random vertical flipping, random rotation, random scaling, random brightness adjustment, random hue-saturation adjustment, and random contrast adjustment.

[0007] To further implement the present invention, the random vertical flipping is performed with a probability of 50%, the random rotation is controlled to rotate within the range of -15° to 15° around the center point, the random scaling is controlled within the range of 0.75 to 1.25 times, the random brightness adjustment is controlled within the range of -40 to 40, the random hue-saturation adjustment is controlled within the range of -20 to 20, the random contrast is controlled within the range of 0.8 to 1.2, and the random Gaussian blur is controlled within the range of 0 to 1.

[0008] To further implement the present invention, the CBS module in step three is composed of a convolutional layer, a batch normalization layer, and a SiLU activation function, and the formula is expressed as: where, is the input feature of the CBS module; SiLU is the SiLU activation function, BN is the batch normalization layer, and conv is the convolutional layer; The FFM module performs feature fusion at different levels to improve the model accuracy and discrimination ability; the FFM module first extracts features of the input feature through a CBS module to obtain the feature , and then is split into and by channel, and is sent to the multi-layer CBS-Block module for feature extraction. Among them, the features extracted by each layer of the CBS-Block module will be combined with a skip connection and and after feature extraction through the multi-layer CBS-Block module are concatenated to restore the original number of channels, so as to achieve the fusion between features at different levels. The formula is expressed as: where, is the feature input to the FFM module; represents splitting the extracted feature by channel into and ; is the feature after fusing features at different levels through residual connections; The operation represents concatenating features for different levels of features along the channel dimension; The GAM module is an attention module that performs attention weighting on features at both the channel and spatial levels, helping the model to better focus on regions and channels that are beneficial for the model to make correct classifications; for the input features Perform a Reshape operation to achieve spatial feature compression, and then respectively pass through a fully connected layer, a ReLU activation layer, and a fully connected layer to learn the channel attention weights. Finally, adjust to the original dimension of the input features through Reshape and add them to the input features to perform weighting and obtain features with channel attention weighting ; Subsequently, for perform spatial attention weighting, that is, for the input respectively pass through a convolutional layer, a batch normalization layer, a ReLu activation layer, a convolutional layer, a batch normalization layer, and a sigmoid activation layer to obtain spatial attention weights, and perform weighted summation with the input to obtain the final ; The formula is expressed as: where is the feature extracted from the previous layer, that is, the feature extracted by the CBS layer; Perform spatial dimension feature compression on the operation features; is a fully connected layer; is the activation function; is the feature input to the GAM layer; is the activation function; The classification head CH is composed of a convolutional layer, a pooling layer, a Dropout layer, and a fully connected layer, and the formula is expressed as: where represents the Dropout operation, represents the global average pooling operation.

[0009] To further implement the present invention, the formula adopted in the self-distillation training in step five is: where represents the L2 distance is the global average pooling operation; represents the intermediate layer features, and in this work, the processed features extracted by the fourth-layer CBS layer are used as ; represents the function; Features extracted from the final layer of the model; is a hyperparameter used to balance and the weight ratio for the final loss, with a default value of 0.5; is to calculate the focal loss for the final logit of the model; represents the features after GAP, FC, and softmax processing; and represent the soft loss function and the hard loss function respectively; Among them, represents the improved cross-entropy in the focal loss; is the output probability score, and the larger it is, the closer it is to the true class; represents taking the logarithm; is a modulation factor used to adjust the imbalance in the number of easy / hard-to-distinguish samples; is a modulation factor used to adjust the imbalance in the number of positive and negative samples.

[0010] To further implement the present invention, the voting strategy described in step six adopts the Voting algorithm. A total of 5 models participate in the voting, and the prediction result of each model for the sample x is (i = 1, 2, 3, 4, 5), and the final prediction result C can be expressed as: Among them, δ is an indicator function used to take the mode, represents the predicted class.

[0011] The beneficial effects of the present invention compared with the prior art are as follows: The present invention aims to quantitatively analyze the experimental data results of mouse CT, HE, Masson, TEM, and IHC by applying a convolutional neural network (CNN). Through methods such as data preprocessing and augmentation, model architecture design, model training, model evaluation, model optimization, and evaluation and analysis of experimental results, a powerful multi-modal learning framework is established, that is, a multi-index picture data analysis model. By quantitatively analyzing the trends between groups, and by showing the correlations between various indicators in the pictures, that is, overall analyzing the correlations between various indicators, for analyzing the treatment response of lung cancer combined with idiopathic pulmonary fibrosis, and evaluating the effectiveness of integrated traditional Chinese and Western medicine treatment.

[0012] The present invention first constructs a new model structure, uses the GAM module, and proposes the FFM feature fusion module. Secondly, the model training adopts the self-knowledge distillation training strategy. Considering the uneven sample distribution (long-tail data) and difficult samples, the Focal loss is used to replace the cross-entropy loss, and a new loss function is designed. During the self-distillation training process, the feature map output by the layer before the classification head in the model is used to guide the feature learning of the lower layer of the model, further helping the model learn discriminative and potential features. Finally, during the inference process, the voting strategy is used, and multi-modal data is used as input at the same time, and the final category is determined according to the final weighted score.

[0013] The present invention adopts a lightweight model structure design, uses the bottleneck design idea to avoid increasing the number of additional parameters; during the training process, the self-distillation strategy is adopted. Different from most models that only use the true value label as the learning target, the present invention can better help the model learn potential information and improve the model accuracy. In addition, the multi-stream network voting strategy is adopted to avoid misjudgment of single-modal data and improve the model accuracy and robustness. Brief Description of the Drawings

[0014] Figure 1 It is the structural composition of the network model in the present invention; Figure 2 It is the flow chart of the self-distillation training method in the present invention; Figure 3 It is the CT result diagram in the experimental example of the present invention. Among them, A is the PR diagram, B is the ROC diagram. In diagrams A and B, Group 1 represents Nacl, Group 2 represents NSCLC+IPF, Group 3 represents PE, Group 4 represents DDP+PFD, and Group 5 represents DDP+PFD+PE; Figure 4 It is the TEM result diagram in the experimental example of the present invention. Among them, A is the PR diagram, B is the ROC diagram. In diagrams A and B, Group 1 represents Nacl, Group 2 represents NSCLC+IPF, Group 3 represents PE, Group 4 represents DDP+PFD, and Group 5 represents DDP+PFD+PE; Figure 5 It is the HE result diagram in the experimental example of the present invention. Among them, A is the PR diagram, B is the ROC diagram. In diagrams A and B, Group 1 represents Nacl, Group 2 represents NSCLC+IPF, Group 3 represents PE, Group 4 represents DDP+PFD, and Group 5 represents DDP+PFD+PE; Figure 6This is the IHC result graph in the experimental examples of the present invention. Among them, A is the PR graph, B is the ROC graph. In graphs A and B, Group1 represents Nacl, Group2 represents NSCLC+IPF, Group3 represents PE, Group4 represents DDP+PFD, and Group5 represents DDP+PFD+PE; Figure 7 This is the MASSON result graph in the experimental examples of the present invention. Among them, A is the PR graph, B is the ROC graph. In graphs A and B, Group1 represents Nacl, Group2 represents NSCLC+IPF, Group3 represents PE, Group4 represents DDP+PFD, and Group5 represents DDP+PFD+PE; Figure 8 This is the localization analysis of each group in the experimental examples of the present invention for indicators such as CT, TEM, and HE. Detailed implementation manners

[0015] The present invention will be further described below in conjunction with the accompanying drawings and detailed implementation manners.

[0016] 1. A multi-modal mouse experimental data trend prediction and analysis method based on deep learning, which is characterized by including the following steps: Step 1: Obtain the pathological section images to be analyzed; Step 2: Offline data augmentation: Preprocess the pathological section images in the offline state. The preprocessing includes random vertical flipping, random rotation, random scaling, random brightness adjustment, random hue and saturation adjustment, random contrast adjustment, and random Gaussian blur. Among them, the random vertical flipping is performed with a probability of 50%, the random rotation is controlled within the range of -15° to 15° and rotates around the center point, the random scaling is controlled within the range of 0.75 to 1.25 times, the random brightness adjustment is controlled within the range of -40 to 40, the random hue and saturation adjustment is controlled within the range of -20 to 20, the random contrast is controlled within the range of 0.8 to 1.2, and the random Gaussian blur is controlled within the range of 0 to 1; Step 3: Construct a network model: As Figure 1 shown, the network model is constructed by stacking several CBS modules, FFM modules, and GAM modules. A classification head CH is set at the end of the model for classification; The CBS module is composed of a convolutional layer, a batch normalization layer, and a SiLU activation function, and the formula expression is: Among them, is the input feature of the CBS module; SiLU is the SiLU activation function, BN is the batch normalization layer, and conv is the convolutional layer; The FFM module performs feature fusion at different levels to improve the accuracy and discrimination ability of the model. First, the input features are passed through a CBS module for feature extraction to obtain features , and then is split into channels as and , and is sent to the multi-layer CBS-Block module for feature extraction. The features extracted by each layer of the CBS-Block module will be combined with a skip connection and and after feature extraction through the multi-layer CBS-Block module to restore the original number of channels, thus realizing the fusion between different-level features. The formula is expressed as: Among them, is the feature input to the FFM module; represents splitting the extracted features by channel into and ; is the feature after fusing different-level features through residual connection; The operation represents concatenating features of different levels according to the channel dimension; The GAM module is an attention module that performs attention weighting on features at the channel and spatial levels respectively, helping the model to better focus on the regions and channels that are beneficial for the model to make correct classifications. For the input feature , a Reshape operation is performed to achieve spatial feature compression, and then channel attention weight learning is carried out through a fully connected layer, a ReLU activation layer, and a fully connected layer respectively. Finally, it is adjusted back to the original dimension of the input feature through Reshape and weighted with the input feature to obtain the feature with channel attention weighting; Subsequently, spatial attention weighting is performed on , that is, for the input , spatial attention weights are obtained through a convolutional layer, a batch normalization layer, a ReLu activation layer, a convolutional layer, a batch normalization layer, and a sigmoid activation layer respectively, and weighted summation is performed with the input to obtain the final ; The formula is expressed as: Among them, is the feature extracted in the previous layer, that is, the feature extracted by the CBS layer; The operation compresses the feature in the spatial dimension; is the fully connected layer; is Activation function; is the feature input to the GAM layer; is Activation function; The classification head CH is composed of a convolutional layer, a pooling layer, a Dropout layer, and a fully connected layer, and the formula is expressed as: where represents the Dropout operation, represents the global average pooling operation; Step 4. Online data augmentation: Send the pictures into the network model. During the data loading phase, online data augmentation is performed on each picture. The online data augmentation includes random vertical flipping, random rotation, random scaling, random brightness adjustment, random hue and saturation adjustment, and random contrast adjustment. Among them, the random vertical flipping is performed with a probability of 50%, the random rotation is controlled to rotate around the center point within the range of -15° to 15°, the random scaling is controlled within the range of 0.75 - 1.25 times, the random brightness adjustment is controlled within the range of -40 to 40, the random hue and saturation adjustment is controlled within the range of -20 to 20, the random contrast is controlled within the range of 0.8 to 1.2, and randomly select pictures from the obtained pictures for mosaic mode splicing. The obtained large pictures enter the network model for training; Step 5. Self-distillation training: As Figure 2 shown, perform self-distillation training on the network model, determine the loss function, and obtain multiple models of different modalities; The formula used for the self-distillation training is: where represents the L2 distance is the global average pooling operation; represents the intermediate layer feature. In this work, the processed feature extracted by the fourth CBS layer is used as ; represents function; is the feature extracted by the final layer of the model; is a hyperparameter used to balance and The weight ratio of the final loss, with a default of 0.5; is to calculate the focal loss for the final logit of the model; represents the feature after GAP, FC, and softmax processing; and respectively represent the soft loss function and the hard loss function; Among them, represents the improved cross-entropy in focal loss; is the output probability score, the larger it is, the closer it is to the true category; represents taking the logarithm; is a modulation factor used to adjust the imbalance in the number of easy / hard-to-distinguish samples; is a modulation factor used to adjust the imbalance in the number of positive and negative samples; Step 6. Voting strategy: The model uses the voting strategy to perform inference on different modality data one by one to obtain the results. The voting strategy uses the Voting algorithm. A total of 5 models participate in the voting. The prediction result of each model for the sample x is (i = 1, 2, 3, 4, 5), and the final prediction result C can be expressed as: Among them, δ is an indicator function used to take the mode, represents the predicted category.

[0017] Experimental example: Experimental settings: All experiments are run on an NVIDIA GeForce RTX 4090 graphics card using the PyTorch framework. Stochastic Gradient Descent (SGD) with Nesterov momentum (0.9) is used as the optimizer. The weight decay is set to 0.0001, the initial learning rate is 0.1, the Batch size is set to 64, and the Warmup training strategy is used in the initial stage of training.

[0018] Experimental results: 1. Detection model accuracy: The data of CT, TEM, HE, IHC, and MASSON indicators of each group are evaluated. From the results in Table 1 and Figures 3 - 7 as shown, in the case of uneven quantity and quality of the experimental data collected in each group, the average value of this model is still greater than 8, indicating that the model has good adaptability and high accuracy.

[0019] 2. Data analysis and trend prediction As shown in the results of Table 2 and Figure 8As shown, compared with the Nacl group, the degrees of significant pathological features, imaging and other indicators of lung cancer complicated with idiopathic pulmonary fibrosis in the NSCLC+IPF group were significantly increased; compared with the NSCLC+IPF group, the PE group, the DDP+PFD group and the DDP+PFD+PE group all showed varying degrees of reduction in the significant pathological features, imaging and other indicators of lung cancer complicated with idiopathic pulmonary fibrosis.

[0020] Note: The parameter w is set by the experience of pathological experts: Nacl group, 0 points; NSCLC+IPF group, 10 points; PE group, 8 points; DDP+PFD group, 6 points; DDP+PFD+PE group, 4 points. The trend score is the score of each group; it is calculated by the weighted sum of the predicted probabilities of each input picture, and the formula can be expressed as: This formula is used for calculation for both the above-mentioned Indictor and Group, where , , , , are the probabilities of each category output after sending the picture into the network model.

[0021] The full names of the English abbreviations mentioned in the text are as follows: Nacl is sham operation; NSCLC+IPF is lung cancer complicated with idiopathic pulmonary fibrosis; PE is perilla; DDP+PFD is cisplatin combined with pirfenidone; DDP+PFD+PE is perilla combined with cisplatin combined with pirfenidone.

Claims

1. A method for predicting and analyzing trends in multimodal mouse experiment data based on deep learning, characterized in that It includes the following steps: Step 1: Obtain the pathological section images to be analyzed; Step 2: Offline data augmentation: Preprocess the pathological section images in the offline state; Step 3: Construct a network model: The network model is constructed by stacking several CBS modules, FFM modules, and GAM modules. A classification head CH is set at the end of the model for classification; Step 4: Online data augmentation: Send the images into the network model. During the data loading phase, online data augmentation is performed on each image, and images are randomly selected from the obtained images for mosaic mode splicing. The resulting large images enter the network model for training; Step 5: Self-distillation training: Perform self-distillation training on the network model, determine the loss function, and obtain multiple models with different modalities; Step 6: Voting strategy: Use the voting strategy for the models to perform inference on the data corresponding to different modalities to obtain the results.

2. The method for predicting and analyzing the trend of multi-modal mouse experimental data based on deep learning according to claim 1, characterized in that: The preprocessing described in Step 2 includes random vertical flipping, random rotation, random scaling, random brightness adjustment, random hue and saturation adjustment, random contrast adjustment, and random Gaussian blur.

3. The method for predicting and analyzing the trend of multi-modal mouse experimental data based on deep learning according to claim 1, wherein: The online data augmentation described in Step 4 includes random vertical flipping, random rotation, random scaling, random brightness adjustment, random hue and saturation adjustment, and random contrast adjustment.

4. The method for predicting and analyzing the trend of multi-modal mouse experimental data based on deep learning according to claim 2 or 3, characterized in that: The random vertical flipping is performed with a probability of 50%. The random rotation is controlled to rotate around the center point within the range of -15° - 15°. The random scaling is controlled within the range of multiples of 0.75 - 1.

25. The random brightness adjustment is controlled within the range of -40 - 40. The random hue and saturation adjustment is controlled within the range of -20 - 20. The random contrast is controlled within the range of 0.8 - 1.

2. The random Gaussian blur is controlled within the range of 0 - 1.

5. The method for predicting and analyzing the trend of multi-modal mouse experimental data based on deep learning according to claim 1, characterized in that: The CBS module described in Step 3 is composed of a convolutional layer, a batch normalization layer, and a SiLU activation function. The formula is expressed as: Among them, is the input feature of the CBS module; SiLU is the SiLU activation function, BN is the batch normalization layer, and conv is the convolutional layer; The FFM module performs feature fusion at different levels to improve the model's accuracy and discrimination ability. First, the input features are passed through a CBS module for feature extraction to obtain features , and then is split into channels as and , and is sent to a multi-layer CBS-Block module for feature extraction. The features extracted by each layer of the CBS-Block module will be concatenated with a skip connection and and after feature extraction through the multi-layer CBS-Block module to restore to the original number of channels, thus realizing the fusion between features at different levels. The formula is expressed as: Among them, is the feature input to the FFM module; represents splitting the extracted features by channel into and ; is the feature after fusing features at different levels through residual connection; The operation means concatenating features at different levels according to the channel dimension; The GAM module is an attention module that performs attention weighting on features at the channel and spatial levels respectively, helping the model to better focus on the regions and channels that are beneficial for the model to make correct classifications; for the input features perform a Reshape operation to achieve spatial feature compression, and then respectively pass through a fully connected layer, a ReLU activation layer, and a fully connected layer to learn the channel attention weights. Finally, adjust the dimension to the original dimension of the input feature through Reshape, and add it to the input feature to perform weighting to obtain features with channel attention weighting ; subsequently, perform spatial attention weighting on , that is, for the input respectively pass through a convolutional layer, a batch normalization layer, a ReLu activation layer, a convolutional layer, a batch normalization layer, and a sigmoid activation layer to obtain spatial attention weights, and perform weighted summation with the input to obtain the final ; the formula is expressed as: Among them, is the feature extracted from the upper layer, that is, the feature extracted by the CBS layer; Performs spatial dimension feature compression on the operation features; is the fully connected layer; is the activation function; is the feature input to the GAM layer; is the activation function; The classification head CH is composed of a convolutional layer, a pooling layer, a Dropout layer, and a fully connected layer. The formula is expressed as: Among them, represents the Dropout operation, represents the global average pooling operation.

6. The method for predicting and analyzing the trend of multi-modal mouse experiment data based on deep learning according to claim 1, wherein: The formula used for the self-distillation training described in Step 5 is: Among them, represents the L2 distance is the global average pooling operation; represents the intermediate layer features. In this work, the processed features extracted by the fourth CBS layer are used as ; represents a function; is the feature extracted by the final layer of the model; is a hyperparameter used to balance and The weight ratio for the final loss, with a default of 0.5; is to calculate the focal loss for the final logit of the model; represents the features after being processed by GAP, FC, and softmax; and represent the soft loss function and the hard loss function respectively; Among them, represents the improved cross-entropy in focal loss; is the output probability score, and the larger it is, the closer it is to the true category; represents taking the logarithm; is a regularization factor used to adjust the imbalance in the number of easy / hard-to-distinguish samples; is a regularization factor used to adjust the imbalance in the number of positive and negative samples.

7. The method for predicting and analyzing the trend of multi-modal mouse experimental data based on deep learning according to claim 1, wherein: The voting strategy described in Step 6 adopts the Voting algorithm. A total of 5 models participate in the voting, and the prediction result of each model for the sample x is (i = 1, 2, 3, 4, 5), and the final prediction result C can be expressed as: Among them, δ is an indicator function used to take the mode, representing the predicted category.