Multi-hop contrast learning node classification method based on feature enhancement

By employing feature mapping, spectral feature enhancement, and multi-hop contrast loss, the problems of feature space interference and graph structure destruction caused by noise addition in existing technologies are solved, thereby improving the accuracy of node classification.

CN120953698APending Publication Date: 2025-11-14GUANGXI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511102775.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing graph contrastive learning methods are prone to excessive interference when adding noise to the feature space, which can damage the graph structure and affect the performance of downstream tasks. Meanwhile, structure enhancement methods may significantly damage the graph structure and reduce the quality of graph embeddings.

Method used

Linear feature mapping is used to reduce dimensionality. Singular value decomposition and incomplete power iteration methods are used to add perturbations at the singular value level to construct an enhanced view of multi-hop neighbor information fusion. The similarity of node representations is calculated by multi-hop contrastive loss, and the classification error is calculated using standard cross-entropy to construct the final training target.

Benefits of technology

Without altering the graph topology, this method enhances the robustness and expressiveness of node features, improves the accuracy of node classification, and outperforms existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953698A_ABST
    Figure CN120953698A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of graph neural networks, in particular to a multi-hop contrast learning node classification method based on feature enhancement, which consists of four key parts, namely feature mapping, spectral feature enhancement, multi-hop view generation and multi-hop contrast loss, and comprises the following steps of: firstly, generating a view with relatively low-dimensional node representation through feature mapping; spectral feature enhancement is then applied to these low-dimensional views, which implicitly adds noise to singular values and rebalances them, resulting in enhanced views. Next, the multi-hop information is aggregated using the adjacency matrix to create a plurality of augmented views. The multi-hop contrast loss is used to improve similarity between positive pairs and reduce similarity between negative pairs in all multi-hop views. Finally, the learned node representation can be applied to downstream tasks, such as node classification. Experimental results show that the performance of the method in a node classification task is superior to that of an existing graph comparison learning method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of graph neural network technology, and specifically to a multi-hop contrastive learning node classification method based on feature enhancement. Background Technology

[0002] Graph neural networks have become a core method for modeling graph-structured data in recent years. Their basic idea is to aggregate neighbor information of nodes through a message-passing mechanism, thereby learning the node representation. Graph contrastive learning is a self-supervised learning method that does not rely on manual labels but is trained by constructing pairs of positive and negative samples.

[0003] Graph contrastive learning techniques that rely on data augmentation can generally be divided into two categories: feature augmentation methods and structure augmentation methods. Feature augmentation methods generate different perspectives by altering the features of nodes in the graph. These modifications can include adding noise to the node feature matrix, masking features, or transforming features in some way. However, directly adding noise to the feature space can result in too many interruptions, which may negatively impact the performance of downstream tasks. Structure augmentation methods create different graph views by modifying the graph's topology while keeping the node features unchanged. These methods generate new graph representations by deleting or adding nodes, perturbing edges, or sampling subgraphs. Most existing structure augmentation techniques generate contrastive views by randomly altering the graph's topology. However, deleting edges or nodes can significantly disrupt the graph's structure, thereby reducing the quality of the graph embedding. Summary of the Invention

[0004] The purpose of this invention is to provide a multi-hop contrastive learning node classification method based on feature enhancement, which injects noise into node features to enhance features without causing excessive interference to the feature space, and integrates multi-hop neighbor information into contrastive learning without changing the graph topology.

[0005] To achieve the above objectives, this invention provides a multi-hop contrastive learning node classification method based on feature enhancement, comprising the following steps:

[0006] Step 1: Reduce the dimensionality of the input features using linear feature mapping;

[0007] Step 2: Perform singular value decomposition on the graphic feature map;

[0008] Step 3: Rebalance the singular values ​​using an incomplete exponential iteration method;

[0009] Step 4: Construct a low-rank perturbation term based on the power iteration result and subtract it from the original feature. Add perturbation at the singular value level to obtain the enhanced feature map.

[0010] Step 5: Normalize the adjacency matrix and, without changing the graph structure, fuse neighbor information to introduce multi-hop structure awareness.

[0011] Step 6: Merge the neighbor information of the current hop and the representation of the previous hop to gradually build a multi-hop view;

[0012] Step 7: Organize all generated views for unified comparison loss calculation;

[0013] Step 8: Measure whether the representations of the same node are similar in different views, and calculate the cosine similarity between the two graphs;

[0014] Step 9: Construct a contrastive loss that makes the same node more similar in different views, and more different in the representations of other nodes;

[0015] Step 10: Average the values ​​across all nodes to obtain the overall contrast loss, thus yielding the average contrast loss between the two views;

[0016] Step 11: Randomly select a central view and compare it with the other views to obtain the final multi-view comparison loss;

[0017] Step 12: Calculate the classification error using standard cross-entropy;

[0018] Step 13: Construct the final training objective.

[0019] Optionally, the linear feature mapping process in step 1 is expressed as follows:

[0020]

[0021] in, Represents the input feature matrix. It is a weight matrix. This represents the bias vector. It is the output matrix after linear transformation.

[0022] Optionally, in step 2, let The singular value decomposition of the graph feature map is as follows:

[0023]

[0024] in, , and It is a unitary matrix, where Includes singular values, representing the energy of each latent direction in the feature space, a matrix. It is diagonal, singular value. Sort in descending order.

[0025] Optionally, in step 5, an iterative method is used to gradually merge neighbor information from different hops, using a normalized adjacency matrix to iteratively aggregate neighbor features;

[0026] The neighbor feature matrix is ​​normalized as follows:

[0027]

[0028] in It is the degree matrix of the graph, which contains a small constant. This is to prevent division by zero.

[0029] Optional, multi-hop view in step 6 Generated using the following update rules:

[0030]

[0031] Among them, all Jump views are all based on the initial view. Sequentially generated, trade-off parameters Control the fusion ratio between consecutively enhanced views to balance the importance of adjacency information in feature updates; This represents the exponential linear unit activation function, while express Activation function.

[0032] Optionally, in step 8, a contrastive learning objective, namely a discriminator, is used to distinguish between the embeddings of the same node and the embeddings of different nodes in different views.

[0033] For each node In view Figure 1 The obtained embedding Acting as an anchor, while the embedding from view 2 Represents positive samples, and Embeddings of other related nodes are considered negative samples.

[0034] Optionally, during step 9 of constructing the contrastive loss, the nodes The contrast loss is expressed as:

[0035]

[0036] in,

[0037] +

[0038] In the formula, Indicates an indicator function if and only if When the value is not equal to 1, the indicator function equals 1. It's a temperature parameter. It is used to measure the similarity between node embeddings and is the value obtained by exponentially scaling the similarity of positive sample node embeddings in two views.

[0039] Optionally, in step 12, during the calculation of the classification error using standard cross-entropy, the cross-entropy loss can be expressed as follows:

[0040]

[0041] in, A node representing a hot code encoding format The true label, and Represents a node Belongs to class The predicted probability, This refers to the total number of classes.

[0042] This invention provides a multi-hop contrastive learning node classification method based on feature enhancement, consisting of four key parts: feature mapping, spectral feature enhancement, multi-hop view generation, and multi-hop contrastive loss. First, feature mapping generates views with relatively low-dimensional node representations. Then, spectral feature enhancement is applied to these low-dimensional views, implicitly adding noise to singular values ​​and rebalancing them to produce enhanced views. Next, adjacency matrices are used to aggregate multi-hop information to create multiple augmented views. Multi-hop contrastive loss is used to increase the similarity between positive pairs and decrease the similarity between negative pairs across all multi-hop views. Finally, the learned node representations can be applied to downstream tasks, such as node classification. Experimental results show that this invention outperforms existing graph contrastive learning methods in node classification tasks. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a schematic diagram illustrating the flowchart of a multi-hop contrastive learning node classification method based on feature enhancement according to the present invention. Detailed Implementation

[0045] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0046] This invention provides a multi-hop contrastive learning node classification method based on feature enhancement, comprising the following steps:

[0047] Step 1: Reduce the dimensionality of the input features using linear feature mapping;

[0048] Step 2: Perform singular value decomposition on the graphic feature map;

[0049] Step 3: Rebalance the singular values ​​using an incomplete exponential iteration method;

[0050] Step 4: Construct a low-rank perturbation term based on the power iteration result and subtract it from the original feature. Add perturbation at the singular value level to obtain the enhanced feature map.

[0051] Step 5: Normalize the adjacency matrix and, without changing the graph structure, fuse neighbor information to introduce multi-hop structure awareness.

[0052] Step 6: Merge the neighbor information of the current hop and the representation of the previous hop to gradually build a multi-hop view;

[0053] Step 7: Organize all generated views for unified comparison loss calculation;

[0054] Step 8: Measure whether the representations of the same node are similar in different views, and calculate the cosine similarity between the two graphs;

[0055] Step 9: Construct a contrastive loss that makes the same node more similar in different views, and more different in the representations of other nodes;

[0056] Step 10: Average the values ​​across all nodes to obtain the overall contrast loss, thus yielding the average contrast loss between the two views;

[0057] Step 11: Randomly select a central view and compare it with the other views to obtain the final multi-view comparison loss;

[0058] Step 12: Calculate the classification error using standard cross-entropy;

[0059] Step 13: Construct the final training objective.

[0060] Please see Figure 1 , Figure 1This diagram illustrates the steps of the feature-enhanced multi-hop contrastive learning node classification method described in this invention. The method consists of four key parts: feature mapping, spectral feature enhancement, multi-hop view generation, and multi-hop contrastive loss. Step 1 is the feature mapping method; steps 2 to 4 are the spectral feature enhancement process; steps 5 to 7 are the multi-hop view generation process; and steps 8 onwards are the multi-hop contrastive loss method. The specific execution steps are further explained below:

[0061] Step 1: Use linear feature mapping to reduce the dimensionality of the input features, thereby reducing feature complexity;

[0062] In deep learning, feature mapping plays a crucial role, typically used to transform input features into a transformed feature space. To simplify the processing, this invention employs a feature mapping operation. To reduce the dimension of the input features from Down to This reduces the complexity of the features. The parameters of the mapping are trainable, allowing the model to dynamically transform the original input features into a more concise and informative feature space.

[0063] (1)

[0064] here, Represents the input feature matrix. It is a weight matrix. This represents the bias vector. It is the output matrix after linear transformation. Through optimization... and Feature mapping operations can adjust feature dimensions (e.g., increase or decrease dimensionality) while extracting more discriminative representations. This improves the model's ability to capture data representations and enhances its predictive accuracy. Furthermore, the mapped features provide more effective representations for subsequent network layers, improving model performance on complex tasks.

[0065] Step 2: Perform singular value decomposition on the graphic feature map;

[0066] In spectral feature enhancement methods, this invention injects noise into singular values ​​in a controlled manner. This approach enables the model to capture key patterns in the data from ∑ different perspectives. By carefully managing the noise injection, excessive interference with the feature space is prevented. Therefore, the model improves its ability to learn more robust representations while maintaining its learning performance.

[0067] Specifically, linear transformation The obtained low-dimensional feature matrix Input to feature enhancement function In this function, random noise is added to the singular values ​​through incomplete exponential iteration to enhance spectral features.

[0068] set up The singular value decomposition of the graph feature map is as follows:

[0069] (2)

[0070] in , and It is a unitary matrix, where Includes singular values, representing the energy of each latent direction in the feature space, a matrix. It is diagonal, singular value. Arranged in descending order. The spectral perturbation strategy of this invention injects controlled noise. Get an enhanced view Due to singular vectors and Since it remains unchanged, the orientation structure of the original feature space is preserved.

[0071] From an information perspective, perturbing low- and mid-frequency singular values ​​can enhance the diversity of the feature spectrum and prevent the model from over-relying on the dominant direction. This spectral diversity helps the model discover structural or semantic information that is often overlooked in the original representation, thus generating richer contrast signals between different views. Unlike direct structural perturbations, spectral perturbations preserve the topological relationships between nodes. Therefore, it enhances the expressiveness and discriminative power of the view while maintaining the semantic stability of the graph.

[0072] Step 3: Rebalance the singular values ​​using an incomplete exponential iteration method;

[0073] This invention uses an incomplete exponential iteration method to rebalance singular values. Initially, the vector... From the standard normal distribution Random sampling is used for each iteration. ,calculate:

[0074] (3)

[0075] Step 4: Construct a low-rank perturbation term based on the power iteration result and subtract it from the original feature. Add perturbation at the singular value level to enhance the robustness and discriminative power of the feature and obtain the enhanced feature map.

[0076] This step involves using feature maps The transpose of the current random vector Multiplication to update random vectors

[0077] After completing the iteration, the feature map is enhanced. The calculation formula is as follows:

[0078] (4)

[0079] here, The number of iterations determines the depth of the enhancement process. It is a feature map Low-rank update, i.e. from Subtract from the middle The relevant parts. yes of The square of the norm.

[0080] The method for generating multi-hop views includes the following steps:

[0081] Step 5: Normalize the adjacency matrix and, without changing the graph structure, fuse neighbor information to introduce multi-hop structure awareness.

[0082] This invention uses an iterative method to progressively merge neighbor information from different hops. At each step, the node representation is updated using the result of the previous iteration. This method allows for capturing a broader range of neighbor relationships. Simultaneously, the process ensures that multi-hop neighborhood information is included for contrastive learning while preserving the original graph topology. Neighbor features are iteratively aggregated using a normalized adjacency matrix. By performing adjacency matrix analysis In the next iteration, the features of multi-hop neighbors were obtained. From equation (4), the basic view was obtained. It represents the enhanced feature matrix. To enhance the stability of the aggregation process, the neighbor feature matrix is ​​normalized as follows:

[0083] (5)

[0084] in It is the degree matrix of the graph, which contains a small constant. This is to prevent division by zero.

[0085] Step 6: Merge the neighbor information of the current hop and the representation of the previous hop to gradually build a multi-hop view;

[0086] A new enhanced view is generated by combining the adjacency matrix and information from the previous view. Multi-hop view It is generated iteratively, with each iteration based on the previous view. Generate a new view This iterative process gradually incorporates additional multi-hop neighbor information. Specifically, multi-hop views...

[0087] Generated using the following update rules:

[0088] (6)

[0089] all Jump views are all based on the initial view. Generated sequentially. Trade-off parameters. Control the fusion ratio between consecutively enhanced views to balance the importance of adjacency information in feature updates. The (exponential linear unit) activation function adds non-linearity, enabling the model to capture more complex feature representations, while express Activation function.

[0090] Step 7: Organize all generated views and use them uniformly for comparative loss calculation;

[0091] Finally, all generated enhanced views are stored in a list. In the middle, it is defined as:

[0092] (7)

[0093] It is a collection of node representations, including multiple different views generated through multi-hop information. The learned representations. This information is then used in downstream tasks. This method effectively incorporates richer multi-hop neighbor information while maintaining the integrity of the original graph topology. This enhances the expressive power of the node representations.

[0094] Step 8: Measure whether the representations of the same node are similar in different views, and calculate the cosine similarity between the two graphs;

[0095] In this method, a contrastive learning objective (i.e., a discriminator) is used to distinguish between the embeddings of the same node and the embeddings of different nodes in different views. For each node... In view Figure 1 The obtained embedding Acting as an anchor, while the embedding from view 2 This represents a positive sample. (And...) Embeddings of other related nodes are considered negative samples. Formally, the comparison target is defined as:

[0096] (8)

[0097] in This represents the cosine similarity.

[0098] Step 9: Construct a contrast loss that makes the same node more similar in different views, and more different in the representations of other nodes;

[0099] node The contrast loss can be written as:

[0100] (9)

[0101] In the method of this invention, negative nodes are not directly sampled. Instead, negative samples are defined based on positive sample pairs. Positive samples are represented by the first term in the denominator of the loss function. From both inter-view and intra-view perspectives, negative samples include other nodes. These are respectively... The first and second items were captured.

[0102] + (10)

[0103] here, Indicates an indicator function if and only if When the value is not equal to 1, the indicator function is equal to 1. It is a temperature parameter.

[0104] Step 10: Average the values ​​across all nodes to form the overall contrast loss, thus obtaining the average contrast loss between the two views;

[0105] Considering the symmetry between views, Figure 1 Total contrast loss between view 1 and view 2 Averaged across all nodes:

[0106] (11)

[0107] Step 11: Randomly select a central view and compare it with the other views to obtain the final multi-view comparison loss;

[0108] Within the framework of this invention, two or more views are available. A view is randomly selected. This serves as the anchor. Then, the total multi-hop contrastive loss is calculated using the average anchor view and all pairwise losses between every other view:

[0109] (12)

[0110] This approach effectively uses multiple views to compute the comprehensive contrastive loss. It enhances the model's ability to learn expressive representations while leveraging multi-hop neighbor information.

[0111] Step 12: Calculate the classification error using standard cross-entropy;

[0112] Node classification loss is typically formulated as cross-entropy loss, which quantifies the difference between the actual label and the predicted probability for each class. Cross-entropy loss can be expressed as follows:

[0113] (13)

[0114] in, A node representing a hot code encoding format The true label, and Represents a node Belongs to class The predicted probability. Here, This refers to the total number of classes.

[0115] Step 13: Construct the final training objective.

[0116] Ultimately, the total loss is calculated as a weighted sum of the multi-hop contrast loss and the node classification loss.

[0117] (14)

[0118] here, Used as a hyperparameter to control the trade-off between contrastive loss and classification loss.

[0119] Furthermore, the present invention is further illustrated through specific embodiments:

[0120] Specifically, the model was trained for 2000 epochs and stopped prematurely after 100 epochs without improvement. The number of iterations was set to t = {0, 1, 2, 4, 6}, and in this embodiment, the number of augmented views was varied between 2 and 10. The method of this invention learns embeddings in an unsupervised manner and uses an L2-regularized logistic regression (LR) classifier for semi-supervised node classification. For Cora, CiteSeer, and PubMed, 20 nodes were randomly selected for training, 500 nodes for validation, and the remaining nodes for testing for each class. For Coauthor-CS and Amazon-Photo, 20 nodes were selected for training, 30 nodes for validation, and the remaining nodes for testing for each category. In both cases, only the validation set was used to tune the hyperparameters of the LR classifier. The C parameter of the LR classifier was selected from the set {1e-4, 1e-3, 1e-2, 0.1, 1, 10, 100}. For each dataset, 20 random splits were performed for training, validation, and testing, respectively. Then, the results of these splits are averaged to report the performance of all algorithms.

[0121] Table 1: Node classification accuracy on five datasets

[0122]

[0123] Table 1 shows the node classification accuracy of the proposed method, Spectral Feature Augmentation-based Multi-Hop Contrastive Learning (SMHCL), on five benchmark graph datasets, and compares its performance with other methods. The results clearly demonstrate the advantages of the semi-supervised node classification method proposed in this invention. SMHCL achieves excellent accuracy on various benchmark datasets. For example, on the Cora dataset, SMHCL outperforms MVGRL by 1.0%. Similarly, on the PubMed dataset, SMHCL outperforms MHVGCL by 0.9%. The strong performance of SMHCL stems from its method of implicitly injecting noise to rebalance feature singular values. This method enhances node features without relying on manually designed augmentations for each dataset. In contrast, manual augmentations in the GCL baseline can severely damage the network topology, leading to invalid embeddings. SMHCL avoids this problem by preserving the original topology. Furthermore, SMHCL aggregates multi-hop neighbor information, enabling the model to extract more valuable insights.

[0124] In summary, the multi-hop contrastive learning node classification method based on feature enhancement proposed in this invention reduces the perturbation to the feature space caused by directly adding noise by injecting noise into the singular values ​​of node features. Multi-hop views are then iteratively generated, capturing information from their multi-hop neighbors. Then, contrastive learning is performed using these enhanced views while maintaining the integrity of the graph structure. Finally, a multi-hop contrastive loss is used, randomly selecting a center view and averaging the contrastive loss across multiple views to improve the learning effect. Experimental results show that SMHCL outperforms existing graph contrastive learning methods in node classification tasks.

[0125] The above description discloses only one or more preferred embodiments of the present invention, and should not be construed as limiting the scope of the present invention. Those skilled in the art will understand that all or part of the processes of the above embodiments can be implemented, and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.

Claims

1. A multi-hop contrastive learning node classification method based on feature enhancement, characterized in that, Includes the following steps: Step 1: Reduce the dimensionality of the input features using linear feature mapping; Step 2: Perform singular value decomposition on the graphic feature map; Step 3: Rebalance the singular values ​​using an incomplete exponential iteration method; Step 4: Construct a low-rank perturbation term based on the power iteration result and subtract it from the original feature. Add perturbation at the singular value level to obtain the enhanced feature map. Step 5: Normalize the adjacency matrix and, without changing the graph structure, fuse neighbor information to introduce multi-hop structure awareness. Step 6: Merge the neighbor information of the current hop and the representation of the previous hop to gradually build a multi-hop view; Step 7: Organize all generated views for unified comparison loss calculation; Step 8: Measure whether the representations of the same node are similar in different views, and calculate the cosine similarity between the two graphs; Step 9: Construct a contrastive loss that makes the same node more similar in different views, and more different in the representations of other nodes; Step 10: Average the values ​​across all nodes to obtain the overall contrast loss, thus yielding the average contrast loss between the two views; Step 11: Randomly select a central view and compare it with the other views to obtain the final multi-view comparison loss; Step 12: Calculate the classification error using standard cross-entropy; Step 13: Construct the final training objective.

2. The multi-hop contrastive learning node classification method based on feature enhancement as described in claim 1, characterized in that, The processing procedure for linear feature mapping in step 1 is expressed as follows: ,in, Represents the input feature matrix. It is a weight matrix. This represents the bias vector. It is the output matrix after linear transformation.

3. The multi-hop contrastive learning node classification method based on feature enhancement as described in claim 2, characterized in that, In step 2, let The singular value decomposition of the graph feature map is as follows: ,in, , and It is a unitary matrix, where Includes singular values, representing the energy of each latent direction in the feature space, a matrix. It is diagonal, singular value. Sort in descending order.

4. The multi-hop contrastive learning node classification method based on feature enhancement as described in claim 3, characterized in that, Step 5 uses an iterative method to gradually merge neighbor information from different hops, and uses a normalized adjacency matrix to iteratively aggregate neighbor features; The neighbor feature matrix is ​​normalized as follows: ,in It is the degree matrix of the graph, which contains a small constant. This is to prevent division by zero.

5. The multi-hop contrastive learning node classification method based on feature enhancement as described in claim 4, characterized in that, Step 6 Multi-hop View Generated using the following update rules: Among them, Jump views are all based on the initial view. Sequentially generated, trade-off parameters Control the fusion ratio between consecutively enhanced views to balance the importance of adjacency information in feature updates; This represents the exponential linear unit activation function, while express Activation function.

6. The multi-hop contrastive learning node classification method based on feature enhancement as described in claim 5, characterized in that, Step 8 uses a contrastive learning objective, namely a discriminator, to distinguish between the embeddings of the same node and the embeddings of different nodes in different views. For each node The embedding obtained in view 1 Acting as an anchor, while the embedding from view 2 Represents positive samples, and Embeddings of other related nodes are considered negative samples.

7. The multi-hop contrastive learning node classification method based on feature enhancement as described in claim 6, characterized in that, In step 9, during the construction of the contrastive loss, the nodes The contrast loss is expressed as: , in, + In the formula, Indicates an indicator function if and only if When the value is not equal to 1, the indicator function equals 1. It's a temperature parameter. It is used to measure the similarity between node embeddings and is the value obtained by exponentially scaling the similarity of positive sample node embeddings in two views.

8. The multi-hop contrastive learning node classification method based on feature enhancement as described in claim 7, characterized in that, In step 12, during the calculation of the classification error using standard cross-entropy, the cross-entropy loss can be expressed as follows: ,in, A node representing a hot code encoding format The true label, and Represents a node Belongs to class The predicted probability, This refers to the total number of classes.