Bearing fault diagnosis method based on space-time convolutional network
Through the space-time convolutional network combined with multi-head attention mechanism and SoftMax function, the problem of insufficient feature extraction in bearing fault diagnosis is solved, and high-accuracy fault recognition and prediction are achieved.
Patent Information
- Application Number
- CN202510483951.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional deep learning algorithms rely on early data processing in bearing fault diagnosis, making it difficult to effectively mine the dependence and overall relationship between data before and after, resulting in low diagnostic accuracy under multiple operating conditions.
The method based on spatiotemporal convolution network is adopted to extract the details and overall characteristics of the vibration signal through time convolution blocks and spatial convolution blocks, and combine the multi-head attention mechanism and SoftMax function for fault diagnosis, and use expanded convolution and residual connection to solve the training difficulties of deep networks.
It significantly improves the accuracy of bearing fault diagnosis, can identify faults in time under complex working conditions, extend the service life of bearings, reduce maintenance costs and ensure production safety.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bearing fault diagnosis, and more particularly to a bearing fault diagnosis method based on a spatio-temporal convolutional network. Background Art
[0002] As one of the crucial components in a rotating machinery system, rolling bearings are widely used in various industrial mechanical devices, including but not limited to high-speed trains, machine tools, and aeroengines. Due to their complex operating conditions and frequent exposure to harsh external environments, rolling bearings are extremely prone to failures. Such failures can not only lead to the shutdown of mechanical equipment but also cause significant economic losses and even endanger personnel safety. Therefore, the Predictive Health Management (PHM) technology based on fault prediction has received extensive attention. Rotating machinery, as one of the core components in the modern industrial system, its efficient and stable operating state is directly related to the reliability, safety, and economic benefits of the entire production system. As an indispensable transmission and support component inside rotating machinery, rolling bearings play a crucial role in transmitting loads, supporting the rotation of rotors, and ensuring the smooth operation of machinery. They not only bear huge mechanical stresses but also need to continuously operate in harsh working environments, such as high temperature, high speed, heavy load, and poor lubrication conditions, which greatly test the performance and lifespan of the bearings. According to statistics, approximately 40% to 50% of the rotating machinery failures are directly or indirectly attributed to bearing failures. Bearing failures not only lead to a decline in mechanical performance but may also cause unplanned shutdowns in severe cases, seriously affecting the production schedule and resulting in huge economic losses. In addition, frequent bearing failures may also pose safety hazards to the production environment and threaten the lives of operators. In the context of intelligent manufacturing and Industry 4.0, achieving early and accurate diagnosis of bearing failures through advanced technical means has become the key to improving the operating efficiency of rotating machinery, preventing unplanned shutdowns, reducing maintenance costs, and ensuring production safety. The realization of this goal not only depends on continuous and accurate monitoring of the bearing operating state but also requires the use of advanced technologies such as big data analysis, artificial intelligence, and machine learning to deeply mine and analyze the monitoring data for early warning and precise identification of bearing failures.
[0003] Deep learning has made great progress in the field of bearing fault diagnosis. However, traditional deep learning algorithms have certain limitations. Traditional deep learning highly relies on pre-data processing. It is necessary to first extract the time-domain or frequency-domain data of the original data so that the deep learning algorithm can well process these features. They do not deeply mine the dependencies between data before and after and the relationship between the whole, resulting in the algorithm not performing well in multi-condition data. Summary of the Invention
[0004] The object of the present invention is to provide a bearing fault diagnosis method with both robustness and accuracy. The technical problem to be solved is to provide a bearing fault diagnosis method based on deep learning, which can extract the temporal and spatial features of vibration signals, and can greatly improve the accuracy of fault diagnosis. Thus, the fault state of the bearing can be predicted in time, the components can be processed and maintained in time before a fault occurs, accidents can be avoided, and the service life of the bearing can be extended.
[0005] To solve the above problems, the technical solution adopted by the present invention is:
[0006] A bearing fault diagnosis method based on a spatio-temporal convolutional network, characterized in that:
[0007] S1. The original data is first normalized; then it is input into the spatio-temporal convolutional network, which contains two branch structures: a temporal convolutional block and a spatial convolutional block.
[0008] S2. The temporal convolutional block uses its large receptive field to extract the forward and backward dependencies of the entire data and sense the detailed features of the data.
[0009] S3. The spatial convolutional block maps the one-dimensional data to a two-dimensional space through polar coordinate transformation to form an image; using the powerful image processing ability of the convolutional neural network to sense the overall features of the data.
[0010] S4. Subsequently, the outputs of the two branches will pass through a multi-head attention mechanism to enable the model to pay more attention to important features; finally, the final output is obtained through a SoftMax layer.
[0011] A further improvement lies in that: S1 specifically includes the following steps:
[0012] S11. Obtain the original degradation data collected by the bearing vibration signal sensor and perform data preprocessing, including removing abnormal points and filling missing points in the data;
[0013] S12. Normalize the original data and scale the data to the range of [0, 1]; normalization can make data with different dimensions or orders of magnitude comparable, thereby improving the training efficiency and prediction accuracy of the model; the formula is as follows:
[0014]
[0015] where X max is the maximum value in the data, and X min is the minimum value in the data.
[0016] A further improvement lies in that: S2 specifically includes the following steps:
[0017] S21. Part of the spatio-temporal convolutional network is the temporal convolutional block. The temporal convolutional block focuses on the detailed features of the data. Inspired by the temporal convolutional network, it uses causal dilated convolution in the temporal convolutional network to capture the detailed features in the fault data with a very large receptive field. Therefore, the subtle differences between different fault data can be compared.
[0018] S22. Dilated convolution allows the convolutional kernel to sample at a fixed interval in the input data, thereby significantly expanding the effective coverage of the convolutional kernel without increasing the computational complexity. This mechanism not only improves the network's ability to model long-term dependencies but also avoids the problem of difficult training caused by the network being too deep. The formula for dilated causal convolution is as follows:
[0019]
[0020] where * represents the convolution operator, d is the dilation factor, k is the filter size, g(i) represents the i-th element in the filter, and u s-d · i represents the (s - d·i)-th element in the input sequence.
[0021] S23. Although dilated causal convolution is introduced, the model may still become very deep in some cases. Although the deep network structure can capture complex features and long-term dependencies, it is also prone to problems such as vanishing gradients or exploding gradients, which affect the training efficiency and performance of the model. To solve this problem, residual connections are adopted.
[0022] S24. Assuming that the input of a neural network layer (or a group of layers) is x and the output after passing through this layer is F(x), then the output y of the residual connection can be expressed as:
[0023] y = F(x) + x
[0024] The further improvement is that: S3 specifically includes the following steps:
[0025] S31. Another part of the spatio-temporal convolutional network is the spatial convolutional block. The spatial convolutional block focuses on the overall features of the data. By converting one-dimensional data into two-dimensional images, the overall trend of the fault data can be accurately grasped.
[0026] S32. After receiving the time series tensor x ∈ R B×L×C after the normalization process in S12, where B represents the batch size, L represents the sequence length, and C represents the number of channels; then the normalized sequence values are mapped to the angular space using the arccosine function, and the polar coordinate angle φ of each value is calculated. The angle calculation formula is as follows:
[0027] φ = arccos(clamp(x norm , -1.0, 1.0))
[0028] Here, the clamp function ensures that the input value is within the range of [-1, 1] to avoid out-of-range values caused by numerical errors; by calculating the angle, the one-dimensional sequence values are converted into points in the angular space.
[0029] S33. Calculate the sum (or difference) of two angles and apply the cosine function to measure the relationship between them; each element of the two-dimensional angle matrix represents the angular relationship between two time steps, and its calculation formula is:
[0030]
[0031] where φ i and φ j are the angular values of different time steps in the sequence respectively. Through the broadcasting mechanism, all pairwise combinations between time steps can be calculated efficiently, forming a two-dimensional angle matrix of size (B, C, L, L).
[0032] S34. By converting the one-dimensional time series data into a two-dimensional image, the convolutional neural network can capture the global patterns and trends of the time series in the two-dimensional space; by combining spatial convolutional blocks, the representation ability of the model for complex time series data can be effectively improved.
[0033] The further improvement lies in: S4 specifically includes the following steps:
[0034] S41. Next, the output of the spatio-temporal convolutional network is fed into the multi-head attention mechanism through a fully connected layer. The multi-head attention mechanism is one of the core modules in the Transformer model, which significantly enhances the model's ability to understand and process complex data; the multi-head attention mechanism pays attention to the input data from different angles and subspaces by processing multiple attention heads in parallel, so as to capture more comprehensive and rich feature information; the calculation process of the multi-head attention mechanism can be summarized as the following steps.
[0035] S42. Step one: Perform a linear transformation on the input tensor X to obtain three tensors: query Q, key K, and value V; this step is usually implemented through three different fully connected layers, namely:
[0036] Q = xW q , K = xW k , V = xW v
[0037] where W q , W k and W v are the weight matrices obtained through training.
[0038] S43. Step 2: Split the query Q, key K, and value V into h heads (i.e., h subspaces), each head having different linear transformation parameters; for each head i, calculate its corresponding scaled dot-product attention output respectively, and the formula is as follows:
[0039] head i = Attention(QW i Q , KW i K , VW i V )
[0040] where W i Q , W i K and W i V are the linear transformation matrices of the query, key, and value of head i respectively; the calculation formula of the scaled dot-product attention is:
[0041]
[0042] where d k is the dimension of the key vector, used to scale the dot product to stabilize the SoftMax function.
[0043] S44. Step 3: Perform a final linear transformation on the concatenated vector to integrate the information from different heads and obtain the final multi-head attention output; this step is usually implemented through a fully connected layer, that is:
[0044] MultiHead(Q, K, V) = Concat(head1, head2,..., head h )W O
[0045] After the input data is carefully processed through the multi-head attention mechanism, the model can fully capture various semantic associations and feature information in the data.
[0046] S45. Subsequently, pass these information-rich outputs to the SoftMax function. The SoftMax function, with its unique normalization ability, converts these outputs into a probability distribution; for an input vector z = [z1, z2,..., z n , the formula of the SoftMax function is as follows:
[0047]
[0048] Finally, the classification result is obtained, and the output is a probability distribution, and the sum of all elements is 1. Description of the Drawings
[0049] Figure 1 is the overall framework diagram of the prediction method of the present invention;
[0050] Figure 2 is the schematic diagram of dilated convolution in the present invention;
[0051] Figure 3 is the residual connection structure diagram in the present invention;
[0052] Figure 4 is the schematic diagram of the spatial convolution block in the present invention; Detailed implementation manners
[0053] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments:
[0054] As Figure 1 shown, a bearing fault diagnosis method based on a spatio-temporal convolutional network includes the following steps:
[0055] S1. Obtain the original data collected by the bearing vibration signal sensor, and perform data preprocessing to obtain complete bearing condition monitoring data; specifically, it includes the following steps;
[0056] S11. After collecting the original data collected by the bearing vibration signal sensor, it is crucial to perform systematic preprocessing on the original data; first, it is necessary to detect outliers in the data and remove all detected outliers; this step can ensure the accuracy and reliability of subsequent analysis; next, for the missing values in the data, appropriate methods should be used to fill them to ensure the integrity and continuity of the data; through these preprocessing steps, more reliable and consistent bearing vibration signal data can be obtained, providing a solid foundation for subsequent analysis and research.
[0057] S12. Perform normalization processing on the original data. The normalization processing maps the original data to the interval [0, 1] according to certain rules through a series of mathematical transformations; it first calculates the minimum and maximum values of each feature, and then linearly transforms the original data to the target interval using these two extreme values; during this process, each data point is assigned a new value within the interval [0, 1] according to its relative position in the original data; this transformation not only retains the relative magnitude relationship of the original data, but also makes all features comparable numerically; the formula is as follows:
[0058]
[0059] where X max is the maximum value in the data, and X min is the minimum value in the data.
[0060] S2. Part of the spatio-temporal convolutional network is the temporal convolutional block. The temporal convolutional block focuses on the detailed features of the data. Inspired by the temporal convolutional network, it uses causal dilated convolution in the temporal convolutional network to capture the detailed features in the fault data with a very large receptive field. Therefore, the subtle differences between different fault data can be compared.
[0061] S21. As Figure 2 shown, dilated convolution allows the convolutional kernel to sample at a fixed interval in the input data, thereby significantly expanding the effective coverage of the convolutional kernel without increasing the computational complexity. This mechanism not only improves the network's ability to model long-term dependencies but also avoids the training difficulty problems caused by the network being too deep. The formula for dilated causal convolution is as follows:
[0062]
[0063] where * represents the convolution operator, d is the dilation factor, k is the filter size, g(i) represents the i-th element in the filter, and u s-d · i represents the (s - d·i)-th element in the input sequence.
[0064] S22. Although dilated causal convolution is introduced, the model may still become very deep in some cases. Although the deep network structure can capture complex features and long-term dependencies, it is also prone to problems such as vanishing gradients or exploding gradients, thus affecting the training efficiency and performance of the model. To solve this problem, residual connections are adopted.
[0065] S23. The schematic diagram of the residual connection is as Figure 3 shown. Suppose the input of a neural network layer (or a group of layers) is x, and the output after passing through this layer is F(x). Then the output y of the residual connection can be expressed as:
[0066] y = F(x) + x
[0067] S3. Another part of the spatio-temporal convolutional network is the spatial convolutional block. As Figure 4 shown, the spatial convolutional block focuses on the overall features of the data. By converting one-dimensional data into two-dimensional images, the overall trend of the fault data can be accurately grasped.
[0068] S31. After receiving the normalized time series tensor x ∈ R B×L×C , where B represents the batch size, L represents the sequence length, and C represents the number of channels. Subsequently, the normalized sequence values are mapped to the angular space using the arccosine function, and the polar coordinate angle φ of each value is calculated. The angle calculation formula is as follows:
[0069] φ = arccos(clamp(x norm, -1.0, 1.0))
[0070] Here, the clamp function ensures that the input value is within the range of [-1, 1] to avoid out-of-range values caused by numerical errors; by calculating the angle, the one-dimensional sequence values are converted into points in the angular space.
[0071] S32. Calculate the sum (or difference) of two angles and apply the cosine function to measure the relationship between them; each element of the two-dimensional angle matrix represents the angular relationship between two time steps, and its calculation formula is:
[0072]
[0073] where, φ i and φ j are the angular values of different time steps in the sequence respectively. Through the broadcasting mechanism, all pairwise combinations between time steps can be calculated efficiently, forming a two-dimensional angle matrix of size (B, C, L, L).
[0074] S33. By converting one-dimensional time series data into two-dimensional images, the convolutional neural network can capture the global patterns and trends of time series in the two-dimensional space; by combining spatial convolutional blocks, the representation ability of the model for complex time series data can be effectively improved.
[0075] S4. Next, the output of the spatio-temporal convolutional network is fed into the multi-head attention mechanism through a fully connected layer. The multi-head attention mechanism is one of the core modules in the Transformer model, which significantly enhances the model's ability to understand and process complex data; the multi-head attention mechanism pays attention to the input data from different angles and subspaces by processing multiple attention heads in parallel, so as to capture more comprehensive and rich feature information; the calculation process of the multi-head attention mechanism can be summarized into the following steps.
[0076] S41. Step one: Perform a linear transformation on the input tensor X to obtain three tensors: query Q, key K, and value V; this step is usually implemented through three different fully connected layers, namely:
[0077] Q = xw q , K = xw k , V = xw v
[0078] where, W q , W k and W v are the weight matrices obtained through training.
[0079] S42. Step 2: Split the query Q, key K, and value V into h heads (i.e., h subspaces), each head having different linear transformation parameters; for each head i, calculate its corresponding scaled dot-product attention output respectively, and the formula is as follows:
[0080] head i = Attention(QW i Q , KW i K , VW i V )
[0081] where W i Q , W i K , and W i V are the linear transformation matrices of the query, key, and value of head i respectively; the calculation formula of the scaled dot-product attention is:
[0082]
[0083] where d k is the dimension of the key vector, which is used to scale the dot product to stabilize the SoftMax function.
[0084] S43. Step 3: Perform a final linear transformation on the concatenated vector to integrate the information from different heads and obtain the final multi-head attention output; this step is usually implemented through a fully connected layer, that is:
[0085] MultiHead(Q, K, V) = Concat(head1, head2,..., head h )W O
[0086] After the input data is carefully processed through the multi-head attention mechanism, the model can fully capture various semantic associations and feature information in the data.
[0087] S44. Subsequently, pass these information-rich outputs to the SoftMax function, and the SoftMax function, with its unique normalization ability, converts these outputs into a probability distribution; for an input vector z = [z1, z2,..., z n , the formula of the SoftMax function is as follows:
[0088]
[0089] Finally, the classification result is obtained, and the output is a probability distribution, and the sum of all elements is 1.
[0090] Embodiment
[0091] A bearing fault diagnosis method based on spatio-temporal convolutional network. In this study, the HUST bearing fault benchmark dataset released by Huazhong University of Science and Technology was used for algorithm verification. This dataset was collected through a Spectra-Quest mechanical fault simulation test bench, which includes a speed control module, a drive motor, a transmission shaft, an acceleration sensor array, a bearing test unit, and a data acquisition system. The dataset covers 8 artificially preset bearing fault modes, including two damage levels of medium / severe and 4 operating conditions. Among them, medium-fault samples are specifically used to construct early fault recognition scenarios. By selecting vibration signals with typical fault characteristics, the detection performance of the proposed algorithm in the fault germination stage was effectively verified.
[0092] The single sampling duration of the HUST experiment is set to 10.2 seconds, and the sampling rate reaches 25.6 kHz, forming a long sequence of vibration data of about 260,000 points. To adapt to the model input requirements, the sliding window technique is used to frame the data stream. The window width is set to 1024 points, and the sliding step is set to 512 points. This overlapping framing strategy effectively balances the time resolution and computational efficiency, retaining both local time-domain features and ensuring the sufficiency of data utilization. As shown in Table 1, the study selected fault samples under two typical operating conditions to construct a training set and assigned independent labels to different damage modes, laying a foundation for subsequent pattern recognition tasks.
[0093] Table 1
[0094]
[0095] When applying the proposed algorithm to the HUST bearing dataset, the evaluation strategy of taking the average value through 10 independent experiments effectively eliminates the randomness influence of the deep learning model. After sufficient training, the loss function of the model converges to the minimum value, and the classification accuracy stabilizes above 99%. Table 2 details the fault recognition accuracy under each operating condition. The results show that the algorithm reaches or approaches a 100% recognition rate for most fault types, with only a small error in Class 35B. It shows that the proposed algorithm exhibits excellent performance in bearing fault feature learning and classification decision-making.
[0096] Table 2
[0097]
[0098] In summary, the present invention proposes a novel spatio-temporal convolutional network, which can effectively extract time and space features from raw data containing a large amount of noise and interference, significantly improving the prediction accuracy and the robustness of the model. This method provides a new approach for the field of bearing remaining fault diagnosis and demonstrates its potential for application under complex operating conditions.
Claims
1. A bearing fault diagnosis method based on a spatio-temporal convolutional network, characterized in that: S1. The original data is first normalized; then it is input into the spatio-temporal convolutional network, which contains two branch structures: a time convolutional block and a spatial convolutional block. S2. The time convolutional block uses its huge receptive field to extract the forward and backward dependencies of the entire data and sense the detailed features of the data. S3. The spatial convolutional block maps one-dimensional data to a two-dimensional space through polar coordinate transformation to form an image; by using the powerful image processing ability of the convolutional neural network, it senses the overall features of the data. S4. Subsequently, the outputs of the two branches will pass through a multi-head attention mechanism to enable the model to pay more attention to important features; finally, it passes through a SoftMax layer to obtain the final output. The further improvement lies in that: S1 specifically includes the following steps: S11. Obtain the original degradation data collected by the bearing vibration signal sensor and perform data preprocessing, including removing abnormal points and filling missing points in the data; S12. Normalize the original data and scale the data to the range of [0, 1]; normalization can make data with different dimensions or magnitudes comparable, thereby improving the training efficiency and prediction accuracy of the model; the formula is as follows: where X max is the maximum value in the data, and X min is the minimum value in the data. The further improvement lies in that: S2 specifically includes the following steps: S21. One part of the spatio-temporal convolutional network is the time convolutional block, which focuses on the detailed features of the data. Inspired by the time convolutional network, it uses causal dilated convolution in the time convolutional network to capture the detailed features in the fault data with an ultra-large receptive field, so that the subtle differences between different fault data can be compared. S22. Dilated convolution allows the convolutional kernel to sample at a fixed interval in the input data, thereby significantly expanding the effective coverage range of the convolutional kernel without increasing the computational complexity; this mechanism not only improves the network's ability to model long-term dependencies but also avoids the problem of training difficulties caused by the network being too deep; the formula for dilated causal convolution is as follows: where * represents the convolution operator, d is the dilation factor, k is the filter size, g(i) represents the i-th element in the filter, and u s-d · i represents the (s - d·i)-th element in the input sequence. S23. Although dilated causal convolution is introduced, the model may still become very deep in some cases. Although the deep network structure can capture complex features and long-term dependencies, it is also prone to problems such as gradient disappearance or gradient explosion, which affect the training efficiency and performance of the model. To solve this problem, residual connections are adopted. S24. Assume that the input of a neural network layer (or a group of layers) is x, and the output after the transformation of this layer is F(x), then the output y of the residual connection can be expressed as: y = F(x) + x The further improvement lies in that: S3 specifically includes the following steps: S31. Another part of the spatio-temporal convolutional network is the spatial convolutional block, which focuses on the overall features of the data; by converting one-dimensional data into a two-dimensional image, the overall trend of the fault data can be accurately grasped. S32. After receiving the time series tensor \(x\in\mathbb{R}\) after the normalization process in S12 B×L×C , where \(B\) represents the batch size, \(L\) represents the sequence length, and \(C\) represents the number of channels; subsequently, the inverse cosine function is used to map the normalized sequence values to the angular space, and the polar coordinate angle \(\varphi\) of each value is calculated; the angle calculation formula is as follows: φ = arccos(clamp(x norm , -1.0, 1.0)) Here, the Clamp function ensures that the input value is within the range of [-1, 1] to avoid out-of-domain values caused by numerical errors; by calculating the angle, the one-dimensional sequence values are converted into points in the angular space. S33. Calculate the sum (or difference) of two angles and apply the cosine function to measure the relationship between them; each element of the two-dimensional angle matrix represents the angular relationship between two time steps, and its calculation formula is: where φ i and φ j are the angular values at different time steps in the sequence, respectively. Through the broadcasting mechanism, all pairwise combinations of all time steps can be efficiently calculated to form a two-dimensional angular matrix of size (B, C, L, L). S34. By converting one-dimensional time series data into a two-dimensional image, the convolutional neural network can capture the global patterns and trends of the time series in the two-dimensional space; by combining spatial convolutional blocks, the representational ability of the model for complex time series data can be effectively improved. The further improvement lies in that: S4 specifically includes the following steps: S41. Next, the output of the spatio-temporal convolutional network is fed into the multi-head attention mechanism through a fully connected layer. The multi-head attention mechanism is one of the core modules in the Transformer model, which significantly enhances the model's ability to understand and process complex data; the multi-head attention mechanism pays attention to the input data from different angles and subspaces by processing multiple attention heads in parallel, so as to capture more comprehensive and rich feature information; the calculation process of the multi-head attention mechanism can be summarized into the following steps. S42. Step 1: Perform a linear transformation on the input tensor X to obtain three tensors: query Q, key K, and value V; this step is usually implemented through three different fully connected layers, namely: Q = xw q , K = xw k , V = xw v Among them, W q , W k and W v are weight matrices obtained through training. S43. Step 2: Split the query Q, key K, and value V into h heads (i.e., h subspaces), and each head has different linear transformation parameters; for each head i, calculate its corresponding scaled dot-product attention output respectively, and the formula is as follows: head i = Attention(QW i Q ,KW i K ,VW i V ) Among them, W i Q , W i K and W i V are the linear transformation matrices of the query, key, and value of head i respectively; the calculation formula for scaled dot-product attention is: where, d k is the dimension of the key vector, which is used to scale the dot product to stabilize the SoftMax function. S44. Step 3: Perform a final linear transformation on the concatenated vector to integrate the information from different heads and obtain the final multi-head attention output; this step is usually implemented through a fully connected layer, namely: MultiHead(Q,K,V)=Concat(head1,head2,…,head h )W O After the multi-head attention mechanism carefully processes the input data, the model can fully capture various semantic associations and feature information in the data. S45. Subsequently, these information-rich outputs are passed to the SoftMax function, which, with its unique normalization ability, converts these outputs into a probability distribution; for an input vector z = [z1, z2,..., z n , the formula of the SoftMax function is as follows: Finally, the classification result is obtained, and the output is a probability distribution, and the sum of all elements is 1.