Blood smear image classification method and system based on knowledge distillation and graph neural network

Through the method based on knowledge distillation and graph neural network, the problem of high computational cost of blood smear image classification in resource-limited environments and difficulty in processing time sequence data is solved, and efficient and accurate blood smear image classification in resource-limited environments is achieved.

CN120472213AActive Publication Date: 2025-08-12GUANGZHOU MEDICAL UNIV

Patent Information

Application Number
CN202510556961.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The prior art is expensive to process blood smear images and is difficult to deploy to environments with limited resources. The existing knowledge distillation model cannot effectively process timing data, resulting in poor classification performance of blood smear images.

Method used

Using a method based on knowledge distillation and graph neural network, the training feature matrix is generated by obtaining the training data set for preprocessing, and iterative training is performed using feature distillation losses, classification losses and timing consistency losses between the teacher model and the student model to optimize the student model to ensure that it can be classified stably in a resource-limited environment.

Benefits of technology

It reduces the cost of model calculation, improves the accuracy and stability of blood smear image classification, can be deployed in resource-limited environments, handles irregular time series data, and ensures consistency of time dimensions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472213A_ABST
    Figure CN120472213A_ABST
Patent Text Reader

Abstract

The invention discloses a blood smear image classification method and system based on knowledge distillation and a graph neural network, and the method comprises the steps: obtaining a training data set, and carrying out the preprocessing, and obtaining a training feature matrix; inputting the training feature matrix into a pre-trained teacher model, outputting teacher features and a teacher prediction result, inputting the training feature matrix into a student model, and outputting student features and a student prediction result; calculating characteristic distillation loss, classification loss and time sequence consistency loss, generating student model loss, and optimizing a student model; and when a to-be-processed blood smear image and a test period are obtained, preprocessing the to-be-processed blood smear image and the test period, inputting the to-be-processed blood smear image and the test period into the trained student model and outputting a classification result. According to the method, the training difficulty is reduced through knowledge distillation, so that the stability and effect of a blood smear image processing result are improved, deployment can be realized in an environment with limited resources, and the practicability of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a blood smear image classification method, system, terminal and computer-readable storage medium based on knowledge distillation and graph neural network. Background Art

[0002] Acute myeloid leukemia (AML) is a disease of the blood and bone marrow characterized by the rapid growth of abnormal cells that interfere with the function of normal blood cells. Advances in deep learning technology and the application of graph neural networks (GNNs) are enabling researchers to more accurately identify and classify AML blood smear images and associated test data.

[0003] However, when graph neural networks and deep learning technologies are currently used to process the corresponding blood smear images, they are difficult to deploy in resource-limited environments due to the high computational cost. In addition, most existing knowledge distillation models cannot process time series data, resulting in poor performance in the classification task of blood smear images, thus affecting the classification of blood smear images.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide a blood smear image classification method, system, terminal and computer-readable storage medium based on knowledge distillation and graph neural network, aiming to solve the problem that when graph neural network and deep learning technology are used to process the corresponding blood smear images in the existing technology, the computational cost is high and it is difficult to deploy them in a resource-limited environment. In addition, most of the existing knowledge distillation models cannot process time series data, resulting in poor performance in the classification task of blood smear images, thereby affecting the classification of blood smear images.

[0006] To achieve the above objectives, the present invention provides a blood smear image classification method based on knowledge distillation and graph neural network, which includes the following steps:

[0007] Obtaining a training data set, and preprocessing the training blood smear images of the training samples in the training data set to obtain a training feature matrix;

[0008] Inputting the training feature matrix into a pre-trained teacher model, outputting teacher features and teacher prediction results, inputting the training feature matrix into a student model, outputting student features and student prediction results;

[0009] Calculating feature distillation loss based on the teacher features and the student features, calculating classification loss based on the student prediction results and the labels in the training samples, calculating temporal consistency loss based on the teacher prediction results and the student prediction results, generating a student model loss based on the feature distillation loss, the classification loss, and the temporal consistency loss, and optimizing the student model based on the student model loss;

[0010] The student model is iteratively trained. When the training meets the preset requirements, the trained student model is output. When the blood smear image and the inspection cycle to be processed are obtained, the blood smear image to be processed and the inspection cycle are pre-processed and input into the trained student model to output the classification result.

[0011] Optionally, preprocessing the training blood smear images of the training samples in the training data set to obtain a training feature matrix specifically includes:

[0012] Obtaining a test cycle, inputting the training blood smear image and the test cycle into a timing alignment module for timing alignment to obtain an image sequence;

[0013] The image sequence and the training blood smear image are input into a graph structure construction module to generate a training feature matrix, wherein the training feature matrix includes a node feature matrix and an edge feature matrix.

[0014] Optionally, inputting the training feature matrix into a pre-trained teacher model and outputting teacher features and teacher prediction results specifically includes:

[0015] Input the training feature matrix into the graph attention network of the pre-trained teacher model, and use the features output by the second layer of the graph attention network as teacher features;

[0016] A time dimension is added to the features of each time point output by the graph attention network to obtain the corresponding teacher final feature representation, and the teacher final feature representation is input into the bidirectional gated recurrent unit of the pre-trained teacher model to generate positive and negative latent states. Based on the positive and negative latent states, the output layer and the feedforward neural network, the teacher prediction results are generated.

[0017] Optionally, inputting the training feature matrix into a student model and outputting student features and student prediction results specifically includes:

[0018] Inputting the training feature matrix into the graph convolutional network in the student model, and using the graph convolutional network output as the student feature;

[0019] A time dimension is added to the features of each time point output by the graph convolutional network to obtain the corresponding student final feature representation, and the student final feature representation is input into the gated recurrent unit of the student model to generate a hidden state. Based on the hidden state, the output layer and the feedforward neural network, the student prediction result is generated.

[0020] Optionally, calculating the feature distillation loss based on the teacher features and the student features, calculating the classification loss based on the student prediction results and the labels in the training samples, and calculating the temporal consistency loss based on the teacher prediction results and the student prediction results specifically include:

[0021] According to a pre-constructed cross-layer projection matrix, the teacher features are mapped to the feature space of the student model to obtain projection features;

[0022] Calculating a bulldozer distance according to the projection feature and the student feature, and obtaining the feature distillation loss according to the bulldozer distance;

[0023] Calculate the cross entropy loss based on the student prediction results and the labels in the training samples and use it as the classification loss;

[0024] According to the teacher prediction result and the student prediction result, dynamic time warping is used to calculate the distance between the trained teacher model and the student model, and an optimal path is found according to the dynamic time warping algorithm;

[0025] According to the found optimal path, the cosine similarity loss is calculated and used as the temporal consistency loss.

[0026] In addition, to achieve the above objectives, the present invention also provides a blood smear image classification system based on knowledge distillation and graph neural network, wherein the blood smear image classification system based on knowledge distillation and graph neural network includes:

[0027] a preprocessing module, configured to obtain a training data set, and preprocess the training blood smear images of the training samples in the training data set to obtain a training feature matrix;

[0028] A model output module is used to input the training feature matrix into a pre-trained teacher model, output teacher features and teacher prediction results, input the training feature matrix into a student model, and output student features and student prediction results;

[0029] a loss calculation module, configured to calculate a feature distillation loss based on the teacher features and the student features, calculate a classification loss based on the student prediction results and the labels in the training samples, calculate a temporal consistency loss based on the teacher prediction results and the student prediction results, generate a student model loss based on the feature distillation loss, the classification loss, and the temporal consistency loss, and optimize the student model based on the student model loss;

[0030] The application module is used to iteratively train the student model. When the training meets the preset requirements, the trained student model is output. When the blood smear image and the inspection cycle to be processed are obtained, the blood smear image to be processed and the inspection cycle are pre-processed and then input into the trained student model to output the classification result.

[0031] Optionally, the preprocessing module includes:

[0032] an alignment unit, configured to obtain a test period, and input the training blood smear image and the test period into a timing alignment module for timing alignment to obtain an image sequence;

[0033] The node construction unit is used to input the image sequence and the training blood smear image into the graph structure construction module to generate a training feature matrix, wherein the training feature matrix includes a node feature matrix and an edge feature matrix.

[0034] Optionally, the loss calculation module includes:

[0035] a feature distillation loss calculation unit, configured to map the teacher features to the feature space of the student model according to a pre-constructed cross-layer projection matrix to obtain projection features, calculate a bulldozer distance according to the projection features and the student features, and obtain the feature distillation loss according to the bulldozer distance;

[0036] a classification loss calculation unit, configured to calculate a cross entropy loss based on the student prediction results and the labels in the training samples, and use the cross entropy loss as the classification loss;

[0037] a temporal consistency loss calculation unit, configured to calculate the distance between the trained teacher model and the student model using dynamic time warping based on the teacher prediction result and the student prediction result, find an optimal path based on the dynamic time warping algorithm, calculate the cosine similarity loss based on the found optimal path, and use it as the temporal consistency loss;

[0038] The student model loss calculation unit is used to generate a student model loss according to the feature distillation loss, the classification loss and the temporal consistency loss, and optimize the student model according to the student model loss.

[0039] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a blood smear image classification program based on knowledge distillation and graph neural network stored on the memory and runnable on the processor, and when the blood smear image classification program based on knowledge distillation and graph neural network is executed by the processor, the steps of the blood smear image classification method based on knowledge distillation and graph neural network as described above are implemented.

[0040] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a blood smear image classification program based on knowledge distillation and graph neural network, and when the blood smear image classification program based on knowledge distillation and graph neural network is executed by a processor, the steps of the blood smear image classification method based on knowledge distillation and graph neural network as described above are implemented.

[0041] In the present invention, a training data set is obtained, and the training blood smear images of the training samples in the training data set are preprocessed to obtain a training feature matrix; the training feature matrix is input into a pre-trained teacher model, and the teacher features and teacher prediction results are output; the training feature matrix is input into a student model, and the student features and student prediction results are output; feature distillation loss is calculated based on the teacher features and the student features, classification loss is calculated based on the student prediction results and the labels in the training samples, temporal consistency loss is calculated based on the teacher prediction results and the student prediction results, student model loss is generated based on the feature distillation loss, the classification loss and the temporal consistency loss, and the student model is optimized based on the student model loss; the student model is iteratively trained, and when the training meets the preset requirements, the trained student model is output, and when the blood smear image to be processed and the inspection cycle are obtained, the blood smear image to be processed and the inspection cycle are preprocessed and input into the trained student model, and the classification result is output. The present invention ensures not only the similarity of surface features but also the transmission of deep spatial structural information by calculating feature distillation loss, thereby more accurately capturing the intrinsic distribution characteristics of the data and improving the performance of the student model; at the same time, the temporal consistency loss is introduced to effectively process irregular time series data and ensure consistency and continuity in the time dimension; in addition, the present invention trains the corresponding model through knowledge distillation, so that the student model obtained by the present invention does not require high computational costs and can be deployed in more environments, so that stable processing can be achieved in different environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a flowchart of a preferred embodiment of the blood smear image classification method based on knowledge distillation and graph neural network of the present invention;

[0043] Figure 2 Schematic diagram of the training process in the blood smear image classification method based on knowledge distillation and graph neural network of the present invention;

[0044] Figure 3 This is a structural diagram of a preferred embodiment of a blood smear image classification system based on knowledge distillation and graph neural network of the present invention;

[0045] Figure 4 FIG. 4 is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0047] Acute myeloid leukemia (AML) is a disease of the blood and bone marrow characterized by the rapid growth of abnormal cells that interfere with the function of normal blood cells. Advances in deep learning and the application of graph neural networks (GNNs) have enabled researchers to more accurately identify and classify AML blood smear images and related test data. However, current approaches using GNNs and deep learning to process these blood smear images have limitations in handling time series data. In AML clinical scenarios, each case's data often exhibits temporal continuity, requiring models to not only understand static image information but also capture trends over time. Current methods often fail to fully account for these dynamic characteristics. For example, traditional logit distillation methods suffer from a loss of temporal coherence due to single-time-step feature alignment in medical time series data. Furthermore, existing knowledge distillation models suffer from incomplete knowledge transfer and limited adaptability. Throughout the knowledge distillation process, the teacher model's knowledge includes not only the probability distribution of the final output but also the feature representations of the intermediate layers. However, existing knowledge distillation models typically do not include GRU modules. This results in the teacher model being unable to transfer time series features to the student model, resulting in the student model learning incomplete knowledge and struggling to fully absorb and apply it to real-world tasks. Finally, due to the limitations of existing graph neural network applications, existing GCN / GAT solutions primarily process static graph data and lack the ability to model dynamic time series graphs.

[0048] In response to one or more of the above problems, the present invention obtains a training data set, preprocesses the training blood smear images of the training samples in the training data set to obtain a training feature matrix; inputs the training feature matrix into a pre-trained teacher model, outputs the teacher features and the teacher prediction results, inputs the training feature matrix into a student model, outputs the student features and the student prediction results; calculates the feature distillation loss based on the teacher features and the student features, calculates the classification loss based on the student prediction results and the labels in the training samples, calculates the temporal consistency loss based on the teacher prediction results and the student prediction results, generates the student model loss based on the feature distillation loss, the classification loss and the temporal consistency loss, and optimizes the student model based on the student model loss; iteratively trains the student model, and when the training meets the preset requirements, outputs the trained student model; when the blood smear image and the inspection cycle to be processed are obtained, the blood smear image to be processed and the inspection cycle are preprocessed and input into the trained student model to output the classification result.

[0049] The blood smear image classification method based on knowledge distillation and graph neural network described in the preferred embodiment of the present invention is as follows: Figure 1 As shown, the blood smear image classification method based on knowledge distillation and graph neural network includes the following steps:

[0050] Step S10: Acquire a training data set, and pre-process the training blood smear images of the training samples in the training data set to obtain a training feature matrix.

[0051] Specifically, in the present invention, the training data set is a set of training blood smear images and corresponding labels. In the present invention, both the training and the actual reasoning process need to pre-process the corresponding blood smear images.

[0052] Furthermore, the preprocessing of the training blood smear images of the training samples in the training data set to obtain a training feature matrix specifically includes:

[0053] Obtaining a test cycle, inputting the training blood smear image and the test cycle into a timing alignment module for timing alignment to obtain an image sequence;

[0054] The image sequence and the training blood smear image are input into a graph structure construction module to generate a training feature matrix, wherein the training feature matrix includes a node feature matrix and an edge feature matrix.

[0055] Specifically, in the present invention, when performing preprocessing, the processing is performed specifically through two modules, namely, a time series alignment module and a graph structure construction module.

[0056] The time series alignment module employs dynamic time warping (DTW) to address the mismatch between the sampling frequency of blood smear images and the laboratory testing cycle. DTW nonlinearly stretches or compresses the time axis to find the optimal alignment path between the two time series, maximizing their similarity. This ensures that image features accurately correspond to corresponding laboratory test results, even when the sampling intervals of blood smear images are inconsistent, improving the accuracy of data analysis. The graph structure construction module converts cells in the blood smear image into nodes in the graph and defines edge weights based on specific rules. The features of each blood cell node (vi) consist of morphological parameters, staining intensity, and spatial coordinates, which capture key cell properties. Edge weight calculations first determine the connectivity between nodes using the Delaunay triangulation method. This method is a geometric partitioning method that ensures the minimum angle of the generated triangles is maximized, helping to avoid narrow triangles. Furthermore, edge weights are further adjusted based on the similarity of cell types to enhance the representation of relationships between cells of the same type. This better reflects the structural characteristics and interactions within the blood cell population, providing a foundation for subsequent analysis.

[0057] Furthermore, in the timing alignment module, the input data is a blood smear image sequence I = {I1, I2, ..., I N}, where I i is the i-th image, and the inspection period is T={T1,T2,……,T M}, where T j is the jth inspection cycle; the subsequent processing process is: calculate the distance matrix, calculate each image I i and the distance d between each inspection cycle ij , that is, construct an N×M distance matrix D, where D ij =d(I i , T j );Apply DTW algorithm and accumulate distance D cum (i, j) = d(I i , T j )+min(D cum (i-1, j), D cum (i, j-1), D cum (i-1, j-1)), find the optimal path P from (1, 1) to (N, M) to minimize the cumulative distance. The boundary condition is D cum (0, 0) = 0, and other boundary values are initialized to infinity; finally, backtrack to get the optimal path, starting from (N, M), and backtracking to (1, 1) based on the minimum cumulative distance to get the optimal path k is the path length; after alignment, the image sequence is rearranged according to the optimal path P, and the image sequence A={A1, A2, ..., AN}, so that it is aligned with the test period sequence T.

[0058] After the time sequence alignment module is processed, the rearranged image sequence A is input into the graph structure construction module, where the rearranged image sequence A is input, and each image A i Contains the position information P of several blood cells i ={P i1 , P i2 ,……,P iK}, where P ij is the jth blood cell in image A i The processing process of the graph structure building module is as follows: extract the nodes, and convert the position P of each blood cell into ij Considered as a node v in the graph k , therefore, the node set V={V1,V2,……,V n}, where n is the total number of blood cells in all images, Generate edges between nodes according to the Delaunay triangulation algorithm. The result of Delaunay triangulation is a planar graph G = (V, E), where each edge (v i , v j )∈E all satisfy that the circumcircle of any two adjacent triangles does not contain any other vertices, thereby maximizing the minimum angle and avoiding narrow triangles. When a node set V is given, Delaunay triangulation is used to generate an edge set E={(V i ,V j )|W ij}. Where W ij For the edge (v i , v j ) weight, W ij =||V i -V j ||2; Finally, the training feature matrix is obtained, and the training feature matrix includes the node feature matrix And the edge feature matrix D, the node feature matrix H is the features of all nodes obtained, and the edge feature matrix D contains the features corresponding to each edge.

[0059] Step S20: input the training feature matrix into the pre-trained teacher model, output the teacher features and teacher prediction results, input the training feature matrix into the student model, output the student features and student prediction results.

[0060] Specifically, this paper addresses the high cost and low coverage of traditional detection methods in the field of AML diagnosis by proposing a blood smear image classification method based on knowledge distillation and graph neural networks. This method can reduce model complexity and parameter count while ensuring model performance. The lightweight model can operate in resource-constrained environments, reducing hardware requirements and costs while improving system flexibility and deployability. This is achieved by employing a corresponding knowledge distillation method.

[0061] It should be noted that if Figure 2 As shown, the teacher model in the present invention is composed of GAT (Graph Attention Network) and BiGRU (Bidirectional Gated Recurrent Unit), and is then connected to an intermediate layer, an output layer and an FNN (Feed-Forward Network); the student model is composed of GCN (Graph Convolutional Network) and GRU (Gated Recurrent Unit), and is then also connected to an intermediate layer, an output layer and an FNN; and in the present invention, the intermediate layer, output layer and FNN of the teaching model and the student model are shared.

[0062] The teacher-student architecture in the present invention is a dual-stream distillation architecture. In the construction of the teacher model, GAT is used to extract local features from blood smear images. GAT enhances the interaction ability between nodes through the self-attention mechanism, which helps to capture richer spatial relationships; then, BiGRU is used to process the extracted feature sequence to capture dynamic changes and high-order spatial dependencies in the long-range time dimension. GCN is used instead of GAT in the design of the student model because GCN has higher computational efficiency and is more suitable for lightweight requirements. Similarly, the student model also integrates GRU to process time series information to ensure consistency with the teacher model. When training the student model, the corresponding distillation algorithm adopts a dynamic weight distillation mechanism, designs an adaptive temperature coefficient regulator based on the severity of the case, and implements differentiated knowledge transfer technology effects in the feature layer or output layer.

[0063] Furthermore, the step of inputting the training feature matrix into a pre-trained teacher model and outputting teacher features and teacher prediction results also includes:

[0064] Obtaining a first training data set, and preprocessing the training blood smear images in the first training data set to obtain a current training feature matrix;

[0065] Input the current training feature matrix into the teacher model and output the current teacher prediction result;

[0066] Calculating a teacher classification loss based on the current teacher prediction result and the labels in the first training data set, and optimizing the teacher model based on the teacher classification loss;

[0067] The teacher model is iteratively trained, and when the training reaches the first preset requirement, the corresponding teacher model is output as the pre-trained teacher model.

[0068] In the present invention, before the student model is trained accordingly, the teacher model is trained first. After the teacher model is trained by the first training data set, the parameters of the teacher model are fixed, and the student model is trained by the teacher model. Wherein, in the process of training the teacher model, the first training data set training data set adopted can be the same data, which is also a set of training blood smear images and corresponding labels. In the process of training the teacher model, the GAT of the teacher model is initialized using He, and BiGRU is initialized using orthogonal initialization; AdamW is used as the optimizer, and the learning rate is set to 1e-4, and the weight decay is 1e-5; and the teacher model only focuses on the classification loss (main loss), that is, the loss is calculated by the model output results and labels, thereby optimizing the parameters, wherein, in the forward propagation process of the teacher model, the graph structure data is input to the GAT layer, the second layer output of the GAT is used as the feature distillation source, and is passed to the BiGRU to process the time series features, and gradient clipping (the threshold is set to 5.0) is used to update the parameters layer by layer for back propagation.

[0069] In one embodiment of the present invention, the teacher model is pre-trained for 50 epochs with a batch size of 32. Its performance is judged by evaluation metrics including accuracy, precision, recall, F1 score, ROC curve, and AUC value. The first preset requirement is a user-preset requirement, which can be achieved by achieving the required number of training times or training results.

[0070] As for the teacher model, GAT is applied to extract local features from blood smear images; GAT enhances the interaction ability between nodes through the self-attention mechanism, so that each node can dynamically adjust the weight according to the importance of its neighboring nodes, thereby more effectively capturing complex spatial relationships. Subsequently, BiGRU is used to process the feature sequence extracted by GAT. BiGRU can not only capture dynamic changes in the time dimension, but also identify high-order spatial dependencies. The forward and reverse GRUs process the forward and backward directions of the time series data respectively, allowing the model to consider the past and future contextual information at the same time, providing richer background support for the current moment.

[0071] Furthermore, the training feature matrix is input into a pre-trained teacher model, and teacher features and teacher prediction results are output, specifically including:

[0072] Input the training feature matrix into the graph attention network of the pre-trained teacher model, and use the features output by the second layer of the graph attention network as teacher features;

[0073] A time dimension is added to the features of each time point output by the graph attention network to obtain the corresponding teacher final feature representation, and the teacher final feature representation is input into the bidirectional gated recurrent unit of the pre-trained teacher model to generate positive and negative latent states. Based on the positive and negative latent states, the output layer and the feedforward neural network, the teacher prediction results are generated.

[0074] Furthermore, in the process of training the student model, after obtaining the training feature matrix, the training feature matrix is input into the pre-trained teacher model, and the teacher features and teacher prediction results are output, which specifically include: GAT extracting local features and BiGRU processing feature sequences.

[0075] GAT is a method based on graph neural network (GNN) that enhances the interaction ability between nodes through the self-attention mechanism. Each node can dynamically adjust the weight according to the importance of its neighboring nodes, thereby more effectively capturing complex spatial relationships. The input graph structure G = (V, E), V is the node set, E is the edge set, and each node v i ∈V has a eigenvector For each pair of nodes (v i , v j ), attention coefficient e ij =LeakyReLU(a T [Wh i ||Wh j ]), where W is a shared linear transformation matrix used to transform node features; a is a learnable attention vector; || represents a vector concatenation operation. Then for each node v i Normalize the attention coefficient of neighbor nodes Where N(i) represents the node v i Next, update the node feature h' based on the normalized attention coefficient i =σ(∑ j∈N(i) α ij Wh j ), where σ is the activation function. In order to improve the stability of the model, a multi-head attention mechanism is used in the model. Each head K They are all calculation results of independent attention mechanisms.

[0076] Bidirectional Gated Recurrent Unit (BiGRU) is a bidirectional recurrent neural network (RNN) that can consider the forward and backward directions of time series data at the same time, allowing the model to utilize both past and future contextual information. The feature sequence after the previous step of GAT processing is H (2) ={h′1,h′2,……,h′ N}, where h′ i is the feature vector of the i-th node. The forward GRU processes the forward part of the feature sequence. is the hidden state of the forward GRU at time t. The reverse GRU processes the backward part of the feature sequence. is the hidden state of the reverse GRU at time t. Finally, the positive and negative hidden states are combined [.;.] represents a vector concatenation operation.

[0077] That is, in the present invention, the input of the teacher model is the preprocessed graph structure G = (V, E), that is, the training feature matrix, where V is the node set, that is, the node feature matrix, and each node v i Corresponding to a feature vector F is the feature dimension. E is the edge set, that is, the edge feature set, which represents the connection relationship between nodes, using the feature matrix Indicates that N is the number of nodes. In the present invention, the node feature matrix is input to GAT And the feature matrix D of the edge, after a layer of GAT, each node will generate a new feature representation F1 is the output feature dimension of the first layer of GAT, and the feature representation is input to the second layer; after passing through the second layer of GAT, each node will generate the final feature representation F2 is the output feature dimension of the second layer GAT, H (2) The shape of is (N, F2). Each GAT is followed by BiGRU, but before inputting into BiGRU, the GAT output needs to be converted into a format suitable for sequence processing, that is, the time dimension T is added to the features of each time point output by the graph attention network. Then the data shape input to BiGRU is (T, N, F2), and the feature sequence can be expressed as It can also be expressed as H (2) ={h′1,h′2,……,h′ N Afterwards, in BiGRU, the forward GRU processes the sequence data from the first time step to the last time step in sequence to generate the forward hidden state; the reverse GRU processes the sequence data from the first time step to the last time step in sequence to generate the reverse hidden state; BiGRU splices the forward and reverse hidden states together to obtain the final output 2W is the output dimension of BiGRU.

[0078] Therefore, in the present invention, for the training feature matrix processing process, such as Figure 2 As shown in the figure, the data of the three time nodes are input into GAT respectively, and the output of the second layer in the GAT processing corresponding to each time point is used as the corresponding teacher feature; then each GAT output is added to the time dimension and input into BiGRU, and at the same time, the output of the BiGRU at the previous time point is combined in each BiGRU, and finally the obtained result is input into the output layer, and then the final teacher prediction result is output through FNN.

[0079] The Middle Layer, located between the input and output layers, uses linear transformations and nonlinear activation functions to learn features such as different cell types, nuclear morphology, and intercellular relationships in a blood smear. This helps identify valuable features for AML diagnosis and removes unnecessary details, thereby improving the model's diagnostic accuracy and reducing computational burden. Furthermore, the Dropout regularization technique incorporated into the Middle Layer prevents overfitting and improves model generalization. The Output Layer, located before the FFN, reduces dimensionality through linear transformation, given that the output features of the teacher and student models may be high-dimensional, while the FFN requires lower-dimensional input. Furthermore, the Output Layer facilitates a smooth transition from the teacher model to the student model, alleviating the performance gap between the two models and making it easier for the student model to mimic the teacher's behavior, thus facilitating a smooth transition. An FNN is a feedforward neural network, which has two main functions: First, the final layer of the FFN acts as a classifier. Since the final output of this model is the classification result of the blood smear image, which corresponds to a three-category classification result, the final layer of the FFN maps the features to a three-dimensional vector. In addition, the FFN can increase the nonlinearity of the model through multiple fully connected layers and activation functions, thereby better capturing complex patterns in the data.

[0080] Furthermore, the step of inputting the training feature matrix into the student model and outputting student features and student prediction results specifically includes:

[0081] Inputting the training feature matrix into the graph convolutional network in the student model, and using the graph convolutional network output as the student feature;

[0082] A time dimension is added to the features of each time point output by the graph convolutional network to obtain the corresponding student final feature representation, and the student final feature representation is input into the gated recurrent unit of the student model to generate a hidden state. Based on the hidden state, the output layer and the feedforward neural network, the student prediction result is generated.

[0083] Furthermore, in the present invention, for the student model, it mainly includes GCN and GRU processing processes, and the GCN in the student model is a simplified version of GCN. The design adopts a simplified version of GCN, which is a lightweight graph convolutional neural network architecture that reduces computational complexity and the number of parameters. In order to further reduce computing requirements, the adjacency matrix update frequency of the student model is set to one-third of that of the teacher model, which helps to reduce unnecessary computational overhead while maintaining the ability to learn important spatial relationships. In addition, the time series distillation unit in the student model uses the hidden state of GRU to perform sliding window correlation constraints with the teacher model, which allows the student model to gradually approach the performance of the teacher model during the learning process, especially in the processing of time series data. The sliding window method can better capture the dynamic patterns within the sequence and realize effective knowledge transfer. In this way, the teacher model and the student model work together to ensure performance and achieve effective resource utilization.

[0084] Among them, the simplified version of the GCN graph convolution operation can be expressed as: in is the adjacency matrix with self-loops added, is the corresponding degree matrix, H is the input feature matrix, that is, the training feature matrix, and W is the learnable weight matrix.

[0085] GRU processes the feature sequence and performs temporal distillation. Specifically, the feature sequence Z after GCN processing is Z = {z1,z2,……,z N}, that is, the output corresponding to the graph convolutional network, where z i is the feature vector of the i-th node, processed by forward GRU, and h t =GRU(h t-1 , z t ), h t is the hidden state of GRU at time t; then the sliding window method is used to calculate the correlation between the hidden state of the student model and the hidden state of the teacher model Among them S teacher,t;t+w represents the hidden state of the teacher model in the time window [t, t+w], S student,t;t+w represents the hidden state of the student model in the time window [t, t+w]. T is the temperature parameter used to smooth the softmax output. w is the size of the sliding window.

[0086] That is, in this invention, the input to GCN is the node feature matrix and the normalized edge feature matrix That is, the training feature matrix. After a layer of GCN, each node will generate a new feature representation E1 is the output feature dimension of the first layer of GCN. After passing through GCN, each node will generate the final feature representation Z (n) The shape (N, E n ), E n It is the output feature dimension of the nth layer GCN (n is 2-3 layers according to the situation). Before inputting into GRU, the output Z of GCN is converted into (n) Convert to a format suitable for sequence processing, that is, add the time dimension T, and the data shape input to GRU is (T, N, E n ). GRU processes the sequence data from the first time step to the last time step and generates hidden states v is the output dimension of GRU.

[0087] Therefore, in the present invention, the training feature matrix processing process of the student model is as follows: Figure 2 As shown in the figure, the data of the three time nodes are input into GCN respectively, and then each GCN output is added to the time dimension and input into GRU, and at the same time, each GRU is combined with the output of the GRU at the previous time point, and finally the result is input into the output layer, and then the final teacher prediction result is output through FNN.

[0088] Furthermore, this invention implements dynamic temperature regulation. Temperature is an inherent hyperparameter in knowledge distillation, controlling the smoothness of the teacher model's output. Adjusting the temperature influences the softness or hardness of the knowledge learned by the student model from the teacher model.

[0089] To adapt to cases of varying severity, a dynamic temperature adjustment mechanism, also known as a dynamic weight distillation mechanism, is employed. Specifically, a temperature coefficient τ = σ(MLP(criticality index)) is set, where MLP is a multilayer perceptron network that dynamically adjusts the temperature value based on the criticality index of the blood smear image. Higher temperature values result in smoother soft labels, while lower temperatures cause the soft labels to approach the hard labels, i.e., sharpen them. This mechanism helps provide more precise guidance in critical situations, enhances the model's adaptability and personalized service capabilities, and strengthens the learning effect of the student model.

[0090] In dynamic temperature control, the user's pre-determined criticality index x is input and the temperature coefficient is calculated using a multi-layer perceptron (MLP): τ = σ(MLP(x)), where σ is the activation function. In addition, a softmax operation is performed on the output of the teacher model and the temperature coefficient is applied: Higher temperatures make soft labels smoother, while lower temperatures make them closer to hard labels. Specifically, when the temperature T is high, the probability distribution of the softmax output becomes smoother and the differences between different categories are reduced, making it easier for the student model to learn the "soft labels" of the teacher model; when the temperature T = 1, the softmax output is the same as the standard case, that is, hard labels; when the temperature T is low, the probability distribution of the softmax output becomes sharper and closer to the hard labels.

[0091] Step S30: Calculate feature distillation loss based on the teacher features and the student features, calculate classification loss based on the student prediction results and the labels in the training samples, calculate temporal consistency loss based on the teacher prediction results and the student prediction results, generate student model loss based on the feature distillation loss, the classification loss and the temporal consistency loss, and optimize the student model based on the student model loss.

[0092] Specifically, when the student model is trained in the present invention, the parameter initialization of the GCN of the student model is to inherit part of the weights of the teacher model GAT using transfer learning. The GRU is randomly initialized, the optimizer is SGD, the momentum is set to 0.9, the learning rate is 3e-4, and the update frequency is 1 / 3 of the teacher model. The forward propagation process is to input the graph structure data to the GCN layer and pass it to the GRU to process the time series features. In addition, the cross entropy loss (Cross Entropy Loss) is used as the main loss function, that is, the classification loss; the output sequence of the teacher model and the student model is aligned using DTW, and the cosine similarity loss after dynamic regularization is calculated as the time series consistency loss; the feature distillation loss is to apply the MSE loss (mean square error loss) between the second layer output of GAT and the final output of GCN. The student model considers classification loss, feature distillation loss and time series consistency loss at the same time. Set the student model loss L student = 0.6 classification loss + 0.3 feature distillation loss + 0.1 temporal consistency loss.

[0093] Furthermore, the calculating of feature distillation loss based on the teacher features and the student features, the calculating of classification loss based on the student prediction results and the labels in the training samples, and the calculating of temporal consistency loss based on the teacher prediction results and the student prediction results specifically include:

[0094] According to a pre-constructed cross-layer projection matrix, the teacher features are mapped to the feature space of the student model to obtain projection features;

[0095] Calculating a bulldozer distance according to the projection feature and the student feature, and obtaining the feature distillation loss according to the bulldozer distance;

[0096] Calculate the cross entropy loss based on the student prediction results and the labels in the training samples and use it as the classification loss;

[0097] According to the teacher prediction result and the student prediction result, dynamic time warping is used to calculate the distance between the trained teacher model and the student model, and an optimal path is found according to the dynamic time warping algorithm;

[0098] According to the found optimal path, the cosine similarity loss is calculated and used as the temporal consistency loss.

[0099] The step of calculating the feature distillation loss based on the teacher features and the student features specifically includes:

[0100] According to a pre-constructed cross-layer projection matrix, the teacher features are mapped to the feature space of the student model to obtain projection features;

[0101] A bulldozer distance is calculated according to the projection feature and the student feature, and the feature distillation loss is obtained according to the bulldozer distance.

[0102] In this paper, the feature distillation process occurs between the second-layer output of the GAT and the final output of the GCN. By establishing a cross-layer projection matrix, the high-level features extracted by the GAT are mapped into the feature space of the GCN. The goal is to minimize the Wasserstein distance between the two feature representations, also known as the bulldozer distance, which is an effective measure of the difference in probability distributions. This not only preserves the spatial relationship of the original features but also enables the student model to learn more detailed structural information, thereby improving its generalization ability.

[0103] Specifically, first obtain the second layer output H of the teacher model (2) And the final output Z of the student model, and then use a linear transformation matrix P to map the features of the teacher model to the feature space of the student model H P =H (2) P, then use Wasserstein distance to measure the difference L between the two feature distributions distill =W c (H p ,Z), where W c is the Wasserstein distance, defined as where ∏(μ,v) is the set of joint probability distributions from μ to v, and c(x,y) is the cost function, usually expressed as the Euclidean distance

[0104] Furthermore, the calculating of the temporal consistency loss according to the teacher's prediction result and the student's prediction result specifically includes:

[0105] According to the teacher prediction result and the student prediction result, dynamic time warping is used to calculate the distance between the trained teacher model and the student model, and an optimal path is found according to the dynamic time warping algorithm;

[0106] According to the found optimal path, the cosine similarity loss is calculated and used as the temporal consistency loss.

[0107] In this paper, a sequence alignment loss function based on DWT is designed to ensure the consistency of time series data. This loss function aims to minimize the difference between the time series output by the teacher model and the student model, even if the sampling frequency or rate of the two models are not synchronized. This method is particularly suitable for processing complex and irregular time series data such as blood smear images because it allows nonlinear alignment and ensures accurate matching in the time dimension.

[0108] The temporal consistency loss module is implemented after the teacher model and the student model generate their outputs. The pre-processed input data is forward propagated through the teacher model to obtain the output of the teacher model, i.e., the teacher prediction result S teacher , the same input data is forward propagated through the student model to obtain the output S of the student model student Use Dynamic Time Warping (DTW) to calculate the distance L between the output of the teacher model and the student model DTW =DTW(S teacher , S student ) and the DTW algorithm finds the optimal path to minimize the cumulative distance, that is, to minimize the difference between the time series output by the teacher model and the student model. Finally, the regularized cosine similarity loss is calculated as the temporal consistency loss.

[0109] The specific steps are: define a cost matrix D, Where D[i][j] represents the distance between the i-th point in the teacher model output sequence and the j-th point in the student model output sequence. Then use dynamic programming to fill in the cost matrix and calculate the cumulative distance C[i][j] = D[i][j] + min(C[i-1][j], C[i][j]-1], C[i-1][j-1]), and get the minimum cumulative distance path L DTW= C[n][m], which is the optimal alignment path. n and m are the lengths of the output sequences of the teacher model and the student model, respectively.

[0110] Step S40: iteratively train the student model. When the training meets the preset requirements, the trained student model is output. When the blood smear image to be processed and the test cycle are obtained, the blood smear image to be processed and the test cycle are pre-processed and input into the trained student model to output the classification result.

[0111] Specifically, in the present invention, when the iterative training of the student model reaches a preset condition, the trained student model is output, wherein the preset condition is set by the user, including reaching the number of training times or the model accuracy meeting the requirements.

[0112] When the blood smear image to be processed and the inspection cycle are obtained, the blood smear image to be processed and the inspection cycle are pre-processed and then input into the trained student model, and a classification result is output, wherein the classification result includes three types: low, medium and high, indicating the situation of the corresponding blood smear image, and the corresponding result is given to the processing personnel, who further processes the blood smear image according to the classification result.

[0113] The present invention can capture dependencies in time series while processing graph-structured data, which is particularly useful for data sets with time and space dimensions. The present invention can construct a more comprehensive feature representation, thereby improving classification performance; compared to simple feature mapping, the GAT+BiGRU teacher model proposed in this paper not only considers node-level features, but also strengthens the connection between important nodes through the attention mechanism, providing richer contextual information for student model learning. GAT can assign different weights to each node, so that the model can adjust its focus according to specific application scenarios. At the same time, BiGRU can help the model better understand patterns in long sequences, which is crucial for learning long-term dependencies; using GCN as part of the student model can significantly reduce computational overhead while maintaining a certain accuracy, making the present invention not only applicable to static graphs, but also extendable to dynamic graph scenarios, increasing the versatility and flexibility of the model.

[0114] The present invention obtains a training data set, preprocesses the training blood smear images of the training samples in the training data set to obtain a training feature matrix; inputs the training feature matrix into a pre-trained teacher model, outputs the teacher features and the teacher prediction results, inputs the training feature matrix into a student model, outputs the student features and the student prediction results; calculates feature distillation loss based on the teacher features and the student features, calculates classification loss based on the student prediction results and the labels in the training samples, calculates temporal consistency loss based on the teacher prediction results and the student prediction results, generates student model loss based on the feature distillation loss, the classification loss and the temporal consistency loss, and optimizes the student model based on the student model loss; iteratively trains the student model, and when the training meets preset requirements, outputs the trained student model; when the blood smear image to be processed and the inspection cycle are obtained, the blood smear image to be processed and the inspection cycle are preprocessed and input into the trained student model, and the classification result is output. The present invention ensures not only the similarity of surface features but also the transmission of deep spatial structural information by calculating feature distillation loss, thereby more accurately capturing the intrinsic distribution characteristics of the data and improving the performance of the student model; at the same time, the temporal consistency loss is introduced to effectively process irregular time series data and ensure consistency and continuity in the time dimension; in addition, the present invention trains the corresponding model through knowledge distillation, so that the student model obtained by the present invention does not require high computational costs and can be deployed in more environments, so that stable processing can be achieved in different environments.

[0115] Furthermore, if Figure 3 As shown, based on the above-mentioned blood smear image classification method based on knowledge distillation and graph neural network, the present invention also provides a blood smear image classification system based on knowledge distillation and graph neural network, wherein the blood smear image classification system based on knowledge distillation and graph neural network includes:

[0116] A preprocessing module 31 is used to obtain a training data set and preprocess the training blood smear images of the training samples in the training data set to obtain a training feature matrix;

[0117] A model output module 32 is used to input the training feature matrix into a pre-trained teacher model, output teacher features and teacher prediction results, and input the training feature matrix into a student model, output student features and student prediction results;

[0118] a loss calculation module 33, configured to calculate a feature distillation loss based on the teacher features and the student features, calculate a classification loss based on the student prediction results and the labels in the training samples, calculate a temporal consistency loss based on the teacher prediction results and the student prediction results, generate a student model loss based on the feature distillation loss, the classification loss, and the temporal consistency loss, and optimize the student model based on the student model loss;

[0119] The application module 34 is used to iteratively train the student model. When the training meets the preset requirements, the trained student model is output. When the blood smear image to be processed and the inspection cycle are obtained, the blood smear image to be processed and the inspection cycle are pre-processed and input into the trained student model to output the classification result.

[0120] The pre-processing module comprises:

[0121] an alignment unit, configured to obtain a test period, and input the training blood smear image and the test period into a timing alignment module for timing alignment to obtain an image sequence;

[0122] The node construction unit is used to input the image sequence and the training blood smear image into the graph structure construction module to generate a training feature matrix, wherein the training feature matrix includes a node feature matrix and an edge feature matrix.

[0123] The model output module includes a teacher prediction submodule and a student prediction submodule.

[0124] The teacher prediction submodule includes:

[0125] A teacher feature output unit, configured to input the training feature matrix into the graph attention network of the pre-trained teacher model, and use the features output by the second layer of the graph attention network as teacher features;

[0126] The teacher prediction result generation unit is used to add a time dimension to the features of each time point output by the graph attention network to obtain the corresponding teacher final feature representation, and input the teacher final feature representation into the bidirectional gated recurrent unit of the pre-trained teacher model to generate positive and negative latent states, and generate teacher prediction results based on the positive and negative latent states, the output layer and the feedforward neural network.

[0127] The student prediction submodule includes:

[0128] A student feature output unit, configured to input the training feature matrix into the graph convolutional network in the student model and output the graph convolutional network as the student feature;

[0129] The student prediction result generation unit is used to add a time dimension to the features of each time point output by the graph convolutional network to obtain the corresponding student final feature representation, and input the student final feature representation into the gated recurrent unit of the student model to generate a hidden state. The student prediction result is generated based on the hidden state, the output layer and the feedforward neural network.

[0130] The loss calculation module includes:

[0131] a feature distillation loss calculation unit, configured to map the teacher features to the feature space of the student model according to a pre-constructed cross-layer projection matrix to obtain projection features, calculate a bulldozer distance according to the projection features and the student features, and obtain the feature distillation loss according to the bulldozer distance;

[0132] a classification loss calculation unit, configured to calculate a cross entropy loss based on the student prediction results and the labels in the training samples, and use the cross entropy loss as the classification loss;

[0133] a temporal consistency loss calculation unit, configured to calculate the distance between the trained teacher model and the student model using dynamic time warping based on the teacher prediction result and the student prediction result, find an optimal path based on the dynamic time warping algorithm, calculate the cosine similarity loss based on the found optimal path, and use it as the temporal consistency loss;

[0134] The student model loss calculation unit is used to generate a student model loss according to the feature distillation loss, the classification loss and the temporal consistency loss, and optimize the student model according to the student model loss.

[0135] Furthermore, if Figure 4 As shown, based on the above-mentioned blood smear image classification method and system based on knowledge distillation and graph neural network, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 4 Only some of the components of the terminal are shown, but it should be understood that implementation of all of the shown components is not required, and more or fewer components may be implemented instead.

[0136] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Furthermore, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a blood smear image classification program 40 based on knowledge distillation and graph neural network is stored on the memory 20, and the blood smear image classification program 40 based on knowledge distillation and graph neural network can be executed by the processor 10, thereby realizing the blood smear image classification method based on knowledge distillation and graph neural network in the present invention.

[0137] In some embodiments, the processor 10 can be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 20, such as executing the blood smear image classification method based on knowledge distillation and graph neural network.

[0138] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch screen, etc. The display 30 is used to display information on the terminal and to display a visual user interface.

[0139] In one embodiment, when the processor 10 executes the blood smear image classification program 40 based on knowledge distillation and graph neural network in the memory 20, the steps of the above blood smear image classification method based on knowledge distillation and graph neural network are implemented.

[0140] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a blood smear image classification program based on knowledge distillation and graph neural network, and when the blood smear image classification program based on knowledge distillation and graph neural network is executed by a processor, the steps of the blood smear image classification method based on knowledge distillation and graph neural network as described above are implemented.

[0141] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal comprising the element.

[0142] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes in the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0143] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A blood smear image classification method based on knowledge distillation and graph neural network, characterized in that: The blood smear image classification method based on knowledge distillation and graph neural network includes: Obtaining a training data set, and preprocessing the training blood smear images of the training samples in the training data set to obtain a training feature matrix; Inputting the training feature matrix into a pre-trained teacher model, outputting teacher features and teacher prediction results, inputting the training feature matrix into a student model, outputting student features and student prediction results; Calculating feature distillation loss based on the teacher features and the student features, calculating classification loss based on the student prediction results and the labels in the training samples, calculating temporal consistency loss based on the teacher prediction results and the student prediction results, generating a student model loss based on the feature distillation loss, the classification loss, and the temporal consistency loss, and optimizing the student model based on the student model loss; The student model is iteratively trained. When the training meets the preset requirements, the trained student model is output. When the blood smear image and the inspection cycle to be processed are obtained, the blood smear image to be processed and the inspection cycle are pre-processed and input into the trained student model to output the classification result.

2. The blood smear image classification method based on knowledge distillation and graph neural network according to claim 1 is characterized in that: The preprocessing of the training blood smear images of the training samples in the training data set to obtain a training feature matrix specifically includes: Obtaining a test cycle, inputting the training blood smear image and the test cycle into a timing alignment module for timing alignment to obtain an image sequence; The image sequence and the training blood smear image are input into a graph structure construction module to generate a training feature matrix, wherein the training feature matrix includes a node feature matrix and an edge feature matrix.

3. The blood smear image classification method based on knowledge distillation and graph neural network according to claim 1 is characterized in that The step of inputting the training feature matrix into a pre-trained teacher model and outputting teacher features and teacher prediction results specifically includes: Input the training feature matrix into the graph attention network of the pre-trained teacher model, and use the features output by the second layer of the graph attention network as teacher features; A time dimension is added to the features of each time point output by the graph attention network to obtain the corresponding teacher final feature representation, and the teacher final feature representation is input into the bidirectional gated recurrent unit of the pre-trained teacher model to generate positive and negative latent states. Based on the positive and negative latent states, the output layer and the feedforward neural network, the teacher prediction results are generated.

4. The blood smear image classification method based on knowledge distillation and graph neural network according to claim 1 is characterized in that The step of inputting the training feature matrix into the student model and outputting student features and student prediction results specifically includes: Inputting the training feature matrix into the graph convolutional network in the student model, and using the graph convolutional network output as the student feature; A time dimension is added to the features of each time point output by the graph convolutional network to obtain the corresponding student final feature representation, and the student final feature representation is input into the gated recurrent unit of the student model to generate a hidden state. Based on the hidden state, the output layer and the feedforward neural network, the student prediction result is generated.

5. The blood smear image classification method based on knowledge distillation and graph neural network according to claim 1, characterized in that: The calculating of feature distillation loss based on the teacher features and the student features, the calculating of classification loss based on the student prediction results and the labels in the training samples, and the calculating of temporal consistency loss based on the teacher prediction results and the student prediction results specifically include: According to a pre-constructed cross-layer projection matrix, the teacher features are mapped to the feature space of the student model to obtain projection features; Calculating a bulldozer distance according to the projection feature and the student feature, and obtaining the feature distillation loss according to the bulldozer distance; Calculate the cross entropy loss based on the student prediction results and the labels in the training samples and use it as the classification loss; According to the teacher prediction result and the student prediction result, dynamic time warping is used to calculate the distance between the trained teacher model and the student model, and an optimal path is found according to the dynamic time warping algorithm; According to the found optimal path, the cosine similarity loss is calculated and used as the temporal consistency loss.

6. A blood smear image classification system based on knowledge distillation and graph neural network, characterized in that: The blood smear image classification system based on knowledge distillation and graph neural network includes: a preprocessing module, configured to obtain a training data set, and preprocess the training blood smear images of the training samples in the training data set to obtain a training feature matrix; A model output module is used to input the training feature matrix into a pre-trained teacher model, output teacher features and teacher prediction results, input the training feature matrix into a student model, and output student features and student prediction results; a loss calculation module, configured to calculate a feature distillation loss based on the teacher features and the student features, calculate a classification loss based on the student prediction results and the labels in the training samples, calculate a temporal consistency loss based on the teacher prediction results and the student prediction results, generate a student model loss based on the feature distillation loss, the classification loss, and the temporal consistency loss, and optimize the student model based on the student model loss; The application module is used to iteratively train the student model. When the training meets the preset requirements, the trained student model is output. When the blood smear image and the inspection cycle to be processed are obtained, the blood smear image to be processed and the inspection cycle are pre-processed and then input into the trained student model to output the classification result.

7. The blood smear image classification system based on knowledge distillation and graph neural network according to claim 6 is characterized in that: The pre-processing module comprises: an alignment unit, configured to obtain a test period, and input the training blood smear image and the test period into a timing alignment module for timing alignment to obtain an image sequence; The node construction unit is used to input the image sequence and the training blood smear image into the graph structure construction module to generate a training feature matrix, wherein the training feature matrix includes a node feature matrix and an edge feature matrix.

8. The blood smear image classification system based on knowledge distillation and graph neural network according to claim 6, characterized in that: The loss calculation module includes: a feature distillation loss calculation unit, configured to map the teacher features to the feature space of the student model according to a pre-constructed cross-layer projection matrix to obtain projection features, calculate a bulldozer distance according to the projection features and the student features, and obtain the feature distillation loss according to the bulldozer distance; a classification loss calculation unit, configured to calculate a cross entropy loss based on the student prediction results and the labels in the training samples, and use the cross entropy loss as the classification loss; a temporal consistency loss calculation unit, configured to calculate the distance between the trained teacher model and the student model using dynamic time warping based on the teacher prediction result and the student prediction result, find an optimal path based on the dynamic time warping algorithm, calculate the cosine similarity loss based on the found optimal path, and use it as the temporal consistency loss; The student model loss calculation unit is used to generate a student model loss according to the feature distillation loss, the classification loss and the temporal consistency loss, and optimize the student model according to the student model loss.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a blood smear image classification program based on knowledge distillation and graph neural network stored in the memory and runnable on the processor. When the blood smear image classification program based on knowledge distillation and graph neural network is executed by the processor, the steps of the blood smear image classification method based on knowledge distillation and graph neural network as described in any one of claims 1 to 5 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a blood smear image classification program based on knowledge distillation and graph neural network. When the blood smear image classification program based on knowledge distillation and graph neural network is executed by a processor, the steps of the blood smear image classification method based on knowledge distillation and graph neural network as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Traffic speed prediction method based on jump map attention gating circulation network

    CN113298319A

  • Weak supervision video time sequence behavior positioning method based on knowledge distillation

    CN113591731A

  • Flood forecasting method based on distributed adaptive graph attention network

    CN116702959A

  • Knowledge distillation-based graph neural network model compression method and system

    CN118643861A

  • Heterogeneous distillation method based on graph neural network

    CN118673329A

Cited By

  • Deep learning model-based willingness prediction method and device, equipment and medium

    CN120975839A

  • Ultrasonic image discrimination perception pre-training method based on cooperative training framework

    CN121213578A

  • An Ultrasonic Image Discriminative Perception Pre-training Method Based on a Collaborative Training Framework

    CN121213578B