Intelligent recognition method of visibly formed components in urine sediment based on HAT-MixNet
By using HAT-MixNet network and data enhancement technology to identify urine sediment components in urine testing instruments, the problems of insufficient accuracy and low working efficiency of urine sediment components in the prior art have been solved, and higher recognition accuracy and efficiency have been achieved.
Patent Information
- Application Number
- CN202411272858.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-09-11
AI Technical Summary
The urine sediment in existing urine testing instruments has problems such as insufficient accuracy and low working efficiency in identification of the formation fraction samples.
Using an intelligent identification method based on HAT-MixNet, the urine sediment formation samples were processed through data augmentation technology, and a new training sample was generated using Mixup technology, and feature visualization and model optimization were performed in combination with HeatMAP heat map technology.
It improves the identification accuracy and working efficiency of the urine sediment formation sample, reduces the model variance, enhances the robustness and generalization ability, and provides more accurate and objective identification results.
Smart Images

Figure CN119091227B_ABST
Abstract
Description
Technical Field
[0001] The method of the invention belongs to the field of intelligent processing and deep learning of medical images, and in particular relates to a method for intelligently identifying shaped components of urine sediment based on HAT-MixNet. Background Art
[0002] In recent years, with the rapid development of technologies such as artificial intelligence and big data and the increasing demand for medical care, medical instruments are constantly developing and improving in the direction of intelligence. At present, medical instruments have been widely used in many fields, such as biochemical detection systems, health monitoring, medical testing, etc., and are a must-have for many medical institutions. Intelligent medical testing instruments will develop in a smarter and more advanced direction based on the existing detection performance, realize more accurate and efficient diagnosis and treatment services, and assist doctors to better understand the occurrence, development and post-healing treatment plan design of diseases.
[0003] Urine formed element detection refers to the analysis of cells, casts, crystals and other components in urine, so as to clarify the diagnosis of related diseases. It is a very important part of medical examination. It can help doctors diagnose urinary system diseases, assess the severity of diseases, monitor treatment effects, etc. Urine formed elements include cells, casts, crystals, etc.; cells include white blood cells, red blood cells, epithelial cells, etc. Analysis of red blood cells can be used to identify the causes of various hematuria, and analysis of white blood cells can be used to diagnose urinary tract infections; analysis of casts can be used to diagnose diseases such as glomerular damage, tubular damage, urinary tract infections and renal function damage; analysis of crystals can help diagnose urinary stones, urinary tract infections, liver diseases, etc.; in short, analysis of different formed elements in urine can help provide reference and guidance for clinical diagnosis of urinary system diseases and kidney diseases.
[0004] However, traditional medical image recognition technology mainly relies on image processing and feature extraction methods. By preprocessing, segmenting and extracting features from medical images, key information in the image is extracted to achieve image recognition. The number of medical images collected by current medical instruments during the inspection process is huge, and there are many impurities in the samples (accounting for nearly 60% of the total sample volume). The characteristics of some categories of cells are extremely similar. The number of samples that can be obtained from a few categories is limited, and the sizes of samples of different categories vary greatly. Therefore, the development of an accurate and efficient new method for identifying formed components in urine sediment is an urgent need for the current research and development of medical testing instruments.
[0005] In recent years, the rapid development of deep learning technology and the demand for intelligent medical testing equipment have led to the gradual application of intelligent networks in the field of medical cell analysis and identification, bringing new development opportunities to this field. It analyzes tasks in a data-driven manner and can automatically learn relevant model features and data characteristics from large-scale data sets of specific problems, thereby achieving high-precision sample identification. Compared with traditional image processing methods, deep learning technology can implicitly and automatically learn high-level abstract features of data directly from data samples, so as to make correct decisions when detecting new data, achieve higher accuracy, and have stronger feature expression and generalization capabilities.
[0006] However, there are still many problems that deserve further study and solution in the identification of formed components in urine sediment. First, medical image datasets often have problems such as difficulty in labeling and data imbalance, which directly affect the training effect and generalization ability of deep network models. Secondly, the diversity and complexity of urine tangible samples make the design and training of models more difficult. Therefore, the identification of formed component samples in urine testing instruments currently has problems of insufficient accuracy and low work efficiency. Summary of the invention
[0007] The present invention provides a method for intelligently identifying formed components of urine sediment based on HAT-MixNet, so as to solve the problems of insufficient accuracy and low working efficiency in the identification of formed component samples in current urine testing instruments.
[0008] The technical solution adopted by the present invention comprises the following steps:
[0009] (1) The cell samples used include ten categories: red blood cells, white blood cells, mucus threads, casts, pseudohyphae yeasts, spores, salts, crystals, sperm, and impurities;
[0010] (2) Perform data enhancement on all grayscale medical cell samples, use Mixup technology to obtain data-enhanced grayscale medical cell samples as the data set samples used in the experiment, and divide the data set into a training set and a validation set;
[0011] (3) Import the enhanced grayscale medical cell samples into the HAT-MixNet network for training;
[0012] (4) The grayscale medical cell samples trained by the HAT-MixNet network are transformed into feature-visualized heat maps through the HeatMAP thermal map technology. Finally, after the cell sample types are accurately identified, the trained HAT-MixNet model is obtained;
[0013] (5) After obtaining the trained HAT-MixNet model, a grayscale medical cell sample image is input into the model to automatically identify the type of cells.
[0014] In step (2) of the present invention, the data set is divided into a training set and a validation set in a ratio of 9:1, the batch size is set to 4, the number of training epochs is set to 70, the learning rate value range is set to 0 to 0.001, and the experimental attenuation coefficient is set to between 0.0001 and 0.001.
[0015] The training process of the grayscale medical cell sample image in step (3) of the present invention when passing through the HAT-MixNet network is as follows:
[0016] 1) Read in a grayscale medical cell sample;
[0017] 2) The size of the cell sample to be read needs to be set to 224×224 and the number of feature channels to 3;
[0018] 3) Then, it goes to the Hat2d convolution layer module with a 4×4 convolution kernel k4 and a step size of 4, and then goes through a layer normalization Layer Norm to output a cell sample feature image with a size of 56×56 and a feature channel number of 96;
[0019] 4) Then, after three HAT-MixNet block processing modules, the first stage of feature processing is completed. After the first stage, the cell sample is still a feature image with a size of 56×56 and the number of feature channels is set to 96;
[0020] 5) The above feature image is processed by a downsampling layer and then a HAT-MixNet block processing module with 192 feature channels. After three repeated processings, the second stage of feature processing is completed. At this time, the output size is 28×28 and the number of feature channels is set to 192.
[0021] 6) The obtained feature image is put into the downsampling layer and then processed by the HAT-MixNet block processing module with a feature channel number of 384. After 9 repeated processing, the third stage of feature processing is completed. At this time, the output size is 14×14 and the number of feature channels is set to 384.
[0022] 7) After the feature map is put into the downsampling layer, it passes through the HAT-MixNetblock processing module with 768 feature channels. After three repeated processing, the fourth stage of feature processing is completed, and the output feature map is 7×7 in size and 768 in feature channels.
[0023] 8) The obtained feature map is passed through the global average pooling and normalization layer Layer Norm, and then the Linear layer in the PyTorch1.13.1+cu117 library is called to convert the output feature map into 1000 output features.
[0024] The downsample component of the downsample layer in the HAT-MixNet network described in the present invention is composed of a normalization layer Layer Norm and a Hat2d convolution layer module with a convolution kernel size of 2×2 and a step size of 2.
[0025] The processing process of the HAT-MixNet block processing module in the HAT-MixNet network of the present invention is:
[0026] a. Apply a Depthwise Hat2d convolution, set the initial number of channels to 96, set three 7×7 convolution kernels in the Depthwise Hat2d convolution, and set the step size to 1. After the Depthwise Hat2d convolution, perform a layer normalization LayerNorm.
[0027] b. Pass through a Hat2d convolutional layer module with a convolution kernel size of 1×1 and a stride of 1, and then pass through a GELU activation function;
[0028] c. After passing through a Hat2d convolutional layer module with a convolution kernel size of 1×1 and a step size of 1, it goes through the scaling layer LayerScale and the regularization process Drop Path;
[0029] d. Then superimpose the output feature image and the input feature image as a new output feature image.
[0030] The advantages of the present invention are:
[0031] 1. Data enhancement to build a sample set of urine sediment components
[0032] The actual clinical test samples include ten categories such as red blood cells, white blood cells, mucus filaments, casts, pseudohyphae yeast, etc., and each major category is divided into 3 to 5 small categories, a total of 50 different types of urine sediment formed components, individual categories of data volume is very small, can not meet the network training requirements, and also affect the accuracy of network recognition. Therefore, the HAT-MixNet proposed in the present invention adopts the Mixup data enhancement technology in the PyTorch library to solve the problem of sample imbalance. Mixup technology can randomly linearly combine two or more different samples according to a certain ratio to generate new training samples, achieve effective expansion of sample capacity, increase the diversity of data sets, so that the model is exposed to more data in different situations during training, and effectively prevent the occurrence of overfitting. Through the mutual fusion of different samples, the model can better learn the inherent laws and characteristics of the data, rather than just being applicable to specific samples in the training set, thereby reducing the model variance, avoiding the problem of insufficient accuracy caused by sample feature differences, and improving the robustness and generalization ability of the model.
[0033] 2. Design the HAT-MixNet network framework to achieve effective and accurate feature extraction
[0034] When the input cell image is processed in each stage in turn, the size of the feature map continues to decrease, the number of channels increases, and the feature receptive field in the feature map continues to expand. According to the number of repetitions of the four stages, the convolution module can be summarized into four stacking dimensions (3, 3, 9, 3). In the four stages of the convolution module stacking dimension (3, 3, 9, 3), except for the first stage composed of the HAT-MixNet block processing module, the other three stages are composed of the downsampling layer DownSample and the HAT-MixNet block processing module, that is, except for the first stage, the feature map of the image is input into the HAT-MixNet block processing module after downsampling DownSample.
[0035] Compared with the ResNet network framework, the HAT-MixNet network framework uses GELU as the activation function, adopts Depthwise Hat2d convolution, and finally reduces the use of normalization layers (Layer Norm). Batch Norm (BN) is completely replaced by normalization layers Layer Norm (LN), and the downsampling layer (Downsample) consists of a layer normalization (LayerNorm) and a Hat2d convolution layer with a convolution kernel size of 2×2 and a step size of 2. The design of the entire HAT-MixNet model fully considers the effective extraction of microscopic image features of urinary sediment components and network efficiency issues.
[0036] 3. HeatMAP feature visualization assisted training strategy
[0037] HAT-MixNet uses feature visualization to map the output of the network layer to the color-coded HeatMAP, so as to understand the degree of attention paid by the network model to the features of different regions of the input sample data, clarify which features in the image the network uses to accurately identify the microscopic image, and intuitively observe the response of the neural network at different levels, so as to judge whether the network correctly pays attention to the effective features of the image, adjust the auxiliary training strategy, and optimize the network performance. At the same time, the HeatMAP visualization strategy can also help discover anomalies in the network model during the training process, understand the distribution of the network model's prediction probability for microscopic images of different categories of visible components, and discover under what circumstances the network model may make errors and inaccurate predictions, which helps to adjust the network model training in a targeted manner and improve the accuracy and robustness of the network model.
[0038] The present invention is applied to the recognition and classification of various medical images such as various complex cells and impurities, providing technical support for medical image processing problems. It can adaptively extract the features of formed components according to image samples, realize accurate recognition of a large number of formed components of urine sediment with complex morphological features, and provide more accurate and objective recognition results to assist medical personnel in making further clinical diagnosis, thereby improving the accuracy of existing urine analysis and gynecological secretion analysis and the efficiency of medical diagnosis, and has important practical research value for promoting the development of intelligent medical instruments and improving the detection performance of medical instruments. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 This is a microscopic image of some poor quality urine formed element samples;
[0040] Figure 2 It is a sample of some urine formed components after data enhancement using Mixup technology;
[0041] Figure 3 It is the overall structural framework diagram of HAT-MixNet of the present invention;
[0042] Figure 4 It is a structural framework diagram of the HAT-MixNet block processing module of the present invention;
[0043] Figure 5 This is a structural framework diagram of the downsample layer of the present invention;
[0044] Figure 6 This is the confusion matrix diagram of the recognition results of the GhostNet method;
[0045] Figure 7 This is the confusion matrix diagram of the recognition results of the HAT-MixNet method;
[0046] Figure 8 This is the confusion matrix diagram of the recognition result of the ResNet method;
[0047] in Figure 6 , 7 , 01 in 8 is deformed leukocytes, 02 is leukocyte clusters, 03 is squamous epithelial cells, 04 is transitional epithelial cells, 05 is coarse granular casts, 06 is fine granular casts, 07 is waxy casts, 08 is multi-granular blastospores, 09 is few-granular blastospores, 10 is large impurity particles, 11 is small impurity particles, 12 is small oil droplets, 13 is filamentous impurities, 14 is phantom impurities, and 15 is uric acid crystals;
[0048] Fig. 9 It is a visualization of the heat map features of some urine formed component samples;
[0049] Fig.10 This is a comparison chart of the recognition accuracy curves of the three network methods;
[0050] Fig.11 This is a comparison chart of the loss function curves of three network methods. DETAILED DESCRIPTION
[0051] The actual cell samples obtained clinically contain 50 different types of urine sediment formed elements, and the microscopic images of these formed elements are of different sizes, unbalanced quantities, and complex morphologies. The present invention first performs data enhancement processing on the microscopic images of massive clinical samples, as follows:
[0052] (1) Construction of complete medical cell data samples of microscopic images of visceral components in urine sediment
[0053] For deep learning methods, training data samples are very important. The data set samples used in the present invention are actual clinical test samples from many domestic hospitals, including ten categories of red blood cells, white blood cells, mucus filaments, casts, pseudohyphae yeast, spores, salts, crystals, sperm, and impurities. Each image category is divided into 3 to 5 subcategories, with a total of 50 different types of urine sediment formed elements, with a total of hundreds of thousands of images. The sample size is huge, and data preprocessing is still required to construct a complete sample suitable for deep network training;
[0054] (2) Enhanced processing of urine formed component sample data
[0055] First, we screen the sample data with poor quality: the feature differences between individual image categories are not obvious, and some image samples contain incomplete image data, aliasing of different categories, impurity interference, etc. We will remove samples in this case to avoid interference with network training. Figure 1Some image samples with poor quality are given. It can be seen from the figure that a single image contains multiple category samples or there are impurity interference in addition to valid samples. If such images are placed in the training samples of a certain cell category, it will directly affect the judgment results of the network during the learning and training process, and it is easy to mislead the network. Therefore, such sample images should be eliminated and should not be imported into the training process of the model to ensure the effectiveness of the constructed urine sediment formed component training sample set, and provide good data support and guarantee for the subsequent deep network model learning and training. In the actual collected medical microscopic images, the sample size of some categories is insufficient (such as elliptical upright red blood cells, renal tubular epithelial cells, etc.), and the number of samples of different types is unbalanced. Among them, the sample size is larger than 10,000, while the sample size is smaller. There are only more than 100, which can easily lead to overfitting of the trained network model, insufficient accuracy, and insufficient generalization ability. How to achieve sample balance between categories is one of the key points to improve cell recognition performance.
[0056] The present invention adopts data enhancement technology to solve the problems of unbalanced urine formed component samples and insufficient training data for individual types. The so-called data enhancement is to apply a certain technology to randomly transform existing image samples to generate more credible sample data, and ensure that the transformed image will not cause trouble for network model recognition, that is, from the perspective of classification and recognition, it should still belong to the same category as the original image, but it is different from the original image, so that the network model will not use exactly the same image twice during training, ensuring that the network can observe more effective features, thereby improving the generalization ability of the model. The random transformation method for generating new images includes brightness change, horizontal shift, vertical shift, scaling, horizontal flip, vertical flip, and these transformations can also be combined.
[0057] The present invention calls the Mixup data enhancement technology in the PyTorch1.13.1+cu117 library, that is, the Mixup data enhancement technology is applied. The Mixup technology can randomly combine two or more different samples linearly to generate new training samples, thereby expanding the training sample set, generating multiple enhanced samples, performing data expansion on the samples, and using these samples for majority voting or averaging to obtain the final prediction result. Figure 2Some samples after data enhancement are given. It can be seen from the figure that the cell image is expanded through a series of operations such as rotating different angles and enhancing brightness. This data enhancement method can increase the diversity of the data set, allowing the model to be exposed to more different situations during the training process. It can also effectively prevent overfitting. By mixing different samples, the model can better learn the inherent laws and characteristics of the data, rather than just remembering specific samples in the training set, avoiding the problem of insufficient network accuracy and generalization caused by the large difference in the number of samples of different categories, thereby reducing the variance of the network model and improving the robustness and generalization ability of the model.
[0058] (3) Import the enhanced grayscale medical cell samples into the HAT-MixNet network for training; Figures 3-5 The overall framework of HAT-MixNet is given. From the main flow chart on the left, it can be observed that it is based on ResNet. The cell sample first passes through a Hat2d convolution layer module with a 4×4 convolution kernel (k4) and a step size of 4 (s4), and then passes through a normalization layer (Layer Norm). The stacking dimension of the convolution module is set to (3, 3, 9, 3). The main purpose is to recombine the previously extracted features through the convolution kernel to obtain more complex features. In convolutional neural networks, complex features in images can be recombined in some way using simple features. The convolution operation can turn a single feature of an image into a more refined image feature, which is more conducive to image recognition. Secondly, after three HAT-MixNet block processing modules, the feature image passes through a downsampling layer (Downsample) and then a HAT-MixNet block processing module with 192 feature channels, which is repeated three times; the obtained feature image is placed in the downsampling layer (Downsample) and passed through a HAT-MixNet block processing module with 384 feature channels, which is repeated nine times; the feature image is placed in the downsampling layer (Downsample) and passed through a HAT-MixNet block processing module with 768 feature channels, which is repeated three times. When the input microscopic image is processed through each stage in turn, the size of the feature map continues to decrease, the number of channels increases, and the feature receptive field in the feature map continues to expand. Each stage consists of a downsampling layer DownSample and a HAT-MixNet block processing module. Except for the first stage, the feature maps of the image are downsampled and input into the HAT-MixNet block processing module. The HAT-MixNet block processing module and the downsampling layer DownSample are respectively Figure 3The HAT-MixNet block processing module uses Depthwise Hat2d convolution to achieve a better balance between model complexity and accuracy. The core of the HAT-MixNet block processing module is to establish a jumper connection between the previous layer and the next layer, that is, the output feature image and the input feature image are superimposed, which helps the back propagation of the gradient during the training process, so as to train a deeper network model.
[0059] Finally, HAT-MixNet only adds the standard normalization layer Layer Norm (LN) after the depthwise separable convolution layer (Depthwise Hat2d) in the HAT-MixNet block processing module to calculate the mean and variance. This design independent of Batch Size can avoid the impact of the number of samples on LN calculation, reduce the model's dependence on parameter initialization, and speed up model training and improve model accuracy. Except for the first HAT-MixNet block processing module, each of the remaining modules is preceded by a downsampling layer DownSample consisting of an LN layer and a convolution layer to change the size of the feature map, facilitating the subsequent extraction of complex features at different scales. In addition, HAT-MixNet only adds the Gaussian error linear unit (GELU) activation function given in formula (1) after the 1×1 convolution layer in the HAT-MixNet block processing module:
[0060] GELU(X)=xP(X≤x),x~N(0,1)
[0061] where X is a Gaussian random variable with zero mean and unit variance, and P is the probability function, N represents the standard normal distribution.
[0062] The programming environment used in the experiments in this invention is Python 3.7.0, PyTorch1.13.1+cu117 deep learning framework, the hardware environment processor model is Inter (R) Core (TM) i5-11400 @2.6GHz, the graphics card model is NVIDIA GeForce RTX 3060, and the memory is 16GB.
[0063] During the training process, the HAT-MixNet model randomly divides the dataset into a training set and a validation set in a ratio of 9:1, sets the batch size to 4, sets the epoch to 70 (in order to ensure the comparability of the comparison methods, the three network methods all use the same training rounds), sets the learning rate value range to 0 to 0.001, and sets the experimental attenuation coefficient to between 0.0001 and 0.001.
[0064] (4) The grayscale medical cell samples trained by the HAT-MixNet network are transformed into feature-visualized heat maps through the HeatMAP thermal map technology. Finally, after the cell sample types are accurately identified, the trained HAT-MixNet model is obtained;
[0065] (5) After obtaining the trained HAT-MixNet model, a grayscale medical cell sample image is input into the model to automatically identify the type of cells;
[0066] The effects of the present invention are further illustrated below through analysis and comparison of experimental results.
[0067] In order to verify the application effect of the HAT-MixNet proposed in this paper in the actual clinical urine sediment visible component identification, experiments were carried out on 50 different types of urine sediment visible component microscopic images obtained clinically, and the effects were compared and analyzed with the GhostNet and ResNet network methods respectively. Figures 6-8 The confusion matrices of 15 types of cells and impurities that are easily misidentified for the three models are given respectively. Figure 6 is the confusion matrix of GhostNet classification results, Figure 7 is the confusion matrix of ResNet classification results, Figure 8 The confusion matrix of HAT-MixNet classification results. By observing, we can intuitively see which categories of samples are incorrectly classified into other categories. By comparing the three figures, we can find that the data in the confusion matrix of the HAT-MixNet classification results proposed by the present invention is more concentrated on the diagonal, that is, the TP and TN values are relatively high. The larger the TP and TN values, the higher the model classification accuracy, indicating that HAT-MixNet has better classification and recognition effects on various types of samples than the other two network methods. The confusion matrix experimental results prove that the classification accuracy of the HAT-MixNet model is higher, more advanced, and more feasible than other models.
[0068] In addition, HAT-MixNet also uses HeatMAP heat maps to visualize the features extracted by the network and guide network training optimization. Fig. 9Six cell samples and their corresponding heat map features are given. It can be seen from the figure that the heat map clearly shows the spatial distribution or trend of the data by mapping the data to different colors, and shows the high-density and high-intensity areas of the sample features. Taking the cell in the second row and second column as an example, the heat map shows that the area that this cell pays more attention to is the cell edge. By observing the heat map, anomalies or deviations in the training process can be quickly located, which may be the cause of poor model performance, so that the training time of the optimized model is shortened. At the same time, it can ensure that the network selects features that contribute more to the model performance and ignores those insignificant or redundant features, making the trained network more credible. The addition of the HeatMAP heat map makes the HAT-MixNet model more credible and feasible.
[0069] Fig.10 , 11 The accuracy curves and loss function curves of the three methods are given respectively. Fig.10 From the accuracy curves in , we can see that when the total number of rounds is 70, ResNet reaches a peak accuracy of 0.9397 in the 51st round, GhostNet reaches a peak accuracy of 0.9078 in the 68th round, and HAT-MixNet reaches a peak accuracy of 0.9798 in the 51st round, which is significantly higher than the other two models. Fig.11 From the loss curve in , we can see that the loss of the HAT-MixNet model is lower than that of ResNet and GhostNet in each round. The smaller the loss, the higher the recognition accuracy of the network model. It can be seen that the HAT-MixNet proposed in this invention has a significant advantage in recognition accuracy compared with the other two network methods. In addition, Table 1 lists some important parameters of the three network methods respectively.
[0070] Table 1 Comparison of various parameters and performance data of GhostNet, ResNet and HAT-MixNet
[0071] Network Name Epoch Optimal Epoch Avg-Acc Train Time (h) FLOPs(G) Params(M) Time / Sample(ms) GhostNet 70 68 0.842 33.54 0.151 3.964 20.435 ResNet 70 51 0.874 17.73 4.132 23.610 19.454 HAT-MixNet 70 51 0.951 51.61 4.455 27.837 18.630
[0072] From the perspective of parameter quantity, HAT-MixNet has a large total floating point number, indicating that its structure is relatively complex. Secondly, from the perspective of the training time required to reach the optimal state, GhostNet runs for a total of 31.54 hours, ResNet runs for a total of 17.73 hours, and HAT-MixNet runs for a total of 59.61 hours. Although the new method sacrifices the training running time, it improves the recognition accuracy. Finally, from the test time of the three models for a single image under GPU conditions, it can be seen that the prediction time of HAT-MixNet is shorter, indicating that although the training time required for the method of the present invention is longer, it can fully guarantee the test indicators of the instrument in actual application prediction. The method of the present invention uses a variety of convolution kernels to capture features of different scales, which enhances the feature picking ability of the model. At the same time, by replacing a single large convolution kernel with multiple small convolution kernels, the superiority of the model is improved. The above experimental results and data have proven that HAT-MixNet has the advantages of high accuracy and low loss compared with ResNet and GhostNet in the problem of identifying formed components of urine sediment. In summary, the HAT-MixNet proposed in this invention achieves good results in accuracy, reliability and feasibility.
Claims
1. A method for intelligent identification of visibly formed components in urine sediment based on HAT-MixNet, characterized in that: The following steps are involved: (1) The cell samples used include ten categories: red blood cells, white blood cells, mucus threads, casts, pseudohyphae yeasts, spores, salts, crystals, sperm, and impurities; (2) Perform data enhancement on all grayscale medical cell samples, use Mixup technology to obtain data-enhanced grayscale medical cell samples as the data set samples used in the experiment, and divide the data set into a training set and a validation set; (3) Import the enhanced grayscale medical cell samples into the HAT-MixNet network for training. The process is as follows: 1) Read in a grayscale medical cell sample; 2) The size of the cell sample to be read needs to be set to 224×224 and the number of feature channels to 3; 3) Then, it goes to the Hat2d convolution layer module with a 4×4 convolution kernel k4 and a step size of 4, and then goes through a layer normalization Layer Norm to output a cell sample feature image with a size of 56×56 and a feature channel number of 96; 4) Then, after three HAT-MixNet block processing modules, the first stage of feature processing is completed. After the first stage, the cell sample is still a feature image with a size of 56×56 and the number of feature channels is set to 96; 5) The above feature image is processed by a downsampling layer and then a HAT-MixNet block processing module with 192 feature channels. After three repeated processings, the second stage of feature processing is completed. At this time, the output size is 28×28 and the number of feature channels is set to 192. 6) The obtained feature image is put into the downsampling layer and then passed through the HAT-MixNetblock processing module with a feature channel number of 384. After 9 repeated processing, the third stage of feature processing is completed. At this time, the output size is 14×14 and the number of feature channels is set to 384. 7) After the feature map is put into the downsampling layer, it passes through the HAT-MixNet block processing module with 768 feature channels. After three repeated processings, the fourth stage of feature processing is completed, and the feature map with a size of 7×7 and 768 feature channels is output; 8) The obtained feature map is subjected to global average pooling and normalization layer Layer Norm, and then the Linear layer in the PyTorch1.13.1+cu117 library is called to convert the output feature map into 1000 output features; (4) The grayscale medical cell samples trained by the HAT-MixNet network are transformed into feature-visualized heat maps through the HeatMAP thermal map technology. Finally, after the cell sample types are accurately identified, the trained HAT-MixNet model is obtained; (5) After obtaining the trained HAT-MixNet model, a grayscale medical cell sample image is input into the model to automatically identify the type of cells.
2. The method for intelligently identifying shaped components of urine sediment based on HAT-MixNet according to claim 1, characterized in that: In step (2), the data set is divided into a training set and a validation set in a ratio of 9:1, the batch size is set to 4, the number of training epochs is set to 70, the learning rate value range is set to 0 to 0.001, and the experimental attenuation coefficient is set to between 0.0001 and 0.
001.
3. The method for intelligently identifying shaped components of urine sediment based on HAT-MixNet according to claim 1, characterized in that: The downsample component of the downsample layer in the HAT-MixNet network is composed of a normalization layer Layer Norm and a Hat2d convolution layer module with a convolution kernel size of 2×2 and a stride of 2.
4. The method for intelligently identifying shaped components of urine sediment based on HAT-MixNet according to claim 1, characterized in that: The processing process of the HAT-MixNet block processing module in the HAT-MixNet network is: a. Apply a Depthwise Hat2d convolution, set the initial number of channels to 96, set three 7×7 convolution kernels in the Depthwise Hat2d convolution, and set the step size to 1. After the Depthwise Hat2d convolution, perform a layer normalization LayerNorm. b. Pass through a Hat2d convolutional layer module with a convolution kernel size of 1×1 and a stride of 1, and then pass through a GELU activation function; c. After passing through a Hat2d convolutional layer module with a convolution kernel size of 1×1 and a step size of 1, it goes through the scaling layer LayerScale and the regularization process Drop Path; d. Then superimpose the output feature image and the input feature image as a new output feature image.
Citation Information
Patent Citations
Urine visible component recognition method based on improved Alexnet model
CN110473166A
Methods of disease detection and characterization using computational analysis of urine raman spectra
US20210215610A1