Intelligent urinary sediment impurity microscopic image clustering method based on depth self-coding

By using a deep autoencoder and K-means algorithm to preprocess and extract features from urine sediment microscopic images, the problem of identifying and clustering impurity samples in urine microscopic images was solved, improving the accuracy of identification and the speed of clustering, and promoting the development of intelligent medical instruments.

CN120411575AActive Publication Date: 2025-08-01JILIN UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510892031.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-08-01
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently identify and cluster urine sediment impurities in urine micrographs, especially due to the complex morphology, varying sizes, unbalanced quantities, and overlap with valid samples, resulting in low identification accuracy and slow clustering speed.

Method used

A deep autoencoder-based intelligent clustering method for urine sediment impurity microscopic images was adopted, including data augmentation, normalization and standardization preprocessing. The optimal number of clusters was determined by calculating the DBI index, CH index and intra-cluster and inter-cluster distance ratio by adjusting the number of clusters. Feature extraction and clustering were performed by combining deep autoencoder and K-means algorithm.

Benefits of technology

It improves the accuracy and clustering efficiency of microscopic images of urine sediment impurities, enhances the detection performance of medical instruments, and provides a key foundation for the research and development of intelligent medical instruments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411575A_ABST
    Figure CN120411575A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of medical image processing, and provides a urinary sediment impurity microscopic image intelligent clustering method based on depth self-coding, which comprises the following steps: firstly, carrying out data enhancement, normalization and standardization preprocessing on urinary sediment visible component impurity sample microscopic images to improve the generalization of a model; inputting the data into the model, and determining the optimal clustering number by changing the clustering number k and comparing the CH index, the DBI index and the intra-cluster and inter-cluster distance ratio; and inputting the preprocessed image into a depth auto-encoder (DAE) for feature extraction, clustering extracted potential features by using a K-Means algorithm, and iteratively updating a cluster center until convergence. The method lays a key foundation for accurate recognition of subsequent massive effective samples, and has important significance in promoting research and development of intelligent medical instruments and improving the detection performance of the medical instruments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing, and particularly relates to an intelligent clustering method for microscopic images of urinary sediment impurities based on deep autoencoders. Background Art

[0002] With the increasing improvement of people's living standards, medical and health issues have become the focus of social attention, which has promoted the rapid growth of medical data. Medical examination is a key link in the medical diagnosis and treatment process, which can help people deeply understand the occurrence and development of diseases, and also provides an important basis for the design and optimization of disease prognosis treatment plans.

[0003] In the field of medical examination, the analysis of urinary formed elements is of great significance. It analyzes the components such as cells, casts, and crystals in urine, so as to provide key information for the diagnosis of urinary system and kidney diseases, the evaluation of disease severity, and the monitoring of treatment effects. However, there are a large number and diverse types of impurity samples in clinical samples of urinary formed elements, and the morphologies of some impurity samples are extremely similar to those of effective category samples. The existence of these impurities not only increases the complexity of sample analysis, but also brings great difficulties to the accurate identification of subsequent massive effective category samples. Therefore, how to achieve accurate and efficient identification of massive medical microscopic images and effectively apply them in medical inspection instruments has become a key issue affecting the inspection performance of medical instruments, and has also been a technical problem that medical instrument R & D enterprises urgently need to overcome.

[0004] Early medical cell and formed element detection methods mainly relied on microscopic examination, chemical detection, immunoassay, and automated analysis. With the continuous development of technology, medical samples can be detected by instruments after being processed. The instrument reads and analyzes the components in the sample and converts them into analyzable data. In terms of image processing, traditional methods such as edge detection, texture features, morphological filtering, and template matching were often used at first. These methods were often designed for specific tasks and belonged to shallow learning of image features.

[0005] However, the number of medical images collected during the inspection by current medical instruments is huge. The microscopic images of actual urinary sediment samples obtained clinically contain dozens of different types of formed elements, and the morphological differences of these formed elements are relatively large. Moreover, there are many impurities in urinary formed samples, and the number of impurities accounts for nearly 60% of the total sample volume. Some microscopic images of urinary sediment formed element impurity samples are such as Figure 1As shown, these impurities have complex characteristics such as complex morphology, varying sizes, unbalanced quantities, and partial overlap with the effective samples, presenting significant challenges for identification. In addition, the characteristics of some categories of cells are extremely similar, and the number of samples that can be obtained for a small number of categories is limited, making it difficult to accurately extract sample characteristics during clustering. At the same time, additional considerations need to be given to the problem of underfitting caused by a small sample size, as well as the decrease in the overall recognition accuracy resulting from the partial overlap between impurity samples and effective samples. Moreover, the significant differences in the sizes of different categories of samples make it difficult for traditional image recognition methods to achieve ideal recognition effects.

[0006] In recent years, artificial intelligence and deep learning technologies have demonstrated great potential and value in multiple fields. Coupled with the urgent need for the intelligent development of medical testing instruments, intelligent networks have gradually been applied to the field of medical cell analysis and recognition, bringing new development opportunities to this field. However, the diversity and complexity of urine formed samples have made the design and training of models more difficult, and existing clustering methods have problems such as low accuracy and slow clustering speed. For this reason, the present invention proposes an intelligent clustering method for urinary sediment impurity microscopic images based on deep autoencoders. Summary of the Invention

[0007] The purpose of the present invention is to provide an intelligent clustering method for urinary sediment impurity microscopic images based on deep autoencoders, aiming to solve the problems raised in the above background technology.

[0008] The purpose of the present invention is achieved through the following technical solutions: An intelligent clustering method for urinary sediment impurity microscopic images based on deep autoencoders, comprising the following steps: Step 1: Perform data augmentation, normalization, and standardization preprocessing on the microscopic images of urinary sediment formed component impurity samples; Step 2: Calculate performance indicators to determine the optimal number of clusters; By adjusting the number of clusters k , compare the DBI index, CH index, and intra-cluster and inter-cluster distance ratio to determine the optimal number of clusters; Step 3: The deep autoencoder intelligently obtains image features and performs K-means clustering; After determining the optimal number of clusters, use a deep clustering method based on deep autoencoders and K-means to train all impurity samples.

[0009] Furthermore, the specific steps of step 1 are as follows: Step 11: Training data processing, the specific operations are as follows: Data augmentation: For training data, use enhancement strategies such as random cropping, horizontal flipping, color jittering, and random grayscaling; Normalization: Convert the image into a PyTorch tensor Tensor, and at the same time normalize the pixel values [0, 255] to [0, 1]; Step 12: Test data processing; Perform random cropping and normalization; Step 13: Standardization; Standardize the normalized data by Equation 1 so that the pixel values of each channel follow a standard normal distribution: Equation 1: ; where is the pixel value after standardization, is the original pixel value, is the mean used for standardization, is the standard deviation used for standardization.

[0010] Furthermore, in the data augmentation step, randomly select a part of the image and resize it to a specified size by random cropping; flip the image horizontally with a probability of 50%; perform color jitter on the image with a probability of 80% to change brightness, contrast, saturation, and hue; convert the color image to a grayscale image with a probability of 20%.

[0011] Furthermore, in Step 2, the calculation formulas for the DBI index, CH index, and intra-cluster and inter-cluster distance ratio are as follows: Equation 2: ; where is the number of clusters, is the average distance within cluster , is the average distance within cluster , is the distance between cluster and cluster ; Equation 3: ; where is the number of data points in cluster , is the set of data points in cluster , x is the internal data point of cluster , is the centroid of cluster , is the centroid of the overall data set, is the total number of data points; Equation 4: ; where is the intra-cluster and inter-cluster distance ratio, is the cluster and the clusters the Euclidean distance between is the Euclidean distance from the data points inside the cluster to the centroid; represents the number of combinations between any two clusters when k the number of clusters is For the calculated different k values, comprehensively compare the DBI index, CH index, and intra-cluster and inter-cluster distance ratio, and finally determine the optimal number of clusters ; The process of comprehensive comparison is as follows: normalize the DBI index, CH index, and intra-cluster and inter-cluster distance ratio corresponding to different k values. Based on the characteristic that the smaller the DBI index value, the higher the clustering compactness, perform an inverse operation on it; draw a scatter plot of the relationship between the number of clusters k and each performance index, calculate the points where the absolute value of the slope in the curve of each performance index exceeds the preset threshold, merge and de-duplicate the slope mutation points of all indicators, and perform frequency statistics on each number of clusters. If the number of clusters exists in the set of slope mutation points of any performance index, the count value of this number of clusters is incremented by 1; finally, determine the number of clusters with the highest count value as the optimal number of clusters;

[0012] Furthermore, the specific steps of step 3 are as follows: Step 31: Input the preprocessed image with added noise into the DAE for feature extraction; In the encoder part of the DAE, first, the fully connected layer processes the input data, then the data is non-linearly transformed through the ReLU activation function, and then passes through another fully connected layer to map the input data to a low-dimensional latent space; in the decoder part of the DAE, the fully connected layer first receives the latent space data from the encoder, then performs non-linear conversion through the ReLU activation function, and then passes through another fully connected layer, and finally outputs through the Sigmoid activation function to achieve image reconstruction; the model training uses the mean square error loss function, selects the Adam optimizer, and optimizes the network parameters within 5 training epochs; Step 32: Obtain the latent feature vector; After training is completed, extract the encoder part of the DAE to obtain the representation of the input data in the latent space, that is, the latent feature vector; Step 33: Use the K-Means algorithm to cluster the extracted latent features; The K-Means method calculates the Euclidean distance between the samples and the cluster centers and iteratively updates the cluster centers until convergence; set the number of clusters as the optimal number of clusters and initialize with a random seed; Step 34: Perform visual analysis on the clustering results; Use the PCA method to reduce the high-dimensional features to two dimensions and draw a scatter plot to visually display the distribution of each cluster; different colors represent different clustering categories to evaluate the separability of the feature distribution. Step 35: Evaluate the clustering effect. Calculate the CH index and DBI index to measure the inter-cluster distance and intra-cluster compactness; custom calculate the inter-cluster distance and intra-cluster distance, and calculate the ratio of intra-cluster to inter-cluster distance Ratio to evaluate the overall quality of clustering.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: Based on the deep network and clustering model, the present invention proposes a new method that uses a deep autoencoder (DAE) for feature extraction and the K-means method for clustering. First, preprocess the data, including data augmentation, normalization, and standardization, to improve the generalization of the model; then input the data into the model, and by changing the number of clusters k , compare the CH index, DBI index, and the ratio of intra-cluster to inter-cluster distance to determine the optimal number of clusters; then input the preprocessed image into the DAE for feature extraction, and use the K-Means algorithm to cluster the extracted latent features, and iteratively update the cluster centers until convergence. This method lays a key foundation for the accurate recognition of subsequent massive effective samples and is of great significance in promoting the research and development of intelligent medical instruments and improving the detection performance of medical instruments. Description of the Drawings

[0014] Figure 1 It is a microscopic image of some urinary sediment formed component impurity samples.

[0015] Figure 2 It is a flowchart of image preprocessing.

[0016] Figure 3 It is a flowchart for determining the optimal number of clusters.

[0017] Figure 4 It is a flowchart for extracting features from impurity samples and performing clustering.

[0018] Figure 5 It is a flowchart of the method of the present invention.

[0019] Figure 6 It is a microscopic image of some urinary sediment formed component impurity samples after preprocessing; among them, (a) is the original cluster-like image, (b) is the cluster-like image enhanced by using the random horizontal flipping technique, and (c) is the cluster-like image enhanced by using the random cropping technique.

[0020] Figure 7 It is the number of clusters k Relationship curve graph with performance indicators.

[0021] Figure 8 The following are comparison diagrams of the PCA dimensionality reduction scatter plots of five clustering methods; (a) is the K-means PCA dimensionality reduction scatter plot, (b) is the SC PCA dimensionality reduction scatter plot, (c) is the AE PCA dimensionality reduction scatter plot, (d) is the DDC PCA dimensionality reduction scatter plot, and (e) is the DAE PCA dimensionality reduction scatter plot. DETAILED DESCRIPTION

[0022] In order to have a clearer understanding of the technical features, objectives and beneficial effects of the present invention, the technical solution of the present invention is now described in detail below, but it should not be understood as limiting the scope of implementation of the present invention.

[0023] The present invention provides a method for intelligent clustering of urine sediment impurity microscopic images based on deep autoencoding, the flow chart of which is as follows: Figure 5 As shown, the method includes the following steps: Step 1: Perform data augmentation, normalization, and standardization preprocessing on the microscopic images of urine sediment impurities. The goal is to reduce the influence of external factors such as lighting and contrast, allowing the model to focus on semantic features rather than brightness or color distribution. This also accelerates training and improves the convergence rate of gradient descent. The specific steps are as follows (see Figure 2 ): Step 11: Training data processing, the specific operations are as follows: Data augmentation: For training data, random cropping, horizontal flipping, color jittering, and random grayscale enhancement strategies are used. Through random cropping, a part of the image is randomly selected and resized to a specified size ( ) to increase data diversity and allow the model to learn objects of different scales; flip the image horizontally with a probability of 50% to enhance the model's adaptability to different perspectives; jitter the image color with a probability of 80% to change the brightness, contrast, saturation and hue to prevent the model from overfitting the color features; convert the color image to grayscale image with a probability of 20% to enable the model to extract information from the grayscale features.

[0024] Normalization: Convert the image to a PyTorch tensor to facilitate deep learning calculations in PyTorch. At the same time, the pixel value [0, 255] is normalized to [0, 1], making it more suitable for neural network training and gradient calculation.

[0025] Step 12: Test data processing; Perform random cropping and normalization (same as the random cropping and normalization operations in training data processing) to ensure the stability and consistency of the test data.

[0026] Step 13: Standardization; Normalize the data by Equation 1 to make the pixel values of each channel follow a standard normal distribution: Equation 1: ; where is the normalized pixel value, is the original pixel value, is the mean for normalization, is the standard deviation for normalization.

[0027] Step 2: Calculate performance metrics to determine the optimal number of clusters; the specific steps are as follows (see Figure 3 ): Traverse the number of clusters (process different numbers of clusters in sequence), and calculate the three performance metrics of the Davies - Bouldin index (DBI index), Calinski - Harabasz index (CH index), and intra - cluster to inter - cluster distance ratio respectively. The calculation formulas are as follows: Equation 2: ; where is the number of clusters, is the average distance within cluster , is the average distance within cluster , is the distance between cluster and cluster .

[0028] Equation 3: ; where is the number of clusters, is the number of data points in cluster , is the set of data points in cluster , x is the data point within cluster , is the centroid of cluster , is the centroid of the overall data set, is the total number of data points.

[0029] Equation 4: ; where is the intra - cluster to inter - cluster distance ratio, is the number of clusters, is the Euclidean distance between cluster and cluster , is the data point within cluster Euclidean distance from internal data points to the centroid; Indicates the number of clusters as k When it is, the combination number between any two categories.

[0030] The smaller the DBI index, the larger the CH index and the intra-cluster and inter-cluster distance ratio, which means higher clustering tightness and better clustering effect. For the calculated different k Values corresponding to the DBI index, CH index and intra-cluster and inter-cluster distance ratio are comprehensively compared, and finally the optimal number of clusters is determined .

[0031] The process of comprehensive comparison is as follows: Normalize the DBI index, CH index and intra-cluster and inter-cluster distance ratio corresponding to different k Values: Equation 5: ; Where Represents the minimum value in the vector , Represents the maximum value in the vector , Represents the value after normalization; Based on the characteristic that the smaller the DBI index value, the higher the clustering tightness, perform an inverse operation on it; Plot the scatter diagram of the relationship between the number of clusters k And each performance index, calculate the points with large slope changes (that is, the absolute value of the slope exceeds the mean value of the absolute value of the slope change) in each performance index curve, merge and de-duplicate the slope mutation points of all indicators, and perform frequency statistics on each number of clusters - if a certain number of clusters exists in the set of slope mutation points of a certain performance index, the count value of this number of clusters is incremented by 1; Finally, determine the number of clusters with the highest count value as the optimal number of clusters.

[0032] Step 3: The deep autoencoder (DAE) intelligently obtains image features and performs K-means clustering; After determining the optimal number of clusters, train all impurity samples using the deep clustering method based on DAE and K-means. The specific steps are as follows (see Figure 4 ): Step 31: Input the preprocessed image with added noise into the DAE for feature extraction; In the encoder part of the DAE, first, the fully connected layer processes the input data. Subsequently, the ReLU activation function is used to perform a non-linear transformation on the data. Then, after another fully connected layer, the input data is mapped to a low-dimensional latent space. In the decoder part of the DAE, the fully connected layer first receives the latent space data from the encoder, then performs a non-linear transformation through the ReLU activation function. After that, it goes through another fully connected layer, and finally, the output is obtained through the Sigmoid activation function to achieve the reconstruction of the image. This process enables the model to have the ability to denoise while learning the data features. The model is trained using the mean squared error (MSE) loss function and the Adam optimizer to optimize the network parameters within 5 training epochs.

[0033] Step 32: Obtain the latent feature vector; After training is completed, the encoder part of the DAE is extracted to obtain the representation of the input data in the latent space, that is, the latent feature vector. This vector is used as the input for subsequent clustering analysis, which can reduce the redundant information of high-dimensional data and improve the stability of clustering.

[0034] Step 33: Cluster the extracted latent features using the K-Means algorithm; The K-Means method calculates the Euclidean distance between the samples and the cluster centers and iteratively updates the cluster centers until convergence. The number of clusters is set to the optimal number of clusters and initialized with a random seed to ensure the reproducibility of the experiment.

[0035] Step 34: Conduct a visual analysis of the clustering results; The PCA method is used to reduce the high-dimensional features to two dimensions, and a scatter plot is drawn to visually display the distribution of each cluster. Different colors represent different clustering categories to evaluate the separability of the feature distribution.

[0036] Step 35: Evaluate the clustering effect; Calculate the CH index and DBI index to measure the inter-cluster distance and intra-cluster compactness. In addition, custom calculations of the inter-cluster distance and intra-cluster distance are performed, and the ratio Ratio of the intra-cluster and inter-cluster distances is calculated to evaluate the overall quality of the clustering.

[0037] The following describes the specific implementation of the present invention in detail in combination with specific embodiments. Embodiment

[0038] The experimental environment configuration adopted in this embodiment is as follows: The Python version is 3.11.4, the CUDA version is 12.3, the Pytorch version is 2.3.0, the graphics card model is RTX3050, and the video memory is 4GB.

[0039] Select 40,089 microscopic images of clinical urinary sediment formed component impurity samples collected in actuality. These images are from different patients in different hospitals, and each image contains only a single impurity sample or partially overlapping samples. First, preprocess the collected sample images, including data augmentation, normalization, and standardization. Figure 6 shows some microscopic images of urinary sediment formed component impurity samples after preprocessing; among them, (a) is the original cluster-like image, (b) is the cluster-like image enhanced using the random horizontal flipping technique, and (c) is the cluster-like image enhanced using the random cropping technique. It can be seen that the image in (b) is horizontally flipped, and the image in (c) crops out the upper left part, increasing the diversity of data samples. Then adjust the number of clusters. , and determine the optimal number of clusters by comparing the DBI index, CH index, and intra-cluster and inter-cluster distance ratio. . The specific data is shown in Table 1: Table 1 Relationship between the number of clusters k and each performance index

[0040] Normalize the relationship between each performance index. Considering that the DBI index should be as small as possible, so perform an inverse operation on it; plot the scatter diagram of the relationship between the number of clusters and each performance index (as shown in Figure 7 ), calculate the points with relatively large slope changes (that is, the absolute value of the slope exceeds the average value of the absolute value of the slope change) in each performance index. After merging and removing duplicates the slope mutation points of all indicators, perform frequency statistics on each number of clusters - if a certain number of clusters exists in the set of slope mutation points of a certain performance index, the count value of this number of clusters is incremented by 1; finally, find the number of clusters with the highest count value, take it as the optimal number of clusters and output. After calculation, the optimal number of clusters in this embodiment is 6. Determine the optimal number of clusters After that, train all impurity samples using the deep clustering method based on DAE and K-means. The specific operations are carried out according to Steps 31 to 35 of this method, and the number of clusters is set to 6 in Step 33.

[0041] To compare the clustering effects, the present invention also uses the traditional K-means method, Spectral Clustering (SC), Autoencoder (AE), and Deep discriminative clustering (DDC) to perform clustering on the microscopic images of the formed component impurity samples of clinical urinary sediment. To ensure the same comparison conditions, the traditional Kmeans method, SC, AE, and DDC all divide all secretion impurity samples into six categories, and the number of parameter training times is 10 times. The scatter plots of the clustering results of the five clustering algorithms are drawn by PCA dimensionality reduction, as Figure 8 shown, where the abscissa represents the first principal component after PCA dimensionality reduction, the ordinate represents the second principal component after PCA dimensionality reduction, and the color depth represents the clustering category. And the performance indicators of each clustering algorithm are calculated, and the results are shown in Table 2: Table 2 Performance indicators of five clustering algorithms

[0042] Judging from the performance indicator data, the DBI of the DAE algorithm is the smallest, and the CH index and Ratio are the largest. At the same time, observing Figure 8 the PCA dimensionality reduction scatter plots of each clustering algorithm, it can be found that although the scatter distributions of K-means and SC have a certain aggregation trend, the overall is relatively fuzzy, the discrimination between classes is not high, and it is difficult to clearly define different classes; the scatter distribution of AE is messy, and the data points do not have an obvious aggregated cluster structure; the scatter in DDC is too scattered to clearly see the cluster structure of clustering, and the clustering effect is not ideal; the scatter of DAE shows an obvious linear distribution. Compared with other figures, the data points have a certain distribution rule, and to a certain extent, different data regions can be distinguished, and the aggregation trend of the same-class data is relatively more obvious. Generally speaking, the clustering effect of the DAE method is the best, effectively proving the effectiveness of the DAE algorithm in realizing the clustering of urinary sediment impurities.

[0043] The above is only the preferred implementation manner of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicability of the patent.

Claims

1. An intelligent clustering method for microscopic images of urinary sediment impurities based on deep auto-encoding, characterized in that, It includes the following steps: Step 1: Perform preprocessing of data augmentation, normalization, and standardization on the microscopic images of urinary sediment formed component impurity samples; Step 2: Calculate performance indicators to determine the optimal number of clusters; By adjusting the number of clusters k , compare the DBI index, CH index, and intra-cluster and inter-cluster distance ratio to determine the optimal number of clusters; Step 3: The deep autoencoder intelligently extracts image features and performs K-means clustering; After determining the optimal number of clusters, use the deep clustering method based on the deep autoencoder and K-means to train all impurity samples.

2. The intelligent clustering method for microscopic images of urinary sediment impurities based on deep autoencoding according to claim 1, wherein, The specific steps of Step 1 are as follows: Step 11: Training data processing, and the specific operations are as follows: Data augmentation: For training data, use enhancement strategies of random cropping, horizontal flipping, color jittering, and random grayscaling; Normalization: Convert the image into a PyTorch tensor Tensor, and at the same time normalize the pixel values [0, 255] to [0, 1]; Step 12: Test data processing; Perform random cropping and normalization; Step 13: Standardization; Perform standardization processing on the normalized data by Equation 1 to make the pixel values of each channel follow the standard normal distribution: Formula 1: ; wherein is the normalized pixel value, is the original pixel value, is the mean value for normalization, is the standard deviation for normalization.

3. The intelligent clustering method for microscopic images of urinary sediment impurities based on deep auto-encoding according to claim 2, characterized in that, In the data augmentation step, through random cropping, randomly select a part of the image and adjust it to the specified size; flip the image horizontally with a probability of 50%; perform color jittering on the image with a probability of 80% to change brightness, contrast, saturation, and hue; convert the color image to a grayscale image with a probability of 20%.

4. The intelligent clustering method for microscopic images of urinary sediment impurities based on deep auto-encoding according to claim 1, characterized in that, In Step 2, the calculation formulas for the DBI index, CH index, and intra-cluster and inter-cluster distance ratio are as follows: Formula 2: ; wherein is the number of clusters, is the average distance within cluster ; is the average distance within cluster ; is the distance between cluster and cluster ; Formula 3: ; where is the number of clusters the number of data points, is the set of data points of the cluster ; x is the data point inside the cluster ; is the centroid of the cluster ; is the centroid of the overall data set, is the total number of data points; Formula 4: ; Among them is the intra-cluster and inter-cluster distance ratio, is the cluster and the cluster the Euclidean distance between them, is the Euclidean distance from the internal data points of the cluster to the centroid; represents the number of combinations between any two classes when the number of clusters is k ; Comprehensively compare the DBI index, CH index, and intra-cluster and inter-cluster distance ratio corresponding to different k values, and finally determine the optimal number of clusters ; The process of comprehensive comparison is as follows: Normalize the DBI index, CH index, and intra-cluster to inter-cluster distance ratio corresponding to different k values. Based on the characteristic that the smaller the value of the DBI index, the higher the clustering tightness, perform a negation operation on it; Plot the scatter diagram of the relationship between the number of clusters k and each performance index, calculate the points on the curve of each performance index where the absolute value of the slope exceeds the preset threshold. After merging and removing duplicates the slope mutation points of all indexes, perform frequency statistics for each number of clusters. If the number of clusters exists in the set of slope mutation points of any performance index, the count value of this number of clusters is incremented by 1; Finally, determine the number of clusters with the highest count value as the optimal number of clusters.

5. The intelligent clustering method for microscopic images of urinary sediment impurities based on deep autoencoding according to claim 1, characterized in that, The specific steps of Step 3 are as follows: Step 31: Input the preprocessed image with added noise into the DAE for feature extraction; In the encoder part of the DAE, first the fully connected layer processes the input data, then the data is non-linearly transformed through the ReLU activation function, and then passes through another fully connected layer to map the input data to a low-dimensional latent space; in the decoder part of the DAE, first the fully connected layer receives the latent space data from the encoder, then performs non-linear conversion through the ReLU activation function, then passes through another fully connected layer, and finally outputs through the Sigmoid activation function to realize the reconstruction of the image; the model training adopts the mean square error loss function, selects the Adam optimizer, and optimizes the network parameters within 5 training epochs; Step 32: Obtain the latent feature vector; After training is completed, extract the encoder part of the DAE to obtain the representation of the input data in the latent space, that is, the latent feature vector; Step 33: Use the K-Means algorithm to cluster the extracted latent features; The K-Means method calculates the Euclidean distance between the sample and the cluster center, and iteratively updates the cluster center until convergence; set the number of clusters to the optimal number of clusters and initialize with a random seed; Step 34: Perform visual analysis on the clustering results; Use the PCA method to reduce the high-dimensional features to two dimensions and draw a scatter plot to visually display the distribution of each cluster; different colors represent different clustering categories to evaluate the separability of the feature distribution; Step 35: Evaluate the clustering effect; Calculate the CH index and the DBI index to measure the inter-cluster distance and intra-cluster compactness; customize the calculation of the inter-cluster distance and intra-cluster distance, and calculate the ratio Ratio of the intra-cluster distance to the inter-cluster distance to evaluate the overall quality of clustering.

Citation Information

Patent Citations

  • Automatic identification system for urinary sediment visible components based on support vector machine

    CN101900737A

  • Method and device used for processing to-be-processed block of urine sediment image

    CN105095901A

  • Power consumption behavior portrait generation method and system based on DAE network features

    CN113191453A

  • Gynecological secretion impurity clustering method based on IUT-ResNet

    CN117726841A