A bone marrow cell image intelligent detection system and method

Cell features are extracted by combining Canny edge detection algorithm and local binary mode technology, and the distinction accuracy value of the CNN model is calculated through polynomial regression model, the sample number is dynamically adjusted and the weighted loss function is used, which solves the misclassification problem of leukemia cells and normal cells in the prior art, and improves diagnostic accuracy and robustness.

CN119762489BActive Publication Date: 2025-06-06SHANDONG UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510266775.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

The prior art has misclassification problems when distinguishing leukemia cells from normal cells, especially when the leukemia cell morphology is similar to that of normal cells or the sample diversity is insufficient, it is difficult to accurately distinguish CNN models.

Method used

By combining Canny's edge detection algorithm and local binary mode technology, the boundary ambiguity outliers and color difference drift values ​​of the cells were extracted as input features, and the distinction accuracy values ​​of the CNN model were calculated through the polynomial regression model. For inaccurately distinguishing results, the sample size is dynamically adjusted and the weighted loss function is used for classification.

Benefits of technology

The accuracy of CNN model in leukemia cell morphology recognition is improved, misclassification caused by insufficient sample diversity is reduced, the model's adaptability to difficult-to-distinguish cells is enhanced, and diagnostic accuracy and robustness are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119762489B_ABST
    Figure CN119762489B_ABST
Patent Text Reader

Abstract

The present invention discloses a bone marrow cell image intelligent detection system and method, and specifically relates to the technical field of image detection. By acquiring bone marrow cell images and annotating and preprocessing them, important features such as cell boundary fuzziness and color difference are extracted, and these features are used as inputs, a machine learning model is used to calculate the accuracy of a CNN model in distinguishing leukemia cell morphology. For inaccurate distinction results, the present invention further optimizes the model training process by dynamically adjusting the number of samples, especially for cell morphology and endoplasmic features at different detection stages, thereby effectively solving the problem of misclassification of leukemia cells and normal cells in the case of similar morphology, and improving the overall accuracy and robustness of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image detection, and in particular to a bone marrow cell image intelligent detection system and method thereof. Background Art

[0002] Intelligent detection of bone marrow cell images refers to the use of computer vision and deep learning technology to automatically analyze and identify bone marrow cell images, thereby assisting medical personnel in diagnosing diseases, especially in the early detection of blood diseases and cancer. By classifying, counting and extracting features of cells in bone marrow sample images, this technology can quickly identify cell types, morphological changes and other characteristics to help doctors assess the health status of patients. The advantages of intelligent detection are that it improves detection efficiency, reduces human errors, and can process large amounts of image data.

[0003] For example, researchers have developed methods based on convolutional neural networks (CNNs) to detect leukemia cells in the bone marrow. These systems can automatically identify abnormal cells in bone marrow images, such as leukemia cells or other blood disease cells, by training large-scale labeled data sets, thereby providing doctors with more accurate diagnostic support. This technology has been used in automated blood test equipment in some hospitals and research institutions.

[0004] The prior art has the following deficiencies:

[0005] Leukemia cells and normal cells may have significant differences in their morphology under a microscope, but sometimes these differences may not be obvious or may vary from patient to patient. Some types of leukemia cells, such as acute myeloid leukemia, may be morphologically similar to normal cells, making it difficult for the CNN model to distinguish. In addition, if the CNN model does not see enough diverse samples (including various variant morphologies) during training, it may misclassify leukemia cells when it encounters morphologically atypical cells. Summary of the invention

[0006] The purpose of the present invention is to provide a bone marrow cell image intelligent detection system and method thereof to solve the deficiencies in the background technology.

[0007] In order to achieve the above object, the present invention provides the following technical solution: a method for intelligent detection of bone marrow cell images, comprising the following steps:

[0008] S1: Obtain bone marrow cell image data, label the cells in each image into different categories, and preprocess the labeled image data;

[0009] S2: Calculate the outlier of the boundary blurriness of cells in the image through the Canny edge detection algorithm as the first input feature, and extract the color difference drift value of the cell endoplasm through the local binary pattern technique as the second input feature;

[0010] S3: Convert the first input feature and the second input feature into feature vectors, calculate the accuracy value of the CNN model for distinguishing the morphology of leukemia cells through a machine learning model, and divide the discrimination result of the CNN model into an accurate discrimination result and an inaccurate discrimination result according to the calculation result;

[0011] S4: For the inaccurate discrimination result, dynamically adjust the number of samples with different cell morphologies and endoplasm characteristics at different detection stages, and classify them in combination with the labels during training.

[0012] Preferably, in S1, the cell annotation method includes manual annotation, semi-automatic annotation, and automatic annotation based on deep learning.

[0013] Preferably, in S1, the image data preprocessing steps include image segmentation, denoising, data augmentation, and image normalization processing.

[0014] Preferably, in S2, calculating the outlier of the boundary blurriness of cells in the image through the Canny edge detection algorithm is specifically as follows:

[0015] Perform Gaussian filtering on the original image, where G(x,y) is the value of the Gaussian filter, and x and y are the coordinates in the image; the image after Gaussian filtering is obtained by convolution with the Gaussian filter: Calculate the gradients of the image in the horizontal and vertical directions using the Sobel operator, and obtain and , calculate the gradient magnitude and direction : ; ; Perform non-maximum suppression on the gradient map, only retain the pixel points with large gradient amplitudes. For each pixel, find the two neighboring pixels along the gradient direction, and retain the local maximum as the edge. Set the high threshold Th and the low threshold Tl. If the gradient intensity G(x,y)>Th, the pixel is a strong edge; if G(x,y)<Tl, the pixel is a non-edge; if Tl≤G(x,y)≤Th, the pixel is a weak edge;

[0016] Quantify the boundary blurriness by calculating the overlap degree between the edges detected by Canny and the cell contour, and use the intersection over union to measure the overlap degree between the edge and the cell contour as the outlier of the boundary blurriness , and the expression is: ; Where: A is the cell contour area, B is the edge area obtained by Canny edge detection, is the intersection area of ​​the contour and the edge, It is the union area of ​​the contour and the edge.

[0017] Preferably, the color difference drift value of the cell endoplasm is extracted by using the local binary pattern technology, specifically: the cell image is converted into a grayscale image, a window is selected around the target pixel for comparison, and the 8 neighboring pixels of the current pixel are set as , where the neighborhood pixels are arranged in a clockwise direction; for each neighborhood pixel , compare it with the central pixel and generate a binary value. If the neighborhood pixel is larger than the central pixel, the binary value is 1, and if the neighborhood pixel is smaller than the central pixel, the binary value is 0. Arrange the binary values ​​in clockwise order to form an 8-bit binary number. After calculating the binary value of each pixel in the image, generate a binary histogram of the entire image, select two different areas of the image, calculate the binary histograms H1 and H2 of the two areas respectively, and then calculate the difference between them through the chi-square distance, and use the calculation result as the color difference drift value.

[0018] Preferably, in S3, the first input feature and the second input feature are converted into a feature vector, and the accuracy value of the CNN model for distinguishing the leukemia cell morphology is calculated by a machine learning model, specifically:

[0019] The boundary fuzziness anomalies and color difference drift values ​​are converted into comprehensive feature vectors, which are used as inputs of the machine learning model. The machine learning model is trained, and the accuracy value of the CNN model in distinguishing the leukemia cell morphology is determined according to the model output results. The machine learning model is a polynomial regression model.

[0020] Preferably, the distinction result of the CNN model is divided into accuracy distinction result and inaccuracy distinction result according to the calculation result, specifically:

[0021] The accuracy value of the CNN model in distinguishing leukemia cell morphology is compared with the accuracy reference threshold value pre-set according to historical data. If the accuracy value of the CNN model in distinguishing leukemia cell morphology is greater than or equal to the pre-set accuracy reference threshold value, it means that the CNN model has high accuracy in distinguishing leukemia cell morphology, and the distinction result of the CNN model is classified as an accurate distinction result; if the accuracy value of the CNN model in distinguishing leukemia cell morphology is less than the pre-set accuracy reference threshold value, it means that the accuracy of the CNN model in distinguishing leukemia cell morphology is low, and the distinction result of the CNN model is classified as an inaccurate distinction result.

[0022] Preferably, in S4, for inaccurate distinction results, the number of samples with different cell morphology and endoplasmic characteristics at different detection stages is dynamically adjusted, and the misclassified samples are weighted by introducing a weighted loss function, and the weighted loss function calculation formula is: ;in, is the weight of sample i. If the prediction accuracy of the sample is low, increase its weight , if the prediction accuracy is high, reduce its weight , is the true label of sample i, is the predicted value of sample i.

[0023] The present invention also provides a bone marrow cell image intelligent detection system, including a data acquisition and annotation module, a feature extraction module, a machine learning and classification module, and a dynamic adjustment and recalibration module;

[0024] Data acquisition and labeling module: acquire bone marrow cell image data, label the cells in each image into different categories, and pre-process the labeled image data;

[0025] Feature extraction module: The Canny edge detection algorithm is used to calculate the abnormal value of the boundary fuzziness of the cells in the image as the first input feature, and the local binary pattern technology is used to extract the color difference drift value of the cell endoplasm as the second input feature;

[0026] Machine learning and classification module: converting the first input feature and the second input feature into a feature vector, calculating the accuracy value of the CNN model in distinguishing the leukemia cell morphology through the machine learning model, and dividing the distinction result of the CNN model into an accurate distinction result and an inaccurate distinction result according to the calculation result;

[0027] Dynamic adjustment and recalibration module: For inaccurate differentiation results, the number of samples with different cell morphology and endoplasmic characteristics at different detection stages is dynamically adjusted, and classified based on the labels during training.

[0028] In the above technical solution, the technical effects and advantages provided by the present invention are:

[0029] 1. The present invention combines the Canny edge detection algorithm and the local binary pattern technology to extract the cell boundary fuzziness anomaly and the color difference drift value as input features, thereby enhancing the expression ability of cell image features. Through the training of the polynomial regression model, the accuracy of the CNN model in leukemia cell morphology recognition is further improved, avoiding misclassification due to insufficient sample diversity.

[0030] 2. The present invention also performs effective weighted processing on misclassified samples by dynamically adjusting the number of samples at different detection stages and combining the introduction of weighted loss functions. In this way, the system can adjust the weights of samples according to their prediction accuracy, improving the model's adaptability to difficult-to-distinguish cells, thereby ensuring higher diagnostic accuracy and robustness, especially when facing cells with variant morphologies, enabling accurate classification, which has significant clinical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0032] Figure 1 The figure is a flow chart of the method of the present invention.

[0033] Figure 2 It is a system module diagram of the present invention. DETAILED DESCRIPTION

[0034] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0035] Example 1, please refer to Figure 1 As shown, the method for intelligent detection of bone marrow cell images described in this embodiment includes the following steps:

[0036] S1: Obtain bone marrow cell image data, label the cells in each image into different categories, and preprocess the labeled image data;

[0037] S2: The Canny edge detection algorithm is used to calculate the abnormal value of the boundary fuzziness of the cells in the image as the first input feature, and the local binary pattern technology is used to extract the color difference drift value of the cell endothelium as the second input feature;

[0038] S3: converting the first input feature and the second input feature into a feature vector, calculating the accuracy value of the CNN model in distinguishing the leukemia cell morphology through the machine learning model, and dividing the distinction result of the CNN model into an accurate distinction result and an inaccurate distinction result according to the calculation result;

[0039] S4: For the inaccurate discrimination results, the number of samples with different cell morphology and endoplasmic features at different detection stages was dynamically adjusted and classified in combination with the labels during training.

[0040] The acquisition of bone marrow cell image data usually requires acquisition through microscopy or other imaging technologies (such as digital pathology, flow cytometry, etc.). The following are some common acquisition techniques:

[0041] Use a high-resolution microscope to photograph bone marrow samples and obtain two-dimensional images of cells. Usually, fluorescence microscopy, phase contrast microscopy, or inverted microscopy is used to collect cell images. After staining in the bone marrow (such as Giemsa staining, Wright staining, etc.), different types of cells can be clearly distinguished. Usually, they are saved as high-resolution image files (such as PNG, JPEG, TIFF, etc.). These images may have different sizes, resolutions, and color modes (RGB, grayscale, etc.).

[0042] Use a digital pathology scanner to scan tissue slices and convert the slice images into digital format. This technology can obtain high-resolution tissue images, which are convenient for remote storage, transmission and analysis. The generated large image files are usually huge. Common data formats include SVS, NDPI, MRXS, etc., which are suitable for processing large-scale pathology image data sets.

[0043] High-throughput analysis of bone marrow cells is performed by flow cytometry, where cells are excited and imaged one by one by a laser beam in the flow cytometer. This technology can provide high-throughput data of cells, but it does not directly provide images. Instead, cell image data is generated through further processing of flow cytometric data. The output data of flow cytometry is usually quantitative data, but by combining it with other techniques (such as microscopic image fusion), cell images can be generated.

[0044] A very critical step in learning (especially supervised learning). It provides the "real labels" for the training of the CNN model - that is, the classification of different cells in each image (such as normal cells, leukemia cells, etc.). Common cell labeling methods include:

[0045] Manually annotate each cell in the image using professional medical image analysis software (such as ImageJ, QuPath, Fiji, etc.). These tools provide a variety of annotation methods, and you can draw polygons, circles, rectangles, and other areas on the image to annotate different types of cells. It is usually necessary to label each cell as a different category (for example: normal cells, leukemia cells, other abnormal cells, etc.). This task requires experienced pathologists to ensure the accuracy of the labels.

[0046] Use semi-automatic annotation tools based on traditional computer vision or deep learning methods. These tools can help analysts automatically detect cell outlines or regions, and then manually check and adjust the annotation results. Common tools include CellProfiler, Labelbox, etc. After automatic or semi-automatic annotation generates preliminary annotations, experts will check whether the cells are accurately labeled and correct mislabeling.

[0047] Use pre-trained deep learning models or use artificial intelligence for preliminary labeling. Common methods include using deep learning segmentation models such as U-Net and Mask R-CNN for cell segmentation and labeling. These models can quickly detect and label cell categories in a large number of images. After automatic labeling, pathologists need to conduct post-confirmation and correction to ensure accuracy.

[0048] Image data preprocessing is an important step to ensure that the CNN model can learn efficiently. Image preprocessing aims to improve image quality, reduce data noise, and adapt the data to the input requirements of CNN. Common image preprocessing steps are as follows:

[0049] Perform cell segmentation on the acquired image to extract each cell from the background or other cells. Common methods include threshold-based segmentation, edge detection algorithms (such as the Canny operator), and deep learning segmentation networks (such as U-Net, Mask R-CNN, etc.).

[0050] In order to enhance the generalization ability of CNN models, images are usually augmented with data, such as rotation, translation, scaling, flipping, etc., to simulate different cell states and shooting conditions. Commonly used data augmentation libraries include Keras, TensorFlow, PyTorch, etc.

[0051] Normalize the image and scale the pixel values ​​to a uniform range (usually [0, 1] or [-1, 1]) to improve the stability of the CNN training process. A common practice is to divide the pixel value of each channel of the RGB image by 255, or use mean-standard deviation normalization.

[0052] Use denoising techniques (such as median filtering, Gaussian filtering, bilateral filtering, etc.) to eliminate noise in the image, especially impurities in the dyeing process and image noise in the shooting process.

[0053] Scale the images to make all images of the same size (e.g. 256x256 or 512x512) to ensure that the images fed into the network are of the same size. You can use interpolation algorithms (e.g. bilinear interpolation, nearest neighbor interpolation, etc.) to achieve image resizing.

[0054] S2. In image processing and cell analysis, boundary blur outliers usually refer to areas in an image where the quality of the cell edges does not meet the expected standards, such as discontinuous edges, overly blurred edges, or overly rough edges. Through the Canny edge detection algorithm, the boundaries of cells can be effectively extracted, and by calculating the boundary blur outliers, the edge quality can be evaluated.

[0055] Perform Gaussian filtering on the original image to remove noise and smooth the image. The formula for the Gaussian filter is: ; where G(x,y) is the value of the Gaussian filter, σ is the standard deviation of the Gaussian distribution, which controls the smoothness of the filter, and x and y are the coordinates in the image; the image after Gaussian filtering is obtained by convolving with the Gaussian filter: ; in the formula, represents the pixel value of the image at position (x,y); use the Sobel operator to calculate the gradients of the image in the horizontal and vertical directions, respectively obtaining and , calculate the gradient magnitude and direction : ; ;

[0056] Perform non-maximum suppression on the gradient map, only retaining the pixels with large gradient magnitudes. For each pixel, find the two neighboring pixels along the gradient direction and retain the local maximum as the edge. Set a high threshold Th and a low threshold Tl, and determine whether each pixel belongs to the edge through these two thresholds.

[0057] If the gradient magnitude G(x,y)>Th, the pixel is a strong edge; if G(x,y)<Tl, the pixel is a non-edge; if Tl≤G(x,y)≤Th, the pixel is a weak edge.

[0058] Quantify the boundary blur by calculating the overlap between the edges detected by Canny and the cell contours. Use the intersection over union (IoU) to measure the overlap between the edge and the cell contour, and use it as the boundary blur outlier , and the expression is: ; where: A is the cell contour area, B is the edge area obtained by Canny edge detection, is the intersection area between the contour and the edge, is the union area between the contour and the edge. If the IoU value is low, it means that the edge and the cell contour do not match, indicating a higher boundary blur.

[0059] Extracting the color difference drift value of the cell's internal substance through the local binary pattern technology is a common technology in image processing. LBP is mainly used to describe the texture characteristics of the image. In cell analysis, it can be used to extract the color difference inside the cell and its change trend to help identify various properties of the cell.

[0060] Before applying LBP, it is usually necessary to preprocess the image. Convert the cell image to a grayscale image , for further processing. Then a smoothing process (such as Gaussian filtering) can be performed to remove noise.

[0061] Select a window (e.g. 3x3) around the target pixel for comparison. Set the current pixel to I(x,y) and its 8 neighboring pixels to , where the neighborhood pixels are arranged in a clockwise direction. For each neighborhood pixel (i=1,2,...,8), compare it with the central pixel, generate a binary value, if the neighboring pixel is larger than the central pixel, the binary value is 1, if the neighboring pixel is smaller than the central pixel, the binary value is 0; arrange the binary values ​​in clockwise order to form an 8-bit binary number. For example, suppose the binary number obtained is , then the decimal value corresponding to the binary number is the LBP value of the pixel. After calculating the LBP value of each pixel in the image, the LBP histogram of the entire image is generated. The histogram reflects the distribution of different textures in the image and can effectively extract the texture features of the cytoplasm.

[0062] After extracting the texture features of the cell image through the local binary pattern (LBP), the color difference drift value of the cell cytoplasm can be further calculated. This is usually measured by comparing the distribution of LBP values ​​in different regions. The color difference drift value reflects the degree of color (grayscale) change of the cytoplasm in space and describes the changes in the cytoplasm.

[0063] Calculate the color difference drift value. For example, select two different areas of the image, calculate the LBP histograms H1 and H2 of the two areas respectively, and then calculate the difference between them through the chi-square distance, and use the calculation result as the color difference drift value.

[0064] S3: The first input feature and the second input feature are converted into feature vectors, and the accuracy value of the CNN model for distinguishing the leukemia cell morphology is calculated through the machine learning model, specifically:

[0065] For example, the boundary fuzziness anomalies and color difference drift values ​​can be converted into comprehensive feature vectors, and the comprehensive feature vectors can be used as the input of the machine learning model. The machine learning model uses each group of comprehensive feature vectors to predict the accuracy value label of the CNN model for distinguishing the leukemia cell morphology as the prediction target, and takes minimizing the sum of the prediction errors of the accuracy value labels of all CNN models for distinguishing the leukemia cell morphology as the training target. The machine learning model is trained until the sum of the prediction errors converges and the model training is stopped. The accuracy value of the CNN model for distinguishing the leukemia cell morphology is determined according to the model output results, wherein the machine learning model is a polynomial regression model.

[0066] The accuracy value of the CNN model in distinguishing the morphology of leukemia cells is obtained by obtaining the corresponding function expression from the comprehensive feature vector training data of the trained machine learning model: ; In the formula, is the output function of the model, is the boundary fuzziness outlier, GH is the accuracy value of the CNN model in distinguishing the leukemia cell morphology, and w is the accuracy value of the CNN model in distinguishing the leukemia cell morphology.

[0067] According to the calculation results, the distinction results of the CNN model are divided into accuracy distinction results and inaccuracy distinction results, specifically:

[0068] The accuracy value of the CNN model in distinguishing leukemia cell morphology is compared with the accuracy reference threshold value pre-set according to historical data. If the accuracy value of the CNN model in distinguishing leukemia cell morphology is greater than or equal to the pre-set accuracy reference threshold value, it means that the CNN model has high accuracy in distinguishing leukemia cell morphology, and the distinction result of the CNN model is classified as an accurate distinction result; if the accuracy value of the CNN model in distinguishing leukemia cell morphology is less than the pre-set accuracy reference threshold value, it means that the accuracy of the CNN model in distinguishing leukemia cell morphology is low, and the distinction result of the CNN model is classified as an inaccurate distinction result.

[0069] S4: For the inaccurate discrimination results, the number of samples with different cell morphology and endoplasmic features at different detection stages was dynamically adjusted and classified in combination with the labels during training.

[0070] Different cell types (such as normal cells, leukemia cells, and other abnormal cells) may have different morphological characteristics at different detection stages. For example, the morphology of leukemia cells may be very similar to that of normal cells, making it difficult for the CNN model to accurately distinguish them. In this case, the number of samples in the training data can be dynamically adjusted based on the accuracy of the model's prediction to enhance the model's learning ability for samples that are difficult to distinguish.

[0071] Adjust sample weights based on prediction error: For samples with low prediction accuracy (i.e., inaccuracy in distinguishing results), increase the weights of these samples in training to strengthen their contribution to model training.

[0072] Dynamically increase or decrease samples based on feature distribution: During the training process, the feature distribution of each cell sample can be calculated, such as color difference drift value, boundary fuzziness, etc., and samples with large deviations or different features can be sampled incrementally (oversampling), or some samples can be removed (undersampling) to improve the model's ability to distinguish difficult samples.

[0073] In order to dynamically adjust the model according to the labels during training, we can introduce a weighted loss function to weight the misclassified samples, thereby guiding the model to pay more attention to these difficult-to-distinguish samples. The weighted loss function calculation formula is: ;in, is the weight of sample i. If the prediction accuracy of the sample is low, its weight is increased; if the prediction accuracy is high, its weight is decreased. is the true label of sample i, is the predicted value of sample i. The specific weight It can be determined by the misclassification rate (such as error rate, model uncertainty measure, etc.): ;in, It is the prediction error or uncertainty measure of the model for sample i, which can usually be estimated by the confidence of the model or the variance of the prediction results.

[0074] In combination with labels during training and dynamically adjusted samples, it is also necessary to adjust the label strategy in a timely manner during the training process to ensure that the model obtains support from more representative and difficult samples during training. The following classification methods can be used:

[0075] For samples with inaccurate distinction results, if their predicted categories do not match the true labels, the robustness of the model to such samples can be improved by adjusting the label weights (such as by smoothing the labels).

[0076] In addition to weighted loss functions and label updates, data enhancement techniques can also be combined to generate more difficult samples (such as transformation, rotation, noise addition, etc.), or the SMOTE algorithm (Synthetic Minority Over-sampling Technique) can be used to generate new samples to balance the imbalance between categories. Feature selection is performed based on the feature distribution of samples at different detection stages, retaining those features that contribute more to the classification results, and optimizing the features used in the model training process to reduce unnecessary noise.

[0077] In this application, for inaccurate differentiation results, by dynamically adjusting the number of samples with different cell morphology and endoplasmic characteristics at different times, the classification performance of the CNN model on difficult-to-distinguish samples can be effectively improved. By adjusting the weight of the training data, using a weighted loss function, combining data enhancement technology, and dynamic label updating, the model can be helped to learn difficult-to-classify cell morphologies more accurately, and ultimately optimize the accuracy of the classification results.

[0078] In this embodiment, the accuracy of the CNN model in distinguishing the leukemia cell morphology is improved through multiple steps. First, the bone marrow cell image data is obtained and the cells in each image are labeled by category, and the image is preprocessed at the same time. Then, the Canny edge detection algorithm is used to calculate the boundary fuzziness anomaly of the cells in the image as the first input feature, and the local binary pattern technology is combined to extract the color difference drift value of the cell endoplasm as the second input feature. Subsequently, these two features are converted into feature vectors, and the accuracy value of the CNN model for distinguishing the leukemia cell morphology is calculated by the machine learning model, and the distinction results of the CNN model are divided into accuracy distinction results and inaccuracy distinction results according to the prediction results. Finally, for the inaccuracy distinction results, the number of samples of cell morphology and endoplasmic features at different detection stages is dynamically adjusted, and the classification is combined with the labels during training to improve the classification accuracy of the model on difficult samples.

[0079] Example 2, please refer to Figure 2 As shown, the bone marrow cell image intelligent detection system described in this embodiment includes a data acquisition and annotation module, a feature extraction module, a machine learning and classification module, and a dynamic adjustment and recalibration module;

[0080] Data acquisition and labeling module: acquire bone marrow cell image data, label the cells in each image into different categories, and pre-process the labeled image data;

[0081] Feature extraction module: The Canny edge detection algorithm is used to calculate the abnormal value of the boundary fuzziness of the cells in the image as the first input feature, and the local binary pattern technology is used to extract the color difference drift value of the cell endoplasm as the second input feature;

[0082] Machine learning and classification module: converting the first input feature and the second input feature into a feature vector, calculating the accuracy value of the CNN model in distinguishing the leukemia cell morphology through the machine learning model, and dividing the distinction result of the CNN model into an accurate distinction result and an inaccurate distinction result according to the calculation result;

[0083] Dynamic adjustment and recalibration module: For inaccurate differentiation results, the number of samples with different cell morphology and endoplasmic characteristics at different detection stages is dynamically adjusted, and classified based on the labels during training.

[0084] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.

[0085] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0086] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0087] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A method for intelligent detection of bone marrow cell images, characterized in that: The following steps are involved: S1: Obtain bone marrow cell image data, label the cells in each image into different categories, and preprocess the labeled image data; S2: The Canny edge detection algorithm is used to calculate the abnormal value of the boundary fuzziness of the cells in the image as the first input feature, and the local binary pattern technology is used to extract the color difference drift value of the cell endoplasm as the second input feature, specifically: Perform Gaussian filtering on the original image. G(x, y) is the value of the Gaussian filter, where x and y are the coordinates in the image; the image after Gaussian filtering is obtained by convolving with the Gaussian filter: Calculate the gradients of the image in the horizontal and vertical directions using the Sobel operator, and obtain respectively and , calculate the gradient magnitude and direction : ; ; Perform non-maximum suppression on the gradient map, only retaining the pixels with large gradient magnitudes. For each pixel, find the two neighboring pixels along the gradient direction, and retain the local maximum as the edge. Set a high threshold Th and a low threshold Tl. If the gradient magnitude > Th, the pixel is a strong edge; if < Tl, the pixel is a non-edge; if Tl ≤ ≤ Th, the pixel is a weak edge; Quantify the boundary blur by calculating the overlap between the edges detected by Canny and the cell contours, and use the intersection over union to measure the overlap between the edges and the cell contours as the boundary blur outlier , and the expression is: ; where: A is the cell contour area, B is the edge area obtained by Canny edge detection, is the intersection area of the contour and the edge, is the union area of the contour and the edge; S3: converting the first input feature and the second input feature into a feature vector, calculating the accuracy value of the CNN model in distinguishing the leukemia cell morphology through the machine learning model, and dividing the distinction result of the CNN model into an accurate distinction result and an inaccurate distinction result according to the calculation result; S4: For the inaccurate discrimination results, the number of samples with different cell morphology and endoplasmic features at different detection stages was dynamically adjusted and classified in combination with the labels during training.

2. The method for intelligent detection of bone marrow cell images according to claim 1, characterized in that: In S1, the cell labeling method includes manual labeling, semi-automatic labeling and automatic labeling based on deep learning.

3. The method for intelligent detection of bone marrow cell images according to claim 2, characterized in that: In S1, the image data preprocessing step includes image segmentation, denoising, data enhancement and image normalization processing.

4. The method for intelligent detection of bone marrow cell images according to claim 3, characterized in that: The local binary pattern technique is used to extract the color difference drift value of the cell endoplasm. Specifically, the cell image is converted into a grayscale image, a window is selected around the target pixel for comparison, and the 8 neighboring pixels of the current pixel are set as , where the neighborhood pixels are arranged in a clockwise direction; for each neighborhood pixel , compare it with the central pixel and generate a binary value. If the neighborhood pixel is larger than the central pixel, the binary value is 1, and if the neighborhood pixel is smaller than the central pixel, the binary value is 0; the binary values ​​are arranged in clockwise order to form an 8-bit binary number; after calculating the binary value of each pixel in the image, a binary histogram of the entire image is generated, and two different areas of the image are selected, and the binary histograms H1 and H2 of the two areas are calculated respectively, and then the difference between them is calculated by the chi-square distance, and the calculation result is used as the color difference drift value.

5. The method for intelligent detection of bone marrow cell images according to claim 4, characterized in that: In S3, the first input feature and the second input feature are converted into feature vectors, and the accuracy value of the CNN model for distinguishing the leukemia cell morphology is calculated through the machine learning model, specifically: The boundary fuzziness anomalies and color difference drift values ​​are converted into comprehensive feature vectors, which are used as inputs of the machine learning model. The machine learning model is trained, and the accuracy value of the CNN model in distinguishing the leukemia cell morphology is determined according to the model output results. The machine learning model is a polynomial regression model.

6. The method for intelligent detection of bone marrow cell images according to claim 5, characterized in that: According to the calculation results, the distinction results of the CNN model are divided into accuracy distinction results and inaccuracy distinction results, specifically: The accuracy value of the CNN model in distinguishing the leukemia cell morphology is compared with the accuracy reference threshold value pre-set according to historical data. If the accuracy value of the CNN model in distinguishing the leukemia cell morphology is greater than or equal to the pre-set accuracy reference threshold value, it means that the CNN model has high accuracy in distinguishing the leukemia cell morphology. At this time, the distinction result of the CNN model is classified as an accuracy distinction result. If the accuracy value of the CNN model in distinguishing the leukemia cell morphology is less than a preset accuracy reference threshold, it means that the accuracy of the CNN model in distinguishing the leukemia cell morphology is low. At this time, the distinction result of the CNN model is classified as an inaccurate distinction result.

7. The method for intelligent detection of bone marrow cell images according to claim 6, characterized in that: In S4, for the inaccurate distinction results, the number of samples with different cell morphology and endoplasmic characteristics at different detection stages is dynamically adjusted. By introducing a weighted loss function, the misclassified samples are weighted. The weighted loss function calculation formula is: ;in, is the weight of sample i. If the prediction accuracy of the sample is low, increase its weight , if the prediction accuracy is high, reduce its weight , is the true label of sample i, is the predicted value of sample i.

8. A bone marrow cell image intelligent detection system, used to implement a bone marrow cell image intelligent detection method according to any one of claims 1 to 7, characterized in that: It includes data acquisition and annotation module, feature extraction module, machine learning and classification module, and dynamic adjustment and recalibration module; Data acquisition and labeling module: acquire bone marrow cell image data, label the cells in each image into different categories, and pre-process the labeled image data; Feature extraction module: The Canny edge detection algorithm is used to calculate the abnormal value of the boundary fuzziness of the cells in the image as the first input feature, and the local binary pattern technology is used to extract the color difference drift value of the cell endoplasm as the second input feature; Machine learning and classification module: converting the first input feature and the second input feature into a feature vector, calculating the accuracy value of the CNN model in distinguishing the leukemia cell morphology through the machine learning model, and dividing the distinction result of the CNN model into an accurate distinction result and an inaccurate distinction result according to the calculation result; Dynamic adjustment and recalibration module: For inaccurate differentiation results, the number of samples with different cell morphology and endoplasmic characteristics at different detection stages is dynamically adjusted, and classified based on the labels during training.

Citation Information

Patent Citations

  • Analysis method based on ASCUS cervical sample pathological result

    CN117809301A

  • Method, system and device for detecting security risk features of border tree and storage medium

    CN119131600A