Pathological recognition analysis method and system based on blood detection sample

By combining multi-scale morphological contour extraction and deep learning analysis with multi-level feature fusion, the problems of low cell segmentation accuracy and weak ability to capture subtle differences in blood cell pathology testing are solved, achieving highly accurate and automated pathology testing.

CN121998951APending Publication Date: 2026-05-08XINXIANG CENTER HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XINXIANG CENTER HOSPITAL
Filing Date
2026-01-28
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing methods for blood cell pathology testing suffer from low cell segmentation accuracy, weak ability to capture subtle differences, and incomplete pathological quantification, resulting in insufficient detection accuracy and automation.

Method used

By employing multi-scale morphological contour precise extraction, deep learning analysis of fine structural features, and multi-level feature fusion and quantitative modeling, combined with image segmentation, edge enhancement and boundary smoothing, a deep learning model is used to perform in-depth analysis of potential abnormal cell regions and generate quantitative pathological feature descriptors.

Benefits of technology

It improves the accuracy and automation of cellular pathological abnormality detection, generates more discriminative quantitative pathological feature descriptors, and realizes the accurate identification of subtle morphological abnormalities in cell populations and the systematic and quantitative assessment of the overall pathological state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998951A_ABST
    Figure CN121998951A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent medical treatment, and discloses a pathological recognition analysis method and system based on a blood detection sample. The method comprises the following steps: acquiring a blood smear microscopic image and extracting a cell contour; performing texture analysis on the interior of the contour to obtain morphological characteristics; judging a potential abnormal region based on the features; carrying out local enhancement on the abnormal region and extracting subtle difference features; based on the difference features, obtaining quantitative pathological descriptors through a deep learning model; identifying abnormal cells according to a matching result of the descriptors and the feature library; and comprehensively analyzing according to an identification result to obtain pathological quantitative indexes. Through a mode of combining local enhancement, a deep learning model and feature library matching, the problems of low blood smear abnormal cell recognition precision and poor efficiency are solved, and the accuracy and automation level of pathological analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent healthcare, and in particular to a method and system for pathological identification and analysis based on blood test samples. Background Technology

[0002] Blood cell morphology analysis is fundamental to clinical laboratory and pathological diagnosis, playing a crucial role in the screening, diagnosis, and efficacy evaluation of hematological diseases, infectious diseases, and malignant tumors. Traditionally, cell morphology observation and abnormality identification have relied on manual microscopic examination, a cumbersome, subjective, and inefficient process. With the development of digital pathology and artificial intelligence technologies, automated cell analysis methods based on image processing have become an important research direction for improving the objectivity, standardization, and scalability of pathological diagnosis.

[0003] Currently, existing methods in the field of image processing-based cellular pathological anomaly detection still have significant limitations. First, in cell segmentation and morphological contour extraction, due to issues such as cell adhesion, overlap, uneven staining, and complex background noise in blood smear images, commonly used methods such as threshold segmentation and edge detection are prone to boundary breaks, oversimplification of contours, or missegmentation, leading to inaccurate initial cell morphological contour extraction and directly affecting the accuracy of subsequent feature analysis. Second, in morphological feature analysis and anomaly judgment, most methods rely on single or superficial features (such as cell area and roundness), lacking the ability to capture subtle morphological differences such as the central pale-stained area of ​​red blood cells and the lobulation of white blood cell nuclei. Furthermore, the feature extraction and anomaly judgment processes are often fragmented, lacking a systematic fusion analysis framework from local details to the global image, making it difficult to quantify the distribution patterns and relationships between cell populations, resulting in low sensitivity in identifying potential heterogeneous lesions (such as myelodysplastic syndromes). Third, in terms of pathological quantification and model generalization, existing methods mostly use traditional machine learning classifiers or fixed threshold rules. When faced with clinical scenarios with diverse cell morphologies and a broad spectrum of lesions, the models have poor adaptability, and the quantitative indicators are not strongly correlated with clinical pathological significance, making it difficult to generate comprehensive quantitative reports with diagnostic value.

[0004] To address the above shortcomings, this application combines multi-scale morphological contour extraction, deep learning analysis of fine structural features, and multi-level feature fusion with quantitative modeling. This solves the technical problems of low cell segmentation accuracy, weak ability to capture subtle differences, and incomplete pathological quantification in existing methods, thereby improving the accuracy, automation, and clinical auxiliary diagnostic value of cytopathological abnormality detection. Summary of the Invention

[0005] This application provides a pathological identification and analysis method and system based on blood test samples, which solves the technical problems of low cell segmentation accuracy, weak ability to capture subtle differences, and incomplete pathological quantification in existing methods, and improves the accuracy, automation, and clinical auxiliary diagnostic value of cellular pathological abnormality detection.

[0006] In a first aspect, this application provides a method for pathological identification and analysis based on blood test samples, the method comprising:

[0007] Step S101: Obtain a blood smear microscopic image of the blood test sample, extract cell region boundary information from the blood smear microscopic image, and obtain preliminary cell morphology outline;

[0008] Step S102: Perform texture analysis on the internal region of the preliminary cell morphology outline to obtain cell morphology distribution characteristics;

[0009] Step S103: Perform a comprehensive analysis of the cell morphology distribution characteristics to determine whether there are any potential abnormal cell regions;

[0010] Step S104: Perform local image enhancement processing on the potential abnormal cell region and extract a set of subtle difference features;

[0011] Step S105: Based on the set of subtle differences in features, a deep learning model is used to perform in-depth analysis on the potential abnormal cell regions to obtain a quantitative pathological feature descriptor.

[0012] Step S106: Based on the degree of matching between the quantified pathological feature descriptor and the preset feature library, determine whether there is a pathological abnormal signal and output the abnormal cell identification result;

[0013] Step S107: Based on the abnormal cell identification results, perform comprehensive analysis on the blood smear microscopic images to obtain comprehensive pathological quantitative indicators.

[0014] Secondly, this application provides a pathological identification and analysis system based on blood test samples, used to implement the aforementioned pathological identification and analysis method based on blood test samples, the system comprising:

[0015] The image acquisition module is used to acquire a blood smear microscopic image of a blood test sample, extract cell region boundary information from the blood smear microscopic image, and obtain a preliminary cell morphology outline.

[0016] The texture analysis module is used to perform texture analysis on the internal region of the preliminary cell morphology outline to obtain cell morphology distribution characteristics.

[0017] An anomaly detection module is used to comprehensively analyze the cell morphology distribution characteristics to determine whether there are potential abnormal cell regions.

[0018] The local enhancement module is used to perform local image enhancement processing on the potential abnormal cell region and extract a set of subtle difference features.

[0019] The deep analysis module, based on the set of subtle differences in features, combines a deep learning model to perform deep analysis on the potentially abnormal cell regions and obtain a quantitative pathological feature descriptor.

[0020] The matching and judgment module is used to determine whether there is a pathological abnormal signal and output the abnormal cell identification result based on the degree of matching between the quantified pathological feature descriptor and the preset feature library.

[0021] The comprehensive evaluation module is used to perform comprehensive analysis on the blood smear microscopic images based on the abnormal cell identification results to obtain comprehensive pathological quantitative indicators.

[0022] This application proposes a method and system for pathological identification and analysis based on blood test samples, which solves the technical problems of low cell segmentation accuracy, weak ability to capture subtle differences, and incomplete pathological quantification in existing methods, thereby improving the accuracy, automation, and clinical auxiliary diagnostic value of cytopathological abnormality detection. Compared with the prior art, the beneficial effects of the technical solution of this application are at least as follows:

[0023] First, by combining image segmentation, edge enhancement, and boundary smoothing, continuous and complete preliminary cell morphological contours can be accurately extracted from blood smear images with complex backgrounds and cell adhesion, effectively improving the accuracy of cell region segmentation and the reliability of contour description.

[0024] Secondly, by adopting a strategy that combines multi-feature fusion analysis with preset anomaly judgment conditions, the system comprehensively analyzes the multi-dimensional morphological distribution characteristics of the central pale-stained area of ​​red blood cells and the number of lobes in the nucleus of white blood cells, thereby achieving accurate and automated identification of potentially abnormal cell regions and enhancing the detection sensitivity of subtle morphological abnormalities in cell populations.

[0025] Third, by introducing a deep learning model to perform in-depth analysis and feature weight optimization on subtle differences in potential abnormal regions (such as the lobulation angle of the neutrophils and the morphology of vacuoles inside the lightly stained areas), it is possible to extract and construct more discriminative quantitative pathological feature descriptors, thereby improving the model's ability to characterize complex pathological features and its generalization performance.

[0026] Fourth, based on the matching judgment between quantitative pathological feature descriptors and preset feature libraries, and combined with multi-feature fusion and weighted analysis of the overall image, comprehensive pathological quantitative indicators can be generated, realizing a systematic and quantitative assessment from local cell abnormalities to the global pathological state of the image, providing more comprehensive and objective data support for clinical diagnosis. Attached Figure Description

[0027] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is a flowchart illustrating the pathological identification and analysis method based on blood test samples in this application;

[0029] Figure 2 This is a comparison chart of ROC curves in this application;

[0030] Figure 3 This is a comparison chart of the recognition accuracy in this application;

[0031] Figure 4 This is a comparison chart of the anomaly detection sensitivity in this application;

[0032] Figure 5 This is a comparison chart of the overall performance scores in this application;

[0033] Figure 6 This is a schematic diagram of the structure of a pathological identification and analysis system based on blood test samples according to this application. Detailed Implementation

[0034] This application provides a method and system for pathological identification and analysis based on blood test samples. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0035] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the pathological identification and analysis method based on blood test samples in this application includes:

[0036] Step S101: Obtain a blood smear microscopic image of the blood test sample, extract cell region boundary information from the blood smear microscopic image, and obtain preliminary cell morphology outline.

[0037] In one specific embodiment, step S101 includes the following steps:

[0038] The blood smear was photographed in multiple regions and at multiple magnifications using a microscopic imaging device to obtain microscopic images of the blood smear.

[0039] Image processing algorithms are used to segment blood smear microscopic images and extract the boundary information of cell regions;

[0040] Based on the extracted boundary information, the Canny edge detection algorithm is used for edge enhancement and connection processing to generate a continuous and complete preliminary cell morphology outline.

[0041] Specifically, blood smears are photographed in multiple regions and at multiple magnifications using a microscopic imaging device to obtain microscopic images of the blood smears. The images cover the edges, center, and multiple randomly selected sampling areas of the smear. For example, the magnification can be set to three gradients: 10x, 40x, and 100x. At 10x magnification, the overall cell distribution image is acquired; at 40x magnification, the basic morphological details of the cells are captured; and at 100x magnification, the fine internal structure of the cells is focused. The data from different regions and at different magnifications form a multi-dimensional image set.

[0042] Furthermore, an image processing algorithm was used to segment the blood smear microscopic image to extract the boundary information of cell regions. The image processing algorithm used was a threshold segmentation algorithm, which separates cells from the background based on the difference in pixel gray values. First, the image gray-level histogram was calculated, and then the Oates thresholding method was used to determine the segmentation threshold. The Oates thresholding formula is as follows: ,in The inter-class variance corresponding to the threshold k. The percentage of background pixels after threshold k segmentation. Foreground pixel ratio The average grayscale value of the background pixels. The average grayscale value of the foreground pixels is used as the threshold value. The value of k that maximizes the inter-class variance is selected as the segmentation threshold. Pixels with grayscale values ​​higher than the threshold are marked as cell regions, and pixels with grayscale values ​​lower than the threshold are marked as background regions. This process initially separates cells from the background and extracts the boundary information of cell regions. This process effectively reduces the interference of background noise on boundary extraction and alleviates the problem of boundary blurring.

[0043] Based on the extracted boundary information, the Canny edge detection algorithm is used for edge enhancement and connection processing to generate continuous and complete preliminary cell morphology contours. The Canny edge detection algorithm sequentially performs four steps: Gaussian filtering, gradient calculation, non-maximum suppression, and double thresholding. Gaussian filtering uses a Gaussian kernel function with a standard deviation of 1.6 to perform convolution operations on the image, smoothing the image while preserving edge information and reducing the impact of noise on edge detection. Gradient calculation uses the Sobel operator to calculate the gradient values ​​in the horizontal and vertical directions respectively. and gradient magnitude gradient direction This determines the intensity and direction of the edge; during non-maximum suppression, the gradient magnitudes of adjacent pixels along the gradient direction are compared, and only the pixel with the largest gradient magnitude is retained as an edge candidate point, while redundant non-edge points are eliminated, thus refining the edge lines; dual thresholding is used to set a high threshold. and low threshold The gradient magnitude is greater than The pixels are determined to be strong edges, smaller than Pixels that are directly discarded, due to and If a pixel is connected to a strong edge, it is retained as an edge point; otherwise, it is discarded. This step strengthens and connects the edges, integrates the initially extracted discontinuous boundary information into continuous edge lines, solves the problem of difficult-to-distinguish boundaries caused by irregular shapes and cell overlap, and enables the morphological features of individual cells to be accurately delineated, ultimately generating a continuous and complete morphological outline for each red blood cell.

[0044] Step S102: Perform texture analysis on the internal region of the preliminary cell morphology outline to obtain cell morphology distribution characteristics.

[0045] In one specific embodiment, step S102 includes the following steps:

[0046] Texture analysis was performed on the internal region of the preliminary cell morphology outline to extract the shape distribution characteristics and color uniformity distribution characteristics of the lightly stained area in the center of the red blood cells within the outline.

[0047] After quantifying the shape distribution characteristics and color uniformity distribution characteristics, principal component analysis is applied to reduce the dimensionality and generate low-dimensional feature vectors.

[0048] Calculate the Euclidean distance between the low-dimensional feature vector and the preset standard feature vector, and quantify the degree of abnormality in cell morphology based on the Euclidean distance;

[0049] A contour tracking algorithm is used to identify the nuclear boundary of leukocytes. The number of nuclear lobes in each leukocyte is analyzed and statistically analyzed through connection components to generate the distribution characteristics of the number of nuclear lobes.

[0050] The degree of abnormality was weighted and fused with the distribution characteristics of the number of lobes in the nucleus to obtain the cell morphology distribution characteristics.

[0051] Specifically, texture analysis is performed on the internal region of the preliminary cell morphology outline. Relevant features are extracted using a gray-level co-occurrence matrix and a color histogram. Preferably, the gray-level co-occurrence matrix is ​​set with a distance parameter d = 1 pixel and angles θ = 0°, 45°, 90°, and 135°. The shape distribution characteristics of the central pale-stained region of the red blood cell are quantified by calculating the co-occurrence probability of pixel pairs in different directions within the matrix. The shape distribution characteristics include ellipticity and boundary smoothness. The ellipticity e is determined by fitting an elliptical model. The calculation yielded that, The pixel value of the semi-major axis of the ellipse. The values ​​represent the short semi-axis pixel values. Boundary smoothness is obtained by calculating the mean of contour curvature. Curvature calculation employs the boundary chain code tracking method, averaging the local curvature obtained by calculating the directional changes of adjacent pixels. Color uniformity distribution features are based on color histogram analysis, statistically analyzing the variance and mean of pixel colors within the light-colored areas. The variance reflects the dispersion of color distribution, while the mean characterizes the overall color level, forming a quantitative description of color uniformity distribution features. Addressing the problem that existing methods in the background technique struggle to capture subtle cellular features and easily miss potential lesion signals, this texture analysis process, through multi-parameter settings and precise calculations, achieves detailed extraction of key morphological features of red blood cells, solving the technical problem of incomplete feature capture in traditional analysis.

[0052] The extracted shape distribution features and color uniformity distribution features are quantized, converting parameters such as ellipticity, boundary smoothness, color variance, and color mean into numerical sequences. Ellipticity values ​​are mapped to [0,1], boundary smoothness is standardized to [0,1], and color variance and mean are normalized to [0,1] according to the image color channel range, forming a multidimensional original feature vector. Principal component analysis is then applied to the original feature vector for dimensionality reduction. First, the covariance matrix C is calculated. Where m is the sample size. For a single sample, the original feature vector. To find the mean of the original eigenvectors, solve for the eigenvalues ​​and corresponding eigenvectors of the covariance matrix. Sort the eigenvalues ​​in descending order and select the eigenvectors corresponding to the first n eigenvalues ​​to construct a projection matrix P. The selection of n satisfies that the cumulative variance contribution rate is ≥90%. Multiply the original eigenvectors by the projection matrix to obtain the low-dimensional eigenvectors. This achieves data dimensionality compression while retaining key feature information, reducing the complexity of subsequent calculations.

[0053] The Euclidean distance between the low-dimensional feature vector and the preset standard feature vector is calculated. The preset standard feature vector is derived from the feature statistics of a large number of healthy blood samples, obtained by performing the same feature extraction and dimensionality reduction process on the healthy samples. The Euclidean distance calculation formula is as follows: , Let j be the j-th element of the low-dimensional eigenvector. For the j-th element of the preset standard feature vector, a distance threshold is set based on the Euclidean distance to quantify the degree of abnormality in cell morphology. Preferably, ,when When the anomaly level is quantified as 0~0.3; when When the anomaly level is quantified as 0.3~0.7; when When the degree of abnormality is quantified, it is 0.7~1.0, forming an intuitive result of the degree of abnormality quantification, which solves the problem that traditional methods are difficult to quantify the degree of cell morphological abnormality.

[0054] A contour tracking algorithm is used to identify the nuclear boundary of leukocytes. Starting from the initial pixel in the leukocyte nuclear region, the algorithm traverses adjacent pixels clockwise, determining boundary points based on differences in pixel grayscale values. Preferably, a grayscale difference threshold of 15 is set; when the grayscale difference between adjacent pixels exceeds the threshold, it is considered a boundary point, continuing until the algorithm returns to the initial pixel to form a complete boundary. The number of lobes in the nucleus of each leukocyte is statistically analyzed using connected component analysis. The identified leukocyte nuclear boundary region is binarized, with a grayscale threshold T=100. Pixels above the threshold are marked as 1, and those below are marked as 0. A connected component labeling algorithm is applied to group adjacent pixels marked as 1 into the same connected component. Each connected component corresponds to one nucleus lobe. The number of connected components represents the number of lobes in a single leukocyte. After traversing all leukocytes, the mean and variance of the number of lobes are calculated to generate a distribution feature of the number of lobes. This process accurately captures the lobulation features of leukocyte nuclei, overcoming the deficiency in traditional analysis that neglects the morphological characteristics of leukocytes.

[0055] The degree of anomaly is weighted and fused with the distribution characteristics of the number of nuclear lobes. Preferably, the weight of the degree of anomaly is set. Weight of the mean in the distribution characteristics of the number of lobes in the nuclear lobe Weights of variance The fusion formula is:

[0056]

[0057] Where A is the quantification value of the degree of abnormality. This represents the average number of lobes in the nuclear lobes. The variance of the number of nuclear lobes. This represents the maximum variance of the number of nucleolobes predefined in a given set of values. This formula integrates the abnormality level associated with erythrocytes with the distribution characteristics of the number of nucleolobes in leukocytes to obtain cell morphology distribution characteristics. For example, in a blood sample, the Euclidean distance D between the low-dimensional feature vector and the predefined standard feature vector is 1.8, the abnormality level is quantified as 0.5, and the mean number of nucleolobes in leukocytes is... =3.2, variance =0.8, =2.0, obtained through weighted fusion calculation This cell morphology distribution feature comprehensively reflects the morphological status of red blood cells and white blood cells, providing a comprehensive basis for subsequent identification of potential abnormal cell regions. It effectively solves the technical problem of existing methods relying on a single cell morphology feature analysis, leading to insufficient accuracy in abnormal cell identification. For example, in the analysis of another blood sample, the Euclidean distance D between the low-dimensional feature vector and the preset standard feature vector was 2.6, the abnormality quantification was 0.8, and the average number of lobes in the white blood cell nuclei was... =4.5, variance =1.5, =2.0, calculated using the fusion formula. This value was significantly higher than that of previous samples. The corresponding image showed extremely poor uniformity of red blood cell color and an abnormally increased number of lobed nuclei of white blood cells. Further testing confirmed that the sample had a blood system lesion, verifying the effectiveness and reliability of the fusion method in describing cell morphology distribution characteristics.

[0058] Step S103: Conduct a comprehensive analysis of the cell morphology distribution characteristics to determine whether there are any potential abnormal cell regions.

[0059] In one specific embodiment, step S103 includes the following steps:

[0060] Extract the shape distribution characteristics, color uniformity distribution characteristics, and nuclear lobe number distribution characteristics from the cell morphology distribution characteristics;

[0061] The KL divergence between the shape distribution characteristics and the normal cell morphology distribution was calculated, the entropy value of the color uniformity distribution characteristics was calculated, and the Poisson distribution was used to fit the distribution characteristics of the number of nuclear lobes and the chi-square statistic was calculated.

[0062] A preset anomaly detection threshold is set. If the KL divergence is greater than the preset divergence threshold, the entropy value is less than the preset entropy value threshold, and the chi-square statistic is greater than the preset chi-square threshold, then the preset anomaly detection condition is met.

[0063] The regions where cells that meet the preset abnormality judgment conditions are located are identified as potential abnormal cell regions.

[0064] Specifically, this study extracts shape distribution features, color uniformity distribution features, and leukocyte lobe number distribution features from cell morphology distribution characteristics. Shape distribution features include quantitative data on ellipticity and boundary smoothness; color uniformity distribution features are represented by color variance and color mean; and leukocyte lobe number distribution features are reflected by the mean and variance of the number of leukocyte lobes. These three types of features reflect cell state from three dimensions: cell morphology and structure, color appearance, and leukocyte lobule distribution, providing multi-dimensional data support for subsequent anomaly identification. Addressing the problem that existing methods in the background technology struggle to handle complex cell morphological changes and insufficient capture of subtle features, leading to inaccurate abnormal cell identification, this study achieves a comprehensive consideration of abnormal cell states through multi-feature integrated analysis, overcoming the limitations of single-feature analysis.

[0065] The KL divergence between the shape distribution characteristics and the normal cell morphology distribution was calculated. The normal cell morphology distribution was established based on the statistical characteristics of the shape distribution of a large number of healthy blood samples. The KL divergence is used to measure the degree of difference between the two distributions, and the calculation formula is as follows: Where log represents the logarithmic function with any positive base, The probability distribution of the current shape distribution characteristics. The probability distribution of normal cell shape is calculated by iterating through all values ​​of both distributions and summing the products of the logarithmic ratios of the corresponding probabilities and the current distribution probability. The larger this value, the more significant the difference between the current shape distribution and the normal distribution. Entropy is then calculated for the color uniformity distribution characteristics. Entropy reflects the degree of disorder in the color distribution, and the calculation formula is as follows: ,in, To determine the probability of each color value in the color uniformity distribution characteristic, the entropy value is calculated by statistically analyzing the probability distributions corresponding to the color variance and color mean, and then substituting these probabilities into the formula. A smaller entropy value indicates a more uniform color distribution, while a larger entropy value indicates a more chaotic color distribution. For the distribution characteristic of the number of kernicterus segments, a Poisson distribution is fitted and the chi-square statistic is calculated. First, it is assumed that the number of kernicterus segments follows a Poisson distribution. ,in The parameters of the Poisson distribution are obtained using the maximum likelihood estimation method. , , This refers to the number of white blood cells in the sample. The number of nucleolobes in a single white blood cell, based on the solution obtained. A Poisson distribution model was constructed to determine the probability values ​​corresponding to different numbers of nuclear lobes. The actual distribution of nuclear lobes was then divided into several groups, each covering a continuous interval of nuclear lobe numbers. For example, groups with 1-2 nuclear lobes were designated as the first group, 3-4 as the second group, and 5 or more as the third group. The observation frequency of each group was then counted. The expected frequency of each group is calculated using the Poisson distribution model. Finally, the chi-square statistic is used to test the goodness of fit between the actual distribution and the Poisson distribution model. The formula for calculating the chi-square statistic is as follows: The larger the chi-square statistic, the worse the fit between the actual distribution and the model, and the more abnormal the distribution of the number of lobes in the nucleus.

[0066] Preset anomaly judgment thresholds are established, combining clinical data and statistical results from a large number of samples. Preferably, the preset divergence threshold is set to 1.2, the preset entropy threshold to 2.5, and the preset chi-square threshold to 10.83. If the KL divergence is greater than the preset divergence threshold, it indicates that the current cell shape distribution differs significantly from normal cells, indicating morphological abnormalities. If the entropy value is less than the preset entropy threshold, it indicates that the cell color distribution is too uniform or there are abnormally uniform areas, which does not conform to the normal cell color distribution pattern. If the chi-square statistic is greater than the preset chi-square threshold, it means that the distribution of lobule number deviates significantly from the normal Poisson distribution, indicating an abnormality in the number of lobules. Only when all three conditions are met simultaneously is the preset anomaly judgment condition satisfied. This multi-condition joint judgment method effectively reduces false abnormal signals caused by misjudgment of a single feature, improves the accuracy of anomaly judgment, and solves the technical problems of single anomaly judgment criteria and high misjudgment rate in traditional methods.

[0067] The regions containing cells that meet preset anomaly criteria are identified as potential abnormal cell regions. By locating the coordinates of these cells in the blood smear microscopic image, a rectangular region extending 5 pixels horizontally and vertically outward from the center of each cell is delineated as the potential abnormal cell region. Basic information such as the boundary coordinates and the number of cells contained within this region is recorded to define the scope for subsequent extraction of subtle difference features. For example, in a blood sample, the KL divergence between the shape distribution feature and the normal cell morphology distribution is 1.5, greater than the preset divergence threshold of 1.2; the entropy value of the color uniformity distribution feature is 2.0, less than the preset entropy threshold of 2.5; and the chi-square statistic of the lobule number distribution feature is 12.3, greater than the preset chi-square threshold of 10.83. These conditions meet the preset anomaly criteria, and the coordinate range of the cell is defined as such. to The rectangular region is identified as a potential abnormal cell region. For example, in the analysis of another blood sample, the KL divergence of cells in a certain region is 1.1, which does not reach the preset divergence threshold. Although the entropy value and chi-square statistic both meet the corresponding threshold conditions, it is still not identified as a potential abnormal cell region. This strict judgment logic ensures the accuracy of potential abnormal cell region identification, provides a reliable target region for subsequent in-depth analysis, and effectively solves the problems of ambiguous localization and inaccurate range of potential abnormal regions in existing methods.

[0068] Step S104: Perform local image enhancement processing on the potentially abnormal cell region and extract a set of subtle difference features.

[0069] In one specific embodiment, performing step S104 includes the following steps:

[0070] An adaptive histogram equalization method is used to perform local image enhancement processing on potentially abnormal cell regions to highlight cell detail features;

[0071] Based on the enhanced image of potential abnormal cell regions, the boundaries of the central pale-stained area and the cytoplasmic boundary of red blood cells are identified, and the vertical distance from each boundary point to the adjacent boundary is calculated to form a width sequence.

[0072] Statistical analysis was performed on the width sequences to calculate the mean width and standard deviation, and a distribution histogram was constructed to obtain the width distribution characteristics of the central pale-stained region and the cytoplasmic transition region of erythrocytes.

[0073] Denoise the leukocyte nucleoid regions in potentially abnormal cell regions and label each nucleoid;

[0074] Locate the nuclear lobe junctions and calculate their positions, curvatures, and junction angle distributions to generate morphological features of leukocyte nuclear lobe junctions;

[0075] The width distribution features and the morphological features of the lobes connecting the core and lobes are spliced ​​and integrated to form a set of subtle difference features.

[0076] Specifically, an adaptive histogram equalization method is used to perform local image enhancement processing on potential abnormal cell regions. This method divides the potential abnormal cell regions into multiple 8×8 pixel sub-blocks, calculates the histogram for each sub-block independently, and performs equalization operations. By limiting the contrast threshold to 2.0, noise amplification caused by excessive local enhancement is avoided, while improving the clarity of cell edges, textures, and other detailed features. Addressing the problems of insufficient capture of subtle cell features and low anomaly recognition accuracy due to image noise interference in background techniques, this local enhancement processing can accurately highlight differences in the internal structure of cells, laying the foundation for subsequent subtle feature extraction and overcoming the technical shortcomings of traditional global enhancement in taking local details into account.

[0077] Based on the enhanced image of potential abnormal cell regions, the Sobel edge detection operator is used to identify the boundaries of the central pale-stained area and the cytoplasmic boundary of erythrocytes. The Sobel operator sets 3×3 convolution kernels in both the horizontal and vertical directions. For example, the horizontal convolution kernel is [[1,2,1],[0,0,0],[-1,-2,-1]], and the vertical convolution kernel is [[1,0,-1],[2,0,-2],[1,0,-1]]. By calculating the pixel gradient magnitude and gradient direction, preferably setting the gradient magnitude threshold to 30, pixels with gradient magnitudes greater than the threshold are marked as boundary points, thus outlining the continuous boundaries of the pale-stained area and the cytoplasm. The vertical distance from each boundary point to the adjacent boundary is calculated using a distance transformation algorithm. The distance transformation uses Euclidean distance, traversing all pixels on the two boundaries to solve for the shortest vertical distance between each pair, forming a width sequence. Each value in this sequence corresponds to the width value at a certain position in the transition area between the pale-stained area and the cytoplasm, directly reflecting the morphological changes of the transition area. Statistical analysis was performed on the width sequence to calculate the mean width and standard deviation. The width sequence was then divided into 10 intervals to construct a distribution histogram. The frequency of width values ​​in each interval was statistically analyzed to form the width distribution characteristics between the central pale-stained region and the cytoplasmic transition zone of erythrocytes. This feature comprehensively quantifies the overall level, dispersion, and distribution pattern of the transition zone width through statistics and histograms, solving the problem of incomplete feature description caused by traditional methods relying only on a single width value.

[0078] Denoising of the leukocyte nucleoid region within the potentially abnormal cell region is performed using morphological opening operations with 3×3 rectangular structuring elements. First, an erosion operation is performed on the nucleoid region image to remove minute noise points, followed by a dilation operation to restore the main shape of the nucleoid. Both erosion and dilation operations involve traversing image pixels and adjusting pixel values ​​based on the matching between the structuring element and the local image region. Each nucleoid is labeled using a connected component labeling algorithm. Preferably, a pixel grayscale threshold of 120 is set, and pixels above this threshold are labeled as foreground pixels. The foreground pixels are traversed, and adjacent foreground pixels are divided into the same connected component. Each connected component corresponds to a nucleoid and is assigned a unique label value, achieving precise separation and labeling of the nucleoids and providing clear nucleoid region data for subsequent connectivity point analysis.

[0079] The connection points between nuclei and lobes are located and relevant parameters are calculated. The boundary coordinates of each nucleus and lobe are obtained using a contour tracking algorithm. The intersection points of different nucleus and lobe boundaries are found by traversing these boundary coordinates. When the distance between pixels on two nucleus and lobe boundaries is less than 2 pixels, the point is determined to be a connection point, and its two-dimensional coordinates (x, y) are recorded. The curvature of the connection point is calculated using a curve fitting method. A quadratic polynomial fitting is performed on 10 adjacent boundary points around the connection point, and the fitting equation is: By solving the first derivative and second derivative Using the curvature formula Calculate the curvature value at the connection point. Calculate the connection angle using the vector dot product, and construct unit vectors for the extension directions of the core leaf, with the connection point as the origin. and Connection angle By statistically analyzing the curvature and connection angle of all connection points, morphological features of leukocyte nucleolobular connections are generated. These features accurately capture structural anomalies in nucleolobular connections, compensating for the shortcomings of traditional analysis in not paying enough attention to the details of nucleolobular connections.

[0080] The width distribution features and the morphological features of the lobule connections are combined and integrated to form a set of subtle difference features. The width distribution features are represented as frequency vectors of average width, standard deviation, and distribution histogram, while the morphological features of the lobule connections are represented as connection point coordinates, curvature sequences, and connection angle sequences. The two types of features are converted into numerical vectors of a unified dimension and then concatenated in sequence to form a set of subtle difference features. For example, in a region of potentially abnormal cells, the average width of the width sequence is 3.2 pixels, the standard deviation is 0.5 pixels, the frequency vector of the 10 intervals of the distribution histogram is [2,5,8,12,15,13,9,6,3,1], the coordinates of the nucleolobular junctions are (50,60) and (75,82), the curvature sequence is [0.12, 0.18], and the junction angle sequence is [45°, 60°]. The set of subtle difference features formed after splicing contains all the above numerical information, comprehensively reflecting the subtle morphological differences between red blood cells and white blood cells, providing rich feature support for subsequent in-depth analysis, and effectively solving the technical problem of existing methods having a single feature dimension and difficulty in distinguishing subtle pathological changes.

[0081] Step S105: Based on the set of subtle differences in features, and combined with a deep learning model, perform in-depth analysis on potential abnormal cell regions to obtain quantitative pathological feature descriptors.

[0082] In one specific embodiment, step S105 includes the following steps:

[0083] The set of subtle difference features is organized into a standardized vector form and input into a pre-trained deep learning model;

[0084] The potential abnormal cell region is divided into multiple sub-region images focusing on the local structure of cells by a deep learning model. Edge and texture information is extracted by a convolutional neural network and then pooled to reduce dimensionality, generating a preliminary feature map.

[0085] The distribution characteristics of the lobulation angle of leukocyte nuclei and the morphological characteristics of vacuoles inside the central pale-stained area of ​​erythrocytes were extracted from the preliminary feature map.

[0086] The distribution characteristics of the lobulation angle of the nuclear lobe and the morphological characteristics of the vacuoles inside the lightly stained area were spliced ​​and integrated to form an initial quantitative pathological feature descriptor.

[0087] Principal component analysis was applied to the initial quantitative pathological feature descriptor to extract the main dimensions, the variance contribution rate of each dimension was calculated, and the feature weight distribution was generated.

[0088] The components of the initial quantified pathological feature descriptor are weighted and optimized according to the feature weight distribution to obtain the final quantified pathological feature descriptor.

[0089] Specifically, the set of subtle morphological features is organized into a standardized vector form. This set includes the width distribution characteristics of the central pale-stained region and the cytoplasmic transition zone of erythrocytes (mean width, standard deviation, and frequency vector of the distribution histogram) and the morphological characteristics of the nuclear lobe connections of leukocytes (connection point coordinates, curvature sequence, and connection angle sequence). These two types of features reflect subtle morphological differences in cells from two core dimensions: the structure of the erythrocyte transition zone and the nuclear lobe connection status of leukocytes, providing a comprehensive feature foundation for in-depth analysis. The standardization process uses Z-score standardization to transform all feature components to a uniform scale with a mean of 0 and a variance of 1. This avoids bias in deep learning model training or feature extraction due to excessive differences in the numerical range between feature dimensions, ensuring the consistency and comparability of the input data.

[0090] The deep learning model employs an architecture combining convolutional neural networks and fully connected layers. The training process uses labeled abnormal cell samples from blood smears as training data. These samples cover potential abnormal cell regions corresponding to various blood diseases, including leukemia, iron deficiency anemia, and megaloblastic anemia. Each sample is labeled with key pathological features such as nucleolobular lobulation angle and vacuolar morphology. Before training, the training and validation sets are divided in an 8:2 ratio. The Adam optimizer is used, with an initial learning rate of 0.001. When the validation set loss does not decrease for five consecutive epochs, the learning rate is reduced to 1 / 10 of its original value. The batch size is set to 64, and the number of iterations is 80 epochs. The mean squared error loss function is used. Backpropagation is used to adjust the convolutional kernel weights and fully connected layer parameters, ensuring that the model's feature extraction accuracy on the validation set remains stable above 92%. After training, the model is adapted to meet the requirements for extracting subtle pathological features of blood cells. To address the issues of weak integration between deep learning models and hematological pathology scenarios and poor ability to capture subtle morphological features in the background technology, this specialized training process enables the model to accurately adapt to business scenarios, solving the technical defects of traditional models such as poor generalization ability and difficulty in focusing on key pathological features.

[0091] The deep learning model divides potential abnormal cell regions into multiple sub-region images focusing on local cell structures. For example, the sub-region size is set to 32×32 pixels, and the entire potential abnormal region is traversed with a sliding stride of 16 pixels. This ensures that the sub-regions can focus on single cell structures (such as pale red blood cell areas or white blood cell nuclei) while also covering all subtle structures within the region, avoiding the omission of key features. The convolutional neural network contains 6 convolutional layers, 4 pooling layers, and 3 fully connected layers. All convolutional layers use 3×3 convolutional kernels with ReLU activation. The number of output channels for the first to sixth convolutional layers are 64, 128, 256, 256, 512, and 512, respectively. The pooling layers use 2×2 max pooling with a stride of 2. Convolutional operations capture local features such as edges and textures in the sub-region images. The pooling layers reduce feature dimensionality while retaining key information. The fully connected layers map high-dimensional convolutional features into one-dimensional feature vectors, ultimately generating a preliminary feature map. Each channel of the feature map corresponds to a specific pathological feature, and the pixel value represents the intensity of that feature.

[0092] The angular distribution features of leukocyte nuclei and the morphological features of vacuoles within the central pale-stained region of erythrocytes were extracted from the preliminary feature map. When extracting the angular distribution features of leukocyte nuclei, the Hough line detection algorithm was used to identify the linear structures of the nuclei in the feature map, the angle between adjacent linear structures was calculated, the angle distribution was statistically analyzed, and a 10-interval angle histogram was constructed. The frequency vector of the histogram was used as the angular distribution feature. When extracting the vacuole morphological features, an adaptive threshold segmentation algorithm (the threshold is dynamically adjusted based on the local pixel grayscale mean) was used to segment the vacuole region from the feature map, and the roundness of the vacuoles was calculated. S is the area of ​​the cavitation pixel, L is the perimeter of the cavitation, ellipticity, and area ratio. , The area of ​​the cavitation bubble. (This refers to the total area of ​​the lightly stained region). These parameters are integrated into multidimensional vacuolar morphological features. These two types of features precisely focus on the fine structures within the cell that are highly correlated with pathological abnormalities, making up for the shortcomings of traditional methods in mining deep cellular morphological features.

[0093] The distribution features of lobular segmentation angles (10-dimensional frequency vector) and the morphological features of vacuoles within the pale-stained areas (3-dimensional parameter vector) are sequentially concatenated to form a 13-dimensional initial quantitative pathological feature descriptor. Each element in the vector corresponds to a specific pathologically relevant feature parameter, directly linking the intrinsic relationship between abnormal cell morphology and pathological changes. Principal component analysis is applied to the initial quantitative pathological feature descriptor to extract the main dimensions. First, the covariance matrix H of the descriptor vector is calculated. , For the sample size, For a single sample, the initial descriptor vector. Given the mean vector of all samples, solve for the eigenvalues ​​of the covariance matrix. The eigenvectors are sorted by eigenvalue from largest to smallest. The eigenvectors corresponding to the top Q eigenvalues ​​(cumulative variance contribution rate ≥ 95%) are selected to construct a projection matrix. The initial descriptor vectors are multiplied by the projection matrix to complete dimensionality reduction. At the same time, the variance contribution rate of each principal component is calculated. The feature weight distribution is generated, and the weight values ​​intuitively reflect the degree of contribution of the corresponding principal components to the judgment of pathological abnormalities.

[0094] The components of the initial quantified pathological feature descriptor are weighted and optimized according to the feature weight distribution. The weighting formula is as follows: , The variance contribution rate of the i-th principal component is also known as the weight of the feature component. The standardized value of the i-th feature component is used to strengthen the representation of key features for pathological judgment and weaken irrelevant interference features, thus obtaining the final quantitative pathological feature descriptor. For example, in the initial descriptor of a potential abnormal cell region, the first principal component variance contribution rate of the lobule angular distribution is 65%, and the principal component variance contribution rate corresponding to vacuolar roundness is 20%. After weighted optimization, the weights of these two components in the final descriptor are 0.65 and 0.20, respectively, and the sum of the weights of the remaining components is 0.15, enabling the descriptor to accurately focus on the core pathological features. For example, in the analysis of potential abnormal regions in iron deficiency anemia samples, the final quantitative pathological feature descriptor shows that the vacuolar area accounts for 30% of the weight, and the lobule angular distribution accounts for 50%. The matching degree with the preset anemia cell feature library is significantly higher than that of normal cells, providing a highly discriminative quantitative basis for subsequent pathological abnormality signal judgment and solving the technical problems of feature redundancy and low pathological correlation in traditional quantitative descriptors.

[0095] Step S106: Based on the degree of matching between the quantified pathological feature descriptor and the preset feature library, determine whether there is a pathological abnormal signal and output the abnormal cell identification result.

[0096] In one specific embodiment, step S106 includes the following steps:

[0097] The cosine similarity algorithm is used to calculate the degree of matching between the quantitative pathological feature descriptor and the corresponding feature vector in the preset feature library;

[0098] If the matching degree is lower than the preset matching threshold, it is determined that there is a pathological abnormality signal;

[0099] Based on the location coordinates of potential abnormal cell regions associated with pathological abnormal signals, preliminary abnormal cell identification results are generated.

[0100] Based on the preliminary abnormal cell identification results, the curvature and smoothness parameters of the lobule edge of the nuclear lobe were extracted from the leukocyte region using an edge detection algorithm to obtain the morphological features of the lobule edge.

[0101] The boundary of the central pale-stained region of red blood cells was delineated by the region segmentation method, and the ratio of the pale-stained region area to the total cell area was calculated to obtain the distribution characteristics of the pale-stained region area ratio.

[0102] The morphological characteristics of the lobular margins and the distribution characteristics of the lightly stained areas are quantitatively described and integrated to generate complete abnormal cell identification results.

[0103] Specifically, a cosine similarity algorithm is used to calculate the matching degree between the quantified pathological feature descriptor and the corresponding feature vector in a preset feature library. The preset feature library is constructed by collecting pathological feature data from a large number of healthy blood samples, including feature vectors related to the central pale-stained area of ​​normal red blood cells and feature vectors related to the lobulation of the nucleus of normal white blood cells. Each feature vector has the same dimension as the quantified pathological feature descriptor and is a low-dimensional vector after dimensionality reduction through principal component analysis. The cosine similarity algorithm measures the matching degree by calculating the cosine value of the angle between two vectors, ranging from [-1, 1]. The closer the value is to 1, the higher the matching degree; the closer it is to -1, the lower the matching degree. Addressing the problem of low matching accuracy in abnormal cell identification due to reliance on single features in the background technology, this algorithm combines multi-dimensional pathological features for matching, solving the technical deficiency of traditional methods in insufficient characterization of complex pathological features.

[0104] The preset matching threshold is determined through statistical analysis of a large number of clinical samples. Preferably, it is set to 0.85. If the cosine similarity calculation result is lower than this threshold, it indicates that the cell morphology corresponding to the quantified pathological feature descriptor differs significantly from the normal cell morphology, and a pathological abnormality signal is determined to exist. If the calculation result is higher than or equal to this threshold, it is determined to be a normal cell with no pathological abnormality signal. The location coordinates of the potential abnormal cell region are associated with the pathological abnormality signal. The location coordinates are recorded synchronously when the potential abnormal cell region is determined in step S103, including the pixel coordinates of the upper left and lower right corners of the region. These coordinates are bound to the pathological abnormality signal to generate a preliminary abnormal cell identification result. This result also includes the intensity value of the abnormal signal (represented by the difference between the cosine similarity calculation result and the preset threshold), providing a location basis and an abnormality degree reference for subsequent feature extraction.

[0105] Based on the preliminary abnormal cell identification results, the Canny edge detection algorithm was used to extract the curvature and smoothness parameters of the lobular edges of the leukocyte region. Preferably, the Gaussian filter standard deviation of the Canny edge detection algorithm was set to 1.2, the low threshold to 50, and the high threshold to 150. This algorithm accurately delineates the edge contours of the leukocyte lobes and obtains the coordinate sequence of the edge pixels. When calculating the curvature, a sliding window processing method was applied to the edge coordinate sequence, with the window size set to 5 pixels. A quadratic polynomial fitting was performed on the pixels within the window, and the fitting equation was obtained as follows: By solving the first derivative and second derivative Using the curvature formula The curvature values ​​of the connection points are calculated, and the curvature value of the center pixel of each window is calculated. The mean and standard deviation of all curvature values ​​are used as the curvature parameters of the lobular edge. The smoothness parameter is obtained by calculating the average distance from the pixel of the edge contour to the fitted curve. The smaller the average distance, the higher the smoothness. These parameters together constitute the morphological features of the lobular edge, accurately capturing the structural abnormalities of leukocyte lobulation.

[0106] The boundary of the central pale-stained region of red blood cells is delineated using an adaptive threshold segmentation method. This method dynamically adjusts the segmentation threshold based on the local gray-level mean of the red blood cell region, as shown in the formula below. ,in This represents the average gray level of a local area. The standard deviation of gray levels in a local area. To adjust the coefficient, the red blood cell region is divided into a lightly stained region and a cytoplasmic region using this threshold. The number of pixels in the lightly stained region (S1) and the total number of pixels in the red blood cell (S2) are counted, and the ratio of the area of ​​the lightly stained region to the total cell area is calculated. Simultaneously, the distribution of this proportion within the entire potentially abnormal cell region is statistically analyzed, and the mean, standard deviation, and median of the proportion are calculated to form the distribution characteristics of the pale-stained area. This characteristic can effectively reflect the morphological abnormalities of the central pale-stained area of ​​red blood cells. For example, the proportion of the pale-stained area of ​​red blood cells in patients with iron deficiency anemia is usually significantly higher than that in normal cells.

[0107] The morphological characteristics of the leaflet margins and the distribution characteristics of the lightly stained area are quantitatively described. The quantitative indicators of the leaflet margin morphological characteristics include the mean curvature, standard deviation of curvature, and mean smoothness. The quantitative indicators of the distribution characteristics of the lightly stained area include the mean proportion, standard deviation of proportion, and median proportion. These quantitative indicators are integrated into a multidimensional feature vector in sequence. At the same time, the position coordinates and abnormal signal intensity values ​​in the preliminary abnormal cell identification results are correlated to generate a complete abnormal cell identification result. For example, the cosine similarity of a potentially abnormal cell region is 0.72, which is lower than the preset threshold of 0.85, indicating the presence of a pathological abnormal signal. The associated location coordinates are (120, 130) to (180, 190). The mean curvature of the leukocyte nucleus lobule edge extracted by the Canny edge detection algorithm is 0.08, with a standard deviation of 0.03 and a mean smoothness of 2.1 pixels. The mean proportion of the central pale-stained area of ​​erythrocytes is 0.65, with a standard deviation of 0.07 and a median proportion of 0.63. Integrating these data forms a complete abnormal cell identification result. For instance, in a leukemia sample, the complete abnormal cell identification result shows that the mean curvature of the nucleus lobule edge is 0.15, significantly higher than the 0.05 of normal cells, and the mean proportion of the pale-stained area is 0.70, higher than the 0.30 of normal cells. These quantitative data provide a precise description of abnormal cell morphology for clinical diagnosis, solving the technical problems of vague abnormal cell identification results and lack of quantitative evidence in traditional methods.

[0108] Step S107: Based on the abnormal cell identification results, perform comprehensive analysis on the blood smear microscopic images to obtain comprehensive pathological quantitative indicators.

[0109] In one specific embodiment, step S107 includes the following steps:

[0110] Based on the abnormal cell identification results, the area distribution characteristics of the central lightly stained area of ​​red blood cells in the blood smear microscopic image were extracted, and the morphological characteristics of the lobed arrangement of white blood cell nuclei were also extracted.

[0111] The area distribution features are converted into normalized vectors, the average distance and standard deviation of the arrangement pattern features are calculated, and the two types of features are fused by weighted summation.

[0112] Based on the fusion processing results, the distribution characteristics of the interlobular gaps of the nuclei and the width distribution characteristics of the transition zone between the lightly stained region and the cytoplasmic region are calculated to generate a comprehensive feature set.

[0113] After standardizing the comprehensive feature set, it is input into a preset linear regression model to calculate the pathological correlation coefficient of each feature and generate a feature weight allocation scheme.

[0114] Based on the feature weight allocation scheme, the features in the comprehensive feature set are weighted and summed to generate a comprehensive pathological quantitative index.

[0115] Specifically, based on the location coordinates, abnormal signal intensity, and quantification characteristics in the abnormal cell identification results, the corresponding regions of the blood smear microscopic images are traversed to extract the area distribution characteristics of the central pale-stained area of ​​erythrocytes and the lobed arrangement morphology characteristics of leukocyte nuclei. When extracting the area distribution characteristics of the central pale-stained area of ​​erythrocytes, the boundaries of the pale-stained area are delineated using an adaptive threshold segmentation method based on the erythrocyte regions marked in the abnormal cell identification results. The segmentation threshold is dynamically adjusted by combining the local gray-level mean and gray-level standard deviation of the erythrocyte region, with an adjustment coefficient set to -0.6. This threshold accurately distinguishes the pale-stained area from the cytoplasmic region. The number of pixels in each pale-stained area of ​​erythrocytes is counted as area data. The area data of all abnormally associated erythrocytes are summarized to form an area sequence, which corresponds to the basic data for subsequent normalization processing. When extracting the morphological features of leukocyte nucleolobule arrangement, for the leukocyte region marked in the abnormal cell identification results, a contour tracking algorithm is used to obtain the boundary pixel coordinates of each leukocyte nucleolobule. The average of all boundary pixel coordinates is taken as the coordinates of the nucleolobule center point. Then, the straight-line distance between the center points of each nucleolobule within the same leukocyte is calculated to form a distance sequence, providing data support for subsequent calculations of mean distance and standard deviation. This process solves the problem in the background technique where pathological analysis only focuses on local abnormalities and lacks global feature integration. By associating abnormality identification results to extract two types of core features, it achieves the linkage analysis of local abnormalities and overall distribution features.

[0116] The area sequence corresponding to the area distribution features is converted into a normalized vector. The min-max normalization method is used to scale each area data point using the minimum and maximum values ​​in the area sequence, ensuring that the numerical range of each element in the normalized vector is mapped to [0,1]. The dimension remains consistent with the length of the area sequence, eliminating interference from differences in area values ​​in subsequent fusion. The mean distance and standard deviation are calculated for the distance sequence corresponding to the arrangement pattern features. The mean distance is the sum of all values ​​in the distance sequence divided by the total number of distances in the sequence. The standard deviation reflects the dispersion of each value in the distance sequence from the mean distance. The obtained mean distance and standard deviation serve as quantitative indicators of the arrangement pattern features and are converted into a two-dimensional vector. A weighted summation method is used to fuse the two types of features. For example, the weight of each element in the normalized area distribution vector is set to 0.4, and the weight of the average distance in the two-dimensional vector of arrangement pattern is set to 0.4 and the weight of the standard deviation is set to 0.2. By multiplying the normalized area data with the corresponding weight element by element, and then superimposing it with the weighted result of the arrangement pattern feature, a fused feature vector is obtained. This vector carries information on both the red blood cell area distribution and the white blood cell arrangement pattern, thus solving the problem of analytical bias caused by feature fragmentation in traditional methods.

[0117] Based on the fused feature vector, the distribution characteristics of the interlobular gaps in the nuclei and the width distribution characteristics of the pale-stained area and the cytoplasmic transition zone are further calculated to generate a comprehensive feature set. When calculating the interlobular gap distribution characteristics, based on the morphological data in the fused feature vector, the distance sequence between the center points of each nucleus lobe within the same leukocyte is divided into intervals. For example, five intervals are set: [0,2), [2,4), [4,6), [6,8), and [8,+∞) pixels. The frequency of distance values ​​within each interval is counted to form a frequency vector. Simultaneously, the skewness and kurtosis of the gap distance are calculated. Skewness is obtained by dividing the sum of the cubes of the differences between each value in the distance sequence and the average distance by the product of the total number of distance sequences and the cube of the standard deviation, and is used to measure the degree of asymmetry in the distance distribution. Kurtosis is obtained by dividing the sum of the fourth powers of the differences between each value in the distance sequence and the average distance by the product of the total number of distance sequences and the cube of the standard deviation, and is used to measure the steepness of the distance distribution. The frequency vector, skewness, and kurtosis are integrated into the interlobular gap distribution feature vector. When calculating the width distribution characteristics of the light-stained area and the cytoplasmic transition zone, based on the red blood cell area distribution data in the fused feature vector, the vertical distance between each pixel point between the light-stained area and the cytoplasmic boundary is calculated using a distance transformation algorithm to form a width sequence. The average width, standard deviation, and median of the width sequence are then calculated. The average width is the sum of all values ​​in the width sequence divided by the total number of width sequences. The standard deviation reflects the dispersion of each value in the width sequence from the average width. The median serves as a supplement to the average width to avoid interference from extreme values. The average width, standard deviation, median, and the frequency vector of the width sequence distribution histogram are integrated into a width distribution feature vector of the light-stained area and the cytoplasmic transition zone. The two types of feature vectors are concatenated with the fused feature vector to generate a comprehensive feature set. The dimensions are determined according to the actual sample data. For example, for a sample containing 50 red blood cells and 30 white blood cells, the comprehensive feature set has 87 dimensions, achieving comprehensive coverage of multi-dimensional pathological features and addressing the technical deficiency of existing methods in capturing subtle features.

[0118] The comprehensive feature set was standardized using Z-scores, resulting in a mean of 0 and a standard deviation of 1 for each feature, ensuring that all features participated in the model calculation on the same scale. The standardized comprehensive feature set was then input into a pre-defined linear regression model, trained on a large number of labeled clinical samples. The training samples included healthy samples and pathological samples such as those from iron deficiency anemia and leukemia, totaling 1000 groups. Each group contained the corresponding comprehensive feature set and a clinical pathological score (range 0-10). During model training, the optimizer used stochastic gradient descent with a learning rate of 0.002, 100 iterations, a batch size of 32, and mean squared error as the loss function. ,in For actual clinical pathology scores, The model is used to predict scores, where N is the number of samples. Training stops when the loss on the validation set shows no decrease for five consecutive rounds, and the model parameters are saved. The pathological correlation coefficients of each feature are calculated using the trained model; these are the regression coefficients in the linear regression equation. The regression equation is... ,in, to For the standardized comprehensive feature set, The intercept is... to The pathological correlation coefficients for each feature are represented by these coefficients. The larger the absolute value of the coefficient, the stronger the correlation between the feature and the pathological state. After normalizing each coefficient, a feature weight allocation scheme is generated.

[0119] Based on the feature weight allocation scheme, the features in the comprehensive feature set are weighted and summed to generate a comprehensive pathological quantitative index. The calculation formula is as follows: ,in The weight of the i-th feature is... Let be the standardized i-th feature value. The index range is mapped to [0, 10], with higher values ​​indicating more severe pathological abnormalities. For example, in a blood sample, the weights for the skewness of the interlobular space are 0.25, the weight for the normalized value of the pale stained area is 0.22, and the weight for the standard deviation of the transition zone width is 0.18. The corresponding standardized feature values ​​are 1.8, 2.3, and 1.5, respectively. The weighted sum of other features is 1.2. The calculated comprehensive pathological quantitative index is: Based on clinical criteria, this value corresponds to mild pathological abnormalities, suggesting early signs of iron deficiency anemia. This process uses a linear regression model to correlate features with clinicopathological states, addressing the problem of weak correlation between quantitative indicators and clinical significance in existing methods. The generated comprehensive pathological quantitative indicators provide objective data support for clinical diagnosis, while simultaneously achieving a systematic assessment from local cellular abnormalities to the overall pathological state, improving the automation and standardization of pathological analysis.

[0120] Please see Figure 2 , Figure 2 The ROC curve comparison chart shows the true positive rate performance of the traditional threshold segmentation method, the traditional value method, and the proposed method at different false positive rates. The ROC curve of the proposed method consistently ranks higher than the other two traditional methods, achieving a true positive rate close to 0.99 when the false positive rate is only 0.05. In contrast, the true positive rates of the traditional threshold segmentation method and the traditional feature extraction method are 0.91 and 0.96, respectively, when the false positive rate is 0.05. This demonstrates that the proposed method exhibits superior sensitivity and specificity in identifying abnormal pathological signals. It can effectively control false positives while accurately capturing true positive abnormal signals, solving the problem of high false positives caused by traditional methods relying on single thresholds or superficial features. This provides a more reliable signal differentiation capability for clinical pathological identification.

[0121] Please see Figure 3 , Figure 3 The comparison chart shows the overall recognition accuracy of the traditional threshold segmentation method, the traditional feature extraction method, and the proposed method. The proposed method achieves an accuracy of 0.92, higher than the 0.78 of the traditional feature extraction method and the 0.72 of the traditional threshold segmentation method, demonstrating its advantage in overall accuracy for blood cell pathology recognition. This is because the proposed method avoids the problems of "inaccurate contour extraction" in the traditional threshold segmentation method and "insufficient capture of subtle differences" in the traditional feature extraction method, ultimately achieving improved recognition accuracy.

[0122] Please see Figure 4 , Figure 4 The chart comparing the sensitivity of anomaly detection shows the detection sensitivity of traditional threshold segmentation, traditional feature extraction, and our proposed method for potential abnormal cell regions. Our proposed method achieves an anomaly detection sensitivity as high as 0.89, while the traditional feature extraction method achieves 0.78, and the traditional threshold segmentation method only achieves 0.65. This demonstrates that our proposed method has a stronger ability to capture subtle pathological anomalies and can effectively reduce the omission of potential lesion signals. The core reason for this is that our proposed method overcomes the limitations of traditional methods' single feature analysis, solves the technical deficiency of traditional methods in failing to capture subtle lesions, and improves the sensitivity of anomaly detection.

[0123] Please see Figure 5 , Figure 5 The chart showing the comprehensive performance scores of the traditional threshold segmentation method, the traditional feature extraction method, and the proposed method is presented. The proposed method scored 91 points, the traditional feature extraction method scored 76 points, and the traditional threshold segmentation method scored only 63 points. This demonstrates that the proposed method comprehensively surpasses traditional methods in overall performance, overcoming the core shortcomings of traditional methods such as low cell segmentation accuracy, weak capture of subtle differences, and incomplete pathological quantification.

[0124] Please see Figure 6 The pathological identification and analysis system based on blood test samples in the embodiments of this application are described below. The pathological identification and analysis system based on blood test samples includes:

[0125] The image acquisition module is used to acquire blood smear microscopic images of blood test samples, extract cell region boundary information from the blood smear microscopic images, and obtain preliminary cell morphology contours.

[0126] The texture analysis module is used to perform texture analysis on the internal region of the preliminary cell morphology outline to obtain cell morphology distribution characteristics.

[0127] The anomaly detection module is used to comprehensively analyze the cell morphology distribution characteristics to determine whether there are potential abnormal cell regions.

[0128] The local enhancement module is used to perform local image enhancement processing on potentially abnormal cell regions and extract a set of subtle difference features;

[0129] The deep analysis module, based on a set of subtle difference features, combines a deep learning model to perform in-depth analysis of potentially abnormal cell regions and obtain quantitative pathological feature descriptors.

[0130] The matching and judgment module is used to determine whether there are abnormal pathological signals and output the abnormal cell identification results based on the degree of matching between the quantified pathological feature descriptor and the preset feature library.

[0131] The comprehensive evaluation module is used to perform comprehensive analysis of blood smear microscopic images based on the results of abnormal cell identification to obtain comprehensive pathological quantitative indicators.

[0132] Through the collaborative efforts of the aforementioned components, this system constructs a fully automated analysis system covering the entire process from blood smear image acquisition, feature extraction, anomaly localization to pathological quantification. It achieves closed-loop processing of blood cell pathology identification, from capturing local details to comprehensive global assessment, solving the problems of module fragmentation, analysis fragmentation, and weak clinical relevance in existing technologies. This significantly improves the accuracy, efficiency, and standardization of pathological identification.

[0133] The image acquisition module serves as the data entry point, acquiring microscopic images of blood smears through a multi-region, multi-magnification imaging strategy. It then combines threshold segmentation and the Canny edge detection algorithm to extract cell region boundaries and enhance contours, generating continuous and complete preliminary cell morphology contours. The texture analysis module, based on the preliminary cell morphology contours output by the image acquisition module, delves deeper into the texture information within the contours. It simultaneously extracts the shape distribution features, color uniformity distribution features, and leukocyte nucleus lobe number distribution features of the central pale-stained area of ​​erythrocytes. After quantization and principal component analysis for dimensionality reduction, weighted fusion is used to generate cell morphology distribution features. The anomaly detection module takes the cell morphology distribution features output by the texture analysis module and comprehensively judges the shape distribution, color uniformity, and nucleus lobe number distribution features through KL divergence calculation, entropy analysis, and chi-square statistic test, accurately locating potential abnormal cell regions that meet the abnormality criteria. The local enhancement module targets the potential abnormal cell regions located by the anomaly identification module. It employs adaptive histogram equalization for local image enhancement, highlighting detailed features within the cells. Simultaneously, it extracts the width distribution features of the central pale-stained area and the cytoplasmic transition zone of erythrocytes, as well as the morphological features of the nuclear lobe segmentation of leukocytes, stitching these together to form a set of subtle difference features. The deep analysis module, relying on the subtle difference feature set output by the local enhancement module, processes it and inputs it into a pre-trained deep learning model. Through a convolutional neural network and a fully connected layer architecture, it extracts the edge and texture information of sub-regions, generating preliminary feature maps. It then focuses on the lobulation angle distribution of the nuclear lobe and the vacuolar morphology features of the pale-stained area. Through principal component analysis for dimensionality reduction and weighted optimization, it generates highly recognizable quantitative pathological feature descriptors. The matching and judgment module calculates the cosine similarity between the quantitative pathological feature descriptors output by the deep analysis module and the normal cell feature vectors in a pre-set feature library. Combined with a pre-set matching threshold, it determines the pathological abnormality signal, associates the coordinates of potential abnormal regions to generate preliminary identification results, and then extracts the morphological features of the nuclear lobe edges and the area ratio features of the pale-stained area through edge detection and region segmentation, integrating them into a complete abnormal cell identification result. The comprehensive assessment module, serving as the system's output terminal, extracts red blood cell area distribution and white blood cell arrangement morphology features based on the abnormal cell identification results from the matching judgment module. Through fusion processing, comprehensive feature set construction, and linear regression model analysis, it generates comprehensive pathological quantitative indicators. Each module, through hierarchical data transmission and complementary functional collaboration, forms a complete analysis mechanism covering "image acquisition - feature extraction - anomaly localization - deep analysis - matching judgment - comprehensive assessment." This achieves automation, precision, and standardization in hematological pathology identification, effectively avoiding the subjectivity and tediousness of manual microscopy, and providing reliable technical support for the early screening, diagnosis, and efficacy evaluation of hematological diseases.

[0134] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the methods and systems described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0135] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for pathological identification and analysis based on blood test samples, characterized in that, Includes the following steps: Step S101: Obtain a blood smear microscopic image of the blood test sample, extract cell region boundary information from the blood smear microscopic image, and obtain preliminary cell morphology outline; Step S102: Perform texture analysis on the internal region of the preliminary cell morphology outline to obtain cell morphology distribution characteristics; Step S103: Perform a comprehensive analysis of the cell morphology distribution characteristics to determine whether there are any potential abnormal cell regions; Step S104: Perform local image enhancement processing on the potential abnormal cell region and extract a set of subtle difference features; Step S105: Based on the set of subtle differences in features, a deep learning model is used to perform in-depth analysis on the potential abnormal cell regions to obtain a quantitative pathological feature descriptor. Step S106: Based on the degree of matching between the quantified pathological feature descriptor and the preset feature library, determine whether there is a pathological abnormal signal and output the abnormal cell identification result; Step S107: Based on the abnormal cell identification results, perform comprehensive analysis on the blood smear microscopic images to obtain comprehensive pathological quantitative indicators.

2. The method according to claim 1, characterized in that, Step S101 includes: The blood smear was photographed in multiple regions and at multiple magnifications using a microscopic imaging device to obtain microscopic images of the blood smear. The blood smear microscopic image is segmented using an image processing algorithm to extract the boundary information of the cell regions; Based on the extracted boundary information, the Canny edge detection algorithm is used for edge enhancement and connection processing to generate a continuous and complete preliminary cell morphology outline.

3. The method according to claim 2, characterized in that, Step S102 includes: Texture analysis was performed on the internal region of the preliminary cell morphology outline to extract the shape distribution characteristics and color uniformity distribution characteristics of the lightly stained area in the center of the red blood cells within the outline. After quantizing the shape distribution features and the color uniformity distribution features, principal component analysis is applied to reduce the dimensionality and generate low-dimensional feature vectors. Calculate the Euclidean distance between the low-dimensional feature vector and the preset standard feature vector, and quantify the degree of abnormality in cell morphology based on the Euclidean distance; A contour tracking algorithm is used to identify the nuclear boundary of leukocytes. The number of nuclear lobes in each leukocyte is analyzed and statistically analyzed through connection components to generate the distribution characteristics of the number of nuclear lobes. The degree of abnormality is weighted and fused with the distribution characteristics of the number of lobes in the nucleus to obtain the cell morphology distribution characteristics.

4. The method according to claim 1, characterized in that, Step S103 includes: Extract the shape distribution features, color uniformity distribution features, and nuclear lobe number distribution features from the cell morphology distribution features; Calculate the KL divergence between the shape distribution feature and the normal cell morphology distribution, calculate the entropy value of the color uniformity distribution feature, and fit the nuclear lobe number distribution feature with a Poisson distribution and calculate the chi-square statistic. A preset anomaly detection threshold is set. If the KL divergence is greater than the preset divergence threshold, the entropy value is less than the preset entropy value threshold, and the chi-square statistic is greater than the preset chi-square threshold, then the preset anomaly detection condition is met. The regions where cells that meet the preset abnormality judgment conditions are located are identified as potential abnormal cell regions.

5. The method according to claim 1, characterized in that, Step S104 includes: An adaptive histogram equalization method is used to perform local image enhancement processing on the potential abnormal cell regions to highlight cell detail features; Based on the enhanced image of the potential abnormal cell region, the boundaries of the central pale-stained area and the cytoplasmic boundary of the red blood cells are identified, and the vertical distance from each boundary point to the adjacent boundary is calculated to form a width sequence. Statistical analysis was performed on the width sequence to calculate the average width and standard deviation and construct a distribution histogram to obtain the width distribution characteristics of the central pale-stained region and the cytoplasmic transition region of red blood cells. The leukocyte nucleoid regions in the potentially abnormal cell regions are denoised and each nucleoid is labeled. Locate the nuclear lobe junctions and calculate their positions, curvatures, and junction angle distributions to generate morphological features of leukocyte nuclear lobe junctions; The width distribution features and the morphological features of the lobes are spliced ​​and integrated to form a set of subtle difference features.

6. The method according to claim 1, characterized in that, Step S105 includes: The set of subtle difference features is organized into a standardized vector form and input into a pre-trained deep learning model; The deep learning model divides the potential abnormal cell region into multiple sub-region images focusing on local cell structures. Edge and texture information are extracted by a convolutional neural network and pooled for dimensionality reduction to generate a preliminary feature map. Extract the distribution characteristics of the lobulation angle of the nucleus lobe of leukocytes and the morphological characteristics of the vacuoles inside the central pale-stained area of ​​erythrocytes from the preliminary feature map; The distribution characteristics of the lobulation angle of the nuclear lobe and the morphological characteristics of the vacuoles inside the lightly stained area are spliced ​​and integrated to form an initial quantitative pathological feature descriptor. Principal component analysis was applied to the initial quantitative pathological feature descriptor to extract the main dimensions, the variance contribution rate of each dimension was calculated, and the feature weight distribution was generated. The components of the initial quantified pathological feature descriptor are weighted and optimized according to the feature weight distribution to obtain the final quantified pathological feature descriptor.

7. The method according to claim 1, characterized in that, Step S106 includes: The cosine similarity algorithm is used to calculate the degree of matching between the quantified pathological feature descriptor and the corresponding feature vector in the preset feature library; If the matching degree is lower than the preset matching threshold, it is determined that there is a pathological abnormality signal; Based on the location coordinates of the potential abnormal cell regions associated with the pathological abnormal signals, preliminary abnormal cell identification results are generated. Based on the preliminary abnormal cell identification results, the curvature and smoothness parameters of the lobular edge of the leukocyte region are extracted using an edge detection algorithm to obtain the morphological features of the lobular edge. The boundary of the central pale-stained region of red blood cells was delineated by the region segmentation method, and the ratio of the pale-stained region area to the total cell area was calculated to obtain the distribution characteristics of the pale-stained region area ratio. The morphological characteristics of the lobes of the nuclear lobes and the distribution characteristics of the area ratio of the lightly stained regions are quantitatively described and integrated to generate complete abnormal cell identification results.

8. The method according to claim 1, characterized in that, Step S107 includes: Based on the abnormal cell identification results, the area distribution characteristics of the lightly stained central area of ​​red blood cells in the blood smear microscopic image are extracted, and the morphological characteristics of the lobed arrangement of white blood cell nuclei are also extracted. The area distribution features are converted into normalized vectors, the average distance and standard deviation of the arrangement pattern features are calculated, and the two types of features are fused using a weighted summation method. Based on the fusion processing results, the distribution characteristics of the interlobular gaps of the nuclei and the width distribution characteristics of the transition zone between the lightly stained region and the cytoplasmic region are calculated to generate a comprehensive feature set. After standardizing the comprehensive feature set, it is input into a preset linear regression model to calculate the pathological correlation coefficient of each feature and generate a feature weight allocation scheme. The features in the comprehensive feature set are weighted and summed according to the feature weight allocation scheme to generate a comprehensive pathological quantitative index.

9. A pathological identification and analysis system based on blood test samples, used to implement the pathological identification and analysis method based on blood test samples as described in any one of claims 1 to 8, characterized in that, The pathological identification and analysis system based on blood test samples includes: The image acquisition module is used to acquire a blood smear microscopic image of a blood test sample, extract cell region boundary information from the blood smear microscopic image, and obtain a preliminary cell morphology outline. The texture analysis module is used to perform texture analysis on the internal region of the preliminary cell morphology outline to obtain cell morphology distribution characteristics. An anomaly detection module is used to comprehensively analyze the cell morphology distribution characteristics to determine whether there are potential abnormal cell regions. The local enhancement module is used to perform local image enhancement processing on the potential abnormal cell region and extract a set of subtle difference features. The deep analysis module, based on the set of subtle differences in features, combines a deep learning model to perform deep analysis on the potentially abnormal cell regions and obtain a quantitative pathological feature descriptor. The matching and judgment module is used to determine whether there is a pathological abnormal signal and output the abnormal cell identification result based on the degree of matching between the quantified pathological feature descriptor and the preset feature library. The comprehensive evaluation module is used to perform comprehensive analysis on the blood smear microscopic images based on the abnormal cell identification results to obtain comprehensive pathological quantitative indicators.