Imaging flow cytometry cell detection method based on improved model
By employing high-collimation illumination, autofocus, and multi-scale feature fusion of the PA-YOLO improved model, the problems of low detection accuracy and efficiency in imaging flow cytometers have been solved, enabling refined feature extraction and efficient classification of small-sized cells, and adapting to complex backgrounds and dynamic imaging requirements.
Patent Information
- Application Number
- CN202511268403.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2026-01-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing imaging flow cytometry cell detection methods suffer from low detection accuracy and efficiency in multi-scale, high-density, and dynamic detection, especially in terms of insufficient feature extraction capability for small-sized cells, and their reliance on manually set parameters leads to poor adaptability.
The system employs high-collimation illumination and high-speed acquisition combined with automatic focusing to select the optimal frame. It uses an improved PA-YOLO model for multi-scale feature fusion and refined convolution. The system utilizes multi-scale feature fusion of neck structure and PConv deformable convolution module for cell feature extraction and classification. Combined with validity verification and error correction, the entire process requires no manual parameter setting.
It significantly improves the accuracy and efficiency of cell detection, adapts to changes in cell morphology and complex backgrounds, ensures the reliability and accuracy of results, solves the problem of incomplete multi-scale feature extraction in traditional methods, and reduces the false positive rate.
Smart Images

Figure CN121305149A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of model analysis, in particular to an imaging flow cytometer cell detection method based on an improved model. BACKGROUND
[0002] In the fields of life science research, clinical diagnosis and biopharmaceuticals, cell detection, as the core technology for analyzing cell morphology, quantity and functional characteristics, directly determines the reliability of research conclusions and the accuracy of clinical decisions. Imaging flow cytometers, with the advantages of "high-speed imaging + flow analysis", can realize rapid imaging and multi-parameter analysis of a large number of cells at the single-cell level, and have become a key device in the field of cell detection.
[0003] In addition, Chinese patent CN119942540A discloses a blood cell image detection method based on a YOLOv8 improved model, which includes: obtaining blood cell image data to be detected, inputting it into the trained YOLOv8 improved model, and obtaining blood cell detection results; the YOLOv8 improved model includes a backbone network backbone, a neck network neck and a detection head head. The present application introduces Swin Transformer module and SimSPPF module in the backbone of the existing YOLOv8 model, which can reduce the model parameters and improve the real-time detection speed of the model; the SGE attention mechanism and SGE_C2f module are introduced in the neck of the existing YOLOv8 model, and the MFN module is also introduced, which can improve the feature extraction ability of blood cells while keeping the model parameter unchanged; the present application greatly improves the YOLOv8 model parameter quantity and detection index, and realizes the balance between light weight and high performance compared with the existing YOLOv8 model, which can significantly improve the accuracy of blood cell detection. In addition, the existing imaging flow cytometer cell detection method also includes a threshold segmentation method based on traditional machine vision and a detection method based on a basic deep learning model, wherein the threshold segmentation method relies on manual setting of image segmentation parameters, and has poor adaptability to cell morphology changes and complex backgrounds, and the detection accuracy is easily disturbed; the method based on the basic deep learning model can improve the feature extraction ability, but most of them are not optimized for the "multi-scale, high-density and dynamic" characteristics of cell detection, such as lack of fine feature learning mechanism for small size cells, or not solving the problem of image clarity screening in high-speed imaging. SUMMARY
[0004] The application provides an imaging flow cytometer cell detection method based on an improved model, which can lay a high-quality data foundation through high-collimation illumination and high-speed acquisition, solve the high-speed imaging definition problem through automatic focusing and optimal frame screening, cover different size cell characteristics through multi-scale feature fusion, finely convolve and strengthen small size and dense cell representation, and ensure reliable results through verification and correction, so that the method does not need manual parameter setting, adapts to cell morphology changes and complex backgrounds, meets multi-scale, high-density and dynamic detection requirements, and greatly improves detection precision and efficiency. Introduce the cell sample to be detected into an imaging flow cytometer system integrated with a microfluidic chip, so as to provide high-collimation and high-brightness white light Kohler illumination through a custom-developed scale light condensation reflection cup, and simultaneously start a high-speed camera to continuously collect images of cells in a high-speed flow state, so as to generate an initial cell image sequence; Apply an automatic digital focusing algorithm to the initial cell image sequence, so as to evaluate the definition of each frame of image in real time based on a multi-scale image definition evaluation function, and screen out a cell image frame set with an optimal focal plane; Input the cell image frame set with the optimal focal plane into a preset PA-YOLO improved model, so as to perform multi-dimensional extraction and fusion on corresponding cell image features in the cell image frame set through a multi-scale feature fusion neck structure introduced in the PA-YOLO improved model, and generate a multi-scale cell feature atlas; Perform fine feature learning on the multi-scale cell feature atlas through a PConv part deformable convolution module in the PA-YOLO improved model, strengthen the feature representation of small size, dense and morphologically similar cells, obtain enhanced cell feature data, perform cell target positioning and classification through a detection head of the PA-YOLO improved model based on the enhanced cell feature data, and output cell detection results, including cell type, quantity and morphological parameter information; Verify the effectiveness of the cell detection results, compare and analyze the original image data collected by the imaging flow cytometer system, correct detection errors, and generate a cell detection report.
[0005] In the present application, high collimation and high brightness illumination is provided by means of the scale spotlight reflection cup, combined with high-speed camera to collect continuous images, laying a high-quality data foundation for subsequent detection, avoiding feature loss caused by uneven lighting or imaging blur; through the automatic digital focusing algorithm to filter the optimal image frame of the focal plane, solving the problem of image clarity difference in high-speed imaging, replacing manual judgment and eliminating detection errors caused by poor focusing; secondly, the improved PA-YOLO model adopts a multi-scale feature fusion neck structure, breaking through the limitations of traditional models in extracting incomplete multi-scale cell features, effectively covering different size cell features; the PConv deformable convolution module strengthens the feature representation of small size, dense and morphologically similar cells, making up for the defects of the basic deep learning model in the detection ability of special morphological cells, and improving the classification and positioning accuracy; finally, through effectiveness verification and error correction, further guaranteeing the reliability of the results, the overall process does not need manual parameter setting, adapts to cell morphological changes and complex background, and meets the dynamic high-speed imaging demand, greatly improving the accuracy and efficiency of cell detection.
[0006] As preferred, the scale spotlight reflection cup developed by customization provides high collimation and high brightness white light Kohler illumination, including the following steps: Start the light source module of the scale spotlight reflection cup to generate an initial white light beam; Focus and collimate the initial white light beam through the spotlight structure of the scale spotlight reflection cup to form a high-collimation parallel light beam; Direct the parallel light beam into the Kohler illumination light path, adjust the diaphragm assembly in the Kohler illumination light path, and make the parallel light beam uniformly irradiate on the cell sample flowing through the detection area of the microfluidic chip to generate an illumination light field meeting the high-speed imaging demand; Apply the illumination light field to the high-speed flowing cell sample to make the cells form cell images with high contrast and high signal-to-noise ratio on the imaging plane of the high-speed camera.
[0007] The invention fundamentally solves the problems of poor collimation, uneven brightness and low contrast of cell images in traditional imaging flow cytometry by the collaborative design of customized scale concentrating reflector cup and Kohler illumination optical path, and provides high-quality light field support for subsequent cell image acquisition. First, after the initial white light beam is generated by starting the light source module, the scale concentrating reflector cup can accurately focus and collimate the light beam. The special scale reflecting surface can reduce light scattering (traditional concentrating cup is prone to light divergence due to insufficient smoothness of the reflecting surface), convert the divergent initial light beam into a highly collimated parallel light beam, ensure low energy loss and high direction consistency of the light beam during propagation, and avoid local over-brightness or over-darkness of cells caused by light beam inclination. Second, the parallel light beam is introduced into the Kohler illumination optical path and the diaphragm assembly is adjusted to achieve uniform illumination of the light beam. Kohler illumination controls the size of the light beam cross section through the diaphragm to uniformly cover the microfluidic chip detection area (traditional illumination often has a "halo effect" of bright center and dark edges), so that consistent illumination intensity can be obtained even when cells pass through different positions of the detection area at high speed, avoiding the loss of cell features (such as cell membrane texture and intracellular particles) caused by uneven illumination. Finally, the uniform illumination light field acts on the high-speed flowing cell sample, forming high-contrast and high-signal-to-noise ratio images on the imaging plane. High contrast ensures clear boundaries between cells and background (such as distinguishing transparent cell contours from complex background impurities), and high signal-to-noise ratio reduces the interference of illumination noise on cell details (such as avoiding noise being mistaken for intracellular particles), providing a "clean" image basis for subsequent automatic focusing selection and cell feature extraction, effectively solving the problems of blurred cell images and difficult feature identification under traditional illumination, and greatly improving the quality of the initial cell image sequence.
[0008] As a preferred embodiment, the application of the automatic digital focusing algorithm to the initial cell image sequence includes the following steps: Extracting pixel gray scale information of each frame of image from the initial cell image sequence to construct an image gray scale matrix; Calculating the sharpness evaluation function value based on the image gray scale matrix, which includes the Laplacian variance value and the frequency energy distribution value; Comparing the sharpness evaluation function values of each frame of image to determine the image frame with the maximum sharpness evaluation function value as the optimal candidate frame of the focus plane in the sequence; Performing local area sharpness verification on the candidate frame to further calculate the focusing quality by analyzing the sharpness of the cell edge contour, and selecting the optimal cell image frame set of the focus plane by comparing with the preset focusing quality threshold.
[0009] This invention solves the problems of "low efficiency, large subjective error, and difficulty in handling the sharpness differences of multiple frames in high-speed imaging" in traditional cell detection by using automated, multi-dimensional sharpness evaluation and screening logic, ensuring that subsequent cell detection uses image frames of optimal quality. First, pixel grayscale information is extracted from the initial image sequence and a grayscale matrix is constructed, providing a quantitative analysis basis for sharpness evaluation. Traditional manual screening relies solely on visual observation of "whether it is clear" and lacks objective data support, while the grayscale matrix can convert the brightness changes of the image into calculable values, avoiding the bias of subjective judgment (such as different operators having different standards for "sharpness"). Secondly, by combining the Laplacian variance and frequency domain energy distribution values to calculate the sharpness evaluation function, a multi-dimensional assessment of image sharpness is achieved: the Laplacian variance reflects the sharpness of image edges (the larger the variance, the sharper the edges, such as the more obvious the transition between cell outlines and background), while the frequency domain energy distribution value captures high-frequency details of the image (the higher the high-frequency energy, the richer the details such as intracellular particles and cell membrane texture). Combining the two avoids the limitations of a single indicator (e.g., using only the Laplacian variance may mistakenly classify noisy images as sharp). Furthermore, by comparing the evaluation function values to determine candidate frames and performing local region sharpness verification (analyzing the sharpness of cell edges), the accuracy of the selection is further improved—traditional methods only select the globally sharpest frame, which may result in "globally sharp but locally blurry cells," while local verification ensures that the edges of each cell in the candidate frame are sufficiently sharp (e.g., avoiding some cells being well-focused while others are out of focus in a certain frame). Combined with a preset threshold, the optimal set of frames is selected to ensure that each frame used in subsequent detection meets the high-quality requirements. The entire process requires no manual intervention and can quickly process a large number of image frames generated by high-speed imaging (such as hundreds of frames per second), which not only improves screening efficiency but also ensures the objectivity and consistency of the results, laying a high-quality image foundation for subsequent cell feature extraction and detection.
[0010] Preferably, the calculation of the sharpness evaluation function value based on the image grayscale matrix includes the following steps: Multi-scale Gaussian filtering is applied to the image grayscale matrix to obtain smooth image matrices at different scales; For smooth image matrices at different scales, calculate their Laplacian second derivative matrices, and then sum the squares of all elements in the Laplacian second derivative matrices to obtain the Laplacian variance values for the corresponding scales. The image grayscale matrix is transformed to the frequency domain space to obtain the frequency domain spectrum matrix. High-frequency components in the frequency domain spectrum matrix are extracted by setting a high-frequency threshold. The sum of the energy of the high-frequency components is calculated to obtain the frequency domain energy distribution value at the corresponding scale. By combining the Laplacian variance values at various scales and the frequency domain energy distribution values, a sharpness evaluation function value for the image frame is generated.
[0011] This invention employs a multi-scale processing and multi-index fusion method for image sharpness evaluation, addressing the problems of insufficient coverage of multi-scale features in cell images, susceptibility to noise interference, and biased evaluation results in traditional sharpness calculations. This significantly improves the accuracy and reliability of image sharpness assessment. First, multi-scale Gaussian filtering of the image grayscale matrix effectively separates features at different scales. Cell images contain both large-scale cell outlines and small-scale intracellular particles and cell membrane textures. Traditional single-scale evaluation tends to overlook small-scale details (e.g., focusing only on cell outline sharpness while ignoring the blurring of intracellular particles). Multi-scale filtering (e.g., using 3×3, 5×5, and 7×7 filter kernels) can correspond to features of different sizes, ensuring that sharpness at each scale is included in the evaluation scope and avoiding evaluation bias caused by a single scale. Secondly, by calculating the Laplacian variance and frequency domain energy distribution values at each scale, we achieved dual verification of sharpness in terms of both "detail sharpness" and "high-frequency richness": the second derivative of Laplacian can enhance image edges (such as the boundary between cells and background, and the edges of intracellular particles), and the larger the variance of its sum of squares, the sharper the edges (i.e., the clearer the features at that scale); after converting the image to the frequency domain, the high-frequency components correspond to the detailed information of the image (such as cell membrane folds, intracellular particles), and the high-frequency components are extracted and the energy sum is calculated to quantify the richness of details (the higher the high-frequency energy, the clearer the details). The combination of the two avoids the limitations of a single indicator (such as the Laplacian variance may be artificially high due to noise, while the frequency domain energy can help determine whether it is real detail). Finally, the final evaluation function value is generated by combining the indicators at various scales, which can comprehensively reflect the image sharpness at different scales. For example, if an image is sharp at a large scale (cell outlines) but blurry at a small scale (intracellular particles), a single-scale evaluation may determine it as "sharp," but after multi-scale synthesis, it will be found that its small-scale indicators are low, thus giving a more objective evaluation result. This refined evaluation method can accurately distinguish between "true sharpness" (sharp features at all scales) and "false sharpness" (sharpness at a local or single scale), providing a precise quantitative basis for selecting the optimal image frame for the focal plane, and effectively avoiding subsequent cell detection errors caused by inaccurate sharpness evaluation (such as misjudging blurry small cells as impurities, or missing sharp small cells).
[0012] Preferably, the step of performing local region sharpness verification on candidate frames to further calculate their focus quality by analyzing the sharpness of cell edge contours, and filtering out the set of cell image frames with the optimal focal plane by comparing with a preset focus quality threshold, includes the following steps: The candidate frames are divided into local regions, and multiple key local regions containing cell targets are selected. The cell edge contour data of each key local region is extracted. The cell edge contour data is processed by an edge detection algorithm, and the gray-scale change rate of the edge contour is calculated. This gray-scale change rate is used to characterize the sharpness of the cell edge contour. The focus quality score of the candidate frame is calculated based on the sharpness of the cell edge contour, and the focus quality score is compared with the preset focus quality threshold. If the focus quality score is greater than or equal to the preset focus quality threshold, the candidate frame is determined to be the optimal image frame for the focal plane. If the focus quality score is less than the preset focus quality threshold, the image frame with the second largest sharpness evaluation function value is selected from the initial cell image sequence as a new candidate frame, and the above local region sharpness verification steps are repeated until image frames that meet the focus quality requirements are selected, thereby forming a set of cell image frames with optimal focal plane.
[0013] This invention solves the problem of traditional focal plane filtering that "only focuses on global sharpness and ignores local cell blur" by using a refined local region sharpness verification and cyclical screening mechanism. This ensures that every cell in the final set of selected image frames has a high-quality focus effect, laying a solid foundation for subsequent feature extraction and detection. First, key local regions are divided into candidate frames and cell edge contour data is extracted, moving the sharpness verification from the "global level" to the "individual cell level." Traditional methods only judge image quality through global sharpness indicators (such as overall grayscale variance), which may result in "global sharpness but local cells being out of focus" (e.g., most cells in a frame are sharp, but small cells in the edge region have blurred contours due to focusing deviation). By focusing on the local region containing the cell target, the focus state of each cell can be accurately locked, avoiding the obscuring of local problems by global evaluation. Secondly, the grayscale change rate is calculated using an edge detection algorithm to characterize edge sharpness, providing a quantitative basis for focus quality. The grayscale change rate of cell edges directly reflects the focusing effect (when in focus, the grayscale transition between cells and background is steep, with a high change rate; when out of focus, the transition is smooth, with a low change rate). This quantitative evaluation avoids the subjectivity of manual judgment and ensures a consistent evaluation standard for focus quality across different cells and frames. Furthermore, a preset threshold and a cyclical screening mechanism are set to construct a closed-loop process of "screening-verification-supplementation": if the local focus quality of a candidate frame is substandard, a second-best frame is immediately re-selected for verification, rather than being discarded directly, avoiding the omission of high-quality images due to local defects in a single candidate frame. For example, if the initial candidate frame has the highest global sharpness, but three small cells in the edge region are out of focus (grayscale change rate below the threshold), the system will automatically select a second-best frame for re-verification until an image frame with "high global sharpness and all local cells are sharp" is found. This mechanism effectively solves the limitations of traditional screening methods that rely on "one-time judgment without remedial measures," ensuring that in the final set of optimal focal plane image frames, every cell in each frame meets the focusing requirements (sharp edges and complete details), completely eliminating subsequent detection errors caused by local defocusing (such as misjudging small defocused cells as impurities or missing morphological abnormalities in defocused cells), and significantly improving the reliability of cell image data.
[0014] Preferably, the multi-dimensional extraction and fusion of cell image features within the cell image frame set by introducing multi-scale feature fusion of the neck structure includes the following steps: The set of cell image frames with the optimal focal plane is input into the backbone feature extraction network of the PA-YOLO improved model to obtain the original feature maps at different levels; Upsampling and downsampling are performed on the original feature map to generate feature maps at multiple scales; By using multi-scale feature fusion and skip connections and feature splicing operations in the neck structure, feature maps of multiple scales are fused to enhance the correlation between them. The fused feature map is subjected to convolutional operations for dimensionality adjustment and feature optimization, generating a multi-scale cell feature map containing multi-scale cell information.
[0015] This invention addresses the shortcomings of traditional deep learning models, such as incomplete extraction of multi-scale cellular features and weak feature correlation, by integrating multi-scale features into the hierarchical processing and correlation enhancement of the neck structure. This achieves deep integration of cellular image features, providing rich feature support for subsequent accurate detection. First, the optimal image frame is input into the backbone network to obtain original feature maps at different levels, covering cellular features from macroscopic to microscopic. The shallow feature maps (e.g., layer 2) capture low-level features such as cell edges and contours, while the deep feature maps (e.g., layer 5) extract high-level features such as intracellular particle distribution and cell membrane texture. This hierarchical extraction avoids the limitations of traditional models with their "single feature level" (e.g., extracting only shallow contour features while missing deep detail features), ensuring no multi-scale cellular information is missed. Secondly, multi-scale feature maps are generated through upsampling and downsampling to adapt to the feature requirements of cells of different sizes. In cell detection, there are scenarios where "large-sized cells (such as macrophages) and small-sized cells (such as lymphocytes) coexist." Traditional fixed-scale feature maps are difficult to adapt to both types of cells at the same time (large-scale feature maps are prone to losing details of small cells, while small-scale feature maps are difficult to capture the global morphology of large cells). Multi-scale mapping maps (such as 16×16, 32×32, and 64×64) can correspond to different sizes of cells respectively, ensuring that the features of each size of cell can be accurately extracted. Furthermore, skip connections and feature concatenation operations enhance feature correlation, breaking the problem of "hierarchical isolation" in traditional feature fusion. Skip connections directly associate shallow, low-dimensional features with deep, high-dimensional features (e.g., concatenating shallow edge features into deep detail features), avoiding the loss of edge information caused by multiple convolutions in deep features. Feature concatenation integrates the advantages of features at different scales (e.g., combining the global shape of large-scale features with the local details of small-scale features), so that the fused features contain both the "overall outline of the cell" and subtle information such as "the location of intracellular particles and cell membrane folds." Finally, convolution operations optimize feature dimensionality and quality, ensuring the consistency and effectiveness of multi-scale feature maps. The number of feature channels is adjusted through 1×1 convolution to avoid dimensional chaos caused by multi-scale feature concatenation. Key features are further enhanced through 3×3 convolution (e.g., highlighting the morphological differences of abnormal cells). The resulting multi-scale cell feature map can comprehensively and accurately represent multi-dimensional cell information, providing sufficient feature basis for the subsequent differentiation of small-sized, morphologically similar cells, effectively solving the pain point of "incomplete multi-scale cell feature extraction" in traditional models.
[0016] Preferably, the refinement of multi-scale cell feature maps through the PConv partially deformable convolutional module in the PA-YOLO improved model includes the following steps: The multi-scale cell feature map is divided into multiple local feature regions, and each local feature region corresponds to a set of convolutional kernel parameters. The PConv deformable convolution module dynamically adjusts the sampling position and weight coefficients of the convolution kernel parameters based on the cell morphology distribution characteristics of local feature regions. By performing convolution operations on local feature regions using adjusted convolution kernel parameters, targeted local cell features can be extracted. By integrating the extraction results of all local feature regions, enhanced cellular feature data that can accurately characterize subtle cellular morphological differences are obtained.
[0017] This invention addresses the shortcomings of traditional convolution, such as insufficient adaptation to cell morphological diversity and weak ability to extract subtle features, by dynamically adjusting the parameters of the deformable convolution module in PConv and extracting local features in a targeted manner. This enables refined learning of cell features and significantly improves the accuracy of cell detection. First, the multi-scale feature map is divided into local feature regions and matched with dedicated convolution kernel parameters, breaking the limitation of the traditional convolution's "globally uniform convolution kernel." Traditional fixed convolution kernels (such as 3×3) have poor adaptability to irregularly shaped cells (such as deformed red blood cells and adherent cancer cells), which can easily lead to biases in local feature extraction (such as misjudging the concave shape of red blood cells as a normal circle). However, after dividing by region, each local region (such as the cell edge region and the intracellular granule region) can correspond to an independent convolution kernel, ensuring that the convolution operation can focus on specific features within the region (such as focusing on contour changes in the edge region and focusing on density distribution in the granule region). Secondly, the PConv module dynamically adjusts the sampling position and weight coefficients of the convolution kernel based on cell morphology distribution to achieve "on-demand extraction." By analyzing the cell morphology of local regions (e.g., cells in one region are distributed in elongated strips, while in another region they are distributed in clusters), the module automatically adjusts the sampling point position (e.g., sampling points in elongated regions are distributed along the long axis, while sampling points in clustered regions are evenly distributed) and weights (e.g., assigning high weights to sampling points at cell edges and low weights to background regions). This avoids the invalid feature acquisition caused by the traditional convolution's "fixed sampling position and equal weights" (e.g., acquiring a large amount of background noise or missing subtle morphological differences in cells). Furthermore, targeted convolution operations extract local cell features, enhancing the representation of subtle morphological differences. For example, when detecting "normal lymphocytes and abnormal lymphocytes with similar morphology," traditional convolution struggles to distinguish subtle differences in nucleus size. The PConv module, however, can dynamically adjust the convolution kernel for the nucleus region (e.g., reducing the sampling interval and increasing weights) to accurately extract grayscale changes and size information at the nucleus edges, significantly amplifying the feature differences between the two. Finally, enhanced feature data is generated by integrating local features to ensure the integrity and accuracy of cell features. By integrating the targeted features of each local region (edge contour, nucleus size, particle density, etc.), the resulting feature data can not only fully cover the global morphology of the cell, but also accurately capture subtle differences (such as tiny protrusions on the cell surface and changes in the number of intracellular particles). This effectively solves the problem of insufficient characterization of small, dense, and morphologically similar cells in traditional models, providing highly recognizable feature support for subsequent cell localization and classification, and significantly reducing the misclassification rate (such as avoiding misclassifying abnormal cells as normal cells or missing small abnormal cells).
[0018] Preferably, the dynamic adjustment of the sampling position and weight coefficients of the convolution kernel parameters includes the following steps: Feature gradients are calculated for local feature regions to determine the distribution direction of cell edges and textures, thereby generating corresponding feature gradient information. An offset parameter is generated based on the feature gradient information. This offset parameter is used to adjust the sampling coordinates of the convolution kernel in the spatial dimension. Different weighting coefficients are assigned to each sampling point of the convolution kernel based on the importance of cell pixels in the local feature region; The offset parameter and weight coefficients are applied to the original convolution kernel to form a deformed convolution kernel that adapts to the local feature region.
[0019] This invention solves the problems of traditional fixed convolutional kernels, such as "rigid sampling positions, equal weight distribution, and inability to adapt to dynamic changes in cell morphology," through a complete process design of "feature gradient analysis - offset generation - weight allocation - deformable convolutional kernel construction," achieving precise matching between the convolutional kernel and local cell features. First, the feature gradient of the local region is calculated, and the distribution direction of cell edges and textures is determined, providing a "targeted basis" for parameter adjustment. The distribution direction of cell edges and textures directly reflects key feature areas (such as the long axis direction of elongated cells and the aggregation direction of intracellular particles). Traditional convolutional kernels do not consider this directional information, easily sampling in non-critical areas, leading to feature omissions. Feature gradient information, however, can accurately pinpoint the directions that need to be focused on, ensuring that subsequent parameter adjustments revolve around the core features. Secondly, offset parameters are generated based on gradient information to adjust the sampling coordinates, allowing the convolutional kernel to "actively fit" the cell morphology. For example, for concave red blood cells, the offset parameters can shift the sampling points of the convolutional kernel towards the concave edge, avoiding sampling points falling in featureless background areas. Compared to traditional fixed sampling coordinates (such as sampling only at fixed positions around the center of the convolutional kernel), this dynamic offset can capture cell morphological details more comprehensively. Furthermore, weight coefficients are assigned according to the importance of cell pixels, achieving "high weight for key features and low weight for background noise"—pixels at the cell edge contribute far more to morphological recognition than background pixels. Traditional equal weighting of convolutional kernels dilutes the influence of key features, while differentiated weight allocation can strengthen the feature contribution of key pixels such as edges and textures (e.g., giving edge pixels twice the weight of background pixels), reducing noise interference. Finally, deformable convolution kernels adapted to local regions are constructed, so that each local region has its own dedicated convolution kernel. The parameters of the convolution kernels for different local regions (such as cell edge regions and intracellular granule regions) are different, which avoids the feature extraction deviation caused by "one-size-fits-all" parameter settings, greatly improves the accuracy of local cell feature extraction, and provides higher quality feature data for subsequent cell classification and recognition.
[0020] Preferably, generating the offset parameter based on feature gradient information includes the following steps: Calculate the gradient values of each pixel in the local feature region in the horizontal and vertical directions, and construct the gradient matrix; Gaussian smoothing is applied to the gradient matrix to suppress noise interference, resulting in a smoothed gradient matrix. Based on the magnitude and direction of the gradient values in the smooth gradient matrix, the regions of significant changes in cell characteristics are determined. Within regions of significant change, offset parameters for the convolution kernel sampling positions are generated according to a preset scaling factor.
[0021] This invention solves the problems of traditional offset generation being "susceptible to noise interference and inaccurate localization of salient features" through refined gradient calculation and smoothing, ensuring that the offset parameters accurately point to the key feature regions of cells, providing a reliable basis for adjusting the sampling position of the convolution kernel. First, the gradient values in the horizontal and vertical directions of each pixel are calculated and a gradient matrix is constructed to quantitatively represent changes in cell features—the gradient value reflects the degree of drastic change in pixel grayscale (large gradient values for edge pixels and small gradient values for background pixels), and the gradient direction reflects the direction of feature change (e.g., the gradient direction of horizontal edges is vertical). This quantification matrix avoids the subjective judgment of feature regions in traditional methods, making the localization of salient feature regions systematic. Second, the gradient matrix is Gaussian smoothed to effectively suppress noise interference—imaging noise (such as random grayscale changes caused by illumination fluctuations) often exists in cell images. This noise generates false gradient values, misleading the judgment of salient feature regions. Gaussian smoothing, through weighted averaging of neighboring pixels, can smooth the false gradients generated by noise (e.g., reducing the high gradient values of isolated noise points) while retaining the high gradient values of real cell edges, ensuring that the smoothed gradient matrix accurately reflects the true feature distribution of cells. Furthermore, based on the magnitude and direction of the gradient values in the smoothing gradient matrix, regions of significant change are identified, precisely pinpointing key cellular features—regions with large gradient values and consistent directions (such as continuous cell edges) are considered salient feature regions. Traditional methods may misclassify noisy regions as salient feature regions due to a lack of distinction between gradient magnitude and direction. This step, however, uses a dual assessment (magnitude + direction) to eliminate noise interference and accurately locate regions requiring focused sampling (such as cell edges and nucleus boundaries). Finally, offset parameters are generated within these salient feature regions, ensuring that parameter adjustments focus on core features—offset parameters are generated only for salient feature regions, avoiding wasted sampling resources in featureless background areas. Compared to the traditional method of generating offsets across the entire region, this significantly improves the targeting and efficiency of parameter adjustments, laying the foundation for accurate sampling of subsequent convolutional kernels.
[0022] Preferably, generating the offset parameter of the convolution kernel sampling position according to a preset scaling factor includes the following steps: Set a base offset range, which is determined based on the resolution of the imaging flow cytometer system and the average size of the cells; The gradient value of each pixel in the smooth gradient matrix is normalized to obtain the gradient normalization value. Multiply the gradient normalization value by the base offset range to obtain the candidate offset for each pixel. Cluster analysis was performed on the candidate offsets, and the candidate offsets with the highest frequency were selected as the offset parameters of the convolution kernel in the region of significant change.
[0023] This invention solves the problems of traditional offset generation, such as "lack of standard range, discrete parameters, and inability to adapt to system and cell characteristics," through a process of "basic range setting - gradient normalization - candidate offset generation - clustering screening." It ensures that the offset parameters both meet the hardware limitations of the imaging system and accurately adapt to cell size, improving the effectiveness of convolution kernel sampling. First, a basic offset range is set based on the imaging system resolution and the average cell size to avoid sampling failure caused by offsets that are "too large or too small." The imaging system resolution determines the minimum sampling accuracy (e.g., when the resolution is 1 μm / pixel, an offset less than 0.5 μm is meaningless), and the average cell size determines the maximum offset range (e.g., when the cell diameter is 10 μm, an offset exceeding 5 μm will cause sampling points to exceed the cell's range). Traditional methods do not consider these hardware and sample characteristics, easily generating invalid offsets (e.g., excessively large offsets cause sampling points to fall outside the cell). The basic range setting in this step ensures that the offset parameters are within a reasonable range, guaranteeing sampling effectiveness. Secondly, the gradient normalization value is multiplied by the base range to generate candidate offsets, thus establishing a correlation between the offset and the importance of the feature. The gradient normalization value reflects the importance of a pixel in the feature region (pixels with a normalization value of 1 are core feature points, and those with a normalization value of 0.5 are secondary feature points). Multiplying it by the base range makes the candidate offset corresponding to the core feature point closer to the upper limit of the base range (e.g., when the base range is 0-2 pixels, the candidate offset for the core feature point is 2 pixels, and for the secondary feature point it is 1 pixel). This ensures that the convolution kernel sampling points are preferentially offset towards the core feature points, improving the sampling targeting. Furthermore, cluster analysis is used to screen high-frequency candidate offsets, addressing the issue of parameter dispersion. Candidate offsets may have discrete values due to local pixel differences (e.g., candidate offsets might be 1.8, 2.0, or 2.2 pixels within the same salient region). Directly using discrete values would lead to chaotic sampling positions for the convolution kernel. Cluster analysis groups similar candidate offsets into one category and selects the most frequent value (e.g., 2.0 pixels) as the final offset, ensuring that the offset parameters are consistent and represent the needs of most feature points in the region, avoiding discontinuous feature extraction caused by scattered sampling positions. Finally, the generated offset parameters not only meet the system hardware limitations but also adapt to cell size and feature importance, ensuring that the convolution kernel sampling points accurately fall on key cell feature regions. This significantly improves the consistency and accuracy of local cell feature extraction, providing more reliable feature support for subsequent cell morphology recognition.
[0024] The present invention has the following specific beneficial effects: (1) By coordinating the design of a customized lighting system and high-speed acquisition, the problems of "uneven illumination and blurred images" in traditional imaging flow cytometers are solved from the source, providing high-quality initial data for subsequent detection. First, the focusing structure of the scaly reflector cup can convert the initial white light beam into highly collimated parallel light, reducing beam scattering and avoiding the problem of cells being too bright or too dark in some areas due to poor collimation in traditional reflector cups. Combined with the Köhler illumination path adjustment aperture, the beam can uniformly cover the detection area of the microfluidic chip, so that even if the cells are flowing at high speed, consistent illumination can be obtained, eliminating the interference of "halo effect" on cell features. Second, the high-speed camera continuously acquires flowing cells, which can capture the dynamic morphology of cells (such as the instantaneous state of deformed cells), avoiding the omission of key information in traditional static acquisition. The final generated initial cell image sequence has both high contrast (clear cell and background boundaries) and high signal-to-noise ratio (reduced illumination noise), ensuring that details such as cell texture and intracellular particles are completely preserved, laying a high-quality data foundation for subsequent focusing screening and feature extraction, and effectively solving the pain points of "poor image quality and lack of dynamic information" in traditional acquisition.
[0025] (2) By using an automatic digital focusing algorithm, the problem of "large differences in image sharpness and low efficiency of manual screening" in high-speed imaging is solved, ensuring that all image frames used for detection are at the optimal focal plane. Traditional manual screening relies on subjective judgment, which is prone to errors due to fatigue or inconsistent standards. This step, however, is based on multi-scale sharpness evaluation functions (such as Laplacian variance and frequency domain energy) to quantitatively evaluate the sharpness of each image frame. The Laplacian variance reflects edge sharpness, and the frequency domain energy captures the richness of details. The combination of the two indicators avoids the one-sidedness of a single standard. At the same time, the algorithm processes a large number of image frames in real time (hundreds of frames per second), which is far more efficient than manual screening. Furthermore, by verifying local regions (analyzing the sharpness of cell edges), the case of "globally clear but locally out of focus" is eliminated. The set of frames with the optimal focal plane ensures that the cell outlines and intracellular structures in each image are clearly distinguishable, providing accurate image input for subsequent feature extraction and significantly reducing detection errors caused by image blur.
[0026] (3) By leveraging the multi-scale feature fusion of the neck structure in the PA-YOLO improved model, the limitation of "incomplete multi-scale cell feature extraction" in traditional deep learning models is overcome, and deep integration of features is achieved. Traditional models often have difficulty adapting to scenarios where "large-sized cells (such as macrophages) and small-sized cells (such as lymphocytes) coexist" due to fixed-scale feature maps. In this step, feature maps of different levels are first obtained through the backbone network (shallow layers capture contours, deep layers extract details), and then multi-scale mapping maps are generated through upsampling / downsampling. Finally, the correlation between features at each scale is strengthened through skip connections and feature splicing (such as fusing shallow edge features with deep particle features). The fused multi-scale cell feature map contains both global morphological information of cells and retains subtle features such as intracellular particle distribution and cell membrane folds. It can comprehensively cover the feature requirements of cells of different sizes and shapes, providing sufficient feature basis for the subsequent differentiation of small-sized and similarly shaped cells, and avoiding the classification error caused by "one-sided feature extraction" in traditional models.
[0027] (4) By leveraging the synergistic effect of the PConv deformable convolution module and the detection head, the problem of "difficulty in identifying small / dense / morphologically similar cells" in traditional detection is solved, significantly improving detection accuracy. First, the PConv module dynamically adjusts the convolution kernel parameters according to the region of the multi-scale feature map—optimizing the sampling position and weight based on the cell morphology distribution (e.g., elongated, clumped), enhancing the subtle features of small cells (e.g., lymphocytes) (e.g., nucleus size), and distinguishing morphologically similar cells (e.g., normal and abnormal lymphocytes), avoiding the feature extraction bias of traditional fixed convolution kernels. Second, the detection head accurately realizes cell localization (defining cell boundaries) and classification (determining cell type) based on enhanced feature data, while outputting key information such as quantity and morphological parameters (e.g., diameter, roundness). Compared with traditional methods, this step can effectively reduce the missed detection of small cells and the misjudgment of similar cells, significantly improving detection accuracy and meeting the clinical demand for "precision and quantification" in cell detection.
[0028] (5) By validating and correcting errors, the problems of "low reliability and lack of traceability" in traditional detection are solved, ensuring the accuracy and credibility of the output report. First, the detection results are compared with the original image data to verify whether the cell localization is accurate and whether the classification is reasonable (e.g., checking whether the labeled abnormal cells actually have morphological abnormalities), and to eliminate model misjudgments (e.g., misjudging noise as cells). Second, for errors found in the comparison (e.g., missed detection of small cells), the detection results are corrected in combination with the features of the original image to ensure data integrity. Finally, the generated detection report includes cell type, quantity, morphological parameters and original image evidence, which not only facilitates clinicians to trace the basis for judgment, but also provides quantitative support for subsequent diagnosis. The whole process forms a closed loop of "detection-verification-correction", avoiding the risk of relying solely on the model output, greatly improving the reliability of the results, and meeting the requirements of "rigor and traceability" for medical testing. Attached Figure Description
[0029] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Fig. 1 This is a schematic diagram of the steps of the imaging flow cytometer cell detection method based on the improved model of the present invention. Fig. 2 This is a schematic diagram of the network structure of the PA-YOLO improved model of the present invention. Detailed Implementation
[0030] The present invention will be further described below with reference to the accompanying drawings and embodiments, but this should not be construed as limiting the present invention.
[0031] To achieve the above objectives, please refer to Figs. 1-2 This invention provides a method for cell detection using imaging flow cytometry based on an improved model, comprising the following steps: S01: The cell sample to be tested is introduced into an imaging flow cytometer system integrated with a microfluidic chip. A custom-developed scaly reflective cup provides high-collimation and high-brightness white light Köhler illumination. At the same time, a high-speed camera is activated to continuously acquire images of the cells in high-speed flow and generate an initial cell image sequence. In this embodiment of the invention, the white blood cell sample to be tested (concentration 1×10⁻⁶) is used. 6(Nucleic acid cells / mL, including neutrophils and lymphocytes) are introduced into the integrated microfluidic chip imaging flow cytometer system through the microfluidic chip sample introduction channel (100 μm inner diameter, 10 m / s flow rate). A custom-developed scaly reflective cup (composed of 12 curved reflective lenses, 98% reflectivity, 60° cone angle) activates the light source module (3 650nm LEDs, total power 15W), generating an initial white light beam (color temperature 5500K, luminous flux 1800 lm). After being focused (divergence angle compressed to 5°) and collimated (plano-convex lens focal length 20 mm) by the focusing structure, the light is introduced into the Köhler illumination path. Adjusting the field stop (20 mm aperture) and aperture stop (15 mm aperture) creates a uniform illumination field (intensity 15000 lux, uniformity ≥90%), which vertically illuminates the microfluidic chip detection area (18 mm × 10 mm). Simultaneously, a high-speed camera (resolution 1920×1080, frame rate 1000fps, exposure time 1μs, CMOS sensor pixel size 3.45μm×3.45μm) is activated to continuously acquire images of the rapidly flowing cells, generating an initial cell image sequence of 30,000 frames within 30 seconds. Each frame covers 1-2 cells, with an image grayscale range of 0-255, a cell region grayscale of 80-150, and a background grayscale of 200-255, ensuring that the sequence completely records the dynamic process of cells flowing through the detection area.
[0032] S02: Apply an automatic digital focusing algorithm to the initial cell image sequence to evaluate the sharpness of each frame in real time based on a multi-scale image sharpness evaluation function, and select the set of cell image frames with the optimal focal plane. In this embodiment of the invention, by applying an automatic digital focusing algorithm to the initial cell image sequence, a simplified grayscale matrix of 200×200 is first constructed frame by frame using an image grayscale extraction tool (only retaining the effective area of the central cell). Then, a multi-scale sharpness evaluation function is used to calculate the sharpness of each frame: a smoothing matrix is generated by using Gaussian filtering at three scales (3×3 kernel σ=0.8, 5×5 kernel σ=1.2, 7×7 kernel σ=1.6), and the Laplacian variance values (scale 10.214, scale 20.131, scale 30.079) and frequency domain energy distribution values (scale 10.295, scale 20.236, scale 30.160) at each scale are calculated. The overall sharpness value of each frame is obtained by weighting according to the scale weights (0.4, 0.3, 0.3). The overall sharpness values of 30,000 frames were sorted in descending order, and the top 100 frames were selected as candidate frames. Local region verification was then performed on the candidate frames: five key regions were divided, including the cell nucleus region and the cell membrane region. Edge contour data was extracted (the cell nucleus edge contains 80 pixels), and the grayscale change rate was calculated (average 40 for the cell protrusion region). The focus quality score was calculated according to the region weight (0.3 for the protrusion region and 0.25 for the cell membrane region). A preset threshold of 25 was set, and 30 frames with a score ≥25 were selected (e.g., the 12,000th frame has a score of 27.95 and the 15,000th frame has a score of 26.8). These frames formed the set of cell image frames with the optimal focal plane, ensuring that the cell edges of each image in the set are clear (contrast ≥80) and the signal-to-noise ratio ≥30dB.
[0033] S03: Input the set of cell image frames with the optimal focal plane into the preset PA-YOLO improved model, so as to extract and fuse the corresponding cell image features in the set of cell image frames in multiple dimensions by introducing multi-scale feature fusion of the neck structure, and generate a multi-scale cell feature map. In this embodiment of the invention, a set of 30 cell image frames with optimal focal plane (200×200×3 channels per frame) is input into a preset PA-YOLO improved model. The model's backbone feature extraction network first extracts features through a Conv module (3×3 cores, outputting 64 channels), a PCov module (grouped into 8 groups, outputting 128 channels), and a C2f module (n=3), generating original feature maps of 32×32×256 (shallow layer), 16×16×512 (middle layer), and 8×8×1024 (deep layer). Multi-scale feature fusion of the neck structure processes the original feature maps: the 8×8 feature map is upsampled to 16×16 through transposed convolution (4×4 kernel, stride 2) and concatenated with the middle layer feature map channels (16×16×1024); the 32×32 feature map is downsampled to 16×16 through convolution (3×3 kernel, stride 2) and concatenated again (16×16×1536); at the same time, the 8×8 and 32×32 feature maps are also concatenated with the corresponding scale feature maps respectively. All concatenated feature maps are adjusted for the number of channels through 1×1 convolution (16×16×512, 8×8×1024, 32×32×256) to generate a multi-scale cell feature map. In the map, the 32×32 scale focuses on cell membrane protrusions, the 16×16 scale focuses on cell nuclei, and the 8×8 scale focuses on the overall semantics of the cell, ensuring the complete fusion of multi-dimensional features.
[0034] S04: Utilize the deformable convolutional module of the PConv part in the PA-YOLO improved model to perform refined feature learning on multi-scale cell feature maps, enhance the feature representation of small, dense, and morphologically similar cells, and obtain enhanced cell feature data; Based on the enhanced cell feature data, perform cell target localization and classification through the detection head of the PA-YOLO improved model, and output cell detection results, including cell type, quantity, and morphological parameter information; In this embodiment of the invention, the deformable convolutional module of the PConv part of the PA-YOLO improved model is used to perform refined feature learning on multi-scale cell feature maps. The 32×32 map is divided into 64 4×4 local regions, and the pixel gradient of each region is calculated (using the Prewitt operator, Gx, Gy). The gradient matrix is constructed and Gaussian smoothed (3×3 kernel σ=0.8). Significant change regions are identified (|G|≥0.6), and offset parameters (e.g., -0.37 pixels in the protrusion region) and weight coefficients (0.15 in the core region) are generated to form deformable convolutional kernels (3×3 kernel, sampling points are offset along the gradient direction). Local features are extracted by convolution of the regions. The 16×16 and 8×8 maps are processed in the same way to obtain enhanced cell feature data (16×16×1392 channels). The model detection head (including classification and regression branches) processes enhanced feature data: the classification branch determines cell types (neutrophil probability ≥0.85, lymphocyte probability ≥0.8) through a fully connected layer (1392-dimensional input, 2-class output); the regression branch predicts cell bounding boxes (coordinate deviation ≤2 pixels) and morphological parameters (length and width errors ≤0.1μm). The final cell detection results are: 18 neutrophils (length 12-15μm, width 8-10μm, number of processes 5-7) and 12 lymphocytes (length 8-10μm, width 6-8μm, number of processes ≤2), with a statistical error of ≤5% and a type recognition accuracy of ≥92%.
[0035] S05: Verify the validity of the cell detection results, compare and analyze them with the raw image data acquired by the imaging flow cytometer system, correct detection errors, and generate a cell detection report.
[0036] In this embodiment of the invention, the validity of the cell detection results is verified by using a result verification tool to extract 20% of the detection results (6 cells) and comparing them with the original image data acquired by the imaging flow cytometer system: In the neutrophil detection results, the cell bounding box in frame 12000 has an overlap of ≥90% with the cell outline in the original image, and the error between the morphological parameter (length 14μm) and the measured value in the original image (14.2μm) is ≤1.4%; for lymphocytes, the bounding box overlap in frame 15000 is ≥88%, and the error between the morphological parameter (length 9μm) and the original value (8.9μm) is ≤1.1%. For one neutrophil detection result with an error exceeding 5% (bounding box offset of 3 pixels), the features are re-extracted from the original image (adjusting the offset parameter to -0.35 pixels), and the overlap is improved to 91% after correction. Integrating all validated and corrected results, a cell detection report is generated: the sample contains 18 neutrophils (60%) and 12 lymphocytes (40%). The average length of neutrophils is 13.5μm and the average width is 9.2μm, while the average length of lymphocytes is 8.8μm and the average width is 7.1μm. The report includes typical original cell images and detection bounding box annotations to ensure accurate and traceable detection results.
[0037] Furthermore, the provision of high-collimation and high-brightness white Köhler illumination through a custom-developed scale-like focusing reflector cup includes the following steps: The light source module of the scale-shaped light-concentrating reflector cup is activated to generate an initial white light beam; In this embodiment of the invention, the light source module built into the scale-like focusing reflector is activated. This module uses three high-power LED beads with a wavelength of 650nm (5W power per bead, 120lm / W luminous efficiency), arranged in an equilateral triangle at the bottom of the reflector (15mm spacing, 30mm from the apex). A constant current drive power supply (1.2A output current, 4.2V voltage) powers the LED beads, controlling the light source module to output an initial white light beam (color temperature 5500K, luminous flux 1800lm, beam divergence angle 45°). During startup, a power meter (measurement range 0-10W, accuracy ±0.01W) monitors the beam power in real time. Once the power stabilizes at 15W (total power of the three LED beads) and the fluctuation is ≤0.5W, the initial white light beam is considered generated successfully, ensuring stable light source output and providing a uniform initial light signal for subsequent focusing processing.
[0038] Furthermore, the initial white light beam is focused and collimated by the focusing structure of the scale-shaped focusing reflector cup to form a highly collimated parallel beam; In this embodiment of the invention, the initial white light beam is processed using the focusing structure of the scale-like focusing reflector cup (composed of 12 curved reflective lenses, each with a radius of curvature of 8mm, a reflectivity of 98%, an inter-lens angle of 15°, and an overall conical shape with a cone angle of 60°). The initial beam enters from the bottom of the reflector cup and is reflected multiple times by the 12 curved lenses (each lens reflects once, for a total of 12 reflections), compressing the beam with a divergence angle of 45° to within 5°. Simultaneously, it is collimated by a plano-convex lens (focal length 20mm, aperture 30mm, transmittance 99%) at the top of the reflector cup, ensuring that the beam wavefront deviation is ≤λ / 4 (λ=650nm), ultimately forming a highly collimated parallel beam (beam diameter 25mm, beam uniformity ≥90%, intensity distribution deviation ≤5%). The collimation of the beam was tested using a laser interferometer (measurement accuracy 0.1μm) to ensure that the beam spot offset was ≤0.5mm within a 1-meter propagation distance, thus meeting the incident requirements of the subsequent Köhler illumination optical path.
[0039] Furthermore, a parallel beam is introduced into the Köhler illumination optical path, and the aperture component in the Köhler illumination optical path is adjusted so that the parallel beam is uniformly irradiated onto the cell sample flowing through the detection area of the microfluidic chip, generating an illumination light field that meets the requirements of high-speed imaging. In this embodiment of the invention, a highly collimated parallel beam is introduced into the Köhler illumination optical path through an optical fiber conduit (25mm in diameter, 500mm in length, 98% transmittance). This optical path comprises three core aperture components: a condenser lens (50mm focal length, 35mm aperture), a field of view aperture (adjustable aperture range 5-25mm), and an aperture stop (adjustable aperture range 3-20mm). The field of view aperture is adjusted to 20mm to ensure the beam covers the microfluidic chip detection area (detection area size 18mm × 10mm), preventing beam overflow and energy waste. The aperture stop aperture is adjusted to 15mm to control the beam incident angle, ensuring uniform illumination within the detection area (light intensity distribution deviation ≤3%). Cell samples (concentration 1×10⁻⁶) are then placed within the microfluidic chip detection area. 6 When the light beam (cells / mL, flow rate 10m / s) flows through the sample, it illuminates the sample perpendicularly. The intensity of the illumination field is monitored by a light intensity detector (measurement range 0-20000 lux, accuracy ±100 lux) to ensure that the light field intensity is stable at 15000 lux, generating a uniform illumination field that meets the requirements of high-speed imaging (frame rate 1000fps) and avoiding blurry cell imaging due to uneven light intensity.
[0040] Furthermore, the illumination field is applied to the high-speed flowing cell sample, so that the cells form a cell image with high contrast and high signal-to-noise ratio on the imaging plane of the high-speed camera.
[0041] In this embodiment of the invention, an illumination light field is applied to a high-speed flowing cell sample (flowing through the detection area in 0.1 ms). The cells scatter and absorb the light beam (the intensity of scattered light is positively correlated with cell size, and the intensity of absorbed light is positively correlated with cytochrome content), forming a light signal carrying cell morphology information. This light signal is focused by an imaging objective (focal length 40 mm, numerical aperture 0.65, magnification 20x) and transmitted to the imaging plane of a high-speed camera (resolution 1920×1080, frame rate 1000 fps, exposure time 1 μs). The camera converts the light signal into an electrical signal using a photoelectric conversion element (CMOS sensor, pixel size 3.45 μm × 3.45 μm, quantum efficiency 65%). After processing by a signal amplification circuit (gain 20 dB, noise figure 1.5 dB), a cell image is generated. Image quality was assessed using image analysis tools: contrast (grayscale difference between cells and background ≥80) and signal-to-noise ratio (≥30dB) to ensure that each frame clearly presented the outline of cells, cell nuclei, and organelle details, providing high-quality image data for subsequent cell detection and analysis.
[0042] Furthermore, applying the automatic digital focusing algorithm to the initial cell image sequence includes the following steps: Pixel grayscale information of each frame of the initial cell image sequence is extracted to construct an image grayscale matrix; In this embodiment of the invention, pixel grayscale information is extracted frame by frame from a previously generated initial cell image sequence (500 frames in total, resolution 1920×1080, grayscale range of 0-255 per frame, covering the complete process of white blood cells and red blood cells flowing through the detection area) using an image grayscale extraction tool. For each frame, the grayscale value of each pixel is read in row and column order (e.g., the grayscale value of the pixel in the first row and first column is 120, and the grayscale value of the pixel in the first row and second column is 122), constructing a 1920-row × 1080-column image grayscale matrix (matrix elements take values from 0-255, data type is 8-bit unsigned integer). During the extraction process, a region segmentation strategy is adopted, retaining only the effective cell region of 200×200 pixels in the center of the image (excluding edge background interference), and reconstructing a simplified 200×200 grayscale matrix from the grayscale values of this region (e.g., in the simplified matrix of the 100th frame of the white blood cell sample, the grayscale values of the cell region are concentrated in 80-150, and the background region is concentrated in 200-255). The grayscale consistency verification tool was used to check that each frame's simplified matrix had no missing pixels (complete 200×200 elements) and no abnormal jumps in grayscale values (grayscale difference between adjacent pixels ≤20), ensuring that the grayscale matrix could truly reflect the grayscale distribution differences between cells and the background, providing an accurate data foundation for subsequent sharpness calculations.
[0043] Furthermore, a sharpness evaluation function is calculated based on the image grayscale matrix, including the Laplacian variance and the frequency domain energy distribution. In this embodiment of the invention, based on a constructed 200×200 simplified grayscale matrix, the Laplacian variance and frequency domain energy distribution values are calculated using a sharpness evaluation tool. When calculating the Laplacian variance, a 3×3 Laplacian operator (with coefficients [0,1,0;1,-4,1;0,1,0]) is used to convolve the grayscale matrix, resulting in an edge-enhanced matrix (the absolute value of the convolution result is large in the cell edge region and small in the background region). Then, the variance of all elements in this matrix is calculated (the larger the variance, the sharper the edges): for example, the variance of the matrix after convolution for the white blood cell sample is 125 in frame 100, 80 in frame 200, and 150 in frame 300. When calculating the frequency domain energy distribution value, the grayscale matrix is first transformed to the frequency domain using a Fast Fourier Transform (resulting in a 200×200 frequency domain matrix, where high-frequency components correspond to cell details and low-frequency components correspond to the background). Then, the total energy of the high-frequency regions (frequency ≥ 100Hz, corresponding to the edge 10% of the matrix) in the frequency domain matrix is calculated (energy = sum of the squares of the absolute values of the frequency domain elements), and the ratio of this sum to the total energy of the entire frequency domain is used as the frequency domain energy distribution value (the larger the ratio, the richer the details): 35% high-frequency energy in frame 100, 20% in frame 200, and 40% in frame 300. All calculation results are rounded to integers, and the accuracy of the convolution operation and Fourier transform is verified (calculation error ≤ 5% through standard resolution image testing) to ensure that the two evaluation indicators can effectively quantify differences in image sharpness.
[0044] Furthermore, the sharpness evaluation function values of each frame are compared, and the frame with the largest sharpness evaluation function value is determined as the candidate frame with the optimal focal plane in the sequence. In this embodiment of the invention, a frame filtering tool is used to sort the sharpness evaluation function values (Laplacian variance + frequency domain energy distribution value × 10, with weights based on calibrated from 100 sets of standard focal plane images to ensure balanced contributions) of 500 frames of images in descending order. The overall sharpness value for each frame is calculated as follows: Frame 100: 125 + 35 × 10 = 475; Frame 200: 80 + 20 × 10 = 280; Frame 300: 150 + 40 × 10 = 550; Frame 400: 110 + 30 × 10 = 410; Frame 500: 90 + 25 × 10 = 340. After sorting, the overall sharpness value of 550 for Frame 300 is the highest, and this frame is determined to be a candidate frame with the optimal focal plane. During the selection process, a numerical stability check was implemented: if the difference in the overall sharpness value of the top 3 frames was ≤10 (e.g., frame 300: 550, frame 301: 548, frame 302: 545), then these 3 frames were simultaneously listed as candidate frames (to avoid the influence of single-frame errors). In this case, the difference between frame 300 and frame 301 was 2, and the difference between frame 302 and frame 302 was 5, so frames 300-302, a total of 3 frames, were listed as the candidate frame set. Through repeated calculation verification (recalculating the overall sharpness value 3 times for the same frame image, with a deviation ≤2), it was ensured that the selected candidate frames possessed the highest sharpness, providing high-quality candidate objects for subsequent local verification.
[0045] Furthermore, the candidate frames are subjected to local region sharpness verification to further calculate their focus quality by analyzing the sharpness of the cell edge contours, and the set of cell image frames with the optimal focal plane is selected by comparing with the preset focus quality threshold.
[0046] In this embodiment of the invention, edge sharpness analysis is performed on the selected candidate frames 300-302 using a local sharpness verification tool. For each candidate frame's 200×200 effective area, it is divided into 1600 local regions (each region being 5×5 pixels) using a 5×5 grid. Cell edge contours are extracted for each local region (using the Canny edge detection algorithm, with a threshold set to 80-150 to ensure only cell edges are retained). The sharpness of the edge contours is calculated (sharpness = difference in grayscale between pixels on both sides of the edge ÷ edge width; the larger the grayscale difference and the smaller the width, the higher the sharpness). The average sharpness of all local regions in each frame is calculated: frame 300: average sharpness 45 (edge grayscale difference 45, width 1 pixel); frame 301: 42; frame 302: 38. The preset focus quality threshold was 40 (determined through testing 50 sets of clear cell images; a sharpness average of ≥40 was considered acceptable for focus). Frames 300 (45≥40) and 301 (42≥40) were selected to meet the threshold, while frame 302 (38<40) was discarded. Further checks were performed on the cell integrity of the frames meeting the criteria: frame 300 clearly showed the boundary between the white blood cell nucleus (grayscale value 100-120) and cytoplasm (grayscale value 130-150); frame 301 showed slight edge blurring (grayscale difference of 40 at the cell nucleus boundary). Frame 300 was ultimately retained as the cell image frame with the optimal focal plane. If multiple frames met the criteria, the frame with the highest average sharpness was selected to generate the set of cell image frames with the optimal focal plane (in this case, a single frame; complex samples may have multiple frames), ensuring that subsequent cell detection is based on the clearest image frame for analysis.
[0047] Furthermore, the calculation of the sharpness evaluation function value based on the image grayscale matrix includes the following steps: Multi-scale Gaussian filtering is applied to the image grayscale matrix to obtain smooth image matrices at different scales; In this embodiment of the invention, a multi-scale Gaussian filtering tool is used to process the constructed 200×200 simplified image grayscale matrix (taking the 300th frame of the white blood cell sample as an example, the grayscale of the cell region is 80-150, and the background is 200-255). Three filtering scales are set, with corresponding Gaussian kernel sizes and standard deviations as follows: scale 1 (3×3 kernel, σ=0.8), scale 2 (5×5 kernel, σ=1.2), and scale 3 (7×7 kernel, σ=1.6). The kernel elements are calculated according to the Gaussian function (e.g., the value of the center element of the 3×3 kernel is 0.225, the value of the adjacent elements is 0.194, and the value of the edge elements is 0.121). The grayscale matrix is convolved with Gaussian kernels at each scale sequentially: After convolution at scale 1, noise within 1 pixel in the image is smoothed (background grayscale fluctuations are reduced from ±5 to ±2), while preserving cell edge details; after convolution at scale 2, 2-3 pixel noise is smoothed (grayscale uniformity within cells is improved by 15%), while edges remain clear; after convolution at scale 3, 4-5 pixel noise is smoothed (grayscale difference in the background area ≤3), while cell outlines remain intact. Each scale convolution generates a corresponding 200×200 smoothed image matrix, which is verified using grayscale distribution analysis tools: the smoothed matrices at each scale have no edge distortion (grayscale transition difference between edge pixels ≤10) and no loss of detail (cell nuclei are still distinguishable), ensuring that smoothed images at different scales can cover image features from fine to coarse, providing a foundation for subsequent multi-scale sharpness calculations.
[0048] Furthermore, for the smoothed image matrix at different scales, the second derivative matrix of Laplacian is calculated, and then all elements in the second derivative matrix of Laplacian are squared and summed to obtain the Laplacian variance value at the corresponding scale. In this embodiment of the invention, the Laplacian second derivative matrix is calculated using a second derivative calculation tool based on the generated three-scale smoothed image matrices. A unified 3×3 Laplacian operator (coefficient matrix [0,1,0;1,-4,1;0,1,0]) is used to perform pixel-by-pixel convolution operations on the smoothed matrices at each scale: After convolution of the scale 1 smoothed matrix, the absolute values of the second derivative in the cell edge region are concentrated between 20-35 (e.g., 32 for cell nucleus edges), and concentrated between 0-5 in the background region; after convolution of the scale 2 smoothed matrix, the edge region values are 15-28 (25 for cell nucleus edges), and 0-4 in the background region; after convolution of the scale 3 smoothed matrix, the edge region values are 10-22 (18 for cell nucleus edges), and 0-3 in the background region. For the second derivative matrix at each scale, the sum of squares of all elements is calculated (sum of squares = Σ|element value|²): the sum of squares for scale 1 is 8560, for scale 2 it is 5230, and for scale 3 it is 3180. Dividing the sum of squares by the total number of matrix elements (200 × 200 = 40000) yields the Laplacian variance values for the corresponding scales: scale 10.214, scale 20.131, and scale 30.079 (rounded to three decimal places). A derivative accuracy verification tool is used to check that the second derivative of the same pixel changes continuously (without jumps) across different scales, ensuring that the variance values reflect the differences in edge sharpness at different scales.
[0049] Furthermore, the image grayscale matrix is transformed to the frequency domain space to obtain the frequency domain spectrum matrix. High-frequency components in the frequency domain spectrum matrix are extracted by setting a high-frequency threshold, and the sum of the energy of the high-frequency components is calculated to obtain the frequency domain energy distribution value at the corresponding scale. In this embodiment of the invention, the previously generated 200×200 original grayscale matrix and three scale-smoothed matrices are converted to the frequency domain using a frequency domain transformation tool. Each matrix is then converted into a 200×200 frequency domain spectrum matrix using a Fast Fourier Transform (real and imaginary parts are stored separately, with a frequency range of 0-100Hz; the center of the matrix corresponds to low frequencies, and the edges correspond to high frequencies). A uniform high-frequency threshold of 60Hz is set (determined through testing with 50 sets of cell images; components above this frequency correspond to cell details such as cell membrane texture and nuclear granules). High-frequency components with frequencies ≥60Hz are extracted from each frequency domain matrix (corresponding to elements in the 20% edge region of the matrix). The total energy of the high-frequency components is calculated (energy = Σ(real part² + imaginary part²)): 12500 for the original grayscale matrix, 9800 for scale 1 smoothed matrix, 7200 for scale 2, and 34500 for scale 3. Dividing the sum of high-frequency energies at each scale by the sum of the full-frequency energy of the corresponding frequency domain matrix (original matrix full-frequency energy 35800, scale 133200, scale 230500, scale 328100), we obtain the frequency domain energy distribution values for the corresponding scales: original matrix 0.349, scale 10.295, scale 20.236, scale 30.160 (rounded to three decimal places). Verification using a frequency domain energy verification tool shows that the proportion of high-frequency energy decreases with increasing scale (consistent with smoothing filter characteristics), ensuring that the energy distribution values accurately reflect the richness of image details at different scales.
[0050] Furthermore, by combining the Laplacian variance values and frequency domain energy distribution values at each scale, a sharpness evaluation function value for the image frame is generated.
[0051] In this embodiment of the invention, a multi-scale comprehensive evaluation tool is used to weight and fuse the Laplacian variance values and frequency domain energy distribution values at three scales to generate a sharpness evaluation function value for the image frame. Based on the contribution of each scale to sharpness (scale 1 reflects fine details, weight 0.4; scale 2 reflects medium details, weight 0.3; scale 3 reflects coarse outlines, weight 0.3), the weighted sharpness values for each scale are calculated: Scale 1 (0.214×0.4+0.295×0.4) = 0.2036 (the weight of the frequency domain energy distribution value is consistent with the Laplacian variance value), Scale 2 (0.131×0.3+0.236×0.3) = 0.1101, Scale 3 (0.079×0.3+0.160×0.3) = 0.0717. The weighted sharpness values at each scale are summed to obtain the total sharpness evaluation function value: 0.2036 + 0.1101 + 0.0717 = 0.3854 (rounded to four decimal places). Simultaneously, the frequency domain energy distribution value of the original grayscale matrix (0.349) is introduced for correction (correction coefficient 0.2, increasing by 0.0698 after correction), resulting in a final total evaluation value of 0.3854 + 0.0698 = 0.4552. Comparative verification shows that this value is higher than that of the 301st frame (0.4215) and 302nd frame (0.3987) in the same sequence, consistent with the comprehensive sharpness ranking results. This ensures that the multi-scale comprehensive evaluation can more comprehensively reflect the image focal plane quality, providing a more reliable quantitative basis for subsequent selection of the optimal focal plane frame.
[0052] Furthermore, the step of performing local region sharpness verification on candidate frames to further calculate their focus quality by analyzing the sharpness of cell edge contours, and filtering out the set of cell image frames with the optimal focal plane by comparing them with a preset focus quality threshold, includes the following steps: The candidate frames are divided into local regions, and multiple key local regions containing cell targets are selected. The cell edge contour data of each key local region is extracted. In this embodiment of the invention, the determined candidate frame (the 300th frame of the white blood cell sample, a 200×200 pixel effective area, with the cell located in the central area, and a grayscale of 80-150) is divided into five key local areas according to the cell structure characteristics using a region segmentation tool: the cell nucleus area (40×40 pixels in the center, grayscale of 100-120), the cell membrane area (a 20-pixel wide annular area surrounding the cell nucleus, grayscale of 120-135), the cell protrusion area (three 15×30 pixel areas extending from the cell, grayscale of 135-150), the near background area (a 10-pixel wide area around the cell, grayscale of 150-180), and the far background area (a 30×30 pixel area at the edge, grayscale of 200-255), ensuring coverage of the cell core structure and the background reference area. Cell edge contour data for each key region was extracted using an edge extraction tool (based on a gradient operator with a gradient threshold of 15): the edge of the cell nucleus region was a closed circular contour (containing 80 edge pixels, with a spacing of 1-2 pixels between adjacent pixels), the edge of the cell membrane region was an irregular ring contour (containing 120 edge pixels), and the edge of the cell protrusion region was a linear contour (each protrusion contained 30 edge pixels). The coordinates (e.g., the coordinates of a point on the edge of the cell nucleus (100, 100)) and grayscale value of each edge pixel were recorded to ensure that the edge data fully reflects the contour features of each cell structure, providing accurate data for subsequent sharpness calculation.
[0053] Furthermore, the cell edge contour data is processed by an edge detection algorithm to calculate the gray-scale change rate of the edge contour, which is used to characterize the sharpness of the cell edge contour. In this embodiment of the invention, the edges of each key region are processed using an edge detection tool (Canny edge detection algorithm, low threshold 80, high threshold 150) based on the extracted edge contour data, and high confidence edge pixels are retained (false edge points with gray-scale difference <10 are removed, and the retention rate is ≥90%). For each edge pixel, calculate its grayscale change rate along the contour normal direction (change rate = |average grayscale within 5 pixels on one side of the edge - average grayscale within 5 pixels on the other side| ÷ pixel spacing): Cell nucleus edge point (100, 100), average grayscale of the left 5 pixels is 105, average grayscale of the right 5 pixels is 125, spacing is 1 pixel, change rate = |105 - 125| ÷ 1 = 20; Cell membrane edge point (80, 100), average grayscale of the left side is 130, average grayscale of the right side is 160, change rate is 30; Cell protrusion edge point (120, 80), average grayscale of the left side is 140, average grayscale of the right side is 180, change rate is 40; Near background edge point (150, 100), average grayscale of the left side is 180, average grayscale of the right side is 220, change rate is 40; Far background area has no effective edge, change rate is 0. The average grayscale change rate of edge pixels in each key region was statistically analyzed: 18 for the cell nucleus region, 25 for the cell membrane region, 35 for the cell protrusion region, and 38 for the near background region. This ensures that the change rate can quantify the grayscale difference between the two sides of the edge, intuitively reflecting the edge sharpness.
[0054] Furthermore, the focus quality score of the candidate frame is calculated based on the sharpness of the cell edge contour, and the focus quality score is compared with the preset focus quality threshold. In this embodiment of the invention, a scoring calculation tool is used to calculate the focus quality score of candidate frames based on the edge grayscale change rate of each key region. Weights are assigned according to the importance of each region: cell protrusion region (reflecting fine details, weight 0.3), cell membrane region (reflecting cell outline, weight 0.25), near background region (reference benchmark, weight 0.2), cell nucleus region (reflecting core structure, weight 0.2), and far background region (no contribution, weight 0.05). The weighted average score is calculated as: 35×0.3+25×0.25+38×0.2+18×0.2+0×0.05=10.5+6.25+7.6+3.6=27.95 (rounded to two decimal places). A preset focus quality threshold of 25 is set (determined through testing 100 sets of clear cell images; a score ≥25 indicates clearly distinguishable cell edges). Comparing the candidate frame score of 27.95 with the threshold of 25, 27.95 ≥ 25, thus initially determining that the candidate frame meets the focus quality requirements. By repeatedly calculating and verifying (repeatedly dividing the same frame into regions and calculating the score with a deviation ≤ 0.5), the scoring results are ensured to be stable and reliable, avoiding the impact of single calculation errors on the judgment.
[0055] Furthermore, if the focus quality score is greater than or equal to the preset focus quality threshold, the candidate frame is determined to be the optimal image frame for the focal plane; if the focus quality score is less than the preset focus quality threshold, the image frame with the second largest sharpness evaluation function value is selected from the initial cell image sequence as a new candidate frame, and the above local region sharpness verification steps are repeated until image frames that meet the focus quality requirements are selected, thereby forming a set of cell image frames with optimal focal plane.
[0056] In this embodiment of the invention, if a candidate frame score is greater than or equal to a threshold, it is directly determined as the optimal image frame for the focal plane. If the score is less than the threshold (e.g., frame 302 has a score of 23.8 < 25), a frame iteration tool is used to select the image frame with the second largest sharpness evaluation function value (frame 301, with a comprehensive sharpness value of 548, second only to frame 300's 550) from the initial cell image sequence (500 frames) as a new candidate frame. The above steps are repeated for the new candidate frame: dividing it into 5 key regions (the cell position deviates from frame 300 by ≤ 5 pixels), extracting edge contour data (78 cell nucleus edge points), calculating the grayscale change rate (average 32 in the cell protrusion area), and weighting the score to 25.6 (≥ 25), which meets the threshold requirement, thus determining frame 301 as the optimal image frame for the focal plane. If the new candidate frame still does not meet the requirement (e.g., a score of 24.5), the second largest frame (frame 299, with a comprehensive sharpness value of 540) is selected for repeated verification until a frame with a score ≥ 25 is selected. Ultimately, this case selected two frames, the 300th and 301st, to form the set of cell image frames with the optimal focal plane, ensuring that all frames in the set meet the focusing quality requirements and providing high-quality image data for subsequent cell detection.
[0057] Furthermore, the multi-dimensional extraction and fusion of cell image features corresponding to the cell image frame set by introducing multi-scale feature fusion of the neck structure includes the following steps: The set of cell image frames with the optimal focal plane is input into the backbone feature extraction network of the PA-YOLO improved model to obtain the original feature maps at different levels; In this embodiment of the invention, a set of cell image frames with optimal focal plane (taking 10 frames of 200×200 pixels containing white blood cells and red blood cells, each with 3 channels) is input into the backbone feature extraction network of the PA-YOLO improved model. First, after passing through a 64D×64D×3 input layer, the image frames enter the Conv module (3×3 kernel size, stride 1, padding 1, 64 output channels) for preliminary feature extraction, resulting in a 64×64×64 feature map. Next, this feature map flows into the PCov module (using grouped convolution, 8 groups, 3×3 kernel, stride 1, 128 output channels) to extract richer local features, generating a 64×64×128 feature map. Subsequently, the features are processed through the C2f module (containing n Bottleneck structures, n=3, each Bottleneck consisting of Conv, Split, Concat, etc., with 128 input and output channels each) to perform multiple convolutions and fusions, resulting in a 64×64×128 feature map. Then, through subsequent modules such as the PCov module (output channels 256) and the C2f module (n=6), downsampling is continuously performed (achieved through convolutions with a stride of 2), ultimately yielding original feature maps at different levels, including feature maps with sizes of 32×32×256, 16×16×512, and 8×8×1024. These feature maps correspond to the shallow, middle, and deep features of cells, respectively. Shallow features preserve details such as cell edges, while deep features contain semantic information such as cell type. Through the convolution and fusion operations of each module, features of different levels of abstraction are gradually extracted from the image.
[0058] Furthermore, the original feature map is upsampled and downsampled to generate feature maps at multiple scales; In this embodiment of the invention, the original feature maps of different levels obtained in step S21 are subjected to upsampling and downsampling. For a deep feature map with a size of 8×8×1024, an upsampling tool (using transposed convolution, kernel size 4×4, stride 2, padding 1, and output channels 512) is used to enlarge its size to 16×16×512, generating a feature map that matches the scale of the mid-level feature map (16×16×512). For a shallow feature map with a size of 32×32×256, a downsampling tool (kernel size 3×3, stride 2, padding 1, and output channels 512) is used to reduce its size to 16×16×512, also generating a feature map that matches the scale of the mid-level feature map. Simultaneously, for the mid-layer feature map (16×16×512), an 8×8×1024 feature map is generated through downsampling (3×3 kernel size, 2 stride, 1 padding, 1024 output channels), which is associated with the scale of the deep feature map. An 32×32×256 feature map is generated through upsampling (transposed convolution, 4×4 kernel size, 2 stride, 1 padding, 256 output channels), which is associated with the scale of the shallow feature map. This generates feature maps at multiple scales (32×32, 16×16, 8×8), with each scale focusing on structural features of different cell sizes, providing a foundation for subsequent multi-scale fusion.
[0059] Furthermore, by using multi-scale feature fusion and skip connections and feature splicing operations in the neck structure, feature maps of multiple scales are fused to enhance the correlation between each feature map. In this embodiment of the invention, feature maps of multiple scales are fused in the neck structure using skip connections and feature concatenation operations. First, the upsampled 16×16×512 deep feature map is concatenated with the original mid-level feature map (16×16×512) via skip connections, merging them in the channel dimension to obtain a 16×16×1024 feature map. This operation fuses the semantic information of the deep features with the structural information of the mid-level features. Next, the downsampled 16×16×512 shallow feature map is also concatenated with the concatenated feature map via channel concatenation, generating a 16×16×1536 feature map. Simultaneously, for the 8×8 scale, the 8×8×1024 feature map obtained by downsampling the mid-layer feature map is concatenated with the original deep feature map (8×8×1024) to generate an 8×8×2048 feature map; for the 32×32 scale, the 32×32×256 feature map obtained by upsampling the mid-layer feature map is concatenated with the original shallow feature map (32×32×256) to generate a 32×32×512 feature map. Through these skip connections (directly passing features from different levels) and feature concatenation (merging features across channel dimensions), the correlation between the various feature maps is strengthened, so that the fused features contain both shallow detailed information and deep semantic information, enabling a more comprehensive description of cell features.
[0060] Furthermore, the fused feature map undergoes dimensionality adjustment and feature optimization through convolution operations to generate a multi-scale cell feature map containing multi-scale cell information.
[0061] In this embodiment of the invention, convolution operations are performed on the fused feature map to adjust the dimensions and optimize the features. Taking a 16×16×1536 fused feature map as an example, the Conv module (3×3 kernel size, stride 1, padding 1, output channels 512) is used to perform a convolution operation, adjusting the number of channels from 1536 to 512. Simultaneously, through sliding calculation of the convolution kernel, the features are further optimized to extract more representative cell features. Similarly, for an 8×8×2048 fused feature map, the Conv module (3×3 kernel, stride 1, padding 1, output channels 1024) is used to adjust the number of channels and optimize the features; for a 32×32×512 fused feature map, the Conv module (3×3 kernel, stride 1, padding 1, output channels 256) is used. After these convolution operations, a multi-scale cell feature map containing multi-scale cell information is generated. The 32×32×256 feature map focuses on small-scale cell structures (such as tiny protrusions on the cell membrane), the 16×16×512 feature map focuses on medium-scale cell structures (such as the shape of the cell nucleus), and the 8×8×1024 feature map focuses on large-scale cell semantics (such as cell type), providing rich and accurate feature basis for subsequent cell detection.
[0062] Furthermore, the refinement of multi-scale cellular feature maps through the PConv deformable convolutional module in the PA-YOLO improved model includes the following steps: The multi-scale cell feature map is divided into multiple local feature regions, and each local feature region corresponds to a set of convolutional kernel parameters. In this embodiment of the invention, the generated multi-scale cell feature atlas (including three scales: 32×32×256, 16×16×512, and 8×8×1024, taking white blood cell features as an example) is divided into local feature regions according to cell morphology and structure using a region segmentation tool. The 32×32 scale atlas (focusing on fine structures such as cell membrane protrusions) is divided into 4×4 pixel local regions, totaling 64 regions. These include 12 cell membrane protrusion regions (corresponding to the edge regions of the atlas, with feature value fluctuations ≥0.6), 20 smooth cell membrane regions (the central ring region of the atlas, with feature value fluctuations of 0.3-0.6), and 32 intercellular space regions (the background region of the atlas, with feature value fluctuations <0.3). Each region corresponds to a set of initial convolution kernel parameters (3×3 kernel, stride 1, padding 1, and weight coefficients of 0.11). 16×16 scale atlas (focusing on cell nucleus morphology): Divided into 8×8 pixel local regions, a total of 4: 1 nucleus region (atlas center, mean eigenvalue ≥ 0.8), 1 nuclear membrane region (surrounding the nucleus, mean eigenvalue 0.5-0.8), and 2 cytoplasmic regions (peripheral regions, mean eigenvalue 0.2-0.5). The corresponding initial convolution kernel is 5×5, with a stride of 1 and padding of 2. 8×8 scale atlas (focusing on overall cell semantics): Divided into 8×8 pixel single local regions (covering the entire atlas), corresponding to an initial convolution kernel of 7×7, with a stride of 1 and padding of 3. During the segmentation process, the region boundaries are determined by eigenvalue thresholds (calibrated based on 100 sets of leukocyte feature atlases) to ensure that each region accurately corresponds to a specific cell morphological structure, providing a targeted target for subsequent convolution kernel adjustments.
[0063] Furthermore, the PConv deformable convolution module dynamically adjusts the sampling position and weight coefficients of the convolution kernel parameters based on the cell morphology distribution characteristics of local feature regions. In this embodiment of the invention, the PConv deformable convolution module reads the cell morphology distribution features of each local feature region (such as the feature value gradient direction of the 32×32 scale cell membrane protrusion region and the feature value distribution density of the 16×16 scale cell nucleus region) and dynamically adjusts the convolution kernel parameters. For the 32×32 scale cell membrane protrusion region (feature value gradient is concentrated along the protrusion extension direction): the sampling position of the initial 3×3 kernel is shifted by 1 pixel along the gradient direction (the original (0,0) sampling point is shifted to (0,1)), and the weight coefficients are adjusted to 0.2 for the kernel element in the protrusion extension direction and 0.05 for the vertical direction (total weight sum is 1), enhancing the capture of protrusion morphology; for the smooth cell membrane region (uniform feature value gradient): the sampling position is not shifted, and the weight coefficients are adjusted to 0.15 for the central element and 0.09 for the peripheral element, preserving the overall features of the smooth region; for the intercellular space region (low and irregular feature values): the sampling position is shifted by 0.5 pixels, and the weight coefficients are all 0.11, reducing background interference. The 16×16 scale nucleus region (high and concentrated feature density): the 5×5 nucleus sampling position is offset by 1 pixel towards the center, with a weight of 0.15 for the central 3×3 nucleus elements and 0.05 for the peripheral elements; the nuclear membrane region (feature density distributed in a ring): the sampling position is offset by 0.8 pixels along the ring tangent direction, with a weight of 0.12 for the tangent direction and 0.08 for the radial direction. During the adjustment process, the module calculates the offset and weights using a morphological feature-parameter mapping table (preset with 1000 sets of morphological-parameter correspondences) to ensure that the parameter adjustment accurately matches the regional morphological features.
[0064] Furthermore, by performing convolution operations on local feature regions using the adjusted convolution kernel parameters, targeted local cell features are extracted. In this embodiment of the invention, convolution operations are performed on each local feature region using adjusted convolution kernel parameters. For the 32×32 scale cell membrane protrusion region: an offset 3×3 kernel (sampling points (0,1), weight 0.2) is used to convolve the 4×4 region pixel-by-pixel, extracting the length and width features of the protrusions (in the convolution result, the feature value of the protrusion region is ≥0.7, and that of the non-protrusion region is ≤0.3), generating a 4×4×256 local feature map for each protrusion region; for the smooth cell membrane region: a 3×3 kernel with enhanced central weights is used to extract the cell membrane curvature features (the feature value fluctuation in the convolution result for the smooth region is <0.2). For the 16×16 scale cell nucleus region: a 5×5 kernel with a central offset is used to extract the roundness and nucleolus number features of the cell nucleus (in the convolution result, the feature value of the nucleolus region is ≥0.9, and that of the nucleoplasm region is 0.5-0.9), generating an 8×8×512 local feature map; for the nuclear membrane region: a ring-weighted kernel is used to extract the nuclear membrane thickness features (the gradient of the feature value in the convolution result reflects the thickness difference). 8×8 scale full atlas region: Using 7×7 kernels, the overall cell size and morphological symmetry features are extracted (the distribution of feature values in the convolution result reflects symmetry, and the deviation of symmetrical regions is ≤0.1), generating an 8×8×1024 local feature map. All convolution operations use zero padding to preserve the region size, ensuring that the extracted local features correspond precisely to the original region positions.
[0065] Furthermore, the extraction results of all local feature regions are integrated to obtain enhanced cellular feature data that can accurately characterize subtle cellular morphological differences.
[0066] In this embodiment of the invention, the extraction results of each local feature region are fused using a feature integration tool. 64 local feature maps at the 32×32 scale are stitched together to form a complete 32×32×256 map, and the number of channels is compressed to 128 using 1×1 convolution, preserving key features such as cell membrane protrusions and smooth areas; 4 local feature maps at the 16×16 scale are stitched together to form a 16×16×512 map, compressing the number of channels to 256, highlighting features of the cell nucleus and nuclear membrane; 1 local feature map at the 8×8 scale directly retains 8×8×1024 channels. The integrated atlases from the three scales were matched by resolution (32×32 atlas downsampled to 16×16 and 8×8, 8×8 atlas upsampled to 16×16), and stitched together along the channel dimension (final channel count for the 16×16 scale: 128+256+1024=1392). Weighted fusion was then performed using an attention mechanism (0.3 weight for cell membrane protrusions, 0.4 for the nucleus, and 0.3 for the overall semantic region) to generate 16×16×1392 enhanced cell feature data. In this data, the deviation of leukocyte cell membrane protrusion feature values is ≤0.05, the error of cell nucleus roundness is ≤0.03, and the deviation of cell symmetry is ≤0.02. This data can accurately characterize the subtle morphological differences of different leukocytes (such as neutrophils and lymphocytes) (neutrophil protrusion count ≥5, lymphocyte protrusion count ≤2), providing highly recognizable features for subsequent cell classification and detection.
[0067] Furthermore, the dynamic adjustment of the sampling position and weight coefficients of the convolution kernel parameters includes the following steps: Feature gradients are calculated for local feature regions to determine the distribution direction of cell edges and textures, thereby generating corresponding feature gradient information. In this embodiment of the invention, the feature gradients are calculated using a gradient calculation tool for the 32×32 scale cell membrane protrusion region (a 4×4 pixel local region with feature values [0.6,0.8,0.9,0.7;0.5,0.7,0.8,0.6;0.4,0.6,0.7,0.5;0.3,0.5,0.6,0.4]) and the 16×16 scale cell nucleus region (an 8×8 pixel local region with a central 4×4 feature value ≥0.8). Using the Sobel operator (horizontal operator [[-1,0,1],[-2,0,2],[-1,0,1]], vertical operator [[-1,-2,-1],[0,0,0],[1,2,1]]), the horizontal gradient Gx and vertical gradient Gy are calculated for each pixel: For the cell membrane protrusion pixel (0,2) (feature value 0.9), Gx = 0.9 - 0.7 = 0.2, Gy = 0.8 - 0.6 = 0.2, gradient direction θ = arctan(Gy / Gx) = 45° (along the direction of protrusion extension), gradient magnitude |G| = √(0.2² + 0.2²) = 0.28; for the central pixel (4,4) of the cell nucleus region (feature value 0.95), Gx = 0.95 - 0.85 = 0.1, Gy = 0.95 - 0.85 = 0.1, gradient direction θ = 45° (along the tangent direction of the nuclear membrane), magnitude 0.14. The gradient direction distribution of each local region is statistically analyzed: 80% of the gradient direction in the cell membrane protrusion region is concentrated between 30° and 60° (consistent with the protrusion extension), and 75% of the gradient direction in the cell nucleus region is concentrated between 0° and 90° (consistent with the ring distribution of the nuclear membrane). Feature gradient information containing Gx, Gy, θ, and |G| for each pixel is generated to ensure that the gradient information accurately reflects the distribution direction and intensity of the cell edge and texture.
[0068] Furthermore, an offset parameter is generated based on the feature gradient information. This offset parameter is used to adjust the sampling coordinates of the convolution kernel in the spatial dimension. In this embodiment of the invention, offset parameters for the sampling coordinates of the convolution kernel are generated using an offset calculation tool based on previously obtained feature gradient information. The offset range is set to -1 to 1 pixel (step size 0.1), and the offset magnitude is positively correlated with the gradient magnitude (the larger the magnitude, the larger the offset), with the offset direction consistent with the gradient direction. For the 32×32 scale cell membrane protrusion region: for the 9 sampling points of the 3×3 convolution kernel (center (0,0), periphery (±1,0), (0,±1), (±1,±1)), the gradient magnitude of the center sampling point is 0.28, and the offset is set to (0.3,0.3) (along the 45° direction); the gradient magnitude of the (1,0) sampling point is 0.25, and the offset is (0.25,0.25); the gradient magnitude of the (0,1) sampling point is 0.26, and the offset is (0.26,0.26). The offsets for the remaining sampling points are generated proportionally to the magnitude, and all offsets are within the range of -1 to 1. For the 16×16 scale cell nucleus region: the gradient magnitude of the central sampling point of the 5×5 convolution kernel is 0.14, with an offset of (0.15, 0.15) (along the 45° direction); the magnitude of the sampling point at (1,0) is 0.12, with an offset of (0.12, 0.12); the magnitude of the sampling point at (0,1) is 0.13, with an offset of (0.13, 0.13); and the offset of the edge sampling points is set to (0.05, 0.05) because the gradient magnitude is small (<0.1). The generated offset parameters are stored in the format of "sampling point index - offset (x, y)" (e.g., central sampling point (0,0) → (0.3, 0.3)) to ensure that the offset of each sampling point is adapted to the gradient features of the corresponding region.
[0069] Furthermore, based on the importance of cell pixels in local feature regions, different weight coefficients are assigned to each sampling point of the convolution kernel; In this embodiment of the invention, a weighting tool is used to assign weight coefficients (total weights sum to 1) to each sampling point of the convolution kernel based on the importance of cell pixels in the local feature region (determined by the correlation between gradient magnitude and cell structure). For the 32×32 scale cell membrane protrusion region: sampling points with gradient magnitude ≥ 0.25 (center, (1,0), (0,1)) correspond to the core region of the cell protrusion, with a weight coefficient of 0.15; sampling points with magnitude 0.2-0.25 ((1,1), (-1,0)) correspond to the protrusion edge, with a weight of 0.12; sampling points with magnitude < 0.2 ((-1,-1), (1,-1), (-1,1)) correspond to the non-protrusion region, with a weight of 0.08. Calculation verification: 0.15×3+0.12×2+0.08×4=0.45+0.24+0.32=1.01 (fine-tuned to 1.0). For the 16×16 scale nucleus region: sampling points within the central 4×4 area (gradient amplitude ≥ 0.12) correspond to the nucleus core, with a weight of 0.12; sampling points on the periphery (amplitude 0.1-0.12) correspond to the nuclear membrane, with a weight of 0.10; and sampling points on the edge (amplitude < 0.1) correspond to the cytoplasm, with a weight of 0.08. The sum of the weights of the 25 sampling points for the 5×5 nucleus is: 16×0.12 + 6×0.10 + 3×0.08 = 1.92 + 0.6 + 0.24 = 2.76 (normalized proportionally to 1.0, with the core sampling point weight adjusted to 0.086, the nuclear membrane to 0.072, and the cytoplasm to 0.058). During the weight allocation process, calibration was performed using cell structure annotation data (100 sets of nucleus and cell membrane annotations) to ensure that sampling points in important areas have higher weights, thus improving the targeting of feature extraction.
[0070] Furthermore, the offset parameter and weight coefficients are applied to the original convolution kernel to form a deformed convolution kernel adapted to the local feature region.
[0071] In this embodiment of the invention, a convolution kernel reconstruction tool is used to apply offset parameters and weight coefficients to the original convolution kernel to form a deformed convolution kernel adapted to the local feature region. For the 32×32 scale cell membrane protrusion region, the original 3×3 convolution kernel (all weights 0.11): the original coordinates of each sampling point (e.g., center (0,0)) are superimposed with an offset (0.3,0.3) to obtain new sampling coordinates (0.3,0.3); simultaneously, the original weights are replaced with assigned weights (center 0.15), generating a deformed convolution kernel with sampling points offset at a 45° angle, and concentrated weights in the core region. For the 16×16 scale cell nucleus region, the original 5×5 convolution kernel (all weights 0.04): the sampling coordinates are superimposed with an offset (center (0,0)→(0.15,0.15)), and the weights are replaced with normalized values (center 0.086). After deformation, the sampling points are offset towards the center of the cell nucleus, with higher weights in the core region. The kernel verification tool was used to check that the offset of the sampling points of the deformable convolution kernel was within the range of -1 to 1, the weight sum was 1.0, and the distribution of sampling points was consistent with the cell edge and texture direction (along 45° in the cell membrane protrusion area and along the nuclear membrane tangent in the cell nucleus area). This ensured that the deformable convolution kernel could accurately adapt to the cell morphology of the local feature region and provide optimized convolution kernel parameters for subsequent feature extraction.
[0072] Furthermore, the generation of offset parameters based on feature gradient information includes the following steps: Calculate the gradient values of each pixel in the local feature region in the horizontal and vertical directions, and construct the gradient matrix; In this embodiment of the invention, a gradient matrix is constructed by using a gradient calculation tool to calculate the horizontal gradient Gx and vertical gradient Gy of each pixel in the 32×32 scale cell membrane protrusion region (4×4 pixel local region, eigenvalue matrix [[0.6,0.8,0.9,0.7],[0.5,0.7,0.8,0.6],[0.4,0.6,0.7,0.5],[0.3,0.5,0.6,0.4]]) and the 16×16 scale cell nucleus region (8×8 pixel local region, central 4×4 eigenvalue matrix [[0.8,0.85,0.9,0.85],[0.82,0.92,0.95,0.88],[0.8,0.9,0.93,0.86],[0.78,0.85,0.88,0.8]]). Using the Prewitt operator (horizontal operator [[-1,0,1],[-1,0,1],[-1,0,1]], vertical operator [[-1,-1,-1],[0,0,0],[1,1,1]]), for the cell membrane protrusion pixel (0,0) (feature value 0.6): Gx=(0.8+0.9+0.7)-(0.5+0.4+0.3)=2.4-1.2=1.2, G y=(0.3+0.5+0.6)-(0.6+0.8+0.9)=1.4-2.3=-0.9; Pixel (0,2) (feature value 0.9): Gx=(0.7+0.6+0.5)-(0.8+0.7+0.6)=1.8-2.1=-0.3, Gy=(0.5+0.6+0.4)-(0.9+0.8+0.7)=1.5-2.4=-0.9. The Gx and Gy values of all pixels are organized into 4×4 horizontal gradient matrices (e.g., the Gx matrix for the protrusion region [[1.2,1.0,0.8,-0.3],[1.1,0.9,0.7,-0.4],[1.0,0.8,0.6,-0.5],[0.9,0.7,0.5,-0.6]]) and 4×4 vertical gradient matrices, respectively. Similarly, 8×8 horizontal and vertical gradient matrices are generated for the cell nucleus region to ensure that the gradient matrix can completely record the horizontal and vertical gradient values of each pixel, reflecting the direction and intensity of feature changes.
[0073] Furthermore, the gradient matrix is Gaussian smoothed to suppress noise interference and obtain a smoothed gradient matrix. In this embodiment of the invention, the gradient matrix constructed in step one is processed using a Gaussian smoothing tool to suppress noise interference. The Gaussian kernel parameters are set as follows: a 3×3 Gaussian kernel (σ=0.8, kernel element values [[0.075,0.124,0.075],[0.124,0.204,0.124],[0.075,0.124,0.075]]) is used for the 32×32 scale cell membrane protrusion region, and a 5×5 Gaussian kernel (σ=1.2, kernel element values 0.087 at the center, 0.083 adjacently, and 0.023 at the edge) is used for the 16×16 scale cell nucleus region. Convolve the gradient matrix with a Gaussian kernel: For the pixel (0,0) (Gx=1.2) in the horizontal gradient matrix of the cell membrane protrusion region, the convolved value is: 1.2×0.075+1.0×0.124+0.8×0.075+1.1×0.124+0.9×0.204+0.7×0.124+1.0×0.075+0.8×0.124+0.6×0.075=0.09+0.124+0.06+0.1364+0.183 6 + 0.0868 + 0.075 + 0.0992 + 0.045 = 0.90, rounded to two decimal places; the vertical gradient matrix is similarly convolved to obtain the smoothed horizontal gradient matrix (protrusion region [[0.90, 0.82, 0.75, -0.32], [0.85, 0.78, 0.70, -0.35], [0.80, 0.72, 0.65, -0.38], [0.75, 0.68, 0.60, -0.40]]) and the smoothed vertical gradient matrix. Verification using a noise detection tool shows that the difference between adjacent pixels in the smoothed gradient matrix is ≤0.1 (the difference in the original matrix is ≤0.3), and the noise suppression rate is ≥60%, ensuring that the smoothed gradient matrix retains the core gradient information while eliminating high-frequency noise interference.
[0074] Furthermore, based on the magnitude and direction of the gradient values in the smooth gradient matrix, the regions of significant changes in cell characteristics are determined; In this embodiment of the invention, regions of significant change in cell features are determined using a region recognition tool based on a smooth gradient matrix. The gradient magnitude |G| = √(Gx² + Gy²) of each pixel is calculated, and a threshold value is set (calibrated based on 100 sets of cell gradient data: 0.6 for cell membrane protrusions and 0.4 for the nucleus). Pixels with |G| ≥ the threshold are classified as regions of significant change. 32×32 scale cell membrane protrusion region: In the smooth gradient matrix, pixel (0,0)|G|=√(0.90²+(-0.85)²)=√(0.81+0.7225)=√1.5325≈1.24≥0.6, pixel (0,3)|G|=√((-0.32)²+(-0.28)²)=√(0.1024+0.0784)=√0.1808≈0.42<0.6. The region of significant change is the 3×3 pixel range in the upper left corner of the protrusion region (containing 9 pixels, all with |G|≥0.6), corresponding to the core structure of the cell membrane protrusion. 16×16 scale cell nucleus region: Pixels with |G|≥0.4 are concentrated in the central 5×5 range (containing 25 pixels), corresponding to the cell nucleus and nuclear membrane region, and are identified as the region of significant change. Verification using structural matching tools showed that the overlap between significantly changed regions and the actual cell structures (protrusions, nucleus) was ≥90%, ensuring that the region division accurately corresponded to the physical location of the significant changes in cell characteristics.
[0075] Furthermore, within regions of significant change, offset parameters for the convolution kernel sampling positions are generated according to a preset scaling factor.
[0076] In this embodiment of the invention, offset parameters of the convolution kernel sampling positions are generated using an offset generation tool within a significantly varying region, according to a preset scaling factor. The scaling factor is set (calibrated based on the correlation between gradient magnitude and sampling offset: 0.3 for the cell membrane protrusion region, 0.2 for the cell nucleus region), the offset magnitude = scaling factor × gradient magnitude, and the offset direction is consistent with the gradient direction (θ = arctan(Gy / Gx)). Within the significantly varying region of the 32×32 scale cell membrane protrusion region, the pixels corresponding to the 9 sampling points of the 3×3 convolution kernel are: the pixel corresponding to the center sampling point is (1,1) (|G| = √(0.78² + (-0.72)²) = √(0.6084 + 0.5184) = √1.1268 ≈ 1.06, θ = arctan((-0.72) / 0.78) ≈ -42°), the offset magnitude = 0.3 × 1.06 ≈ 0.32, offset (x,y)=(0.32×cos(-42°),0.32×sin(-42°))≈(0.24,-0.21); (1,0) sampling point corresponds to pixel (1,0) (|G|=0.85, θ=-45°), offset≈(0.3×0.85×cos(-45°),0.3×0.85×sin(-45°))≈(0.18,-0.18). Within the 16×16 scale cell nucleus region with significant changes, the offset of the 5×5 convolution kernel sampling point =0.2×|G|×(cosθ,sinθ), such as the offset of the center sampling point ≈(0.15,0.15). All offsets are limited to the range of -1 to 1 pixel, generating a "sampling point-offset (x,y)" parameter table to ensure that the offsets can accurately adapt to the feature gradients of significantly changing regions, providing a basis for adjusting the sampling position of subsequent deformable convolution kernels.
[0077] Furthermore, generating the offset parameter of the convolution kernel sampling position according to the preset scaling factor includes the following steps: Set a base offset range, which is determined based on the resolution of the imaging flow cytometer system and the average size of the cells; In this embodiment of the invention, the basic offset range is determined based on the parameters of the imaging flow cytometer system and the average cell size. The system resolution is 0.1 μm / pixel (i.e., 1 pixel corresponds to an actual cell size of 0.1 μm). The average length of leukocyte protrusions detected in the 32×32 scale cell membrane protrusion region is 2 μm (corresponding to 20 pixels) and the width is 0.5 μm (corresponding to 5 pixels). To ensure that the offset can cover the morphological changes of the protrusions and does not exceed the structural range, the basic offset range for this region is set to -0.5 to 0.5 pixels (corresponding to an actual size of -0.05 to 0.05 μm, accounting for 10% of the protrusion width, avoiding excessive offset that would cause sampling to deviate from the protrusion structure). The average diameter of leukocyte nuclei detected in the 16×16 scale cell nucleus region is 8 μm (corresponding to 80 pixels) and the nuclear membrane thickness is 0.3 μm (corresponding to 3 pixels). The basic offset range is set to -0.3 to 0.3 pixels (corresponding to an actual size of -0.03 to 0.03 μm, accounting for 10% of the nuclear membrane thickness, ensuring that the sampling focuses on the nuclear membrane features). Verification using system calibration tools showed that the basic offset range was ≤10% of the cell structure size, and within the allowable range of system imaging accuracy (±0.02μm), ensuring that the offset could both adapt to cell morphology and comply with system hardware limitations.
[0078] Furthermore, the gradient value of each pixel in the smooth gradient matrix is normalized to obtain the gradient normalized value. In this embodiment of the invention, the gradient magnitude |G| of each pixel in the obtained smooth gradient matrix is normalized, and a normalization tool is used to map the gradient magnitude to the interval between 0 and 1. The range of |G| in the smooth gradient matrix of the 32×32 scale cell membrane protrusion region is 0.42 to 1.24. The normalized value is calculated using the formula "normalized value = (|G| - minimum value) / (maximum value - minimum value)": for pixel (0,0), |G| = 1.24, normalized value = (1.24 - 0.42) / (1.24 - 0.42) = 0.82 / 0.82 = 1.0; for pixel (0,3), |G| = 0.42, normalized value = 0; for pixel (1,1), |G| = 1.06, normalized value = (1.06 - 0.42) / (1.24 - 0.42) / (1.24 - 0.42) = 0.82 / 0.82 = 1.0. 0.42) / 0.82=0.64 / 0.82≈0.78. The |G| range of the smoothed gradient matrix in the 16×16 scale cell nucleus region is 0.25 to 0.98. For pixel (4,4), |G|=0.98 (maximum value), normalized value=1.0; for pixel (0,0), |G|=0.25 (minimum value), normalized value=0; for pixel (3,3), |G|=0.72, normalized value=(0.72-0.25) / (0.98-0.25)=0.47 / 0.73≈0.64. All normalized values are rounded to two decimal places and are checked for range (ensuring all values are in the range of 0 to 1, with a deviation ≤0.01) to ensure that the normalized values can uniformly reflect the relative magnitude of the gradient, providing a unified weight basis for subsequent offset calculations.
[0079] Furthermore, the gradient normalization value is multiplied by the basic offset range to obtain the candidate offset for each pixel. In this embodiment of the invention, a candidate offset for each pixel is generated by combining the gradient normalization value with the basic offset range using an offset calculation tool. The basic offset range is divided symmetrically by positive and negative directions, with the positive direction corresponding to the gradient direction and the negative direction corresponding to the opposite direction of the gradient. The formula is "candidate offset = normalization value × upper limit of basic offset (positive direction)" or "candidate offset = -normalization value × lower limit of basic offset (negative direction)". For the 32×32 scale cell membrane protrusion region, the upper limit of the basic offset is 0.5 pixels. Pixel (0,0) has a normalized value of 1.0 and a gradient direction of -42° (negative direction), so the candidate offset is -1.0×0.5=-0.5 pixels. Pixel (1,1) has a normalized value of 0.78 and a gradient direction of -42°, so the candidate offset is -0.78×0.5≈-0.39 pixels. Pixel (0,3) has a normalized value of 0, so the candidate offset is 0. For the 16×16 scale cell nucleus region, the upper limit of the basic offset is 0.3 pixels. Pixel (4,4) has a normalized value of 1.0 and a gradient direction of 45° (positive direction), so the candidate offset is 1.0×0.3=0.3 pixels. Pixel (3,3) has a normalized value of 0.64 and a gradient direction of 45°, so the candidate offset is 0.64×0.3≈0.19 pixels. All candidate offsets are rounded to two decimal places and limited to the range of the base offset (e.g., candidate offsets for protrusion areas are all between -0.5 and 0.5 pixels) to ensure that the candidate offsets are both related to the gradient intensity and comply with the preset range limit.
[0080] Furthermore, cluster analysis was performed on the candidate offsets, and the candidate offsets with the highest frequency of occurrence were selected as the offset parameters of the convolution kernel in the region of significant change.
[0081] In this embodiment of the invention, cluster analysis is performed on the candidate offsets of all pixels within the significantly changing region using a clustering tool (K-means algorithm, K value set to 3, calibrated based on 100 sets of cell offset data). The candidate offsets for the significantly changing region (9 pixels) of the 32×32 scale cell membrane protrusion area are: -0.5, -0.45, -0.39, -0.42, -0.38, -0.40, -0.35, -0.43, -0.37. After clustering, three clusters were obtained: cluster 1 (-0.5, -0.45, frequency 2), cluster 2 (-0.42, -0.43, -0.40, frequency 3), and cluster 3 (-0.39, -0.38, -0.37, -0.35, frequency 4). Cluster 3 had the highest frequency (4 times). The average offset of cluster 3 (-0.39-0.38-0.37-0.35) / 4 ≈ -0.37 pixels was taken as the offset parameter of the convolution kernel in this region. After clustering the candidate offsets of the 16×16 scale cell nucleus region with significant changes (25 pixels), cluster 2 (0.18, 0.19, 0.20, 0.17, 0.16, frequency 5) had the highest frequency. The average value ≈ 0.18 pixels was taken as the offset parameter. By verifying the frequency (the highest cluster frequency accounts for ≥30% of the total number of pixels), we ensure that the selected offset parameters can represent the gradient characteristics of most pixels in the region, providing a stable basis for adjusting the sampling position of the deformable convolution kernel. Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.
[0082] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A cell detection method based on an improved model using imaging flow cytometry, characterized in that, Includes the following steps: The cell sample to be tested is introduced into an imaging flow cytometer system integrated with a microfluidic chip. A custom-developed scaly reflector cup provides high-collimation and high-brightness white light Köhler illumination. At the same time, a high-speed camera is activated to continuously acquire images of the cells in a high-speed flow state, generating an initial cell image sequence. An automatic digital focusing algorithm is applied to the initial cell image sequence to evaluate the sharpness of each frame in real time based on a multi-scale image sharpness evaluation function, and the set of cell image frames with the optimal focal plane is selected. The set of cell image frames with the optimal focal plane is input into the preset PA-YOLO improved model. By introducing multi-scale feature fusion of the neck structure, the corresponding cell image features in the set of cell image frames are extracted and fused in multiple dimensions to generate a multi-scale cell feature map. The deformable convolutional module of the PConv part in the PA-YOLO improved model is used to perform refined feature learning on multi-scale cell feature maps, which enhances the feature representation of small, dense and morphologically similar cells, and obtains enhanced cell feature data. Based on the enhanced cell feature data, the detection head of the PA-YOLO improved model is used to locate and classify cell targets, and output cell detection results, including cell type, number and morphological parameter information. The validity of the cell detection results is verified, and the results are compared and analyzed with the raw image data acquired by the imaging flow cytometer system to correct detection errors and generate a cell detection report.
2. The imaging flow cytometry cell detection method based on the improved model according to claim 1, characterized in that, The method of providing high collimation and high brightness white Köhler illumination through a custom-developed scale-like focusing reflector cup includes the following steps: The light source module of the scale-shaped light-concentrating reflector cup is activated to generate an initial white light beam; The initial white light beam is focused and collimated by the focusing structure of the scale-shaped light-concentrating reflector cup to form a highly collimated parallel beam; A parallel beam is introduced into the Köhler illumination optical path, and the aperture component in the Köhler illumination optical path is adjusted so that the parallel beam is uniformly irradiated onto the cell sample flowing through the detection area of the microfluidic chip, generating an illumination light field that meets the requirements of high-speed imaging. An illumination field is applied to a high-speed flowing cell sample, causing the cells to form a cell image with high contrast and high signal-to-noise ratio on the imaging plane of a high-speed camera.
3. The imaging flow cytometry cell detection method based on the improved model according to claim 1, characterized in that, The application of the automatic digital focusing algorithm to the initial cell image sequence includes the following steps: Pixel grayscale information of each frame of the initial cell image sequence is extracted to construct an image grayscale matrix; The sharpness evaluation function is calculated based on the image grayscale matrix, including the Laplacian variance and the frequency domain energy distribution. The sharpness evaluation function values of each frame are compared, and the frame with the largest sharpness evaluation function value is determined as the candidate frame with the best focal plane in the sequence. Local region sharpness verification is performed on candidate frames to further calculate their focus quality by analyzing the sharpness of cell edge contours, and the set of cell image frames with the optimal focal plane is selected by comparing with a preset focus quality threshold.
4. The imaging flow cytometry cell detection method based on the improved model according to claim 3, characterized in that, The calculation of the sharpness evaluation function value based on the image grayscale matrix includes the following steps: Multi-scale Gaussian filtering is applied to the image grayscale matrix to obtain smooth image matrices at different scales; For smooth image matrices at different scales, calculate their Laplacian second derivative matrices, and then sum the squares of all elements in the Laplacian second derivative matrices to obtain the Laplacian variance values for the corresponding scales. The image grayscale matrix is transformed to the frequency domain space to obtain the frequency domain spectrum matrix. High-frequency components in the frequency domain spectrum matrix are extracted by setting a high-frequency threshold. The sum of the energy of the high-frequency components is calculated to obtain the frequency domain energy distribution value at the corresponding scale. By combining the Laplacian variance values at various scales and the frequency domain energy distribution values, a sharpness evaluation function value for the image frame is generated.
5. The imaging flow cytometry cell detection method based on the improved model according to claim 4, characterized in that, The step of performing local region sharpness verification on candidate frames to further calculate their focus quality by analyzing the sharpness of cell edge contours, and then filtering out the set of cell image frames with the optimal focal plane by comparing them with a preset focus quality threshold, includes the following steps: The candidate frames are divided into local regions, and multiple key local regions containing cell targets are selected. The cell edge contour data of each key local region is extracted. The cell edge contour data is processed by an edge detection algorithm, and the gray-scale change rate of the edge contour is calculated. This gray-scale change rate is used to characterize the sharpness of the cell edge contour. The focus quality score of the candidate frame is calculated based on the sharpness of the cell edge contour, and the focus quality score is compared with the preset focus quality threshold. If the focus quality score is greater than or equal to the preset focus quality threshold, the candidate frame is determined to be the optimal image frame for the focal plane. If the focus quality score is less than the preset focus quality threshold, the image frame with the second largest sharpness evaluation function value is selected from the initial cell image sequence as a new candidate frame, and the above local region sharpness verification steps are repeated until image frames that meet the focus quality requirements are selected, thereby forming a set of cell image frames with optimal focal plane.
6. The imaging flow cytometry cell detection method based on the improved model according to claim 1, characterized in that, The method of multi-dimensionally extracting and fusing cell image features corresponding to the cell image frame set by introducing multi-scale feature fusion of the neck structure includes the following steps: The set of cell image frames with the optimal focal plane is input into the backbone feature extraction network of the PA-YOLO improved model to obtain the original feature maps at different levels; Upsampling and downsampling are performed on the original feature map to generate feature maps at multiple scales; By using multi-scale feature fusion and skip connections and feature splicing operations in the neck structure, feature maps of multiple scales are fused to enhance the correlation between them. The fused feature map is subjected to convolutional operations for dimensionality adjustment and feature optimization, generating a multi-scale cell feature map containing multi-scale cell information.
7. The imaging flow cytometry cell detection method based on the improved model according to claim 1, characterized in that, The process of using the PConv deformable convolutional module in the PA-YOLO improved model to perform refined feature learning on multi-scale cell feature maps includes the following steps: The multi-scale cell feature map is divided into multiple local feature regions, and each local feature region corresponds to a set of convolutional kernel parameters. The PConv deformable convolution module dynamically adjusts the sampling position and weight coefficients of the convolution kernel parameters based on the cell morphology distribution characteristics of local feature regions. By performing convolution operations on local feature regions using adjusted convolution kernel parameters, targeted local cell features can be extracted. By integrating the extraction results of all local feature regions, enhanced cellular feature data that can accurately characterize subtle cellular morphological differences are obtained.
8. The imaging flow cytometry cell detection method based on the improved model according to claim 7, characterized in that, The dynamic adjustment of the sampling position and weight coefficients of the convolution kernel parameters includes the following steps: Feature gradients are calculated for local feature regions to determine the distribution direction of cell edges and textures, thereby generating corresponding feature gradient information. An offset parameter is generated based on the feature gradient information. This offset parameter is used to adjust the sampling coordinates of the convolution kernel in the spatial dimension. Different weighting coefficients are assigned to each sampling point of the convolution kernel based on the importance of cell pixels in the local feature region; The offset parameter and weight coefficients are applied to the original convolution kernel to form a deformed convolution kernel that adapts to the local feature region.
9. The imaging flow cytometry cell detection method based on the improved model according to claim 8, characterized in that, The process of generating offset parameters based on feature gradient information includes the following steps: Calculate the gradient values of each pixel in the local feature region in the horizontal and vertical directions, and construct the gradient matrix; Gaussian smoothing is applied to the gradient matrix to suppress noise interference, resulting in a smoothed gradient matrix. Based on the magnitude and direction of the gradient values in the smooth gradient matrix, the regions of significant changes in cell characteristics are determined. Within regions of significant change, offset parameters for the convolution kernel sampling positions are generated according to a preset scaling factor.
10. The imaging flow cytometry cell detection method based on the improved model according to claim 9, characterized in that, The process of generating the offset parameters of the convolution kernel sampling position according to a preset scaling factor includes the following steps: Set a base offset range, which is determined based on the resolution of the imaging flow cytometer system and the average size of the cells; The gradient value of each pixel in the smooth gradient matrix is normalized to obtain the gradient normalization value. Multiply the gradient normalization value by the base offset range to obtain the candidate offset for each pixel. Cluster analysis was performed on the candidate offsets, and the candidate offsets with the highest frequency were selected as the offset parameters of the convolution kernel in the region of significant change.
Citation Information
Patent Citations
Blood cell image detection method based on YOLOv8 improved model
CN119942540A
Cited By
Sputum smear compliance detection method based on fragmentation processing and AI model
CN121810666A
Strain morphology screening system and method based on image technology
CN121937447A
Stacked cell hierarchy judgment method and system based on multi-focal-length definition analysis
CN122157255A
Microfluidic droplet array visual detection system and method based on fermentation strain screening
CN122385618A