YOLOv8 cryptococcus polymorphic detection and counting method and system based on hypergraph calculation

Through the YOLOv8 model based on hypergraph calculation, combined with data enhancement and attention mechanism, the problems of low efficiency and high false detection rate in Cryptococci detection and counting are solved, and adaptive detection and automated counting of Cryptococci multimorphic Cryptococci are realized, improving the accuracy and efficiency of detection and counting.

CN120279014AActive Publication Date: 2025-07-08NANCHANG UNIV

Patent Information

Application Number
CN202510757179.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The prior art is inefficient and has a high error detection rate in Cryptococci detection and counting, which cannot meet clinical needs. This is mainly due to significant differences in Cryptococci morphology and insufficient image data, making it difficult to achieve accurate Cryptococci detection and counting.

Method used

The YOLOv8 model based on hypergraph calculation is used to amplify the Cryptococcal image dataset through offline and online data augmentation, and an attention mechanism and hypergraph calculation are introduced into the model to enhance the perception of Cryptococcal multimorphicity, and to combine density map estimation to achieve automated counting.

Benefits of technology

It improves the accuracy of Cryptococcus detection and counting accuracy, can provide clinical diagnostic information in a timely and effective manner, and expands intelligent detection and counting applied to other opportunistic fungal infections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279014A_ABST
    Figure CN120279014A_ABST
Patent Text Reader

Abstract

The invention discloses a YOLOv8 cryptococcus polymorphic detection and counting method and system based on hypergraph calculation, the method uses a YOLOv8 model based on hypergraph calculation to carry out cryptococcus polymorphic detection and counting, and the method comprises the following steps: obtaining a cryptococcus image to be detected and counted; the CSPDarknet backbone network is used for carrying out feature extraction on the cryptococcus polymorphic image; the neck network based on hypergraph calculation performs adaptive detection on cryptococcus with different morphological characteristics, and introduces an attention mechanism to enhance the perception ability of the model to a target; the head network generates a density estimation graph; the adaptive density fusion network aggregates coarse-grained cryptococcus density information; optimizing density estimation based on a loss function of distributed supervision; and outputting detection and counting results. According to the method, adaptive detection of the cryptococcus with different morphological characteristics is realized through hypergraph calculation, automatic counting of the cryptococcus is realized through counting of the density map estimation graph, and the accuracy of feature counting of significant morphological differences and dense small targets presented by the cryptococcus image is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and medical image detection, and more specifically, to a method and system for detecting and counting multi-morphological cryptococcus based on hypergraph computing in YOLOv8. Background Art

[0002] Cryptococcal infections can lead to diseases such as cryptococcal meningitis and pulmonary cryptococcosis. In recent years, the prevalence of such diseases has increased significantly, and the atypical symptoms presented by such diseases result in extremely high subjective misdiagnosis rates and treatment delay rates, leading to a cure rate of only 40%. Therefore, how to accurately and timely diagnose cryptococcal infections is an urgent task in clinical practice.

[0003] Currently, it is possible to determine whether a patient has cryptococcosis and the severity of the patient's illness by identifying the content of cryptococcus in the patient's cerebrospinal fluid, providing a timely and effective diagnostic plan for subsequent treatment. However, the manual microscopy method used clinically for detecting and counting cryptococcus has low efficiency, and due to the significant morphological differences of cryptococcus and the presence of a large number of dense small targets, the misdetection rate is extremely high, making the existing detection and counting methods for cryptococcus unable to meet the needs of clinical medicine.

[0004] In recent years, the field of medical artificial intelligence has made breakthroughs in microscopic image analysis. In particular, deep learning methods based on feature engineering have demonstrated superior performance in tasks such as blood cell classification and pathological section analysis. However, the detection and counting of cryptococcus face special challenges: 1) Due to the requirements of medical experts for marking and the huge costs and manpower consumed in medical imaging, it is impossible to construct a standard database that can contain a large number of cryptococcus images; 2) The morphological characteristics of cryptococcus are significantly different; 3) There are a large number of cases containing dense small targets in cryptococcus images. Summary of the Invention

[0005] In view of the above limitations of the prior art or the need for improved technologies, the present invention proposes a method and system for detecting and counting multi-morphological cryptococcus based on hypergraph computing in YOLOv8. Based on the characteristics of significant morphological differences of cryptococcus and less image data, a data augmentation method that can enrich cryptococcus image data is designed to expand the cryptococcus image dataset. Hypergraph computing is incorporated into the YOLOv8 model to achieve adaptive detection of cryptococcus with different morphological characteristics. An attention mechanism is introduced into the neck network to enhance the model's perception ability of multi-morphological cryptococcus, and an automated counting of cryptococcus is achieved through density map estimation.

[0006] In a first aspect, the present invention provides a YOLOv8 polymorphic cryptococcal detection and counting method based on hypergraph computing, using a YOLOv8 model based on hypergraph computing to detect and count polymorphic cryptococci. According to the characteristics of polymorphic cryptococcal images, before model training, permanent amplification data is generated by offline data enhancement, and a negative sample data set is added to form a total data set. During the model training process, random amplification data is generated by online data enhancement. Based on the data set after online data enhancement, a YOLOv8 model including an input module, a CSPDarknet backbone network, a neck network based on hypergraph computing, a head network, an adaptive density fusion network and an output module is constructed, wherein the neck network based on hypergraph computing is composed of four units: a semantic collection unit, a hypergraph computing unit, a semantic scattering unit and an attention-enhanced bottom-up unit. The YOLOv8 polymorphic cryptococcal detection and counting method based on hypergraph computing comprises: In the input module, an image of cryptococci to be detected and counted is obtained; In the CSPDarknet backbone network, feature extraction is performed on cryptococcal images to obtain multi-scale feature maps for cryptococcal images; In the neck network based on hypergraph computing, cryptococci with different morphological characteristics are adaptively detected, and the perception ability of polymorphic cryptococci is enhanced by introducing an attention mechanism; In the head network, two 3×3 convolution blocks and one 1×1 convolution layer are used in each head to generate a coarse-grained density map and a density estimation map at the first resolution level; In the adaptive density fusion network, the coarse-grained density map is gradually refined in an adaptive fusion manner at the density level, and the coarse-grained cryptococcal density information is introduced into the density estimation map at the first resolution level; A multi-level distributed supervision strategy is adopted to optimize the density maps of each scale generated by the adaptive density fusion network through a combined loss function. In the output module, the cryptococcal detection and counting results are output.

[0007] In a second aspect, the present invention provides a YOLOv8 polymorphic cryptococcal detection and counting system based on hypergraph computing, which is used to implement the aforementioned YOLOv8 polymorphic cryptococcal detection and counting method based on hypergraph computing. The system has a YOLOv8 model based on hypergraph computing, and the YOLOv8 model includes: An input module, used for acquiring an image of cryptococci to be detected and counted; CSPDarknet backbone network, used to extract features from cryptococcal images to obtain multi-scale feature maps for cryptococcal images; The neck network based on hypergraph computing consists of four units: a semantic collection unit, a hypergraph computing unit, a semantic scattering unit, and an attention-enhanced bottom-up unit, which is used for the adaptive detection of cryptococci with different morphological features, and enhances the perception ability of multi-morphological cryptococci by introducing an attention mechanism; The head network is used to generate a coarse-grained density map and a density estimation map at the first resolution level by using two 3×3 convolutional blocks and a 1×1 convolutional layer for each head; The adaptive density fusion network is used to gradually refine the coarse-grained density map in an adaptive fusion manner at the density level, and introduce the coarse-grained cryptococci density information into the density estimation map at the first resolution level; The combined loss function based on distributed supervision is used to adopt a multi-level distributed supervision strategy to optimize the density maps of each scale generated by the adaptive density fusion network through the combined loss function; The output module is used to output the detection and counting results of cryptococci.

[0008] In summary, compared with the prior art by the above technical solutions conceived in the present invention, a method and system for detecting and counting multi-morphological cryptococci based on hypergraph computing of YOLOv8 provided by the present invention mainly have the following beneficial effects: The present invention realizes the automatic detection and counting of cryptococci by constructing a YOLOv8 model based on hypergraph computing, providing an effective reference model for the accurate identification and counting of cryptococci; Based on the characteristics of significant morphological differences and less image data of cryptococci, the present invention generates a large number of cryptococci images through offline data augmentation methods such as random cropping, rotation, and horizontal flipping to expand the cryptococci image dataset, and adds online data augmentation methods such as perspective transformation and vertical flipping on the basis of the built-in data augmentation method of the YOLOv8 model to enrich the cryptococci image data, improving the model's recognition ability for different morphological cryptococci in clinical practice; Based on the characteristics of significant morphological differences of cryptococci, the present invention models and learns the high-order correlation between visual features through a YOLOv8 model integrating hypergraph computing to obtain features with high-order perception ability, realizing the adaptive detection of cryptococci with different morphological features; Aiming at the characteristics of a large number of dense small targets in cryptococci images, the present invention enhances the model's perception ability for small targets by introducing an attention mechanism in the neck network based on hypergraph computing, and at the same time realizes the automatic counting of cryptococci images by counting the density estimation map. The designed adaptive density fusion network and the combined loss function based on distributed supervision improve the counting accuracy of the model, providing timely and effective detection information for clinical diagnosis; The YOLOv8 model designed by the present invention based on hypergraph computing can not only be used for accurately and quickly detecting and counting Cryptococcus, but also can be extended to the intelligent detection and counting of other opportunistic fungal infections, which is conducive to large-scale popularization and application on clinical detection devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 It is a flow diagram of a method for detecting and counting multi-form Cryptococcus of YOLOv8 based on hypergraph computing provided by the present invention; Figure 2 It is a structural diagram of a YOLOv8 model based on hypergraph computing provided by the present invention; Figure 3 It is a structural diagram of a neck network based on hypergraph computing provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0010] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0011] First, some technical terms in the present invention are explained.

[0012] The density level refers to the semantic abstraction degree and spatial resolution of the density map generated by the network at different depths or branches. The density level of the deep layer or low-resolution branch is more "coarse" and is used to capture the global target distribution; the density level of the shallow layer or high-resolution branch is more "fine" and is used to capture local details. The division of the density level can be reflected by the downsampling multiple. The density level with a downsampling rate in the range of 1 / 16 to 1 / 32 is a low-density level, and the density level with a downsampling rate in the range of 1 / 2 to 1 / 8 is a high-density level.

[0013] The coarse-grained density map refers to the density map generated at the low-resolution (high downsampling rate) level, and the downsampling rate range is from 1 / 16 to 1 / 32.

[0014] The coarse-grained Cryptococcus density information refers to the Cryptococcus density information in the coarse-grained density map.

[0015] The resolution level refers to the resolution level of the density estimation map. The present invention divides the density estimation map into two levels, namely the first resolution level and the second resolution level. Among them, the downsampling rate in the range of 1 / 2 to 1 / 8 is the first resolution level, and the downsampling rate in the range of 1 / 16 to 1 / 32 is the second resolution level.

[0016] Next, please refer toFigures 1 to 3 , a method for detecting and counting polymorphic Cryptococcus neoformans based on hypergraph computing provided by the present invention, uses a YOLOv8 model based on hypergraph computing for detecting and counting polymorphic Cryptococcus neoformans, which is applicable to the problem of Cryptococcus neoformans image detection and counting.

[0017] Based on the characteristics of significant morphological differences and less image data of Cryptococcus neoformans, in order to improve the model's recognition ability of different morphological Cryptococcus neoformans clinically, before model training, permanent amplified data is generated through offline data augmentation, and negative sample datasets are added to form a total dataset. During model training, random amplified data is generated through online data augmentation.

[0018] Among them, aiming at the characteristic of less Cryptococcus neoformans image data, for each Cryptococcus neoformans image in the original Cryptococcus neoformans dataset, an offline data augmentation method is adopted to enrich the Cryptococcus neoformans image data. The offline data augmentation method includes random cropping, rotation, and flipping. The specific operations are as follows: for each Cryptococcus neoformans image in the original Cryptococcus neoformans dataset, the first type of new Cryptococcus neoformans image is generated by the offline data augmentation method of random cropping. The size of random cropping can be 640×640, and the present invention does not limit the specific size of random cropping; For each Cryptococcus neoformans image in the original Cryptococcus neoformans dataset, the second type of new Cryptococcus neoformans image is generated by the offline data augmentation method of rotation. The rotation angle can be 90 degrees, and the present invention does not limit the rotation angle and rotation direction; For each Cryptococcus neoformans image in the original Cryptococcus neoformans dataset, the third type of new Cryptococcus neoformans image is generated by the offline data augmentation method of horizontal flipping; The three types of new Cryptococcus neoformans images generated are combined with the original Cryptococcus neoformans dataset to amplify the Cryptococcus neoformans dataset.

[0019] The present invention divides the combined dataset into a training set, a validation set, and a test set according to a ratio of 8:1:1. The present invention does not limit the division ratio. For example, the division ratio can also be set to 7:2:1.

[0020] Aiming at the characteristic of significant morphological differences of Cryptococcus neoformans, after each Cryptococcus neoformans image is loaded into the YOLOv8 model to be trained, perspective transformation and vertical flipping are added as online data augmentation methods on the basis of the built-in data augmentation method of the model to enrich the Cryptococcus neoformans image data.

[0021] After obtaining the total Cryptococcus neoformans dataset, a YOLOv8 model including an input module, a CSPDarknet backbone network, a neck network based on hypergraph computing, a head network, an adaptive density fusion network, and an output module is constructed. Among them, the neck network based on hypergraph computing is composed of four units: a semantic collection unit, a hypergraph computing unit, a semantic scattering unit, and a bottom-up unit with enhanced attention.

[0022] A method for detecting and counting Cryptococcus neoformans in multiple forms based on hypergraph computing of the present invention includes steps S1 to S7.

[0023] S1. In the input module, obtain the Cryptococcus neoformans image to be detected and counted.

[0024] S2. In the CSPDarknet backbone network, perform feature extraction on the Cryptococcus neoformans image to obtain a multi-scale feature map for the Cryptococcus neoformans image.

[0025] Among them, the CSPDarknet backbone network is stacked by 5 blocks in a bottom-up manner. Among them, the first block consists of a CBS unit; the second, third, and fourth blocks have the same structure, each consisting of a CBS unit and a C2f unit; the fifth block consists of a CBS unit, a C2f unit, and an SPPF unit; the CBS unit consists of a convolutional layer with a kernel size of 3, a stride of 2, and a padding of 1, a batch normalization layer, and a SiLU activation function layer; the C2f unit consists of two 1×1 convolutional blocks, a Concat splicing layer, and multiple Bottleneck units; the SPPF unit consists of two 1×1 convolutional blocks, three max pooling layers with a kernel size of 5, a stride of 1, and a padding of 2, and a Concat splicing layer.

[0026] Among them, the step of performing feature extraction on the Cryptococcus neoformans image to obtain a multi-scale feature map for the Cryptococcus neoformans image specifically includes sub-steps S21 to S24.

[0027] S21. In the CBS unit, perform a 2-fold downsampling operation on the Cryptococcus neoformans image to obtain a feature map. S22. In the C2f unit, aggregate multi-scale features by splicing the output feature maps and input feature maps of each Bottleneck unit together, and at the same time perform a convolutional operation to compress the feature map, reducing the computational amount while maintaining or enhancing the expressive ability of the model.

[0028] S23. In the SPPF unit, perform a multi-scale pooling operation on the feature map to capture information of different-sized receptive fields, reducing the computational complexity while maintaining performance. The multi-scale pooling operation includes sub-steps S231 to S234.

[0029] S231. Reduce the channel dimension of the input feature map to half through a 1×1 convolutional block.

[0030] S232. The feature map with the channel dimension reduced to half is successively subjected to two-dimensional operations through three max pooling layers to obtain three feature maps with receptive fields of 5×5, 9×9, and 13×13 respectively; S233. Use a Concat concatenation layer to concatenate the three feature maps obtained from the max pooling 2D operation with the feature map whose channel dimension is reduced by half along the channel dimension to obtain feature information with different-sized receptive fields. S234. Use a 1×1 convolutional block to compress the channel dimension of the feature map after channel dimension concatenation to a specified number. S24. The cryptococcus image passes through five blocks in sequence to obtain five feature maps of different scales.

[0031] S3. In the neck network based on hypergraph computing, adaptively detect cryptococcus with different morphological features and enhance the perception ability of multi-morphological cryptococcus by introducing an attention mechanism.

[0032] Among them, the steps of adaptively detecting cryptococcus with different morphological features and enhancing the perception ability of multi-morphological cryptococcus by introducing an attention mechanism specifically include: The semantic collection unit performs information fusion on the multi-scale feature maps of the CSPDarknet backbone network in the semantic space. The hypergraph computing unit uses the hypergraph computing method to model and learn the potential high-order correlations between visual features to generate features with high-order perception ability, which fuse high-order structural information and semantic information. The semantic scattering unit scatters the high-order structural information into three feature maps of different scales. The attention-enhanced bottom-up unit transfers the detailed information of the shallow feature map to the deep feature map layer by layer through a bottom-up lateral connection path to make up for the missing detailed information in the deep feature map, and introduces an attention mechanism to extract the attention area to enhance the model's perception ability of multi-morphological cryptococcus.

[0033] Next, the sub-steps specifically executed by the semantic collection unit, hypergraph computing unit, semantic scattering unit, and attention-enhanced bottom-up unit will be described respectively.

[0034] S31. The semantic collection unit executes sub-steps 311 to 314.

[0035] S311. Perform 8-fold, 4-fold, and 2-fold downsampling operations on the output feature maps of the first, second, and third blocks of the CSPDarknet backbone network respectively to obtain three feature maps with the same resolution level as the output feature map of the fourth block of the CSPDarknet backbone network.

[0036] S312. Perform a 2-fold upsampling operation on the output feature map of the fifth block of the CSPDarknet backbone network to obtain a feature map with the same resolution level as the output feature map of the fourth block of the CSPDarknet backbone network.

[0037] S313. Concatenate the four feature maps obtained from the upsampling and downsampling operations along the channels with the output feature map of the 4th block of the CSPDarknet backbone network to obtain a feature map that fuses the high-level semantic information of the deep feature map and the detailed information of the shallow feature map.

[0038] S314. Compress the channels of the fused feature map through a 1×1 convolutional block.

[0039] S32. The hypergraph computing unit executes the hypergraph construction sub-step and the hypergraph convolution sub-step.

[0040] S321. Hypergraph construction sub-step: The hypergraph is defined by its vertex set and the overall hyperedge set ; Map the grid-based visual features to the vertex set of the hypergraph, where the feature vector at each spatial position corresponds to a vertex ; To model the neighborhood relationships in the semantic space, use a distance threshold to construct a neighborhood sphere for each vertex, and use the set of vertices within this neighborhood sphere as a hyperedge ; The formula corresponding to the overall hyperedge set is: , , In the above formula, represents the neighboring vertex of vertex in the vertex set , represents the Euclidean distance between vertex and its neighboring vertex .

[0041] It should be noted that in the calculation, the hypergraph is usually represented by its incidence matrix .

[0042] S322. Hypergraph convolution sub-step: To facilitate the transmission of high-order information on the hypergraph structure, use typical spatial-domain hypergraph convolution and additional residual connections for two-stage hypergraph information transmission of vertex features to output the feature map after hypergraph calculation; The first-stage hypergraph information transmission refers to aggregating vertices with similar attributes into hyperedges to generate high-order semantic representations; The second-stage hypergraph information transmission refers to diffusing the semantic information of the hyperedges back to the vertices to update the vertex features to fuse the global context; The formulas corresponding to the two-stage hypergraph information transmission are as follows: , , In the above formula, represents the aggregated feature vector of the hyperedge after the first-stage hypergraph information transmission, and represents the vertex set of the hyperedge , represents the feature vector of the vertex , represents the feature vector of the vertex after the second-stage hypergraph information transmission, and represents the set of incident hyperedges of the vertex Define the hypergraph convolution operations in these two stages, namely the first-stage hypergraph information transmission and the second-stage hypergraph information transmission, as: , In the above formula, represents the hypergraph convolution operation, represents the input feature map, represents the incidence matrix of the hypergraph, represents the inverse matrix of the diagonal matrix of the vertices, and represents the inverse matrix of the diagonal matrix of the hyperedge , and

[0043] S33. The semantic scattering unit executes sub-steps S331 to S332.

[0044] S331. Perform a two-fold upsampling and a two-fold downsampling operation on the output feature map of the hypergraph calculation unit.

[0045] S332. Concatenate the feature maps obtained by upsampling, hypergraph calculation, and downsampling with the output feature maps of the 3rd, 4th, and 5th blocks of the CSPDarknet backbone network respectively, and then pass the three feature maps of different scales obtained by concatenation through a C2f unit, a C2f unit, and a 1×1 convolution block respectively to further enhance the features.

[0046] S34. The bottom-up unit with attention enhancement executes sub-steps S341 to S343.

[0047] S341. Through a bottom-up lateral connection path, layer by layer transfer the detailed information of the shallow feature map with a downsampling rate of 1 / 8 in the output feature map of the semantic scattering unit to the deep feature map with a downsampling rate of 1 / 32 in the output feature map of the semantic scattering unit to make up for the missing detailed information of the deep feature map.

[0048] S342. Add a convolutional block attention unit before each downsampling to extract the attention region, so as to enhance the model's perception ability of Cryptococcus in multiple forms.

[0049] Among them, the convolutional block attention unit is composed of a channel attention sub-unit and a spatial attention sub-unit in series, and its calculation process is as follows: , , , , In the above formula, represents the channel attention weight matrix, represents the input feature map, represents the average pooling operation, represents the maximum pooling operation, represents the convolution operation with a kernel size of 1 for learning the dependencies between channels, represents the convolution operation with a kernel size of 7 for capturing wide-area spatial context information, represents the sigmoid activation function, represents the operation of element-wise multiplication of the input feature map and the attention weight matrix, represents the operation of concatenating two feature maps along the channels, represents the output feature map of the channel attention sub-unit, represents the spatial attention weight matrix, represents the output feature map of the spatial attention sub-unit.

[0050] S343. Obtain three feature maps with different scales after attention enhancement, which are respectively used to detect small-target Cryptococcus, medium-target Cryptococcus, and large-target Cryptococcus.

[0051] S4. In the head network, use two 3×3 convolutional blocks and a 1×1 convolutional layer in each head to generate the density estimation map.

[0052] Among them, the head network includes three heads, and each head is composed of two 3×3 convolutional blocks and a 1×1 convolutional layer.

[0053] The steps of using two 3×3 convolutional blocks and a 1×1 convolutional layer in a single head to generate the density estimation map specifically include sub-steps S41 to S43.

[0054] S41. Use the three different-scale feature maps output by the neck network based on hypergraph computing as input feature maps. The input feature maps are compressed to 128 channels through the first convolutional block with a kernel size of 3, a stride of 1, and a padding of 1, and then pass through the ReLU activation function to obtain a 128-dimensional intermediate feature map. .

[0055] S42. The 128-dimensional intermediate feature map passes through the second convolutional block with a kernel size of 3, a stride of 1, and a padding of 1 to further reduce the number of channels to 64, and then passes through the ReLU activation function to obtain a 64-dimensional refined feature map. .

[0056] S43. The 64-dimensional refined feature map is compressed to 1 channel through a 1×1 convolutional layer to obtain three density estimation maps with different resolution levels. , , .

[0057] S5. In the adaptive density fusion network, gradually refine the coarse-grained density map in an adaptive fusion manner at the density level, and introduce the coarse-grained cryptococcus density information into the density estimation map of the first resolution level.

[0058] The step of gradually refining the coarse-grained density map in an adaptive fusion manner at the density level and introducing the coarse-grained cryptococcus density information into the density estimation map of the first resolution level specifically includes sub-steps S51 to S52.

[0059] Sub-step S51. In the adaptive density fusion network, use the gradient optimization method to obtain the weight coefficients of different density estimation maps during the density fusion process.

[0060] Sub-step S52. Gradually aggregate the coarse-grained cryptococcus density information into the density estimation map of the first resolution level through an adaptive fusion method. The specific expression is as follows: , In the above formula, represents the density estimation maps with different resolutions generated by the head network, represents the layer corresponding to the density estimation map , , , represent the three density estimation maps with different resolutions generated by the head network, represents the cryptococcus density estimation map after adaptive fusion, represents bilinear interpolation, represents the weight coefficients of different density estimation maps during the density fusion process.

[0061] S6. Adopt a multi-level distributed supervision strategy to optimize the density maps of each scale generated by the adaptive density fusion network through a combined loss function.

[0062] Among them, the combined loss includes a counting loss, an optimal transport loss, and a total variation loss, and its calculation steps include sub-steps S61 to S64.

[0063] S61. Calculate the counting loss: The counting loss is calculated by using the L1 loss between the true count and the predicted count to ensure that the predicted density estimation map is consistent with the true count in terms of global count. The specific expression is as follows: , In the above formula, represents the counting loss on the density estimation map , represents the level corresponding to the density estimation map , represents the true density map with the same size as the density estimation map , represents the pixel value corresponding to the pixel position in the density estimation map , represents the pixel value corresponding to the pixel position in the true density map , represents the abscissa of the pixel position , represents the ordinate of the pixel position .

[0064] S62. Calculate the optimal transport loss: The optimal transport loss is calculated by using the Sinkhorn algorithm to find the distribution difference between the predicted Cryptococcus distribution and the true Cryptococcus distribution that minimizes the transport cost, so as to solve the problem of local distribution mismatch of the density estimation map. Its calculation steps include sub-steps S621 to S622.

[0065] S621. Regard the density map as a probability distribution: Normalize the density estimation map and the true density map so that their sum is 1.

[0066] S622. Calculate the transport cost matrix: Define the square of the Euclidean distance between pixels as the transport cost matrix , and its formula is: , In the above formula, represents the true density map The pixel position The corresponding pixel value Indicates the abscissa of the pixel position at Indicates the ordinate of the pixel position at

[0067] S623. Solve the optimal transport plan: Under the transport constraint conditions, the optimal transport plan is obtained by minimizing the total transport cost .

[0068] Among them, the formula corresponding to minimizing the total transport cost is , In the above formula Indicates the optimal transport loss on the density estimation map ; Indicates the transport plan from the pixel position of the density estimation map to the pixel position of the true density map .

[0069] The formula for the transport constraint conditions is , , In the above formula Indicates that the total transport volume emitted from the pixel position is equal to the predicted density value at that position Indicates that the total transport volume reaching the pixel position is equal to the true density value at that position; S624. Use the Sinkhorn algorithm for efficient approximate solution

[0070] S63. Calculate the total variation loss: The total variation loss is calculated by using the L1 norm of the gradients of the density estimation map in the horizontal and vertical directions to enhance the local smoothness of the predicted density estimation map and reduce noise. The specific expression is as follows , In the above formula Indicates the total variation loss on the density estimation map .

[0071] S64. Calculate the combined loss: Based on the counting loss, optimal transport loss, and total variation loss, calculate the combined loss. The expression of the combined loss is as follows , In the above formula Indicates the combined loss and represent two adjustable hyperparameters.

[0072] S7. In the output module, output the cryptococcus detection and counting results.

[0073] Among them, the step of outputting the cryptococcus detection and counting results specifically includes sub-steps S71 to S73.

[0074] S71. Sum the density estimation map obtained by the improved YOLOv8 model to obtain the cryptococcus counting result in the cryptococcus image.

[0075] S72. If the cryptococcus counting result is less than 1, it is considered that no cryptococcus is detected in the cryptococcus image, and the output cryptococcus counting result is 0.

[0076] S73. If the cryptococcus counting result is greater than or equal to 1, it is considered that cryptococcus is detected in the cryptococcus image, and output the cryptococcus counting result.

[0077] The present invention also provides a YOLOv8 multi-morphology cryptococcus detection and counting system based on hypergraph computing, which is used to implement the aforementioned YOLOv8 multi-morphology cryptococcus detection and counting method based on hypergraph computing. The system has a YOLOv8 model based on hypergraph computing, and the YOLOv8 model includes: An input module, which is used to obtain the cryptococcus image to be detected and counted; A CSPDarknet backbone network, which is used to extract features from the cryptococcus image to obtain a multi-scale feature map for the cryptococcus image; A neck network based on hypergraph computing, which consists of four units: a semantic collection unit, a hypergraph computing unit, a semantic scattering unit, and an attention-enhanced bottom-up unit, and is used to adaptively detect cryptococcus with different morphological features and enhance the perception ability of multi-morphology cryptococcus by introducing an attention mechanism; A head network, which is used to generate a coarse-grained density map and a density estimation map at the first resolution level by using two 3×3 convolutional blocks and a 1×1 convolutional layer in each head; An adaptive density fusion network, which is used to gradually refine the coarse-grained density map in an adaptive fusion manner at the density level and introduce the coarse-grained cryptococcus density information into the density estimation map at the first resolution level; A combined loss function based on distributed supervision, which is used to adopt a multi-level distributed supervision strategy to optimize the density maps at each scale generated by the adaptive density fusion network through the combined loss function; An output module, which is used to output the cryptococcus detection and counting results.

[0078] Through the detailed description of the above-mentioned YOLOv8 multi-morphological cryptococcus detection and counting method based on hypergraph computing, it can be seen that the YOLOv8 multi-morphological cryptococcus detection and counting system based on hypergraph computing of The specific structure, working process, etc. will not be elaborated any further.

[0079] In this embodiment, a YOLOv8 model based on hypergraph computing is trained and generated using a cryptococcus private dataset. The private dataset contains both cryptococcus images and other similar fungal images, and the corresponding proportional relationship is 7:3.

[0080] To deeply elaborate on the superiority of the YOLOv8 model based on hypergraph computing, the mean absolute error (MAE) and mean squared error (MSE) are used as evaluation metrics. The evaluation metric MAE can reflect the counting accuracy and precision, and the evaluation metric MSE can measure the generalization performance and robustness. The evaluation results are shown in Table 1. It can be seen from the results in Table 1 that compared with the traditional YOLOv8s model, the YOLOv8 model based on hypergraph computing in this embodiment has a lower mean absolute error (MAE) and mean squared error (MSE), which reflects the counting accuracy and generalization performance of the YOLOv8 model based on hypergraph computing in this embodiment. It can be considered that it has a good performance in the detection and counting of cryptococcus images.

[0081] Table 1 Image recognition results of different models

[0082] A YOLOv8 multi-morphological cryptococcus detection and counting method provided by the present invention designs a data augmentation method based on random cropping, rotation, flipping, and perspective transformation for the characteristics of less cryptococcus image data, significant morphological differences, and a large number of dense small targets. It effectively expands the cryptococcus dataset, effectively integrates hypergraph computing and the attention mechanism into the traditional YOLOv8 model, realizes the adaptive detection of cryptococcus with different morphological features, enhances the model's perception ability of multi-morphological cryptococcus, and also realizes the automatic counting of cryptococcus images through the counting of the density estimation map, providing a systematic solution for the timely and accurate identification and counting of cryptococcus.

[0083] Those skilled in the art can easily understand that the above is only a preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for detecting and counting polymorphic cryptococcus in YOLOv8 based on hypergraph computing, characterized in that, Use a YOLOv8 model based on hypergraph computing for the detection and counting of Cryptococcus neoformans in multiple forms. In view of the characteristics of Cryptococcus neoformans images in multiple forms, before model training, permanent augmented data is generated through offline data augmentation, and negative sample datasets are added to form a total dataset. During model training, random augmented data is generated through online data augmentation; Based on the dataset after online data augmentation, a YOLOv8 model is constructed, which includes an input module, a CSPDarknet backbone network, a neck network based on hypergraph computing, a head network, an adaptive density fusion network, and an output module. Among them, the neck network based on hypergraph computing is composed of four units: a semantic collection unit, a hypergraph computing unit, a semantic scattering unit, and an attention-enhanced bottom-up unit; The method for detecting and counting Cryptococcus neoformans in multiple forms based on the YOLOv8 model using hypergraph computing includes: In the input module, obtain the Cryptococcus neoformans image to be detected and counted; In the CSPDarknet backbone network, extract features from the Cryptococcus neoformans image to obtain a multi-scale feature map for the Cryptococcus neoformans image; In the neck network based on hypergraph computing, adaptively detect Cryptococcus neoformans with different morphological features, and enhance the perception ability of Cryptococcus neoformans in multiple forms by introducing an attention mechanism; In the head network, use two 3×3 convolutional blocks and a 1×1 convolutional layer for each head to generate a coarse-grained density map and a density estimation map at the first resolution level; In the adaptive density fusion network, gradually refine the coarse-grained density map in an adaptive fusion manner at the density level, and introduce the coarse-grained Cryptococcus neoformans density information into the density estimation map at the first resolution level; Adopt a multi-level distributed supervision strategy to optimize the density maps at each scale generated by the adaptive density fusion network through a combined loss function; In the output module, output the detection and counting results of Cryptococcus neoformans.

2. The method according to claim 1, wherein The steps of augmenting the Cryptococcus neoformans dataset through offline and online data augmentation methods specifically include: In view of the small amount of Cryptococcus neoformans image data, for each Cryptococcus neoformans image in the original Cryptococcus neoformans dataset, use offline data augmentation methods such as random cropping, rotation, and horizontal flipping to enrich the Cryptococcus neoformans image data; In view of the significant morphological differences of Cryptococcus neoformans, when each Cryptococcus neoformans image is loaded into the YOLOv8 model to be trained, online data augmentation methods such as perspective transformation and vertical flipping are added on the basis of the built-in data augmentation method of the model to enrich the Cryptococcus neoformans image data.

3. The method according to claim 1, wherein The CSPDarknet backbone network is stacked by 5 blocks in a bottom-up manner. Among them, the first block consists of a CBS unit; the second, third, and fourth blocks have the same structure, each consisting of a CBS unit and a C2f unit; the fifth block consists of a CBS unit, a C2f unit, and an SPPF unit; the CBS unit consists of a convolutional layer with a kernel size of 3, a stride of 2, and a padding of 1, a batch normalization layer, and a SiLU activation function layer; the C2f unit consists of two 1×1 convolutional blocks, a Concat splicing layer, and multiple Bottleneck units; the SPPF unit consists of two 1×1 convolutional blocks, three max pooling layers with a kernel size of 5, a stride of 1, and a padding of 2, and a Concat splicing layer; the step of extracting features from the cryptococcus image to obtain a multi-scale feature map for the cryptococcus image specifically includes: In the CBS unit, a 2-fold downsampling operation is performed on the cryptococcus image to obtain a feature map; In the C2f unit, the output feature maps and input feature maps of each Bottleneck unit are spliced together to aggregate multi-scale features, and at the same time, a convolution operation is performed to compress the feature map, reducing the computational amount while maintaining or enhancing the expressive ability of the model; In the SPPF unit, a multi-scale pooling operation is performed on the feature map to capture information on receptive fields of different sizes, reducing the computational complexity while maintaining performance. The multi-scale pooling operation is as follows: The channel dimension of the input feature map is reduced to half through a 1×1 convolutional block; The feature map with the reduced channel dimension is successively subjected to two-dimensional operations through three max pooling layers to obtain three feature maps with receptive fields of 5×5, 9×9, and 13×13 respectively; Through the Concat splicing layer, the three feature maps obtained from the max pooling two-dimensional operation are spliced with the feature map with the reduced channel dimension along the channel dimension to obtain feature information with receptive fields of different sizes; The channel dimension of the feature map after channel dimension splicing is compressed to a specified number through a 1×1 convolutional block; The cryptococcus image passes through five blocks in sequence to obtain five feature maps of different scales.

4. The method according to claim 1, wherein The step of adaptively detecting cryptococcus with different morphological features and enhancing the perception ability of multi-morphological cryptococcus by introducing an attention mechanism specifically includes: The semantic collection unit performs information fusion on the multi-scale feature maps of the CSPDarknet backbone network in the semantic space; The hypergraph calculation unit uses the hypergraph calculation method to model and learn the potential high-order correlations between visual features to generate features with high-order perception ability, which integrates high-order structural information and semantic information; The semantic scattering unit scatters the high-order structural information into three feature maps of different scales; The bottom-up unit with enhanced attention transfers the detailed information of the shallow feature maps layer by layer to the deep feature maps through a bottom-up lateral connection path to make up for the missing detailed information in the deep feature maps, and introduces an attention mechanism to extract the attention area to enhance the model's perception ability of Cryptococcus neoformans in multiple forms.

5. The method according to claim 4, wherein The specific sub-steps executed by the semantic collection unit, the hypergraph calculation unit, the semantic scattering unit, and the bottom-up unit with enhanced attention are as follows: The semantic collection unit executes the following sub-steps: In the first step, perform 8-fold, 4-fold, and 2-fold downsampling operations on the output feature maps of the first, second, and third blocks of the CSPDarknet backbone network respectively to obtain three feature maps with the same resolution level as the output feature map of the fourth block of the CSPDarknet backbone network; In the second step, perform 2-fold upsampling on the output feature map of the fifth block of the CSPDarknet backbone network to obtain a feature map with the same resolution level as the output feature map of the fourth block of the CSPDarknet backbone network; In the third step, concatenate the four feature maps obtained by the upsampling and downsampling operations along the channel with the output feature map of the fourth block of the CSPDarknet backbone network to obtain a feature map that fuses the high-level semantic information of the deep feature map and the detailed information of the shallow feature map; In the fourth step, perform channel compression on the fused feature map through a 1×1 convolutional block; The hypergraph calculation unit executes the following sub-steps: Hypergraph construction sub-step: Hypergraph is defined by its vertex set and the overall hyperedge set ; map the grid-based visual features to the vertex set of the hypergraph , where the feature vector at each spatial position corresponds to a vertex ; to model the neighborhood relationship in the semantic space, use a distance threshold to construct a neighborhood sphere for each vertex , and take the set of vertices within this neighborhood sphere as a hyperedge ; the formula corresponding to the overall hyperedge set is as follows: , , In the above formula, represents the vertex in the vertex set of the neighboring vertices, represents the vertex and its neighboring vertices the Euclidean distance between; Hypergraph convolution sub-step: To facilitate the transfer of high-order information on the hypergraph structure, use typical spatial-domain hypergraph convolution and additional residual connections to perform two-stage hypergraph information transfer on the vertex features to output the feature map after hypergraph calculation; The first-stage hypergraph information transfer refers to aggregating vertices with similar attributes into hyperedges to generate high-order semantic representations; the second-stage hypergraph information transfer refers to diffusing the semantic information of the hyperedges back to the vertices to update the vertex features to fuse the global context; the formulas corresponding to the two-stage hypergraph information transfer are as follows: , , In the above formula, represents the aggregated feature vector of the hyperedge after the first-stage hypergraph information transmission, represents the vertex set of the hyperedge represents the vertex feature vector, represents the feature vector of the vertex after the second-stage hypergraph information transmission, represents the vertex associated hyperedge set;​​ Define the hypergraph convolution operations of the first-stage hypergraph information transfer and the second-stage hypergraph information transfer as: , In the above formula, denotes the hypergraph convolution operation, denotes the input feature map, denotes the incidence matrix of the hypergraph, denotes the inverse matrix of the diagonal matrix of vertices, denotes the hyperedge the inverse matrix of the diagonal matrix of denotes the transpose matrix of the incidence matrix of the hypergraph; The semantic scattering unit executes the following sub-steps: In the first step, perform 2-fold upsampling and 2-fold downsampling on the output feature map of the hypergraph calculation unit; In the second step, concatenate the feature maps obtained by upsampling, hypergraph calculation, and downsampling with the output feature maps of the third, fourth, and fifth blocks of the CSPDarknet backbone network respectively, and then pass the three feature maps of different scales obtained by concatenation through a C2f unit, a C2f unit, and a 1×1 convolutional block respectively to further enhance the features; The bottom-up unit with enhanced attention executes the following sub-steps: In the first step, through a bottom-up lateral connection path, transfer the detailed information of the shallow feature map with a downsampling rate of 1 / 8 in the output feature map of the semantic scattering unit layer by layer to the deep feature map with a downsampling rate of 1 / 32 in the output feature map of the semantic scattering unit to make up for the missing detailed information in the deep feature map; In the second step, a convolutional block attention unit is added before each downsampling to extract the attention region, so as to enhance the model's perception ability of Cryptococcus in multiple forms; Step 3: Obtain three feature maps of different scales with enhanced attention, which are respectively used to detect cryptococcus with small targets, medium targets, and large targets. Among them, if the width and height of the annotation box of cryptococcus are set as and , and the width and height of the image are and , then the cryptococcus with small targets refers to the cryptococcus that satisfies , the cryptococcus with medium targets refers to the cryptococcus that satisfies , and the cryptococcus with large targets refers to the cryptococcus that satisfies ; Among them, the convolutional block attention unit is composed of a channel attention subunit and a spatial attention subunit in series, and its calculation process is as follows: , , , , In the above formula, represents the channel attention weight matrix, represents the input feature map, represents the average pooling operation, represents the max pooling operation, represents the convolution operation with a kernel size of 1 for learning channel dependencies, represents the convolution operation with a kernel size of 7 for capturing wide-area spatial context information, represents the sigmoid activation function, represents the operation of element-wise multiplication of the input feature map and the attention weight matrix, represents the concatenation operation of two feature maps along the channel, represents the output feature map of the channel attention sub-unit, represents the spatial attention weight matrix, represents the output feature map of the spatial attention sub-unit.

6. The method according to claim 1, wherein The head network includes three heads, each head is composed of two 3×3 convolutional blocks and a 1×1 convolutional layer. The steps of generating the density estimation map by using the two 3×3 convolutional blocks and a 1×1 convolutional layer in a single head specifically include: Take the three different-scale feature maps output by the neck network based on hypergraph computing as the input feature maps. The input feature maps are compressed to 128 channels through the first convolutional block with a kernel size of 3, a stride of 1, and a padding of 1, and then passed through the ReLU activation function to obtain a 128-dimensional intermediate feature map ; 128-dimensional intermediate feature map Through the second convolutional block with a kernel size of 3, a stride of 1, and a padding of 1, the number of channels is further reduced to 64, and then through the ReLU activation function, a 64-dimensional refined feature map is obtained ; 64-dimensional refined feature map The number of channels is compressed to 1 through a 1×1 convolutional layer, and three density estimation maps with different resolution levels are obtained , , .

7. The method according to claim 1, wherein The steps of gradually refining the coarse-grained density map in an adaptive fusion manner at the density level and introducing the coarse-grained Cryptococcus density information into the density estimation map at the first resolution level specifically include the following sub-steps: In the adaptive density fusion network, a gradient optimization method is used to obtain the weight coefficients of different density estimation maps during the density fusion process; The coarse-grained Cryptococcus density information is gradually aggregated into the density estimation map at the first resolution level in an adaptive fusion manner, and the specific expression is as follows: , In the above formula, represents the density estimation maps with different resolutions generated by the head network, represents the level corresponding to the density estimation map , , represent three density estimation maps with different resolutions generated by the head network, represents the cryptococcus density estimation map after adaptive fusion, represents bilinear interpolation, represents the weight coefficients of different density estimation maps in the density fusion process.​ 8. The method according to claim 1, wherein The combined loss includes a counting loss, an optimal transport loss, and a total variation loss, and its calculation steps include: Calculating the counting loss: The counting loss is calculated by using the L1 loss between the true count and the predicted count to ensure that the predicted density estimation map is consistent with the true count in terms of global count. The specific expression is as follows: , In the above formula, represents the counting loss on the density estimation graph , represents the level corresponding to the density estimation graph , represents the true density graph with the same size as the density estimation graph , represents the pixel value corresponding to the pixel position in the density estimation graph , represents the pixel value corresponding to the pixel position in the true density graph , represents the abscissa at the pixel position , represents the ordinate at the pixel position ; Calculating the optimal transport loss: The optimal transport loss is calculated by finding the distribution difference between the predicted Cryptococcus distribution and the true Cryptococcus distribution with the minimum transport cost through the Sinkhorn algorithm to solve the problem of local distribution mismatch in the density estimation map. Its calculation steps are as follows: First step, regard the density map as a probability distribution: normalize the density estimation map and the true density map so that their sum is 1; Step 2: Calculate the transmission cost matrix: Define the square of the Euclidean distance between pixels as the transmission cost matrix , and its formula is: , In the above formula, represents the true density map at the pixel position corresponding to the pixel value at that position, represents the abscissa at the pixel position at that position, represents the ordinate at the pixel position at that position; Step 3, Solve the optimal transportation plan: Under the transportation constraints, solve for the optimal transportation plan by minimizing the total transportation cost ; where the formula corresponding to minimizing the total transportation cost is: , In the above formula, represents the optimal transport loss on the density estimation map ; represents the transport plan from the pixel position of the density estimation map to the pixel position of the true density map ; The formula for the transport constraint condition is: , , In the above formula, represents that the total transmission amount emitted from the pixel position is equal to the predicted density value at that position, represents that the total transmission amount reaching the pixel position is equal to the true density value at that position; In the fourth step, the Sinkhorn algorithm is used for efficient approximate solution; Calculating the total variation loss: The total variation loss is calculated by using the L1 norm of the gradients of the density estimation map in the horizontal and vertical directions to enhance the local smoothness of the predicted density estimation map and reduce noise. The specific expression is as follows: , In the above formula, represents the total variational loss on the density estimation graph ; Calculating the combined loss: Based on the counting loss, the optimal transport loss, and the total variation loss, calculate the combined loss. The expression of the combined loss is as follows: , In the above formula, represents the combined loss, and represent two adjustable hyperparameters.

9. The method according to claim 1, characterized in that, The steps of outputting the Cryptococcus detection and counting results specifically include: Sum the density estimation maps obtained by the improved YOLOv8 model to obtain the Cryptococcus count result in the Cryptococcus image; If the Cryptococcus count result is less than 1, it is considered that no Cryptococcus is detected in the Cryptococcus image, and the output Cryptococcus count result is 0; If the Cryptococcus count result is greater than or equal to 1, it is considered that Cryptococcus is detected in the Cryptococcus image, and the Cryptococcus count result is output.

10. A YOLOv8 multi-morphological cryptococcus detection and counting system based on hypergraph computing, which is used to implement the YOLOv8 multi-morphological cryptococcus detection and counting method according to any one of claims 1-9, characterized in that, The system has a YOLOv8 model based on hypergraph computing, and the YOLOv8 model includes: An input module for acquiring the Cryptococcus image to be detected and counted; A CSPDarknet backbone network for extracting features from the Cryptococcus image to obtain a multi-scale feature map for the Cryptococcus image; The neck network based on hypergraph computing is composed of four units: a semantic collection unit, a hypergraph computing unit, a semantic scattering unit, and an attention-enhanced bottom-up unit, which is used for the adaptive detection of cryptococcus with different morphological features, and enhances the perception ability of multi-morphological cryptococcus by introducing an attention mechanism; The head network is used to generate a coarse-grained density map and a density estimation map at the first resolution level by using two 3×3 convolutional blocks and a 1×1 convolutional layer for each head; The adaptive density fusion network is used to gradually refine the coarse-grained density map in an adaptive fusion manner at the density level, and introduce the coarse-grained cryptococcus density information into the density estimation map at the first resolution level; The combined loss function based on distributed supervision is used to adopt a multi-level distributed supervision strategy to optimize the density maps at each scale generated by the adaptive density fusion network through the combined loss function; The output module is used to output the cryptococcus detection and counting results.

Citation Information

Patent Citations

  • In-carriage crowd counting method based on position enhancement and multi-scale fusion network

    CN113887489A

  • Cell counting method and device, terminal equipment and medium

    CN115861275A

  • Vehicle counting method based on pyramid density perception attention network

    CN116311091A

  • Cryptococcus image recognition method based on deep convolutional neural network

    CN118247784A

  • Traffic flow prediction method based on space-time causal attention network

    CN119107794A

Cited By

  • Work clothes wearing identification method based on hypergraph calculation

    CN121170505A

  • Fall behavior identification method based on hypergraph depth feature fusion

    CN121236821A

  • Layout hot spot detection method based on multi-feature representation learning and storage medium thereof

    CN121615585A