YOLOv8 polymorphic cryptococcal detection and counting method and system based on hypergraph computing
Through the YOLOv8 model based on hypergraph calculation, combined with data enhancement and attention mechanism, the problems of low efficiency and high false detection rate in Cryptococci detection and counting are solved, and adaptive detection and accurate counting of Cryptococci multimorphic Cryptococci are realized, which is suitable for clinical diagnosis.
Patent Information
- Application Number
- CN202510757179.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The prior art has low efficiency and high error detection rate in Cryptococcal detection and counting, which cannot meet clinical needs, mainly due to significant differences in Cryptococcal morphology and insufficient image data.
The YOLOv8 model based on hypergraph calculation is used to amplify the Cryptococcal image dataset through offline and online data augmentation, and an attention mechanism is introduced into the model to realize adaptive detection and counting of different morphologies of Cryptococcal species.
The model's recognition ability and counting accuracy of multimorphic Cryptococci is improved, and the automated detection and counting of Cryptococci is realized, providing timely and effective detection information for clinical diagnosis.
Smart Images

Figure CN120279014B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and medical image detection technology, and more specifically, to a YOLOv8 polymorphic Cryptococcus detection and counting method and system based on hypergraph computing. Background Art
[0002] Cryptococcal infection can cause diseases such as cryptococcal meningitis and pulmonary cryptococcosis. The prevalence of these diseases has increased significantly in recent years. Their atypical symptoms lead to high rates of subjective misdiagnosis and delayed treatment, resulting in a cure rate of only 40%. Therefore, accurate and timely diagnosis of cryptococcal infection is a pressing clinical challenge.
[0003] Currently, the presence of Cryptococcus in a patient's cerebrospinal fluid can be used to determine whether the patient has cryptococcal disease and the severity of the disease, providing a timely and effective diagnostic plan for subsequent treatment. However, the manual microscopic examination method used in clinical practice to detect and count Cryptococci is inefficient, and the significant morphological differences of Cryptococci and the presence of a large number of densely packed small targets lead to a very high false positive rate, making existing methods for detecting and counting Cryptococci unable to meet clinical medical needs.
[0004] In recent years, the field of medical artificial intelligence has made breakthrough progress in microscopic image analysis. In particular, deep learning methods based on feature engineering have demonstrated superior performance in tasks such as blood cell classification and pathological section analysis. However, the detection and counting of Cryptococcus face special challenges:
[0005] 1) Due to the need for medical expert labeling and the huge cost and manpower required for medical imaging, it is impossible to build a standard database that can contain a large number of Cryptococcus images;
[0006] 2) The morphological characteristics of Cryptococcus differ significantly;
[0007] 3) There are a large number of densely packed small objects in cryptococcal images. Summary of the Invention
[0008] In response to the above limitations of the existing technology or the need for improved technology, the present invention proposes a YOLOv8 polymorphic Cryptococcus detection and counting method and system based on hypergraph computing. Based on the significant morphological differences of cryptococci and the limited image data, a data enhancement method that can enrich cryptococcal image data is designed to expand the cryptococcal image dataset. Hypergraph computing is integrated into the YOLOv8 model to achieve adaptive detection of cryptococci with different morphological characteristics. An attention mechanism is introduced into the neck network to enhance the model's perception of polymorphic cryptococci. Automated cryptococcal counting is achieved through density map estimation.
[0009] In a first aspect, the present invention provides a YOLOv8 polymorphic cryptococcal detection and counting method based on hypergraph computing, which uses a YOLOv8 model based on hypergraph computing to detect and count polymorphic cryptococci. According to the characteristics of polymorphic cryptococcal images, before model training, permanent amplification data is generated by offline data enhancement, and a negative sample data set is added to form a total data set. During the model training process, random amplification data is generated by online data enhancement. Based on the data set after online data enhancement, a YOLOv8 model is constructed, which includes an input module, a CSPDarknet backbone network, a neck network based on hypergraph computing, a head network, an adaptive density fusion network, and an output module. The neck network based on hypergraph computing is composed of four units: a semantic collection unit, a hypergraph computing unit, a semantic scattering unit, and an attention-enhanced bottom-up unit. The YOLOv8 polymorphic cryptococcal detection and counting method based on hypergraph computing includes:
[0010] In the input module, an image of cryptococci to be detected and counted is obtained;
[0011] In the CSPDarknet backbone network, feature extraction is performed on the cryptococcal image to obtain a multi-scale feature map for the cryptococcal image;
[0012] In the neck network based on hypergraph computing, Cryptococcus with different morphological characteristics is adaptively detected, and the perception ability of polymorphic Cryptococcus is enhanced by introducing an attention mechanism;
[0013] In the head network, two 3×3 convolution blocks and one 1×1 convolution layer are used in each head to generate a coarse-grained density map and a density estimation map at the first resolution level;
[0014] In the adaptive density fusion network, the coarse-grained density map is gradually refined in an adaptive fusion manner at the density level, and the coarse-grained cryptococcal density information is introduced into the density estimation map at the first resolution level;
[0015] A multi-level distributed supervision strategy is adopted to optimize the density maps of each scale generated by the adaptive density fusion network through a combined loss function.
[0016] In the output module, the Cryptococcus detection and counting results are output.
[0017] In a second aspect, the present invention provides a YOLOv8 polymorphic cryptococcal detection and counting system based on hypergraph computing, which is used to implement the aforementioned YOLOv8 polymorphic cryptococcal detection and counting method based on hypergraph computing. The system has a YOLOv8 model based on hypergraph computing, and the YOLOv8 model includes:
[0018] An input module is used to obtain an image of cryptococci to be detected and counted;
[0019] CSPDarknet backbone network, used to extract features from cryptococcal images to obtain multi-scale feature maps for cryptococcal images;
[0020] The neck network based on hypergraph computing is composed of four units: semantic collection unit, hypergraph computing unit, semantic scattering unit, and attention-enhanced bottom-up unit. It is used to adaptively detect cryptococci with different morphological characteristics and enhances the perception of polymorphic cryptococci by introducing an attention mechanism.
[0021] The head network is used to generate a coarse-grained density map and a density estimation map of the first resolution level using two 3×3 convolution blocks and one 1×1 convolution layer in each head;
[0022] An adaptive density fusion network is used to gradually refine the coarse-grained density map in an adaptive fusion manner at the density level, and introduce the coarse-grained cryptococcal density information into the density estimation map at the first resolution level;
[0023] A combined loss function based on distributed supervision is used to optimize the density maps of each scale generated by the adaptive density fusion network using a multi-level distributed supervision strategy;
[0024] Output module, used to output cryptococcal detection and counting results.
[0025] In summary, compared with the prior art, the above technical solutions conceived by the present invention provide a YOLOv8 polymorphic cryptococcal detection and counting method and system based on hypergraph computing, which has the following beneficial effects:
[0026] The present invention realizes the automatic detection and counting of Cryptococcus by constructing the YOLOv8 model based on hypergraph computing, providing an effective reference model for the accurate identification and counting of Cryptococcus.
[0027] Based on the significant morphological differences of Cryptococcus and the limited image data, this paper generates a large number of Cryptococcus images through offline data enhancement methods such as random cropping, rotation, and horizontal flipping to expand the Cryptococcus image dataset. In addition, the YOLOv8 model's built-in data enhancement methods are added with online data enhancement methods such as perspective transformation and vertical flipping to enrich the Cryptococcus image data, thereby improving the model's ability to recognize different morphological forms of Cryptococcus in clinical practice.
[0028] Based on the significant morphological differences of Cryptococcus, the present invention integrates the YOLOv8 model with hypergraph computing to model and learn high-order correlations between visual features to obtain features with high-order perception capabilities, thus achieving adaptive detection of Cryptococcus with different morphological characteristics.
[0029] In view of the fact that cryptococcal images contain a large number of densely packed small targets, this paper introduces an attention mechanism into the neck network based on hypergraph computing to enhance the model's ability to perceive small targets. It also implements automated counting of cryptococcal images by counting density estimation maps. The designed adaptive density fusion network and the combined loss function based on distributed supervision improve the model's counting accuracy, providing timely and effective detection information for clinical diagnosis.
[0030] The YOLOv8 model designed in the present invention based on hypergraph computing can not only be used to accurately and quickly detect and count Cryptococcus, but can also be extended to the intelligent detection and counting of other opportunistic fungal infections, which is conducive to large-scale promotion and application in clinical testing equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 A simplified flowchart of the YOLOv8 polymorphic Cryptococcus detection and counting method based on hypergraph computing provided by the present invention;
[0032] Figure 2 A structural diagram of the YOLOv8 model based on hypergraph computing provided by the present invention;
[0033] Figure 3 A structural diagram of a neck network based on hypergraph computing provided by the present invention. DETAILED DESCRIPTION
[0034] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.
[0035] First, some technical terms in the present invention are explained.
[0036] Density level refers to the degree of semantic abstraction and spatial resolution of the density maps generated by the network at different depths or branches. Deeper layers or low-resolution branches have a coarser density level, used to capture global object distribution; shallower layers or high-resolution branches have a finer density level, used to capture local details. Density levels can be categorized by the downsampling factor: a downsampling ratio between 1 / 16 and 1 / 32 is considered low density, while a downsampling ratio between 1 / 2 and 1 / 8 is considered high density.
[0037] Coarse-grained density maps refer to density maps generated at low-resolution (high downsampling rate) levels, with downsampling rates ranging from 1 / 16 to 1 / 32.
[0038] The coarse-grained cryptococcal density information refers to the cryptococcal density information in the coarse-grained density map.
[0039] The resolution level refers to the resolution grade of the density estimation map. The present invention divides the density estimation map into two levels: the first resolution level and the second resolution level. The downsampling rate in the range of 1 / 2 to 1 / 8 is the first resolution level, and the downsampling rate in the range of 1 / 16 to 1 / 32 is the second resolution level.
[0040] See below. Figures 1 to 3 The present invention provides a YOLOv8 polymorphic cryptococcal detection and counting method based on hypergraph computing, which uses a YOLOv8 model based on hypergraph computing to detect and count polymorphic cryptococci, which is suitable for cryptococcal image detection and counting problems.
[0041] Based on the significant morphological differences of Cryptococcus and the limited image data, in order to improve the model's ability to recognize different clinical morphologies of Cryptococcus, before model training, permanent amplification data were generated through offline data augmentation, and negative sample datasets were added to form the total dataset. During the model training process, random amplification data were generated through online data augmentation.
[0042] Given the limited amount of cryptococcal image data, offline data augmentation is used to enrich each cryptococcal image in the original cryptococcal dataset. This offline data augmentation method includes random cropping, rotation, and flipping. Specifically, the following operations are performed: Each cryptococcal image in the original cryptococcal dataset is subjected to random cropping to generate a first-class new cryptococcal image. The random cropping size can be 640×640, and the present invention does not impose any restrictions on the specific size of the random cropping.
[0043] A second type of new cryptococcal image is generated by rotating each cryptococcal image in the original cryptococcal dataset using offline data enhancement. The rotation angle can be 90 degrees. The present invention does not limit the rotation angle or rotation direction.
[0044] A third type of new Cryptococcus images is generated by using offline data augmentation with horizontal flipping for each Cryptococcus image in the original Cryptococcus dataset.
[0045] The generated three types of new Cryptococcus images were merged with the original Cryptococcus dataset to expand the Cryptococcus dataset.
[0046] The present invention divides the merged data set into a training set, a validation set, and a test set in a ratio of 8:1:1. The present invention does not limit the division ratio. For example, the division ratio can also be set to 7:2:1.
[0047] Given the significant morphological differences of Cryptococcus, after each Cryptococcus image is loaded into the YOLOv8 model to be trained, online data augmentation methods such as perspective transformation and vertical flipping are added to the model's built-in data augmentation methods to enrich the Cryptococcus image data.
[0048] After obtaining the Cryptococcus dataset, a YOLOv8 model was constructed, consisting of an input module, a CSPDarknet backbone network, a neck network based on hypergraph computing, a head network, an adaptive density fusion network, and an output module. The neck network based on hypergraph computing consists of four units: a semantic collection unit, a hypergraph computing unit, a semantic scattering unit, and an attention-enhanced bottom-up unit.
[0049] The present invention provides a YOLOv8 polymorphic cryptococcal detection and counting method based on hypergraph computing, comprising steps S1 to S7.
[0050] S1. In the input module, obtain the cryptococcal image to be detected and counted.
[0051] S2. In the CSPDarknet backbone network, feature extraction is performed on the cryptococcal image to obtain a multi-scale feature map for the cryptococcal image.
[0052] Among them, the CSPDarknet backbone network consists of 5 blocks stacked in a bottom-up manner, among which the first block consists of a CBS unit; the second, third, and fourth blocks have the same structure, each consisting of a CBS unit and a C2f unit; the fifth block consists of a CBS unit, a C2f unit, and an SPPF unit; the CBS unit consists of a convolutional layer with a kernel size of 3, a stride of 2, and a padding of 1, a batch normalization layer, and a SiLU activation function layer; the C2f unit consists of two 1×1 convolutional blocks, a Concat splicing layer, and multiple Bottleneck units; the SPPF unit consists of two 1×1 convolutional blocks, three maximum pooling layers with a kernel size of 5, a stride of 1, and a padding of 2, and a Concat splicing layer.
[0053] The step of extracting features from the cryptococcal image to obtain a multi-scale feature map for the cryptococcal image specifically includes sub-steps S21 to S24.
[0054] S21. In the CBS unit, perform a 2x downsampling operation on the Cryptococcus image to obtain a feature map.
[0055] S22. In the C2f unit, the output feature maps and input feature maps of each Bottleneck unit are spliced together to aggregate multi-scale features, and convolution operations are performed to compress the feature maps, thereby reducing the amount of computation while maintaining or enhancing the expressive power of the model.
[0056] S23. In the SPPF unit, a multi-scale pooling operation is performed on the feature map to capture information of receptive fields of different sizes, thereby reducing computational complexity while maintaining performance. The multi-scale pooling operation includes sub-steps S231 to S234.
[0057] S231, reduce the channel dimension of the input feature map to half through a 1×1 convolution block.
[0058] S232, the feature map after the channel dimension is reduced to half is sequentially subjected to three maximum pooling layers for two-dimensional operation to obtain three feature maps with receptive fields of 5×5, 9×9, and 13×13 respectively;
[0059] S233, concatenating the three feature maps obtained by the maximum pooling two-dimensional operation and the feature map with the channel dimension reduced to half along the channel dimension through a Concat splicing layer to obtain feature information with different sizes of receptive fields;
[0060] S234, compressing the channel dimension of the feature map after the channel dimension is spliced to a specified number through a 1×1 convolution block;
[0061] S24. The cryptococcal image passes through five blocks in sequence to obtain five feature maps of different scales.
[0062] S3. In the neck network based on hypergraph computing, Cryptococcus with different morphological characteristics is adaptively detected, and the perception ability of polymorphic Cryptococcus is enhanced by introducing the attention mechanism.
[0063] The steps of adaptively detecting cryptococci with different morphological characteristics and enhancing the perception of polymorphic cryptococci by introducing an attention mechanism include:
[0064] The semantic collection unit performs information fusion on the multi-scale feature maps of the CSPDarknet backbone network in the semantic space;
[0065] The hypergraph computing unit uses hypergraph computing to model and learn the potential high-order correlations between visual features to generate features with high-order perceptual capabilities. This feature integrates high-order structural information and semantic information.
[0066] The semantic scattering unit scatters high-order structural information into three feature maps of different scales;
[0067] The attention-enhanced bottom-up unit transfers the detailed information of the shallow feature map to the deep feature map layer by layer through a bottom-up lateral connection path to make up for the missing detailed information in the deep feature map, and introduces the attention mechanism to extract the attention area to enhance the model's perception of polymorphic Cryptococcus.
[0068] Below, the specific sub-steps of the semantic collection unit, hypergraph calculation unit, semantic scattering unit, and attention-enhanced bottom-up unit will be described respectively.
[0069] S31 . The semantic collection unit executes sub-steps 311 to 314 .
[0070] S311. Perform downsampling operations on the output feature maps of the 1st, 2nd, and 3rd blocks of the CSPDarknet backbone network by 8, 4, and 2 times, respectively, to obtain three feature maps with the same resolution level as the output feature map of the 4th block of the CSPDarknet backbone network.
[0071] S312. Perform a 2x upsampling operation on the output feature map of the 5th block of the CSPDarknet backbone network to obtain a feature map with the same resolution level as the output feature map of the 4th block of the CSPDarknet backbone network.
[0072] S313. The four feature maps obtained by the upsampling and downsampling operations are spliced along the channel with the output feature map of the fourth block of the CSPDarknet backbone network to obtain a feature map that integrates the high-level semantic information of the deep feature map and the detailed information of the shallow feature map.
[0073] S314. Perform channel compression on the fused feature map through a 1×1 convolution block.
[0074] S32. The hypergraph computing unit executes the hypergraph construction sub-step and the hypergraph convolution sub-step.
[0075] S321, Hypergraph construction sub-step: Hypergraph By its vertex set and the overall hyperedge set To define; map the grid-based visual features to the vertex set of the hypergraph , where the eigenvector of each spatial position corresponds to a vertex ; To model domain relationships in the semantic space, a distance threshold is used Construct a neighborhood sphere for each vertex , and take the vertex set in the neighborhood ball as a hyperedge ; Overall hyperedge set The corresponding formula is:
[0076] ,
[0077] ,
[0078] In the above formula, Represents a vertex In the vertex set The neighboring vertices in Represents a vertex and its neighboring vertices The Euclidean distance between .
[0079] It should be noted that in the calculation, the hypergraph Usually by its incidence matrix express.
[0080] S322, Hypergraph Convolution Sub-step: To facilitate the transmission of high-order information on the hypergraph structure, a two-stage hypergraph information transfer is performed on vertex features using typical spatial domain hypergraph convolution and additional residual connections to output a feature graph after hypergraph calculation; the first stage of hypergraph information transfer refers to aggregating vertices with similar attributes into hyperedges to generate high-order semantic representations; the second stage of hypergraph information transfer refers to diffusing the semantic information of hyperedges back to vertices and updating vertex features to integrate global context; the corresponding formulas for the two-stage hypergraph information transfer are as follows:
[0081] ,
[0082] ,
[0083] In the above formula, Represents the hyperedge after the first stage of hypergraph information transmission The aggregate feature vector of Represents a hyperedge The vertex set of Represents a vertex The eigenvector of Represents the vertices after the second stage of hypergraph information transfer The eigenvector of Represents a vertex The set of associated hyperedges of ;
[0084] The hypergraph convolution operations of the first stage hypergraph information transfer and the second stage hypergraph information transfer are defined as:
[0085] ,
[0086] In the above formula, represents the hypergraph convolution operation, represents the input feature map, represents the incidence matrix of the hypergraph, represents the inverse matrix of the vertex angle matrix, Represents a hyperedge The inverse matrix of the diagonal matrix, Represents the transposed matrix of the incidence matrix of the hypergraph.
[0087] S33. The semantic scattering unit executes sub-steps S331 to S332.
[0088] S331. Perform two-fold upsampling and two-fold downsampling operations on the output feature map of the hypergraph computing unit.
[0089] S332. The feature maps obtained by upsampling, hypergraph calculation, and downsampling are spliced with the output feature maps of the 3rd, 4th, and 5th blocks of the CSPDarknet backbone network, respectively. Then, the three spliced feature maps of different scales are passed through a C2f unit, a C2f unit, and a 1×1 convolution block respectively to further enhance the features.
[0090] S34. The attention-enhanced bottom-up unit executes sub-steps S341 to S343.
[0091] S341. Through a bottom-up lateral connection path, the detail information of the shallow feature map with a downsampling rate of 1 / 8 in the output feature map of the semantic scattering unit is transferred layer by layer to the deep feature map with a downsampling rate of 1 / 32 in the output feature map of the semantic scattering unit to make up for the missing detail information of the deep feature map.
[0092] S342. A convolutional block attention unit is added before each downsampling to extract the attention area to enhance the model's perception of polymorphic Cryptococcus.
[0093] Among them, the convolution block attention unit is composed of a channel attention sub-unit and a spatial attention sub-unit in series, and its calculation process is as follows:
[0094] ,
[0095] ,
[0096] ,
[0097] ,
[0098] In the above formula, represents the channel attention weight matrix, represents the input feature map, represents the average pooling operation, represents the maximum pooling operation, represents a convolution operation with a kernel size of 1 for learning inter-channel dependencies, represents a convolution operation with a kernel size of 7 for capturing wide-area spatial context information, represents the sigmoid activation function, Represents the element-by-element multiplication of the input feature map and the attention weight matrix, Indicates the concatenation operation of two feature maps along the channel, represents the output feature map of the channel attention subunit, represents the spatial attention weight matrix, Represents the output feature map of the spatial attention sub-unit.
[0099] S343. After attention enhancement, three feature maps of different scales are obtained, which are used to detect small target cryptococci, medium target cryptococci, and large target cryptococci, respectively.
[0100] S4. In the head network, two 3×3 convolution blocks and one 1×1 convolution layer are used in each head to generate the density estimation map.
[0101] The head network consists of three heads, each of which consists of two 3×3 convolutional blocks and one 1×1 convolutional layer.
[0102] The step of generating a density estimation map using two 3×3 convolution blocks and one 1×1 convolution layer in a single head specifically includes sub-steps S41 to S43.
[0103] S41. The three feature maps of different scales output by the neck network based on hypergraph calculation are used as input feature maps. The input feature map is compressed to 128 channels through the first convolution block with a kernel size of 3, a stride of 1, and a padding of 1. Then, a 128-dimensional intermediate feature map is obtained through the ReLU activation function. .
[0104] S42, 128-dimensional intermediate feature map The number of channels is further reduced to 64 through the second convolution block with a kernel size of 3, a stride of 1, and a padding of 1. Then, after the ReLU activation function, a 64-dimensional refined feature map is obtained. .
[0105] S43, 64-dimensional refined feature map The number of channels is compressed to 1 through a 1×1 convolution layer, and three density estimation maps with different resolution levels are obtained. 、 、 .
[0106] S5. In the adaptive density fusion network, the coarse-grained density map is gradually refined in an adaptive fusion manner at the density level, and the coarse-grained cryptococcal density information is introduced into the density estimation map at the first resolution level.
[0107] The step of gradually refining the coarse-grained density map in an adaptive fusion manner at the density level and introducing the coarse-grained cryptococcal density information into the density estimation map at the first resolution level specifically includes sub-steps S51 and S52.
[0108] Sub-step S51: In the adaptive density fusion network, a gradient optimization method is used to obtain weight coefficients of different density estimation graphs in the density fusion process.
[0109] Sub-step S52: gradually aggregate the coarse-grained cryptococcal density information into a density estimation map at the first resolution level by adaptive fusion. The specific expression is as follows:
[0110] ,
[0111] In the above formula, Represents the density estimation maps with different resolutions generated by the head network, Representation and density estimation plots The corresponding level, 、 、 Shows three density estimation maps with different resolutions generated by the head network. It represents the Cryptococcus density estimation map after adaptive fusion, represents bilinear interpolation, Represents the weight coefficients of different density estimation maps during density fusion.
[0112] S6. A multi-level distributed supervision strategy is adopted to optimize the density maps of each scale generated by the adaptive density fusion network through a combined loss function.
[0113] The combined loss includes counting loss, optimal transmission loss and total variation loss, and its calculation steps include sub-steps S61 to S64.
[0114] S61. Calculate the count loss: The count loss is calculated by using the L1 loss between the true count and the predicted count to ensure that the predicted density estimation map is consistent with the true count in terms of global count. The specific expression is as follows:
[0115] ,
[0116] In the above formula, Represented in the density estimation plot The counting loss on Representation and density estimation plots The corresponding level, Representation and density estimation plots True density maps of the same size, Represents a density estimation plot In pixel position The pixel value corresponding to the location, Represents the true density map Middle pixel position The pixel value corresponding to the location, Indicates pixel position The horizontal coordinate at Indicates pixel position The vertical coordinate of .
[0117] S62. Calculate the optimal transmission loss: The optimal transmission loss is calculated by finding the distribution difference between the predicted cryptococcal distribution and the actual cryptococcal distribution that minimizes the transmission cost using the Sinkhorn algorithm to solve the problem of local distribution mismatch in the density estimation map. The calculation steps include sub-steps S621 to S622.
[0118] S621. Treating Density Maps as Probability Distributions: Normalized Density Estimation Maps and the true density map Make the sum equal to 1.
[0119] S622, calculate the transmission cost matrix: define the square of the Euclidean distance between pixels as the transmission cost matrix , the formula is:
[0120] ,
[0121] In the above formula, Represents the true density map Middle pixel position The pixel value corresponding to the location, Indicates pixel position The horizontal coordinate at Indicates pixel position The vertical coordinate of .
[0122] S623. Solve the optimal transmission plan: Under the transmission constraints, the optimal transmission plan is obtained by minimizing the total transmission cost. .
[0123] Among them, the formula corresponding to minimizing the total transmission cost is:
[0124] ,
[0125] In the above formula, Represented in the density estimation plot Optimal transmission loss on ; Represents the density estimation graph Pixel position To the true density map Pixel position transmission plan.
[0126] The formula for the transmission constraint is:
[0127] ,
[0128] ,
[0129] In the above formula, Indicates the pixel position The total amount of transmission sent is equal to the predicted density value at that location, Indicates the pixel position reached The total transmission amount is equal to the true density value at that location;
[0130] S624, use the Sinkhorn algorithm to obtain an efficient approximate solution.
[0131] S63. Calculate the total variation loss: The total variation loss is calculated by using the L1 norm of the gradient of the density estimation map in the horizontal and vertical directions to enhance the local smoothness of the predicted density estimation map and reduce noise. The specific expression is as follows:
[0132] ,
[0133] In the above formula, Represented in the density estimation plot The total variational loss on .
[0134] S64. Calculate the combined loss: Based on the counting loss, optimal transmission loss, and total variational loss, calculate the combined loss. The expression of the combined loss is as follows:
[0135] ,
[0136] In the above formula, represents the combined loss, and denotes two tunable hyperparameters.
[0137] S7. In the output module, the Cryptococcus detection and counting results are output.
[0138] The step of outputting the cryptococcal detection and counting results specifically includes sub-steps S71 to S73.
[0139] S71. Sum the density estimation map obtained by the improved YOLOv8 model to obtain the cryptococcal count result in the cryptococcal image.
[0140] S72. If the cryptococcal count result is less than 1, it is considered that no cryptococci are detected in the cryptococcal image, and the cryptococcal count result is output as 0.
[0141] S73. If the cryptococcal count result is greater than or equal to 1, it is considered that cryptococci are detected in the cryptococcal image, and the cryptococcal count result is output.
[0142] The present invention also provides a YOLOv8 polymorphic cryptococcal detection and counting system based on hypergraph computing, which is used to implement the aforementioned YOLOv8 polymorphic cryptococcal detection and counting method based on hypergraph computing. The system has a YOLOv8 model based on hypergraph computing, and the YOLOv8 model includes:
[0143] An input module is used to obtain an image of cryptococci to be detected and counted;
[0144] CSPDarknet backbone network, used to extract features from cryptococcal images to obtain multi-scale feature maps for cryptococcal images;
[0145] The neck network based on hypergraph computing is composed of four units: semantic collection unit, hypergraph computing unit, semantic scattering unit, and attention-enhanced bottom-up unit. It is used to adaptively detect cryptococci with different morphological characteristics and enhances the perception of polymorphic cryptococci by introducing an attention mechanism.
[0146] The head network is used to generate a coarse-grained density map and a density estimation map of the first resolution level using two 3×3 convolution blocks and one 1×1 convolution layer in each head;
[0147] An adaptive density fusion network is used to gradually refine the coarse-grained density map in an adaptive fusion manner at the density level, and introduce the coarse-grained cryptococcal density information into the density estimation map at the first resolution level;
[0148] A combined loss function based on distributed supervision is used to optimize the density maps of each scale generated by the adaptive density fusion network using a multi-level distributed supervision strategy;
[0149] Output module, used to output cryptococcal detection and counting results.
[0150] Through the detailed description of the YOLOv8 polymorphic cryptococcal detection and counting method based on hypergraph computing, it can be seen that the YOLOv8 polymorphic cryptococcal detection and counting system based on hypergraph computing of The specific structure, working process, etc. will not be elaborated on here.
[0151] In this embodiment, a private dataset of Cryptococcus is used to train and generate a YOLOv8 model based on hypergraph computing. The private dataset contains images of Cryptococcus and other similar fungi, with a corresponding ratio of 7:3.
[0152] To further demonstrate the superiority of the YOLOv8 model based on hypergraph computing, mean absolute error (MAE) and mean square error (MSE) are used as evaluation metrics. The MAE metric reflects counting precision and accuracy, while the MSE metric measures generalization performance and robustness. The evaluation results are shown in Table 1. As can be seen from the results in Table 1, the YOLOv8 model based on hypergraph computing in this embodiment has lower mean absolute error (MAE) and mean square error (MSE) than the traditional YOLOv8s model. This demonstrates the counting accuracy and generalization performance of the YOLOv8 model based on hypergraph computing in this embodiment, and it can be considered that it performs well in the problem of detecting and counting cryptococci images.
[0153] Table 1 Image recognition results of different models
[0154]
[0155] The present invention provides a YOLOv8 polymorphic cryptococcal detection and counting method based on hypergraph computing. In view of the characteristics of cryptococcal image data being relatively small, morphological differences being significant, and the presence of a large number of dense small targets, a data enhancement method based on random cropping, rotation, flipping, and perspective transformation is designed. This effectively expands the cryptococcal dataset and effectively integrates hypergraph computing and attention mechanism into the traditional YOLOv8 model. This realizes adaptive detection of cryptococci with different morphological characteristics, enhances the model's perception of polymorphic cryptococci, and realizes automatic counting of cryptococcal images by counting density estimation maps, providing a systematic solution for the timely and accurate identification and counting of cryptococci.
[0156] It will be easily understood by those skilled in the art that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A YOLOv8 polymorphic cryptococcal detection and counting method based on hypergraph computing, characterized in that: A YOLOv8 model based on hypergraph computing is used to detect and count polymorphic cryptococci. Based on the characteristics of polymorphic cryptococcal images, offline data augmentation is used to generate permanent amplified data before model training, and a negative sample dataset is added to form a total dataset. During model training, online data augmentation is used to generate random amplified data. Based on the dataset after online data augmentation, a YOLOv8 model is constructed, which includes an input module, a CSPDarknet backbone network, a neck network based on hypergraph computing, a head network, an adaptive density fusion network, and an output module. The neck network based on hypergraph computing consists of four units: a semantic collection unit, a hypergraph computing unit, a semantic scattering unit, and an attention-enhanced bottom-up unit. The YOLOv8 polymorphic cryptococcal detection and counting method based on hypergraph computing includes: In the input module, an image of cryptococci to be detected and counted is obtained; In the CSPDarknet backbone network, feature extraction is performed on the cryptococcal image to obtain a multi-scale feature map for the cryptococcal image; In the neck network based on hypergraph computing, Cryptococcus with different morphological characteristics is adaptively detected, and the perception ability of polymorphic Cryptococcus is enhanced by introducing an attention mechanism; In the head network, two 3×3 convolution blocks and one 1×1 convolution layer are used in each head to generate a coarse-grained density map and a density estimation map at the first resolution level; In the adaptive density fusion network, the coarse-grained density map is gradually refined in an adaptive fusion manner at the density level, and the coarse-grained cryptococcal density information is introduced into the density estimation map at the first resolution level; A multi-level distributed supervision strategy is adopted to optimize the density maps of each scale generated by the adaptive density fusion network through a combined loss function. In the output module, the Cryptococcus detection and counting results are output.
2. The method according to claim 1, wherein The steps to augment the Cryptococcus dataset using both offline and online data augmentation methods include: Considering the small amount of Cryptococcus image data, we enriched the Cryptococcus image data by using offline data augmentation methods such as random cropping, rotation, and horizontal flipping for each Cryptococcus image in the original Cryptococcus dataset. Given the significant morphological differences of Cryptococcus, after each Cryptococcus image is loaded into the YOLOv8 model to be trained, online data augmentation methods such as perspective transformation and vertical flipping are added to the model's built-in data augmentation methods to enrich the Cryptococcus image data.
3. The method according to claim 1, wherein The CSPDarknet backbone network is composed of 5 blocks stacked in a bottom-up manner, wherein the first block consists of a CBS unit; the second, third, and fourth blocks have the same structure, each consisting of a CBS unit and a C2f unit; the fifth block consists of a CBS unit, a C2f unit, and an SPPF unit; the CBS unit consists of a convolutional layer with a kernel size of 3, a stride of 2, and a padding of 1, a batch normalization layer, and a SiLU activation function layer; the C2f unit consists of two 1×1 convolutional blocks, a Concat splicing layer, and multiple Bottleneck units; the SPPF unit consists of two 1×1 convolutional blocks, three maximum pooling layers with a kernel size of 5, a stride of 1, and a padding of 2, and a Concat splicing layer; the step of extracting features from the cryptococcal image to obtain a multi-scale feature map for the cryptococcal image specifically includes: In the CBS unit, a 2-fold downsampling operation is performed on the cryptococcal image to obtain a feature map; In the C2f unit, the output feature maps and input feature maps of each Bottleneck unit are spliced together to aggregate multi-scale features, and a convolution operation is performed to compress the feature map, thereby reducing the amount of calculation while maintaining or enhancing the expressive power of the model; In the SPPF unit, a multi-scale pooling operation is performed on the feature map to capture information of receptive fields of different sizes, reducing computational complexity while maintaining performance. The multi-scale pooling operation is as follows: Reduce the channel dimension of the input feature map to half through a 1×1 convolution block; The feature map after the channel dimension is reduced to half is sequentially passed through three maximum pooling layers for two-dimensional operation to obtain three feature maps with receptive fields of 5×5, 9×9, and 13×13 respectively; The three feature maps obtained by the maximum pooling two-dimensional operation and the feature map with the channel dimension reduced to half are spliced along the channel dimension through the Concat splicing layer to obtain feature information with different receptive fields; The channel dimension of the feature map after channel dimension splicing is compressed to the specified number through a 1×1 convolution block; The cryptococcal image passes through five blocks in sequence to obtain five feature maps of different scales.
4. The method according to claim 1, wherein The steps of adaptively detecting cryptococci with different morphological characteristics and enhancing the perception of polymorphic cryptococci by introducing an attention mechanism specifically include: The semantic collection unit performs information fusion on the multi-scale feature maps of the CSPDarknet backbone network in the semantic space; The hypergraph computing unit uses hypergraph computing methods to model and learn the potential high-order correlations between visual features to generate features with high-order perceptual capabilities. This feature integrates high-order structural information and semantic information. The semantic scattering unit scatters high-order structural information into three feature maps of different scales; The attention-enhanced bottom-up unit transfers the detailed information of the shallow feature map to the deep feature map layer by layer through a bottom-up lateral connection path to make up for the missing detailed information in the deep feature map, and introduces the attention mechanism to extract the attention area to enhance the model's perception of polymorphic Cryptococcus.
5. The method according to claim 4, wherein The semantic collection unit, the hypergraph calculation unit, the semantic scattering unit, and the attention-enhanced bottom-up unit specifically perform the following sub-steps: The semantic collection unit performs the following sub-steps: In the first step, the output feature maps of the 1st, 2nd, and 3rd blocks of the CSPDarknet backbone network are downsampled by 8, 4, and 2 times respectively to obtain three feature maps with the same resolution level as the output feature map of the 4th block of the CSPDarknet backbone network; The second step is to perform a 2x upsampling operation on the output feature map of the 5th block of the CSPDarknet backbone network to obtain a feature map with the same resolution level as the output feature map of the 4th block of the CSPDarknet backbone network; In the third step, the four feature maps obtained by upsampling and downsampling operations are spliced along the channel with the output feature map of the fourth block of the CSPDarknet backbone network to obtain a feature map that integrates the high-level semantic information of the deep feature map and the detailed information of the shallow feature map; The fourth step is to perform channel compression on the fused feature map through a 1×1 convolution block; The hypergraph computing unit performs the following sub-steps: Hypergraph construction sub-steps: Hypergraph By its vertex set and the overall hyperedge set To define; map the grid-based visual features to the vertex set of the hypergraph , where the eigenvector of each spatial position corresponds to a vertex ; To model domain relationships in the semantic space, a distance threshold is used Construct a neighborhood sphere for each vertex , and take the vertex set in the neighborhood ball as a hyperedge ; Overall hyperedge set The corresponding formula is: , , In the above formula, Represents a vertex In the vertex set The neighboring vertices in Represents a vertex and its neighboring vertices The Euclidean distance between Hypergraph convolution sub-step: To facilitate the transfer of high-order information on the hypergraph structure, a two-stage hypergraph information transfer is performed on vertex features using typical spatial domain hypergraph convolution and additional residual connections to output the feature map after hypergraph calculation; The first stage of hypergraph information transfer refers to aggregating vertices with similar attributes into hyperedges to generate high-order semantic representations. The second stage of hypergraph information transfer refers to diffusing the semantic information of hyperedges back to vertices and updating vertex features to integrate global context. The corresponding formulas for the two-stage hypergraph information transfer are as follows: , , In the above formula, Represents the hyperedge after the first stage of hypergraph information transmission The aggregate feature vector of Represents a hyperedge The vertex set of Represents a vertex The eigenvector of Represents the vertices after the second stage of hypergraph information transfer The eigenvector of Represents a vertex The set of associated hyperedges of ; The hypergraph convolution operations of the first stage hypergraph information transfer and the second stage hypergraph information transfer are defined as: , In the above formula, represents the hypergraph convolution operation, represents the input feature map, represents the incidence matrix of the hypergraph, represents the inverse matrix of the vertex angle matrix, Represents a hyperedge The inverse matrix of the diagonal matrix, The transposed matrix representing the incidence matrix of the hypergraph; The semantic scattering unit performs the following sub-steps: The first step is to perform two-fold upsampling and two-fold downsampling operations on the output feature map of the hypergraph computing unit; In the second step, the feature maps obtained by upsampling, hypergraph calculation, and downsampling are spliced with the output feature maps of the 3rd, 4th, and 5th blocks of the CSPDarknet backbone network respectively. Then, the three spliced feature maps of different scales are passed through a C2f unit, a C2f unit, and a 1×1 convolution block respectively to further enhance the features; The attention-enhanced bottom-up unit performs the following substeps: In the first step, through a bottom-up lateral connection path, the detailed information of the shallow feature map with a downsampling rate of 1 / 8 in the output feature map of the semantic scattering unit is transferred layer by layer to the deep feature map with a downsampling rate of 1 / 32 in the output feature map of the semantic scattering unit to make up for the missing detailed information in the deep feature map; In the second step, a convolutional block attention unit is added before each downsampling to extract the attention area to enhance the model's perception of polymorphic Cryptococcus; In the third step, three feature maps of different scales are obtained after attention enhancement, which are used to detect small target cryptococci, medium target cryptococci, and large target cryptococci respectively. If the width and height of the annotation box of cryptococci are and , the width and height of the image are and , then the small target Cryptococcus refers to the one that meets Cryptococcus, the target Cryptococcus refers to those that meet Cryptococcus, the target Cryptococcus refers to the Cryptococcus that meets Cryptococcus; The convolutional block attention unit is composed of a channel attention sub-unit and a spatial attention sub-unit in series, and its calculation process is as follows: , , , , In the above formula, represents the channel attention weight matrix, represents the input feature map, represents the average pooling operation, represents the maximum pooling operation, represents a convolution operation with a kernel size of 1 for learning inter-channel dependencies, represents a convolution operation with a kernel size of 7 for capturing wide-area spatial context information, represents the sigmoid activation function, Represents the element-by-element multiplication of the input feature map and the attention weight matrix, Indicates the concatenation operation of two feature maps along the channel, represents the output feature map of the channel attention subunit, represents the spatial attention weight matrix, Represents the output feature map of the spatial attention sub-unit.
6. The method according to claim 1, wherein The head network includes three heads, each of which is composed of two 3×3 convolution blocks and one 1×1 convolution layer. The step of generating a density estimation map using the two 3×3 convolution blocks and one 1×1 convolution layer in a single head specifically includes: The three feature maps of different scales output by the neck network based on hypergraph calculation are used as input feature maps. The input feature map is compressed to 128 channels through the first convolution block with a kernel size of 3, a stride of 1, and a padding of 1. Then, a 128-dimensional intermediate feature map is obtained by the ReLU activation function. ; 128-dimensional intermediate feature map The number of channels is further reduced to 64 through the second convolution block with a kernel size of 3, a stride of 1, and a padding of 1. Then, after the ReLU activation function, a 64-dimensional refined feature map is obtained. ; 64-dimensional refined feature map The number of channels is compressed to 1 through a 1×1 convolution layer, and three density estimation maps with different resolution levels are obtained. 、 、 .
7. The method according to claim 1, wherein The step of gradually refining the coarse-grained density map in an adaptive fusion manner at the density level and introducing the coarse-grained cryptococcal density information into the density estimation map at the first resolution level specifically includes the following sub-steps: In the adaptive density fusion network, the gradient optimization method is used to obtain the weight coefficients of different density estimation graphs in the density fusion process; The coarse-grained cryptococcal density information is gradually aggregated into the density estimation map at the first resolution level by adaptive fusion. The specific expression is as follows: , In the above formula, Represents the density estimation maps with different resolutions generated by the head network, Representation and density estimation plots The corresponding level, 、 、 Represents three density estimation maps with different resolutions generated by the head network. It represents the Cryptococcus density estimation map after adaptive fusion, represents bilinear interpolation, Represents the weight coefficients of different density estimation maps during density fusion.
8. The method according to claim 1, wherein The combined loss includes counting loss, optimal transmission loss and total variation loss, and its calculation steps include: Calculate the count loss: The count loss is calculated by using the L1 loss between the true count and the predicted count to ensure that the predicted density estimation map is consistent with the true count in terms of global count. The specific expression is as follows: , In the above formula, Represented in the density estimation plot The counting loss on Representation and density estimation plots The corresponding level, Representation and density estimation plots True density maps of the same size, Represents a density estimation plot In pixel position The pixel value corresponding to the location, Represents the true density map Middle pixel position The pixel value corresponding to the location, Indicates pixel position The horizontal coordinate at Indicates pixel position The vertical coordinate of Calculating the optimal transmission loss: The optimal transmission loss is calculated by finding the distribution difference between the predicted cryptococcal distribution and the actual cryptococcal distribution that minimizes the transmission cost using the Sinkhorn algorithm to solve the problem of local distribution mismatch in the density estimation map. The calculation steps are as follows: The first step is to treat the density map as a probability distribution: Normalized density estimate map and the true density map Make the sum equal to 1; The second step is to calculate the transmission cost matrix: define the square of the Euclidean distance between pixels as the transmission cost matrix , the formula is: , In the above formula, Represents the true density map Middle pixel position The pixel value corresponding to the location, Indicates pixel position The horizontal coordinate at Indicates pixel position The vertical coordinate of The third step is to solve the optimal transmission plan: Under the transmission constraints, the optimal transmission plan is obtained by minimizing the total transmission cost. ; Among them, the formula corresponding to minimizing the total transmission cost is: , In the above formula, Represented in the density estimation plot Optimal transmission loss on ; Represents the density estimation graph Pixel position To the true density map Pixel position transmission plan; The formula for the transmission constraint is: , , In the above formula, Indicates the pixel position The total amount of transmission sent is equal to the predicted density value at that location, Indicates the pixel position reached The total transmission amount is equal to the true density value at that location; The fourth step is to use the Sinkhorn algorithm to obtain an efficient approximate solution; Calculate the total variation loss: The total variation loss is calculated by using the L1 norm of the gradient of the density estimation map in the horizontal and vertical directions to enhance the local smoothness of the predicted density estimation map and reduce noise. The specific expression is as follows: , In the above formula, Represented in the density estimation plot The total variational loss on ; Calculating the combined loss: Based on the counting loss, the optimal transmission loss, and the total variational loss, the combined loss is calculated. The expression of the combined loss is as follows: , In the above formula, represents the combined loss, and denotes two tunable hyperparameters.
9. The method according to claim 1, wherein The step of outputting cryptococcal detection and counting results specifically includes: The density estimation map obtained by the improved YOLOv8 model is summed to obtain the cryptococcal count result in the cryptococcal image; If the cryptococcal count result is less than 1, it is considered that no cryptococci are detected in the cryptococcal image, and the output cryptococcal count result is 0; If the cryptococcal count result is greater than or equal to 1, it is considered that cryptococci are detected in the cryptococcal image, and the cryptococcal count result is output.
10. A YOLOv8 polymorphic cryptococcal detection and counting system based on hypergraph computing, used to implement the YOLOv8 polymorphic cryptococcal detection and counting method based on hypergraph computing according to any one of claims 1 to 9, characterized in that: The system has a YOLOv8 model based on hypergraph computing, and the YOLOv8 model includes: An input module is used to obtain an image of cryptococci to be detected and counted; CSPDarknet backbone network, used to extract features from cryptococcal images to obtain multi-scale feature maps for cryptococcal images; The neck network based on hypergraph computing is composed of four units: semantic collection unit, hypergraph computing unit, semantic scattering unit, and attention-enhanced bottom-up unit. It is used to adaptively detect cryptococci with different morphological characteristics and enhances the perception of polymorphic cryptococci by introducing an attention mechanism. The head network is used to generate a coarse-grained density map and a density estimation map of the first resolution level using two 3×3 convolution blocks and one 1×1 convolution layer in each head; An adaptive density fusion network is used to gradually refine the coarse-grained density map in an adaptive fusion manner at the density level, and introduce the coarse-grained cryptococcal density information into the density estimation map at the first resolution level; A combined loss function based on distributed supervision is used to optimize the density maps of each scale generated by the adaptive density fusion network using a multi-level distributed supervision strategy; Output module, used to output cryptococcal detection and counting results.
Citation Information
Patent Citations
Cell counting method and device, terminal equipment and medium
CN115861275A
Traffic flow prediction method based on space-time causal attention network
CN119107794A