Precise detection and localization method of cerebral microbleeds combined with rpn and anatomical information
By combining the detection and positioning methods of RPN and anatomical information, and using spatial overlap and structural feature similarity to adjust the confidence of candidate regions, the problems of low efficiency and high false positive rate in cerebral microbleed detection in existing technologies are solved, and efficient and accurate cerebral microbleed detection and positioning are achieved.
Patent Information
- Application Number
- CN202510986803.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-17
AI Technical Summary
Existing methods for detecting cerebral microbleeds are inefficient and highly subjective, making it difficult to accurately distinguish CMBs from their analogs. Deep learning methods have high computational costs and high false positive rates, and cannot effectively solve this problem.
A cerebral microbleed detection and localization method that combines RPN and anatomical information is proposed. Through the candidate region detection module and anatomical localization module, the confidence of the candidate region is adjusted by using spatial overlap and structural feature similarity. The model is trained through a multi-stage optimization strategy to reduce the false positive rate and improve detection accuracy.
It achieves efficient automated detection of cerebral microbleeds, reduces false positive rates, improves detection accuracy and clinical value, ensures that candidate regions match anatomical structures, and reduces computational complexity.
Smart Images

Figure CN120563487B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiment of the present application relates to the technical field of medical image processing, and particularly relates to a cerebral microbleed precise detection and positioning method combining RPN and anatomical information. BACKGROUND
[0002] Cerebral microbleeds (CMBs) are small blood products deposited in brain tissue, which are usually associated with various cerebrovascular diseases such as cognitive decline, cerebral hemorrhage and cerebral infarction. The detection of CMBs is of great significance for early diagnosis and treatment of diseases. At present, the detection of CMBs mainly relies on MRI (Magnetic Resonance Imaging) technology, especially the SWI (Susceptibility Weighted Imaging) image generated by gradient echo sequence. SWI image can clearly show small blood vessels and microbleeds in brain tissue, providing important imaging basis for the detection of CMBs.
[0003] However, the current CMBs detection methods all have some problems and limitations. For example, manual detection is low in efficiency and strong in subjectivity, which is easy to lead to inconsistent results; traditional deep learning-based automatic methods are based on handcrafted features, which are difficult to accurately distinguish CMBs and similar objects; the single-stage detector of deep learning has high computational cost and high false positive rate, while the two-stage detector relies on the detection results of the first stage and needs an additional classification step, which increases the complexity and computational cost. Patent application CN115908381A discloses a method and device for positioning a target region in a brain CT image, and CN120108011A discloses a method for identifying microbleed lesions in brain images using a deep learning algorithm, which cannot solve the above problems.
[0004] Therefore, it is urgent to provide a deep learning-based cerebral microbleed precise detection and positioning model, and to effectively use and train it to reduce the false positive rate and improve the detection efficiency. SUMMARY
[0005] The embodiment of the present application provides a cerebral microbleed precise detection and positioning method combining RPN and anatomical information to solve the above technical problems.
[0006] In a first aspect, an embodiment of the present application provides a method for precise detection and positioning of cerebral microbleeds combined with RPN and anatomical information, which is implemented based on a deep learning-based model for precise detection and positioning of cerebral microbleeds, and the model comprises a candidate region detection module and an anatomical positioning module, wherein the candidate region detection module is configured to detect candidate regions of cerebral microbleeds from brain images of a patient, and the anatomical positioning module is configured to segment brain anatomical regions from the brain images of the patient; and the method comprises:
[0007] obtaining at least one candidate region of cerebral microbleeds from the candidate region detection module and at least one brain anatomical region from the anatomical positioning module based on the brain images of the same patient;
[0008] calculating spatial overlap and structural feature similarity between each candidate region and each intersecting anatomical region;
[0009] adjusting the confidence of a candidate region that does not match according to the spatial overlap and the structural feature similarity and corresponding threshold.
[0010] In a second aspect, an embodiment of the present application provides an electronic device, which comprises:
[0011] one or more processors;
[0012] a memory configured to store one or more programs,
[0013] when the one or more programs are executed by the one or more processors, the one or more processors implement the method for precise detection and positioning of cerebral microbleeds combined with RPN and anatomical information according to any embodiment.
[0014] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, and the program is executed by a processor to implement the method for precise detection and positioning of cerebral microbleeds combined with RPN and anatomical information according to any embodiment.
[0015] In summary, an embodiment of the present application adopts a deep learning-based model for precise detection and positioning of cerebral microbleeds, realizes efficient automatic detection of cerebral microbleeds, reduces the time and labor intensity of manual detection, and reduces the false positive rate in the detection process, thereby improving the accuracy of detection; at the same time, the automatic anatomical positioning of CMBs is realized, and the clinical value of detection is further improved. In particular, the embodiment determines whether the candidate region and the anatomical region match through spatial consistency and feature consistency, and timely eliminates the candidate region that does not match the anatomical structure; and a remedial measure of expanding the region is set for the misidentified candidate region, so as to determine the reasonable attribution of the candidate region as much as possible, and fully realize the anatomical information to realize the error detection and correction of the candidate region.
[0016] Further, to ensure the effective integration of each module during the training process, the embodiment gradually unlocks the module, dynamically adjusts the loss term, and uses a multi-stage optimization strategy of parameter freezing and unfreezing mechanism, so that the focus of training gradually shifts from lesion detection to fine segmentation and anatomical matching, ensuring that the optimization order of each module is reasonable, and also allowing each module to fully exert its own advantages at different stages, ultimately achieving collaborative optimization of the overall system and improving the stability and accuracy of the overall detection. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the specific embodiments or prior art of the present application, the drawings needed in the specific embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0018] Figure 1 is a structural schematic diagram of a brain microhemorrhage precise detection and positioning model based on deep learning provided by an embodiment of the present application;
[0019] Figure 2 is a structural schematic diagram of a brain microhemorrhage candidate region detection module provided by an embodiment of the present application;
[0020] Figure 3 is a structural schematic diagram of a brain microhemorrhage anatomical positioning module provided by an embodiment of the present application;
[0021] Figure 4 is a flowchart of a brain microhemorrhage precise detection and positioning method combining RPN and anatomical information provided by an embodiment of the present application;
[0022] Figure 5 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of the present application.
[0024] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first", "second", "third" are only for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0025] In the description of the present application, it should be noted that unless otherwise explicitly specified and limited, the terms "mounting", "connection", "connection" should be understood broadly, for example, it can be fixed connection, or detachable connection, or integral connection; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, or the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0026] The embodiment provides a brain microbleed precise detection and positioning method combined with RPN and anatomical information, in order to illustrate the method, the deep learning-based brain microbleed precise detection and positioning model used by the method is introduced first. As shown in the figure, Figure 1 The model includes a candidate region detection module and an anatomical positioning module, wherein the candidate region detection module is used to detect the candidate region of brain microbleed according to the brain image of the patient, and the anatomical positioning module is used to segment the anatomical region of the brain according to the brain image of the patient. The data processing process of the two basic modules is described in detail below.
[0027] In a specific embodiment, before starting the data processing process in the model, data preprocessing is first performed. The original MRI data collected (mainly including SWI and phase images) is processed. The processing procedure includes brain extraction, pixel normalization, slice interpolation and data cropping and the like. Specifically, the brain region is separated from the original image by using an automatic brain extraction tool, eliminating the interference of skull and other non-target tissues; then the min-max normalization method is used to scale the gray value of each voxel to 0 to 1, thereby reducing the influence of different scanners or scanning parameters; in view of the problem of insufficient number of slices caused by scanning thickness and the like, the image is expanded to a predetermined number of layers in the z direction through slice interpolation; finally, according to the preset network input size, the preprocessed image is cropped to obtain a fixed size three-dimensional data. These preprocessing steps not only ensure the uniformity of the data, but also provide sufficient guarantee for the efficient training of the subsequent network.
[0028] The pre-processed image is input into a candidate region detection module as shown in Figure 2 The image is first input into a detection network in the candidate region detection module, which uses an improved 3D U-Net as a basic architecture to extract multi-scale feature information through an encoder and a decoder path. The encoder part gradually extracts high-level semantic information hidden in the image through a series of 3D convolution, batch normalization and activation function operations, while the decoder part restores the spatial resolution through deconvolution or upsampling operations, and fuses the low-level detail features in the encoder with the high-level semantic information of the decoder through a skip connection. It is worth noting that in order to directly realize the positioning and classification of candidate lesions, the present embodiment introduces the design concept of RPN (Region Proposal Network) on the basis of 3D U-Net, and the number of output channels N out is determined by the following formula:
[0029]
[0030] wherein, N cls represents the number of categories (usually 2, i.e. CMB and non-CMB), and N cd represents the dimension of the center coordinates of the bounding box, which is 3 in the present method because it is three-dimensional data. N cd Through this design, the network can obtain the category information and accurate spatial position information of the lesion in one forward propagation process, avoiding the error accumulation and computational redundancy caused by the separation of candidate region generation and classification in the traditional two-stage detection.
[0031] To further improve the detection ability of the network for small lesions, the present embodiment introduces a FFM (Feature Fusion Module) in the detection network. The design concept of this module is to effectively fuse feature maps from different levels and different scales in the network, so as to ensure the preservation of low-level detail features while ensuring high-level semantic information. By inserting convolution layers at multiple places in the decoder to reduce the dimension of the feature channels, and then adding the feature maps from different levels point by point after upsampling to a uniform size, the FFM can significantly enhance the sensitivity of the network to local subtle changes, thereby more accurately distinguishing real lesions from false positive signals.
[0032] Meanwhile, the collected original MRI data enters the anatomical positioning module, which mainly uses the prior anatomical information of the brain structure to perform secondary screening on the candidate regions output by the detection module, so as to exclude those anatomically unreasonable false positive detection results, and ensure that the final output lesion positioning has high reliability and clinical significance.
[0033] Specifically, in combination with Figure 3 , the anatomical positioning module adopts a segmentation network based on 3D U-Net to divide the anatomical regions of the preprocessed original MRI images (including SWI and T1-MPRAGE sequences). In the data preprocessing stage, similar to the detection network, the original image is subjected to brain extraction, normalization, slice interpolation and fixed size cropping operations to ensure the consistency of the input data. In order to solve the problem of loss of position information in the cropping process, the embodiment specially constructs three auxiliary input tensors, which correspond to the absolute coordinate information of the image in the x, y and z directions respectively, and are normalized and input into the segmentation network together with the main image data. In this way, the network can not only learn the texture and structure features in the image, but also obtain spatial position information, so that the division of each anatomical region of the brain is more accurate.
[0034] In terms of network structure, the anatomical positioning module still adopts 3D U-Net as the basic architecture, and introduces residual units to improve the problem of gradient disappearance in the training process of deep network. The encoder part of the network extracts multi-scale high-dimensional features from the original image through the cascade operation of multi-layer 3D convolution, batch normalization and activation function; and the decoder part fuses the low-level detail features and high-level semantic features in the encoder through upsampling and skip connection, and gradually restores the spatial resolution of the image. Finally, the network outputs a multi-channel segmentation map, where each channel represents a probability map of an anatomical region. Generally, the embodiment divides the brain into four categories: cerebral lobe, deep structure, subdural region and CMBs-free region, and uses the softmax function to classify each voxel, thereby obtaining the probability distribution of each voxel belonging to each category. It should be noted that the network output here only retains 4 macro channels; but in the specific anatomical region training / labeling, 27 sub-labels of FreeSurfer-Aseg are used, i.e. 27 anatomical regions are segmented; in the inference stage, it is automatically mapped to 4 categories according to the following corresponding structure, which ensures accuracy and simplifies output:
[0035] 1) Cerebral lobe: including Cerebral Cortex (cerebral cortex) and Subcortical White Matter (subcortical white matter) and other cortical and subcortical white matter regions.
[0036] 2) Deep structures: including Thalamus (Thalamus), Caudate (Caudate), Putamen (Putamen), Pallidum (Pallidum) and other basal nuclei and adjacent nuclei.
[0037] 3) Subtentorial region: including Brain Stem (Brain Stem) and Cerebellum (Cerebellum cortex and white matter of cerebellum).
[0038] 4) CMB-free region: including lateral ventricle, fourth ventricle and other ventricular system, cerebrospinal fluid cavity, main blood vessels / dura mater and skull-air region and other anatomical sites.
[0039] Based on the above two basic modules, Figure 4 is a flowchart of a brain microbleeding precise detection and positioning method combining RPN and anatomical information provided by the embodiment of the present application, which is used to match and analyze the outputs of the above two basic modules, and output the final CMB detection result. The method is executed by an electronic device. As shown in Figure 4 , the method specifically includes:
[0040] S110, obtaining brain microbleeding candidate regions of brain images of the same patient obtained by the candidate region detection module, and brain anatomical regions obtained by the anatomical positioning module.
[0041] Combining Figure 1 , the phase image and the SWI image from the same patient are input into the candidate region modeling module, at least one CMB candidate region in the image can be obtained; at the same time, the SWI image or the T1-MPRAGE image from the patient is input into the anatomical positioning module, and the anatomical region segmentation result (including the above 27 anatomical regions) of the image can be obtained. The elements at the same position in the input image and the output image of the two modules correspond to the same brain position of the patient.
[0042] S120, calculating the spatial overlap degree and the structural feature similarity of each candidate region and each anatomical region.
[0043] The same candidate region may intersect with multiple anatomical regions, so the embodiment calculates the spatial overlap degree and the structural feature similarity of the candidate region and each intersecting anatomical region for the same candidate region.
[0044] Wherein, the spatial overlap degree is used to measure the spatial consistency of the candidate region and each anatomical region intersecting with it. Optionally, the spatial overlap degree (IoU, Intersection over Union) is calculated as follows:
[0045]
[0046] Wherein, Rcand represents a candidate region, R seg represents an anatomical region, represents R cand represents the intersection volume of R seg and R represents R cand represents the union volume of R seg and R
[0047] The structural feature similarity is used to measure the consistency of the structural features (including texture features, gradient features, etc.) of the candidate region and each anatomical region intersected by the candidate region. Optionally, the structural feature similarity is calculated as follows:
[0048]
[0049] wherein, represents the structural feature vector of the i-th candidate region, i represents the structural feature vector of any anatomical region intersected by the candidate region.
[0050] S130, according to the spatial overlap degree and the structural feature similarity, and the corresponding threshold, the confidence of the unmatched candidate region is adjusted.
[0051] The embodiment sets a threshold for the spatial overlap degree and the structural feature similarity, respectively. Optionally, for any candidate region, if the candidate region does not match all the anatomical regions (the spatial overlap degree is less than the overlap threshold, and the structural feature similarity is less than the display threshold), the confidence of the candidate region is adjusted.
[0052] In particular, if a candidate region does not match each anatomical region intersected by the candidate region, it is also possible that the candidate region is assigned to the wrong anatomical region. At this time, if the confidence of the candidate region is high (such as higher than a set threshold), the candidate region can be regionally expanded, and the spatial overlap degree and the structural feature similarity of the expanded candidate region and each anatomical region intersected by the candidate region are recalculated. If the expanded candidate region still does not match each anatomical region intersected by the candidate region, the confidence of the candidate region is adjusted; if the expanded candidate region matches a certain anatomical region intersected by the candidate region, the candidate region plays a role of correction and reasonable attribution.
[0053] The above describes the confidence of the candidate region Figure 1 The method for using the deep learning-based brain microbleed precise detection and positioning model shown adjusts the confidence of each candidate region through spatial overlap degree and feature similarity, and fully realizes the fusion of the RPN and the anatomical information. In combination with the method for using the deep learning-based brain microbleed precise detection and positioning model, the embodiment further proposes a model training method, which further improves the detection accuracy of the model by introducing an anatomical loss function and a phased training.
[0054] The training method adopts a step-by-step unlocking module + dynamic adjustment of a loss term + a parameter freezing and unfreezing mechanism, so that the training focus gradually shifts from lesion detection to fine segmentation and anatomical matching, ensures that the optimization order of each module is reasonable, and finally realizes the collaborative optimization of the overall system. In a specific embodiment, the detailed steps and control of the training focus of each stage are as follows:
[0055] In the initial stage, the candidate region detection module is pre-trained. The training goal of this stage is to mainly learn global features and optimize the generation of candidate regions, and to improve the initial accuracy of the detection module. It mainly focuses on overall feature extraction and rough detection. At this time, the detection module uses a CNN (Convolutional Neural Network) to construct a multi-scale feature representation, and generates preliminary lesion candidate regions through a region candidate network. In order to enhance the recognition ability of small lesions, a multi-scale fusion strategy is adopted to make full use of feature information at different levels.
[0056] In order to achieve the above goal, only the candidate region detection module is enabled in this training stage, and the anatomical positioning module is frozen (i.e., the network parameters are unchanged), and only detection-related loss terms are used L det Optimize the detection network. In order to facilitate differentiation and description, the training of the candidate region detection module in this stage is referred to as one training in this embodiment. Optionally, the detection-related loss terms include:
[0057] L det = L cls + L reg
[0058] Among them, L cls is the classification loss of the candidate region, which is used to ensure the classification accuracy of the lesion and the background; L reg is the candidate box regression loss, which is used to optimize the position accuracy of the candidate region.
[0059] At the same time, a larger learning rate η init (such as 1e-3) is adopted in this stage to accelerate convergence and improve the generalization ability of the detection module.
[0060] In the middle stage, the HSPL (Hard Sample Prototype Learning) mechanism is introduced, and the anatomical positioning module is gradually enabled. The training goal of this stage is to improve the detection ability of the candidate region detection module for complex lesions (such as low-contrast and small-size lesions) using HSPL, and to optimize the anatomical positioning module to improve the segmentation accuracy of the anatomical region. The optimization focus gradually shifts to the high-order structure perception learning of the candidate region detection module and the anatomical positioning module. The high-order structure perception learning is achieved through the HSPL mechanism, which enhances local features combined with attention mechanisms to improve the recognition ability of low-contrast lesions. The anatomical positioning module uses an improved 3D U-Net network to perform whole-brain structure segmentation on MRI images and generate probability distribution maps for each anatomical region.
[0061] To achieve the above goal, the HPSL mechanism is first introduced in this stage. Traditional detection methods often struggle to effectively distinguish real lesions from artifacts in the feature space when faced with pseudo-positive samples with high similarity. This problem is particularly pronounced in brain microbleed detection, as lesions are extremely small and have similar appearance characteristics to surrounding anatomical structures such as cortical blood vessels or calcification lesions. The core idea of the HSPL mechanism is to "pull" and "push" the constraints on positive samples (real brain microbleeds) and negative samples (pseudo-positive) in the feature space by constructing specific prototype vectors, so that the model can better focus on hard-to-distinguish hard samples.
[0062] Specifically, in the HSPL mechanism, first, the sliding window technique is used to crop the feature map generated by the detection network, retaining the local information in the image. Let a certain local region be X After forward propagation, the obtained feature map is denoted as f ( X ), and the number of channels of the feature map is denoted as N ch , and the depth, width, and height are denoted as d, w, and h, respectively. For each local region, according to whether it contains a real brain microbleed sample, the processing method is divided into two cases: if the region contains a real lesion, the corresponding artificial annotation coordinates e l can be directly used; if the region does not contain a real lesion, the position with the highest prediction probability in the region is taken as the candidate pseudo-positive point, i.e., calculating
[0063]
[0064] where represents the detection probability of each position e in the local region, S and the set of all positions in the region. In this way, each local region can extract a key coordinatec , for the extraction of subsequent feature vectors.
[0065] Subsequently, by collecting the response values at the coordinates f(X) on each channel, a feature vector c is formed. V c .
[0066] During the training process, two prototype vectors are constructed for the regions containing real lesions and not containing real lesions, respectively, denoted as M a and M b , where M a represents the prototype of real cerebral microbleeds, M b and the prototype of false positive samples. These two prototype vectors, as learnable parameters, will be updated constantly during the backpropagation process, thus representing the center positions of their respective classes in the feature space.
[0067] To achieve the goal of hard sample differentiation, a concentration loss function L con is designed.
[0068]
[0069] where · represents the Euclidean distance, n is a set boundary parameter, usually taking the value of 1. The core of this formula is to minimize the distance between V c and M a , while maximizing the distance between V c and M b , so that the feature vectors of real lesions are more concentrated in the feature space, while false positive samples are relatively dispersed. Such constraints not only help the model optimize feature representation during the training phase, but also effectively reduce the misjudgment rate of false positives during inference.
[0070] The introduction of the HSPL mechanism makes the entire detection network no longer need to design a separate classification module during single-stage end-to-end training, thus greatly reducing the overall complexity and computational burden of the system. At the same time, HSPL further makes up for the shortcomings of single-stage detection in fine-grained sample differentiation by focusing on learning hard samples, providing more accurate candidate region information for the subsequent dissection localization module.
[0071] During the training process, the loss function corresponding to the HSPL mechanism and the loss of the main detection network participate in the overall back propagation together to form a joint training mechanism. L final It can be expressed as:
[0072]
[0073] in, L cls is the classification loss implemented using FocalLoss, L reg is the positioning loss calculated using the bounding box regression technique (similar to the regression strategy in YOLO-v2), and L con This is the above-mentioned concentration loss. These are empirical weight parameters, respectively. Through extensive experimental tuning, the optimal balance is achieved, ensuring that each loss plays its due role throughout the training process. This joint loss strategy not only ensures the accuracy of the detection module in both category determination and position regression, but also effectively improves the ability to distinguish difficult-to-distinguish samples through HSPL.
[0074] Furthermore, global and local feature matching loss functions can be introduced during the training process. L HSPL ,and L final Together, the candidate region detection module is trained again after the first training. Optional:
[0075]
[0076] in, and Respectively represent i The loss is used to measure the consistency of the model's structural perception of the same lesion region in multi-scale paths, thereby improving the model's robustness in distinguishing complex lesions and ensuring the consistency of lesions at different scales across the detection module.
[0077] Correspondingly, the total loss function = L final + L HSPL .
[0078] After the detection module is robust, reduce the HSPL-related loss term in the loss function L H ( L H =L HSPL+L con ) of the proportion of the part of the anatomical positioning module, using a weakly supervised strategy, first only using the candidate region provided by the RPN for preliminary segmentation, rather than directly optimizing the global segmentation; that is, using the secondary trained candidate region detection module output RPN, and unlocking the parameters of the region in the anatomical positioning module corresponding to the RPN, and performing anatomical region segmentation on the RPN. At the same time, update the loss function, add the anatomical segmentation loss term L seg :
[0079]
[0080] wherein, L dice is the Dice Loss, which improves the segmentation accuracy of the anatomical structure; L CE is the cross-entropy loss, which ensures the accuracy of pixel-level classification.
[0081] This stage reduces the learning rate to η mind , so that the model can learn the characteristics more stably. At the same time, through the dynamic adjustment of the loss term, gradually increase the weight of L seg , so that the training focus gradually shifts from detection to anatomical segmentation. Still keep the detection module involved in training, but reduce its learning rate, reduce large-scale updates, and ensure overall stability. Optionally, the loss function can be constructed as follows L 1:
[0082]
[0083] Optionally, set the weight: ,
[0084] wherein, represents the time when the training is started, L 1, represents time, represents a hyperparameter to achieve a smooth transition of training focus, and optionally, can take a value between 3 and 8, which can be adjusted according to the convergence speed. This weight setting allows the weight coefficient to transition smoothly over time. Initially (detection task is still the main task). As the training progresses, gradually decreases, gradually increases, so that the anatomical matching loss gradually takes effect. Of course, in order to improve the training rate, L H can also be weakened or removed.
[0085] This phase uses the same sample pairs to jointly train the two basic modules. This not only improves training efficiency but also synchronizes their learning phases, facilitating rapid convergence of subsequent anatomical matching. For ease of distinction and description, this phase of training the anatomical localization module is referred to as a single training run.
[0086] Later stages: Global optimization and mutual constraints are achieved between the two basic modules (the candidate region detection module and the anatomical localization module). The training goal of this stage is to enable the detection module and the anatomical localization module to work together to improve the final accuracy of lesion detection and ensure the rationality of the anatomical structure. The system further optimizes the synergy between detection and segmentation through global parameter adjustment and information exchange between modules. Specifically, the system utilizes a refined feature matching strategy to automatically adjust the spatial position and anatomical features of the candidate region, ensuring that the final lesion localization results are both consistent with the image characteristics and anatomically reasonable. In addition, a dynamic weight adjustment strategy is used to optimize the balance between target detection loss and anatomical matching loss, allowing the detection module and anatomical localization module to achieve optimal performance within the overall framework.
[0087] In order to achieve the above goals, in this stage, the anatomical positioning module after the training is completely unlocked, and the anatomical matching loss function between the candidate region and the anatomical region is added, and the fine-tuned candidate region detection module and the anatomical positioning module after the training are further fine-tuned. L anatomy Including spatial consistency loss function L spatial And structural feature matching loss function L feat , which can be expressed as:
[0088]
[0089] in, i represents the index of the candidate region, N represents the number of candidate regions, and Respectively represent i candidate regions and their corresponding anatomical regions, IoU represents the spatial overlap, and Respectively represent i The structural feature vectors of the candidate regions and the structural feature vectors of their corresponding anatomical regions. L spatial The IoU value is used to ensure that the candidate region and the corresponding segmented region are close in space, preventing the lesion detection results from deviating from the anatomical region; L featTo ensure the high-dimensional feature similarity between the candidate region and the corresponding segmentation region, the final lesion localization result is not only consistent with the image feature, but also maintains the anatomical rationality.
[0090] At this time, the following loss function can be constructed:
[0091]
[0092] And the weight setting: Let the detection loss L det And the anatomical matching loss L anatomy Balance optimization. Ensure that the final prediction result is consistent with both the image feature and the anatomical rationality.
[0093] At the same time, further reduce the learning rate to η final (1e−5), only fine-tune, and improve overall performance. At this time, the gradient backtracking strategy: allow the gradient of the detection module to be backpropagated to the HSPL module, but do not directly affect the anatomical localization module to ensure that the segmentation accuracy is not disturbed by the update of the detection module. Of course, in order to improve the training rate, L H It can also be weakened or removed.
[0094] Further, in the calculation of L spatial And L feat If a candidate region corresponds to multiple anatomical segmentation regions R i , a weighted matching strategy can be used to calculate the matching score of each anatomical segmentation region S match ( R i ):
[0095]
[0096] Wherein, represents the spatial overlap degree of the candidate region and each anatomical region, represents the structural feature similarity of the candidate region and each anatomical region, and respectively represent the corresponding weights.
[0097] The anatomical region with the highest score is taken as the final matching region R final :
[0098]
[0099] Further, in the selection of the optimizer, the application mainly adopts an optimization method based on stochastic gradient descent, supplemented by a momentum and learning rate decay mechanism. In the initial stage, in order to enable the model to quickly capture the main features in the data and achieve faster convergence, a relatively high initial learning rate is set, and at the same time, by introducing the momentum mechanism, the gradient update process is effectively smoothed, and the training instability problem caused by gradient fluctuation is reduced. The momentum mechanism considers the gradient direction of the previous step at each parameter update, thereby accelerating the convergence speed to some extent and helping to reduce the interference of local minima on the training process.
[0100] As the training progresses, the model gradually reduces the learning rate in the later stage through the learning rate decay mechanism, so that the training process gradually tends to be stable, avoiding the oscillation and overfitting phenomenon caused by a too high learning rate. In order to further stabilize the training process, the network parameters are initialized using the He initialization method, which initializes according to the variance of the weight distribution in the network structure, effectively maintaining the consistency of the variance of the input data of each layer, and avoiding the problems of gradient disappearance or gradient explosion. In addition, the application combines the batch normalization technology in the network structure, which standardizes the data distribution before entering the activation function at each layer, thereby maintaining the stability of the data distribution, accelerating the convergence of the network, and reducing the sensitivity to the initial parameter setting.
[0101] In addition, during the training process at each stage, the application adopts cross-validation and early stopping mechanism to prevent overfitting and ensure the generalization ability of the model on the validation set.
[0102] In summary, the embodiment adopts a brain microbleed precise detection and positioning model based on deep learning, realizes efficient automatic detection of brain microbleeds, reduces the time and labor intensity of manual detection, and reduces the false positive rate in the detection process, improves the accuracy of detection; at the same time, the automatic anatomical positioning of CMBs is realized, which further improves the clinical value of detection. In particular, the embodiment determines whether the candidate region and the anatomical region match by spatial consistency and feature consistency, and timely eliminates the candidate region that does not match the anatomical structure; and sets a remedy measure of expanding the region for the misidentified candidate region, to determine the reasonable attribution of the candidate region as much as possible, and fully realize the anatomical information to realize the error detection and correction of the candidate region.
[0103] Further, in order to ensure the effective fusion of each module during the training process, the embodiment adopts an end-to-end training strategy. In the forward propagation process of the entire network, the data successively passes through preprocessing, feature extraction, HSPL feature optimization, candidate region generation, and loss calculation, and finally outputs the class and position information of the lesion detection. In the backward propagation, the gradient information of each module is transmitted to each other, and the network parameters are updated together, so that the entire system reaches the global optimum.
[0104] In particular, in the end-to-end processing flow, the detection module and the anatomical localization module respectively undertake the tasks of lesion candidate region generation and brain structure fine segmentation, and there is a close cooperative relationship between the two. Specifically, the detection module generates the center coordinates and class information of the candidate region through feature extraction and region proposal network (RPN), providing preliminary positioning basis for the anatomical localization module. The anatomical localization module uses an improved 3D U-Net network to perform whole brain segmentation on the MRI image, generates a probability map of each anatomical region, and classifies the candidate region provided by the detection module according to the anatomical structure to ensure that the detected lesions meet the anatomical structure characteristics.
[0105] In addition, through the multi-stage optimization strategy of gradually unlocking the module + dynamically adjusting the loss term + parameter freezing and unfreezing mechanism, the training focus gradually shifts from lesion detection to fine segmentation and anatomical matching, ensuring that the optimization order of each module is reasonable, and also allowing each module to fully exert its respective advantages at different stages, ultimately realizing the collaborative optimization of the overall system and improving the stability and accuracy of the overall detection.
[0106] Overall, through innovative design in the aspects of optimizer design, learning rate scheduling, and multi-stage training, the embodiment establishes a stable and efficient end-to-end training framework, enabling the model to maintain excellent detection ability and positioning accuracy when facing complex and variable MRI data, and has high clinical applicability and promotional value.
[0107] Figure 5 A structural schematic diagram of an electronic device provided by the embodiment of the present application is shown in FIG. 1, which includes a processor 60, a memory 61, an input device 62, and an output device 63. The number of processors 60 in the device can be one or more, and one processor 60 is taken as an example in the embodiment. Figure 5 The processor 60, the memory 61, the input device 62, and the output device 63 in the device can be connected through a bus or other means, and the connection through a bus is taken as an example in the embodiment. Figure 5 Figure 5 The memory 61 is a computer readable storage medium, which can be used to store software programs, computer executable programs, and modules, such as program instructions / modules corresponding to the method for precise detection and positioning of cerebral microbleeds combined with RPN and anatomical information in the embodiment of the present application. The processor 60 executes various functions and data processing of the device by running the software programs, instructions, and modules stored in the memory 61, that is, implements the method for precise detection and positioning of cerebral microbleeds combined with RPN and anatomical information.
[0108]
[0109] The memory 61 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal. Furthermore, the memory 61 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, the memory 61 may further include memory remotely located relative to the processor 60, and these remote memories may be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0110] The input device 62 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The output device 63 may include a display device such as a display screen.
[0111] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method for accurately detecting and locating cerebral microbleeds combining RPN and anatomical information according to any embodiment is implemented.
[0112] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device.
[0113] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0114] The program code embodied on the computer readable media can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0115] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0116] Finally, it should be noted that the above-mentioned embodiments are merely used to illustrate the technical solutions of the present application, instead of limiting the present application; although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that: they can still make modifications to the technical solutions recorded in the above-mentioned embodiments, or make equivalent replacements to some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present application.
Claims
1. A method for accurately detecting and localizing cerebral microbleeds by combining RPN and anatomical information, based on a deep learning-based model for accurately detecting and localizing cerebral microbleeds. The model includes a candidate region detection module and an anatomical localization module, wherein: The candidate region detection module is used to detect candidate regions of cerebral microbleeds based on the patient's brain image, and the anatomical positioning module is used to segment brain anatomical regions based on the patient's brain image; the method is characterized in that it includes: Acquire, from a brain image of the same patient, at least one candidate brain microbleed region obtained by the candidate region detection module, and at least one brain anatomical region obtained by the anatomical positioning module; Calculate the spatial overlap and structural feature similarity between each candidate region and the intersecting anatomical regions; According to the spatial overlap and structural feature similarity, as well as the corresponding threshold, the confidence of the unmatched candidate regions is lowered; The model is trained in the following way: Freezing the anatomical localization module, and training the candidate region detection module once using a classification loss function and a bounding box regression loss function; The HSPL mechanism is used to perform secondary training on the candidate region detection module after the first training. The global and local feature matching loss functions, as well as the concentration loss function, are added to the training. The candidate region detection module after secondary training is used to output the candidate region, and the parameters corresponding to the candidate region in the anatomical positioning module are unlocked to perform anatomical region segmentation on the candidate region; an anatomical segmentation loss function is added, and the unlocked parameters in the anatomical positioning module are trained according to the segmentation results, while fine-tuning the candidate region detection module; The trained anatomical positioning module is completely unlocked, and an anatomical matching loss function between the candidate region and the anatomical region is added to continue optimizing the trained anatomical positioning module and the fine-tuned candidate region detection module.
2. The method according to claim 1, characterized in that The step of lowering the confidence of the unmatched candidate regions according to the spatial overlap, structural feature similarity, and corresponding thresholds includes: If the spatial overlap between a candidate region and each intersecting anatomical region is less than an overlap threshold, and the structural feature similarity between the candidate region and each intersecting anatomical region is less than a similarity threshold, the confidence of the candidate region is lowered.
3. The method according to claim 2, characterized in that If the spatial overlap between a candidate region and each intersecting anatomical region is less than an overlap threshold, and the structural feature similarity between the candidate region and each intersecting anatomical region is less than a similarity threshold, lowering the confidence of the candidate region includes: If the spatial overlap between a candidate region and each intersecting anatomical region is less than an overlap threshold, and the structural feature similarity between the candidate region and each intersecting anatomical region is less than a similarity threshold, the candidate region is expanded; Calculate the spatial overlap and structural feature similarity between the expanded candidate region and the intersecting anatomical regions; Returning to the operation of lowering the confidence of the unmatched candidate regions based on the spatial overlap and structural feature similarity, as well as the corresponding thresholds.
4. The method according to claim 1, wherein The anatomical matching loss function Including spatial consistency loss function L spatial And structural feature matching loss function L feat , expressed as: in, i represents the index of the candidate region, N represents the number of candidate regions, and Respectively represent i candidate regions and their corresponding anatomical regions, IoU represents the spatial overlap, and Respectively represent i The structural feature vectors of the candidate regions and the structural feature vectors of their corresponding anatomical regions.
5. The method according to claim 4, characterized in that exist L spatial and L feat In the calculation of R cand Corresponding to multiple anatomical segmentation regions R i , using a weighted matching strategy to calculate the matching score of each anatomical segmentation region : in, represents the spatial overlap between the candidate region and each anatomical region, represents the similarity of the structural features between the candidate region and each anatomical region, and Represent the corresponding weights respectively; The anatomical region with the highest score is obtained as the anatomical region finally corresponding to the candidate region.
6. The method according to claim 1, characterized in that The adding of the anatomical segmentation loss function, training the unlocked parameters in the anatomical positioning module according to the segmentation results, and fine-tuning the candidate region detection module at the same time, include: Construct the following loss function L 1: in, L det represents the detection loss function, including the classification loss function and the bounding box regression loss function; L H represents the relevant loss function of the HSPL mechanism; L seg represents the segmentation loss function; Set the weights: , in, Indicates passing L 1Start training time, Indicates time, represents hyperparameters to achieve a smooth transition of training focus.
7. The method according to claim 6, characterized in that The anatomical matching loss function between the candidate region and the anatomical region is added to further optimize the trained anatomical positioning module and the fine-tuned candidate region detection module, including: Construct the following loss function: in, L anatomy represents the anatomical matching loss function; set up , to achieve global optimization.
8. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method for accurately detecting and locating cerebral microbleeds combining RPN and anatomical information as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that A computer program is stored thereon, which, when executed by a processor, implements the method for accurately detecting and locating cerebral microbleeds combining RPN and anatomical information as described in any one of claims 1-7.
Citation Information
Patent Citations
Method for carrying out micro-hemorrhage focus identification on brain image by using deep learning algorithm
CN120108011A
Cerebral microhemorrhage automatic detection method and system based on deep learning
CN110956634A
Method, device and equipment for positioning target area in brain CT image
CN115908381A