Craniocerebral ct image analysis method and system
By building an improved Net neural network segmentation model, combining attention mechanism and double-layer twin encoder decoder architecture, the noise and artifact problems of brain tumor image segmentation in the existing technology are solved, and high-precision craniocerebral CT image analysis is realized, reducing the risk of misdiagnosis, and is suitable for promotion in multiple medical places.
Patent Information
- Application Number
- CN202510463056.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing automatic segmentation method of craniocerebral CT images has problems such as noise, low contrast and artifacts when processing brain tumor images, resulting in insufficient recognition accuracy, difficult to meet real-time needs, and easy to miss diagnosis and misdiagnosis.
A modified Net neural network segmentation model is built, combining attention mechanism and a two-layer twin encoder decoder architecture, and high-precision brain tumor segmentation is achieved through preprocessing and optimization of the sum of Dice loss and cross entropy.
It improves the accuracy and accuracy of craniocerebral CT imaging analysis, reduces artificial errors, and is suitable for promotion and application in more medical places. It is simple to operate and easy to achieve.
Smart Images

Figure CN120374558A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of image segmentation recognition and analysis, and particularly to a method and system for analyzing brain CT images. Background Art
[0002] Brain CT scanning is a non-invasive imaging technique that generates cross-sectional images of the brain through X-ray and computer reconstruction techniques. CT scans have high resolution and density resolution capabilities, and can clearly display intracranial tissue structures, such as brain parenchyma, blood vessels, bone, and calcification foci. Common types of CT scans include plain scans, enhanced scans, dynamic scans, perfusion imaging, and three-dimensional reconstruction, etc. For example, enhanced scans can improve the contrast of diseased tissues by injecting contrast agents, thereby diagnosing diseases more accurately.
[0003] Image segmentation and processing: Utilize the gray-scale differences of CT images to extract information of specific regions (such as brain tissue, ventricles, calcification foci, etc.) through algorithms. For example, an automated method based on genetic algorithms can efficiently segment brain regions and calculate the volume of ventricles.
[0004] With the rapid development of deep learning technology, automatic segmentation of brain tumor images has become an emerging research direction in the field of medical imaging. However, due to the complexity of brain tumor images, such as problems like noise, low contrast, and artifacts, as well as the difficulty in meeting the real-time requirements, the existing automatic segmentation methods perform unsatisfactorily when dealing with these challenges.
[0005] Therefore, designing a method for analyzing brain CT images, proposing an efficient and accurate automatic segmentation algorithm for brain tumor segmentation tasks, improving the accuracy of recognition, reducing the errors caused by human subjectivity, and the risks of missed diagnosis and misdiagnosis, and saving manpower and material resources, is an important research content for those skilled in the art. Summary of the Invention
[0006] In view of this, it is necessary to provide a method for analyzing brain CT images, which can improve the accuracy of recognition, reduce invasiveness, accurately locate under ultrasound guidance, reduce errors and complications, be easy to promote and apply, relatively simple to operate, and suitable for popularization and application in more medical settings.
[0007] The present application provides a method and system for analyzing brain CT images. The method includes obtaining brain CT image data, preprocessing the brain CT image data, constructing an improved Net neural network segmentation model by combining an attention mechanism on the basis of a Net neural network segmentation model, inputting the brain CT image into the improved Net neural network segmentation model, and obtaining the output result of the Net neural network segmentation model. The Net neural network segmentation model is used to predict parameters and target region segmentation results based on the input brain CT image data and output them. By constructing an improved Net neural network segmentation model including a double-layer Siamese encoder and decoder architecture, pixel-level classification results can be obtained, and the positions and ranges of abnormal regions can be marked, which can improve the accuracy and precision of brain CT image analysis. Using the sum of the improved Dice loss and the improved cross-entropy as the loss function, the evaluation results are more accurate.
[0008] In a first aspect, an embodiment of the present application provides a method for analyzing brain CT images, the method including:
[0009] Obtaining brain CT image data;
[0010] Preprocessing the brain CT image data;
[0011] On the basis of a Net neural network segmentation model, constructing an improved Net neural network segmentation model by combining an attention mechanism;
[0012] Inputting the brain CT image into the improved Net neural network segmentation model, and obtaining the output result of the Net neural network segmentation model. The Net neural network segmentation model is used to predict parameters and target region segmentation results based on the input brain CT image data and output them.
[0013] Optionally, in an implementation manner of the first aspect of the present invention, the preprocessing the brain CT image data includes:
[0014] Clustering the gray values of each pixel of the brain CT image by the FCM algorithm, and dividing them into gray matter, white matter, cerebrospinal fluid, and hemorrhage regions;
[0015] Using morphological dilation and erosion methods to remove the skull part.
[0016] Optionally, in an implementation manner of the first aspect of the present invention, the constructing an improved Net neural network segmentation model by combining an attention mechanism on the basis of a Net neural network segmentation model includes:
[0017] The initial improved Net neural network segmentation model includes a double-layer twin encoder and decoder architecture, which uses two double-layer structures with the same structure. The encoder and decoder each have a double-layer design and perform the cranial CT image segmentation steps in parallel to obtain the segmentation results of the available features of cranial CT images in different modalities, and then average and fuse the segmentation results, and use the average value of the available features in different modalities as the fusion result;
[0018] Among them, the initial improved Net neural network segmentation model extracts the high-level semantic features of the image through the encoder respectively, and restores the high-level semantic features to the resolution of the original image through the decoder, and generates the segmentation result;
[0019] Among them, taking the U-Net network as the benchmark model, a double-layer twin encoder-decoder structure is used to capture multi-scale features, and spatial details are retained through skip connections; among them, the encoder adopts an asymmetric structure using a hybrid architecture of U-Net and CNN to extract local and global features, and introduces an improved residual attention mechanism to increase the weight of the tumor part of the cranial CT image in the network model, so that the network model focuses on the tumor part area and suppresses the interference of irrelevant tissues;
[0020] The decoder consists of a CNN architecture. An improved residual attention mechanism is added in the upsampling step to improve the segmentation accuracy. An attention gate is incorporated into the skip connection, and the obtained attention map is concatenated with the feature map obtained by upsampling the decoder to fuse the low-level and high-level features to achieve a more refined segmentation. The concatenated feature map after the skip connection is subjected to two convolution operations to enhance the feature extraction ability of the network for cranial CT images, and then a 1×1 convolution is performed for dimensionality reduction to reduce the number of network parameters.
[0021] To aggregate image features from different modalities and handle the possibility of missing one or more modalities, the average value of the available features from different modalities can be used as the first fusion result. That is, the average features of the fusion are obtained.
[0022] Optionally, in an implementation manner of the first aspect of the present invention, the architecture of the improved residual attention mechanism includes:
[0023] On the basis of U-Net, improved residual blocks and attention residual blocks are added. This architecture consists of a contraction module, a central module, and an expansion module. Residual blocks are used in the contraction module to reduce information loss; an attention mechanism is introduced in the expansion module to focus on key regions through gating and attention blocks; the attention mechanism enables the network to emphasize important features, enabling it to focus on key regions, and the integration of residual connections helps to reduce information loss during the training process, ensuring more efficient and stable learning, optimizing feature representation, and alleviating the network degradation phenomenon;
[0024] Among them, the arrow from the contraction module to the expansion module is called a skip connection. The purpose of the skip connection is to store additional information from important features. The upsampling output is combined with the skip connection to increase the localization of features, thereby creating a finer output at deeper levels of the model and enhancing the final prediction;
[0025] In the contraction module of each downsampling block, the residual technique is implemented to improve the extraction efficiency and reduce information loss. Among them, in each downsampling block, the input tensor passes through two modules. One module applies two filters for two-dimensional convolution and configures the rectified linear unit (ReLu) activation function. The other module tensor passes through a single convolution to form a shortcut module. The output results of the two modules are added together and then ReLu activation is performed.
[0026] Optionally, in an implementation manner of the first aspect of the present invention, the improved Net neural network segmentation model further includes a training process:
[0027] The initial improved Net neural network segmentation model is deeply trained using a training set. The number of filters and the learning rate are initialized. The best combination of hyperparameters of the model is adjusted while the model learns the features and patterns of the training data. By evaluating the performance of the model on a validation set, the learning rate and optimizer parameters are adjusted. By evaluating its performance on unseen data using a test set, the improved Net neural network segmentation model is obtained;
[0028] Among them, when adjusting the hyperparameter learning rate in the deep learning model, the learning rate controls the step size of parameter updates, determines the speed at which model parameters are updated in each iteration, and determines the step size of the model's update along the gradient direction in the parameter space. If the validation set loss fluctuates greatly, the learning rate needs to be reduced; if the convergence is too slow, it is appropriately increased, and the Adam adaptive optimizer is used for automatic adjustment;
[0029] The batch size affects the stability of gradient calculation. A larger batch size can accelerate training but occupies more memory; a smaller batch size can enhance generalization ability but increases noise. Adjust the number of samples used to update model parameters in each training iteration, and balance according to hardware conditions and task requirements, so that the model performs forward propagation, calculates the loss based on a batch of samples, and updates the model parameters through backpropagation;
[0030] In each training batch, each feature is normalized so that its mean is close to 0 and the standard deviation is close to 1, reducing internal covariate shift, accelerating convergence, and improving the stability of the model;
[0031] After obtaining a stable loss function Loss, the performance is evaluated using the accuracy ACC, Dice score, and precision PC index, and a line chart is generated through a visualization module.
[0032] Optionally, in an implementation manner of the first aspect of the present invention, obtaining the stable loss function Loss includes:
[0033] Adding an adaptive mechanism to the Dice loss and the cross-entropy loss function to obtain an improved Dice loss and an improved cross-entropy;
[0034] Using the sum of the improved Dice loss and the improved cross-entropy as the loss function, and solving and evaluating the model by minimizing the loss function. The formula of the loss function is:
[0035]
[0036] , where N c 、N r respectively represent the normalized coefficient pixel points, P i 0 represents the probability that the belonging area belongs to the i-th class, P i * represents the value determined according to the label of the belonging area, P i * ∈{0,1}, λ c 、λ r respectively represent the adjustment weights, respectively represent the predicted classes relative to the true positive TP, false positive FP, and false negative FN class probabilities. i and j represent pixel points, and ε represents the smoothing factor.
[0037] Optionally, in an implementation manner of the first aspect of the present invention, inputting the cranial CT image into the improved Net neural network segmentation model to obtain the output result of the Net neural network segmentation model includes:
[0038] Outputting the segmented area through the improved Net neural network segmentation model to obtain a pixel-level classification result, and marking the position and range of the abnormal area.
[0039] In a second aspect, an embodiment of the present application provides a cranial CT image analysis system, which is applied to the cranial CT image analysis method as described in the first aspect, and is characterized in that it includes:
[0040] A data acquisition module for acquiring cranial CT image data;
[0041] A data processing module for preprocessing the cranial CT image data;
[0042] A model construction module for constructing an improved Net neural network segmentation model by combining an attention mechanism on the basis of the Net neural network segmentation model;
[0043] An auxiliary diagnosis module, configured to input the cranial CT image into the improved Net neural network segmentation model, and obtain the output result of the Net neural network segmentation model, where the Net neural network segmentation model is used to predict parameters and target region segmentation results based on the input cranial CT image data and output them.
[0044] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0045] A processor;
[0046] A memory for storing instructions executable by the processor;
[0047] Wherein, when the processor is configured to execute the instructions, it implements the cranial CT image analysis method according to any one of claims 1 to 7.
[0048] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a program, and the program instructs the device to execute the cranial CT image analysis method according to any one of claims 1 to 7.
[0049] The present application provides a cranial CT image analysis method and system. By acquiring cranial CT image data; preprocessing the cranial CT image data; on the basis of the Net neural network segmentation model, constructing an improved Net neural network segmentation model by combining an attention mechanism; inputting the cranial CT image into the improved Net neural network segmentation model, and obtaining the output result of the Net neural network segmentation model, where the Net neural network segmentation model is used to predict parameters and target region segmentation results based on the input cranial CT image data and output them.
[0050] Advantageous effects:
[0051] (1) By constructing an improved Net neural network segmentation model including a double-layer siamese encoder and decoder architecture, a pixel-level classification result is obtained, and the position and scope of the abnormal region are marked, which can improve the accuracy and precision of cranial CT image analysis.
[0052] (2) Using the sum of the improved Dice loss and the improved cross-entropy as the loss function, solving and evaluating the model by minimizing the loss function, optimizing and improving the loss function of the model, and the evaluation result is more accurate.
[0053] (3) Easy to promote and apply: Since no surgery is required and the operation is relatively simple, it is suitable for promotion and application in more medical places. Description of the Drawings
[0054] Figure 1Schematic diagram of the process of a brain CT image analysis method provided by an embodiment of the present application.
[0055] Figure 2 Overall structure diagram of an improved Net neural network segmentation model provided by an embodiment of the present application.
[0056] Figure 3 Encoder-decoder network architecture diagram provided by an embodiment of the present application.
[0057] Figure 4 Architecture diagram of an improved residual attention mechanism provided by an embodiment of the present application.
[0058] Figure 5 Schematic diagram of the modules of a brain CT image analysis system provided by an embodiment of the present application.
[0059] Figure 6 Schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0060] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application.
[0061] It should be noted that "at least one" in the embodiments of the present application means one or more, and multiple means two or more. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0062] It should be noted that in the embodiments of the present application, terms such as "first" and "second" are only used for the purpose of distinguishing descriptions and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order. Features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the embodiments of the present application, words such as "exemplary" or "for example" are used to mean as an example, illustration or explanation. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.
[0063] All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts are within the scope of protection of this application.
[0064] Embodiment 1
[0065] This application provides a method and system for analyzing cranial CT images. By acquiring cranial CT image data, preprocessing the cranial CT image data, constructing an improved Net neural network segmentation model by combining an attention mechanism on the basis of the Net neural network segmentation model, and inputting the cranial CT image into the improved Net neural network segmentation model to obtain the output result of the Net neural network segmentation model, the Net neural network segmentation model is used to predict parameters and the target area segmentation result based on the input cranial CT image data and output. By constructing an improved Net neural network segmentation model including a double-layer Siamese encoder and decoder architecture, a pixel-level classification result can be obtained, marking the position and scope of the abnormal area, which can improve the accuracy and precision of cranial CT image analysis. Using the sum of the improved Dice loss and the improved cross-entropy as the loss function, the evaluation result is more accurate.
[0066] Figure 1 It is a schematic flowchart of the method for analyzing cranial CT images provided by an embodiment of this application.
[0067] As Figure 1 shown, a method for analyzing cranial CT images includes:
[0068] S1: Acquire cranial CT image data.
[0069] It can be understood that in this embodiment, acquiring cranial CT image data follows legal and compliant paths.
[0070] The data sources include public data sets, such as open-source medical image databases, or hospitals or medical institutions, such as clinical collaborations, to ensure the legal source of the data and protect patient privacy: the data needs to be anonymized (removing sensitive information such as names and IDs). The medical image data format is usually in DICOM format and requires professional software (such as ITK-SNAP, 3D Slicer) to view or process.
[0071] The data acquisition methods include: Patient position and positioning line: The patient takes the supine position on a hard board bed, and the head is fixed on the CT scan pillow support to maintain the midline position to reduce motion artifacts. Common scanning baselines include the orbitomeatal line (the line connecting the external auditory meatus and the outer canthus of the eye) and the supraorbitalmeatal line (the line connecting the midpoint of the upper edge of the eyebrow and the external auditory meatus), among which the orbitomeatal line is the most commonly used, with accurate positioning and an ideal scanning range.
[0072] S2: Preprocess the cranial CT image data.
[0073] It can be understood that in this embodiment, the preprocessing of the cranial CT image data includes:
[0074] Cluster the gray values of each pixel in the cranial CT image through the FCM algorithm, and divide them into gray matter, white matter, cerebrospinal fluid, and hemorrhage areas;
[0075] Use morphological dilation and erosion methods to remove the skull part.
[0076] Use algorithms (such as FCM clustering) for intracranial structure segmentation, exclude skull interference, and improve the accuracy of lesion recognition. Deep learning models (such as convolutional neural networks) can be used for automated feature extraction and classification.
[0077] Specifically, on the basis of the above preprocessing, basic preprocessing processes need to be carried out, including data cleaning: removing salt-and-pepper noise through a median filter, reducing high-frequency noise for smoothing the image through a Gaussian filter, and removing noise points in the image through a mean filter.
[0078] Removing noise is the primary link in data cleaning. Cranial CT images may be affected by noise interference from imaging devices, scanning processes, or transmission processes. This noise exists in the form of speckle, artifacts, motion artifacts, etc., seriously affecting the quality of the image and the accurate display of the cranial structure by the image. Therefore, removing noise is a key step in improving the image quality in the data cleaning step. Filters are the most commonly used tools for removing noise. This invention uses a median filter, a Gaussian filter, and a mean filter. The median filter is used to remove salt-and-pepper noise, the Gaussian filter is used to smooth the image and reduce high-frequency noise, and the mean filter is used to remove noise points in the image.
[0079] Perform image screening and elimination. The images to be eliminated include: images with too low quality, incomplete image content, serious artifacts in the image, and images that do not meet the research objectives.
[0080] Through such screening and elimination steps, it can be preliminarily ensured that all the images used are of high quality, complete in content, and meet the research objectives, which is of great significance for improving the accuracy and efficiency of subsequent image processing and analysis. Data cleaning is an essential step in the panoramic data preprocessing of cranial CT images. Through effective data cleaning, the quality of cranial CT image data can be significantly improved, providing basic data guarantee for subsequent image segmentation and analysis tasks. This is of great significance for improving the accuracy and stability of cranial CT image data segmentation, as well as enhancing the model performance and generalization ability.
[0081] Image resampling can also be performed through bilinear interpolation. Since the acquisition devices, parameters, and resolutions of cranial CT image data may vary, this may result in differences in the size and resolution of the obtained images. Therefore, in order to ensure the consistency and fairness of subsequent processing steps, image resampling is performed so that all images have the same size and resolution.
[0082] The basic idea of image resampling is to change the size and resolution of an image through interpolation or sampling methods. Interpolation is a method of predicting unknown pixel values based on known pixel values. Several modern mainstream image resampling methods (all set to shrink by 10%) include bilinear interpolation, nearest neighbor interpolation, bicubic interpolation, and Lanczos interpolation. Tags are marked for specific regions or objects in the image, and tag format conversion and tag resampling are performed.
[0083] A tag (or marker) is a description of a specific region or object in an image, used to indicate specific structures in the image, such as gray matter, white matter, cerebrospinal fluid, hemorrhage regions, etc. The tag processing in this study mainly includes two parts: tag generation and tag preprocessing.
[0084] The generated tags need to be preprocessed to adapt to subsequent image processing tasks. The tag preprocessing in this study includes tag format conversion and tag resampling. However, tag generation is a time-consuming task that requires basic professional knowledge and is sometimes affected by subjective factors. Through effective tag processing, high-quality and consistent tags can be generated, laying a good foundation for model training and validation.
[0085] The image is enhanced by a contrast enhancement method, which includes one or more of linear contrast stretching, non-linear contrast stretching, and histogram equalization. The edges and details of the structures in the cranial CT image are enhanced by sharpening the image, and the texture and shape are enhanced by filters.
[0086] Data augmentation increases the size and diversity of the training set by creating variant forms of the original data, thereby improving the generalization ability of the model and reducing the risk of overfitting. The data augmentation method used in this invention includes geometric transformation and random noise injection. Geometric transformation is a common data augmentation method, including translation, rotation, scaling, and flipping operations. These operations can change the position, orientation, and scale of the image, but do not change the content of the image. For example, rotating a cranial CT image can simulate different shooting angles and increase the robustness of the model to angle changes. Random noise injection is a more advanced data augmentation method that adds random noise to the image to simulate the noise that may occur in the actual image acquisition process. However, some data augmentation methods are not fully applicable to cranial CT images.
[0087] S3: Based on the Net neural network segmentation model, an improved Net neural network segmentation model is constructed by combining an attention mechanism.
[0088] It is understandable that the role of the residual block: The residual block is one of the core components in the Net neural network segmentation model. Through residual connections, the residual block can retain the detailed information in the original input feature map. This is very important for image segmentation tasks because the segmentation results require good spatial details and boundary information. Feature extraction: The convolutional layers in each residual block can effectively extract the local features of the image. The stacking of multiple residual blocks allows the network to extract higher-level semantic features, thus better performing the image segmentation task.
[0089] The role of the encoder: By successive downsampling and feature extraction, the input image is transformed into a feature representation rich in semantics.
[0090] The role of the decoder: Restore the feature map extracted by the encoder to the resolution of the original image and generate the final segmentation result. Resolution restoration: The decoder achieves the gradual restoration of the resolution through upsampling layers. Through multiple upsamplings, the decoder can gradually restore the resolution of the feature map to be close to that of the original image, thus restoring the details and boundaries of the image.
[0091] Figure 2 This is the overall structure diagram of the improved Net neural network segmentation model provided by an embodiment of the present application. As Figure 2 shown, in this embodiment, on the basis of the Net neural network segmentation model, an improved Net neural network segmentation model is constructed by combining an attention mechanism, including:
[0092] The initial improved Net neural network segmentation model includes a double-layer twin encoder and decoder architecture, and this architecture uses two double-layer structures with the same structure, namely the first encoder and decoder architecture and the second encoder and decoder architecture. The encoder and decoder each have a double-layer design and perform the brain CT image segmentation steps in parallel to obtain the segmentation results of the available features of different modalities of brain CT images, and then average and fuse the segmentation results, and use the average value of the available features of different modalities as the fusion result;
[0093] Among them, the initial improved Net neural network segmentation model extracts the high-level semantic features of the image through the encoder respectively, and restores the high-level semantic features to the resolution of the original image through the decoder and generates the segmentation result.
[0094] Among them, taking the U-Net network as the benchmark model, a double-layer twin encoder-decoder structure is used to capture multi-scale features, and skip connections are used to retain spatial details; among them, the encoder adopts an asymmetric structure using a hybrid architecture of U-Net and CNN to extract local and global features, and an improved residual attention mechanism is introduced to increase the weight of the tumor part in the brain CT image in the network model, so that the network model focuses on the tumor part area and suppresses the interference of irrelevant tissues.
[0095] Figure 3 This is the encoder-decoder network architecture diagram provided by an embodiment of the present application. Specifically, as Figure 3 shown, the decoder consists of a CNN architecture. An improved residual attention mechanism is added in the upsampling step to improve the segmentation accuracy. An attention gate is incorporated into the skip connection, and the obtained attention map is concatenated with the feature map obtained by upsampling the decoder to fuse low-level and high-level features, thereby achieving a more refined segmentation. The concatenated feature map after the skip connection is subjected to two convolutional operations to enhance the network's feature extraction ability for cranial CT images, and then a 1×1 convolution is performed for dimensionality reduction to reduce the number of network parameters.
[0096] Figure 4 This is the architecture diagram of the improved residual attention mechanism provided by an embodiment of the present application. As Figure 4 shown, specifically, the architecture of the improved residual attention mechanism includes:
[0097] An improved residual block and an attention residual block are added on the basis of U-Net. This architecture consists of a contraction module, a central module, and an expansion module. The residual block is used in the contraction module to reduce information loss; the attention mechanism is introduced in the expansion module to focus on key regions through gating and attention blocks; the attention mechanism enables the network to emphasize important features, enabling it to focus on key regions. The integration of residual connections helps reduce information loss during the training process, ensuring more efficient and stable learning, optimizing feature representation, and alleviating the network degradation phenomenon;
[0098] Among them, the arrow from the contraction module to the expansion module is called a skip connection. The purpose of the skip connection is to store additional information from important features. The upsampling output is connected to the skip connection to increase the localization of features, thereby creating a more refined output at a deeper level of the model and enhancing the final prediction;
[0099] In the contraction module of each downsampling block, the residual technique is implemented to improve the extraction efficiency and reduce information loss. Among them, in each downsampling block, the input tensor passes through two modules. One module applies two filters for two-dimensional convolution and configures the rectified linear unit ReLu activation function. The other module tensor passes through one convolution to form a shortcut module. The output results of the two modules are added and activated by ReLu.
[0100] Specifically, as Figure 4As shown, it consists of a contraction module (left), a bottleneck (central), and an expansion module (right). The arrow from the contraction module to the expansion module is called a skip connection. The purpose of the skip connection is to store additional information from important features, where the upsampled output is connected to the skip connection to increase the localization of features, resulting in enhanced final predictions. In this work, a more complex U-Net architecture was adopted, using attention and residual techniques. The attention mechanism enables the network to emphasize important features, allowing it to focus on key regions. The integration of residual connections helps reduce information loss during the training process, ensuring more efficient and stable learning. The model was further enhanced by implementing Monte Carlo dropout technique.
[0101] The architecture of the model is designed around two downsampling blocks, a bottleneck, and two upsampling blocks. In the contraction module part of each downsampling block, the residual technique was implemented to improve efficiency, thereby reducing information loss. This means that in each downsampling block, its input tensor passes through two modules. In the main module, two 2D convolutions with a filter size of 3 by 3 were applied, each followed by a rectified linear unit (ReLu) activation function. At the same time, in the secondary module, the tensor undergoes one convolution to form a shortcut module. After these steps, the outputs of the two modules are added together and ReLu activated.
[0102] The improved Net neural network segmentation model also includes a training process:
[0103] The initial improved Net neural network segmentation model is deeply trained with a training set, initializing the number of filters and the learning rate, adjusting the optimal combination of hyperparameters of the model while the model learns the features and patterns of the training data, evaluating the performance of the model on the validation set, adjusting the learning rate and optimizer parameters, and evaluating its performance on unseen data through the test set to obtain the improved Net neural network segmentation model;
[0104] Among them, adjusting the learning rate, a hyperparameter in the deep learning model, the learning rate controls the step size of parameter updates, determines the speed at which model parameters are updated in each iteration, and determines the step size of the model's update along the gradient direction in the parameter space. If the validation set loss fluctuates greatly, the learning rate needs to be reduced; if the convergence is too slow, it should be appropriately increased, and it is automatically adjusted using the Adam adaptive optimizer;
[0105] The batch size affects the stability of gradient calculation. A larger batch size can accelerate training but occupies more memory; a smaller batch size can enhance generalization ability but increases noise. Adjust the number of samples used to update model parameters in each training iteration, balance according to hardware conditions and task requirements, and enable the model to perform forward propagation, calculate loss based on a batch of samples, and update model parameters through backpropagation;
[0106] In each training batch, each feature is normalized so that its mean is close to 0 and its standard deviation is close to 1, reducing internal covariate shift, accelerating convergence, and improving the stability of the model;
[0107] After obtaining a stable loss function Loss, the performance is evaluated using the accuracy ACC, Dice score, and precision PC index, and a line chart is generated through a visualization module.
[0108] Specifically, the obtaining of the stable loss function Loss includes:
[0109] An adaptive mechanism is added based on the Dice loss and cross-entropy loss function to obtain an improved Dice loss and an improved cross-entropy;
[0110] The sum of the improved Dice loss and the improved cross-entropy is used as the loss function, and the model is solved and evaluated by minimizing the loss function. The loss function formula is:
[0111]
[0112] , where N c 、N r respectively represent the normalized coefficient pixel points, P i 0 represents the probability that the belonging region belongs to the i-th class, P i * represents the value determined according to the label of the belonging region, P i * ∈{0,1}, λ c 、λ r respectively represent the adjustment weights, respectively represent the predicted classes relative to the true positive TP, false positive FP, and false negative FN class probabilities, i and j represent pixel points, and ε represents a smoothing factor.
[0113] S4: Input the cranial CT image into the improved Net neural network segmentation model to obtain the output result of the Net neural network segmentation model. The Net neural network segmentation model is used to predict parameters and target region segmentation results based on the input cranial CT image data and output them.
[0114] Specifically, in this embodiment, the inputting of the cranial CT image into the improved Net neural network segmentation model to obtain the output result of the Net neural network segmentation model includes:
[0115] The segmented region is output through the improved Net neural network segmentation model to obtain a pixel-level classification result, and the position and range of the abnormal region are marked.
[0116] Example 2
[0117] As Figure 5 shown, the present application provides a cranial CT image analysis system, which is applied to the cranial CT image analysis method described in Example 1, and includes: a data acquisition module 11, a data processing module 12, a model construction module 13, and an auxiliary diagnosis module 14.
[0118] It can be understood that, in this embodiment, the data acquisition module 11 is used to acquire cranial CT image data.
[0119] It can be understood that, in this embodiment, the data processing module 12 is used to preprocess the cranial CT image data.
[0120] It can be understood that, in this embodiment, the model construction module 13 is used to construct an improved Net neural network segmentation model by combining an attention mechanism on the basis of the Net neural network segmentation model.
[0121] It can be understood that, in this embodiment, the auxiliary diagnosis module 14 is used to input the cranial CT image into the improved Net neural network segmentation model to obtain the output result of the Net neural network segmentation model, and the Net neural network segmentation model is used to predict parameters and target region segmentation results based on the input cranial CT image data and output them.
[0122] The present application provides a cranial CT image analysis method and system, which includes acquiring cranial CT image data; preprocessing the cranial CT image data; constructing an improved Net neural network segmentation model by combining an attention mechanism on the basis of the Net neural network segmentation model; inputting the cranial CT image into the improved Net neural network segmentation model to obtain the output result of the Net neural network segmentation model, and the Net neural network segmentation model is used to predict parameters and target region segmentation results based on the input cranial CT image data and output them. By constructing an improved Net neural network segmentation model including a double-layer twin encoder and decoder architecture, pixel-level classification results are obtained, and the positions and ranges of abnormal regions are marked, which can improve the accuracy and precision of cranial CT image analysis. Using the sum of the improved Dice loss and the improved cross-entropy as the loss function, the evaluation results are more accurate.
[0123] Figure 6 is an electronic device provided by an embodiment of the present application. As Figure 6 shown, the electronic device at least includes the following parts: a processor 101, a memory 100, a communication interface 103, and a bus 102.
[0124] In an embodiment of the present application, the memory 100 is used to store executable instructions for the processor 101, and when the processor 101 is configured to execute the instructions, it implements the Figure 5 device module for analyzing cranial CT images shown.
[0125] In an embodiment of the present application, a computer-readable storage medium includes instructions that direct the device to execute the method according to the first aspect. For example, the instructions direct the device to execute the Figure 1 method for predicting device status and risk assessment based on artificial intelligence shown in the process steps.
[0126] The program operating in the electronic device according to an embodiment of the present application can be a program that controls a central processing unit (CPU) and the like to implement the functions of the above-described embodiments related to a solution of the present invention (a program that causes a computer to function). Then, the information processed by these devices is temporarily stored in a random access memory (RAM) during its processing, and then stored in various ROMs such as a read-only memory (FlashROM), a hard disk drive (HDD), etc. It is read out, corrected, and written by the CPU as needed.
[0127] It should be noted that a part of the electronic device of the above-described embodiment can also be implemented by a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and implemented by reading the program recorded on the recording medium into the computer and executing it.
[0128] It should be noted that the "computer" mentioned here refers to a computer built into an electronic device, which is a computer including hardware such as an OS and peripheral devices. In addition, the "computer-readable recording medium" refers to a removable medium such as a floppy disk, a magneto-optical disk, a ROM, a CD-ROM, etc., and a storage device such as a hard disk built into a computer.
[0129] Moreover, the "computer-readable recording medium" can include: a medium that stores a program dynamically for a short time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line; a medium that stores a program for a fixed time, such as a volatile memory inside a computer of a server or a client in this case. In addition, the above program can be a part of the program for implementing the above functions, and can also be a program that can implement the above functions by combining with a program already recorded in a computer.
[0130] In addition, the electronic device in the above-described embodiment can also be implemented as an aggregate (device group) composed of multiple devices. Each device constituting the device group may have all or part of the functions or functional blocks of the electronic device in the above-described embodiment. As the device group, it is sufficient to have all the functions or functional blocks of the electronic device.
[0131] Those of ordinary skill in the art should recognize that the above embodiments are only used to illustrate the present application, rather than to limit the present application. As long as appropriate changes and variations made to the above embodiments fall within the scope of the spirit of the present application, they fall within the scope of protection required by the present application.
Claims
1. A method for analyzing cranial CT images, characterized in that, The method includes: Obtaining cranial CT image data; Preprocessing the cranial CT image data; Based on the Net neural network segmentation model, constructing an improved Net neural network segmentation model by combining an attention mechanism; Inputting the cranial CT image into the improved Net neural network segmentation model to obtain the output result of the Net neural network segmentation model, where the Net neural network segmentation model is used to predict parameters and the target region segmentation result based on the input cranial CT image data and output them.
2. The method for analyzing cranial CT images according to claim 1, characterized in that, The preprocessing of the cranial CT image data includes: Clustering the gray value of each pixel of the cranial CT image through the FCM algorithm and dividing it into gray matter, white matter, cerebrospinal fluid, and hemorrhage regions; Using morphological dilation and erosion methods to remove the skull part.
3. A method for analyzing cranial CT images according to claim 2, characterized in that, The constructing of the improved Net neural network segmentation model by combining an attention mechanism based on the Net neural network segmentation model includes: The initial improved Net neural network segmentation model includes a double-layer twin encoder and decoder architecture, and this architecture uses two double-layer structures with the same structure. The encoder and decoder each have a double-layer design and perform the cranial CT image segmentation steps in parallel to obtain the segmentation results of the available features of different modalities of the cranial CT image, and then average and fuse the segmentation results, and use the average value of the available features of different modalities as the fusion result; Among them, the initial improved Net neural network segmentation model extracts the high-level semantic features of the image through the encoder respectively, and restores the high-level semantic features to the resolution of the original image through the decoder and generates the segmentation result; Among them, taking the U-Net network as the benchmark model, capturing multi-scale features through a double-layer twin encoder-decoder structure, and retaining spatial details through skip connections; among them, the encoder adopts an asymmetric structure using a hybrid architecture of U-Net and CNN, extracts local and global features, and introduces an improved residual attention mechanism to increase the weight of the tumor part of the cranial CT image in the network model, so that the network model focuses on the tumor part area and suppresses the interference of irrelevant tissues; The decoder consists of a CNN architecture. An improved residual attention mechanism is added in the upsampling step to improve the segmentation accuracy. An attention gate is incorporated in the skip connection, and the obtained attention map is spliced with the feature map obtained by upsampling the decoder to fuse low-level and high-level features to achieve a more refined segmentation. The spliced feature map after the skip connection is subjected to two convolutional operations to enhance the feature extraction ability of the network for the cranial CT image, and then a 1×1 convolution is performed for dimensionality reduction to reduce the number of network parameters.
4. A method for analyzing cranial CT images according to claim 3, characterized in that, The architecture of the improved residual attention mechanism includes: Based on U-Net, improved residual blocks and attention residual blocks are added. This architecture consists of a contraction module, a central module, and an expansion module. Residual blocks are used in the contraction module to reduce information loss; an attention mechanism is introduced in the expansion module, and the key regions are focused through gating and attention blocks; the attention mechanism enables the network to emphasize important features, enabling it to concentrate on key regions. The integration of residual connections helps reduce information loss during the training process, ensuring more efficient and stable learning, optimizing feature representation, and alleviating the network degradation phenomenon; Among them, the arrow from the contraction module to the expansion module is called a skip connection. The purpose of the skip connection is to store additional information from important features. Among them, the upsampling output is connected to the skip connection to increase the localization of features, thereby creating a finer output at a deeper level of the model and enhancing the final prediction; In the contraction module of each downsampling block, the residual technique is implemented to improve the extraction efficiency and reduce information loss. Among them, in each downsampling block, the input tensor passes through two modules. One module applies two filters for two-dimensional convolution and configures the rectified linear unit ReLU activation function. The other module tensor passes through one convolution to form a shortcut module. The output results of the two modules are added and ReLU activation is performed.
5. The method for analyzing cranial CT images according to claim 4, wherein, The improved Net neural network segmentation model also includes a training process: The initial improved Net neural network segmentation model is deeply trained through the training set. The number of filters and the learning rate are initialized. The best combination of hyperparameters of the model is adjusted during the model's learning of the features and patterns of the training data. By evaluating the performance of the model on the validation set, the learning rate and optimizer parameters are adjusted. By evaluating its performance on unseen data through the test set, the improved Net neural network segmentation model is obtained; Among them, one hyperparameter learning rate in the deep learning model is adjusted. The learning rate controls the step size of parameter updates, controls the speed at which model parameters are updated in each iteration, and determines the step size of the model's update along the gradient direction in the parameter space. If the validation set loss fluctuates greatly, the learning rate needs to be reduced; if the convergence is too slow, it is appropriately increased and automatically adjusted using the Adam adaptive optimizer; The batch size affects the stability of gradient calculation. A larger batch size can accelerate training but occupies more memory; a smaller batch size can enhance generalization ability but increase noise. Adjust the number of samples used to update model parameters in each training iteration, and balance according to hardware conditions and task requirements, so that the model performs forward propagation, calculates the loss based on a batch of samples, and updates model parameters through backpropagation; In each training batch, each feature is normalized so that its mean is close to 0 and the standard deviation is close to 1, reducing internal covariate shift, accelerating convergence, and improving the stability of the model; After obtaining a stable loss function Loss, the performance is evaluated by the accuracy ACC, Dice score, and precision PC index, and a line chart is generated through the visualization module.
6. A method for analyzing cranial CT images according to claim 5, characterized in that, The obtaining of the stable loss function Loss includes: An adaptive mechanism is added to the Dice loss and cross-entropy loss functions to obtain an improved Dice loss and an improved cross-entropy; The sum of the improved Dice loss and the improved cross-entropy is used as the loss function, and the model is solved and evaluated by minimizing the loss function. The formula of the loss function is: , Among them, N c and N r respectively represent the normalized coefficient pixel points, P i 0 represents the probability that the belonging area belongs to the i-th class, P i * represents the value determined according to the label of the belonging area, P i * ∈{0, 1}, λ c and λ r respectively represent the adjustment weights, respectively represent the predicted classes relative to the true positive TP, false positive FP, and false negative FN class probabilities, i and j represent pixel points, and ε represents the smoothing factor.
7. A method for analyzing cranial CT images according to claim 6, characterized in that, Inputting the cranial CT image into the improved Net neural network segmentation model to obtain the output result of the Net neural network segmentation model includes: Outputting the segmented region through the improved Net neural network segmentation model to obtain a pixel-level classification result, and marking the position and scope of the abnormal region.
8. A cranial CT image analysis system, applied to the cranial CT image analysis method according to any one of claims 1 to 7, characterized in that, Including: A data acquisition module for acquiring cranial CT image data; A data processing module for preprocessing the cranial CT image data; A model construction module for constructing an improved Net neural network segmentation model by combining the attention mechanism on the basis of the Net neural network segmentation model; An auxiliary diagnosis module for inputting the cranial CT image into the improved Net neural network segmentation model to obtain the output result of the Net neural network segmentation model. The Net neural network segmentation model is used to predict parameters and target region segmentation results based on the input cranial CT image data and output them.
9. An electronic device, characterized in that, Including: A processor; A memory for storing processor-executable instructions; Wherein, when the processor is configured to execute the instructions, it implements the cranial CT image analysis method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program, and the program instructs the device to execute the cranial CT image analysis method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Cerebral hemorrhage segmentation method and system based on multi-model combination
CN112348796A
Improved U-Net brain tumor segmentation method based on attention mechanism and multi-scale feature fusion
CN115424103A
Whole-brain clinical target region segmentation method based on series double-attention U-Net network
CN116934772A
TransUNet image reconstruction algorithm and system based on hyper-parameter optimization
CN118172492A
MRI brain tumor image segmentation method and system
CN118351130A
Cited By
Craniocerebral image cutting method, craniocerebral image segmentation model training method and related equipment
CN121767282A