Craniocerebral image cutting method, craniocerebral image segmentation model training method and related equipment
By combining the encoder and decoder of a deep learning segmentation model, rapid, accurate, and automated cropping of the skull and cerebral cortex is achieved, solving the problems of time-consuming, labor-intensive, and subjective cropping results in existing technologies, and improving the efficiency and accuracy of cranial image cropping.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-03-31
AI Technical Summary
In existing technologies, brain image cropping relies on manual operation, which is time-consuming, inefficient, and the cropping results are highly subjective and difficult to standardize. This results in a time-consuming and labor-intensive process for cerebral vascular examination, with frequent segmentation errors.
A deep learning segmentation model is adopted, which combines encoder and decoder to achieve fast, accurate and automated cropping of the skull and cerebral cortex regions. Multi-level feature extraction and feature fusion are used to generate segmentation results that include the cranial mask, ensuring accurate identification of brain parenchyma and vascular regions.
It significantly reduces cropping time from several minutes to seconds, while ensuring segmentation accuracy and complete preservation of vascular structures across different scanning ranges, thus improving the efficiency and quality of maximum intensity projection image generation.
Smart Images

Figure CN121767282A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer and communication technology, and more specifically, to a method for cropping cranial images, a method for training cranial image segmentation models, and related equipment. Background Technology
[0002] The cerebral vascular examination process requires manual cropping of skull and cortical signals. This involves technicians repeatedly changing perspectives to remove portions of the skull and cerebral cortex before generating a final Maximum Intensity Projection (MIP) image for medical diagnosis. This cropping process is time-consuming and labor-intensive, requiring multiple perspective changes and region selections. Furthermore, the cropping results rely on the technician's subjective judgment, making segmentation errors prone to occur and difficult to standardize. Therefore, a fast and accurate skull and cortical cropping algorithm is proposed. This algorithm reduces cropping time (from 3 minutes to 2 seconds) while ensuring segmentation accuracy, thereby guaranteeing the quality of the generated MIP image. Summary of the Invention
[0003] The embodiments of this application provide a method and related equipment for cropping cranial images, aiming to solve the technical problems of long time consumption, low efficiency, strong subjectivity of cropping results and difficulty in standardization caused by manual operation when cropping cranial images in the prior art.
[0004] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0005] According to one aspect of the embodiments of this application, a method for cropping a cranial image is provided, comprising: acquiring a cranial image to be processed; inputting the cranial image to be processed into a cranial image segmentation model to obtain a segmentation result, the segmentation result including a cranial mask, wherein the non-basal surface of the cranial mask is a brain parenchyma region, and the basal surface is a regular region containing all target blood vessels; cropping the cranial image to be processed based on the segmentation result to obtain a cropped result; wherein the cranial image segmentation model includes an encoder and a decoder; the step of inputting the cranial image to be processed into the cranial image segmentation model to obtain a segmentation result specifically includes: inputting the cranial image to be processed into the encoder for a predetermined number of downsampling and feature extractions to obtain a target feature image; inputting the target feature image into the decoder for a predetermined number of upsamplings and then performing feature fusion to obtain a segmentation result.
[0006] According to one aspect of the embodiments of this application, a training method for a cranial image segmentation model is provided, comprising: acquiring multiple cranial images, labeling each cranial image to form cranial image samples, each cranial image sample being labeled with a corresponding mask-like segmentation label, wherein the non-basal surface of the cranial image is the brain parenchyma region, and the basal surface of the cranial image is a regular region containing all target blood vessels; inputting the cranial image samples one by one into a cranial image segmentation model to obtain segmentation results; updating the parameters of the cranial image segmentation model according to the output segmentation results and the segmentation labels until a predetermined termination condition is reached, ending the training, and obtaining a trained cranial image segmentation model; wherein, the cranial image segmentation model includes an encoder and a decoder; the step of inputting the cranial image samples one by one into the cranial image segmentation model to obtain segmentation results specifically includes: inputting the cranial image samples to be processed one by one into the encoder for a predetermined number of downsampling and feature extractions to obtain a target feature image; inputting the target feature image into the decoder for a predetermined number of upsamplings and then performing feature fusion to obtain a segmentation result.
[0007] According to one aspect of the embodiments of this application, a cranial image cropping device is provided. The cranial image cropping device includes: an image acquisition module for acquiring a cranial image to be processed; an image segmentation module for inputting the cranial image to be processed into a cranial image segmentation model to obtain a segmentation result, the segmentation result including a cranial mask, wherein the non-basal surface of the cranial mask is a brain parenchyma region, and the basal surface is a regular region containing all target blood vessels; and an image cropping module for cropping the cranial image to be processed based on the segmentation result to obtain a cropping result; wherein the cranial image segmentation model includes an encoder and a decoder; and the image segmentation module specifically includes: an encoder submodule for inputting the cranial image to be processed into the encoder for downsampling and feature extraction a predetermined number of times to obtain a target feature image; and a decoder submodule for inputting the target feature image into the decoder for upsampling a predetermined number of times and then performing feature fusion to obtain a segmentation result.
[0008] According to one aspect of the embodiments of this application, a training device for a cranial image segmentation model is provided. The cranial image cropping device includes: a sample generation module, used to acquire multiple cranial images, and label each cranial image to form cranial image samples. Each cranial image sample is labeled with a corresponding segmentation label in the form of a mask. In the segmentation label, the non-basal surface is the brain parenchyma region, and the basal surface is a regular region containing all target blood vessels. A sample input module is used to input the cranial image samples one by one into the cranial image segmentation model to obtain segmentation results. A parameter update module is used to update the parameters of the cranial image segmentation model according to the output segmentation results and the segmentation labels until a predetermined termination condition is reached, thereby ending the training and obtaining a trained cranial image segmentation model. The cranial image segmentation model includes an encoder and a decoder. The sample input module specifically includes: a downsampling submodule, used to input the cranial image samples to be processed one by one into the encoder for a predetermined number of downsampling and feature extractions to obtain a target feature image; and an upsampling submodule, used to input the target feature image into the decoder for a predetermined number of upsamplings and then perform feature fusion to obtain a segmentation result.
[0009] According to one aspect of the embodiments of this application, a computer program product is provided, including one or more computer programs that, when executed by one or more processors, implement the method described in the above embodiments.
[0010] According to one aspect of the embodiments of this application, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the above embodiments.
[0011] According to one aspect of the embodiments of this application, an electronic device is provided, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to perform the method described in the above embodiments.
[0012] In some embodiments of this application, the technical solutions provided employ a deep learning segmentation model with innovative structure trained by a specific data strategy, which enables rapid, accurate, and automated cropping of the skull and cerebral cortex regions. This significantly improves the efficiency and quality of subsequent maximum intensity projection (MIP) image generation, reduces cropping time from several minutes to seconds, and ensures segmentation accuracy across different scanning ranges.
[0013] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings: Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown.
[0015] Figure 2 A flowchart illustrating a brain image cropping method provided in an embodiment of this application is shown.
[0016] Figure 3 It shows that according to Figure 2 A flowchart illustrating a specific implementation of step S200 in the cranial image cropping method shown in the corresponding embodiment.
[0017] Figure 4 It shows that according to Figure 3 A flowchart illustrating a specific implementation of step S210 in the cranial image cropping method shown in the corresponding embodiment.
[0018] Figure 5 The diagram shows a flowchart illustrating a training method for a cranial image segmentation model provided in an embodiment of this application.
[0019] Figure 6 A schematic diagram of the structure of a cranial image cropping device provided in an embodiment of this application is shown.
[0020] Figure 7 This illustration shows a schematic diagram of the structure of a training device for a cranial image segmentation model provided in an embodiment of this application.
[0021] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0022] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.
[0023] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0024] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0025] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0026] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown.
[0027] like Figure 1 As shown, the system architecture may include terminal devices (such as...) Figure 1 The device shown includes one or more of a smartphone 101, tablet 102, and portable computer 103 (which could also be a desktop computer, etc.), a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal device and the server 105. The network 104 can include various connection types, such as wired communication links, wireless communication links, etc.
[0028] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, there can be any number of terminal devices, networks, and servers. For example, server 105 could be a server cluster composed of multiple servers.
[0029] Users can interact with server 105 via network 104 using terminal devices to receive or send messages, etc. Server 105 can be a server providing various services. For example, a user uploads a brain image to be processed to server 105 using terminal device 103 (or terminal device 101 or 102). Server 105 can input the brain image to be processed into a brain image segmentation model to obtain a segmentation result, which includes a brain mask. Based on the segmentation result, the brain image to be processed is cropped to obtain a cropped result. The brain image segmentation model includes an encoder and a decoder. The process of inputting the brain image to be processed into the brain image segmentation model to obtain a segmentation result specifically includes: inputting the brain image to be processed into the encoder for a predetermined number of downsampling and feature extractions to obtain a target feature image; inputting the target feature image into the decoder for a predetermined number of upsamplings and then performing feature fusion to obtain a segmentation result.
[0030] It should be noted that the cranial image cropping method provided in this application embodiment is generally executed by server 105, and correspondingly, the cranial image cropping device is generally disposed in server 105. However, in other embodiments of this application, the terminal device may also have similar functions to the server, thereby executing the cranial image cropping scheme provided in this application embodiment.
[0031] The implementation details of the technical solutions in the embodiments of this application are described in detail below: Figure 2 A flowchart of a cranial image cropping method according to an embodiment of this application is shown. This cranial image cropping method can be executed by a server, which may be... Figure 1 The server shown. (Refer to...) Figure 2 As shown, this brain image cropping method includes at least the following: S100, acquire the brain image to be processed.
[0032] S200, the brain image to be processed is input into the brain image segmentation model to obtain the segmentation result. The segmentation result includes a brain mask, in which the non-basal surface of the brain mask is the brain parenchyma region, and the basal surface is a regular region containing all target blood vessels.
[0033] S300, based on the segmentation result, the brain image to be processed is cropped to obtain the cropping result.
[0034] In the embodiments of this application, a structurally innovative deep learning segmentation model trained with a specific data strategy is employed to achieve rapid, accurate, and automated cropping of the skull and cerebral cortex regions. This significantly improves the efficiency and quality of subsequent maximum intensity projection image generation, reducing cropping time from several minutes to seconds, while ensuring segmentation accuracy across different scanning ranges. Simultaneously, by introducing a regular region mask on the basal surface of the skull, the segmentation stage distinguishes between non-basal and basal surfaces using a cranial mask. Furthermore, a mask design containing regular regions of all target blood vessels is introduced on the basal surface, effectively preventing vascular edges from being truncated during segmentation. This also ensures the complete preservation of blood vessels even in complex skull base structures, guaranteeing the overall appearance of the vascular structure.
[0035] In S100, the cranial image is three-dimensional volume data obtained by 3D-TOF magnetic resonance scanning, and the data includes signals of tissues such as brain parenchyma, blood vessels, skull, and cortex.
[0036] In practice, a magnetic resonance imaging (MRI) device can be used to scan the head or head and neck region to obtain three-dimensional image data in DICOM format or voxel matrix form. Before image input, preprocessing operations can be performed, including spatial resampling, grayscale normalization, and noise suppression. Spatial resampling ensures that all input data has a uniform voxel resolution; grayscale normalization adjusts pixel values from different devices and scanning conditions to a uniform intensity range; noise suppression can be achieved through Gaussian filtering or nonlocal mean filtering to reduce high-frequency noise interference. The preprocessed cranial image is then used as the image to be processed and input into the subsequent segmentation model for automatic processing.
[0037] In S200, the cranial image segmentation model is based on a deep convolutional neural network structure, which consists of an encoder and a decoder.
[0038] The encoder performs layer-by-layer downsampling and feature extraction on the input cranial image to extract multi-scale spatial and semantic features. The decoder performs layer-by-layer upsampling and reconstruction on the high-dimensional features extracted by the encoder, restoring the image to the same spatial dimensions as the input image and generating pixel-level segmentation results. This model achieves automatic differentiation between brain parenchyma and non-brain parenchyma regions (including the skull and cerebral cortex) through deep feature learning from 3D-TOF images. The output segmentation results include a cranial mask, where the target region of the mask corresponds to the brain parenchyma and vascular regions. For non-basal skull surfaces, the mask region is the brain parenchyma region; for basal skull surfaces, the mask region is a regular region containing all target vessels. This regular region is generated through boundary completion and morphological constraints to ensure the complete preservation of vascular regions in the complex anatomical environment of the skull base.
[0039] Specifically, in some embodiments, the specific implementation of step S200 can be found in [reference needed]. Figure 3 . Figure 3 It is based on Figure 2 A detailed description of step S200 in the cranial image cropping method shown in the corresponding embodiment: In the cranial image cropping method, the cranial image segmentation model includes an encoder and a decoder, and step S200 may include the following steps: S210, the brain image to be processed is input into the encoder for a predetermined number of downsampling and feature extraction to obtain the target feature image.
[0040] S220, the target feature image is input into the decoder and upsampled a predetermined number of times, and then feature fusion is performed to obtain the segmentation result.
[0041] This embodiment achieves automatic identification and removal of the skull and cerebral cortex in 3D-TOF images without manual intervention. It has high segmentation accuracy, is applicable to different scanning ranges and different equipment conditions, and the cropping process is efficient and fast, with an overall processing time of less than 2 seconds. It provides accurate input data for subsequent MIP generation and cerebral vascular visualization, thereby improving the efficiency of clinical diagnosis.
[0042] In S210, the encoder internally incorporates multiple feature extraction units, each containing a convolutional layer, a batch normalization layer, and an activation layer. The convolutional layer extracts local spatial features of the image, the batch normalization layer stabilizes the network training process, and the activation layer introduces non-linear feature representation capabilities. During execution, the encoder performs downsampling operations sequentially a predetermined number of times. Downsampling can be achieved through convolution or pooling operations with a stride of 2, gradually reducing the spatial size of the feature map and gradually increasing the number of channels, thereby obtaining higher-level abstract semantic information.
[0043] For example, the first downsampling mainly captures the texture and boundary features of the image, while the second and subsequent downsampling gradually extract the structural differences between the brain parenchyma and the skull.
[0044] By combining multi-level convolution extraction with downsampling, a target feature image with multi-scale features can be formed. The target feature image retains the main structural information of brain tissue and suppresses noise and redundant information.
[0045] Ultimately, the target feature image output by the encoder serves as the input to the decoder, providing basic feature support for subsequent upsampling and feature fusion.
[0046] Specifically, in some embodiments, the specific implementation of step S210 can be found in [reference needed]. Figure 4 . Figure 4 It is based on Figure 3Detailed description of step S210 in the cranial image cropping method shown in the corresponding embodiment: In the cranial image cropping method, the encoder includes a first encoder and a second encoder. The first encoder contains a first convolutional block with a number corresponding to the predetermined number of times. Step S210 may include the following steps: S212, the brain image to be processed is input into the first encoder, and each of the convolutional blocks performs downsampling and feature extraction on the brain image to be processed in sequence to obtain a preliminary feature image.
[0047] S214, the preliminary feature image is input into the second encoder for multi-head self-attention processing to obtain the target feature image.
[0048] In this embodiment, the encoder includes a first encoder and a second encoder. The first encoder is used to extract multi-level local spatial features from the brain image to be processed; the second encoder is used to calculate global dependencies based on the preliminary feature image through a multi-head self-attention mechanism, and further generate a target feature image with global contextual information. By setting up two-level modules, the first encoder and the second encoder, the comprehensive extraction of local detail features and global contextual features is realized, thereby improving the accuracy and completeness of brain parenchyma region identification.
[0049] The multi-head self-attention mechanism in this embodiment can capture long-distance feature associations across the entire image range, exhibiting higher recognition accuracy, especially at the junction of brain parenchyma and skull. The multi-layer convolutional structure of the first encoder can preserve rich edge information and texture features. By introducing a dual-encoding structure, the model can achieve stable segmentation performance under different scanning ranges and different MRI equipment conditions. Compared with traditional single-convolutional network structures, this method significantly improves segmentation accuracy in complex regions such as the basal surface of the skull, effectively reducing false positives and false negatives.
[0050] In S212, the first encoder includes multiple first convolutional blocks corresponding to a predetermined number of iterations. Each first convolutional block consists of a downsampling unit and a feature extraction unit. The downsampling unit is used to reduce the spatial resolution of the feature map through stride convolution or pooling operations to expand the receptive field. The feature extraction unit includes multiple convolutional layers, batch normalization layers, and activation layers to extract texture features, shape edges, and tissue structure information from the input image.
[0051] In practice, the first convolutional block typically consists of four layers, corresponding to four downsampling operations. The first convolutional block is used to extract shallow edge and density variation features, while the second to fourth convolutional blocks progressively extract deep semantic features, including the intensity distribution differences between the brain parenchyma and the skull.
[0052] After processing by each level of convolutional blocks, the spatial size of the image is reduced layer by layer, while the number of channels is increased layer by layer, forming a comprehensive multi-scale feature representation. The output feature maps are then concatenated or residually connected to form the output of the first encoder, i.e., the preliminary feature image.
[0053] The initial feature images contain boundary information and local detail features of brain parenchyma regions, but due to the limitations of the local receptive field of the convolution kernel, they still lack global structural dependency information. Therefore, further global feature enhancement is required through a second encoder.
[0054] Specifically, in some embodiments, the specific implementation of step S212 can be found in the following embodiments. This embodiment is based on... Figure 4 According to the detailed description of step S212 in the cranial image cropping method shown in the corresponding embodiment, the first encoder in the cranial image cropping method includes: a 2D spatial branch, a 3D sequence branch, and a cross-context fusion module. The 2D spatial branch includes a first convolutional block. Step S212 may include the following steps: The brain image to be processed is sliced and then input into the 2D spatial branch. Each of the first convolutional blocks sequentially downsamples and extracts features from each slice of the brain image to be processed, thereby obtaining the slice feature image corresponding to each slice.
[0055] The brain image to be processed is input into the 3D sequence branch, and the contextual features of the brain image sequence are extracted to obtain a sequence feature image.
[0056] The cross-context fusion module uses one of the slice feature images and the sequence feature images as a query and the other as a key and value for cross-attention processing to obtain a preliminary feature image.
[0057] In this embodiment, the first encoder includes a 2D spatial branch, a 3D sequence branch, and a cross-context fusion module. The 2D spatial branch extracts the planar spatial structural features of each slice, the 3D sequence branch extracts the contextual relationships between slices, and the cross-context fusion module enables information interaction and weighted fusion between the two types of features, thereby obtaining a preliminary feature image that combines local details and global context. The 2D branch captures texture details within slice layers, while the 3D branch captures spatial continuity between layers, enabling multi-scale feature complementarity. Furthermore, the cross-context fusion module performs cross-attention calculations, effectively integrating local and global information and improving the model's recognition accuracy for complex areas of the skull base. Because the fused features include slice edge and volume continuity information, the segmentation accuracy of the brain parenchyma and skull boundary is significantly improved; the model performs stably under different scanning ranges (head, head and neck), reducing segmentation interruptions or excessive rejection.
[0058] Specifically, firstly, the input 3D cranial image is sliced along a predetermined axis to obtain multiple consecutive 2D image slices. Each slice represents the signal distribution of the 3D image at a specific level. Each slice is sequentially input into a 2D spatial branch, which consists of multiple first convolutional blocks, each containing a downsampling unit and a feature extraction unit. Through multi-level convolution stacking, each slice can extract multi-level spatial texture features and boundary features. For example, shallow convolution captures the gray-level gradient changes of blood vessel edges and brain parenchyma, while deep convolution identifies signal differences between the cortex and the skull. Finally, each slice generates a corresponding slice feature image, and all slice feature images are sequentially arranged to form a slice feature image sequence, fully representing the detailed structural information of the 3D image in the planar direction.
[0059] Simultaneously, unsliced, complete 3D cranial images are directly input into the 3D sequence branch. The 3D sequence branch employs 3D convolution operations, with the convolution kernel sliding along the x, y, and z dimensions to capture structural continuity and tissue relationships between adjacent layers. Through multi-layer 3D convolution, the model can identify the course of cerebral blood vessels across different layers and their spatial extension characteristics in the skull base region. The output of the 3D sequence branch is a sequence feature image containing rich inter-layer contextual semantic information, reflecting the overall spatial relationships of different anatomical structures in the brain.
[0060] It should be noted that 3D sequence branches and 2D spatial branches can be executed synchronously or asynchronously, and this application does not impose any restrictions on this.
[0061] The cross-context fusion module models the interdependencies between two types of features based on a cross-attention mechanism. Specifically, in this module, either the slice feature image or the sequence feature image is designated as the query, and the other is designated as the key and value. The module calculates the relevance weights between the query and the key, and then performs a weighted summation of the values, thereby achieving selective information transfer.
[0062] When slice features are used as queries, the model can enhance the response of slice features to distant structures based on the global context information provided by sequence features; when sequence features are used as queries, the model can refine the global feature representation by utilizing the local detail information of slice features.
[0063] Through this interaction mechanism, the model simultaneously preserves planar details and 3D structural information, forming a fused preliminary feature image. This preliminary feature image combines high-resolution 2D features with 3D contextual features, providing high-quality input for subsequent processing by the second encoder and decoder.
[0064] It should be noted that in practical engineering implementation, the 2D spatial branch and the 3D sequence branch can run in parallel to reduce computation time. The parameters of the cross-attention mechanism can be automatically learned through end-to-end training, without the need for manual setting of weight ratios. The preliminary feature image output by the first encoder is directly passed to the second encoder, maintaining the continuity of the overall structure and ensuring that the model can complete the entire feature extraction process in a single forward propagation.
[0065] In S214, the second encoder includes a multi-head self-attention computation module for calculating the correlation between each location in the feature map and other locations. This mechanism allows the model to simultaneously focus on long-distance structural associations during segmentation, thereby improving the ability to identify blurred boundaries or fine structural regions that are difficult to recognize by traditional convolutional structures.
[0066] The core computational process of the multi-head self-attention mechanism is as follows: the initial feature image is flattened into a feature sequence, with each pixel corresponding to a feature vector; this feature sequence is mapped to query (Q), key (K), and value (V) vectors respectively; the similarity matrix between Q and K is calculated for each attention head, and a weighted output is formed by weighting the Q and K vectors; the outputs of multiple attention heads are concatenated and then merged into a weighted feature sequence through a linear transformation. After multi-head self-attention processing, the output features can integrate global information from different spatial locations, thus forming a feature representation with contextual understanding capabilities.
[0067] The output of the second encoder is the target feature image, which, while preserving the original local texture details, comprehensively expresses the global structural information. This feature image will be input into the decoder to generate the segmentation result.
[0068] Specifically, in some embodiments, the specific implementation of step S214 can be found in the following embodiments. This embodiment is based on... Figure 4 According to the detailed description of step S214 in the cranial image cropping method shown in the corresponding embodiment, in the cranial image cropping method, the second encoder includes: a multi-head self-attention layer, a normalization layer, and a feedforward network. Step S214 may include the following steps: The preliminary feature image is flattened into a sequence of feature vectors, where each element of the feature vector sequence corresponds to the feature vector of a pixel in the preliminary feature image.
[0069] The feature vector sequence is fed into the multi-head self-attention layer to compute multiple attention heads in parallel, resulting in a weighted feature sequence, which is a feature sequence weighted by global information.
[0070] The weighted feature sequence and the feature vector sequence are added together through the normalization layer via residual connection, and then layer normalization is performed to obtain the first normalization result.
[0071] The first normalized result is fed into the feedforward network for feature transformation, and then input into the normalization layer for residual connection addition and layer normalization to obtain the second normalized result.
[0072] The second normalization result is reshaped to obtain the target feature image.
[0073] In this embodiment, the second encoder includes a multi-head self-attention layer, a normalization layer, and a feedforward network. The multi-head self-attention layer calculates global dependencies between pixels in the feature space; the normalization layer standardizes the feature distribution after residual connections to prevent gradient instability; and the feedforward network performs nonlinear feature transformations after global modeling, enhancing the model's ability to express complex structural features. Thus, the multi-head self-attention layer in this embodiment can simultaneously capture structural relationships at different scales from multiple feature subspaces, giving the model global perception capabilities; residual connections and layer normalization prevent feature degradation and gradient explosion, ensuring efficient model convergence; and the feedforward network enhances feature transformation capabilities, giving the output features stronger semantic discriminative power. Through global dependency modeling, the model can achieve accurate segmentation in regions with similar gray levels, such as the junction of brain parenchyma and skull. This structure maintains high stability and strong adaptability under different data sources and scanning conditions.
[0074] Specifically, the initial input feature image is reorganized into a one-dimensional feature sequence. In this step, each pixel position in the initial feature image corresponds to a feature vector, and the dimension of the feature vector is equal to the number of features of that pixel in the channel direction. The purpose of the flattening operation is to transform the two-dimensional (or three-dimensional) feature map into a sequence form so that the correlation between features at different positions can be calculated in the subsequent attention layer. The flattened feature sequence constitutes the input of the second encoder.
[0075] The role of the multi-head self-attention layer is to simultaneously compute the dependencies between input features from multiple subspaces. Each attention head performs the following computation process: mapping the input feature sequence into query, key, and value vectors respectively; calculating the relevance weights between the query and the key to represent the global dependency between each feature; and performing a weighted summation of the value vectors according to the relevance weights to obtain the weighted feature output.
[0076] Multiple attention heads perform the above calculations in parallel. Each attention head captures the structural relationships of the image from different feature dimensions. Finally, the outputs of each attention head are concatenated and subjected to a linear transformation to generate a weighted feature sequence. This weighted feature sequence contains the dependency information of the initial feature image in a global scope and is a feature representation after global weighting.
[0077] The weighted feature sequence is residually summed with the initial input feature vector sequence. This residual connection maintains the continuity of input information and prevents feature degradation caused by deep structures. The summed residuals are then fed into a normalization layer for layer-level normalization, outputting the first normalized result. This normalization operation improves the stability and convergence speed of model training by standardizing the feature distribution. The first normalized result is then fed into a feedforward network. The feedforward network typically consists of two linear transformation layers and one non-linear activation layer, used to perform non-linear mapping and dimensionality transformation on the features along the channel dimension to enhance the model's expressive power. The output of the feedforward network is summed a second time with its input, and then fed back into the normalization layer for layer-level normalization, yielding the second normalized result. Through continuous residual and normalization operations, gradient vanishing is prevented and the model's feature stability is maintained.
[0078] The second normalization result is reshaped according to the spatial dimensions of the initial feature image, restoring it to a two-dimensional (or three-dimensional) spatial structure. The feature image after this reconstruction is the target feature image. This image remains consistent with the initial feature image in the spatial dimension, but contains global information and optimized feature representation in the semantic dimension.
[0079] It should be noted that the second encoder can be directly cascaded with the first encoder in the main segmentation model. During model training, the multi-head self-attention layer, normalization layer, and feedforward network all participate in gradient updates, and the entire model is trained end-to-end. During inference, after inputting the initial feature image, only one forward propagation is needed to output the target feature image, providing high-quality feature input for the subsequent decoder.
[0080] In practical implementation, the first encoder and the second encoder can be implemented as consecutive modules in the same neural network structure. The feature tensor output by the first encoder is directly input into the second encoder without manual intervention. During the model training phase, the dual-encoder structure can be jointly optimized through end-to-end training; during the inference phase, the encoder as a whole can obtain the target feature image by performing a single forward propagation, without the need for step-by-step operations.
[0081] In S220, the decoder consists of multiple upsampling units, each including an upsampling layer, a convolutional layer, a batch normalization layer, and an activation layer. The upsampling layer can progressively restore spatial dimensions through transposed convolution or interpolation operations. During each upsampling process, the decoder fuses the feature map of the current layer with the feature map of the corresponding layer in the encoder. Feature fusion can be achieved through concatenation or weighted summation, aiming to simultaneously preserve high-level semantic information and low-level spatial details, thereby improving the accuracy of segmentation boundaries. After the last upsampling layer, the model performs a convolution operation on the fused feature map, generating a segmentation result with the same dimensions as the original input image. The segmentation result is a multi-channel image, with one channel representing the probability map of the brain parenchyma and vascular regions, and another channel corresponding to the background region. Thresholding can convert the probability map into a binary mask, obtaining a cranial mask.
[0082] Specifically, in some embodiments, the specific implementation of step S220 can be found in the following embodiments. This embodiment is based on... Figure 4 A detailed description of step S220 in the cranial image cropping method shown in the corresponding embodiment: In the cranial image cropping method, the decoder includes a gated attention network, a feature fusion layer, and a second convolutional block with a number corresponding to the predetermined number of times. Step S220 may include the following steps: After the target feature image is upsampled sequentially by each of the second convolutional blocks, it is filtered and weighted by the gated attention network to obtain the attention coefficient map.
[0083] The attention coefficient map and each of the preliminary feature images are fused in the feature fusion layer to obtain the fused feature result.
[0084] The fusion feature results are post-processed to obtain the segmentation results.
[0085] In this embodiment, the decoder includes a gated attention network, a feature fusion layer, and a number of second convolutional blocks corresponding to the predetermined number of times. The number of second convolutional blocks corresponds to the number of downsampling operations of the encoder, used for layer-by-layer upsampling to restore spatial resolution. The gated attention network is used to filter and weight features after upsampling. The feature fusion layer fuses the weighted upsampled features with the preliminary feature image output by the encoder. The gated attention network can automatically determine important feature regions based on contextual information, significantly improving the accuracy of brain parenchyma segmentation. The feature fusion layer achieves semantic and detail complementarity of multi-layer features, resulting in better preservation of boundary continuity and local details. Through the attention filtering mechanism, spurious responses from the skull and cortex are effectively suppressed, reducing missegmentation. The segmentation results remain stable under different scanning parameters and different patient data. Combined with the gated mechanism and multi-scale fusion, the segmentation accuracy, speed, and generalization ability are significantly superior to traditional decoding structures.
[0086] This embodiment introduces a gated attention network to adaptively weight the upsampled feature map and combines it with multi-source feature integration of the feature fusion layer to achieve selective enhancement of key structural regions, thereby improving the segmentation accuracy of brain parenchyma and vascular regions.
[0087] Specifically, each second convolutional block within the decoder sequentially performs upsampling operations. Each convolutional block includes an upsampling layer, a convolutional layer, a batch normalization layer, and an activation layer. The upsampling layer gradually restores the spatial size of the feature map to the original input image size through transposed convolution or interpolation reconstruction. The convolutional layer is used to extract structural details from the upsampled features. The normalization layer stabilizes the feature distribution, and the activation layer introduces non-linear expressive power. During the multi-level upsampling process, the decoder can gradually recover the spatial details compressed by the encoder, achieving a smooth mapping from high semantic features to low semantic space.
[0088] After each upsampling operation, the generated feature map is fed into the gated attention network.
[0089] The role of this network is to adaptively calculate the importance of each position and channel in the feature map in order to select feature regions that contribute to the segmentation task.
[0090] The gated attention network consists of a channel attention submodule and a spatial attention submodule. The channel attention submodule is used to calculate the global weights of each feature channel, while the spatial attention submodule is used to determine the positional response of key regions in the spatial dimension.
[0091] The network first calculates the average and maximum responses of the input features along the channel dimension, then generates channel weights through fully connected layers and activation functions. Subsequently, combining the channel weighting results, it calculates an attention map along the spatial dimension, obtaining the weight coefficients for each spatial location. The final output is an attention coefficient map, where each coefficient represents the importance of the feature at that location. By performing a dot product operation between the upsampled feature map and the attention coefficient map, the network achieves weighted filtering of features, suppressing irrelevant feature responses and highlighting features from key regions such as brain parenchyma and blood vessels.
[0092] The weighted feature map, processed by the gated attention network, is input to the feature fusion layer. The function of the feature fusion layer is to integrate the upsampled features with the preliminary feature image of the corresponding layer. Fusion methods can include concatenation fusion and weighted fusion. Concatenation fusion concatenates the weighted features with the preliminary feature image along the channel dimension; weighted fusion performs a weighted sum based on attention coefficients to balance the importance of different feature sources. The fused result simultaneously contains the semantic information of deep features and the spatial boundary information of shallow features, thus achieving feature complementarity at both the detail and semantic levels. This fusion mechanism ensures that in structurally complex regions such as the skull base, the model can maintain the continuity of blood vessels while avoiding misidentification of the skull as brain tissue.
[0093] The fused feature results are processed through convolutional smoothing and activation function transformation to generate the final segmentation result. Convolutional smoothing removes spurious response noise during the fusion process, while the activation function limits the output to a predetermined range for subsequent generation of a binary mask. The final output segmentation result is a cranial mask, where pixel values of the brain parenchyma and vascular regions are labeled as foreground, and the skull and cerebral cortex regions are labeled as background. This mask is directly used for cropping the original image, achieving automatic extraction and interference-free display of brain tissue.
[0094] It should be noted that in practical engineering implementation, the decoder can adopt a hierarchical symmetrical structure, forming a one-to-one correspondence with the encoder. A gated attention network can be embedded in each upsampling stage to ensure that features from different layers are weighted and filtered; the feature fusion layer fuses the output features of the encoder at each corresponding stage. The entire decoding process is completed automatically in a single forward propagation, requiring no manual intervention.
[0095] In S300, pixel-level cropping is performed on the original brain image to be processed. The cropping operation includes the following steps: performing logical operations on the mask in the segmentation result and the original image according to their positional correspondence; retaining pixels in the mask marked as brain parenchyma and blood vessels; and removing or setting to zero pixels in the mask marked as non-brain parenchyma regions (corresponding to skull, cortex, and background signals).
[0096] The cropped image retains only the brain parenchyma and vascular regions, effectively eliminating interference from the skull and cortex on subsequent vascular visualization. The final cropped result can be directly used to generate MIP images, providing doctors with clear visualizations of cerebral blood vessels.
[0097] Compared to traditional manual cropping methods, the embodiments of this application can complete the entire automatic cropping process in approximately 2 seconds, significantly shortening operation time and reducing human error. Verification has shown that this method maintains high segmentation accuracy and stability across different scanning ranges of the head and head and neck.
[0098] Here, this application also proposes a training method for a cranial image segmentation model, for training the cranial image segmentation model described above. Figure 5 A flowchart illustrating a training method for a cranial image segmentation model according to an embodiment of this application is shown. This training method can also be executed by a server, which can also be... Figure 1 The server shown. (Refer to...) Figure 5 As shown, this brain image cropping method includes at least the following: S600: Acquire multiple cranial images, and label each cranial image to form cranial image samples. Each cranial image sample is labeled with a corresponding mask-like segmentation label. In the segmentation label, the non-basal surface is the brain parenchyma region, and the basal surface is a regular region containing all target blood vessels.
[0099] S700, the cranial image samples are input one by one into the cranial image segmentation model to obtain the segmentation results.
[0100] S800: Based on the output segmentation results and the segmentation labels, update the parameters of the cranial image segmentation model until a predetermined termination condition is met, end the training, and obtain the trained cranial image segmentation model.
[0101] This embodiment provides an efficient training method to complement the aforementioned cranial image segmentation model. This training method aims to fully utilize limited and imbalanced medical image data to train a robust and accurate segmentation model. This embodiment aims to establish a segmentation model with multi-scale feature extraction and high generalization ability through a systematic sample construction, network training, and parameter update process, achieving automatic identification and accurate segmentation of brain tissue structures.
[0102] In the S600, multiple 3D-TOF cranial images are acquired from magnetic resonance imaging (MRI) devices, covering different scan ranges (head and head-neck) and imaging parameters from different devices. Each cranial image is manually annotated to form cranial image samples. Annotation is performed using medical image segmentation tools such as 3Dslicer and ITK-SNAP. The annotation method uses the brain parenchyma and vascular regions as target areas, generating a mask image of the same size as the input image, with each pixel value corresponding to either the foreground or background. During annotation, the boundary of the basal surface of the skull is defined by the inner edge of the skull plate as the posterior boundary and the posterior wall of the maxillary sinus and the posteromedial border of the temporalis muscle as the anterior boundary, ensuring that the annotated area completely covers the brain parenchyma and major vascular structures. Finally, each cranial image sample includes the original image and a mask label pair for subsequent model training.
[0103] Specifically, in some embodiments, the specific implementation of step S600 can be found in the following embodiments. This embodiment is based on... Figure 5According to the detailed description of step S600 in the cranial image cropping method shown in the corresponding embodiment, step S600 in the cranial image cropping method may include the following steps: Multiple cranial images are acquired, and each cranial image is labeled to form an original image sample.
[0104] The original image sample is augmented to obtain an augmented image sample.
[0105] The original image samples and the enhanced image samples are normalized to obtain cranial image samples.
[0106] In this embodiment, by introducing data augmentation and intensity normalization mechanisms during the sample construction stage, the number of samples is increased while ensuring the consistency of data distribution, thereby improving the robustness and adaptability of model training.
[0107] The normalization operation in this embodiment effectively eliminates grayscale differences between different scanning parameters, and the data enhancement increases the sample size several times, preventing overfitting. Through translation, rotation, and flipping operations, the model can learn brain features from different angles. Noise perturbation of the samples ensures that the model maintains high accuracy in actual clinical noise environments. Optimization of sample quality and balance makes the network training process easier to converge, and the loss function decreases smoothly.
[0108] Specifically, multiple 3D-TOF cranial images were acquired from a magnetic resonance imaging (MRI) system. The data could include combined head and neck scans to ensure sample diversity and structural coverage. Each 3D-TOF cranial image was manually annotated with mask labels for the brain parenchyma and vascular regions. Annotation was performed using medical image processing software, adhering to the principle of complete coverage of the vascular region within the brain parenchyma, with the posterior boundary of the basal surface of the skull base defined by the inner edge of the skull plate, and the anterior boundary defined by the posterior wall of the maxillary sinus and the posteromedial border of the temporalis muscle. Simultaneously, the boundaries were to be continuous and smooth, avoiding the inclusion of non-target tissues.
[0109] After annotation, the original images are paired with their corresponding mask labels and stored to form the original image sample set. Each sample contains two parts: the input image and the segmentation label, which are used for model training.
[0110] Without altering the semantic content of the image, spatial and grayscale transformations are performed on the original image samples. These transformations can include translation, rotation, flipping, scaling, and noise perturbation.
[0111] Among them, translation transformation randomly moves the image position in the three coordinate axes to simulate patient positional deviation; rotation transformation rotates the image in three dimensions at random angles to enhance the model's adaptability to changes in direction; flip transformation mirrors the image along any axis to expand the sample space; scaling transformation enlarges or reduces the image to simulate structural changes under different imaging ranges; and noise perturbation adds Gaussian or Poisson noise to the image to improve the model's resistance to imaging noise.
[0112] The enhancement operation is performed simultaneously on the original image and its corresponding mask label to maintain pixel-level correspondence. The number of enhanced image samples can be flexibly set according to training requirements, for example, generating 5 to 10 enhanced samples for each original sample. The enhanced samples constitute the enhanced image sample set, which exhibits high diversity in spatial morphology, grayscale distribution, and noise characteristics.
[0113] To eliminate training bias caused by inconsistent pixel value ranges under different imaging conditions, grayscale normalization is performed on all samples. The normalization operation is calculated as follows:
[0114] in, The original image in three-dimensional coordinates Pixel value at; and These are the minimum and maximum pixel values in the image, respectively. These are the normalized pixel values.
[0115] After normalization, the pixel value range of all images is unified to the [0,1] interval, ensuring a consistent model input distribution. The normalized original image samples and the enhanced image samples together constitute the final training sample set, namely, the brain image samples. This sample set has a sufficient number of samples to meet the training requirements of deep networks, a uniform grayscale distribution to avoid bias caused by specific samples in network training, and diverse structural features to improve the robustness and generalization performance of the model.
[0116] It should be noted that sample augmentation and normalization can be performed automatically during the data preprocessing stage, specifically implemented by the image processing script or the data loading module in the model training framework. During training, each iteration randomly draws batches of input from the original and augmented samples to ensure balanced sample usage. Samples and mask labels are stored in voxel format and can be directly input into the 3D convolutional network structure for training.
[0117] In some embodiments of this application, prior to S700, the training method further includes: The brain image samples are classified to obtain brain image samples corresponding to the skull base region.
[0118] The brain image samples corresponding to the skull base region are processed to improve their sampling rate.
[0119] In this embodiment, the frequency of skull base samples participating in the training process is increased, enabling the model to fully learn the surface features of the skull base and reduce missed and misclassified blood vessels. After adjusting the sampling rate, the contribution of samples from different regions during training is more balanced, avoiding model bias towards easily segmented areas. By focusing training on difficult-to-segment regions, the model maintains stable performance under different patients and different scanning ranges. Because the information of the skull base region is fully learned, the overall convergence speed of the model is accelerated, and the loss function decreases more stably. The trained model has a higher resistance to noise and gray-level unevenness.
[0120] In some embodiments of this application, class-weighted sampling can be used to increase the sampling rate of the skull base region.
[0121] Specifically, the first step is to classify all training image slices. This can be broadly categorized into two types: basal surface slices and non-basal surface slices. This is typically done during data preparation and can be done by labeling the images based on their Z-axis position in a 3D sequence or anatomical features.
[0122] Since the basal surface is a minority class, its sampling rate needs to be increased, so it can be given a higher weight (e.g., a weight of 5.0); while the non-basal surface is a majority class, so it can be given a lower weight (e.g., a weight of 1.0). In this way, the probability of a basal surface sample being selected during sampling will be 5 times that of a non-basal surface sample.
[0123] During the data loading phase, this list of weights is used to construct a weighted random sampler. This way, each time a training batch is generated, the sampler performs biased sampling from the entire dataset according to the set weights, ensuring that each batch contains a sufficient number of skull basal surface images.
[0124] In some embodiments of this application, oversampling can be used to increase the sampling rate in the skull base region.
[0125] Specifically, similarly, all skull base image slices are first identified. Then, these skull base images and their corresponding masks are directly copied from the dataset. For example, if there are 900 non-skull base images and only 100 skull base images, these 100 skull base images can be copied 8 times, resulting in a dataset containing 900 skull base images (100 original images + 800 copied images). Standard random sampling is then performed on the copied dataset. Since the number of skull base samples has been greatly increased, their probability of being naturally selected also increases.
[0126] To prevent the model from overfitting due to seeing identical duplicate samples, oversampling is usually combined with data augmentation. That is, each time a sample is copied, it is subjected to a random data augmentation operation (such as slight rotation, translation, brightness adjustment, etc.). In this way, although the model sees the same cranial basal surface data each time, it is slightly different at the pixel level, which helps to improve the model's generalization ability.
[0127] In some embodiments of this application, layered sampling can be used to increase the sampling rate in the skull base region.
[0128] Specifically, the dataset is grouped into basal skull planes and non-basal skull planes. Instead of randomly sampling from the entire data pool when constructing each training batch, the number of samples from each category that must be included in each batch is specified. For example, a batch of size 16 can be defined, which must contain 8 images randomly sampled from the basal skull plane group and 8 images randomly sampled from the non-basal skull plane group. Following this rule, samples are drawn from each group and combined to form the final training batch.
[0129] By combining any one or more of the above methods, the goal of increasing the sampling rate of the skull base region can be effectively achieved. The ultimate effect is to give special attention to these difficult and rare samples during the training process, thereby ensuring that the model can also obtain excellent segmentation performance in these key regions.
[0130] In S700, the model extracts features, reconstructs features, and segments the input samples through a forward propagation process, outputting a segmentation result with the same size as the input image.
[0131] Specifically, in some embodiments, the specific implementation of step S700 can be found in the following embodiments. This embodiment is based on... Figure 5 A detailed description of step S700 in the cranial image cropping method shown in the corresponding embodiment: In the cranial image cropping method, the cranial image segmentation model includes an encoder and a decoder, and step S700 may include the following steps: Each brain image sample to be processed is input into the encoder for a predetermined number of downsampling and feature extraction operations to obtain the target feature image.
[0132] The feature image is input into the decoder and upsampled a predetermined number of times before feature fusion is performed to obtain the segmentation result.
[0133] In this embodiment, the brain image samples that have undergone all the above processing are input into the brain image segmentation model in batches. As mentioned above, the image samples first enter the encoder for a predetermined number of downsampling and multi-dimensional feature extraction to obtain the target feature image; then, the target feature image is input into the decoder, and after a predetermined number of upsampling and feature fusion, the model outputs a predicted segmentation result, i.e., a segmentation probability map.
[0134] Specifically, each brain image sample is input into the encoder, which performs a predetermined number of downsampling operations. Each level of the encoder contains convolutional layers, batch normalization layers, and nonlinear activation layers to extract structural features at different scales. Through multi-level feature extraction, the encoder outputs a target feature image, which contains high-dimensional semantic information and global contextual relationships of the brain parenchyma and vascular regions.
[0135] The target feature image is input into the decoder. The decoder restores spatial resolution through progressive upsampling and uses a gated attention network to weight and filter the upsampled features. In the feature fusion layer, the upsampled features are fused with features from each layer of the encoder, preserving local details and enhancing boundary structures. After multiple levels of upsampling, a segmentation result, i.e., a segmentation probability image, is output with the same size as the original image. Each pixel value in the segmentation result represents the probability distribution of that location belonging to the brain parenchyma and vascular regions.
[0136] In S800, after the model forward propagates, the output segmentation result is compared with the corresponding labeled mask to calculate the error.
[0137] In some embodiments, the error metric uses the Dice loss function to measure the degree of overlap between the predicted region and the ground truth labeled region. The loss function is calculated as follows:
[0138] in, For the predicted results, For real masking.
[0139] The gradient is calculated using the backpropagation algorithm, and the model weight parameters are updated using optimizers such as Adam and SGD. The update process is performed iteratively across the entire sample set until a predetermined termination condition is met. The predetermined termination condition may include any one or a combination of the following: the loss function value decreases by less than a preset threshold for several consecutive iterations, the validation set segmentation accuracy reaches a set target, or a preset number of training epochs is reached.
[0140] When the termination condition is met, the training is terminated, and the trained cranial image segmentation model is obtained.
[0141] The following describes an embodiment of the apparatus described in this application, which can be used to execute the cranial image cropping method and the cranial image segmentation model training method described in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the cranial image cropping method and the cranial image segmentation model training method described in the above embodiments of this application.
[0142] Figure 6 A block diagram of a cranial image cropping device according to an embodiment of this application is shown.
[0143] Reference Figure 6 As shown, a cranial image cropping device 600 according to an embodiment of this application includes an image acquisition module 610, an image segmentation module 620, and an image cropping module 630.
[0144] The image acquisition module 610 is used to acquire a brain image to be processed; the image segmentation module 620 is used to input the brain image to be processed into a brain image segmentation model to obtain a segmentation result, wherein the non-basal surface of the brain mask is the brain parenchyma region, the basal surface is a regular region containing all target blood vessels, and the segmentation result includes the brain mask; the image cropping module 630 is used to crop the brain image to be processed based on the segmentation result to obtain a cropping result.
[0145] The above-mentioned brain image segmentation model includes an encoder and a decoder; the image segmentation module 620 specifically includes an encoder submodule 621 and a decoder submodule 622.
[0146] The encoder submodule 621 is used to input the brain image to be processed into the encoder for a predetermined number of downsampling and feature extraction to obtain the target feature image; the decoder submodule 622 is used to input the target feature image into the decoder for a predetermined number of upsampling and feature fusion to obtain the segmentation result.
[0147] The specific details of each module in the above-mentioned cranial image cropping device have been described in detail in the corresponding cranial image cropping method, so they will not be repeated here.
[0148] In one application scenario, a cranial image cropping device can be deployed in, for example... Figure 1 In the system architecture shown.
[0149] During system operation, the image acquisition module 610 receives 3D-TOF cranial images from the scanning device; the image segmentation module 620 performs automatic segmentation and outputs a cranial mask; and the image cropping module 630 automatically crops the skull and cortex based on the mask. This system can automatically segment and crop cranial structures without manual intervention, improving image processing efficiency, reducing the annotation burden on doctors, and providing reliable support for clinical diagnosis and preoperative planning.
[0150] Figure 7 A block diagram of a training apparatus for a cranial image segmentation model according to an embodiment of this application is shown.
[0151] Reference Figure 7 As shown, a training device 700 for a cranial image segmentation model according to an embodiment of this application includes a sample generation module 710, a sample input module 720, and a parameter update module 730.
[0152] The sample generation module 710 acquires multiple cranial images and labels each image to form cranial image samples. Each cranial image sample is labeled with a corresponding mask-like segmentation label. In the segmentation label, the non-basal surface represents the brain parenchyma region, and the basal surface represents a regular region containing all target blood vessels. The sample input module 720 inputs the cranial image samples one by one into the cranial image segmentation model to obtain segmentation results. The parameter update module 730 updates the parameters of the cranial image segmentation model based on the output segmentation results and the segmentation labels until a predetermined termination condition is met, ending the training and obtaining a trained cranial image segmentation model. The above-mentioned brain image segmentation model includes an encoder and a decoder; the sample input module 720 specifically includes a downsampling submodule 721 and an upsampling submodule 722.
[0153] The downsampling submodule 721 is used to input the brain image samples to be processed one by one into the encoder for downsampling and feature extraction a predetermined number of times to obtain the target feature image; the upsampling submodule 722 is used to input the target feature image into the decoder for upsampling a predetermined number of times and then perform feature fusion to obtain the segmentation result.
[0154] The specific details of each module in the training device of the above-mentioned cranial image segmentation model have been described in detail in the training method of the corresponding cranial image segmentation model, so they will not be repeated here.
[0155] In one application scenario, the training device for the cranial image segmentation model can be applied to, for example... Figure 1 In the system architecture described above, a brain image segmentation model is trained. During system operation, the medical image acquisition terminal automatically constructs training samples through the sample generation module 710; the training server calls the sample input module 720 to perform encoding and decoding network calculations, and the parameter update module 730 adjusts the model parameters in real time; after training is completed, the obtained model can be deployed in a clinical image post-processing system to achieve automatic recognition and cropping of the skull and cerebral cortex in 3D-TOF images, thereby reducing manual intervention and improving image processing efficiency.
[0156] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0157] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0158] Figure 8 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.
[0159] It should be noted that, Figure 8 The computer system of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0160] like Figure 8 As shown, the computer system includes a Central Processing Unit (CPU) 1801, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1802 or programs loaded from storage portion 1808 into Random Access Memory (RAM) 1803, such as performing the methods described in the above embodiments. Various programs and data required for system operation are also stored in RAM 1803. The CPU 1801, ROM 1802, and RAM 1803 are interconnected via bus 1804. An Input / Output (I / O) interface 1805 is also connected to bus 1804.
[0161] The following components are connected to I / O interface 1805: an input section 1806 including a keyboard, mouse, etc.; an output section 1807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1808 including a hard disk, etc.; and a communication section 1809 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1809 performs communication processing via a network such as the Internet. A drive 1810 is also connected to I / O interface 1805 as needed. Removable media 1811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1810 as needed so that computer programs read from them can be installed into storage section 1808 as needed.
[0162] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1809, and / or installed from removable medium 1811. When the computer program is executed by central processing unit (CPU) 1801, it performs various functions defined in the system of this application.
[0163] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0164] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0165] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0166] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.
[0167] This application also provides a computer program product storing at least one instruction, which is loaded and executed by the processor as described above. Figures 1-5 The method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 1-5 The specific details of the illustrated embodiments will not be elaborated here.
[0168] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0169] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the methods according to the embodiments of this application.
[0170] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0171] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A method for cropping cranial images, characterized in that, include: Acquire brain images to be processed; The brain image to be processed is input into the brain image segmentation model to obtain the segmentation result. The segmentation result includes a brain mask, in which the non-basal surface of the brain mask is the brain parenchyma region, and the basal surface is a regular region containing all target blood vessels. Based on the segmentation results, the brain image to be processed is cropped to obtain the cropping result; The cranial image segmentation model includes an encoder and a decoder; The step of inputting the brain image to be processed into the brain image segmentation model to obtain the segmentation result specifically includes: The brain image to be processed is input into an encoder for a predetermined number of downsampling and feature extraction operations to obtain a target feature image. The target feature image is input into the decoder and upsampled a predetermined number of times before feature fusion is performed to obtain the segmentation result.
2. The cranial image cropping method as described in claim 1, characterized in that, The encoder includes a first encoder and a second encoder, wherein the first encoder contains a first convolutional block, and the number of the first convolutional blocks corresponds to the predetermined number of times; The step of inputting the brain image to be processed into an encoder for a predetermined number of downsampling and feature extraction operations to obtain a target feature image specifically includes: The brain image to be processed is input into the first encoder, and each of the first convolutional blocks sequentially performs downsampling and feature extraction on the brain image to be processed to obtain a preliminary feature image. The preliminary feature image is input into the second encoder for multi-head self-attention processing to obtain the target feature image.
3. The cranial image cropping method as described in claim 2, characterized in that, The first encoder includes: a 2D spatial branch, a 3D sequence branch, and a cross-context fusion module, wherein the 2D spatial branch contains a first convolutional block; The step of inputting the brain image to be processed into the first encoder, and having each of the first convolutional blocks sequentially downsample and extract features from the brain image to obtain a preliminary feature image, specifically includes: The brain image to be processed is sliced and then input into the 2D spatial branch. Each of the first convolutional blocks performs downsampling and feature extraction on each slice of the brain image to be processed in sequence to obtain the slice feature image corresponding to each slice. The brain image to be processed is input into the 3D sequence branch, and the sequence context features of the brain image to be processed are extracted to obtain a sequence feature image; The cross-context fusion module uses one of the slice feature images and the sequence feature images as a query and the other as a key and value for cross-attention processing to obtain a preliminary feature image.
4. The cranial image cropping method as described in claim 2, characterized in that, The second encoder includes: a multi-head self-attention layer, a normalization layer, and a feedforward network; The step of inputting the preliminary feature image into the second encoder for multi-head self-attention processing to obtain the target feature image specifically includes: The preliminary feature image is flattened into a feature vector sequence, where each element of the feature vector sequence corresponds to the feature vector of a pixel in the preliminary feature image; The feature vector sequence is fed into the multi-head self-attention layer to compute multiple attention heads in parallel, resulting in a weighted feature sequence, which is a feature sequence weighted by global information. The weighted feature sequence and the feature vector sequence are added by residual connection through the normalization layer and then layer normalization is performed to obtain the first normalization result. The first normalized result is fed into the feedforward network for feature transformation, and then input into the normalization layer for residual connection addition and layer normalization to obtain the second normalized result; The second normalization result is reshaped to obtain the target feature image.
5. The cranial image cropping method as described in claim 2, characterized in that, The decoder includes a gated attention network, a feature fusion layer, and a second convolutional block with a number corresponding to the predetermined number of iterations; there are multiple preliminary feature images; The step of inputting the target feature image into the decoder for upsampling a predetermined number of times and then performing feature fusion to obtain the segmentation result specifically includes: After each of the second convolutional blocks sequentially upsamples the target feature image, it is filtered and weighted by the gated attention network to obtain an attention coefficient map. In the feature fusion layer, the attention coefficient map and each of the preliminary feature images are fused to obtain the fused feature result; The fusion feature results are post-processed to obtain the segmentation results.
6. A training method for a cranial image segmentation model, characterized in that, include: Multiple cranial images are acquired, and each cranial image is labeled to form cranial image samples. Each cranial image sample is labeled with a corresponding mask-like segmentation label. In the segmentation label, the non-basal surface is the brain parenchyma region, and the basal surface is a regular region containing all target blood vessels. The cranial image samples are input one by one into the cranial image segmentation model to obtain the segmentation results; Based on the output segmentation results and the segmentation labels, the parameters of the cranial image segmentation model are updated until the predetermined termination condition is met, the training ends, and a trained cranial image segmentation model is obtained. The cranial image segmentation model includes an encoder and a decoder; The step of inputting the cranial image samples one by one into the cranial image segmentation model to obtain the segmentation result specifically includes: The brain image samples to be processed are input one by one into the encoder for downsampling and feature extraction a predetermined number of times to obtain the target feature image; The target feature image is input into the decoder and upsampled a predetermined number of times before feature fusion is performed to obtain the segmentation result.
7. The training method for the cranial image segmentation model as described in claim 6, characterized in that, The process of acquiring multiple cranial images and labeling each cranial image to form a cranial image sample specifically includes: Multiple cranial images are acquired, and each cranial image is labeled to form an original image sample; The original image sample is augmented to obtain an augmented image sample; The original image samples and the enhanced image samples are normalized to obtain cranial image samples.
8. The training method for the cranial image segmentation model as described in claim 6, characterized in that, Before inputting the cranial image samples one by one into the cranial image segmentation model to obtain the segmentation results, the method further includes: The brain image samples are classified to obtain brain image samples corresponding to the skull base region; The brain image samples corresponding to the skull base region are processed to improve their sampling rate.
9. The training method for the cranial image segmentation model as described in claim 8, characterized in that, The process of processing the cranial image samples corresponding to the skull base region to improve their sampling rate specifically includes: Increase the weight of the cranial image samples corresponding to the skull base region to generate a weight list; A weighted random sampler is constructed using the weight list to perform biased sampling from the entire dataset according to the set weights, thereby improving the sampling rate of cranial image samples corresponding to the skull base region.
10. A cranial image cropping device, characterized in that, The cranial image cropping device includes: Image acquisition module, used to acquire images of the brain to be processed; The image segmentation module is used to input the brain image to be processed into the brain image segmentation model to obtain the segmentation result. The segmentation result includes a brain mask, in which the non-basal surface of the brain mask is the brain parenchyma region, and the basal surface is a regular region containing all target blood vessels. The image cropping module is used to crop the brain image to be processed based on the segmentation result to obtain the cropping result; The cranial image segmentation model includes an encoder and a decoder; The image segmentation module specifically includes: The encoder submodule is used to input the brain image to be processed into the encoder for a predetermined number of downsampling and feature extraction to obtain the target feature image; The decoder submodule is used to input the target feature image into the decoder for upsampling a predetermined number of times and then perform feature fusion to obtain the segmentation result.
11. A training device for a cranial image segmentation model, characterized in that, The cranial image cropping device includes: The sample generation module is used to acquire multiple cranial images, label each cranial image to form cranial image samples, and each cranial image sample is labeled with a corresponding mask-like segmentation label. In the segmentation label, the non-basal surface is the brain parenchyma region, and the basal surface is a regular region containing all target blood vessels. The sample input module is used to input the cranial image samples one by one into the cranial image segmentation model to obtain the segmentation result; The parameter update module is used to update the parameters of the cranial image segmentation model according to the output segmentation result and the segmentation label until a predetermined termination condition is reached, thereby ending the training and obtaining a trained cranial image segmentation model. The cranial image segmentation model includes an encoder and a decoder; The sample input module specifically includes: The downsampling submodule is used to input the brain image samples to be processed one by one into the encoder for downsampling and feature extraction a predetermined number of times to obtain the target feature image; The upsampling submodule is used to input the target feature image into the decoder for upsampling a predetermined number of times and then perform feature fusion to obtain the segmentation result.
12. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 9.
13. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 9.
14. A computer program product, characterized in that, It includes one or more computer programs that, when executed by one or more processors, implement the method of any one of claims 1 to 9.