An image segmentation method, device, storage medium and electronic equipment
By using a cascaded bone segmentation method, the area above the shoulders is first cropped out, and then high-precision segmentation and fusion are performed, which solves the problem of imprecise bone tissue segmentation in CTA image data and achieves high-precision and efficient bone tissue segmentation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING WANDONG MEDICAL TECH CO LTD
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-29
Smart Images

Figure CN122115470A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an image segmentation method, apparatus, storage medium and electronic device. Background Technology
[0002] Computed tomography angiography (CTA) is an important medical imaging technique. It is a non-invasive angiography technique that uses computer-aided three-dimensional reconstruction to provide a good understanding of the blood vessels in the body without causing trauma. It has important applications in clinical diagnosis and treatment, especially in the diagnosis and treatment of diseases such as intracranial aneurysms, vascular stenosis, and arteriovenous malformations.
[0003] However, segmentation of bone tissue and blood vessels has always been a challenge in the processing of CTA image data. Traditional methods typically combine thresholding, region growing, and dilatational erosion to segment bone tissue in CTA image data. However, this approach has several drawbacks, such as insufficiently fine segmentation of the resulting bone tissue, a tendency to missegment intracranial and intracervical blood vessels, difficulty in handling special cases such as bone calcification and scaffolds, and the need for multiple iterations, which consumes a significant amount of time. Summary of the Invention
[0004] This application provides an image segmentation method, apparatus, storage medium, and electronic device, which can solve the above-mentioned problems. The technical solution is as follows: In a first aspect, embodiments of this application provide an image segmentation method, the method comprising: Acquire CTA image data; The image data is subjected to a first bone segmentation process to obtain a first bone segmentation result; The image data is cropped based on the first bone segmentation result to obtain first cropped image data that only includes the area above the shoulders. The first cropped image data is subjected to a second bone segmentation process to obtain a second bone segmentation result. The first bone segmentation result and the second bone segmentation result are fused to obtain the fused bone result.
[0005] Secondly, embodiments of this application provide an image segmentation apparatus, the apparatus comprising: The data acquisition module is used to acquire CTA image data; The first segmentation module is used to perform a first bone segmentation process on the image data to obtain a first bone segmentation result; The image cropping module is used to crop the image data according to the first bone segmentation result to obtain first cropped image data that only includes the area above the shoulders; The second segmentation module is used to perform second bone segmentation processing on the first cropped image data to obtain the second bone segmentation result. The result fusion module is used to fuse the first bone segmentation result and the second bone segmentation result to obtain the fused bone result.
[0006] Thirdly, embodiments of this application provide a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the above-described method steps.
[0007] Fourthly, embodiments of this application provide a computer program product that stores multiple instructions adapted for loading by a processor and executing the above-described method steps.
[0008] Fifthly, embodiments of this application provide an electronic device that may include: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the above-described method steps.
[0009] The beneficial effects of the technical solutions provided in some embodiments of this application include at least the following: In this application, a first bone segmentation process is performed on the CTA image data to obtain a first bone segmentation result. Further, the image data is cropped based on the first bone segmentation result to obtain first cropped image data that includes only the area above the shoulders. This application crops the image data based on the first bone segmentation result without manual annotation or the use of fixed coordinates. The cropping range dynamically adapts to subjects of different body types and scanning ranges, effectively preserving image data of the head and neck bone region above the shoulders, eliminating interference from the chest and abdominal structures, and providing input information focused on the head and neck bone structures for the next stage of bone segmentation processing. Furthermore, a second bone segmentation process is performed on the first cropped image data to obtain a second bone segmentation result. The first and second bone segmentation results are then fused to obtain a fused bone result. This application cascades two bone segmentation processes with different accuracies and inputs. The first stage locates a small region of interest in the image data, and the second stage performs high-precision bone segmentation on the small cropped image. The results of the two bone segmentation processes are then fused to obtain a more accurate fused bone result, which significantly improves the segmentation accuracy of head and neck bone tissues. This avoids problems such as missed or incorrect segmentation caused by receptive field or resolution limitations in single-stage bone segmentation processing, and can greatly improve inference efficiency and data generalization ability. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of the architecture of an image segmentation method provided in an embodiment of this application; Figure 2 This is a schematic flowchart of an image segmentation method provided in an embodiment of this application; Figure 3 This is a schematic diagram of a process for obtaining a first bone segmentation result provided in an embodiment of this application; Figure 4 This is a schematic diagram of a process for obtaining a second bone segmentation result provided in an embodiment of this application; Figure 5 This is a schematic diagram of a process for obtaining fusion bone results provided in an embodiment of this application; Figure 6 This is a flowchart illustrating an image segmentation method provided in an embodiment of this application; Figure 7 This is a schematic flowchart of an image segmentation method provided in an embodiment of this application; Figure 8 This is a schematic flowchart illustrating how to obtain the first bone segmentation result and the cervical vertebrae bone result according to an embodiment of this application; Figure 9 This is a schematic diagram of a process for obtaining second cropped image data provided in an embodiment of this application; Figure 10 This is a schematic diagram of a process for obtaining blood vessel segmentation results provided in an embodiment of this application; Figure 11 This is a schematic flowchart of an image segmentation method provided in an embodiment of this application; Figure 12 This is a schematic diagram of a process for obtaining a second bone segmentation result provided in an embodiment of this application; Figure 13 This is a schematic flowchart of an image segmentation method provided in an embodiment of this application; Figure 14 This is a schematic diagram of the structure of an image segmentation device provided in an embodiment of this application; Figure 15 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this application, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0014] The present application will now be described in detail with reference to specific embodiments.
[0015] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the features, information, and data involved in this application were all obtained under full authorization.
[0016] like Figure 1 As shown, Figure 1 This is a schematic diagram of the architecture of an image segmentation method provided in an embodiment of this application. Figure 1 It includes at least a server 101 that performs the image segmentation method, and also includes multiple electronic devices 102 for uploading CTA image data. It is understood that... Figure 1 The number of servers and electronic devices shown is for illustrative purposes only, and this application does not impose any limitations on them.
[0017] The aforementioned server 101 can be a standalone server device, such as a rack-mount, blade, tower, or cabinet-type server device, or a workstation, mainframe, or other hardware device with strong computing power; it can also be a server cluster composed of multiple servers. The servers in the service cluster can be composed in a symmetrical manner, where each server is functionally and hierarchically equivalent in the transaction chain, and each server can provide services to the outside world independently. Providing services independently can be understood as not requiring the assistance of other servers.
[0018] For example, a server can be multiple physical servers, each hardware-independent. Alternatively, a server can be multiple virtual servers deployed within the same hardware resource pool. Virtual server deployment methods include, but are not limited to, VMware, VirtualBoX, and Virtual PC.
[0019] It is understood that server 101 also possesses other service capabilities and functions to complete the tasks described in the following embodiments. For example, server 101 also provides portal services, resource management services, and CI / CD services, etc.
[0020] Electronic device 102 includes, but is not limited to: wearable devices, handheld devices, personal computers, tablets, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. Electronic devices may have different names in different networks, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), and electronic devices in 5G networks or future evolved networks.
[0021] In this embodiment, the electronic device 102 may also be equipped with a display device. The display device can be any device capable of displaying information, such as a cathode ray tube display (CR), a light-emitting diode display (LED), an electronic ink screen, a liquid crystal display (LCD), or a plasma display panel (PDP). For example, a user can use the display device on the electronic device 102 to send CTA image data to the server 101 and view the fusion bone segmentation results sent by the server 101.
[0022] Multiple electronic devices and multiple servers can communicate through communication links established by communication protocols. For example, the network can be a wireless network or a wired network. Wireless networks include, but are not limited to, cellular networks, wireless LANs, infrared networks, or Bluetooth networks. Wired networks include, but are not limited to, Ethernet, universal serial bus (USB), or controller area networks. In one or more embodiments of the specification, technologies and / or formats including Hyper-Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network (such as target compressed packets). Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0023] In one embodiment, such as Figure 2 The diagram shown is a flowchart illustrating an image segmentation method provided in an embodiment of this application. This method can be implemented using a computer program and can run on an image segmentation device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application.
[0024] Specifically, the image segmentation method includes: S101. Obtain CTA image data.
[0025] Acquire the CT angiography (CTA) image data to be processed. CTA image data is three-dimensional volumetric data acquired by CT scanning equipment during the arterial phase after intravenous injection of iodine contrast agent into the subject. It can simultaneously present high-contrast information of vascular structures and surrounding tissues (such as bone and soft tissue). CTA image data is usually stored in DICOM (Digital Imaging and Communications in Medicine) format, containing one or more consecutive two-dimensional cross-sectional image slices, which together form a three-dimensional voxel grid. The gray value of each voxel reflects the X-ray attenuation coefficient of the corresponding tissue.
[0026] In this application, the CTA image data is preprocessed after acquisition. For example, a window width and window level suitable for joint bone and blood vessel analysis are used, such as a window width of 800 HU and a window level of 300 HU. The image data is truncated and offset, and the offset data is divided by the window width for normalization, mapping all pixel values to the [0, 1] interval. This preprocessing preserves the relative contrast between bone tissue and blood vessels in the original CTA image data and meets the numerical range requirements of the subsequent bone segmentation model, which is beneficial for the convergence of the bone segmentation model.
[0027] S102. Perform the first bone segmentation process on the image data to obtain the first bone segmentation result.
[0028] First, the image data undergoes initial bone segmentation to quickly extract all regions that could potentially belong to bone tissue, obtaining preliminary bone tissue distribution information as the first bone segmentation result. This first bone segmentation result is essentially a three-dimensional binary mask with the same spatial dimensions as the input CTA image data.
[0029] This application utilizes a first bone segmentation model to perform initial bone segmentation on image data. The first bone segmentation model can employ a semantic segmentation network based on an encoder-decoder structure, such as U-Net, 3D U-Net, V-Net, ResUNet, or nnU-Net. This first bone segmentation model effectively learns the high-density features of bone tissue in CTA image data and preserves spatial details through skip connections, thereby generating preliminary first bone segmentation results.
[0030] For example, this application adopts a first bone segmentation model based on the ResUnet3D network architecture. Compared with the Unet network architecture, this network architecture adds a residual mechanism, which can effectively alleviate the problems of gradient vanishing and difficulty in training that arise as the number of network layers increases.
[0031] like Figure 3 As shown, Figure 3 This is a schematic flowchart of obtaining a first bone segmentation result provided in an embodiment of this application. When inferring CTA image data 201 based on the first bone segmentation model, the CTA image data 201 is cut along the data height and divided into CTA image data 2011 and CTA image data 2012.
[0032] Then, bilinear interpolation was used to interpolate CTA image data 2011 and CTA image data 2012 to a size of 224×176×176, and they were input into the first bone segmentation model for inference to obtain two initial first bone segmentation results.
[0033] Furthermore, the two initial first bone segmentation results are mapped to the original dimensions corresponding to CTA image data 2011 and CTA image data 2012 respectively using the nearest neighbor interpolation method. The two initial first bone segmentation results are then stitched together in the Z-axis direction to merge them into one, resulting in a complete first bone segmentation result 202.
[0034] Furthermore, the first bone segmentation result 202 can be post-processed, such as by performing connected component analysis, to remove isolated noise points and small connected components in the first bone segmentation result 202.
[0035] S103. Based on the first bone segmentation result, the image data is cropped to obtain the first cropped image data that only includes the area above the shoulder.
[0036] Using the first bone segmentation results obtained in the first stage, considering the large difference in cross-sectional area between head and neck bone tissues (such as skull and cervical vertebrae) and chest bone tissues (such as thoracic vertebrae, ribs, and scapula), and the difficulty in accurately predicting head and neck bone tissues when segmenting head and neck bone tissues and chest bone tissues at the same time, the image data of the area above the shoulders is cropped out so that head and neck bone tissues can be inferred separately.
[0037] Based on the first bone segmentation result, locate the segmentation layer representing the shoulder in the first bone segmentation result. Along the Z-axis direction from the top of the head to the bottom, the image data above this segmentation layer is the part of the image data that needs to be retained (i.e., the first cropped image data), and the image data below this segmentation layer is the part of the image data that needs to be cropped out.
[0038] In one embodiment, the method for determining the position of the uppermost edge of the shoulder in the first bone segmentation result may be: calculating the cross-sectional area of each layer of bone tissue along the Z-axis of the first bone segmentation result, for example, characterizing the cross-sectional area by the number of non-zero voxels; traversing along the Z-axis direction from the top of the head to the bottom, determining the target layer that meets the preset area mutation condition, for example, if the area of the lower layer is several times that of the current layer, such as 1.5 times, then the current layer is determined to be the segmentation layer that characterizes the position of the uppermost edge of the shoulder.
[0039] In another embodiment, in addition to cropping the image data along the Z-axis based on the shoulder position in the first bone segmentation result, the image data can also be cropped along the X-axis and Y-axis based on the location of bone tissue in the X-axis and Y-axis directions in the first bone segmentation result, which includes the area above the shoulder. This further reduces the number of voxels included in the first cropped image data, which is beneficial to the efficiency of the second bone segmentation process based on the first cropped image data in the second stage.
[0040] S104. Perform second bone segmentation processing on the first cropped image data to obtain the second bone segmentation result.
[0041] After obtaining the first cropped image data containing only the area above the shoulders, a second bone segmentation process is performed on the first cropped image data to achieve high-precision segmentation of the head and neck bone tissue. This step, as the fine segmentation stage, forms a coarse-fine cascade architecture with the first bone segmentation process based on all image data, aiming to perform reasoning and segmentation of the head and neck bone tissue with a pursuit of detail fidelity and accurate boundaries.
[0042] The second bone segmentation process can also be performed using a second bone segmentation model based on a deep learning-based semantic segmentation model. The second bone segmentation model can employ the same or different network structures as the first bone segmentation model, such as 3D U-Net, ResUNet, or Attention U-Net. Preferably, a model with higher resolution input, a deeper network, or a more complex attention mechanism can be deployed as the second bone segmentation model to improve segmentation performance without significantly increasing computational burden.
[0043] like Figure 4 As shown, Figure 4 This is a schematic flowchart illustrating a method for obtaining a second bone segmentation result according to an embodiment of this application. The first cropped image data 203 is subjected to second bone segmentation processing to obtain a second bone segmentation result 204. The second bone segmentation result 204, compared to the first bone segmentation result 202, only includes bone tissue in the area above the shoulder.
[0044] S105. The first bone segmentation result and the second bone segmentation result are fused to obtain the fused bone result.
[0045] The first bone segmentation result 202 is a preliminary bone mask covering the chest and abdomen area, which has good global connectivity and robustness, but it may have problems such as blurred boundaries and missing small structures in the head and neck bone region. The second bone segmentation result is a high-precision bone mask only for the area above the shoulders, which can reconstruct the skull, cervical vertebrae and other structures.
[0046] like Figure 5 As shown, Figure 5 This is a schematic flowchart illustrating a method for obtaining fused bone results according to an embodiment of this application. The first bone segmentation result 202 and the second bone segmentation result 204 are fused to obtain the fused bone result.
[0047] The process of fusing the first and second bone segmentation results can be as follows: First, align the first and second bone segmentation results to the same coordinate system in three-dimensional space. Then, fuse the probability maps corresponding to the first and second bone segmentation results respectively to obtain a fused probability map that includes the confidence levels of multiple voxels. Infer from the fused probability map to determine the probability that each voxel belongs to bone tissue. After the inference is completed, binarize the fused probability map to obtain the fused bone result representing the bone tissue included in the image data.
[0048] In one embodiment, the first bone segmentation result and the second bone segmentation result are fused based on a first weight corresponding to the first bone segmentation result and a second weight corresponding to the second bone segmentation result to obtain a fused bone result. The second weight is greater than the first weight.
[0049] In this embodiment, to improve the accuracy and robustness of the fused bone results, a weighted fusion strategy is used to fuse the first and second bone segmentation results. This is because the first bone segmentation result is derived from a coarse segmentation of the entire image, which, although covering a wide area, has low confidence in the head and neck bone region. The second bone segmentation result, on the other hand, is based on cropped data from the first cropped image, and has higher segmentation confidence and detail fidelity in the area above the shoulders. Therefore, the second bone segmentation result is given a higher weight during the fusion process. For example, the first weight corresponding to the first bone segmentation result is 0.4, and the second weight corresponding to the second bone segmentation result is 0.6.
[0050] For example, the weight corresponding to the second bone segmentation result decreases along the Z-axis from the top of the head to the bottom, while the weight corresponding to the first bone segmentation result increases along the Z-axis from the top of the head to the bottom. This is manifested in the area where the top of the skull is located, the second weight corresponding to the second bone segmentation result is 0.8, the first weight corresponding to the first bone segmentation result is 0.2, the second weight corresponding to the second bone segmentation result is 0.6, the first weight corresponding to the first bone segmentation result is 0.4, and the first weight corresponding to the first bone segmentation result is 1 in the area below the shoulders.
[0051] In this embodiment, by performing weighted fusion processing on the first bone segmentation result and the second bone segmentation result, the segmentation accuracy of the head and neck region can be improved, the bone structure of the head and neck region can be effectively restored, and the structural integrity of all bone tissues can be guaranteed at the same time, and a smooth transition can be achieved in the shoulder to avoid hard boundary artifacts.
[0052] In another embodiment, after obtaining the fused bone result, bone tissue removal processing is performed on the image data based on the fused bone result to remove the areas containing bone tissue from the image data, resulting in image data excluding bone tissue. Both the fused bone result and the image data excluding bone tissue are sent to an electronic device for display, allowing users to view the fused bone result and the image data excluding bone tissue as needed, meeting the usage needs of complex scenarios.
[0053] In this application, image data is cropped based on the first bone segmentation result without manual annotation or fixed coordinates. The cropping range dynamically adapts to subjects with different body types and scanning ranges, effectively preserving image data of the head and neck bone region above the shoulders, eliminating interference from the chest and abdominal structures, and providing input information focused on the head and neck bone structure for the next stage of bone segmentation processing. Furthermore, a second bone segmentation process is performed on the first cropped image data to obtain a second bone segmentation result, which is then fused to obtain a fused bone result. This application cascades two bone segmentation processes with different accuracies and inputs. The first stage locates a small region of interest in the image data, and the second stage performs high-precision bone segmentation on the small cropped image. The two bone segmentation results are then fused to obtain a more accurate fused bone result, significantly improving the segmentation accuracy of head and neck bone tissue. This avoids problems such as missed or incorrect segmentation caused by receptive field or resolution limitations in single-stage bone segmentation processing, and can greatly improve inference efficiency and data generalization ability.
[0054] In one embodiment, such as Figure 6 The diagram shown is a flowchart illustrating an image segmentation method provided in an embodiment of this application. This method can be implemented using a computer program and can run on an image segmentation device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application.
[0055] Specifically, the image segmentation method includes: S201. Obtain CTA image data.
[0056] See S101 above; it will not be repeated here.
[0057] S202. Perform the first bone segmentation process on the image data to obtain the first bone segmentation result.
[0058] See S102 above, which will not be repeated here.
[0059] S203. Traverse the first bone segmentation results along the Z-axis, compare the span of the first bone segmentation results in adjacent layers in the X-axis direction, and obtain the first Z-axis coordinates corresponding to the span growth meeting the preset conditions.
[0060] Specifically, find the maximum value Z in the Z-axis direction in the first bone segmentation result maX and the minimum value Z min , traverse downward along the direction from the top of the head to the bottom, that is, traverse along the Z-axis from Z max to Z min traverse, take the current layer i and the next layer i + 1, compare the number of voxels of the bone tissue in the first bone segmentation result in the X-axis direction in layer Zi and layer Zi+1, that is, compare the sizes of Xi and Xi+1 in the X-axis direction. When K × Xi < Xi+1, at this time the span growth meets the preset condition, and obtain the first Z-axis coordinate corresponding to the current layer i. K is a coefficient greater than 1, for example, K is 1.5.
[0061] In another embodiment, it is also possible to traverse upward along the direction from the bottom to the top of the head, compare the spans of the first bone segmentation result in the X-axis direction in adjacent layers, and determine that the span growth meets the preset condition when K × Xi < Xi+1, and obtain the first Z-axis coordinate corresponding to the current layer i.
[0062] S304. Take the first Z-axis coordinate as the upper edge position of the shoulder, and crop the image data to obtain the first cropped image data including only the area above the shoulder.
[0063] Locate the segmentation layer in multiple image data based on the first Z-axis coordinate. Along the Z-axis direction from the top of the head to the bottom, the image data above this segmentation layer is the first cropped image data obtained by cropping.
[0064] In one embodiment, take the first Z-axis coordinate as the upper edge position of the shoulder, extract the target bone segmentation result corresponding to the bone structure in the area above the shoulder from the first bone segmentation result; expand the spatial boundary of the area where the target bone segmentation result is located to obtain the cropping boundary; crop the image data based on the cropping boundary to obtain the first cropped image data including only the area above the shoulder.
[0065] Extract the target bone segmentation result corresponding to the bone structure in the area above the shoulder from the first bone segmentation result, and determine the maximum value X max and the minimum value X min in the X-axis direction, the maximum value Y max and the minimum value Y min in the Y-axis coordinate direction, as well as the maximum value in the Z-axis direction and the first Z-axis coordinate, and determine the spatial boundary by combining the above multiple coordinates. Further expand the above spatial boundary, for example, expand outward by 20 voxels, to obtain the cropping boundary. The first Z-axis coordinate may not be expanded outward. Further crop the image data based on the cropping boundary to obtain the first cropped image data including only the area above the shoulder.
[0066] In this embodiment, after extracting the area above the shoulder, the determined spatial boundary is appropriately expanded to ensure the bone tissue in the area above the shoulder, such as the mandibular ramus, skull base, and atlantoaxial joint, while retaining a small amount of surrounding soft tissue as contextual information. This avoids boundary truncation or structural loss due to excessively tight cutting, and provides sufficient information for the subsequent fine bone segmentation process in the second stage.
[0067] S205. Perform second bone segmentation processing on the first cropped image data to obtain the second bone segmentation result.
[0068] See S104 above; it will not be repeated here.
[0069] S206. The first bone segmentation result and the second bone segmentation result are fused to obtain the fused bone result.
[0070] See S105 above; it will not be repeated here.
[0071] In this application, image data is cropped based on the first bone segmentation result without manual annotation or fixed coordinates. The cropping range dynamically adapts to subjects with different body types and scanning ranges, effectively preserving image data of the head and neck bone region above the shoulders, eliminating interference from the chest and abdominal structures, and providing input information focused on the head and neck bone structure for the next stage of bone segmentation processing. Furthermore, a second bone segmentation process is performed on the first cropped image data to obtain a second bone segmentation result, which is then fused to obtain a fused bone result. This application cascades two bone segmentation processes with different accuracies and inputs. The first stage locates a small region of interest in the image data, and the second stage performs high-precision bone segmentation on the small cropped image. The two bone segmentation results are then fused to obtain a more accurate fused bone result, significantly improving the segmentation accuracy of head and neck bone tissue. This avoids problems such as missed or incorrect segmentation caused by receptive field or resolution limitations in single-stage bone segmentation processing, and can greatly improve inference efficiency and data generalization ability.
[0072] In one embodiment, such as Figure 7 The diagram shown is a flowchart illustrating an image segmentation method provided in an embodiment of this application. This method can be implemented using a computer program and can run on an image segmentation device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application.
[0073] Specifically, the image segmentation method includes: S301. Obtain CTA image data.
[0074] See S101 above; it will not be repeated here.
[0075] S302. Perform the first bone segmentation process on the image data to obtain the first bone segmentation result and the cervical vertebrae result.
[0076] In this embodiment, when performing the first bone segmentation process on the image data using the first bone segmentation model, the training loss function is composed of the bone loss function loss and the cervical vertebrae loss function loss, both of which are cross-entropy losses (bceloss). Considering the size of the bones and cervical vertebrae, the ratio is adjusted to 1:5. The loss function of the first bone segmentation model is shown in the following formula:
[0077] Based on the above formula, the image data is subjected to the first bone segmentation process, resulting in the first bone segmentation result and the cervical vertebrae result. For example... Figure 8 As shown, Figure 8 This is a schematic flowchart illustrating the process of obtaining a first bone segmentation result and a cervical vertebrae result according to an embodiment of this application. When performing inference on image data 301 based on the first bone segmentation model, the CTA image data 301 is segmented along its height, resulting in CTA image data 3011 and CTA image data 3012. Further, CTA image data 3011 and CTA image data 3012 are input into the first bone segmentation model to obtain the first bone segmentation result 302 and the cervical vertebrae result 303, respectively.
[0078] S303. Based on the first bone segmentation result, the image data is cropped to obtain the first cropped image data that only includes the area above the shoulder.
[0079] See S103 above, which will not be repeated here.
[0080] S304. Perform second bone segmentation processing on the first cropped image data to obtain the second bone segmentation result.
[0081] See S104 above; it will not be repeated here.
[0082] S305. The first bone segmentation result and the second bone segmentation result are fused to obtain the fused bone result.
[0083] See S105 above; it will not be repeated here.
[0084] S306. Based on the cervical vertebrae structure and the first bone segmentation result, the image data is cropped to obtain second cropped image data including the target region.
[0085] The target region encompasses the cervical vertebrae and skull, representing a region of interest where bone tissue and blood vessels highly overlap in CTA image data. In other words, false positives due to bone and blood vessel confounding are concentrated in the target region. Therefore, this embodiment locates the target region, including the cervical vertebrae and skull, using the cervical vertebrae results and the first bone segmentation results. When processing the fused bone results subsequently, blood vessels overlapping with the target region in the fused bone results are removed. This significantly reduces false positives for blood vessels in the fused bone results, eliminates the need to focus on blood vessel information outside the target region in the image data, improves the efficiency of blood vessel segmentation, and avoids excessive removal of blood vessels from the fused bone results.
[0086] In one embodiment, such as Figure 9 As shown, Figure 9 This is a schematic diagram of a process for obtaining second cropped image data according to an embodiment of this application. S306 includes the following steps: S306-1. Using the minimum Z-axis coordinate in the cervical vertebrae results as a reference position, move a preset number of voxels downwards along the direction from the top of the head to the bottom to obtain the second Z-axis coordinate, and use the second Z-axis coordinate as the lower boundary of the Z-axis of the target area.
[0087] The minimum Z-axis coordinate in the cervical vertebrae results represents the lowest position of the cervical vertebrae. Then, the data is moved downwards by a preset number of voxels, such as 80 voxels, from the top of the head to the bottom. The second Z-axis coordinate is used as the lower boundary of the Z-axis of the target region to ensure that the second cropped image data including the target region contains sufficient vascular information, thereby enabling better vascular segmentation.
[0088] S306-2. Determine the spatial boundary of the target region based on the first bone segmentation result and the second Z-axis coordinate.
[0089] Based on the region in three-dimensional space where the first bone segmentation result is located, determine the maximum value X' in the X-axis direction. max and minimum value X' min The maximum value Y' in the Y-axis coordinate direction max and minimum value Y' min The spatial boundary of the target area is determined by combining the maximum value of the Z-axis and the second Z-axis coordinate with the above coordinates.
[0090] S306-3. Cropping the image data according to the spatial boundary of the target area to obtain second cropped image data including the target area.
[0091] After determining the spatial boundary of the target region, a preset number of voxels are extended outward, for example, by 20 voxels (the second Z-axis coordinate may not be extended outward), to obtain the cropping boundary of the target region, avoiding boundary truncation or structural loss due to overly tight cropping. Further cropping of the image data is performed based on the cropping boundary of the target region to obtain second cropped image data including the target region.
[0092] S307. Perform blood vessel segmentation processing on the second cropped image data to obtain the blood vessel segmentation result.
[0093] In this embodiment, a blood vessel segmentation model can be used to perform blood vessel segmentation processing on the second cropped image data. For example... Figure 10 As shown, Figure 10 This is a schematic diagram illustrating a process for obtaining vessel segmentation results according to an embodiment of this application. The vessel segmentation model 302 is built based on the U-Net architecture and includes two sub-modules: an encoder and a decoder. The encoder extracts multi-scale features layer by layer through convolution and pooling, while the decoder restores spatial resolution through upsampling and skip connections. It is understood that the vessel segmentation model 302 can also be a model with other architectures; this embodiment is merely illustrative.
[0094] The second cropped image data 301 is divided into multiple continuous and overlapping sub-volumes along the Z-axis, each sub-volume having a size of approximately 128×128×128 voxels. To adapt to the input requirements of the blood vessel segmentation model 302, each slice is uniformly interpolated to a fixed size of 320×192×192 using bilinear interpolation, forming a local image 3011 of the second cropped image data.
[0095] Furthermore, each interpolated local image 3011 is input into the blood vessel segmentation model 302 for independent inference, and finally outputs a blood vessel probability map 303 of the same size as the input. Each voxel value represents the probability that the corresponding position belongs to a blood vessel, and the value range is [0, 1].
[0096] Furthermore, given that the original second cropped image data 301 is divided into multiple segments, the vascular probability maps 303 corresponding to each segment need to be spatially aligned and fused. Specifically, the vascular probability maps of all segments are stitched together according to their positions in the original image, and a weighted average method (e.g., the weights decrease linearly as they get closer to the boundary) is used to smoothly fuse the overlapping areas, thereby generating a complete vascular probability map covering the entire second cropped image data.
[0097] Furthermore, the fused vessel probability map is thresholded (e.g., a threshold of 0.5) to obtain a preliminary binary vessel mask. Subsequently, connected component analysis is performed on the mask to remove small, isolated connected components with a volume smaller than a preset threshold (e.g., 50 voxels) to eliminate the influence of noise points, artifacts, or microcalcifications, thereby further improving the integrity and robustness of the vessel segmentation results.
[0098] After the above processing, the final blood vessel segmentation result 304 is obtained. Blood vessel segmentation result 304 is a three-dimensional binary mask with the same spatial size as the second cropped image data 301, which accurately identifies the location of the main enhanced blood vessels (such as the vertebral artery, internal carotid artery, basilar artery, etc.) in the target area.
[0099] S308. Based on the vascular segmentation results, perform vascular removal processing on the fused bone results to obtain the target bone results.
[0100] To avoid false positives due to slight contraction of the vessel segmentation boundaries or registration errors, a morphological expansion operation is first performed on the vessel segmentation results, with an expansion radius of 1 voxel. This operation slightly expands the vessel area, ensuring coverage of all vessel voxels that might be misidentified as bone. This is particularly suitable for scenarios with closely spaced anatomical structures, such as the vertebral artery passing through the transverse foramen of the cervical vertebrae and the internal carotid artery closely following the skull base.
[0101] Furthermore, the expanded vascular mask is logically inverted to obtain the non-vascular region mask. This non-vascular region mask is then multiplied voxel-by-voxel by the fused bone result. This operation sets all predicted bone voxels that overlap with the expanded vascular region to 0, while the remaining true bone structure remains unchanged. The final output is the target bone result, which fully preserves bone tissue while effectively eliminating false positives caused by enhanced vessels such as intracranial arteries, vertebral arteries, and internal carotid arteries.
[0102] In this application, image data is cropped based on the first bone segmentation result without manual annotation or fixed coordinates. The cropping range dynamically adapts to subjects with different body types and scanning ranges, effectively preserving image data of the head and neck bone region above the shoulders, eliminating interference from the chest and abdominal structures, and providing input information focused on the head and neck bone structure for the next stage of bone segmentation processing. Furthermore, a second bone segmentation process is performed on the first cropped image data to obtain a second bone segmentation result, which is then fused to obtain a fused bone result. This application cascades two bone segmentation processes with different accuracies and inputs. The first stage locates a small region of interest in the image data, and the second stage performs high-precision bone segmentation on the small cropped image. The two bone segmentation results are then fused to obtain a more accurate fused bone result, significantly improving the segmentation accuracy of head and neck bone tissue. This avoids problems such as missed or incorrect segmentation caused by receptive field or resolution limitations in single-stage bone segmentation processing, and can greatly improve inference efficiency and data generalization ability.
[0103] In one embodiment, such as Figure 11 The diagram shown is a flowchart illustrating an image segmentation method provided in an embodiment of this application. This method can be implemented using a computer program and can run on an image segmentation device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application.
[0104] Specifically, the image segmentation method includes: S401. Obtain CTA image data.
[0105] See S101 above; it will not be repeated here.
[0106] S402. Perform the first bone segmentation process on the image data to obtain the first bone segmentation result.
[0107] See S102 above, which will not be repeated here.
[0108] S403. Based on the first bone segmentation result, the image data is cropped to obtain the first cropped image data that only includes the area above the shoulder.
[0109] See S103 above, which will not be repeated here.
[0110] S404. Extract the global features of the first cropped image data to obtain the global feature image.
[0111] Considering the need for precise organ and tissue boundaries in image data, and the requirement for surrounding information to determine issues such as scaffolds and calcifications, a network structure that balances global and local information was constructed. Simultaneously, considering speed requirements, the network structure was pruned and lightweighted while maintaining accuracy, resulting in the network structure FSnet (Full Scale Net) used in this embodiment. Balancing global and local information involves extracting local and global information separately using two simple networks (Fnet model and Snet model), then combining them to consider both. To meet the speed requirements of clinical inference, the original 5-layer ResUNet was pruned while maintaining accuracy, simplifying it into a global feature extraction network with a 2-layer global feature extraction path, and a local feature extraction network with a streamlined 4-layer structure, significantly reducing computational load and runtime.
[0112] Specifically, the FSnet model comprises the Fnet and Snet models, responsible for inferring global and local feature images, respectively. The global feature image represents the entire image data, while the S-graph represents the local feature images cropped from the image chunks. The Fnet model, based on ResUnet, is a 5-layer structure containing an encoder and a decoder. The Snet model is a 4-layer network, also consisting of encoder and decoder sub-modules.
[0113] During the training of the FSnet model, in order to simulate the fluctuations in CT values caused by different scanning equipment or contrast agent concentrations, the brightness of the original sample data volume was adjusted. Since calcification is common in intracranial blood vessels (such as the vertebral artery and basilar artery), in order to enhance the robustness of the FSnet model to such complex scenarios, a point on an intracranial blood vessel was randomly selected as the center and Gaussian blur was applied to locally enhance the brightness, simulating the CT appearance of calcified plaques. This enabled the FSnet model to learn to distinguish between "real calcification" and "false positive bone", thereby enhancing the model's generalization and segmentation ability.
[0114] In this study, brightness variation is used to classify sample data into bone and non-bone tissues. First, the sample data volume is multiplied by a global coefficient, global param, ranging from 0.9 to 1.1. Then, the bone tissue is multiplied by a local coefficient, local param, ranging from 0.8 to 1.2, to enhance the difference between bone and non-bone tissues.
[0115] Since calcification lesions mainly occur in intracranial blood vessels, when simulating calcification data, a point in an intracranial blood vessel is randomly selected, and Gaussian blurring is applied around that point to enhance the local brightness of the blood vessel and simulate calcification lesions.
[0116] The loss function of the FSnet model is composed of the loss of the global Fnet model. f loss of the local SNet model s The composition, in which the ratio is 1:5 (a=1, b=5), includes the following formula:
[0117]
[0118]
[0119] The loss function of the Fnet model is... f Only the BCE loss, which is the binary cross-entropy loss function, is used in the SNet model. s It consists of Dice loss and BCE loss. Dice loss is a loss function based on the degree of region overlap, where c=1 and d=5.
[0120] like Figure 12 As shown, Figure 12 This is a schematic flowchart of obtaining a second bone segmentation result provided in an embodiment of this application. The first cropped image data 4011, including the first cropped image data 401, is linearly interpolated to 224×176×176 to obtain an F-map. The F-map is input into the Fnet model 402, so that the encoder of the Fnet model 402 gradually compresses the spatial size and extracts high-level semantic features through layer-by-layer convolution and pooling operations, and the decoder restores the spatial resolution through upsampling and skip connections, and finally outputs a global feature image 403 of the same size as the input.
[0121] The global feature image 403 not only preserves the spatial layout of the original image but also incorporates cross-regional contextual information, effectively recognizing the overall morphology of bone tissue (such as skull contours and cervical vertebrae alignment) and suppressing the influence of local noise or artifacts. This global feature image 403 serves as a priori guidance for subsequent local inference, providing macroscopic consistency support for refined segmentation.
[0122] S405. Obtain a local image from a local region in the first cropped image data, and obtain a local feature image that is spatially aligned with the local region from the global feature image.
[0123] After obtaining the global feature image 403, multiple local images 4012 (such as small blocks of 96×96×96 voxels) are cropped from the first cropped image data 401, with each local image focusing on the key region of the head and neck bones.
[0124] Simultaneously, a local feature image, spatially aligned with the local image 4012, is extracted from the global feature image 403; that is, a feature patch at the corresponding location. This local feature image inherits global context information, providing semantic guidance for subsequent local inference.
[0125] S406. The local image and the local feature image are stitched together to obtain a stitched image. The stitched image is then input into the local information extraction model for inference to obtain the second bone segmentation result.
[0126] The local image 4012 is concatenated with its corresponding local feature image at the channel level to generate a mosaic image 404. Then, the mosaic image 404 is input into the multi-local information extraction model Snet405 for inference. The Snet model 405, also based on an encoder-decoder structure, gradually reconstructs the fine boundaries of the local bone structure through multi-level feature fusion. The final output is a bone probability map 406 of the same size as the mosaic image.
[0127] To obtain the complete second bone segmentation result 407, this application employs a sequential reasoning mechanism. The above process is sequentially executed on all local regions of the first cropped image data 4012, and the bone probability maps of each region are stitched together and fused to finally generate a three-dimensional bone mask covering the entire first cropped image data 401, which is the second bone segmentation result 407.
[0128] S407. The first bone segmentation result and the second bone segmentation result are fused to obtain the fused bone result.
[0129] See S105 above; it will not be repeated here.
[0130] In this application, image data is cropped based on the first bone segmentation result without manual annotation or fixed coordinates. The cropping range dynamically adapts to subjects with different body types and scanning ranges, effectively preserving image data of the head and neck bone region above the shoulders, eliminating interference from the chest and abdominal structures, and providing input information focused on the head and neck bone structure for the next stage of bone segmentation processing. Furthermore, a second bone segmentation process is performed on the first cropped image data to obtain a second bone segmentation result, which is then fused to obtain a fused bone result. This application cascades two bone segmentation processes with different accuracies and inputs. The first stage locates a small region of interest in the image data, and the second stage performs high-precision bone segmentation on the small cropped image. The two bone segmentation results are then fused to obtain a more accurate fused bone result, significantly improving the segmentation accuracy of head and neck bone tissue. This avoids problems such as missed or incorrect segmentation caused by receptive field or resolution limitations in single-stage bone segmentation processing, and can greatly improve inference efficiency and data generalization ability.
[0131] In one embodiment, such as Figure 13 The diagram shown is a flowchart illustrating an image segmentation method provided in an embodiment of this application. This method can be implemented using a computer program and can run on an image segmentation device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application.
[0132] In this embodiment, CTA image data 501 is acquired, and the image data 501 is subjected to a first bone segmentation process by the first bone segmentation model 502 to obtain the first bone segmentation result and the cervical vertebrae result. This step is described in S302 above.
[0133] Furthermore, the image data 501 is cropped based on the first bone segmentation result to obtain first cropped image data that includes only the area above the shoulders. Then, the first cropped image data is processed by the second bone segmentation model 503 to obtain the second bone segmentation result. This step is described in S404-S406 above.
[0134] Furthermore, based on the cervical vertebrae results and the first bone segmentation results, the image data 501 is cropped to obtain second cropped image data including the target region. Then, the second cropped image data is processed for blood vessel segmentation using the blood vessel segmentation model 504 to obtain the blood vessel segmentation result. This step is described in S306-S307 above.
[0135] Further, based on the vascular segmentation results, the fused bone result is subjected to vascular removal processing to obtain the target bone result 505. This step is described in S307 above.
[0136] This application cascades two bone segmentation processes with different accuracies and inputs, as well as a cascaded vessel removal process. In the first stage, a small region of interest is located on the image data. In the second stage, high-precision bone segmentation is performed on the small cropped image. The results of the two bone segmentation processes are then fused to obtain a more accurate fused bone result. False positives for blood vessels in the fused bone result are also removed. This significantly improves the segmentation accuracy of head and neck bone tissues and avoids problems such as missed or incorrect segmentation caused by receptive field or resolution limitations in single-stage bone segmentation. It can also greatly improve inference efficiency and data generalization ability.
[0137] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0138] Please see Figure 14This illustration shows a schematic diagram of an image segmentation apparatus provided in an exemplary embodiment of this application. The image segmentation apparatus can be implemented as all or part of a device through software, hardware, or a combination of both. The image segmentation apparatus includes a data acquisition module 501, a first segmentation module 502, an image cropping module 503, a second segmentation module 504, and a result fusion module 505.
[0139] Data acquisition module 501 is used to acquire CTA image data; The first segmentation module 502 is used to perform a first bone segmentation process on the image data to obtain a first bone segmentation result; The image cropping module 503 is used to crop the image data according to the first bone segmentation result to obtain first cropped image data that only includes the area above the shoulders; The second segmentation module 504 is used to perform second bone segmentation processing on the first cropped image data to obtain the second bone segmentation result. The result fusion module 505 is used to fuse the first bone segmentation result and the second bone segmentation result to obtain the fused bone result.
[0140] In one or more embodiments, the image cropping module 503 includes: The first trimming unit is used to traverse the first bone segmentation result along the Z-axis, compare the span of the first bone segmentation result in the adjacent layers in the X-axis direction, and obtain the first Z-axis coordinate corresponding to the span growth meeting the preset conditions. The second cropping unit is used to crop the image data by taking the first Z-axis coordinate as the upper edge position of the shoulder, so as to obtain the first cropped image data that only includes the area above the shoulder.
[0141] In one or more embodiments, the second trimming unit includes: The first clipping subunit is used to take the first Z-axis coordinate as the upper edge position of the shoulder and extract the target bone segmentation result corresponding to the bone structure of the area above the shoulder from the first bone segmentation result. The second trimming subunit is used to expand the spatial boundary of the region where the target bone segmentation result is located to obtain the trimming boundary; The third cropping subunit is used to crop the image data based on the cropping boundary to obtain first cropped image data that only includes the area above the shoulders.
[0142] In one or more embodiments, the first segmentation module 502 includes: The cervical spine segmentation module is used to perform the first bone segmentation processing on the image data to obtain the first bone segmentation result and the cervical spine bone result; Image segmentation apparatus, comprising: The second cropping module is used to crop the image data based on the cervical vertebrae results and the first bone segmentation results to obtain second cropped image data including the target region. The blood vessel segmentation module is used to perform blood vessel segmentation processing on the second cropped image data to obtain blood vessel segmentation results; The vessel removal module is used to perform vessel removal processing on the fused bone result based on the vessel segmentation result to obtain the target bone result.
[0143] In one or more embodiments, the second trimming module includes: The third trimming subunit is used to take the minimum Z-axis coordinate in the cervical vertebrae results as the reference position, move a preset number of voxels downward along the direction from the top of the head to the bottom to obtain the second Z-axis coordinate, and use the second Z-axis coordinate as the lower boundary of the Z-axis of the target area. The fourth trimming subunit is used to determine the spatial boundary of the target area based on the first bone segmentation result and the second Z-axis coordinate; The fifth cropping subunit is used to crop the image data according to the spatial boundary of the target region to obtain second cropped image data including the target region.
[0144] In one or more embodiments, the second segmentation module 504 includes: A global extraction unit is used to extract global features from the first cropped image data to obtain a global feature image; The local extraction unit is used to obtain a local image from a local region in the first cropped image data, and to obtain a local feature image that is spatially aligned with the local region from the global feature image; The collage processing unit is used to collage the local image and the local feature image to obtain a collage image, and input the collage image into the local information extraction model for inference to obtain the second bone segmentation result.
[0145] In one or more embodiments, the result fusion module 505 includes: The result fusion unit is used to perform fusion processing on the first bone segmentation result and the second bone segmentation result based on the first weight corresponding to the first bone segmentation result and the second weight corresponding to the second bone segmentation result to obtain a fused bone result; wherein the second weight is greater than the first weight.
[0146] In this application, image data is cropped based on the first bone segmentation result without manual annotation or fixed coordinates. The cropping range dynamically adapts to subjects with different body types and scanning ranges, effectively preserving image data of the head and neck bone region above the shoulders, eliminating interference from the chest and abdominal structures, and providing input information focused on the head and neck bone structure for the next stage of bone segmentation processing. Furthermore, a second bone segmentation process is performed on the first cropped image data to obtain a second bone segmentation result, which is then fused to obtain a fused bone result. This application cascades two bone segmentation processes with different accuracies and inputs. The first stage locates a small region of interest in the image data, and the second stage performs high-precision bone segmentation on the small cropped image. The two bone segmentation results are then fused to obtain a more accurate fused bone result, significantly improving the segmentation accuracy of head and neck bone tissue. This avoids problems such as missed or incorrect segmentation caused by receptive field or resolution limitations in single-stage bone segmentation processing, and can greatly improve inference efficiency and data generalization ability.
[0147] It should be noted that the image segmentation apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when performing the image segmentation method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the image segmentation apparatus and the image segmentation method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.
[0148] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0149] This application also provides a computer storage medium that can store multiple instructions, which are adapted to be loaded and executed by a processor as described above. Figure 1 - Figure 13 The image segmentation method of the illustrated embodiment can be found in the following documentation for its specific execution process: Figure 1 - Figure 13 The specific details of the illustrated embodiments will not be elaborated here.
[0150] This application also provides a computer program product that stores at least one instruction, which is loaded and executed by a processor as described above. Figure 1 - Figure 13 The image segmentation method of the illustrated embodiment can be found in the following documentation for its specific execution process: Figure 1 - Figure 13 The specific details of the illustrated embodiments will not be elaborated here.
[0151] Please see Figure 15This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 15 As shown, the electronic device 700 may include: at least one processor 701, at least one network interface 704, a user interface 703, a memory 705, and at least one communication bus 702.
[0152] The communication bus 702 is used to enable communication between these components.
[0153] The user interface 703 may include a display screen and a camera. Optionally, the user interface 703 may also include a standard wired interface and a wireless interface.
[0154] The network interface 704 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0155] The processor 701 may include one or more processing cores. The processor 701 connects to various parts within the electronic device 700 using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 705, and by calling data stored in the memory 705. Optionally, the processor 701 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 701 may integrate one or a combination of several of the following: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 701 and may be implemented as a separate chip.
[0156] The memory 705 may include random access memory (RAM) or read-only memory. Optionally, the memory 705 may include a non-transitory computer-readable storage medium. The memory 705 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 705 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 705 may also be at least one storage device located remotely from the aforementioned processor 701. Figure 15 As shown, the memory 705, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an image segmentation application.
[0157] exist Figure 15 In the illustrated electronic device 700, the user interface 703 is mainly used to provide an input interface for the user and to acquire user input data; while the processor 701 can be used to call the image segmentation application stored in the memory 705 and specifically perform the following operations: Acquire CTA image data; The image data is subjected to a first bone segmentation process to obtain a first bone segmentation result; The image data is cropped based on the first bone segmentation result to obtain first cropped image data that only includes the area above the shoulders. The first cropped image data is subjected to a second bone segmentation process to obtain a second bone segmentation result. The first bone segmentation result and the second bone segmentation result are fused to obtain the fused bone result.
[0158] In one embodiment, processor 701 performs the cropping of the image data based on the first bone segmentation result to obtain first cropped image data that includes only the area above the shoulders. Specifically, the following is executed: Traverse the first bone segmentation results along the Z-axis, compare the span of the first bone segmentation results in adjacent layers in the X-axis direction, and obtain the first Z-axis coordinate corresponding to the span growth meeting the preset conditions; Using the first Z-axis coordinate as the position of the upper edge of the shoulder, the image data is cropped to obtain a first cropped image data that only includes the area above the shoulder.
[0159] In one embodiment, processor 701 executes the step of using the first Z-axis coordinate as the upper edge position of the shoulder to crop the image data, obtaining first cropped image data that only includes the area above the shoulder. Specifically, the following steps are performed: Using the first Z-axis coordinate as the upper edge position of the shoulder, the target bone segmentation result corresponding to the bone structure above the shoulder is extracted from the first bone segmentation result. The spatial boundary of the region where the target bone segmentation result is located is expanded to obtain the clipping boundary; The image data is cropped based on the cropping boundary to obtain a first cropped image data that includes only the area above the shoulders.
[0160] In one embodiment, the processor 701 performs the first bone segmentation process on the image data to obtain a first bone segmentation result, specifically executing: The image data is subjected to a first bone segmentation process to obtain the first bone segmentation result and the cervical vertebrae result; After processor 701 performs the fusion processing on the first bone segmentation result and the second bone segmentation result to obtain the fused bone result, it further performs: Based on the cervical vertebrae results and the first bone segmentation results, the image data is cropped to obtain second cropped image data including the target region. The second cropped image data is subjected to blood vessel segmentation processing to obtain blood vessel segmentation results; Based on the vessel segmentation results, the fused bone result is subjected to vessel removal processing to obtain the target bone result.
[0161] In one embodiment, processor 701 performs the step of cropping the image data based on the cervical vertebrae results and the first bone segmentation results to obtain second cropped image data including the target region, specifically: Using the minimum Z-axis coordinate in the cervical vertebrae results as a reference position, a preset number of voxels are moved downwards along the direction from the top of the head to the bottom to obtain the second Z-axis coordinate, and the second Z-axis coordinate is used as the lower boundary of the Z-axis of the target area. Based on the first bone segmentation result and the second Z-axis coordinate, determine the spatial boundary of the target region; The image data is cropped according to the spatial boundary of the target region to obtain a second cropped image data including the target region.
[0162] In one embodiment, processor 701 performs the second bone segmentation process on the first cropped image data to obtain a second bone segmentation result, specifically executing: Extract the global features from the first cropped image data to obtain a global feature image; A local image is obtained by acquiring a local region in the first cropped image data, and a local feature image that is spatially aligned with the local region is obtained from the global feature image; The local image and the local feature image are stitched together to obtain a stitched image. The stitched image is then input into a local information extraction model for inference to obtain a second bone segmentation result.
[0163] In one embodiment, the processor 701 performs the fusion processing of the first bone segmentation result and the second bone segmentation result to obtain a fused bone result, specifically: Based on the first weight corresponding to the first bone segmentation result and the second weight corresponding to the second bone segmentation result, the first bone segmentation result and the second bone segmentation result are fused to obtain a fused bone result; wherein, the second weight is greater than the first weight.
[0164] In this application, image data is cropped based on the first bone segmentation result without manual annotation or fixed coordinates. The cropping range dynamically adapts to subjects with different body types and scanning ranges, effectively preserving image data of the head and neck bone region above the shoulders, eliminating interference from the chest and abdominal structures, and providing input information focused on the head and neck bone structure for the next stage of bone segmentation processing. Furthermore, a second bone segmentation process is performed on the first cropped image data to obtain a second bone segmentation result, which is then fused to obtain a fused bone result. This application cascades two bone segmentation processes with different accuracies and inputs. The first stage locates a small region of interest in the image data, and the second stage performs high-precision bone segmentation on the small cropped image. The two bone segmentation results are then fused to obtain a more accurate fused bone result, significantly improving the segmentation accuracy of head and neck bone tissue. This avoids problems such as missed or incorrect segmentation caused by receptive field or resolution limitations in single-stage bone segmentation processing, and can greatly improve inference efficiency and data generalization ability.
[0165] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented. Each of the above methods can be executed by a computer program instructing related hardware. The program corresponding to each method can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The storage medium of the electronic device 700 can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.
[0166] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0167] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application are still within the scope of this application.
Claims
1. An image segmentation method, characterized in that, The method includes: Acquire CTA image data; The image data is subjected to a first bone segmentation process to obtain a first bone segmentation result; The image data is cropped based on the first bone segmentation result to obtain first cropped image data that only includes the area above the shoulders. The first cropped image data is subjected to a second bone segmentation process to obtain a second bone segmentation result. The first bone segmentation result and the second bone segmentation result are fused to obtain the fused bone result.
2. The image segmentation method according to claim 1, characterized in that, The step of cropping the image data based on the first bone segmentation result to obtain first cropped image data that includes only the area above the shoulders includes: Traverse the first bone segmentation results along the Z-axis, compare the span of the first bone segmentation results in adjacent layers in the X-axis direction, and obtain the first Z-axis coordinate corresponding to the span growth meeting the preset conditions; Using the first Z-axis coordinate as the position of the upper edge of the shoulder, the image data is cropped to obtain a first cropped image data that only includes the area above the shoulder.
3. The image segmentation method according to claim 2, characterized in that, The step of cropping the image data by using the first Z-axis coordinate as the upper edge position of the shoulder to obtain first cropped image data that only includes the area above the shoulder includes: Using the first Z-axis coordinate as the upper edge position of the shoulder, the target bone segmentation result corresponding to the bone structure above the shoulder is extracted from the first bone segmentation result. The spatial boundary of the region where the target bone segmentation result is located is expanded to obtain the clipping boundary; The image data is cropped based on the cropping boundary to obtain a first cropped image data that includes only the area above the shoulders.
4. The image segmentation method according to claim 1, characterized in that, The first bone segmentation process on the image data to obtain the first bone segmentation result includes: The image data is subjected to a first bone segmentation process to obtain the first bone segmentation result and the cervical vertebrae result; After fusing the first bone segmentation result and the second bone segmentation result to obtain the fused bone result, the process further includes: Based on the cervical vertebrae results and the first bone segmentation results, the image data is cropped to obtain second cropped image data including the target region. The second cropped image data is subjected to blood vessel segmentation processing to obtain blood vessel segmentation results; Based on the vessel segmentation results, the fused bone result is subjected to vessel removal processing to obtain the target bone result.
5. The image segmentation method according to claim 4, characterized in that, The step of cropping the image data based on the cervical vertebrae results and the first bone segmentation results to obtain second cropped image data including the target region includes: Using the minimum Z-axis coordinate in the cervical vertebrae results as a reference position, a preset number of voxels are moved downwards along the direction from the top of the head to the bottom to obtain the second Z-axis coordinate, and the second Z-axis coordinate is used as the lower boundary of the Z-axis of the target area. Based on the first bone segmentation result and the second Z-axis coordinate, determine the spatial boundary of the target region; The image data is cropped according to the spatial boundary of the target region to obtain a second cropped image data including the target region.
6. The image segmentation method according to claim 1, characterized in that, The second bone segmentation process on the first cropped image data to obtain the second bone segmentation result includes: Extract the global features from the first cropped image data to obtain a global feature image; A local image is obtained by acquiring a local region in the first cropped image data, and a local feature image that is spatially aligned with the local region is obtained from the global feature image; The local image and the local feature image are stitched together to obtain a stitched image. The stitched image is then input into a local information extraction model for inference to obtain a second bone segmentation result.
7. The image segmentation method according to claim 1, characterized in that, The process of fusing the first bone segmentation result and the second bone segmentation result to obtain the fused bone result includes: Based on the first weight corresponding to the first bone segmentation result and the second weight corresponding to the second bone segmentation result, the first bone segmentation result and the second bone segmentation result are fused to obtain a fused bone result; wherein, the second weight is greater than the first weight.
8. An image segmentation apparatus, characterized in that, The device includes: The data acquisition module is used to acquire CTA image data; The first segmentation module is used to perform a first bone segmentation process on the image data to obtain a first bone segmentation result; The image cropping module is used to crop the image data according to the first bone segmentation result to obtain first cropped image data that only includes the area above the shoulders; The second segmentation module is used to perform second bone segmentation processing on the first cropped image data to obtain the second bone segmentation result. The result fusion module is used to fuse the first bone segmentation result and the second bone segmentation result to obtain the fused bone result.
9. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions, which are adapted to be loaded by a processor and executed as method steps as claimed in any one of claims 1 to 7.
10. An electronic device, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed the method steps as claimed in any one of claims 1 to 7.