A nasopharyngeal carcinoma whole-fraction MVCT target region and organ at risk segmentation method
By constructing a structural prior atlas and a planned CT-guided spatiotemporal Transformer network, the problem of discontinuity in MVCT image delineation was solved, achieving accurate, continuous, and stable segmentation of the entire target area and organs at risk in nasopharyngeal carcinoma, thus improving the clinical usability of segmentation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 江西省肿瘤医院(江西省第二人民医院 江西省癌症中心)
- Filing Date
- 2026-07-03
- Publication Date
- 2026-07-31
AI Technical Summary
In the current technology for radiotherapy of nasopharyngeal carcinoma, the low soft tissue contrast and high noise of MVCT images result in a large workload for delineating the target area and organs at risk, and it is difficult to ensure the consistency of the delineation. Furthermore, the existing automatic segmentation methods fail to effectively utilize the temporal continuity of multiple MVCT images of the patient and the high-quality structural information of the planned CT images, resulting in discontinuous segmentation results and displacement of organs at risk.
By constructing a structural prior map, utilizing the target area and organ-at-risk contours on the planned CT images, and combining them with MVCT time-series data, a spatiotemporal Transformer segmentation network guided by the planned CT structural prior is used to perform cross-domain adaptation and adaptive image quality enhancement, generating a direct segmentation probability map and label map for the current MVCT images, and optimizing the segmentation results through short-range and long-range propagation results.
It achieves accurate, continuous, and stable segmentation of the target area and organs at risk in nasopharyngeal carcinoma using full-segment MVCT, reducing reliance on manual delineation and improving the clinical usability and consistency of segmentation.
Smart Images

Figure CN122492734A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image segmentation technology, and specifically relates to a method for segmenting the target area and organs at risk in nasopharyngeal carcinoma using full fractional MVCT. Background Technology
[0002] Nasopharyngeal carcinoma is one of the most common malignant tumors of the head and neck, and radiotherapy is the main treatment for it. Because the primary lesion of nasopharyngeal carcinoma is usually located near the skull base, nasopharynx, and parapharyngeal space, and is adjacent to important organs at risk such as the brainstem, spinal cord, parotid gland, temporal lobe, optic nerve, optic chiasm, eyeball, lens, cochlea, mandible, oral cavity, larynx, pharyngeal constrictor muscles, thyroid gland, and brachial plexus, accurate delineation of the target area and organs at risk is crucial for developing a radiotherapy plan, ensuring target dose coverage, and reducing the risk of complications in normal tissues.
[0003] In nasopharyngeal carcinoma radiotherapy, planning CT images are commonly used for treatment planning. Physicians manually delineate the target area and organs at risk based on the planning CT images, combined with MRI, PET / CT, or clinical examination results. The target area may include PTV1, PTV2, PTVnx, PTVnd, or GTVnx, GTVnd, CTV1, CTV2, or other target areas defined by the clinical treatment protocol. Planning CT images and the physician's manually delineated outlines have high clinical reliability and are crucial for radiotherapy planning and dose assessment.
[0004] Equipment such as helical tomotherapy systems can acquire MVCT images during treatment for patient positioning verification and anatomical observation. Compared to planning CT, MVCT images typically suffer from lower soft tissue contrast, higher noise, unclear boundaries, and unstable grayscale distribution. This results in a massive workload for manually delineating the target area and organs at risk on MVCT images fraction by fraction, and ensuring consistency in delineation is difficult. Nasopharyngeal carcinoma patients usually receive approximately 30 fractions of radiotherapy, although the actual number of fractions may vary depending on the treatment plan. Requiring physicians to delineate the complete target area and organs at risk on every single MVCT image fraction would significantly increase the clinical workload.
[0005] Existing automatic segmentation methods are mostly trained on single-phase images such as planning CT, CBCT, or MRI, typically treating each image as an independent sample for segmentation, ignoring the temporal continuity between multiple MVCT images of the same patient during radiotherapy. In fact, nasopharyngeal carcinoma patients may experience dynamic changes during treatment, including tumor volume changes, parotid gland shrinkage, weight loss, changes in external contour, changes in local edema, tissue displacement, and positioning errors. These changes exhibit certain patterns in consecutive fractionated MVCT images. If only a single MVCT image is segmented independently, the model cannot fully utilize the patient's previous fractionated images and pre-treatment structural information, easily leading to problems such as discontinuities in segmentation results between adjacent fractions, abrupt changes in target area boundaries, and abnormal displacement of endangered organs.
[0006] Furthermore, the difficulty in obtaining high-quality manual annotations for MVCT images on a large scale limits the training and application of deep learning models in automatic MVCT segmentation. Although planning CT images and surgeon-manually drawn contours can provide high-quality structural information, planning CT and MVCT belong to different imaging domains and differ in image grayscale, noise characteristics, spatial resolution, and tissue contrast. How to effectively transfer surgeon contours from planning CT images for automatic segmentation of full-segmentation MVCT remains a key problem that needs to be solved. Summary of the Invention
[0007] Based on this, the present invention provides a method for segmenting the target area and organs at risk in full fractionation MVCT of nasopharyngeal carcinoma. The method aims to make full use of the manual outlines of the planning CT physician, the initial structural information of MVCT before the first treatment, the temporal change pattern of subsequent full fractionation MVCT, and unlabeled MVCT data for automatic segmentation, so as to improve the accuracy, continuity, stability and clinical usability of segmentation of the target area and organs at risk in full fractionation MVCT of nasopharyngeal carcinoma.
[0008] A first aspect of this invention provides a method for segmenting the target region and organs at risk in nasopharyngeal carcinoma using full fractionation MVCT, the method comprising: Acquire planned CT images, target areas and organs at risk delineated on the planned CT images, and full fractionated MVCT images, and construct a structural prior atlas based on the planned CT images, target areas, and organs at risk; The target area and organs at risk delineated on the planned CT images are transformed into initial priors in the MVCT space to obtain initial structural labels; Image quality adaptive enhancement and cross-domain adaptation are performed on full-segment MVCT images, and the enhanced full-segment MVCT images are constructed into full-segment MVCT time series data according to the segment order. The short-range propagation results and long-range propagation results are determined. The short-range propagation results are determined based on the segmentation results obtained in adjacent treatment fractions, while the long-range propagation results are determined based on the initial structural labels. A spatiotemporal Transformer segmentation network guided by planned CT structural priors is constructed, including an MVCT spatial encoder, a structural prior encoder, a long temporal memory module, a spatiotemporal Transformer module, an anatomical relationship constraint module, and a multi-structure segmentation decoder. Using the full-sequence MVCT temporal data, the structural prior atlas, and the initial structural labels, a direct segmentation probability map and a direct segmentation label map of the current MVCT image are generated. Based on the short-range propagation results, long-range propagation results, and the direct segmentation label map, the final segmentation result for the current segment is obtained.
[0009] Furthermore, the step of obtaining the final segmentation result for the current segment based on the short-range propagation results, long-range propagation results, and the direct segmentation label map includes: Calculate the uncertainty value based on the direct segmentation probability diagram, short-range propagation results, and long-range propagation results; The uncertainty value is compared with a threshold, and the final segmentation result of the current segment is subjected to hierarchical processing. Specifically, based on the comparison results of the uncertainty value with at least two preset thresholds, the final segmentation result of the current segmentation is subject to three levels of differentiated processing: When the uncertainty value meets the first preset condition, the segmentation result is automatically accepted; When the uncertainty value meets the second preset condition, the segmentation result is automatically optimized and corrected based on multi-source prior information and historical segmentation information. When the uncertainty value meets the third preset condition, the structure or region corresponding to the segmentation result is marked as a manually reviewed region.
[0010] Furthermore, in the step of constructing a structural priori map based on the planned CT images, target area, and organs at risk, For the k-th structure of the i-th patient, its structural prior representation is as follows: ; in, The label map represents the k-th structure of the i-th patient. This represents the structural volume of the k-th structure in the i-th patient. The structural centroid of the k-th structure in the i-th patient is represented. This represents the structural boundary of the k-th structure in the i-th patient. This represents the structural shape characteristics of the k-th structure in the i-th patient; For any two structures k and l, calculate their spatial relationship: ; in, This represents the distance between the centroids of the k-th and l-th structures in the i-th patient. This represents the relative orientation between the k-th and l-th structures in the i-th patient. This represents the proximity relationship between the k-th structure and the l-th structure in the i-th patient. This indicates the inclusion, intersection, separation, or spatial constraint relationship between the k-th structure and the l-th structure of the i-th patient; Thus, a structural prior map is constructed: ; in, This represents the structural prior map of the i-th patient constructed based on the planned CT scan. This represents the anatomical prior information of the k-th structure in the i-th patient. This represents the spatial relationship between the k-th structure and the l-th structure of the i-th patient.
[0011] Furthermore, in the step of converting the target area and organs at risk delineated on the planned CT image into initial priors in the MVCT space to obtain initial structural labels, Cross-modal registration of the planned CT image and the MVCT0 image yields the spatial transformation relationship from the planned CT space to the MVCT0 space, represented as follows: ; MVCT0 images refer to MVCT images acquired before the first treatment, and pCT refers to the planning CT space. For MVCT0 space; Let the planned CT images of the i-th patient be... The target area and the outline of organs at risk are delineated on the planned CT images. Then, the target area and organ at risk contours on the planned CT image are mapped to the MVCT0 space through the spatial transformation relationship to obtain the initial structural labels on the MVCT0 image:
[0012] ; Where K represents the total number of structures in the target area and organs at risk. This represents the initial target area and organ at risk labels on the MVCT0 image of the i-th patient, i.e., the initial structural labels.
[0013] Furthermore, in the step of performing image quality adaptive enhancement and cross-domain adaptation on the full-segment MVCT images, and constructing the enhanced full-segment MVCT images into full-segment MVCT time-series data according to the segment order, First, an image quality score is calculated for each fraction of the full-fraction MVCT image, expressed as: ; in, This represents the quality score of the t-th fractional MVCT image. This represents an image quality evaluation function, which is calculated based on one or more of the following: noise level, local contrast, edge sharpness, gray-level distribution stability, and structural boundary visibility. This represents the MVCT image corresponding to the t-th treatment fraction for the i-th patient; Subsequently, based on the image quality score, the corresponding enhancement strategy is selected or adjusted, including grayscale normalization, noise suppression, edge-preserving filtering, local contrast enhancement, structural boundary enhancement, and artifact suppression. The enhanced MVCT image is represented as follows: ; in, This represents the t-th fractional MVCT image after enhancement. Represents the image enhancement function; The enhanced MVCT0 to MVCTn images are constructed into full-fraction MVCT time-series data according to the treatment fractionation order, and represented as follows: .
[0014] Furthermore, in the step of determining the short-range propagation result and the long-range propagation result, Image registration is performed on adjacent MVCT images to obtain the spatial transformation field from the (t-1)th fraction to the tth fraction, expressed as: ; Using this spatial transformation field, the final segmentation result set obtained in the previous segment is propagated to the current segment: ; in, This represents the short-range transmission result of the i-th patient from the previous transmission to the current t-th transmission. This represents the set of final segmentation results obtained for the i-th patient in the (t-1)th segment; The initial structural label is propagated directly or through a cascade of multiple adjacent fractional transformation fields to the t-th fraction: ; in, This represents the long-range propagation result obtained from the initial structural tag of patient i to the t-th segment. This represents the long-range spatial transformation from the MVCT0 image to the t-th MVCT image, which is obtained by cascading multiple adjacent transformation fields.
[0015] Furthermore, the MVCT spatial encoder is used to extract three-dimensional spatial features from each enhanced MVCT image, as follows: ; in, Represents a three-dimensional spatial encoder. Represents the spatial features of the t-th fractional MVCT image; The structural prior encoder is used to encode the structural prior information consisting of the structural prior map constructed based on the planned CT and the initial structural labels, and represents it as follows: ; in, This represents a structural prior encoder. Represents prior structural features; In the long-term memory module, a dynamic memory bank is established, represented as follows: ; in, This represents the dynamic memory bank of the i-th patient up to the t-th fraction. This refers to the memory item of the (t-1)th historical segment in the dynamic memory bank; For the current segment t, the most relevant historical structural information about the current image is extracted from the dynamic memory using an attention retrieval mechanism: ; in, Indicates historical memory retrieval features, Represents the attention function; The spatiotemporal Transformer module is used to acquire the spatial features of the full fractional MVCT of the same patient and learn the long-range dependencies between different treatment fractions, represented as: ; in, This refers to the spatiotemporal Transformer module. This represents the spatiotemporal context feature corresponding to the t-th fraction; The anatomical relationship constraint module is used to construct an anatomical relationship diagram based on the spatial relationships between the target area and organs at risk in the prior structural atlas, as shown below:
[0016] Among them, nodes Indicates the target area and organs at risk, border Indicates the distance, direction, proximity, containment, separation, or other spatial relationships between structures; The spatial features, structural prior features, historical memory retrieval features, and spatiotemporal context features of the current fractional MVCT images are fused and represented as follows: ; in, This indicates the spatiotemporal feature fusion module. Indicates the characteristics after fusion; The fused features are input into the multi-structure segmentation decoder to obtain the direct segmentation probability map of the current fractional MVCT image, represented as: ; And obtain the direct segmentation label map from the direct segmentation probability map: ; in, This represents a multi-structure segmentation decoder. This represents the probability map of direct segmentation of the t-th fractional MVCT image of the i-th patient. This represents the direct segmentation label map obtained from the corresponding direct segmentation probability map. This represents the maximum probability index function.
[0017] A second aspect of this invention provides a nasopharyngeal carcinoma full-fraction MVCT target region and organ-at-risk segmentation system for implementing the nasopharyngeal carcinoma full-fraction MVCT target region and organ-at-risk segmentation method provided in the first aspect. The system includes: The first acquisition module is used to acquire the planned CT image, the target area and organs at risk delineated on the planned CT image, and the full fraction MVCT image, and to construct a structural prior atlas based on the planned CT image, the target area and organs at risk. The transformation module is used to transform the target area and organs at risk delineated on the planned CT image into initial priors in the MVCT space to obtain initial structural labels; The first construction module is used to perform adaptive image quality enhancement and cross-domain adaptation on the full-segment MVCT images, and to construct the enhanced full-segment MVCT images into full-segment MVCT time-series data according to the segment order. The determination module is used to determine the short-range propagation results and the long-range propagation results. The short-range propagation results are determined based on the segmentation results obtained in adjacent treatment fractions, while the long-range propagation results are determined based on the initial structure labels. The second construction module is used to construct a spatiotemporal Transformer segmentation network guided by the prior structure of the planned CT, including an MVCT spatial encoder, a structural prior encoder, a long temporal memory module, a spatiotemporal Transformer module, an anatomical relationship constraint module, and a multi-structure segmentation decoder. It generates a direct segmentation probability map and a direct segmentation label map of the current MVCT image through the full-segment MVCT temporal data, the structural prior map, and the initial structural labels. The second acquisition module is used to obtain the final segmentation result of the current segmentation based on the short-range propagation result, the long-range propagation result, and the direct segmentation label map.
[0018] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the nasopharyngeal carcinoma full-fraction MVCT target region and organ at risk segmentation method provided in the first aspect.
[0019] A fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the program to implement the nasopharyngeal carcinoma full-fraction MVCT target region and organ at risk segmentation method provided in the first aspect.
[0020] This invention provides a method for segmenting the target region and organs at risk in nasopharyngeal carcinoma using multi-stage MVCT. This method constructs a structural prior map using pre-treatment planning CT images and manually drawn target region and organ at risk contours. Through cross-modal registration between the planning CT images and the first pre-treatment MVCT0 images, the manually drawn contours on the planning CT images are transferred to the MVCT0 space to obtain initial MVCT0 structural labels. Furthermore, the MVCT0 to MVCTn images are constructed into full-stage MVCT temporal data according to the treatment stage sequence. A spatiotemporal deep learning network guided by the planning CT structural prior generates a direct segmentation probability map and a direct segmentation label map for the current stage MVCT image. Based on short-range propagation results, long-range propagation results, and the direct segmentation label map, the final segmentation result for the current stage is obtained. This achieves automatic segmentation of the target region and organs at risk on all stages of MVCT images, ensuring the accuracy, continuity, stability, and clinical usability of the nasopharyngeal carcinoma full-stage MVCT target region and organ at risk segmentation. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating the implementation of a method for segmenting the target region and organs at risk in nasopharyngeal carcinoma using full fractional MVCT, as provided in Embodiment 1 of the present invention. Figure 2 This is a structural block diagram of a nasopharyngeal carcinoma full-fraction MVCT target region and organ at risk segmentation system provided in Embodiment 2 of the present invention; Figure 3 This is a structural block diagram of an electronic device provided in Embodiment 3 of the present invention. Detailed Implementation
[0022] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.
[0023] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0025] Example 1 Please see Figure 1 , Figure 1 The following is a flowchart illustrating the implementation of a method for segmenting the target region and organs at risk in a full fractional MVCT of nasopharyngeal carcinoma according to Embodiment 1 of the present invention. The method for segmenting the target region and organs at risk in a full fractional MVCT of nasopharyngeal carcinoma specifically includes steps S01 to S06.
[0026] Step S01: Acquire the planned CT image, the target area and organs at risk delineated on the planned CT image, and the full fractional MVCT image, and construct a structural prior atlas based on the planned CT image, the target area and organs at risk.
[0027] Specifically, the target area and organs at risk are manually delineated on the planned CT images. The target area includes one or more of PTV1, PTV2, PTVnx, and PTVnd, and may also include GTVnx, GTVnd, CTV1, CTV2, or other nasopharyngeal carcinoma radiotherapy target areas defined by the clinical treatment protocol. Organs at risk include one or more of the following: brainstem, spinal cord, left parotid gland, right parotid gland, left temporal lobe, right temporal lobe, left optic nerve, right optic nerve, optic chiasm, left eyeball, right eyeball, left lens, right lens, left cochlea, right cochlea, mandible, temporomandibular joint, oral cavity, larynx, pharyngeal constrictor muscle, thyroid gland, brachial plexus, and body surface contour.
[0028] Acquire MVCT images of the same patient before the first treatment and MVCT images acquired in subsequent treatment fractions. The MVCT image acquired before the first treatment is designated MVCT0, and the MVCT images acquired in subsequent treatment fractions are designated MVCT1 to MVCTn. Due to differences in treatment protocols and actual acquisition conditions among patients, the total number of MVCT fractions may vary, typically around 30 fractions, but may also be other fractions determined by the clinical treatment protocol.
[0029] For the i-th patient, the MVCT time-series image is represented as follows: ; in, This represents the MVCT0 image of the i-th patient. This represents the MVCT image corresponding to the t-th treatment fraction for the i-th patient.
[0030] Furthermore, for the k-th structure of the i-th patient, its structural prior representation is as follows: ; in, The label map represents the k-th structure of the i-th patient. This represents the structural volume of the k-th structure in the i-th patient. The structural centroid of the k-th structure in the i-th patient is represented. This represents the structural boundary of the k-th structure in the i-th patient. This represents the structural shape characteristics of the k-th structure in the i-th patient; For any two structures k and l, calculate their spatial relationship: ; in, This represents the distance between the centroids of the k-th and l-th structures in the i-th patient. This represents the relative orientation between the k-th and l-th structures in the i-th patient. This represents the proximity relationship between the k-th structure and the l-th structure in the i-th patient. This indicates the inclusion, intersection, separation, or spatial constraint relationship between the k-th structure and the l-th structure of the i-th patient; Thus, a structural prior map is constructed: ; in, This represents the structural prior map of the i-th patient constructed based on the planned CT scan. This represents the anatomical prior information of the k-th structure in the i-th patient. This represents the spatial relationship between the k-th structure and the l-th structure of the i-th patient.
[0031] This prior structural atlas includes not only labels for each target region and organ at risk, but also the volume, centroid, boundary, shape descriptor, and spatial relationships between structures. It is used for prior guidance, anatomical constraints, abnormal contour recognition, and segmentation result correction in subsequent MVCT segmentation.
[0032] Step S02 involves converting the target area and organs at risk outlined on the planned CT image into initial priors in the MVCT space to obtain initial structural labels.
[0033] It should be noted that cross-modal registration of the planned CT image and the MVCT0 image yields the spatial transformation relationship from the planned CT space to the MVCT0 space, which is represented as follows: ; MVCT0 images refer to MVCT images acquired before the first treatment, and pCT refers to the planning CT space. For MVCT0 space; Let the planned CT images of the i-th patient be... The target area and the outline of organs at risk are delineated on the planned CT images. Then, the target area and organ at risk contours on the planned CT image are mapped to the MVCT0 space through the spatial transformation relationship to obtain the initial structural labels on the MVCT0 image:
[0034] ; Where K represents the total number of structures in the target area and organs at risk. This represents the initial target area and organ at risk labels on the MVCT0 image of the i-th patient, i.e., the initial structural labels.
[0035] In this embodiment of the invention, the initial structural labels of the migrated MVCT0 are evaluated for registration quality. Registration quality evaluation includes image similarity assessment, structural boundary offset assessment, local deformation smoothness assessment, volume change rationality assessment, or rapid manual verification. For regions with high registration quality, the migrated labels are directly used as initial supervisory labels; for regions with low registration quality, morphological correction, boundary smoothing, connected component filtering, or rapid manual correction are performed.
[0036] Through the above steps, the high-quality contours drawn in the planned CT images are transformed into initial structural priors in the MVCT space, reducing the reliance on manual step-by-step drawing in MVCT.
[0037] Step S03: Perform adaptive image quality enhancement and cross-domain adaptation on the full-segment MVCT images, and construct the enhanced full-segment MVCT images into full-segment MVCT time-series data according to the segment order.
[0038] Specifically, firstly, an image quality score is calculated for each fraction of the full-fraction MVCT image, expressed as: ; in, This represents the quality score of the t-th fractional MVCT image. This represents an image quality evaluation function, which is calculated based on one or more of the following: noise level, local contrast, edge sharpness, gray-level distribution stability, and structural boundary visibility. This represents the MVCT image corresponding to the t-th treatment fraction for the i-th patient; Subsequently, based on the image quality score, the corresponding enhancement strategy is selected or adjusted, including grayscale normalization, noise suppression, edge-preserving filtering, local contrast enhancement, structural boundary enhancement, and artifact suppression. The enhanced MVCT image is represented as follows: ; in, This represents the t-th fractional MVCT image after enhancement. Represents the image enhancement function; The enhanced MVCT0 to MVCTn images are constructed into full-fraction MVCT time-series data according to the treatment fractionation order, and represented as follows: ; To address the issue of different number of doses for different patients, variable-length sequence input, sequence completion, sequence truncation, or time masking mechanisms are employed. The time mask is represented as follows: ; Where L is the preset maximum sequence length. This indicates that the t-th fractional MVCT truly exists. This indicates whether the segment is a completed segment or a missing segment. The model only calculates the loss and outputs valid segmentation results for actual MVCT segments.
[0039] Step S04: Determine the short-range propagation result and the long-range propagation result, wherein the long-range propagation result is determined based on the initial structure label.
[0040] It should be noted that image registration is performed on adjacent MVCT images to obtain the spatial transformation field from the (t-1)th fraction to the tth fraction, which is represented as: ; Using this spatial transformation field, the final segmentation result set obtained in the previous segment is propagated to the current segment: ; in, This represents the short-range propagation result of the i-th patient from the previous segment to the current t-th segment, i.e., the short-range candidate structure label set. This represents the set of final segmentation results obtained for the i-th patient in the (t-1)th segment; The initial structural label is propagated directly or through a cascade of multiple adjacent fractional transformation fields to the t-th fraction: ; in, This represents the long-range propagation result obtained from the initial structural label to the t-th step for the i-th patient, i.e., the long-range candidate structural label set. This represents the long-range spatial transformation from the MVCT0 image to the t-th MVCT image, which is obtained by cascading multiple adjacent transformation fields.
[0041] Understandably, short-range propagation results primarily reflect local continuous changes between adjacent segments, characterized by short registration spans and relatively small local deformation errors; long-range propagation results mainly retain high-quality anatomical priors from the initial structural labels of MVCT0, reducing potential error accumulation during successive propagation. Both, along with the direct segmentation results from deep learning, serve as candidate information for the final segmentation of the current segment.
[0042] Step S05: Construct a spatiotemporal Transformer segmentation network guided by the planned CT structure prior, including an MVCT spatial encoder, a structure prior encoder, a long temporal memory module, a spatiotemporal Transformer module, an anatomical relationship constraint module, and a multi-structure segmentation decoder. Using the full-series MVCT temporal data, the structure prior map, and the initial structure labels, generate a direct segmentation probability map and a direct segmentation label map for the current MVCT image.
[0043] In this embodiment of the invention, the MVCT spatial encoder is used to extract three-dimensional spatial features from each enhanced MVCT image, as shown below: ; in, Represents a three-dimensional spatial encoder. The spatial features of the t-th fractional MVCT image are represented by the 3D spatial encoder, which can be a 3D U-Net encoder, residual convolutional encoder, Swin Transformer encoder, convolution-Transformer hybrid encoder or other 3D medical image feature extraction network, used to extract the local boundary features, spatial location features and global context features of the nasopharyngeal carcinoma target area and organs at risk. The structural prior encoder is used to encode the structural prior information consisting of the structural prior map constructed based on the planned CT and the initial structural labels, and represents it as follows: ; in, This represents a structural prior encoder. This represents the structural prior features, which are used to provide the model with the initial location, shape, volume, boundaries, and spatial relationships between structures of the target area and organs at risk before treatment. In the long-term memory module, a dynamic memory bank is established, represented as follows: ; in, This represents the dynamic memory bank of the i-th patient up to the t-th fraction. This refers to the memory item of the (t-1)th historical segment in the dynamic memory bank; For the current segment t, the most relevant historical structural information about the current image is extracted from the dynamic memory using an attention retrieval mechanism: ; in, Indicates historical memory retrieval features, This represents the attention function. The long-term memory module is used to adapt to long-term MVCT sequences of approximately 30 fractions for nasopharyngeal carcinoma, enabling the current fractionation to reference the structural changes in the patient's previous fractions; The spatiotemporal Transformer module is used to acquire the spatial features of the full fractional MVCT of the same patient and learn the long-range dependencies between different treatment fractions, represented as: ; in, This refers to the spatiotemporal Transformer module. This represents the spatiotemporal context features corresponding to the t-th fraction. The spatiotemporal Transformer module establishes the correlation between different fractional MVCT images through a self-attention mechanism, used to capture the continuous changing trends of the target area, parotid gland, body surface contours, and other organs at risk during treatment; The anatomical relationship constraint module is used to construct an anatomical relationship diagram based on the spatial relationships between the target area and organs at risk in the prior structural atlas, as shown below: ; Among them, nodes Indicates the target area and organs at risk, border This indicates the distance, direction, proximity, containment, separation, or other spatial relationships between structures. The anatomical relationship constraint module is used to correct predictions that do not conform to anatomical spatial relationships during model training and inference, such as structural breaks, confusion between left and right structures, significant structural positional shifts, abnormal overlaps, or unreasonable spatial relationships between the target area and organs at risk. The spatial features, structural prior features, historical memory retrieval features, and spatiotemporal context features of the current fractional MVCT images are fused and represented as follows: ; in, This indicates the spatiotemporal feature fusion module. Indicates the characteristics after fusion; The fused features are input into the multi-structure segmentation decoder to obtain the direct segmentation probability map of the current fractional MVCT image, represented as: ; And obtain the direct segmentation label map from the direct segmentation probability map: ; in, This represents a multi-structure segmentation decoder. This represents the probability map of direct segmentation of the t-th fractional MVCT image of the i-th patient. This represents the direct segmentation label map obtained from the corresponding direct segmentation probability map. This represents the maximum probability index function.
[0044] Step S06: Based on the short-range propagation results, long-range propagation results, and direct segmentation label map, obtain the final segmentation result for the current segment.
[0045] Specifically, for the t-th fractional MVCT image of the i-th patient, the aforementioned steps obtain three sets of candidate structural labels: a direct segmentation label map generated by the spatiotemporal Transformer segmentation network, a short-range propagation result obtained from the final segmentation result of the previous fraction, and a long-range propagation result obtained from the initial structural labels of MVCT0. The direct segmentation label map primarily reflects local image features and model recognition results in the current fractional MVCT image; the short-range propagation result primarily reflects continuous structural changes between adjacent fractions; and the long-range propagation result primarily retains high-quality anatomical priors from the initial structural labels of MVCT0. To comprehensively utilize these three types of information, this embodiment of the invention employs an adaptive weight fusion method to obtain the final segmentation result of the current fraction: ; in, .
[0046] The fusion weight , , The weighting is adaptively determined based on the current MVCT image quality, model prediction confidence, registration reliability, structural variation magnitude, and uncertainty level. When the current MVCT image quality is high and the model prediction confidence is high, the weight of the direct segmentation result is increased; when the current MVCT image quality is low or the structural boundaries are unclear, the weights of the short-range and long-range propagation results are increased; when the reliability of adjacent segment registration is low, the weight of the short-range propagation result is decreased; when the long-range registration span is large and the deformation uncertainty is high, the weight of the long-range propagation result is decreased.
[0047] Through the aforementioned adaptive fusion, the model can dynamically balance the current image evidence, the continuity of adjacent sub-segments, and the prior knowledge of the initial structure of MVCT0, thereby obtaining the final segmentation results of the target area and organs at risk on the current sub-segment MVCT image.
[0048] Furthermore, in some other embodiments of the present invention, the step of obtaining the final segmentation result of the current segmentation based on the short-range propagation result, the long-range propagation result, and the direct segmentation label map includes: Based on the direct segmentation probability map, short-range propagation results, and long-range propagation results, the uncertainty value is calculated and expressed as: ; in, This represents a multi-structure probability graph that directly segments the output of the model. Represents probability entropy. This indicates the difference between the direct segmentation results and the short-range propagation results. This indicates the difference between the direct segmentation results and the long-range propagation results. and Indicates the weighting coefficient; The uncertainty value is compared with the threshold, and the final segmentation result of the current segment is processed in a hierarchical manner. Specifically, when the uncertainty value meets the first preset condition, the segmentation result is automatically accepted, that is, when the uncertainty is lower than the first threshold, the segmentation result is automatically accepted. When the uncertainty value meets the second preset condition, the segmentation result is automatically optimized and corrected based on multi-source prior information and historical segmentation information. That is, when the uncertainty is between the first threshold and the second threshold, automatic correction is performed using structural prior, adjacent propagation results, MVCT0 long-range propagation results and anatomical relationship constraints. When the uncertainty value meets the third preset condition, the structure or region corresponding to the segmentation result is marked as a manually reviewed region. That is, when the uncertainty is higher than the second threshold, the corresponding structure or local region is marked as a manually reviewed region. It can be understood that the first threshold is denoted as θ1, the second threshold as θ2, and the uncertainty value is normalized to the range of 0 to 1, where 0 ≤ θ1 < θ2 ≤ 1. The value range of θ1 can be 0.20 to 0.30, and the value range of θ2 can be 0.40 to 0.60. In image quality scoring... In high-quality MVCT scenarios with a resolution ≥0.75, θ1 can be taken as 0.30 and θ2 as 0.60; in scenarios with a resolution ≤0.50, θ1 can be taken as 0.30 and θ2 as 0.60. In medium-quality MVCT scenarios with a resolution <0.75, θ1 can be taken as 0.25, and θ2 can be taken as 0.50; In low-quality MVCT scenarios with an uncertainty value < 0.50, θ1 can be set to 0.20 and θ2 to 0.40. Accordingly, when the uncertainty value < θ1, the first preset condition is met and the segmentation result is automatically accepted; when θ1 ≤ uncertainty value < θ2, the second preset condition is met and automatic optimization correction is triggered; when the uncertainty value ≥ θ2, the third preset condition is met and the corresponding structure or local region is marked as a manually reviewed region.
[0049] If the segmentation result is automatically corrected, the corrected label set is denoted as... This is output as the final segmentation result for the current segment; if no correction is triggered, it is retained. This serves as the final segmentation result for the current stage. Understandably, this step reduces erroneous segmentation caused by low-quality MVCT images and improves the safety of clinical applications.
[0050] It should be noted that during the model training phase, for MVCT1 to MVCTn which lack manual annotation, a teacher-student framework is used for iterative pseudo-label updates. Specifically, the teacher model generates pseudo-labels for unannotated MVCTs: ; The student model is trained using pseudo-labels as supervision signals: ; The teacher model parameters are updated from the student model parameters using an exponential moving average: ; in, Indicates the teacher model parameters, Represents the parameters of the student model. The momentum coefficient is represented by Teacher(·), which represents the teacher model, and Student(·) represents the student model.
[0051] To reduce the impact of erroneous false labels, only regions with low uncertainty are selected for false label supervision: ; in, Indicates a reliable pseudo-label area. denoted by the uncertainty threshold, and v represents the voxel position.
[0052] The pseudo-label loss is: ; This represents the voxel-level loss function. In this way, unannotated subsequent fractional MVCT images can participate in model training, improving the model's adaptability to different treatment fractions, different image qualities, and different anatomical changes.
[0053] In addition, during the model training phase, the total loss function for model training includes: ; in, This represents the segmentation supervision loss based on the initial MVCT0 label or available artificial labels; This represents the weakly supervised loss based on reliable pseudo-labels; This represents the temporal consistency loss between adjacent segments; This represents the structural shape constraint loss; This represents the structural boundary constraint loss; This indicates the loss of anatomical constraint. This represents the uncertainty regularization loss; to This represents the weighting coefficient of each loss term.
[0054] The time-series consistency loss can be expressed as: ; The loss of anatomical constraint can be expressed as: ; in, This represents the spatial relationship between structure k and structure l, calculated from the predicted segmentation results. This indicates the spatial relationship between corresponding structures in the prior map of the planned CT structure.
[0055] Understandably, after training, the target area and organs at risk, MVCT0 and MVCT1 to MVCTn, drawn on the planned CT images of the patient to be segmented, are input into the model. The output is the outline of the target area and organs at risk of each MVCT segment after direct segmentation, short-range propagation, long-range propagation, adaptive fusion, and uncertainty correction. The planned target area label in the output is used to indicate the location and extent of the planned target area that matches the patient's anatomical state in the current MVCT space, and the organs at risk label is used to indicate the location and extent of the corresponding normal tissue structure in the current MVCT space.
[0056] In summary, the nasopharyngeal carcinoma full-stage MVCT target region and organ at-risk segmentation method described in the above embodiments of the present invention constructs a structural prior map using pre-treatment planning CT images and manually drawn target region and organ at-risk contours. Through cross-modal registration between the planning CT images and the first pre-treatment MVCT0 images, the manually drawn contours on the planning CT images are transferred to the MVCT0 space to obtain initial MVCT0 structural labels. Furthermore, the MVCT0 images to MVCTn images are constructed into full-stage MVCT temporal data according to the treatment stage sequence. A spatiotemporal deep learning network guided by the planning CT structural prior generates a direct segmentation probability map and a direct segmentation label map for the current stage MVCT image. Based on short-range propagation results, long-range propagation results, and the direct segmentation label map, the final segmentation result for the current stage is obtained, achieving automatic segmentation of the target region and organ at-risk on all stage MVCT images, and ensuring the accuracy, continuity, stability, and clinical usability of nasopharyngeal carcinoma full-stage MVCT target region and organ at-risk segmentation.
[0057] Example 2 Please see Figure 2 , Figure 2 This is a structural block diagram of a nasopharyngeal carcinoma full-fraction MVCT target region and organ-at-risk segmentation system provided in Embodiment 2 of the present invention. The nasopharyngeal carcinoma full-fraction MVCT target region and organ-at-risk segmentation system 200 includes: a first acquisition module 21, a conversion module 22, a first construction module 23, a determination module 24, a second construction module 25, and a second acquisition module 26, wherein: The first acquisition module 21 is used to acquire the planned CT image, the target area and organs at risk delineated on the planned CT image, and the full fractional MVCT image, and to construct a structural prior map based on the planned CT image, the target area, and the organs at risk. For the k-th structure of the i-th patient, its structural prior is represented as follows: ; in, The label map represents the k-th structure of the i-th patient. This represents the structural volume of the k-th structure in the i-th patient. The structural centroid of the k-th structure in the i-th patient is represented. This represents the structural boundary of the k-th structure in the i-th patient. This represents the structural shape characteristics of the k-th structure in the i-th patient; For any two structures k and l, calculate their spatial relationship: ; in, This represents the distance between the centroids of the k-th and l-th structures in the i-th patient. This represents the relative orientation between the k-th and l-th structures in the i-th patient. This represents the proximity relationship between the k-th structure and the l-th structure in the i-th patient. This indicates the inclusion, intersection, separation, or spatial constraint relationship between the k-th structure and the l-th structure of the i-th patient; Thus, a structural prior map is constructed: ; in, This represents the structural prior map of the i-th patient constructed based on the planned CT scan. This represents the anatomical prior information of the k-th structure in the i-th patient. This represents the spatial relationship between the k-th structure and the l-th structure in the i-th patient; Transformation module 22 is used to transform the target area and organs at risk delineated on the planning CT image into initial priors in the MVCT space, obtain initial structure labels, and perform cross-modal registration between the planning CT image and the MVCT0 image to obtain the spatial transformation relationship from the planning CT space to the MVCT0 space, represented as: ; MVCT0 images refer to MVCT images acquired before the first treatment, and pCT refers to the planning CT space. For MVCT0 space; Let the planned CT images of the i-th patient be... The target area and the outline of organs at risk are delineated on the planned CT images. Then, the target area and organ at risk contours on the planned CT image are mapped to the MVCT0 space through the spatial transformation relationship to obtain the initial structural labels on the MVCT0 image:
[0058] ; Where K represents the total number of structures in the target area and organs at risk. This represents the initial target area and organ at risk labels on the MVCT0 image of the i-th patient, i.e., the initial structural labels; The first construction module 23 is used to perform adaptive image quality enhancement and cross-domain adaptation on the full-segment MVCT images, and to construct the enhanced full-segment MVCT images into full-segment MVCT time-series data according to the segment order. First, the image quality score is calculated for each segment of the full-segment MVCT images, represented as: ; in, This represents the quality score of the t-th fractional MVCT image. This represents an image quality evaluation function, which is calculated based on one or more of the following: noise level, local contrast, edge sharpness, gray-level distribution stability, and structural boundary visibility. This represents the MVCT image corresponding to the t-th treatment fraction for the i-th patient; Subsequently, based on the image quality score, the corresponding enhancement strategy is selected or adjusted, including grayscale normalization, noise suppression, edge-preserving filtering, local contrast enhancement, structural boundary enhancement, and artifact suppression. The enhanced MVCT image is represented as follows: ; in, This represents the t-th fractional MVCT image after enhancement. Represents the image enhancement function; The enhanced MVCT0 to MVCTn images are constructed into full-fraction MVCT time-series data according to the treatment fractionation order, and represented as follows: ; Module 24 is used to determine the short-range propagation result and the long-range propagation result. The long-range propagation result is determined based on the initial structure label. Image registration is performed on adjacent MVCT images to obtain the spatial transformation field from the (t-1)th segment to the tth segment, as shown below: ; Using this spatial transformation field, the final segmentation result set obtained in the previous segment is propagated to the current segment: ; in, This represents the short-range transmission result of the i-th patient from the previous transmission to the current t-th transmission. This represents the set of final segmentation results obtained for the i-th patient in the (t-1)th segment; The initial structural label is propagated directly or through a cascade of multiple adjacent fractional transformation fields to the t-th fraction: ; in, This represents the long-range propagation result obtained from the initial structural tag of patient i to the t-th segment. This represents the long-range spatial transformation from the MVCT0 image to the t-th MVCT image, which is obtained by cascading multiple adjacent transformation fields; The second construction module 25 is used to construct a spatiotemporal Transformer segmentation network guided by the prior knowledge of the planned CT structure. It includes an MVCT spatial encoder, a structural prior encoder, a long temporal memory module, a spatiotemporal Transformer module, an anatomical relationship constraint module, and a multi-structure segmentation decoder. Using the full-segment MVCT temporal data, the structural prior atlas, and the initial structural labels, it generates a direct segmentation probability map and a direct segmentation label map for the current segment of the MVCT image. The MVCT spatial encoder is used to extract three-dimensional spatial features from each enhanced MVCT image, represented as follows: ; in, Represents a three-dimensional spatial encoder. Represents the spatial features of the t-th fractional MVCT image; The structural prior encoder is used to encode the structural prior information consisting of the structural prior map constructed based on the planned CT and the initial structural labels, and represents it as follows: ; in, This represents a structural prior encoder. Represents prior structural features; In the long-term memory module, a dynamic memory bank is established, represented as follows: ; in, This represents the dynamic memory bank of the i-th patient up to the t-th fraction. This refers to the memory item of the (t-1)th historical segment in the dynamic memory bank; For the current segment t, the most relevant historical structural information about the current image is extracted from the dynamic memory using an attention retrieval mechanism: ; in, Indicates historical memory retrieval features, Represents the attention function; The spatiotemporal Transformer module is used to acquire the spatial features of the full fractional MVCT of the same patient and learn the long-range dependencies between different treatment fractions, represented as: ; in, This refers to the spatiotemporal Transformer module. This represents the spatiotemporal context feature corresponding to the t-th fraction; The anatomical relationship constraint module is used to construct an anatomical relationship diagram based on the spatial relationships between the target area and organs at risk in the prior structural atlas, as shown below: ; Among them, nodes Indicates the target area and organs at risk, border Indicates the distance, direction, proximity, containment, separation, or other spatial relationships between structures; The spatial features, structural prior features, historical memory retrieval features, and spatiotemporal context features of the current fractional MVCT images are fused and represented as follows: ; in, This indicates the spatiotemporal feature fusion module. Indicates the characteristics after fusion; The fused features are input into the multi-structure segmentation decoder to obtain the direct segmentation probability map of the current fractional MVCT image, represented as: ; And obtain the direct segmentation label map from the direct segmentation probability map: ; in, This represents a multi-structure segmentation decoder. This represents the probability map of direct segmentation of the t-th fractional MVCT image of the i-th patient. This represents the direct segmentation label map obtained from the corresponding direct segmentation probability map. This represents the maximum probability index function; The second acquisition module 26 is used to obtain the final segmentation result of the current segmentation based on the short-range propagation result, the long-range propagation result, and the direct segmentation label map.
[0059] Furthermore, in other embodiments of the present invention, the nasopharyngeal carcinoma full-fraction MVCT target area and organ at-risk segmentation system further includes: The calculation module is used to calculate the uncertainty value based on the direct segmentation probability map, short-range propagation results, and long-range propagation results; The comparison module is used to compare the uncertainty value with a threshold and perform hierarchical processing on the final segmentation result of the current segmentation. Specifically, based on the comparison results of the uncertainty value with at least two preset thresholds, the final segmentation result of the current segmentation is subject to three levels of differentiated processing: When the uncertainty value meets the first preset condition, the segmentation result is automatically accepted; When the uncertainty value meets the second preset condition, the segmentation result is automatically optimized and corrected based on multi-source prior information and historical segmentation information. When the uncertainty value meets the third preset condition, the structure or region corresponding to the segmentation result is marked as a manually reviewed region.
[0060] Example 3 In another aspect, the present invention also proposes an electronic device, please refer to [link to relevant documentation]. Figure 3 The image shows an electronic device according to Embodiment 3 of the present invention, including a memory 20, a processor 10, and a computer program 30 stored in the memory and executable on the processor. When the processor 10 executes the computer program 30, it implements the above-described method for segmenting the target area and organs at risk of nasopharyngeal carcinoma in full fractionation MVCT.
[0061] In some embodiments, the processor 10 may be a central processing unit (CPU), controller, microcontroller, microprocessor or other data processing chip, used to run program code stored in memory 20 or process data, such as executing access restriction programs.
[0062] The memory 20 includes at least one type of readable storage medium, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 20 can be an internal storage unit of an electronic device, such as the hard disk of the electronic device. In other embodiments, the memory 20 can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. Furthermore, the memory 20 can include both internal and external storage units of the electronic device. The memory 20 can be used not only to store application software and various types of data of the electronic device, but also to temporarily store data that has been output or will be output.
[0063] It should be pointed out that, Figure 3The structure shown does not constitute a limitation on the electronic device. In other embodiments, the electronic device may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0064] This invention also proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for segmenting the target area and organs at risk in full-fractionation MVCT of nasopharyngeal carcinoma.
[0065] Those skilled in the art will understand that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0066] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0067] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0068] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0069] The above embodiments merely illustrate several implementation methods of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this patent should be determined by the appended claims.
Claims
1. A nasopharyngeal carcinoma whole-fraction MVCT target region and organ at risk segmentation method, characterized in that, The method includes: Acquire planned CT images, target areas and organs at risk delineated on the planned CT images, and full fractionated MVCT images, and construct a structural prior atlas based on the planned CT images, target areas, and organs at risk; The target area and organs at risk delineated on the planned CT images are transformed into initial priors in the MVCT space to obtain initial structural labels; Image quality adaptive enhancement and cross-domain adaptation are performed on full-segment MVCT images, and the enhanced full-segment MVCT images are constructed into full-segment MVCT time series data according to the segment order. The short-range propagation results and long-range propagation results are determined. The short-range propagation results are determined based on the segmentation results obtained in adjacent treatment fractions, while the long-range propagation results are determined based on the initial structural labels. A spatiotemporal Transformer segmentation network guided by planned CT structural priors is constructed, including an MVCT spatial encoder, a structural prior encoder, a long temporal memory module, a spatiotemporal Transformer module, an anatomical relationship constraint module, and a multi-structure segmentation decoder. Using the full-sequence MVCT temporal data, the structural prior atlas, and the initial structural labels, a direct segmentation probability map and a direct segmentation label map of the current MVCT image are generated. Based on the short-range propagation results, long-range propagation results, and the direct segmentation label map, the final segmentation result for the current segment is obtained. 2.The nasopharyngeal carcinoma whole-fraction MVCT target region and organ at risk segmentation method according to claim 1, characterized in that, The step of obtaining the final segmentation result for the current segmentation based on the short-range propagation result, the long-range propagation result, and the direct segmentation label map includes: Calculate the uncertainty value based on the direct segmentation probability diagram, short-range propagation results, and long-range propagation results; The uncertainty value is compared with a threshold, and the final segmentation result of the current segment is subjected to hierarchical processing. Specifically, based on the comparison results of the uncertainty value with at least two preset thresholds, the final segmentation result of the current segmentation is subject to three levels of differentiated processing: When the uncertainty value meets the first preset condition, the segmentation result is automatically accepted; When the uncertainty value meets the second preset condition, the segmentation result is automatically optimized and corrected based on multi-source prior information and historical segmentation information. When the uncertainty value meets the third preset condition, the structure or region corresponding to the segmentation result is marked as a manually reviewed region. 3.The nasopharyngeal carcinoma whole-fraction MVCT target region and organ at risk segmentation method according to claim 2, characterized in that, In the step of constructing a priori structural atlases based on planned CT images, target areas, and organs at risk... For the k-th structure of the i-th patient, its structural prior representation is as follows: ; wherein, a label map representing the kth structure of the ith patient, a structure volume representing the kth structure of the ith patient, a structure centroid representing the kth structure of the ith patient, a structure boundary representing the kth structure of the ith patient, a structure shape feature representing the kth structure of the ith patient, For any two structures k and l, calculate their spatial relationship: ; wherein, represents a distance between the kth structure and the lth structure centroid of the ith patient, represents a relative directional relationship between the kth structure and the lth structure of the ith patient, represents a proximity relationship between the kth structure and the lth structure of the ith patient, represents a containment, intersection, separation, or spatial constraint relationship between the kth structure and the lth structure of the ith patient; Thus, a structural prior map is constructed: ; in, This represents the structural prior map of the i-th patient constructed based on the planned CT scan. This represents the anatomical prior information of the k-th structure in the i-th patient. This represents the spatial relationship between the k-th structure and the l-th structure of the i-th patient.
4. The method for segmenting the target region and organs at risk in nasopharyngeal carcinoma using full fractional MVCT according to claim 3, characterized in that, In the step of converting the target area and organs at risk delineated on the planned CT image into initial priors in the MVCT space to obtain initial structural labels... Cross-modal registration of the planned CT image and the MVCT0 image yields the spatial transformation relationship from the planned CT space to the MVCT0 space, represented as follows: ; MVCT0 images refer to MVCT images acquired before the first treatment, and pCT refers to the planning CT space. For MVCT0 space; Let the planned CT images of the i-th patient be... The target area and the outline of organs at risk are delineated on the planned CT images. Then, the target area and organ at risk contours on the planned CT image are mapped to the MVCT0 space through the spatial transformation relationship to obtain the initial structural labels on the MVCT0 image: ; Where K represents the total number of structures in the target area and organs at risk. This represents the initial target area and organ at risk labels on the MVCT0 image of the i-th patient, i.e., the initial structural labels.
5. The method for segmenting the target region and organs at risk in nasopharyngeal carcinoma using full fractional MVCT according to claim 4, characterized in that, In the step of performing adaptive image quality enhancement and cross-domain adaptation on the full-segment MVCT images, and constructing the enhanced full-segment MVCT images into full-segment MVCT time-series data according to the segment order, First, an image quality score is calculated for each fraction of the full-fraction MVCT image, expressed as: ; in, This represents the quality score of the t-th fractional MVCT image. This represents an image quality evaluation function, which is calculated based on one or more of the following: noise level, local contrast, edge sharpness, gray-level distribution stability, and structural boundary visibility. This represents the MVCT image corresponding to the t-th treatment fraction for the i-th patient; Subsequently, based on the image quality score, the corresponding enhancement strategy is selected or adjusted, including grayscale normalization, noise suppression, edge-preserving filtering, local contrast enhancement, structural boundary enhancement, and artifact suppression. The enhanced MVCT image is represented as follows: ; in, This represents the t-th fractional MVCT image after enhancement. Represents the image enhancement function; The enhanced MVCT0 to MVCTn images are constructed into full-fraction MVCT time-series data according to the treatment fractionation order, and represented as follows: 。 6. The method for segmenting the target region and organs at risk in nasopharyngeal carcinoma using full fractional MVCT according to claim 5, characterized in that, In the steps of determining the short-range propagation results and the long-range propagation results Image registration is performed on adjacent MVCT images to obtain the spatial transformation field from the (t-1)th fraction to the tth fraction, expressed as: ; Using this spatial transformation field, the final segmentation result set obtained in the previous segment is propagated to the current segment: ; in, This represents the short-range transmission result of the i-th patient from the previous transmission to the current t-th transmission. This represents the set of final segmentation results obtained for the i-th patient in the (t-1)th segment; The initial structural label is propagated directly or through a cascade of multiple adjacent fractional transformation fields to the t-th fraction: ; in, This represents the long-range propagation result obtained from the initial structural tag of patient i to the t-th segment. This represents the long-range spatial transformation from the MVCT0 image to the t-th MVCT image, which is obtained by cascading multiple adjacent transformation fields.
7. The method for segmenting the target region and organs at risk in nasopharyngeal carcinoma using full fractional MVCT according to claim 6, characterized in that, The MVCT spatial encoder is used to extract three-dimensional spatial features from each enhanced MVCT image, represented as follows: ; in, Represents a three-dimensional spatial encoder. Represents the spatial features of the t-th fractional MVCT image; The structural prior encoder is used to encode the structural prior information consisting of the structural prior map constructed based on the planned CT and the initial structural labels, and represents it as follows: ; in, This represents a structural prior encoder. Represents prior structural features; In the long-term memory module, a dynamic memory bank is established, represented as follows: ; in, This represents the dynamic memory bank of the i-th patient up to the t-th fraction. This refers to the memory item of the (t-1)th historical segment in the dynamic memory bank; For the current segment t, the most relevant historical structural information about the current image is extracted from the dynamic memory using an attention retrieval mechanism: ; in, Indicates historical memory retrieval features, Represents the attention function; The spatiotemporal Transformer module is used to acquire the spatial features of the full fractional MVCT of the same patient and learn the long-range dependencies between different treatment fractions, represented as: ; in, This refers to the spatiotemporal Transformer module. This represents the spatiotemporal context feature corresponding to the t-th fraction; The anatomical relationship constraint module is used to construct an anatomical relationship diagram based on the spatial relationships between the target area and organs at risk in the prior structural atlas, as shown below: ; Among them, nodes Indicates the target area and organs at risk, border Indicates the distance, direction, proximity, containment, separation, or other spatial relationships between structures; The spatial features, structural prior features, historical memory retrieval features, and spatiotemporal context features of the current fractional MVCT images are fused and represented as follows: ; in, This indicates the spatiotemporal feature fusion module. Indicates the characteristics after fusion; The fused features are input into the multi-structure segmentation decoder to obtain the direct segmentation probability map of the current fractional MVCT image, represented as: ; And obtain the direct segmentation label map from the direct segmentation probability map: ; in, This represents a multi-structure segmentation decoder. This represents the probability map of direct segmentation of the t-th fractional MVCT image of the i-th patient. This represents the direct segmentation label map obtained from the corresponding direct segmentation probability map. This represents the maximum probability index function.
8. A MVCT target region and organ-at-risk segmentation system for nasopharyngeal carcinoma, characterized in that, The system for implementing the nasopharyngeal carcinoma full-fraction MVCT target region and organ at-risk segmentation method according to any one of claims 1-7 comprises: The first acquisition module is used to acquire the planned CT image, the target area and organs at risk delineated on the planned CT image, and the full fraction MVCT image, and to construct a structural prior atlas based on the planned CT image, the target area and organs at risk. The transformation module is used to transform the target area and organs at risk delineated on the planned CT image into initial priors in the MVCT space to obtain initial structural labels; The first construction module is used to perform adaptive image quality enhancement and cross-domain adaptation on the full-segment MVCT images, and to construct the enhanced full-segment MVCT images into full-segment MVCT time-series data according to the segment order. The determination module is used to determine the short-range propagation results and the long-range propagation results. The short-range propagation results are determined based on the segmentation results obtained in adjacent treatment fractions, while the long-range propagation results are determined based on the initial structure labels. The second construction module is used to construct a spatiotemporal Transformer segmentation network guided by the prior structure of the planned CT, including an MVCT spatial encoder, a structural prior encoder, a long temporal memory module, a spatiotemporal Transformer module, an anatomical relationship constraint module, and a multi-structure segmentation decoder. It generates a direct segmentation probability map and a direct segmentation label map of the current MVCT image through the full-segment MVCT temporal data, the structural prior map, and the initial structural labels. The second acquisition module is used to obtain the final segmentation result of the current segmentation based on the short-range propagation result, the long-range propagation result, and the direct segmentation label map.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the nasopharyngeal carcinoma full-fraction MVCT target region and organ at risk segmentation method as described in any one of claims 1-7.
10. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the program to implement the nasopharyngeal carcinoma full-fraction MVCT target region and organ at risk segmentation method as described in any one of claims 1-7.