Respiration-based organ motion control method and apparatus

By constructing a deformation constraint loss function and a deep learning generative model, and using a pre-trained organ deformation encoder to stitch together organ masks, the problem of inaccurate organ deformation prediction was solved, and more accurate organ deformation prediction was achieved, meeting the accuracy requirements of radiotherapy.

CN122391296APending Publication Date: 2026-07-14MANTEIA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MANTEIA TECH CO LTD
Filing Date
2026-06-09
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing respiratory motion management methods based on body surface signals are prone to unreasonable deformations such as organ stretching and boundary fragmentation when predicting organ deformation, resulting in inaccurate prediction of organ deformation information, affecting the reliability of target delineation and jeopardizing the credibility of organ dose assessment.

Method used

By using a pre-trained organ deformation encoder to stitch together a reference organ mask and a predicted organ mask, a deformation constraint loss function is constructed. A deep learning generative model is then used to predict the organ delineation mask under the target respiratory phase, reducing the distance between the prediction result and the actual deformation and improving the prediction accuracy.

Benefits of technology

By minimizing the distance between the predicted latent deformation variable distribution and the actual latent deformation variable distribution, abnormal stretching and boundary fragmentation of organs can be avoided, thereby improving the accuracy of organ deformation information prediction and meeting the precision requirements of radiotherapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391296A_ABST
    Figure CN122391296A_ABST
Patent Text Reader

Abstract

The application discloses a kind of organ movement control method and device based on breathing, it is related to medical technology field, the method comprises: using organ deformation encoder, reference organ mask is spliced with predicted organ mask input to obtain predicted deformation latent variable distribution;Reference organ mask is spliced with real organ mask input organ deformation encoder to obtain real deformation latent variable distribution;With minimizing the distance of predicted deformation latent variable distribution and real deformation latent variable distribution as target, construct deformation constraint loss function;The body surface deformation field of target object is obtained;Body surface deformation field is input into the deep learning generation model trained by deformation constraint loss function, and the organ delineation mask of target object under target breathing phase is predicted.The application solves the problem that existing breathing motion management method based on body surface signal can easily lead to unreasonable deformation conditions such as organ stretching and boundary fragmentation in predicted results, thereby causing inaccurate organ deformation information prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical technology, and more specifically, to a method and device for controlling organ movement based on respiration. Background Technology

[0002] During tumor radiotherapy, the patient's respiratory movements cause periodic displacement and deformation of the tumor in the target area. To reduce the impact of respiratory movements on treatment accuracy, a respiratory movement management method based on body surface signals is often used clinically. This involves collecting the patient's body surface movement signals using optical body surface monitoring equipment and using these signals to indirectly estimate the movement status of internal organs.

[0003] However, due to the complex mapping relationship between surface motion and organ deformation, traditional methods often produce predictions of abnormal stretching, fragmented local boundaries, or abrupt changes in deformation between consecutive phases when predicting organ deformation from a reference respiratory phase to a target respiratory phase. These predictions, which do not conform to the actual motion morphology of the organ, make it difficult to meet the accuracy requirements of clinical radiotherapy for organ deformation information, thereby affecting the reliability of target delineation and jeopardizing the credibility of organ dose assessment.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This application provides a respiratory-based organ movement control method and apparatus to at least solve the technical problem that existing respiratory movement management methods based on body surface signals are prone to causing unreasonable deformations such as organ stretching and boundary breakage in the prediction results, resulting in inaccurate prediction of organ deformation information.

[0006] According to one aspect of the embodiments of this application, a respiratory-based organ motion control method is provided, comprising: using a pre-trained organ deformation encoder, concatenating a reference organ mask and a predicted organ mask as input to obtain a predicted deformation latent variable distribution, wherein the predicted deformation latent variable distribution is a vector distribution used to characterize the predicted organ deformation features output by the organ deformation encoder after encoding the reference organ mask and the predicted organ mask; concatenating a reference organ mask and a real organ mask as input to the same organ deformation encoder to obtain a real deformation latent variable distribution, wherein the real deformation latent variable distribution refers to... The vector distribution, which represents the deformation features of real organs, is output by an organ deformation encoder after encoding the reference organ mask and the real organ mask. A deformation constraint loss function is constructed with the goal of minimizing the distance between the predicted deformation latent variable distribution and the real deformation latent variable distribution. The body surface deformation field of the target object is obtained, where the body surface deformation field represents the deformation of the target object's body surface from the reference respiratory phase to the target respiratory phase. The body surface deformation field is input into a deep learning generative model trained by the deformation constraint loss function to predict the organ delineation mask of the target object under the target respiratory phase.

[0007] According to another aspect of the embodiments of this application, a respiratory-based organ motion control device is also provided, comprising: a first splicing processing unit, configured to use a pre-trained organ deformation encoder to splice a reference organ mask and a predicted organ mask into a vector distribution for predicting organ deformation characteristics, wherein the predicted organ deformation latent variable distribution is a vector distribution output by the organ deformation encoder after encoding the reference organ mask and the predicted organ mask; and a second splicing processing unit, configured to splice a reference organ mask and a real organ mask into the same organ deformation encoder to obtain a real organ deformation latent variable distribution, wherein the real organ deformation latent variable distribution refers to the distribution obtained by... The organ deformation encoder encodes a vector distribution that represents the deformation features of the real organ after encoding the reference organ mask and the real organ mask; the function construction unit constructs a deformation constraint loss function with the objective of minimizing the distance between the predicted deformation latent variable distribution and the real deformation latent variable distribution; the acquisition unit acquires the surface deformation field of the target object, where the surface deformation field represents the deformation of the target object's surface from the reference respiratory phase to the target respiratory phase; and the mask processing unit inputs the surface deformation field into the deep learning generative model trained by the deformation constraint loss function to predict the organ delineation mask of the target object under the target respiratory phase.

[0008] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, which stores a computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located executes the above-described respiratory-based organ movement control method.

[0009] According to another aspect of the embodiments of this application, an electronic device is also provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors cause the one or more processors to perform the above-described respiratory-based organ movement control method.

[0010] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program or instructions, which, when executed by a processor, implement the above-described respiratory-based organ movement control method.

[0011] In this embodiment, a pre-trained organ deformation encoder is first used to concatenate a reference organ mask under a reference respiratory phase with a predicted organ mask under the same phase predicted by a deep learning generative model. The concatenated data is then input into the organ deformation encoder to obtain a predicted deformation latent variable distribution. This distribution describes the predicted organ deformation features. Simultaneously, the reference organ mask under the same reference respiratory phase is concatenated with the actual organ mask under the actual target respiratory phase. The resulting data is then input into the same organ deformation encoder to obtain a true deformation latent variable distribution. This true distribution originates from the actual organ mask data and can encode the physiological deformation features of an organ during actual respiration, from the reference respiratory phase to the target respiratory phase. Next, a deformation constraint loss function is constructed with the objective of minimizing the distance between the predicted and true deformation latent variable distributions. Since the true distribution of latent deformation variables represents a deformation pattern that conforms to physiological laws, while the predicted distribution of latent deformation variables is the deformation feature currently output by the deep learning generative model, the distance between the predicted and true distributions of latent deformation variables can be continuously reduced. This helps to force the organ deformation features learned by the deep learning generative model to approximate the statistical laws of real organ deformation. Then, a body surface deformation field is obtained, representing the actual deformation of the target object's body surface from the reference respiratory phase to the target respiratory phase. This surface deformation field is input into the deep learning generative model trained with a deformation constraint loss function, and the deep learning generative model outputs an organ delineation mask under the target respiratory phase. Because deep learning generative models adjust the distribution of predicted deformation latent variables to match the distribution of actual deformation latent variables during training using deformation constraint loss functions, they can output organ deformation results that conform to physiological laws even when faced with new body surface deformation fields. This avoids unreasonable deformations such as abnormal organ stretching and boundary breakage, improves the accuracy of organ deformation information prediction, and solves the technical problem that existing respiratory movement management methods based on body surface signals are prone to predicting unreasonable deformations such as organ stretching and boundary breakage, resulting in inaccurate organ deformation information prediction. Attached Figure Description

[0012] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0013] Figure 1 This is a schematic diagram of an optional respiratory-based organ movement control method according to an embodiment of this application;

[0014] Figure 2 This is a flowchart of an optional organ movement control method based on respiration according to an embodiment of this application;

[0015] Figure 3 This is a schematic diagram of an optional respiratory-based organ movement control device according to an embodiment of this application. Detailed Implementation

[0016] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0017] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0018] According to an embodiment of this application, a method embodiment of an organ movement control method based on respiration is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0019] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of this data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding access points are provided for users to choose to authorize or refuse. For example, interfaces are set up between this system and relevant users or organizations, providing users with corresponding access points to choose to agree to or refuse automated decision-making results; if the user chooses to refuse, the process proceeds to the expert decision-making stage.

[0020] According to the embodiments of this application, a radiotherapy processing system (hereinafter referred to as "processing system") can be used as the execution subject of the respiratory-based organ movement control method of this application embodiment. The system can be a software system or an embedded system combining software and hardware. Of course, the method execution subject in the embodiments of this application can also be other forms of execution subject, such as devices, equipment, etc. It should be known by those skilled in the art that this application does not particularly limit the specific form of the method execution subject.

[0021] Figure 1 This is a schematic diagram of a respiratory-based organ movement control method according to an embodiment of this application, as shown below. Figure 1 As shown, the method includes the following steps:

[0022] Step S101: Using a pre-trained organ deformation encoder, the reference organ mask and the predicted organ mask are concatenated and input to obtain the predicted deformation latent variable distribution. The predicted deformation latent variable distribution is a vector distribution output by the organ deformation encoder after encoding the reference organ mask and the predicted organ mask, which is used to characterize the predicted organ deformation features.

[0023] For example, the reference organ mask can refer to the organ contour data segmented from a medical image at a reference respiratory phase, and the predicted organ mask can refer to the predicted contour data of the same organ at a target respiratory phase, output by a deep learning generation model. The organ deformation encoder can refer to a pre-trained neural network model used to map the organ masks from two different respiratory phases, concatenated along the channel dimension, to a deformation latent variable distribution. This distribution can be represented as a multi-dimensional vector distribution characterizing the organ's deformation features from the reference phase to the target phase. The processing system concatenates the reference organ mask and the predicted organ mask along the channel dimension, inputs the concatenated data into the pre-trained organ deformation encoder, and obtains the predicted deformation latent variable distribution.

[0024] In some embodiments, the organ deformation encoder can employ the encoder portion of a variational autoencoder structure. The deformation latent variable distribution output by this organ deformation encoder can be represented as a multidimensional Gaussian distribution, parameterized by a mean vector and a diagonal covariance vector. The processing system concatenates the reference organ mask and the predicted organ mask and inputs them into the variational autoencoder to obtain the mean vector and diagonal covariance vector of the predicted deformation latent variable distribution. The processing system also concatenates the reference organ mask and the real organ mask and inputs them into the same variational autoencoder to obtain the mean vector and diagonal covariance vector of the real deformation latent variable distribution. The processing system can use KL divergence (Kullback-Leibler divergence, also known as relative entropy) as a distance metric between the predicted and real deformation latent variable distributions, and incorporate this KL divergence as a component of the deformation constraint loss function during the training of the deep learning generative model. This embodiment characterizes the uncertainty of organ deformation in the form of a probability distribution, which is beneficial to improving the statistical stability of deformation constraints and to adapting to the natural fluctuations of organ deformation in different respiratory cycles.

[0025] In other embodiments, the organ deformation encoder can employ a deterministic encoder structure based on metric learning. The deformation latent variable distribution output by this encoder can be represented as a fixed-dimensional feature vector, which serves as a point estimate of the deformation features. The processing system concatenates the reference organ mask and the predicted organ mask and inputs them into the deterministic encoder to obtain the predicted deformation feature vector. The processing system also concatenates the reference organ mask and the real organ mask and inputs them into the same deterministic encoder to obtain the real deformation feature vector. The processing system constructs a deformation constraint loss function with the goal of minimizing the Euclidean distance between the predicted and real deformation feature vectors or maximizing their cosine similarity. This embodiment eliminates the need for explicit modeling of the distribution parameters, which improves computational efficiency and is beneficial for real-time applications with high inference speed requirements. Furthermore, the feature space obtained through metric learning more directly reflects the semantic similarity between organ deformations.

[0026] Optionally, the organ deformation encoder can also be pre-trained using other latent space representation models, including at least one of standard autoencoders, denoising autoencoders, contrastive learning encoders, masked autoencoders, diffusion autoencoders, streaming model encoders, or attention-based deformation encoders. Regardless of the encoder structure used, the common goal of the organ deformation encoder is to learn a low-dimensional deformation representation of the organ across the respiratory phase and to provide deformation distribution constraints during the training phase of the deep learning generative model. In the embodiments of this application, the processing system can refer to any computing device with the capabilities of data acquisition, data processing, model inference, and result output, including but not limited to a radiotherapy control system, a stand-alone image processing workstation, a general-purpose computer, or a server.

[0027] Step S102: The reference organ mask and the real organ mask are concatenated and input into the same organ deformation encoder to obtain the real deformation latent variable distribution. The real deformation latent variable distribution refers to the vector distribution output by the organ deformation encoder after encoding the reference organ mask and the real organ mask, which is used to characterize the deformation features of the real organ.

[0028] For example, the true organ mask can refer to the true contour data of the same organ segmented from medical images at the target respiratory phase. This true organ mask can serve as the primary standard for the supervision signal during the training phase. The true deformation latent variable distribution can be represented as a vector distribution output by the same organ deformation encoder after encoding the reference organ mask and the true organ mask. This vector distribution is used to characterize the true deformation features of the organ from the reference respiratory phase to the target respiratory phase. The processing system concatenates the reference organ mask and the true organ mask along the channel dimension and inputs the concatenated data into the same organ deformation encoder as in the above steps. This organ deformation encoder outputs the true deformation latent variable distribution, which is beneficial for extracting the true organ deformation feature distribution from the standard data using the organ deformation encoder. This true deformation latent variable distribution serves as a reference benchmark for subsequently constructing the deformation constraint loss function, which helps the generative model learn the statistical characteristics of true deformation during training, thereby improving the physiological reliability of the prediction results.

[0029] In some embodiments, the organ deformation encoder may employ the encoder portion of a variational autoencoder structure, and the true deformation latent variable distribution output by the organ deformation encoder may be represented as a multidimensional Gaussian distribution. The processing system inputs the concatenated reference organ mask and the true organ mask into the variational autoencoder, which outputs the mean vector and diagonal covariance vector of the true deformation latent variable distribution.

[0030] In other embodiments, the organ deformation encoder can employ a deterministic encoder structure based on metric learning. The true deformation latent variable distribution output by this encoder can be represented as a fixed-dimensional true deformation feature vector. The processing system concatenates a reference organ mask with a true organ mask and inputs the result into the deterministic encoder. The deterministic encoder outputs the true deformation feature vector, which serves as a point estimate of the true deformation features. In this embodiment, the true deformation feature vector exists in the form of a deterministic vector, which simplifies the subsequent calculation of the deformation constraint loss function, improves training efficiency, and facilitates direct comparison with the predicted deformation feature vector output using the same deterministic encoder structure using Euclidean distance or cosine similarity.

[0031] Step S103: Construct a deformation constraint loss function with the objective of minimizing the distance between the predicted latent deformation variable distribution and the actual latent deformation variable distribution.

[0032] For example, the distance between the predicted latent deformation variable distribution and the actual latent deformation variable distribution can be represented as a non-negative numerical value. This distance quantifies the degree of difference between the two distributions; a smaller value indicates that the two distributions are closer. The deformation constraint loss function can be represented as a function with distance as the independent variable, and its optimization direction is to minimize the distance between the predicted and actual latent deformation variable distributions. The processing system constructs the deformation constraint loss function with the goal of minimizing the distance and incorporates this function as a component of the overall loss function of the deep learning generative model during model training. This embodiment of the application facilitates the introduction of statistical prior knowledge of real organ deformation into the training process of the generative model in the form of minimizing the distribution distance. This helps the predicted organ mask output by the generative model approximate the deformation feature distribution of the real organ mask at the latent deformation variable space level, thereby improving the physiological rationality of the prediction results and reducing the systematic deviation between the predicted deformation and the actual deformation patterns.

[0033] In some embodiments, the processing system can use KL divergence as a distance metric between the predicted and actual distributions of latent deformation variables. The processing system calculates the KL divergence between the predicted and actual distributions of latent deformation variables and uses the calculated KL divergence value as the deformation constraint loss function. The processing system weights and sums the deformation constraint loss function with the overlap loss function and the boundary loss function to obtain the total loss function, and updates the network parameters of the deep learning generative model during backpropagation with the goal of minimizing the total loss function. This embodiment, by constraining the deformation prediction of the generative model through KL divergence between probability distributions, helps to minimize the relative entropy between the predicted and actual distributions of latent deformation variables and facilitates considering both the mean and variance information of the distributions during training.

[0034] In other embodiments, the processing system may also use the Wasserstein distance between the predicted and true deformation latent variable distributions as a distance metric. The processing system calculates the minimum transportation cost required to transform the predicted deformation latent variable distribution into the true distribution by solving an optimal transportation problem, and uses this minimum transportation cost as the deformation constraint loss function. The processing system combines the deformation constraint loss function constructed from the Wasserstein distance with the basic delineation loss function to train the deep learning generative model. This embodiment uses the Wasserstein distance, which is beneficial for providing smoother gradient information than KL divergence, avoiding gradient vanishing or mode collapse problems during training, and providing a meaningful distance metric even when the support sets of the predicted and true deformation latent variable distributions do not overlap.

[0035] Optionally, the distance metric between the predicted latent deformation variable distribution and the actual latent deformation variable distribution can also employ a latent variable quantile excess penalty. The latent variable quantile excess penalty quantifies the difference between the two distributions by calculating the degree to which the quantiles of the predicted latent deformation variable distribution exceed a preset multiple of the corresponding quantiles of the actual latent deformation variable distribution. This distance metric is beneficial for providing a stronger constraint signal when the differences at the tails of the distributions are large.

[0036] Optionally, the organ-level latent deformation latent space profile corresponding to the true deformation latent variable distribution can be represented in the following ways: a mean and covariance model to describe the first and second-order statistical properties of the deformation latent variable distribution; upper and lower bounds of quantiles to define the fluctuation range of normal deformation; cluster centers to divide normal deformation into several typical patterns; a single-class classifier to learn the boundaries of normal deformation; and an anomaly detection model to identify prediction results that deviate from the normal deformation distribution. All of the above representations can serve as alternative reference benchmarks for the true deformation latent variable distribution in the deformation constraint loss function.

[0037] Step S104: Obtain the surface deformation field of the target object, wherein the surface deformation field characterizes the deformation of the target object's surface from the reference respiratory phase to the target respiratory phase.

[0038] For example, the surface deformation field can be represented as a vector field, where each vector corresponds to a three-dimensional displacement vector of a spatial point on the target object's surface between a reference respiratory phase and a target respiratory phase. The surface deformation field describes the degree and direction of deformation of the target object's surface from the reference respiratory phase to the target respiratory phase. The processing system can acquire the surface deformation field of the target object by: acquiring a reference surface point cloud of the target object at the reference respiratory phase and a target surface point cloud at the target respiratory phase; registering the reference and target surface point clouds; and generating the surface deformation field based on the registration result. This embodiment of the application is advantageous in using the surface deformation field as a transition signal connecting external respiratory motion and internal organ motion, and in utilizing the dense three-dimensional spatial information of the surface deformation field to characterize the full-field deformation of the surface under different respiratory modes, thereby providing richer motion guidance features for subsequent deep learning generative models.

[0039] In some embodiments, the processing system can acquire reference surface point clouds and target surface point clouds of the target object under a reference breathing phase and a target breathing phase using a binocular vision camera. The processing system reconstructs a reference surface mask based on the reference surface point clouds and a target surface mask based on the target surface point clouds. The processing system uses a non-rigid registration algorithm to register the reference and target surface masks, calculates the displacement vector from each surface point on the reference surface mask to the corresponding point on the target surface mask, and organizes all displacement vectors into a surface deformation field. This embodiment obtains the surface deformation field through non-rigid registration, which is beneficial for handling the complex non-rigid and nonlinear deformations that occur on the body surface during breathing, and improves the accuracy of the surface deformation field in representing real body surface motion.

[0040] In other embodiments, the processing system acquires a sequence of surface depth images of the target object using a depth sensor or structured light camera, and extracts a reference depth map at a reference respiratory phase and a target depth map at the target respiratory phase from the surface depth image sequence. The processing system converts the reference depth map and the target depth map into 3D point clouds, and calculates the dense displacement field between the reference depth map and the target depth map using an optical flow-based registration method, mapping the dense displacement field into a 3D surface deformation field. This embodiment directly calculates the dense displacement field from the depth images, which helps reduce the computational complexity of point cloud reconstruction and registration, improves the acquisition speed of the surface deformation field, and meets the low-latency processing requirements of real-time radiotherapy control scenarios.

[0041] Step S105: Input the body surface deformation field into the deep learning generative model trained by the deformation constraint loss function to predict the organ delineation mask of the target object under the target respiratory phase.

[0042] For example, a deep learning generative model can be represented as a trained nonlinear mapping network that takes the body surface deformation field as input features and outputs an organ delineation mask at the target respiratory phase. A deep learning generative model trained using a deformation constraint loss function can be represented as a model that uses the deformation constraint loss function for parameter optimization during training. This deformation constraint loss function can constrain the distribution of the latent deformation variables of the predicted organ mask to approximate the distribution of the latent deformation variables of the real organ mask. The processing system inputs the acquired body surface deformation field into the deep learning generative model, which performs feature extraction and mapping on the input body surface deformation field, outputting an organ delineation mask of the target object at the target respiratory phase. This embodiment of the application is advantageous in utilizing the body surface deformation field as an external motion guidance signal, establishing a nonlinear mapping relationship between body surface deformation and internal organ deformation through a deep learning generative model. This facilitates real-time output of organ contours at the target respiratory phase without the need for in vivo imaging equipment, thereby providing real-time anatomical location information for radiotherapy.

[0043] In some embodiments, the deep learning generative model can employ a generative network based on an encoder-decoder structure. This encoder-decoder structure includes an encoder module and a decoder module. The encoder module performs multi-scale downsampling on the input body surface deformation field to extract multi-level deformation features. The decoder module upsamples and reconstructs the multi-level deformation features, outputting an organ delineation mask with the same spatial resolution as the input. The processing system inputs the body surface deformation field into the deep learning generative model with this encoder-decoder structure. This deep learning generative model passes the shallow features from the encoder module to the corresponding levels of the decoder module through skip connections, outputting an organ delineation mask at the target respiratory phase. This embodiment, through the encoder-decoder structure and skip connections, helps to preserve detailed information in the body surface deformation field and improves the spatial resolution and edge accuracy of the output organ delineation mask.

[0044] In other embodiments, the deep learning generative model can employ a flow matching-based generative model, which includes a feature encoder, a continuous deformation path construction module, a velocity field prediction network, and a latent space decoder. The processing system inputs the body surface deformation field into the flow matching model, which performs joint feature encoding on the deformation field to obtain encoded features. It then constructs a continuous deformation path from the encoded features to the feature space corresponding to the organ delineation mask. At multiple consecutive time steps along the continuous deformation path, it predicts the velocity field corresponding to each time step, generates a latent space feature representation based on the velocity field integral, and finally decodes the latent space feature representation to obtain the organ delineation mask under the target respiratory phase. This embodiment generates the organ delineation mask using flow matching technology, which helps maintain the continuity and smoothness of the deformation path during the generation process, and improves the temporal consistency and anatomical rationality of the generated results.

[0045] In some optional embodiments, the body surface deformation field is input into a deep learning generative model trained by a deformation constraint loss function to predict the organ delineation mask of the target object under the target respiratory phase. This includes: the processing system can obtain a reference organ delineation of the target object under a reference respiratory phase, wherein the reference organ delineation includes a segmentation mask of at least one target organ and at least one organ at risk; the reference organ delineation and the body surface deformation field are jointly input into the deep learning generative model, and the deep learning generative model performs deformation prediction on the reference organ delineation based on the body surface deformation field to output the organ delineation mask of the target object under the target respiratory phase.

[0046] For example, reference organ delineation can refer to organ contour data segmented from medical images at a reference respiratory phase. This reference organ delineation includes a segmentation mask for at least one target organ and a segmentation mask for at least one endangered organ. The target organ can refer to a tumor region requiring radiotherapy, and the endangered organ can refer to normal tissue or organs surrounding the tumor that are sensitive to radiation. The processing system inputs both the reference organ delineation and the body surface deformation field into a deep learning generative model. Based on the surface motion guidance information provided by the body surface deformation field, the deep learning generative model simultaneously performs deformation prediction on the target organ segmentation mask and the endangered organ segmentation mask in the reference organ delineation, outputting the organ delineation mask of the target object at the target respiratory phase. The embodiments of this application are advantageous in using the body surface deformation field as a dynamic motion guidance signal to drive a deep learning generation model to synchronously predict the deformation of multiple organ contours in the reference organ delineation. This is beneficial in enabling the output organ delineation mask to simultaneously contain the position and morphological information of the target organ and the organs at risk under the target respiratory phase, thereby providing a more comprehensive anatomical basis for adaptive radiotherapy and facilitating the simultaneous assessment of target dose coverage and radiation risk to surrounding normal organs.

[0047] In some embodiments, the processing system uses a reference organ delineation as a multi-channel static anatomical prior input and a body surface deformation field as a dynamic motion guidance input. During the feature encoding stage of the deep learning generative model, the reference organ delineation and the body surface deformation field are concatenated. For example, the processing system performs independent downsampling feature extraction on the reference organ delineation and the body surface deformation field, obtaining a reference organ delineation feature map and a body surface deformation field feature map. These feature maps are then concatenated along the channel dimension and input into the decoder, which outputs an organ delineation mask under the target respiratory phase. This embodiment, by extracting static anatomical features and dynamic motion features separately and then fusing them, helps maintain the expressive power of each feature and improves the deep learning generative model's ability to model complex deformation relationships.

[0048] In other embodiments, the processing system uses a reference organ delineation as a reference template for spatial transformation and a body surface deformation field as the driving field for the spatial transformation. The deep learning generative model employs a spatial transformation network structure, which estimates a dense displacement field from a reference respiratory phase to a target respiratory phase based on the body surface deformation field. This dense displacement field is then applied to the reference organ delineation, and a mask for the organ delineation under the target respiratory phase is generated through spatial resampling. This embodiment achieves organ delineation deformation prediction through explicit spatial transformation operations, which helps maintain the topological invariance of the organ contour, improves the geometric consistency of the prediction results, and reduces the parameter complexity of training the deep learning generative model.

[0049] In some optional embodiments, obtaining the surface deformation field of the target object includes: the processing system can acquire a reference surface point cloud of the target object under a reference respiratory phase, and a target surface point cloud under a target respiratory phase; reconstruct a reference surface mask based on the reference surface point cloud, and reconstruct a target surface mask based on the target surface point cloud; register the reference surface mask and the target surface mask to obtain the surface deformation field.

[0050] For example, the reference surface point cloud can refer to the set of three-dimensional spatial points on the surface of the target object acquired by an optical acquisition device under a reference respiratory phase, and the target surface point cloud can refer to the set of three-dimensional spatial points on the surface of the target object acquired by the same optical acquisition device under the target respiratory phase. The reference surface mask can refer to a surface mesh model generated by a three-dimensional reconstruction algorithm based on the reference surface point cloud, and the target surface mask can refer to a surface mesh model generated by a three-dimensional reconstruction algorithm based on the target surface point cloud. The surface deformation field can be represented as a vector field, where each vector in the surface deformation field corresponds to a three-dimensional displacement vector between a certain point on the reference surface mask and a corresponding point on the target surface mask. The processing system acquires the reference surface point cloud and the target surface point cloud through a binocular vision camera or a depth sensor. The processing system can use three-dimensional reconstruction technology to convert the reference surface point cloud and the target surface point cloud into a reference surface mask and a target surface mask, respectively. The processing system then calculates the mapping relationship between the reference surface mask and the target surface mask using a registration algorithm to obtain the surface deformation field. The embodiments of this application are advantageous in utilizing dense three-dimensional point cloud information to reconstruct the deformation details of the entire body surface, in converting discrete point cloud data into continuous mesh masks for registration calculation, and in using the body surface deformation field obtained through registration as an external motion guidance signal for subsequent organ deformation prediction.

[0051] For example, the processing system can acquire a reference surface point cloud of the target object at a reference respiratory phase and a target surface point cloud at the target respiratory phase using a binocular vision camera. The reference surface point cloud can be represented as a set of points in three-dimensional space. The point cloud of the target surface can be represented as The processing system uses a visualization toolkit to perform 3D reconstruction of the reference surface point cloud, obtaining the reference surface mask. The target surface mask is obtained by reconstructing the point cloud of the target surface using the same method. A body surface mask is used to represent the three-dimensional morphology of the target object's body surface under a reference respiratory phase or a target respiratory phase. The processing system uses a non-rigid registration algorithm to register the reference body surface mask and the target body surface mask, calculates the displacement vector from each point on the reference body surface mask to the corresponding point on the target body surface mask, and organizes all displacement vectors into a body surface deformation field.

[0052] In some embodiments, the processing system uses a binocular vision camera to simultaneously acquire left and right eye images of the target object's body surface under a reference respiratory phase. It calculates the disparity value of each pixel using a stereo matching algorithm and reconstructs the reference body surface point cloud based on the disparity values. The processing system then reconstructs the target body surface point cloud in the same manner under the target's respiratory phase. The processing system uses a Poisson surface reconstruction algorithm to convert the reference and target body surface point clouds into a reference and a target body surface mask, respectively. The processing system uses a non-rigid iterative nearest-point algorithm to register the reference and target body surface masks, outputting the body surface deformation field. This embodiment acquires point cloud data using binocular vision, which helps improve the accuracy and density of point cloud reconstruction. The non-rigid registration facilitates handling the non-rigid deformation of the body surface during respiration.

[0053] In other embodiments, the processing system uses a structured light depth sensor to continuously acquire depth image sequences of the target object's surface, extracting a reference depth image corresponding to the reference respiratory phase and a target depth image corresponding to the target respiratory phase from the depth image sequence. The processing system backprojects each pixel in the reference depth image onto 3D space according to camera intrinsic parameters to generate a reference surface point cloud, and generates a target surface point cloud in the same way. The processing system uses a moving least squares algorithm to perform surface reconstruction on the reference and target surface point clouds, obtaining a reference surface mask and a target surface mask. The processing system uses a dense registration method based on optical flow fields to calculate the pixel-level displacement field between the reference and target depth images, mapping this displacement field onto 3D space to obtain the surface deformation field. This embodiment directly acquires point cloud data from depth images, which helps reduce the computational complexity of point cloud reconstruction, and the optical flow field-based registration method helps improve the computational efficiency and real-time performance of registration.

[0054] Optionally, embodiments of this application may employ optical imaging equipment to acquire three-dimensional point cloud data of the target object's body surface as an external monitoring signal. This acquisition process can reduce trauma and the introduction of additional ionizing radiation load, thus minimizing invasive risks and additional radiation load. The acquisition of three-dimensional point cloud data and the reconstruction of musculoskeletal structures can reduce the negative impact of changes in the target object's posture. The target object can maintain free breathing during radiotherapy, reducing the need for physical restraints such as breath-holding or abdominal compression. This helps improve clinical compliance while ensuring precise motor control and avoids muscle tremors and maladaptation caused by respiratory resistance.

[0055] In some optional embodiments, the organ deformation encoder is pre-trained through the following steps: the processing system can acquire a first organ mask at a reference respiratory phase and a second organ mask at a target respiratory phase from historical training data; the first and second organ masks are concatenated along the channel dimension and input into the encoder, outputting the distribution of organ deformation latent variables from the reference phase to the target phase; latent variables are sampled from the distribution of deformation latent variables, where the latent variables are latent space vectors characterizing organ deformation features; the latent variables and the first organ mask are input into the decoder, and the decoder reconstructs the organ mask at the target respiratory phase; the reconstruction loss between the reconstructed organ mask and the second organ mask is calculated, and the regularization loss between the distribution of deformation latent variables and the preset prior distribution is calculated; with the goal of minimizing the reconstruction loss and the regularization loss, the encoder and decoder are trained so that the encoder and decoder learn the true deformation distribution of different organs at different respiratory phases, until the training convergence condition is met, the encoder parameters are frozen, and the trained encoder is used as the organ deformation encoder.

[0056] For example, the first organ mask can refer to organ contour data obtained from historical training data at a reference respiratory phase, and the second organ mask can refer to the contour data of the same organ obtained from the same historical training data at a target respiratory phase. The encoder can be represented as a neural network module that maps the input pair of organ masks to a deformation latent variable distribution. The deformation latent variable distribution can be represented as a probability distribution that describes the statistical regularity of the organ's deformation features from the reference phase to the target phase in the latent space. The latent variable can refer to a latent space vector sampled from the deformation latent variable distribution, serving as a compact representation of the deformation features. The decoder can be represented as a neural network module that takes the latent variable and the first organ mask as input to reconstruct the organ mask at the target respiratory phase. The reconstruction loss measures the difference between the organ mask reconstructed by the decoder and the true second organ mask, while the regularization loss constrains the deviation between the deformation latent variable distribution and a preset prior distribution. The processing system aims to minimize the sum of reconstruction loss and regularization loss by jointly training the encoder and decoder, enabling them to learn the true deformation distribution of different organs across different respiratory phases. After training, the encoder parameters are frozen, and the trained encoder is used as the organ deformation encoder. This embodiment of the application is advantageous in encoding organ deformation knowledge into a probability distribution in a low-dimensional latent space, in ensuring the fidelity of deformation information during the encoding and decoding process through reconstruction loss, and in ensuring good structural regularity of the latent variable space through regularization loss. This facilitates subsequent use of the frozen encoder to impose physiologically reasonable constraints on predicted deformation.

[0057] For example, the pre-training process can be represented in the following mathematical form: (referring to the reference breathing phase) First organ mask and target breathing phase Second organ mask After concatenation along the channel dimensions, the data is input into the encoder, which outputs the deformation latent variable distribution. Latent variables are obtained by sampling from the distribution of latent variables of this deformation. latent variables With the first organ mask Input decoder, decoder reconstructs the second organ mask under the target respiratory phase By minimizing the sum of reconstruction loss and regularization loss, the encoder and decoder learn the true deformation distribution of the organ.

[0058] In some embodiments, the encoder and decoder employ a variational autoencoder architecture. The processing system concatenates the first and second organ masks along the channel dimension and inputs them into the encoder. The encoder outputs the mean vector and diagonal covariance vector of the deformation latent variable distribution, which is set as a standard Gaussian distribution as a preset prior distribution. The processing system reparameterizes latent variables from this deformation latent variable distribution to obtain latent variables, concatenates the latent variables with the first organ mask along the channel dimension, and inputs them into the decoder. The decoder outputs the reconstructed organ mask. The processing system calculates the pixel-level cross-entropy loss or Dice loss between the reconstructed organ mask and the second organ mask as the reconstruction loss, and calculates the KL divergence between the deformation latent variable distribution and the standard Gaussian distribution as the regularization loss. The Dice loss can be a loss function based on the Dice similarity coefficient, used to measure the degree of overlap between the predicted organ delineation mask and the real organ delineation mask; a smaller Dice loss value indicates a higher degree of overlap between the two masks. This embodiment learns the probability distribution of organ deformation through a variational inference framework, which is beneficial for capturing the multimodal characteristics of organ deformation and for making the latent variable space continuous and having good interpolation properties.

[0059] In other embodiments, the encoder and decoder employ a reversible generative architecture based on a flow model. The processing system concatenates the first and second organ masks along the channel dimension and inputs them into the encoder. The encoder maps the input to a deformation latent variable distribution through a series of reversible transformations, setting this distribution to an isotropic Gaussian distribution. The processing system samples latent variables from this distribution and inputs them along with the first organ mask into the decoder. The decoder reconstructs the organ mask under the target respiratory phase through the inverse transformation of the encoder. The processing system calculates the mean square error between the reconstructed organ mask and the second organ mask as the reconstruction loss and the negative log-likelihood of the deformation latent variable distribution as the regularization loss. This embodiment achieves accurate latent variable inference and reconstruction through reversible transformations, which helps maintain the bijective relationship between the latent variable space and the data space, and improves the representation efficiency and reconstruction accuracy of deformation coding.

[0060] In some optional embodiments, the distance metric between the predicted latent variable distribution and the true latent variable distribution employs at least one of the following: KL divergence, used to measure the relative entropy difference between the predicted and true latent variable distributions; symmetric KL divergence, the average of the KL divergence and the inverse KL divergence, used to symmetrically measure the difference between the predicted and true latent variable distributions; Wasserstein distance, used to measure the minimum transport cost of transforming one distribution into the other between the predicted and true latent variable distributions; maximum mean difference, used to measure the mean difference between the predicted and true latent variable distributions in the reproducing kernel Hilbert space; cosine distance, used to measure the cosine of the angle between the vectors corresponding to the predicted and true latent variable distributions; and Mahalanobis distance, used to measure the distance between the predicted and true latent variable distributions considering covariance.

[0061] For example, the distance between the predicted latent variable distribution and the true latent variable distribution can be calculated using various metrics. KL divergence can be represented as a measure of the relative entropy difference between two probability distributions; it calculates the amount of information lost when using the predicted latent variable distribution to approximate the true latent variable distribution. Symmetric KL divergence can be represented as the average of the KL divergence and the inverse KL divergence; it symmetrically measures the bidirectional difference between the predicted and true latent variable distributions. The Wasserstein distance can be represented as a measure of the minimum transportation cost required to transform the predicted latent variable distribution into the true latent variable distribution. Maximum mean difference can be represented as a measure of the difference in means between the predicted and true latent variable distributions in the reproducing kernel Hilbert space. Cosine distance can be represented as a measure of the cosine of the angle between the mean vectors of the predicted and true latent variable distributions. Mahalanobis distance can be used to measure the distance between the predicted distribution of latent deformation variables and the actual distribution of latent deformation variables, taking into account the covariance structure. The processing system selects at least one of the above distance metrics to construct the deformation-constrained loss function based on the specific application scenario. The embodiments of this application facilitate the selection of appropriate distance metrics according to the characteristics of the distribution of latent deformation variables in different application scenarios, and facilitate a flexible balance between computational efficiency, gradient smoothness, and the ability to represent distribution differences, thereby improving the guiding effect of the deformation-constrained loss function on the training process.

[0062] In some embodiments, the processing system uses KL divergence as a distance metric between the predicted and true latent deformation variable distributions. When both distributions are multidimensional Gaussian distributions, the KL divergence can be directly calculated using the distribution's mean vector and covariance matrix to obtain a closed-form solution. The processing system calculates the KL divergence value of the predicted latent deformation variable distribution relative to the true latent deformation variable distribution and uses this KL divergence value as the deformation constraint loss function. This embodiment constrains the relative entropy between the two distributions using KL divergence, which is beneficial for making the predicted distribution approximate the true distribution from a probability density perspective during training and for simultaneously constraining the distribution's mean and covariance structure.

[0063] In other embodiments, the processing system uses the Wasserstein distance as a distance metric between the predicted and actual deformation latent variable distributions. The system calculates the minimum transportation cost required to transform the predicted deformation latent variable distribution into the actual distribution by solving an optimal transportation problem, and uses this minimum transportation cost as the deformation constraint loss function. This embodiment's use of the Wasserstein distance helps provide meaningful gradient information even when the support sets of the two distributions do not overlap, avoids gradient vanishing or numerical instability issues that may occur with KL divergence, and improves the stability and convergence of the training process.

[0064] In some optional embodiments, the training process of the deep learning generative model includes the following steps: obtaining an overlap loss function, wherein the overlap loss function is used to constrain the degree of overlap between the predicted organ delineation mask and the real organ delineation mask; obtaining a boundary loss function, wherein the boundary loss function is used to constrain the distance between the boundary of the predicted organ delineation mask and the boundary of the real organ delineation mask; weighted summing of the overlap loss function, the boundary loss function, and the deformation constraint loss function to obtain a total loss function; and iteratively training the neural network with the goal of minimizing the total loss function until the convergence condition is met to obtain the deep learning generative model.

[0065] For example, the overlap loss function can be represented as a loss term that measures the degree of overlap between the predicted organ delineation mask and the real organ delineation mask. This overlap loss function can employ Dice loss or Soft-Intersection over Union loss. The boundary loss function can be represented as a loss term that measures the distance between the boundary of the predicted organ delineation mask and the boundary of the real organ delineation mask. This boundary loss function is used to enhance the model's prediction sensitivity for organ edge regions. The deformation constraint loss function can be represented as a loss term constructed in the preceding steps to constrain the distribution of the predicted deformation latent variables to approximate the distribution of the real deformation latent variables. The processing system performs a weighted summation of the overlap loss function, the boundary loss function, and the deformation constraint loss function to obtain the total loss function. Iterative training of the neural network is then performed with the goal of minimizing the total loss function until convergence is achieved, resulting in a deep learning generative model. The embodiments of this application are beneficial to simultaneously constrain the prediction results of the deep learning generative model from three dimensions: voxel-level overlap accuracy, organ boundary localization accuracy, and overall deformation physiological rationality. This is beneficial to enabling the trained model to generate organ deformation results that conform to physiological laws while maintaining high segmentation accuracy, thereby improving the reliability and usability of the model in clinical radiotherapy scenarios.

[0066] For example, after the organ deformation encoding model is trained, its encoder parameters are frozen, and this encoder is used as a deformation feature constraint during the generative model training phase. This is for the predicted organ mask under the target respiratory phase predicted by the deep learning generative model. The processing system can use reference organ masks. After concatenation with the prediction mask, the result is input into the frozen organ deformation encoder to obtain the distribution of the predicted deformation latent variables, which can be represented as follows: At the same time, a reference organ mask will be used. Compared to real organ masking After concatenation, the data is input into the same encoder to obtain the true distribution of latent deformation variables, which can be represented as follows: ;

[0067] Subsequently, the processing system can use at least one of the following distance metrics: KL divergence, symmetric KL divergence, Wasserstein distance, maximum mean difference, or latent variable mean-variance distance. Calculate the deformation constraint loss to constrain the distribution of latent deformation variables in the prediction. Approximate distribution of latent variables of deformation The deformation distribution constraint loss can be expressed as: ;in, This represents the distance metric for the latent variable distribution. Ultimately, the training loss of the generative model includes the basic delineation loss and the deformation distribution constraint loss, and the total loss function can be expressed as:

[0068] ;

[0069] in, It can represent the overlap loss function, which measures the degree of overlap between the predicted organ delineation mask and the real organ delineation mask. The smaller the value, the better the overlap. It can represent a boundary loss function, used to constrain the distance between the boundary of the predicted organ delineation mask and the boundary of the real organ delineation mask. The smaller the value, the more accurate the boundary localization. It can represent a deformation constraint loss function, which is based on the distance between the predicted latent deformation variable distribution and the actual latent deformation variable distribution, and is used to constrain the predicted organ deformation to conform to the actual physiological laws. and These are preset weighting coefficients.

[0070] In some embodiments, the processing system sets the overlap loss function to the Dice loss function, which calculates the Dice similarity coefficient between the predicted organ delineation mask and the real organ delineation mask, and uses the result of subtracting the Dice similarity coefficient as the loss value. The processing system sets the boundary loss function to a level set-based distance loss function, which calculates the signed distance from the zero level set of the predicted organ delineation mask to the boundary of the real organ delineation mask. The processing system performs a weighted sum of the overlap loss function, boundary loss function, and deformation constraint loss function according to preset weight coefficients, where the weight coefficient of the overlap loss function is set to 1, the weight coefficient of the boundary loss function is set to a value between 0.1 and 1, and the weight coefficient of the deformation constraint loss function is set to a value between 0.01 and 0.1. This embodiment, by setting weight coefficients of different magnitudes, helps to balance the contribution ratio of the three loss functions in the total loss, and helps to avoid one loss function dominating the training process while weakening the constraint effect of other loss functions.

[0071] In other embodiments, the processing system employs a dynamic weight adjustment strategy during training. Initially, the system sets the weight coefficients of the deformation constraint loss function to a low value, allowing the neural network to prioritize learning basic organ segmentation capabilities. As training iterations increase, the system gradually increases the weight coefficients of the deformation constraint loss function, enabling the neural network to gradually enhance its learning of the physiological rationality of deformation after acquiring basic segmentation capabilities. The system monitors the overlap loss function and deformation constraint loss function values ​​on the validation set, and begins to gradually increase the weight coefficients of the deformation constraint loss function when the decline in the overlap loss function value becomes gradual. This embodiment, by dynamically adjusting the weights of the loss function, helps the neural network focus on different types of supervisory signals at different training stages. It helps avoid the model's learning of basic segmentation tasks being negatively impacted by excessively strong deformation constraints in the early stages of training, and also helps strengthen the model's learning effect on physiological deformation patterns in the later stages of training.

[0072] In some optional embodiments, the deep learning generative model is a model built on any of the following architectures: flow matching technology, which generates the target result by constructing a continuous deformation path from the input features to the output feature space and predicting the velocity field on the path; a diffusion model, which generates the target result by progressively adding noise to the data and learning the inverse diffusion process; a generative adversarial network, which generates the target result through game training between the generator and the discriminator; and an encoder and decoder structure, which extracts features through multi-scale downsampling and performs upsampling reconstruction with skip connections to preserve multi-scale detail information. The flow matching generative model predicts the organ delineation mask by performing the following steps in sequence: jointly encoding the surface deformation field and the reference organ mask to obtain joint encoded features; constructing a continuous deformation path from the joint encoded features to the feature space corresponding to the organ delineation mask; predicting the velocity field corresponding to each time step at multiple consecutive time steps on the continuous deformation path; generating a latent space feature representation based on the velocity field integral; and decoding the latent space feature representation to obtain the organ delineation mask.

[0073] For example, flow matching can represent a generative model architecture that generates a target result by constructing a continuous deformation path from the input feature space to the output feature space and predicting the velocity field along that path. A diffusion model can represent a generative model architecture that generates a target result by progressively adding noise to the data and learning an inverse diffusion process. A generative adversarial network (GAN) can represent a generative model architecture that generates a target result through game-like training between a generator and a discriminator. An encoder-decoder architecture can represent a generative model architecture that extracts features through multi-scale downsampling and performs upsampling reconstruction using skip connections to preserve multi-scale detail. When the processing system uses a generative model employing flow matching technology to predict organ delineation masks, it sequentially performs the following steps: jointly encoding the surface deformation field and the reference organ mask to obtain joint encoded features; constructing a continuous deformation path from the joint encoded features to the feature space corresponding to the organ delineation mask; predicting the velocity field corresponding to each time step at multiple consecutive time steps along the continuous deformation path; generating a latent space feature representation based on the velocity field integral; and decoding the latent space feature representation to obtain the organ delineation mask. This embodiment of the application is advantageous in utilizing flow matching technology to generate continuous and smooth deformation paths, and in obtaining generation results with good temporal consistency through velocity field integration, thereby improving the prediction quality and anatomical rationality of the organ delineation mask.

[0074] In some embodiments, the processing system employs a generative model using flow matching technology, wherein the joint feature encoder uses a 3D convolutional neural network to extract multi-scale features from the body surface deformation field and the reference organ mask. The processing system uses the jointly encoded features as the starting point of the continuous deformation path and random sampling points in a standard Gaussian distribution as the feature representation of the target point, constructing a probabilistic path from the starting point to the target point. The processing system samples multiple consecutive time steps at equal intervals within a time interval, predicting the velocity field corresponding to the current time step through a velocity field prediction network at each time step. The processing system integrates the velocity field using an ordinary differential equation solver, obtaining the endpoint feature representation by integrating along the velocity field direction from the starting point; this endpoint feature representation serves as the latent space feature representation. The processing system maps the latent space feature representation to an organ delineation mask through a decoder network. This embodiment achieves continuous transformation from input to output through ordinary differential equation integration, which helps maintain the numerical stability and reversibility of the generation process and improves the fidelity of the generation result.

[0075] In other embodiments, the processing system employs a generative model with an encoder and decoder structure. This encoder and decoder structure includes an encoder part and a decoder part. The encoder part consists of multiple downsampling convolutional blocks, each containing a convolutional layer, a batch normalization layer, and an activation function layer, used to progressively reduce the spatial resolution of the feature map and increase the number of channels. The decoder part consists of multiple upsampling convolutional blocks, each containing a transposed convolutional layer, a batch normalization layer, and an activation function layer. The encoder and decoder parts are connected via skip connections to concatenate the intermediate feature maps from the encoder part's downsampling process with the corresponding scale's upsampling feature maps from the decoder part. The processing system inputs the concatenated surface deformation field and reference organ delineation along the channel dimension into the generative model of the encoder and decoder structure. This generative model outputs an organ delineation mask under the target respiratory phase. This embodiment, through the encoder and decoder structure and skip connections, helps to preserve multi-scale detail information in the input data when generating organ delineation masks, which is beneficial for improving the localization accuracy of organ boundaries and the segmentation effect of small organs.

[0076] In some optional embodiments, the reference organ mask includes a segmentation mask of at least one target organ and at least one organ at risk, and the organ delineation mask includes a simultaneous segmentation mask of the target organ and the organ at risk.

[0077] For example, the target organ can refer to the tumor area requiring radiotherapy, such as a lung tumor, liver tumor, or prostate tumor. The organ at risk can refer to normal tissues and organs surrounding the tumor that are sensitive to radiation; for example, in thoracic radiotherapy, the organ at risk may include the heart, lungs, esophagus, and spinal cord. The reference organ mask includes a segmentation mask for at least one target organ and a segmentation mask for at least one organ at risk. The organ delineation mask includes simultaneous segmentation masks for the target organ and the organ at risk at the target respiratory phase. The processing system takes the reference organ delineation containing multiple organ segmentation masks as input and simultaneously outputs the segmentation masks for the target organ and the organ at risk at the target respiratory phase using a deep learning generation model. The embodiments of this application facilitate simultaneous real-time monitoring of multiple organs, and are beneficial for obtaining the deformation state of surrounding normal organs while predicting the tumor target location. This provides more comprehensive anatomical information for adaptive radiotherapy and facilitates simultaneous assessment of target dose coverage and radiation risk to organs at risk.

[0078] In some embodiments, the organ delineation mask output by the processing system includes a segmentation mask for a target organ and segmentation masks for multiple organs at risk, which may include organs such as the heart, lungs, esophagus, spinal cord, and liver. The processing system encodes the segmentation mask for each organ into an independent channel; that is, the organ delineation mask is a multi-channel three-dimensional mask tensor, with each channel corresponding to the segmentation result of one organ. This embodiment, through multi-channel output, facilitates the deep learning generative model to simultaneously learn the deformation patterns of multiple organs and leverages the correlation of motion between different organs to improve prediction accuracy.

[0079] In other embodiments, the organ delineation mask output by the processing system includes target organs such as primary tumor target areas and metastatic lymph node target areas, while organs at risk are dynamically determined based on the radiotherapy site. The processing system can predefine segmentation masks for multiple candidate organs at risk in the reference organ delineation. During processing, it automatically selects organs at risk whose spatial distance from the target area is less than a preset threshold as those requiring simultaneous prediction, ignoring organs at greater distances. This embodiment helps reduce redundant information in the model output, concentrates computational resources on organs at risk highly correlated with the radiotherapy plan, and improves model inference efficiency.

[0080] In some optional embodiments, after the organ delineation mask is predicted, at least one of the following operations is included: performing beam gating based on the organ delineation mask and pausing beam emission when the tumor location is detected to exceed a preset safety boundary; performing multi-leaf grating tracking based on the organ delineation mask to dynamically adjust the leaf positions so that the beam covers the target area; performing dose accumulation calculation based on the organ delineation mask and determining whether the radiotherapy plan needs to be optimized based on the calculation results.

[0081] For example, beam gating can represent an operation that controls the switching of a radiotherapy beam based on the target organ position in an organ delineation mask, wherein the beam gating operation suspends beam output when the target organ position is detected to be outside a preset safety boundary. Multileaf grating tracking can represent an operation that dynamically adjusts the position of multileaf grating blades based on the target organ contour in an organ delineation mask, wherein the multileaf grating tracking operation is used to match the beam's irradiation field shape with the projected contour of the target organ under the target respiratory phase. Dose accumulation calculation can represent a calculation process that calculates the cumulative radiation dose received by the target organ and organs at risk based on organ delineation masks under multiple respiratory phases, wherein the dose accumulation calculation process determines whether the radiotherapy plan needs optimization based on the calculation results. The processing system sends the organ delineation mask predicted by the deep learning generative model to the processing system in real time, and the processing system performs at least one of the following operations based on the organ delineation mask: beam gating, multileaf grating tracking, or dose accumulation calculation. The embodiments of this application facilitate the direct application of predicted organ delineation masks to various stages of radiotherapy control, enabling the formation of a closed-loop adaptive radiotherapy system from surface monitoring to organ prediction and then to beam control, and improving the response speed and accuracy of radiotherapy to respiratory movements.

[0082] In some embodiments, the processing system simultaneously performs beam gating and multi-leaf grating tracking operations based on the organ delineation mask. The processing system monitors the centroid position of the target organ in the organ delineation mask in real time. When the offset between the centroid position of the target organ and the reference position exceeds a preset safety boundary threshold, the processing system sends a gating pause signal to the accelerator control system to pause beam output. When the offset between the centroid position of the target organ and the reference position is less than or equal to the preset safety boundary threshold, the processing system calculates the leaf position sequence of the multi-leaf grating based on the contour shape of the target organ in the organ delineation mask and sends this leaf position sequence to the multi-leaf grating controller to drive the leaves to dynamically follow the target area movement. This embodiment, through the coordinated operation of gating and tracking, helps to avoid off-target irradiation when the target area movement amplitude is large, and achieves continuous and precise irradiation through dynamic tracking when the target area movement amplitude is small, thus improving treatment efficiency while ensuring treatment safety.

[0083] In other embodiments, the processing system performs dose accumulation calculations based on organ delineation masks. The processing system acquires the organ delineation mask for the target respiratory phase and the beam intensity distribution data corresponding to that respiratory phase, and calculates the instantaneous dose distribution for that respiratory phase. The processing system accumulates the instantaneous dose distributions for multiple respiratory phases according to the duration weight of each respiratory phase to obtain the cumulative dose distribution. The processing system compares the cumulative dose distribution with the planned dose distribution of the reference radiotherapy plan, calculating the dose coverage deviation of the target organs and the dose overshoot deviation of endangered organs. When the dose coverage deviation of the target organs exceeds a preset dose deviation threshold or the dose overshoot deviation of endangered organs exceeds a preset safety threshold, the processing system determines that the radiotherapy plan needs to be optimized and triggers a re-optimization or adjustment of the radiotherapy plan. This embodiment facilitates real-time dose monitoring and feedback evaluation during treatment, facilitates timely detection of intra-fraction dose deviations and adaptive adjustments, and improves the execution accuracy and personalization level of the radiotherapy plan.

[0084] For example, Figure 2 The flowchart illustrates a respiratory-based organ motion control method provided in an embodiment of this application. In the offline phase, the processing system acquires organ masks from historical training data, pre-trains an organ deformation encoder, freezes the encoder parameters after training, and uses the frozen-parameter organ deformation encoder as a deformation constraint during the training phase. During the training phase, the processing system acquires reference and target body surface point clouds, reconstructs reference and target body surface masks respectively, and generates a body surface deformation field as a geometric guidance signal after registration. The processing system inputs the reference organ delineation and the body surface deformation field into a deep learning generative model to obtain a predicted organ mask. Simultaneously, the processing system concatenates the reference organ mask with the real organ mask and inputs it into the frozen organ deformation encoder to obtain the real deformation latent variable distribution; it also concatenates the reference organ mask with the predicted organ mask and inputs it into the same encoder to obtain the predicted deformation latent variable distribution. The processing system constructs a deformation constraint loss function with the objective of minimizing the distance between the real and predicted deformation latent variable distributions, which is used to train the deep learning generative model. During the inference phase, the processing system inputs the reference organ delineation and body surface deformation field into the trained deep learning generative model, and outputs the organ delineation mask under the target respiratory phase.

[0085] See Figure 3According to another aspect of the embodiments of this application, a respiratory-based organ motion control device is also provided, comprising: a first splicing processing unit, configured to use a pre-trained organ deformation encoder to splice a reference organ mask and a predicted organ mask into a vector distribution for predicting organ deformation characteristics, wherein the predicted organ deformation latent variable distribution is a vector distribution output by the organ deformation encoder after encoding the reference organ mask and the predicted organ mask; and a second splicing processing unit, configured to splice a reference organ mask and a real organ mask into the same organ deformation encoder to obtain a real organ deformation latent variable distribution, wherein the real organ deformation latent variable distribution refers to the distribution obtained by... The organ deformation encoder encodes a vector distribution that represents the deformation features of the real organ after encoding the reference organ mask and the real organ mask; the function construction unit constructs a deformation constraint loss function with the objective of minimizing the distance between the predicted deformation latent variable distribution and the real deformation latent variable distribution; the acquisition unit acquires the surface deformation field of the target object, where the surface deformation field represents the deformation of the target object's surface from the reference respiratory phase to the target respiratory phase; and the mask processing unit inputs the surface deformation field into the deep learning generative model trained by the deformation constraint loss function to predict the organ delineation mask of the target object under the target respiratory phase.

[0086] Optionally, the mask processing unit includes: a reference delineation acquisition subunit, used to acquire reference organ delineations of the target object under a reference respiratory phase, wherein the reference organ delineation includes segmentation masks of at least one target organ and at least one organ at risk; and a common input subunit, used to input the reference organ delineation and the body surface deformation field into a deep learning generative model, and to perform deformation prediction on the reference organ delineation based on the body surface deformation field through the deep learning generative model, and output the organ delineation mask of the target object under the target respiratory phase.

[0087] Optionally, the acquisition unit includes: a point cloud acquisition subunit, used to acquire a reference body surface point cloud of the target object under a reference respiratory phase, and a target body surface point cloud under a target respiratory phase; a mask reconstruction subunit, used to reconstruct a reference body surface mask based on the reference body surface point cloud, and to reconstruct a target body surface mask based on the target body surface point cloud; and a registration subunit, used to register the reference body surface mask and the target body surface mask to obtain the body surface deformation field.

[0088] Optionally, the organ deformation encoder is obtained through the following pre-training module, which includes: a historical data acquisition unit, used to acquire the first organ mask of the organ at the reference respiratory phase and the second organ mask at the target respiratory phase from the historical training data; an encoding output unit, used to concatenate the first organ mask and the second organ mask in the channel dimension and input them into the encoder, outputting the distribution of latent deformation variables of the organ from the reference phase to the target phase; a sampling unit, used to sample latent variables from the distribution of latent deformation variables, wherein the latent variables are latent space vectors characterizing the organ deformation features; and a decoding and reconstruction unit, used to convert the latent variables into latent space vectors. The encoder and decoder are input with the first organ mask and the decoder reconstructs the organ mask under the target respiratory phase. The loss calculation unit is used to calculate the reconstruction loss between the reconstructed organ mask and the second organ mask, and to calculate the regularization loss between the distribution of deformation latent variables and the preset prior distribution. The training and freezing unit is used to train the encoder and decoder with the goal of minimizing the reconstruction loss and regularization loss, so that the encoder and decoder learn the true deformation distribution of different organs in different respiratory phases, until the training convergence condition is met, then the encoder parameters are frozen, and the trained encoder is used as the organ deformation encoder.

[0089] Optionally, the distance metric between the predicted latent variable distribution and the actual latent variable distribution employs at least one of the following: a KL divergence module to measure the relative entropy difference between the predicted and actual latent variable distributions; a symmetric KL divergence module to symmetrically measure the difference between the predicted and actual latent variable distributions, wherein the symmetric KL divergence is the average of the KL divergence and the reverse KL divergence; a Wasserstein distance module to measure the minimum transportation cost of transforming one distribution into the other between the predicted and actual latent variable distributions; a maximum mean difference module to measure the mean difference between the predicted and actual latent variable distributions in the reproducing kernel Hilbert space; a cosine distance module to measure the cosine of the angle between the vectors corresponding to the predicted and actual latent variable distributions; and a Mahalanobis distance module to measure the distance between the predicted and actual latent variable distributions considering covariance.

[0090] Optionally, the training process of the deep learning generative model includes the following modules: an overlap loss acquisition module, used to acquire an overlap loss function, wherein the overlap loss function is used to constrain the degree of overlap between the predicted organ delineation mask and the real organ delineation mask; a boundary loss acquisition module, used to acquire a boundary loss function, wherein the boundary loss function is used to constrain the distance between the boundary of the predicted organ delineation mask and the boundary of the real organ delineation mask; a weighted summation module, used to weightedly sum the overlap loss function, the boundary loss function, and the deformation constraint loss function to obtain the total loss function; and an iterative training module, used to iteratively train the neural network with the goal of minimizing the total loss function until the convergence condition is met to obtain the deep learning generative model.

[0091] Optionally, the deep learning generative model is a model built based on any of the following architectures: a flow matching model, which generates the target result by constructing a continuous deformation path from the input features to the output feature space and predicting the velocity field on the path; a diffusion model, which generates the target result by progressively adding noise to the data and learning the inverse diffusion process; a generative adversarial network model, which generates the target result through game training between the generator and the discriminator; and an encoder-decoder structure model, which extracts features through multi-scale downsampling and performs upsampling reconstruction with skip connections to preserve multi-scale detail information. The flow matching model includes the following sub-modules: a joint encoding sub-module, used to jointly encode the surface deformation field and the reference organ mask to obtain joint encoded features; a path construction sub-module, used to construct a continuous deformation path from the joint encoded features to the feature space corresponding to the organ delineation mask; a velocity field prediction sub-module, used to predict the velocity field corresponding to each time step at multiple consecutive time steps on the continuous deformation path; an integral generation sub-module, used to generate a latent space feature representation based on the velocity field integral; and a decoding sub-module, used to decode the latent space feature representation to obtain the organ delineation mask.

[0092] Optionally, the reference organ mask includes a segmentation mask for at least one target organ and at least one organ at risk, and the organ delineation mask includes a simultaneous segmentation mask for the target organ and the organ at risk.

[0093] Optionally, the device further includes at least one of the following operating units: a beam gating unit, used to perform beam gating operation based on the organ delineation mask after the organ delineation mask is predicted, and to pause beam emission when the tumor location is detected to exceed a preset safety boundary; a multi-leaf grating tracking unit, used to perform multi-leaf grating tracking operation based on the organ delineation mask, and dynamically adjust the leaf positions to make the beam cover the target area; and a dose accumulation and optimization unit, used to perform dose accumulation calculation based on the organ delineation mask, and to determine whether the radiotherapy plan needs to be optimized based on the calculation results.

[0094] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, which stores a computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located executes the above-described respiratory-based organ movement control method.

[0095] According to another aspect of the embodiments of this application, an electronic device is also provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors cause the one or more processors to perform the above-described respiratory-based organ movement control method.

[0096] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program or instructions, which, when executed by a processor, implement the above-described respiratory-based organ movement control method.

[0097] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0098] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0099] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0100] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0101] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0102] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0103] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for organ movement control based on respiration, characterized in that, include: Using a pre-trained organ deformation encoder, the reference organ mask and the predicted organ mask are concatenated and input to obtain the predicted deformation latent variable distribution. The predicted deformation latent variable distribution is a vector distribution that is output by the organ deformation encoder after encoding the reference organ mask and the predicted organ mask, and is used to characterize the predicted organ deformation features. The reference organ mask and the real organ mask are concatenated and input into the same organ deformation encoder to obtain the real deformation latent variable distribution. The real deformation latent variable distribution refers to the vector distribution output by the organ deformation encoder after encoding the reference organ mask and the real organ mask, which is used to characterize the deformation features of the real organ. A deformation constraint loss function is constructed with the objective of minimizing the distance between the predicted latent deformation variable distribution and the actual latent deformation variable distribution; Obtain the surface deformation field of the target object, wherein the surface deformation field characterizes the deformation of the target object's surface from a reference respiratory phase to a target respiratory phase; The body surface deformation field is input into a deep learning generative model trained by the deformation constraint loss function to predict the organ delineation mask of the target object under the target respiratory phase.

2. The method according to claim 1, characterized in that, The surface deformation field is input into a deep learning generative model trained using the deformation constraint loss function to predict the organ delineation mask of the target object under the target respiratory phase, including: Obtain a reference organ delineation of the target object under the reference respiratory phase, wherein the reference organ delineation includes a segmentation mask of at least one target organ and at least one organ at risk. The reference organ delineation and the body surface deformation field are input into the deep learning generation model. The deep learning generation model performs deformation prediction on the reference organ delineation based on the body surface deformation field and outputs the organ delineation mask of the target object under the target respiratory phase.

3. The method according to claim 1, characterized in that, Obtain the surface deformation field of the target object, including: The reference body surface point cloud of the target object under the reference respiratory phase and the target body surface point cloud under the target respiratory phase are acquired. A reference body surface mask is obtained by reconstructing the reference body surface point cloud, and a target body surface mask is obtained by reconstructing the target body surface point cloud. The reference surface mask and the target surface mask are registered to obtain the surface deformation field.

4. The method according to claim 1, characterized in that, The organ deformation encoder is obtained through the following pre-training steps: Obtain the first organ mask of the organ in the reference respiratory phase and the second organ mask in the target respiratory phase from the historical training data; The first organ mask and the second organ mask are concatenated in the channel dimension and then input into the encoder to output the deformation latent variable distribution of the organ from the reference phase to the target phase; Latent variables are sampled from the distribution of latent deformation variables, wherein the latent variables are latent space vectors characterizing the deformation features of organs; The latent variable and the first organ mask are input into the decoder, and the decoder reconstructs the organ mask of the organ under the target respiratory phase. Calculate the reconstruction loss between the reconstructed organ mask and the second organ mask, and calculate the regularization loss between the deformation latent variable distribution and the preset prior distribution; With the goal of minimizing the reconstruction loss and the regularization loss, the encoder and the decoder are trained to learn the true deformation distribution of different organs between different respiratory phases. The parameters of the encoder are frozen after the training convergence condition is met, and the trained encoder is used as the organ deformation encoder.

5. The method according to claim 1, characterized in that, The distance metric between the predicted latent deformation variable distribution and the actual latent deformation variable distribution is at least one of the following: KL divergence is used to measure the relative entropy difference between the predicted latent deformation variable distribution and the actual latent deformation variable distribution; Symmetric KL divergence, which is the average of the KL divergence and the reverse KL divergence, is used to symmetrically measure the difference between the predicted latent deformation variable distribution and the actual latent deformation variable distribution. Wasserstein distance is used to measure the minimum transportation cost between the predicted latent deformation variable distribution and the actual latent deformation variable distribution to transform one distribution into the other. The maximum mean difference is used to measure the difference between the mean of the predicted latent deformation variable distribution and the mean of the actual latent deformation variable distribution in the reproducing kernel Hilbert space. Cosine distance is used to measure the cosine value of the angle between the vectors corresponding to the predicted latent deformation variable distribution and the actual latent deformation variable distribution, respectively. Mahalanobis distance is used to measure the distance between the predicted latent deformation variable distribution and the actual latent deformation variable distribution, taking into account covariance.

6. The method according to claim 1, characterized in that, The training process of the deep learning generative model includes the following steps: Obtain the overlap loss function, wherein the overlap loss function is used to constrain the degree of overlap between the predicted organ delineation mask and the real organ delineation mask; Obtain the boundary loss function, wherein the boundary loss function is used to constrain the distance between the boundary of the predicted organ delineation mask and the boundary of the real organ delineation mask; The total loss function is obtained by weighted summing of the overlap loss function, the boundary loss function, and the deformation constraint loss function. With the goal of minimizing the total loss function, the neural network is iteratively trained until the convergence condition is met, thus obtaining the deep learning generative model.

7. The method according to claim 1, characterized in that, The deep learning generative model is a model built based on any of the following architectures: Stream matching technology generates target results by constructing a continuous deformation path from the input feature space to the output feature space and predicting the velocity field along the path; The diffusion model generates the target result by progressively adding noise to the data and learning the inverse diffusion process; Generative adversarial networks (GANs) generate target results through game-like training between a generator and a discriminator. The encoder and decoder structure extracts features through multi-scale downsampling and performs upsampling reconstruction in conjunction with skip connections to preserve multi-scale detail information; The generative model of the streaming matching technique predicts the organ delineation mask through the following steps performed sequentially: The body surface deformation field and the reference organ mask are jointly coded to obtain joint coded features; Construct a continuous deformation path from the joint encoded features to the feature space corresponding to the organ delineation mask; Predict the velocity field corresponding to each time step at multiple consecutive time steps along the continuous deformation path; The latent space feature representation is generated based on the velocity field integral; The latent space feature representation is decoded to obtain the organ delineation mask.

8. The method according to claim 1, characterized in that, The reference organ mask includes a segmentation mask for at least one target organ and at least one organ at risk, and the organ delineation mask includes a synchronous segmentation mask for the target organ and the organ at risk.

9. The method according to claim 1, characterized in that, After the organ delineation mask is obtained through prediction, at least one of the following operations is also included: The beam gating operation is performed according to the organ delineation mask, and the beam emission is paused when the tumor location is detected to exceed the preset safety boundary. Perform multi-leaf grating tracking operation based on the organ delineation mask to dynamically adjust the leaf position so that the beam covers the target area; Dose accumulation calculation is performed based on the organ delineation mask, and the results are used to determine whether the radiotherapy plan needs to be optimized.

10. A respiratory-based organ movement control device, characterized in that, include: The first splicing processing unit is used to use a pre-trained organ deformation encoder to splice a reference organ mask and a predicted organ mask into a predicted deformation latent variable distribution, wherein the predicted deformation latent variable distribution is a vector distribution used to characterize the predicted organ deformation features, output by the organ deformation encoder after encoding the reference organ mask and the predicted organ mask. The second splicing processing unit is used to splice the reference organ mask and the real organ mask and input them into the same organ deformation encoder to obtain the real deformation latent variable distribution. The real deformation latent variable distribution refers to the vector distribution output by the organ deformation encoder after encoding the reference organ mask and the real organ mask, which is used to characterize the deformation features of the real organ. The function construction unit is used to construct a deformation constraint loss function with the objective of minimizing the distance between the predicted latent deformation variable distribution and the actual latent deformation variable distribution; The acquisition unit is used to acquire the surface deformation field of the target object, wherein the surface deformation field characterizes the deformation of the target object's surface from the reference respiratory phase to the target respiratory phase; The mask processing unit is used to input the body surface deformation field into the deep learning generative model trained by the deformation constraint loss function, and predict the organ delineation mask of the target object under the target respiratory phase.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed, the device on which the computer-readable storage medium is located performs the respiratory-based organ movement control method according to any one of claims 1 to 9.

12. An electronic device, characterized in that, It includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to perform the respiratory-based organ movement control method according to any one of claims 1 to 9.

13. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a processor, implement the respiratory-based organ movement control method of any one of claims 1 to 9.