Particle radiotherapy adaptive plan optimization method and related equipment

By constructing a multimodal deep learning network that integrates state representation and physical law constraints, the dose distribution in particle radiotherapy can be evaluated and adjusted in real time, solving the problem of dose deposition deviating from the target area in existing technologies and improving the accuracy and safety of particle radiotherapy.

CN121695435APending Publication Date: 2026-03-20TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610153042.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-03
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing particle radiotherapy methods struggle to achieve rapid and accurate dose distribution adjustments when faced with changes in anatomical structures and equipment deviations, resulting in dose deposition deviating from the preset target area and failing to support true online or quasi-online adaptive therapy.

Method used

By constructing a fusion state representation and utilizing a pre-trained dose optimization model, dose distribution is evaluated in real time based on dynamic treatment log data, and adaptive dose adjustment instructions, including speckle correction or plan re-optimization, are generated. A multimodal deep learning network constrained by physical laws is then used for precise evaluation and adjustment.

Benefits of technology

It enables real-time assessment and automatic adjustment of dose deviation during treatment, improving the accuracy and safety of radiotherapy, reducing treatment risks caused by dose deviation, and supporting dynamic adjustment of the current fraction and subsequent fractions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121695435A_ABST
    Figure CN121695435A_ABST
Patent Text Reader

Abstract

The invention discloses a particle radiotherapy adaptive plan optimization method and related equipment, and relates to the technical field of radiotherapy, and the method comprises the steps: constructing a fusion state representation reflecting the anatomical state and the treatment execution state of a patient at the same time according to the dynamic treatment log data generated by the patient in the particle radiotherapy execution process; inputting the fusion state representation into a pre-trained dose optimization model; the dose optimization model is configured to directly output a quantitative evaluation value of dose distribution generated by the executed part of the current treatment fraction according to the fusion state characterization; and based on the quantitative evaluation value, applying a preset decision rule, automatically generating and executing a self-adaptive dose adjustment instruction used for guiding the non-executed part of the current fractional treatment or the subsequent fractional treatment, thereby being capable of efficiently and accurately evaluating the dose distribution under the actual treatment condition, and improving the accuracy and safety of radiotherapy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of radiotherapy technology, and in particular to an adaptive planning optimization method and related equipment for particle radiotherapy. Background Technology

[0002] In pencil beam therapy, treatment planning is based on scanning beam spots as the basic execution unit. The treatment planning system precisely sets the target location, energy, weight, and scan timing for each beam spot based on the planned CT images, target area, and organ at risk contours, and generates a planned dose distribution that meets clinical dose constraints. However, in actual treatment fractions, the accuracy of the final deposited dose is affected by a variety of complex factors:

[0003] Due to factors such as respiratory movements, organ displacement, differences in daily positioning, tumor regression, or weight changes, the actual anatomical structure (such as tissue density, organ shape, and location) during treatment differs from the CT images used to develop the treatment plan. This variation alters the range (Bragg peak position) and scattering behavior of the particle beam within the body, causing the dose deposition to deviate from the preset target area. Furthermore, due to limitations in the control precision of the particle therapy equipment, system response time, and magnetic field stability, the actual irradiation parameters of the beam spot (including position, energy, weight, and irradiation sequence) may not be able to completely and accurately reproduce the planned settings, thus introducing direct dose delivery errors.

[0004] Existing particle-adaptive radiotherapy methods have the following main limitations when facing the above challenges:

[0005] The mainstream approach relies on re-acquiring patient images (such as CT scans) after a significant dose deviation is detected. Physicians then re-delineate the target area and organs at risk, and perform a complete dose recalculation and optimization within the treatment planning system. This process is time-consuming and labor-intensive, making it impossible to achieve rapid response within or between treatment sessions, and thus difficult to support true online or near-online adaptive therapy.

[0006] Limited Assessment Dimensions and Lack of Coupling: Existing techniques often treat anatomical changes or execution biases in isolation. For example, they may estimate the impact of anatomical changes on dose solely through image deformation or assess execution errors only by analyzing log files. However, the effects of anatomical state and execution bias on the final dose are highly nonlinear and interdependent, and isolated analysis cannot accurately reflect the combined dose consequences of both.

[0007] To accurately assess actual dose, physical principle-based dose calculations (such as Monte Carlo simulations or fast pencil beam algorithms) are required. While these calculations offer high accuracy, they are also computationally complex and time-consuming, making it difficult to meet the clinical needs for real-time or near-real-time dose assessment and decision-making. Summary of the Invention

[0008] Based on the above problems, the purpose of this invention is to provide a method and related equipment for adaptive planning optimization of particle radiotherapy, which can efficiently and accurately evaluate the dose distribution under actual treatment conditions, and can quickly generate decisions to guide treatment adjustments based on this evaluation, so as to overcome the problems of lag, one-sidedness and inefficiency of existing adaptive processes, and improve the accuracy and safety of radiotherapy.

[0009] In a first aspect, the present invention provides a method for adaptive planning optimization of particle radiotherapy, comprising:

[0010] Based on the dynamic treatment log data generated by the patient during particle radiotherapy, such as proton radiotherapy, a fusion state representation that simultaneously reflects the patient's anatomical state and treatment execution state is constructed.

[0011] The fusion state representation is input into a pre-trained dose optimization model; the dose optimization model is configured to directly output a quantitative evaluation value of the dose distribution generated by the portion of the current treatment fraction that has been performed, based on the fusion state representation.

[0012] Based on the quantitative evaluation value, a preset decision rule is applied to automatically generate and execute adaptive dose adjustment instructions to guide the unexecuted portion of the current fraction or subsequent fractions of treatment.

[0013] Preferably, the dynamic treatment log data includes at least the patient's anatomical imaging data, a time-corresponding sequence of treatment plan instructions, and a sequence of treatment device execution feedback; the anatomical imaging data includes planned CT images and target area and organ-at-risk structural information delineated based on the images; the treatment plan instruction sequence includes the target position, target energy, and target weight of multiple scanning beams; and the treatment device execution feedback sequence includes the actual position, actual energy, and actual weight corresponding to the scanning beams.

[0014] Preferably, the dose optimization model adopts a multimodal deep learning network architecture driven by physical constraints, including a dual-stream coding module, a physical information fusion module, and a physical constraint decoding module; the dual-stream coding module includes a parallel anatomical structure coding stream and a beam spot sequence coding stream.

[0015] Preferably, the anatomical structure encoding stream takes the anatomical image data as input and extracts multi-scale anatomical spatial features through a three-dimensional depth feature extraction network;

[0016] The beamspot sequence encoding stream takes the deviation sequence consisting of the target physical quantity and the actual executed physical quantity of each beamspot as input, models the temporal dependency of beamspot deviation through the temporal-spatial feature enhancement module, and enhances the feature expression of beamspot deviation events that have a preset correlation with the final dose distribution deviation.

[0017] Preferably, the physical information fusion module receives features from the dual-stream coding module and modulates the features through a differentiable proton transport approximation unit; the proton transport approximation unit generates a physical rationality attention weight map based on the tissue density and proton energy information implied in the current features, which is used to weight the fused features;

[0018] The physical constraint decoding module receives the modulated fused features, restores the spatial resolution through a three-dimensional deconvolutional network, and integrates a physical law verification layer during the decoding process. The physical law verification layer uses a parameterized proton depth dose curve model and a transverse scattering model to verify and adjust the intermediate feature map, and finally outputs the three-dimensional dose deviation distribution.

[0019] Preferably, the differentiable proton transport approximation unit is configured to learn a mapping function from the proton's initial energy and path organization density to its Bragg peak depth and width parameters, and dynamically generate the physical plausibility attention weight map using the parameters output by the function.

[0020] Preferably, the parameterized model upon which the physical law verification layer is based includes:

[0021] A Bragg peak analytical model based on water equivalent path length correction is used to constrain the dose deposition morphology along the beam spot range direction.

[0022] A transverse dose broadening model based on Fermi-Eitch scattering theory is used to constrain the dose distribution perpendicular to the beam spot direction.

[0023] Preferably, the training method for the dose optimization model includes:

[0024] Multiple fusion state representations from several historical treatment cases are obtained as input samples for the model;

[0025] For each input sample, a composite supervision target is generated, which integrates clinical dose difference signals, physical consistency constraint signals based on the physical laws of proton transport, and cyclic consistency constraint signals to ensure the closed loop of the evaluation logic.

[0026] The dose optimization model is iteratively trained with the training direction of minimizing the difference between the model output and the composite supervision objective.

[0027] During iterative training, the relative weights of the physical consistency constraint signal and the cyclic consistency constraint signal in the optimization objective are dynamically adjusted.

[0028] Preferably, the composite supervision objective is achieved by minimizing a multi-objective loss function, wherein the loss function is:

[0029]

[0030] Wherein, L_data is the data fitting loss term, which uses the dose difference distribution calculated by the fast pencil beam algorithm as the supervision label to calculate the difference between the predicted value and the label; L_physics is the physical consistency loss term, which is used to constrain the dose distribution predicted by the network to conform to the physical laws of proton transport on a macroscopic scale; L_cycle is the cycle consistency loss term; λ_phy is the weight of the physical consistency loss term; and λ_cyc is the weight of the cycle consistency loss term.

[0031] Preferably, the physical law consistency loss term includes a regional energy conservation sub-term, a depth dose curve morphology sub-term, and / or a lateral scattering constraint sub-term:

[0032] The regional energy conservation sub-term is achieved by minimizing the total integral of the predicted dose deviation within the planned target area, and the first difference measure between the weighted deviation statistics of all beam spots;

[0033] The depth dose profile morphology sub-item is achieved by minimizing a second difference measure between the predicted depth dose profile extracted along the beam spot path and the standard Bragg peak profile generated based on the beam spot energy parameterization;

[0034] The lateral scattering constraint is achieved by minimizing a third measure of difference between the predicted dose distribution at a specific depth lateral profile and a reference profile predicted based on the scattering model.

[0035] Preferably, the cycle consistency loss term is calculated in the following manner:

[0036] A reverse reasoning network is constructed, which is trained to infer beam spot execution bias characteristics by taking the sum of the original planned dose distribution and the dose deviation distribution predicted by the dose optimization model as input.

[0037] The cycle consistency loss term is defined as the difference between the predicted execution bias feature output by the inverse inference network and the speckle execution bias feature extracted from the real log data;

[0038] The inverse reasoning network and the dosage optimization model are jointly optimized during training.

[0039] Preferably, the quantitative evaluation value includes:

[0040] The difference between the actual delivered dose distribution and the planned dose per voxel; and / or,

[0041] Dosage statistics for each clinically relevant region, including target coverage metrics and dose-limiting metrics for organs at risk.

[0042] Preferably, the preset decision rules and corresponding adaptive dose adjustment instructions include at least one of the following:

[0043] If the target area dose coverage is lower than the preset minimum coverage requirement, or if the dose to any organ at risk exceeds the preset safety limit, the online adaptive adjustment process is triggered and a speckle correction instruction is generated.

[0044] If, based on the cumulative and trend analysis results of the quantitative assessment values ​​of multiple treatment fractions already performed, it is predicted that the cumulative total dose distribution of the patient's entire treatment course cannot meet the preset clinical overall goal, then the offline adaptive re-optimization process is triggered, and a treatment plan re-optimization instruction is generated.

[0045] Preferably, the beam spot correction command is generated by solving a constrained optimization problem, which aims to minimize the remaining beam spot weight adjustment while ensuring that the adjusted predicted dose meets clinical constraints.

[0046] In a second aspect, the present invention provides a particle radiotherapy adaptive planning optimization device, comprising:

[0047] The feature fusion module is used to construct a fusion state representation that simultaneously reflects the patient's anatomical state and treatment execution state based on the dynamic treatment log data generated by the patient during radiotherapy.

[0048] A quantitative evaluation module is used to input the fusion state representation into a pre-trained dose optimization model; the dose optimization model is configured to directly output a quantitative evaluation value of the dose distribution generated by the portion of the current treatment fraction that has been performed, based on the fusion state representation.

[0049] The optimization and adjustment module is used to automatically generate and execute adaptive dose adjustment instructions to guide the unexecuted portion of the current treatment or subsequent treatments, based on the quantitative evaluation value and by applying preset decision rules.

[0050] Thirdly, the present invention provides a radiotherapy system, comprising:

[0051] The particle radiotherapy adaptive planning optimization device described in this embodiment of the invention.

[0052] Fourthly, the present invention provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of any of the methods of the present invention or the functions of the apparatus of the present invention.

[0053] Fifthly, the present invention provides a computer-readable storage medium storing computer instructions, wherein when a computer reads the computer instructions, the computer performs the steps of any of the methods described in the present invention or the functions of the apparatus described in the present invention.

[0054] Compared with existing technologies, the beneficial effects of this invention include at least the following: Constructing a fusion state representation based on dynamic treatment log data accurately captures real-time changes in anatomical state and beam spot execution deviations, avoiding the disconnect between static planning and actual treatment. The model is calibrated using physical constraints (such as the Bragg peak model and transverse scattering model) and multi-objective loss functions to ensure that dose assessment and adjustment conform to the transmission physics of particles (e.g., protons), allowing for more precise dose focusing on the target area. Real-time output of voxel-level dose differences and key clinical indicators (target coverage, dose to organs at risk) promptly identifies dose deviation risks, avoiding over-irradiation or under-irradiation of the target area. Automatic online / offline adjustment processes are triggered for different deviation scenarios, maximizing the protection of normal tissues and reducing the incidence of complications through beam spot weight correction or plan re-optimization. Dynamic adjustments are supported during treatment (unexecuted portion of the current fraction) and across fractions (subsequent treatment sessions), addressing uncertainties such as anatomical changes and execution deviations without interrupting treatment and re-planning. By integrating multimodal data with prior physical knowledge, the model training takes into account clinical data fitting, consistency with physical laws, and logical closure, thereby improving the assessment accuracy for different patients and treatment scenarios. Attached Figure Description

[0055] Figure 1 This is a schematic diagram of the adaptive planning optimization method for particle radiotherapy according to an embodiment of the present invention. Detailed Implementation

[0056] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided to make the invention more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the drawings denote the same or similar structures, and therefore repeated descriptions of them will be omitted.

[0057] The terms used to express position and direction in this invention are illustrated with reference to the accompanying drawings, but changes can be made as needed, and all such changes are included within the scope of protection of this invention.

[0058] Example 1, refer to Appendix Figure 1 This embodiment provides a method for adaptive planning optimization of particle radiotherapy, including:

[0059] Based on the dynamic treatment log data generated by the patient during particle radiotherapy, such as proton radiotherapy, a fusion state representation that simultaneously reflects the patient's anatomical state and treatment execution state is constructed.

[0060] The fusion state representation is input into a pre-trained dose optimization model; the dose optimization model is configured to directly output a quantitative evaluation value of the dose distribution generated by the portion of the current treatment fraction that has been performed, based on the fusion state representation.

[0061] Based on the quantitative evaluation value, a preset decision rule is applied to automatically generate and execute adaptive dose adjustment instructions to guide the unexecuted portion of the current fraction or subsequent fractions of treatment.

[0062] The dynamic treatment log data includes at least the patient's anatomical imaging data, a time-corresponding sequence of treatment plan instructions, and a sequence of treatment device execution feedback. The anatomical imaging data includes planned CT images and target area and organ-at-risk structural information delineated based on these images. The treatment plan instruction sequence includes the target position, target energy, and target weight of multiple scanning beams. The treatment device execution feedback sequence includes the actual position, actual energy, and actual weight corresponding to the scanning beams. The treatment device can be a particle therapy device, the beam can be a beam of a particle beam, and the particles can be protons, carbon ions, or helium ions. In this embodiment, protons are used as an example to illustrate the technical solution. It should be noted that the technical solution is also applicable to carbon ions or helium ions; the proton mentioned in this specification or claims can be replaced with particles, carbon ions, or helium ions.

[0063] During proton radiotherapy, dynamic treatment log data is collected in real time. This data contains at least three types of core information: first, patient anatomical imaging data (including planning CT images and target area and organ-at-risk structural information delineated based on these images); second, treatment plan instruction sequence (including target location, target energy, and target weight for multiple scanning beam spots); and third, treatment equipment execution feedback sequence (including actual location, actual energy, and actual weight corresponding to each scanning beam spot).

[0064] The collected dynamic treatment log data is analyzed and its features are fused to construct a fused state representation that simultaneously reflects the patient's anatomical state and the treatment execution state. The core of this process is to associate and integrate the spatial structural features contained in the anatomical imaging data with the deviation features of the treatment plan instructions and equipment execution feedback to form a unified dimension and a feature expression that can be recognized by the model.

[0065] The constructed fusion state representation is input into a pre-trained dose optimization model. This model, having been trained in advance, is capable of mapping fusion state information to dose distribution changes and can directly output quantitative assessment values ​​of the dose distribution generated by the portion of the current treatment fraction that has been executed, accurately reflecting the difference between the actual delivered dose and the planned dose.

[0066] Based on the quantitative assessment values ​​output by the model, preset clinical decision rules (with target coverage, dose to organs at risk, and other clinical standards as the core basis) are applied to determine whether the current dose deviation needs to be adjusted. If the adjustment trigger condition is met, an adaptive dose adjustment instruction is automatically generated. This instruction is used to guide the optimization of the unexecuted portion of the current fraction or the subsequent fraction treatment plan, and the adjustment instruction is executed directly to complete the adaptive correction of the treatment plan.

[0067] The effects of the above technical solution are as follows:

[0068] By collecting dynamic data in real time during treatment and integrating anatomical and execution information, dose assessment is made more closely aligned with the patient's actual treatment scenario. This effectively avoids dose deviations caused by changes in anatomical structure or equipment execution errors, ensuring sufficient target area dose coverage and controllable dose to organs at risk. Breaking through the limitations of traditional fixed-plan radiotherapy, it can assess dose deviations in real time and automatically adjust them during treatment. Whether it's incomplete portions of the current fraction or subsequent fractions, targeted corrections can be made, significantly improving adaptability to individual patient differences and dynamic changes during treatment. Pre-set decision rules based on clinical standards ensure the scientific and compliant nature of adaptive adjustments. A closed-loop management logic keeps the entire treatment process under dose monitoring and dynamic correction, significantly reducing treatment risks caused by dose deviations and improving the safety and reliability of particle radiotherapy.

[0069] In one possible implementation, the construction of a fusion state representation that simultaneously reflects the patient's anatomical state and the treatment execution state includes:

[0070] The three-dimensional voxel grid of the planned CT images serves as a unified spatial coordinate system;

[0071] Based on the dynamic treatment log data, in the three-dimensional voxel grid, a deviation feature vector is calculated and assigned to each voxel. This vector encodes at least the statistical information of the positional and energy deviations of all beam spots flowing through the voxel.

[0072] In the three-dimensional voxel grid, each voxel is assigned an anatomical feature vector, which includes at least the voxel's CT Hounsfield Unit (HU) value (i.e., CT value), the organ category code, and the distance to the nearest target boundary.

[0073] The deviation feature vector corresponding to each voxel is concatenated with the anatomical feature vector to form the fusion feature vector of that voxel; the fusion feature vectors of all voxels together constitute the fusion state representation, which is in the form of a multi-channel three-dimensional feature tensor.

[0074] In the specific implementation, the three-dimensional voxel grid of the planned CT image is first selected as a unified spatial coordinate system. This coordinate system naturally carries the patient's anatomical spatial information and provides spatial anchors for the subsequent fusion of the two types of features.

[0075] For each executed beam spot, based on its actual position and energy, the three-dimensional influence area of ​​its energy deposition in vivo (such as the Bragg peak region along the range and its scattering range) can be approximately depicted. For each voxel, based on the treatment execution information in the dynamic treatment log data, all beam spots flowing through or significantly affecting the voxel are identified, and the deviation between the target and actual physical quantities of each scan beam spot is extracted. Through statistical calculation, the core execution deviation information such as the position deviation and energy deviation of all beam spots flowing through each voxel is encoded into a deviation feature vector and assigned to the corresponding voxel, realizing the mapping of execution deviation information to anatomical space.

[0076] On the same voxel grid, each voxel is directly assigned its inherent anatomical and structural features, forming an anatomical feature vector:

[0077] CT Henlein unit value: directly reflects the electron density of tissue and is the most critical basic physical quantity for calculating proton range (Bragg peak position).

[0078] Organ category coding: Indicates the voxel's target area, specific organ at risk, or normal tissue in the form of a classification label (such as unique heat coding), providing important semantic context for subsequent dose assessment.

[0079] Distance to the nearest target boundary: This quantifies the spatial relationship of the voxel relative to the core therapeutic target (target area), helping the model understand that dose changes closer to the target edge are more clinically important than those farther away.

[0080] For each voxel, its deviation feature vector (dynamic execution information) is concatenated with its anatomical feature vector (static structural information) to form a longer and more informative fused feature vector.

[0081] The fused feature vectors of all voxels are stacked together to form a multi-channel three-dimensional feature tensor with dimensions of [height, width, depth, number of feature channels]. This tensor is the final fused state representation.

[0082] The effects of the above technical solution are as follows:

[0083] By using a unified three-dimensional voxel grid coordinate system, the temporal beam spot execution bias information and spatial anatomical feature information are anchored and fused, effectively solving the problem of inaccurate correlation caused by the heterogeneity of the two types of information. This enables the model to directly learn the intrinsic correlation between execution bias and dose distribution changes at specific anatomical locations.

[0084] By constructing and fusing feature vectors for each voxel separately, the voxel-level anatomical details (such as tissue density differences) and local execution bias information (such as the concentration bias of beam spots in specific regions) are fully preserved, avoiding information loss caused by feature coarsening and providing a high-quality feature foundation for the accurate evaluation of subsequent dose distribution.

[0085] The final output multi-channel 3D feature tensor naturally matches the input format of deep learning models (especially 3D networks), eliminating the need for additional feature format conversion and spatial alignment processing by the model. This significantly reduces the preprocessing burden on the model and improves the overall efficiency of dose assessment.

[0086] The fusion of feature vectors encompasses both execution bias and anatomical attributes, enabling the model to accurately identify the impact of different scenarios, such as the same bias in different anatomical regions and different biases in the same anatomical region, on dose distribution, providing crucial support for the accurate output of subsequent quantitative evaluation values.

[0087] In one possible implementation, the dose optimization model employs a multimodal deep learning network architecture driven by physical constraints, including a two-stream coding module, a physical information fusion module, and a physical constraint decoding module; the two-stream coding module includes a parallel anatomical structure coding stream and a beam spot sequence coding stream.

[0088] The anatomical structure encoding stream takes the anatomical image data as input and extracts multi-scale anatomical spatial features through a three-dimensional depth feature extraction network.

[0089] The beamspot sequence encoding stream takes the deviation sequence formed by the target physical quantity and the actual executed physical quantity of each beamspot as input, models the temporal dependency of beamspot deviation through the temporal-spatial feature enhancement module, and enhances the feature expression of beamspot deviation events that have a preset correlation with the final dose distribution deviation; the preset correlation is determined based on the dose influence weight corresponding to the beamspot deviation event.

[0090] The dose influence weight is calculated using a correlation model between the beam spot deviation parameter and the dose distribution deviation; the correlation model is either a physics-driven analytical correlation model or a data-driven statistical correlation model.

[0091] If the correlation model is a physics-driven analytical correlation model, then the model is constructed based on the physical laws of proton energy deposition, including the mapping relationship between beam spot position deviation and water equivalent path length, the mapping relationship between beam spot energy deviation and Bragg peak parameters, and the linear superposition relationship between beam spot weight deviation and dose contribution.

[0092] If the association model is a data-driven statistical association model, the model is trained using clinical treatment data or Monte Carlo simulation data, with beam spot position deviation, energy deviation, and weight deviation as input features, and the dose deviation value of the target area or organs at risk as the output label.

[0093] The preset correlation degree is a preset contribution threshold; the temporal-spatial feature enhancement module identifies beam spot deviation events with dose influence weights greater than the threshold as key deviation events that significantly affect the final dose distribution, and enhances their feature representation.

[0094] The preset correlation degree is achieved through feature weight sorting; the temporal-spatial feature enhancement module identifies the top N beam deviation events in the dose influence weight sort as key deviation events that significantly affect the final dose distribution, and enhances their feature expression; N is a positive integer and can be adjusted according to the actual treatment scenario.

[0095] The three-dimensional depth feature extraction network is a 3D network containing residual connections, and the 3D network is selected from 3DDenseNet, 3DResNet or 3DCNN.

[0096] In a preferred implementation, the temporal-spatial feature enhancement module is composed of a cascaded temporal modeling unit and a spatial attention unit; the temporal modeling unit is a gated recurrent unit (GRU) or a long short-term memory network (LSTM), and the spatial attention unit is a self-attention mechanism.

[0097] The cascading order of the time-space feature enhancement module is as follows: first, the deviation sequence is modeled on a time-dependent basis using a time-series modeling unit, and then the feature representation of key deviation events is identified and enhanced from the feature sequence after time-series modeling using a self-attention mechanism.

[0098] The physical information fusion module receives features from the dual-stream coding module and modulates the features through a differentiable proton transport approximation unit; the proton transport approximation unit generates a physical rationality attention weight map based on the tissue density and proton energy information implied in the current features, which is used to weight the fused features;

[0099] The physical constraint decoding module receives the modulated fused features, restores the spatial resolution through a three-dimensional deconvolutional network, and integrates a physical law verification layer during the decoding process. The physical law verification layer uses a parameterized proton depth dose curve model and a transverse scattering model to verify and adjust the intermediate feature map, and finally outputs the three-dimensional dose deviation distribution map.

[0100] The differentiable proton transport approximation unit is configured to learn a mapping function from the proton's initial energy and path organization density to its Bragg peak depth and width parameters, and dynamically generate the physical plausibility attention weight map using the parameters output by this function.

[0101] The parameterized model upon which the physical law verification layer is based includes:

[0102] A Bragg peak analytical model based on water equivalent path length correction is used to constrain the dose deposition morphology along the beam spot range direction.

[0103] A transverse dose broadening model based on Fermi-Eitch scattering theory is used to constrain the dose distribution perpendicular to the beam spot direction.

[0104] The differentiable proton transport approximation unit (PTAU) maps the initially fused features to physically guided attention weights. Specific implementations include:

[0105] The approximate proton energy and path-relative blocking ability are decoded from the fusion features at spatial locations.

[0106] A nonlinear mapping between approximate proton energy, path-relative stopping power, and Bragg peak depth and width is established using parameterized functions.

[0107] Based on the calculated Bragg peak depth, width and voxel water equivalent depth, a physical rationality attention weight map is generated through Gaussian distribution function, so that voxels with water equivalent depth close to Bragg peak depth are given the highest weight.

[0108] The physical rationality attention weight map is multiplied element-wise with the input features to complete feature focusing based on physical priors, specifically as follows:

[0109] Physical quantity decoding: Two key physical quantities are decoded spatially from the fused features output by the dual-stream coding module using two independent 1×1×1 convolutional layers, realizing the mapping from features to physical quantities.

[0110] Decoding the approximate proton energy E and path-relative stopping power S

[0111] Convolutional layer 1: Outputs the approximate proton energy E at each spatial location (initially a feature map value, without actual physical units);

[0112] Convolutional layer 2: Outputs the path-relative blocking ability S at each spatial location (initially a feature map value, without actual physical units).

[0113] The decoded E and S are standardized to ensure they conform to the range of clinical physical parameters: the approximate proton energy range is normalized to [0,1] using the Sigmoid function and then linearly mapped to the typical therapeutic energy range (e.g., 70-250 MeV); the path-relative blocking ability is normalized to the typical range of biological tissues (e.g., 0.9-1.1) using the Sigmoid function.

[0114] A nonlinear mapping between the proton energy E-path stopping power S and the key parameters of the Bragg peak is established using a learnable parameterized function, as shown in the following formula:

[0115]

[0116]

[0117] Where E is the initial proton energy (a physical quantity decoded and mapped from the features), and S is the path-relative stopping power (the energy loss efficiency of the tissue along the path of the proton beam). For the predicted Prague depth, p and q are trainable exponential parameters used to fit the complex nonlinear relationship between energy, stopping power and Bragg peak parameters, ranging from 1.5 to 2.5 (optimized through training with clinical data to fit the physical characteristics of proton energy deposition). , , These are trainable coefficient parameters used to adjust the fitting accuracy of the Bragg peak depth and width calculation functions, which are learned through training to match the Bragg peak characteristics of real proton beams; the network learns these parameters through training, thereby fitting the complex relationship from energy, stopping power to peak parameters;

[0118] To ensure that the exponential parameters p and q can be updated through backpropagation, the gradient is calculated using logarithmic differentiation, as shown in the following formula:

[0119]

[0120]

[0121] By taking the logarithmic derivative, the gradient of the exponential parameter is transformed into a computable energy × logarithmic energy form, ensuring that the entire PTAU module supports end-to-end training.

[0122] Based on calculations (Bragg Peak depth) and (Bragg peak width), combined with the equivalent water depth calculated by integration along the ray direction, generates a physically plausible attention weight map. The attention weight map uses a high-speed distribution function to make the Bragg peak region (d≈ The voxels of the highest weight are given to enhance the features of physically plausible regions:

[0123]

[0124] voxels Physical rationality attention weight (characterizing the probability that the voxel will physically undergo significant dose deposition); water equivalent depth to the virtual beam inlet; It is a voxel The water equivalent depth to the virtual beam inlet; i, j, k are the spatial coordinates of the voxels, corresponding to the x, y, and z axis coordinates of the three-dimensional voxel grid of the planned CT image, used to locate the specific position of each voxel in the anatomical space; The predicted Bragg peak depth corresponding to this spatial location; This represents the predicted Bragg peak width corresponding to this spatial location.

[0125] The closer the water equivalent depth of a voxel is to the predicted Bragg peak depth, the greater its weight, thus enabling the focus on the characteristics of the Bragg peak region.

[0126] The generated W phy Element-wise multiplication with the PTAU input features enables feature focusing based on physical priors. The gradient of the entire PTAU module can be backpropagated using the chain rule, achieving end-to-end training.

[0127] In one possible implementation, the physical law verification layer extracts longitudinal feature profiles along the beam depth direction and transverse feature profiles perpendicular to the beam direction from the currently generated intermediate feature map, according to the physical projection direction of the proton beam.

[0128] Subsequently, the verification layer inputs the extracted feature profile and the target dose profile calculated by the aforementioned physical model into a lightweight learnable multilayer perceptron. Based on the difference between the two profiles, the multilayer perceptron automatically learns and calculates a feature correction value, which indicates how to adjust the shape of the current feature profile to a target profile that is more in line with physical laws.

[0129] Finally, the verification layer uses a differentiable spatial mapping operation to precisely superimpose the calculated feature correction amount back into the corresponding spatial position in the intermediate feature map, thereby completing the physical verification and adjustment of the feature map and obtaining a new feature map that has been physically normalized.

[0130] The specific implementation is as follows:

[0131] From the intermediate feature map output by the 3D deconvolution, two types of core profiles (key morphologies reflecting dose distribution) are extracted along the physical projection direction of the proton beam:

[0132] Longitudinal feature profile: 1D feature sequence extracted along the beam range direction (depth z direction), corresponding to the depth-dose relationship of proton energy deposition;

[0133] Lateral feature profile: The 2D feature distribution extracted at a specific depth z perpendicular to the beam direction (radial profile with the beam center as the origin), corresponding to the lateral scattering pattern of protons.

[0134] Based on a predefined parametric physical model, a longitudinal Bragg peak curve and a transverse scattering profile are generated as target templates for comparing the physical plausibility of intermediate feature profiles.

[0135] The Bragg peak analytical model based on water equivalent path length correction uses a Gaussian function superimposed with a plateau correction term to construct the longitudinal target dose curve. The longitudinal target curve (Bragg peak depth - dose distribution) simulates the front plateau-peak-sharp drop shape of a real proton beam, as shown in the following formula:

[0136]

[0137] The target dose curve is the longitudinal direction (ideal dose value at depth z, unit: Gy), where z is the depth coordinate along the beam path. The depth of the standard Bragg peak is based on physical principles (reference template, not model prediction). ϵ is the standard Bragg peak width based on physical principles (reference template, not model prediction); ϵ is a small parameter characterizing the plateau region, a plateau region correction parameter (simulating the small amount of energy deposition of a real proton beam before reaching the Bragg peak, avoiding zero dose in the pre-plateau region), typically set to 10. -3 Magnitude;

[0138] The lateral target profile (scattering distribution at a specific depth z) is simulated using a Gaussian distribution based on Fermi-Edge scattering theory, as shown in the following formula:

[0139]

[0140] This is the lateral target dose profile (ideal dose value at depth z and lateral radial direction r); z is the depth coordinate along the beam range direction; r is the lateral radial coordinate (with the beam center as the origin, perpendicular to the beam direction). denoted as the transverse scattering width at depth z (the degree of transverse diffusion of the proton beam at depth z). It is the beam spot intensity normalization coefficient, used to ensure that the integral of the transverse dose distribution is equal to the total beam spot intensity, so that the dose scale of the target curve meets clinical requirements.

[0141] The verification layer inputs the extracted feature profile and the target dose profile calculated by the aforementioned physical model into a lightweight learnable multilayer perceptron to obtain a correction value. This correction value is then mapped back to the corresponding position in the feature map in a differentiable manner (e.g., spatial transformation or addition) to obtain the corrected feature map. The mapping from feature profile to physical bias to correction value is achieved through a lightweight learnable multilayer perceptron (MLP). The specific steps include:

[0142] The extracted longitudinal / lateral feature profiles are concatenated with the corresponding longitudinal target curves / lateral target profiles to form the input vector of the MLP (dimension is profile length × 2).

[0143] Based on the deviation between the input features and the target, the MLP automatically learns and outputs a feature correction value ΔP (consistent with the dimension of the input profile). The physical meaning of ΔP is to adjust the current feature profile to the correction value required to match the target profile.

[0144] By using bilinear interpolation (ensuring gradient propagation), the correction amount ΔP is mapped back to the corresponding spatial location in the intermediate feature map. Element-wise addition is then performed with the original feature map to obtain a physically normalized new feature map.

[0145] All operations of the physical law verification layer (profile extraction, MLP calculation, bilinear interpolation) support gradient backpropagation and can be trained in conjunction with the PTAU module and the dual-stream coding module to ensure that the entire dose optimization model is always constrained by physical laws throughout the feature extraction, fusion and decoding process, and to avoid outputting physically unreasonable dose deviation distributions.

[0146] The working principle of the above technical solution is as follows:

[0147] To achieve rapid and accurate mapping from beam speckle execution deviation and anatomical structural features to dose distribution changes, this embodiment proposes a physical law-constrained multimodal deep learning network. The core of this network is that it does not rely on time-consuming high-precision Monte Carlo simulations to directly generate training labels. Instead, it transforms the underlying physical principles into differentiable and embeddable constraint modules and loss functions, guiding the network to meet clinical real-time requirements while ensuring that its prediction results strictly follow the basic physical laws of proton transport and energy deposition.

[0148] The network adopts a cascaded architecture of dual-stream feature extraction, physical information fusion, and constraint decoding to ensure that the information stream contains both data features and physical priors.

[0149] Anatomical image data (3D images of patients and organ delineations extracted from log files) are input into a 3D DenseNet network with residual connections. This design is optimized for processing high-dimensional medical images with complex spatial structures. DenseNet promotes feature reuse through dense connections and alleviates the training difficulties of deep networks through residual connections. Together, they efficiently extract multi-scale anatomical context features from low-level (edges, textures) to high-level (organ morphology, spatial relationships), and focus on modeling tissue density gradient features closely related to proton range and energy deposition, outputting a compact, semantically rich 3D feature map.

[0150] A beamspot sequence encoding stream is used to process temporal beamspot execution bias sequences. First, a gated recurrent unit (GRU) processes the sequence according to the beamspot execution order, modeling the temporal dependencies of the biases (e.g., a positional bias may affect the energy deposition position of subsequent beamspots). Then, a self-attention mechanism performs a global analysis of the temporal features output by the GRU, dynamically calculating the importance weight of each beamspot bias feature for the final dose prediction. This module not only models long-range dependencies but also identifies critical bias events that, although small, occur at key anatomical locations or interact with specific anatomical structures and have an amplified effect, enhancing their feature representation and suppressing interference from minor or redundant biases. The output of this stream is a beamspot bias feature representation that incorporates temporal context and importance weights.

[0151] The physical information fusion module performs preliminary fusion of features from the dual-stream coding module. Then, a differentiable proton transport approximation unit (PTAU) is introduced; the PTAU is a learnable physics engine simulator that decodes the implicit approximate proton energy and path organization density information corresponding to each spatial location (or feature channel) from the preliminary fused features.

[0152] By learning the core determinants in Monte Carlo calculations, such as the functional relationship between the initial proton energy and the Bragg peak depth and width, and the perturbation model of range caused by tissue inhomogeneity, a differentiable mapping function (typically a small neural network) within PTAU is used to calculate the key physical parameters of the proton: the Bragg peak depth and width. These two parameters directly determine the longitudinal deposition characteristics of the dose in the tissue.

[0153] Specifically, the differentiable computation process of this PTAU unit executes the following steps sequentially: First, using two independent micro-convolutional layers, two key physical quantities corresponding to each three-dimensional spatial location point are parsed and estimated from the input features: one is the approximate energy value that the proton beam may have at that point, and the other is the approximate relative stopping power value of the path the proton beam takes to reach that point. Next, the unit uses a set of internal parameters that can be automatically adjusted during training to convert the decoded approximate energy value and relative stopping power value into two key physical parameters that determine the dose deposition morphology: one is the predicted Bragg peak depth, and the other is the predicted Bragg peak distribution width. In particular, the exponent of the energy value itself is designed as a trainable parameter during the calculation process, so that the entire mapping relationship can learn from actual data and self-optimize, rather than relying solely on fixed theoretical formulas. Then, based on the calculated peak depth and distribution width parameters, and considering the equivalent water thickness accumulated from the human body surface to the current voxel point along the proton beam incident direction, the unit dynamically generates a three-dimensional physical plausibility attention weight map. The weight map is calculated using a Gaussian distribution, with its value peaking at the center of the predicted Bragg peak deposition region and smoothly decreasing with increasing distance from this region. Finally, the unit multiplies the generated physical plausibility weight map point-by-point with the initial input feature map. This operation is equivalent to weighting feature information according to physical laws, highlighting features from areas physically more likely to have significant dose deposition, while weakening features from inappropriate regions. Since all the above steps support gradient calculation, the error gradient of the entire PTAU module can be backpropagated from the output to the input using the chain rule, ensuring that this module can be jointly trained end-to-end with the entire neural network.

[0154] Subsequently, PTAU dynamically generates a physically plausible attention weight map using the calculated depth and width parameters. This map highlights regions where significant dose deposition or changes are more likely to occur physically (e.g., near the expected landing point of the Bragg peak, at high-density gradient interfaces) and assigns higher weights to features in these regions. This process modulates or focuses features based on physical principles.

[0155] The features modulated by PTAU are fed into the physical constraint decoding module, and then gradually upsampled through a 3D deconvolution network to restore the spatial resolution and generate the final three-dimensional dose map.

[0156] At the critical layer of decoding, a Physical Law Verification Layer (PCRL) is integrated. This layer does not directly perform time-consuming Monte Carlo simulations; instead, it utilizes a set of lightweight, parameterized physical equations (such as a modified Bragg peak analytical model and an energy-range relationship based on mass thickness correction) to perform real-time verification and fine-tuning of the dose distribution in the initial decoding, ensuring the output is physically reasonable. The Physical Law Verification Layer includes:

[0157] Longitudinal constraints: A Bragg peak analytical model based on water equivalent path length correction is employed. This model can quickly generate standard Bragg peak curves for a given energy and tissue density. The physical law verification layer extracts a profile along the beam spot incident direction from the intermediate feature map of the decoder and guides its shape to approach the reasonable shape predicted by the analytical model, ensuring that the dose distribution in the depth direction conforms to physical common sense.

[0158] Lateral Constraints: A lateral dose broadening model based on Fermi-Edge scattering theory is employed. This model describes the dose diffusion perpendicular to the beam direction. A physical law verification layer is used to extract a lateral profile at a specific depth, constraining its broadening to match model predictions and ensuring a reasonable lateral dose distribution.

[0159] The role of the validation layer is to fine-tune or normalize the features generated by the network so that they do not deviate from the basic physical laws during the evolution towards the final dose map, thereby ensuring that the output results have basic physical rationality even in extreme cases where the training data is not fully covered.

[0160] The effects of the above technical solution are as follows:

[0161] Through dual-stream encoding, the model can fully extract information from anatomical and temporal biases and their interactions. In particular, the self-attention mechanism enables it to intelligently focus on key biases with significant clinical impact, improving its ability to resolve complex scenarios. Compared to networks that do not distinguish modalities or simply merge inputs, it achieves higher prediction accuracy.

[0162] By introducing a differentiable proton transport approximation unit (PTAU) and a physics verification layer, the core physics knowledge of proton therapy (range-energy-density relationship, Bragg peak morphology, and scattering behavior) is explicitly and structurally embedded into the network's forward propagation process. This transforms the model from a mere data fitter into an intelligent agent that understands physics. The network's internal features and attention maps can be associated with physical concepts, improving the transparency of the model's decision-making process and the reliability of its results.

[0163] Purely data-driven models may produce physically impossible dose predictions (such as Bragg peaks appearing outside the tissue) when encountering extreme anatomical variations or rare combinations of biases outside the training data distribution. This architecture, through the hard physical constraints of PTAU and PCRL, forces the model's output to conform to basic physical laws in macroscopic morphology, greatly enhancing the model's generalization ability and safety in unknown scenarios and reducing the risk of producing absurd results.

[0164] The entire architecture aims to approximate the accuracy of full physical dose calculation (such as Monte Carlo) using the efficient forward inference of deep learning. Both PTAU and PCRL use parameterized fast approximation models with minimal computational overhead, but provide highly effective physical guidance. This allows the model to meet the real-time requirements of online adaptation while providing evaluation accuracy far exceeding that of traditional fast algorithms, achieving a breakthrough in clinical usability.

[0165] Ultimately, the three-dimensional dose deviation distribution map output by the model, due to its combination of data-driven detail restoration capabilities and overall rationality guaranteed by physical laws, becomes a high-quality and highly reliable source of information for triggering subsequent adaptive dose adjustment commands (such as beam position and weight correction). This ensures that the entire adaptive closed-loop system is built on a solid and reliable assessment foundation, thereby enabling safer, more effective, and machine-executable treatment decisions.

[0166] In one possible implementation, the training method for the dose optimization model includes:

[0167] Multiple fusion state representations from several historical treatment cases are obtained as input samples for the model;

[0168] For each input sample, a composite supervision target is generated, which integrates clinical dose difference signals, physical consistency constraint signals based on the physical laws of proton transport, and cyclic consistency constraint signals to ensure the closed loop of the evaluation logic.

[0169] The dose optimization model is iteratively trained with the training direction of minimizing the difference between the model output and the composite supervision objective.

[0170] During iterative training, the relative weights of the physical consistency constraint signal and the cyclic consistency constraint signal in the optimization objective are dynamically adjusted.

[0171] The composite supervision objective is achieved by minimizing a multi-objective loss function, which is:

[0172]

[0173] Wherein, L_data is the data fitting loss term, which uses the dose difference distribution calculated by the fast pencil beam algorithm as the supervision label to calculate the difference between the predicted value and the label; L_physics is the physical consistency loss term, which is used to constrain the dose distribution predicted by the network to conform to the physical laws of proton transport on a macroscopic scale; L_cycle is the cycle consistency loss term; λ_phy is the weight of the physical consistency loss term; and λ_cyc is the weight of the cycle consistency loss term.

[0174] The physical law consistency loss term includes a regional energy conservation sub-term, a depth dose curve morphology sub-term, and / or a lateral scattering constraint sub-term:

[0175] The regional energy conservation sub-term is achieved by minimizing the total integral of the predicted dose deviation within the planned target area, and the first difference measure between the weighted deviation statistics of all beam spots;

[0176] The depth dose profile morphology sub-item is achieved by minimizing a second difference measure between the predicted depth dose profile extracted along the beam spot path and the standard Bragg peak profile generated based on the beam spot energy parameterization;

[0177] The lateral scattering constraint is achieved by minimizing a third measure of difference between the predicted dose distribution at a specific depth lateral profile and a reference profile predicted based on the scattering model.

[0178] The first difference measure is mean square error or absolute error; the second difference measure is the negative value of the correlation coefficient between curves or a shape similarity measure; and the third difference measure is the difference between the KL divergence or profile width parameters between dose distributions.

[0179] The cycle consistency loss term is calculated in the following way:

[0180] A reverse reasoning network is constructed, which is trained to infer beam spot execution bias characteristics by taking the sum of the original planned dose distribution and the dose deviation distribution predicted by the dose optimization model as input.

[0181] The cycle consistency loss term is defined as the difference between the predicted execution bias feature output by the inverse inference network and the speckle execution bias feature extracted from the real log data;

[0182] The inverse reasoning network and the dose optimization model are jointly optimized during training, and the entire system forms a self-consistent closed loop by minimizing the cycle consistency loss term.

[0183] Among them, the Fast Pencil Beam Algorithm (FPBA) is the core engineering algorithm in the field of proton radiotherapy dose calculation. Its core logic is to discretize the continuous proton beam into countless pencil beam elements (micro-beam spots), pre-calculate the dose distribution kernel of the pencil beam in a standard medium (such as a water phantom), and then combine the tissue density information provided by the patient's CT images to perform medium correction and superposition on the dose distribution of each pencil beam element, and finally quickly obtain the three-dimensional dose distribution of the entire target area.

[0184] By constructing a multi-task, multi-constraint collaborative learning framework, the model is guided to deeply internalize the laws of proton physics while fitting high-precision clinical data, ensuring the self-consistency of its internal logical reasoning. It abandons the single supervised learning paradigm and instead shapes a precise and reliable dose assessment agent through carefully designed composite supervision signals and dynamic optimization strategies. The training method for the dose optimization model specifically includes the following steps:

[0185] Constructing a training dataset: The training dataset contains historical data of multiple proton radiotherapy patients. Each patient's data includes at least three-dimensional images of the patient, organ delineation information, beam spot target and actual physical quantity sequence, and planned dose distribution data.

[0186] Data preprocessing: Voxel normalization and region of interest cropping were performed on the patient's 3D images; temporal alignment and standardization were performed on the beam spot deviation sequence; and the clinical dose difference monitoring signal was calculated using the fast pencil beam algorithm.

[0187] Sample construction: The fusion state representation formed by fusing patient anatomical features and beam spot execution deviation features is used as the input sample, and the combination of clinical dose difference, physical law constraint target, and cycle consistency constraint target is used as the label sample;

[0188] Model initialization: Initialize the parameters of the dual-stream coding network, the physical constraint decoder, and the built-in differentiable proton transport approximation module, and set the initial weight coefficients λ_physics and λ_cycle of the multi-objective physical consistency constraint loss function;

[0189] Iterative training: The model is trained iteratively using mini-batch gradient descent. In each iteration, the total loss L_total is calculated, and all trainable parameters of the model are updated through backpropagation. During training, the values ​​of λ_physics and λ_cycle are dynamically adjusted to achieve a balance between clinical accuracy and physical consistency.

[0190] Model validation and selection: The performance of the models during the training process was evaluated using an independent validation dataset. The evaluation metrics included clinical dose prediction error, physical law compliance, and prediction speed. The model with the best overall performance on the validation set and that meets the clinical real-time requirements was selected as the final pre-trained dose optimization model.

[0191] In each training iteration, the model outputs the predicted dose bias distribution during forward propagation. At this point, the loss function comprehensively evaluates the prediction results from three dimensions:

[0192] Data fitting loss (L_data): This directly calculates the error (such as L1 or L2 loss) between the predicted dose and the gold standard dose calculated by the Fast Pencil Beam Analysis (FPBA), ensuring the clinical accuracy of the model output. FPBA, as a high-efficiency physics engine, provides supervisory labels that balance accuracy and speed.

[0193] Physical consistency loss (L_physics): This loss term constrains the prediction results from the perspective of physical principles and includes several sub-terms:

[0194] Regional energy conservation sub-term: Ensures that the total energy change of the predicted dose within the target area matches the total energy change calculated from the actual weighted deviations of all beam spots, thus guaranteeing the physical correctness of the prediction from a macroscopic perspective.

[0195] Depth dose curve morphology sub-item: It is required that the depth dose curve predicted by the model along any beam spot range conforms to the parameterized Bragg peak theoretical curve in morphology, so as to ensure that the longitudinal distribution of dose deposition is reasonable.

[0196] Lateral scattering constraint sub-item: Constrains the broadening of the predicted dose in the lateral profile, consistent with the prediction based on the scattering theory model, and ensures that the lateral shape of the dose distribution is reasonable.

[0197] Cyclic Consistency Loss (L_cycle): By introducing a backward inference network, a closed loop of "dose prediction → bias back-calculation" is constructed. This network attempts to deduce the beam pattern execution bias characteristics that led to the final predicted dose distribution. Cyclic consistency loss measures the difference between this back-calculated bias characteristics and the bias characteristics in the actual log. It forces the forward prediction model to produce logically reversible and self-consistent prediction results, thereby improving the model's internal logical rigor and generalization ability.

[0198] The weight coefficients λ_phy and λ_cyc in the loss function are dynamically adjusted using a course learning strategy: in the early stage of training, a larger λ_phy is set to guide the network to learn basic laws with strong physical constraints; in the middle and later stages of training, λ_phy is gradually reduced and λ_cyc is increased to optimize the fitting accuracy of clinical data and the self-consistency of the model's internal logic.

[0199] The effects of the above technical solution are as follows:

[0200] Compared to traditional deep learning training that only uses data fitting loss, this embodiment introduces a physical consistency loss to directly encode the core physical laws of proton transport (the longitudinal and lateral laws of energy deposition) as the model's optimization objective. This allows the model to not only memorize data patterns but also understand and obey physical laws. Therefore, when faced with novel anatomical structures or rare combinations of biases not found in the training data, the model can still produce physically reasonable and credible predictions based on its internalized physical knowledge, significantly reducing the risk of generating absurd or dangerous predictions. During the training phase, the model uses a physical consistency loss term for global parameter constraints; while during the inference phase, forward feature modulation and output verification are performed through differentiable proton transport approximation units integrated within the network and a physical law verification layer.

[0201] The introduction of cycle consistency loss requires that the model's forward predictions (from bias to dosage) and backward reasoning (from dosage to bias) corroborate each other, forming a logical closed loop. This training forces the model to learn a more fundamental and deterministic causal relationship between dosage and bias, rather than a superficial statistical correlation. The model trained in this way has a more logically coherent prediction process and more interpretable prediction results.

[0202] By adopting a strategy of dynamically adjusting loss weights, the conflict in multi-objective optimization is effectively alleviated, and the model is prevented from getting stuck in local optima or oscillating between different loss terms. This makes the training process more stable, converges faster, and ultimately results in a more balanced and superior model performance.

[0203] In specific applications, the quantitative evaluation values ​​include:

[0204] The difference between the actual delivered dose distribution and the planned dose per voxel; and / or,

[0205] Dosage statistics for each clinically relevant region, including target coverage indicators and dose limitation indicators for organs at risk; these dose statistics are calculated based on the difference between the actual delivered dose distribution and the planned dose per voxel.

[0206] The preset decision rules and corresponding adaptive dose adjustment instructions include at least one of the following:

[0207] If the quantitative evaluation value shows that the target area dose coverage is lower than the preset minimum coverage requirement (e.g., D95 < 40 Gy, meaning that the proportion of the irradiated target area volume with a dose greater than 40 Gy to the total target area volume is less than 95%), or the dose to any organ at risk exceeds the preset safety limit, then the online adaptive adjustment process is triggered, and a beam correction instruction is generated for beam position, weight, etc.; the beam correction instruction is used to adjust the irradiation of the beam in subsequent scans in real time within the current treatment fraction;

[0208] If, based on the cumulative and trend analysis results of the quantitative assessment values ​​of multiple treatment fractions already executed, it is predicted that the cumulative total dose distribution of the patient's entire treatment course will not meet the preset clinical overall goal, then an offline adaptive re-optimization process is triggered, and a treatment plan re-optimization instruction is generated; the treatment plan re-optimization instruction triggers the re-formulation of a complete treatment plan for all subsequent fractions to be executed based on the latest patient anatomical images.

[0209] The beam spot correction command is generated by solving a constrained optimization problem. The constrained optimization problem aims to minimize the remaining beam spot weight adjustment while ensuring that the adjusted predicted dose meets clinical constraints. That is, under the premise of minimizing the change in the remaining beam spot plan weight, the predicted final fractional dose distribution can meet the preset target coverage and organ at risk dose limits when all beam spots are executed.

[0210] The working principle of the above technical solution is as follows:

[0211] The system calculates two levels of assessment values ​​in real time. The first level is the microscopic voxel-level dose difference, providing the most refined error distribution map. The second level aggregates macroscopic clinical dose statistics from the voxel differences. These statistics directly correspond to the goals and limitations of the clinical plan and are the gold standard for judging whether the treatment is safe and effective.

[0212] The calculated clinical statistics are automatically compared with preset safety thresholds derived from clinical protocols (e.g., the proportion of the irradiated target volume to the total target volume at doses greater than 40 Gy is not less than 95%). If a failure to meet the preset safety thresholds is detected, a risk status is identified, and automatic triage is performed based on the timing of the risk occurrence.

[0213] If a risk occurs within a single treatment session, it is immediately determined that the plan for that session has deviated from its intended course, triggering an online adaptive adjustment process. This process aims to correct deviations on-site.

[0214] If, based on the summation and trend extrapolation of dosage results from multiple executed fractions, the predicted total dosage for the entire treatment course will fail to meet the target, then the overall treatment plan needs to be revised, triggering an offline adaptive re-optimization process. This process aims to re-plan the treatment.

[0215] For online adjustment (bead weight correction): the goal is to quickly correct biases with minimal disturbance, modeled as a constrained optimization problem: minimizing the adjustment of the remaining bead weights (optimization objective) while adhering to clinical dose constraints (hard conditions). Here, "minimal" typically means minimizing the sum of squares (L2 norm) or the sum of absolute values ​​(L1 norm) of the adjustment magnitudes, aiming to maintain treatment stability and execution efficiency. After solving this problem, the optimal weight adjustment scheme is automatically distributed to the treatment device for execution, which can be solved using gradient descent.

[0216] For offline re-optimization (plan redo): the goal is to develop the optimal new treatment plan based on the latest information. When this process is triggered, the system retrieves the latest patient images, and the treatment planning system performs a complete and entirely new plan optimization, generating new treatment plans for all subsequent sessions.

[0217] The effects of the above technical solution are as follows:

[0218] By differentiating between online real-time fine-tuning and offline full rework response modes, the system achieves tiered and precise handling of treatment deviations. Minor or localized deviations that can be compensated for within the current session are quickly resolved through efficient weighting adjustments, maximizing the use of treatment opportunities within the current session and avoiding excessive interruptions to the treatment process. Only when medium- to long-term, global risks are predicted is a full rework, which consumes more resources, initiated. This tiered mechanism optimizes the efficiency of clinical resource utilization while ensuring efficacy and safety.

[0219] Online adjustments can be formulated as a constrained optimization problem with the objective of minimizing the adjustment amount. This is not only a mathematically optimal solution, but also directly addresses a key clinical need: to achieve the dosage target while maintaining maximum consistency with the original plan. This ensures that the adjustment plan will not introduce new, unpredictable risks due to over-adjustment, while minimizing the adjustment range also increases the likelihood that the plan will be executed safely and accurately by the treatment device, greatly enhancing clinicians' trust and acceptance of the automated adjustment results.

[0220] By using cumulative and trend analysis based on multi-fraction data to predict total dose, adaptive therapy is upgraded from dealing with existing problems to preventing future problems. This gives treatment management a forward-looking and individualized character, enabling dynamic adjustment of long-term strategies based on the actual response of specific patients to treatment. It provides a powerful closed-loop control tool for achieving truly personalized precision radiotherapy.

[0221] The evaluation timing for the quantitative evaluation value output by the dose optimization model is at least one of the following:

[0222] During the current treatment fractionation of beam spot irradiation, according to the preset beam spot irradiation angle, the dynamic treatment log data corresponding to the beam spot that has been irradiated at the preset angle is analyzed in real time to construct a fusion status characterization and complete quantitative evaluation.

[0223] After all beam spot irradiation is completed in the previous treatment fraction and before the start of the next treatment fraction, a fusion status characterization is constructed and a quantitative assessment is completed based on the complete dynamic treatment log data of the previous treatment fraction.

[0224] By supporting real-time assessment at preset irradiation angles during the current treatment fraction, this embodiment provides proton pencil beam scanning therapy with the ability to dynamically monitor and adjust in real time during a single treatment session. Once the system completes irradiation at a certain angle and real-time assessment detects that the dose deviation has reached a clinical risk threshold (such as insufficient target coverage or excessive dose to endanger organs), it can immediately generate and execute weight adjustment instructions for subsequent beam spots that have not yet been irradiated in the current fraction. This achieves a leap from traditional offline, posterior adaptive to online, real-time adaptive, effectively intercepting and correcting ongoing dose deviations, eliminating errors in their infancy, and greatly improving the accuracy of a single treatment and patient safety.

[0225] By comprehensively evaluating complete fraction data between fractions (i.e., after the end of one fraction and before the beginning of the next), the trend of cumulative dose across multiple fractions can be predicted more accurately. If the prediction indicates that the total dose will deviate from the clinical target, the treatment plan can be re-optimized in a timely manner, allowing for the development of new irradiation protocols for subsequent fractions that are more closely aligned with the current anatomical condition. This upgrades adaptive therapy from a passive response to an active, proactive, and dynamic management throughout the entire treatment course, ensuring that the entire treatment course remains focused on the optimal target.

[0226] Example 2: This example provides a particle radiotherapy adaptive planning optimization device, comprising:

[0227] The feature fusion module is used to construct a fusion state representation that simultaneously reflects the patient's anatomical state and treatment execution state based on the dynamic treatment log data generated by the patient during proton radiotherapy.

[0228] A quantitative evaluation module is used to input the fusion state representation into a pre-trained dose optimization model; the dose optimization model is configured to directly output a quantitative evaluation value of the dose distribution generated by the portion of the current treatment fraction that has been performed, based on the fusion state representation.

[0229] The optimization and adjustment module is used to automatically generate and execute adaptive dose adjustment instructions to guide the unexecuted portion of the current treatment or subsequent treatments, based on the quantitative evaluation value and by applying preset decision rules.

[0230] The dynamic treatment log data includes at least the patient's anatomical imaging data, a time-corresponding sequence of treatment plan instructions, and a sequence of treatment device execution feedback. The anatomical imaging data includes planned CT images and target area and organ-at-risk structural information delineated based on these images. The treatment plan instruction sequence includes the target location, target energy, and target weight of multiple scanning beams. The treatment device execution feedback sequence includes the actual location, actual energy, and actual weight corresponding to the scanning beams.

[0231] In one possible implementation, the feature fusion module is configured to perform the following steps:

[0232] The three-dimensional voxel grid of the planned CT images serves as a unified spatial coordinate system;

[0233] Based on the dynamic treatment log data, in the three-dimensional voxel grid, a deviation feature vector is calculated and assigned to each voxel. This vector encodes at least the statistical information of the positional and energy deviations of all beam spots flowing through the voxel.

[0234] In the three-dimensional voxel grid, each voxel is assigned an anatomical feature vector, which includes at least the voxel's CT Henlein unit value, the organ class code, and the distance to the nearest target boundary.

[0235] The deviation feature vector corresponding to each voxel is concatenated with the anatomical feature vector to form the fusion feature vector of that voxel; the fusion feature vectors of all voxels together constitute the fusion state representation, which is in the form of a multi-channel three-dimensional feature tensor.

[0236] The dose optimization model adopts a multimodal deep learning network architecture driven by physical constraints, including a dual-stream coding module, a physical information fusion module, and a physical constraint decoding module; the dual-stream coding module includes a parallel anatomical structure coding stream and a beam spot sequence coding stream.

[0237] The anatomical structure encoding stream takes the anatomical image data as input and extracts multi-scale anatomical spatial features through a three-dimensional depth feature extraction network.

[0238] The beamspot sequence encoding stream takes the deviation sequence formed by the target physical quantity and the actual executed physical quantity of each beamspot as input, models the temporal dependency of beamspot deviation through the temporal-spatial feature enhancement module, and enhances the feature expression of beamspot deviation events that have a preset correlation with the final dose distribution deviation; the preset correlation is determined based on the dose influence weight corresponding to the beamspot deviation event.

[0239] The dose influence weight is calculated using a correlation model between the beam spot deviation parameter and the dose distribution deviation; the correlation model is either a physics-driven analytical correlation model or a data-driven statistical correlation model.

[0240] If the correlation model is a physics-driven analytical correlation model, then the model is constructed based on the physical laws of proton energy deposition, including the mapping relationship between beam spot position deviation and water equivalent path length, the mapping relationship between beam spot energy deviation and Bragg peak parameters, and the linear superposition relationship between beam spot weight deviation and dose contribution.

[0241] If the association model is a data-driven statistical association model, the model is trained using clinical treatment data or Monte Carlo simulation data, with beam spot position deviation, energy deviation, and weight deviation as input features, and the dose deviation value of the target area or organs at risk as the output label.

[0242] The preset correlation degree is a preset contribution threshold; the temporal-spatial feature enhancement module identifies beam spot deviation events with dose influence weights greater than the threshold as key deviation events that significantly affect the final dose distribution, and enhances their feature representation.

[0243] The preset correlation degree is achieved through feature weight sorting; the temporal-spatial feature enhancement module identifies the top N beam deviation events in the dose influence weight sort as key deviation events that significantly affect the final dose distribution, and enhances their feature expression; N is a positive integer and can be adjusted according to the actual treatment scenario.

[0244] The three-dimensional depth feature extraction network is a 3D network containing residual connections, and the 3D network is selected from 3DDenseNet, 3DResNet or 3DCNN.

[0245] In a preferred implementation, the temporal-spatial feature enhancement module is composed of a cascaded temporal modeling unit and a spatial attention unit; the temporal modeling unit is a gated recurrent unit (GRU) or a long short-term memory network (LSTM), and the spatial attention unit is a self-attention mechanism.

[0246] The cascading order of the time-space feature enhancement module is as follows: first, the deviation sequence is modeled on a time-dependent basis using a time-series modeling unit, and then the feature representation of key deviation events is identified and enhanced from the feature sequence after time-series modeling using a self-attention mechanism.

[0247] The physical information fusion module receives features from the dual-stream coding module and modulates the features through a differentiable proton transport approximation unit; the proton transport approximation unit generates a physical rationality attention weight map based on the tissue density and proton energy information implied in the current features, which is used to weight the fused features;

[0248] The physical constraint decoding module receives the modulated fused features, restores the spatial resolution through a three-dimensional deconvolutional network, and integrates a physical law verification layer during the decoding process. The physical law verification layer uses a parameterized proton depth dose curve model and a transverse scattering model to verify and adjust the intermediate feature map, and finally outputs the three-dimensional dose deviation distribution map.

[0249] The differentiable proton transport approximation unit is configured to learn a mapping function from the proton's initial energy and path organization density to its Bragg peak depth and width parameters, and dynamically generate the physical plausibility attention weight map using the parameters output by this function.

[0250] The parameterized model upon which the physical law verification layer is based includes:

[0251] A Bragg peak analytical model based on water equivalent path length correction is used to constrain the dose deposition morphology along the beam spot range direction.

[0252] A transverse dose broadening model based on Fermi-Eitch scattering theory is used to constrain the dose distribution perpendicular to the beam spot direction.

[0253] The training method for the dose optimization model includes:

[0254] Multiple fusion state representations from several historical treatment cases are obtained as input samples for the model;

[0255] For each input sample, a composite supervision target is generated, which integrates clinical dose difference signals, physical consistency constraint signals based on the physical laws of proton transport, and cyclic consistency constraint signals to ensure the closed loop of the evaluation logic.

[0256] The dose optimization model is iteratively trained with the training direction of minimizing the difference between the model output and the composite supervision objective.

[0257] During iterative training, the relative weights of the physical consistency constraint signal and the cyclic consistency constraint signal in the optimization objective are dynamically adjusted.

[0258] The composite supervision objective is achieved by minimizing a multi-objective loss function, which is:

[0259]

[0260] Wherein, L_data is the data fitting loss term, which uses the dose difference distribution calculated by the fast pencil beam algorithm as the supervision label to calculate the difference between the predicted value and the label; L_physics is the physical consistency loss term, which is used to constrain the dose distribution predicted by the network to conform to the physical laws of proton transport on a macroscopic scale; L_cycle is the cycle consistency loss term; λ_phy is the weight of the physical consistency loss term; and λ_cyc is the weight of the cycle consistency loss term.

[0261] The physical law consistency loss term includes a regional energy conservation sub-term, a depth dose curve morphology sub-term, and / or a lateral scattering constraint sub-term:

[0262] The regional energy conservation sub-term is achieved by minimizing the total integral of the predicted dose deviation within the planned target area, and the first difference measure between the weighted deviation statistics of all beam spots;

[0263] The depth dose profile morphology sub-item is achieved by minimizing a second difference measure between the predicted depth dose profile extracted along the beam spot path and the standard Bragg peak profile generated based on the beam spot energy parameterization;

[0264] The lateral scattering constraint is achieved by minimizing a third measure of difference between the predicted dose distribution at a specific depth lateral profile and a reference profile predicted based on the scattering model.

[0265] The first difference measure is mean square error or absolute error; the second difference measure is the negative value of the correlation coefficient between curves or a shape similarity measure; and the third difference measure is the difference between the KL divergence or profile width parameters between dose distributions.

[0266] The cycle consistency loss term is calculated in the following way:

[0267] A reverse reasoning network is constructed, which is trained to infer beam spot execution bias characteristics by taking the sum of the original planned dose distribution and the dose deviation distribution predicted by the dose optimization model as input.

[0268] The cycle consistency loss term is defined as the difference between the predicted execution bias feature output by the inverse inference network and the speckle execution bias feature extracted from the real log data;

[0269] The inverse reasoning network and the dose optimization model are jointly optimized during training, and the entire system forms a self-consistent closed loop by minimizing the cycle consistency loss term.

[0270] The quantitative evaluation values ​​include:

[0271] The difference between the actual delivered dose distribution and the planned dose per voxel; and / or,

[0272] Dosage statistics for each clinically relevant region, including target coverage indicators and dose limitation indicators for organs at risk; these dose statistics are calculated based on the difference between the actual delivered dose distribution and the planned dose per voxel.

[0273] The preset decision rules and corresponding adaptive dose adjustment instructions include at least one of the following:

[0274] If the target area dose coverage is lower than the preset minimum coverage requirement, or if the dose to any organ at risk exceeds the preset safety limit, an online adaptive adjustment process is triggered, and a speckle correction instruction is generated; the speckle correction instruction is used to correct the speckle position, weight, etc.

[0275] If, based on the cumulative and trend analysis results of the quantitative assessment values ​​of multiple treatment fractions already performed, it is predicted that the cumulative total dose distribution of the patient's entire treatment course cannot meet the preset clinical overall goal, then the offline adaptive re-optimization process is triggered, and a treatment plan re-optimization instruction is generated.

[0276] The beam spot correction command is generated by solving a constrained optimization problem. The constrained optimization problem aims to minimize the remaining beam spot weight adjustment while ensuring that the adjusted predicted dose meets clinical constraints.

[0277] The working principle and beneficial effects of the above technical solution are the same as those in Example 1, and will not be repeated here.

[0278] This invention also provides a radiotherapy system, comprising:

[0279] The particle radiotherapy adaptive planning evaluation and optimization device described in Example 2.

[0280] This invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the method described in Embodiment 1 of this invention or the function of the device described in Embodiment 2 of this invention.

[0281] This invention also provides a computer-readable storage medium for storing a computer program. When the computer program is executed, it implements the steps of the method in Embodiment 1 of this invention or the function of the device in Embodiment 2 of this invention. The specific implementation method is consistent with the implementation method and the technical effect achieved in the above method embodiments, and some contents will not be repeated.

[0282] In this invention, a readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The program product can take the form of any combination of one or more readable media. A readable medium can be a readable signal medium or a readable storage medium. A readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0283] Computer-readable storage media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium capable of sending, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, or any suitable combination thereof. Program code for performing operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on a user computing device, partially on an associated device, as a standalone software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing devices can be connected to user computing devices via any type of network, including local area networks (LANs) or wide area networks (WANs), or they can be connected to external computing devices (e.g., via the Internet using an Internet service provider).

[0284] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the invention without departing from the principles and spirit of the invention, and all such changes should fall within the protection scope of the claims of the present invention.

Claims

1. A method for adaptive planning optimization in particle radiotherapy, characterized in that, include: Based on the dynamic treatment log data generated by the patient during radiotherapy, a fusion state representation that simultaneously reflects the patient's anatomical state and treatment execution state is constructed. The fusion state representation is input into a pre-trained dose optimization model; the dose optimization model is configured to directly output a quantitative evaluation value of the dose distribution generated by the portion of the current treatment fraction that has been performed, based on the fusion state representation. Based on the quantitative evaluation value, a preset decision rule is applied to automatically generate and execute adaptive dose adjustment instructions to guide the unexecuted portion of the current fraction or subsequent fractions of treatment.

2. The adaptive planning optimization method for particle radiotherapy according to claim 1, characterized in that, The dynamic treatment log data includes at least the patient's anatomical imaging data, time-corresponding treatment plan instruction sequences, and treatment device execution feedback sequences; the anatomical imaging data includes planned CT images and target area and organ-at-risk structural information delineated based on the images; the treatment plan instruction sequences include the target positions, target energies, and target weights of multiple scanning beams; the treatment device execution feedback sequences include the actual positions, actual energies, and actual weights corresponding to the scanning beams.

3. The adaptive planning optimization method for particle radiotherapy according to claim 2, characterized in that, The dose optimization model adopts a multimodal deep learning network architecture driven by physical constraints, including a dual-stream coding module, a physical information fusion module, and a physical constraint decoding module; the dual-stream coding module includes a parallel anatomical structure coding stream and a beam spot sequence coding stream.

4. The adaptive planning optimization method for particle radiotherapy according to claim 3, characterized in that, The anatomical structure encoding stream takes the anatomical image data as input and extracts multi-scale anatomical spatial features through a three-dimensional depth feature extraction network. The beamspot sequence encoding stream takes the deviation sequence consisting of the target physical quantity and the actual executed physical quantity of each beamspot as input, models the temporal dependency of beamspot deviation through the temporal-spatial feature enhancement module, and enhances the feature expression of beamspot deviation events that have a preset correlation with the final dose distribution deviation.

5. The adaptive planning optimization method for particle radiotherapy according to claim 3, characterized in that, The physical information fusion module receives features from the dual-stream coding module and modulates the features through a differentiable proton transport approximation unit; the proton transport approximation unit generates a physical rationality attention weight map based on the tissue density and proton energy information implied in the current features, which is used to weight the fused features; The physical constraint decoding module receives the modulated fused features, restores the spatial resolution through a three-dimensional deconvolutional network, and integrates a physical law verification layer during the decoding process. The physical law verification layer uses a parameterized proton depth dose curve model and a transverse scattering model to verify and adjust the intermediate feature map, and finally outputs the three-dimensional dose deviation distribution.

6. The adaptive planning optimization method for particle radiotherapy according to claim 5, characterized in that, The differentiable proton transport approximation unit is configured to learn a mapping function from the proton's initial energy and path organization density to its Bragg peak depth and width parameters, and dynamically generate the physical plausibility attention weight map using the parameters output by this function.

7. The adaptive planning optimization method for particle radiotherapy according to claim 5, characterized in that, The parameterized model upon which the physical law verification layer is based includes: A Bragg peak analytical model based on water equivalent path length correction is used to constrain the dose deposition morphology along the beam spot range direction. A transverse dose broadening model based on Fermi-Eitch scattering theory is used to constrain the dose distribution perpendicular to the beam spot direction.

8. The adaptive planning optimization method for particle radiotherapy according to claim 1, characterized in that, The training method for the dose optimization model includes: Multiple fusion state representations from several historical treatment cases are obtained as input samples for the model; For each input sample, a composite supervision target is generated, which integrates clinical dose difference signals, physical consistency constraint signals based on the physical laws of proton transport, and cyclic consistency constraint signals to ensure the closed loop of the evaluation logic. The dose optimization model is iteratively trained with the training direction of minimizing the difference between the model output and the composite supervision objective. During iterative training, the relative weights of the physical consistency constraint signal and the cyclic consistency constraint signal in the optimization objective are dynamically adjusted.

9. The adaptive planning optimization method for particle radiotherapy according to claim 8, characterized in that, The composite supervision objective is achieved by minimizing a multi-objective loss function, which is: Wherein, L_data is the data fitting loss term, which uses the dose difference distribution calculated by the fast pencil beam algorithm as the supervision label to calculate the difference between the predicted value and the label; L_physics is the physical consistency loss term, which is used to constrain the dose distribution predicted by the network to conform to the physical laws of proton transport on a macroscopic scale; L_cycle is the cycle consistency loss term; λ_phy is the weight of the physical consistency loss term; and λ_cyc is the weight of the cycle consistency loss term.

10. The adaptive planning optimization method for particle radiotherapy according to claim 9, characterized in that, The physical law consistency loss term includes a regional energy conservation sub-term, a depth dose curve morphology sub-term, and / or a lateral scattering constraint sub-term: The regional energy conservation sub-term is achieved by minimizing the total integral of the predicted dose deviation within the planned target area, and the first difference measure between the weighted deviation statistics of all beam spots; The depth dose profile morphology sub-item is achieved by minimizing a second difference measure between the predicted depth dose profile extracted along the beam spot path and the standard Bragg peak profile generated based on the beam spot energy parameterization; The lateral scattering constraint is achieved by minimizing a third measure of difference between the predicted dose distribution at a specific depth lateral profile and a reference profile predicted based on the scattering model.

11. The adaptive planning optimization method for particle radiotherapy according to claim 9, characterized in that, The cycle consistency loss term is calculated in the following way: A reverse reasoning network is constructed, which is trained to infer beam spot execution bias characteristics by taking the sum of the original planned dose distribution and the dose deviation distribution predicted by the dose optimization model as input. The cycle consistency loss term is defined as the difference between the predicted execution bias feature output by the inverse inference network and the speckle execution bias feature extracted from the real log data; The inverse reasoning network and the dosage optimization model are jointly optimized during training.

12. The adaptive planning optimization method for particle radiotherapy according to claim 1, characterized in that, The quantitative evaluation values ​​include: The difference between the actual delivered dose distribution and the planned dose per voxel; and / or, Dosage statistics for each clinically relevant region, including target coverage metrics and dose-limiting metrics for organs at risk.

13. The adaptive planning optimization method for particle radiotherapy according to claim 12, characterized in that, The preset decision rules and corresponding adaptive dose adjustment instructions include at least one of the following: If the target area dose coverage is lower than the preset minimum coverage requirement, or if the dose to any organ at risk exceeds the preset safety limit, the online adaptive adjustment process is triggered and a speckle correction instruction is generated. If, based on the cumulative and trend analysis results of the quantitative assessment values ​​of multiple treatment fractions already performed, it is predicted that the cumulative total dose distribution of the patient's entire treatment course cannot meet the preset clinical overall goal, then the offline adaptive re-optimization process is triggered, and a treatment plan re-optimization instruction is generated.

14. The adaptive planning optimization method for particle radiotherapy according to claim 13, characterized in that, The beam spot correction command is generated by solving a constrained optimization problem. The constrained optimization problem aims to minimize the remaining beam spot weight adjustment while ensuring that the adjusted predicted dose meets clinical constraints.

15. A particle radiotherapy adaptive planning optimization device, characterized in that, include: The feature fusion module is used to construct a fusion state representation that simultaneously reflects the patient's anatomical state and treatment execution state based on the dynamic treatment log data generated by the patient during radiotherapy. A quantitative evaluation module is used to input the fusion state representation into a pre-trained dose optimization model; the dose optimization model is configured to directly output a quantitative evaluation value of the dose distribution generated by the portion of the current treatment fraction that has been performed, based on the fusion state representation. The optimization and adjustment module is used to automatically generate and execute adaptive dose adjustment instructions to guide the unexecuted portion of the current treatment or subsequent treatments, based on the quantitative evaluation value and by applying preset decision rules.

16. A radiotherapy system, characterized in that, include: The particle radiotherapy adaptive planning optimization device according to claim 15.

17. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it performs the steps of the method according to any one of claims 1-14 or implements the function of the device according to claim 15.

18. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions, and when the computer reads the computer instructions, the computer executes the steps of the method according to any one of claims 1-14 or implements the function of the device according to claim 15.