Agent-based beam parameter determination method, device, equipment, medium and program product

By using an agent-based beam parameter determination method, beam parameters are automatically optimized, solving the problems of reliance on manual intervention and low efficiency of gradient descent iterative optimization in existing technologies. This achieves efficient and precise beam parameter adjustment, improving the reliability and accuracy of heavy ion radiotherapy.

CN122245615APending Publication Date: 2026-06-19CAS ION MEDICAL TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CAS ION MEDICAL TECHNOLOGY CO LTD
Filing Date
2026-01-28
Publication Date
2026-06-19

Smart Images

  • Figure CN122245615A_ABST
    Figure CN122245615A_ABST
Patent Text Reader

Abstract

This disclosure provides a method, apparatus, device, and medium for determining beam parameters based on an intelligent agent, applicable to the fields of physics and computer science. The method includes: inputting currently acquired clinical data into a target intelligent agent, enabling the agent to construct a state space based on the clinical data, select a target beam adjustment action within the state space, and output an optimal beam weight combination; wherein the target beam adjustment action is obtained after evaluation based on a pre-set reward function, which is constructed based on clinical evaluation indicators; the beam weight combination is a parameter setting that the target intelligent agent believes achieves the optimal dose distribution in the current state; and generating a beam parameter adjustment command based on the beam weight combination output by the target intelligent agent, the beam parameter adjustment command being used to adjust the beam parameters of the treatment device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of physics and computer science, specifically to the application of physics and computer technology in the field of radiotherapy, and more specifically to a method, apparatus, device, medium, and program product for determining beam parameters based on intelligent agents. Background Technology

[0002] Heavy ion radiotherapy has become an important means of treating malignant tumors due to its advantages such as high precision and minimal damage to normal tissues. The adjustment of machine parameters is the core link in the design of heavy ion radiotherapy plans, and its accuracy directly determines the treatment effect and patient safety.

[0003] Currently, the mainstream methods for adjusting beam parameters in heavy ion radiotherapy mainly rely on a combination of manual intervention and gradient descent iterative optimization. These parameter adjustment methods have the following problems:

[0004] The process of manually adjusting parameters lacks a unified objective standard, relies on personal clinical experience, is inefficient and highly subjective; the optimization process and the final evaluation process of gradient descent iterative optimization are independent of each other, the optimization process cannot directly respond to the needs of the evaluation indicators, and can only be optimized by manually adjusting the loss function parameters multiple times, repeatedly performing the "optimization-evaluation" cycle, which not only further increases the complexity of operation, but also cannot guarantee that the optimization result will necessarily meet the evaluation requirements, reducing the reliability of beam parameter adjustment. Summary of the Invention

[0005] This application provides a method, apparatus, device, medium, and program product for determining beam parameters based on intelligent agents.

[0006] According to a first aspect of this application, a beam parameter determination method based on an intelligent agent is provided, comprising: inputting currently acquired clinical data into a target intelligent agent, so that the target intelligent agent constructs a state space based on the clinical data and selects a target beam adjustment action in the state space, and outputs an optimal beam weight combination; wherein, the target beam adjustment action is obtained after evaluation based on a pre-set reward function, and the reward function is constructed based on clinical evaluation indicators; the beam weight combination is a parameter setting that the target intelligent agent believes achieves the best dose distribution in the current state; and generating a beam parameter adjustment command based on the beam weight combination output by the target intelligent agent, the beam parameter adjustment command being used to adjust the beam parameters of the treatment device.

[0007] According to an embodiment of this application, the target agent is trained using training data. The training process includes: constructing an initial agent with beam parameter adjustment as the target task, the interaction environment of the initial agent being set as a heavy ion radiotherapy and treatment planning software environment; inputting training data into the initial agent, which then selects a beam adjustment action based on the current state space; the training data includes heavy ion radiotherapy cases under different conditions; calculating the reward value corresponding to the beam adjustment action using a reward function; the reward value being used to quantify the effect of the beam adjustment action on the treatment target; updating the parameters of the initial agent based on the reward value, so that the agent learns the correlation between different beam adjustment actions and reward values; repeating the above operations N times until the initial agent converges, and determining the converged initial agent as the target agent.

[0008] According to an embodiment of this application, the original treatment data is input to an initial agent, which selects a beam adjustment action based on the current state space. This includes: in response to the agent receiving the input training data, parsing the state space corresponding to the training data; the state space includes structural information, current dose distribution, and clinical constraints; and the agent determining the beam weight adjustment action based on the current state information.

[0009] According to an embodiment of this application, calculating the reward value corresponding to a beam adjustment action using a reward function includes: calculating the actual dose distribution corresponding to the beam adjustment action; and determining the reward value corresponding to the beam weight adjustment action based on the actual dose distribution using the reward function.

[0010] According to an embodiment of this application, updating the parameters of the initial agent based on the reward value includes: comparing the difference information between the expected reward of the current action and the actual reward obtained; and adjusting the parameters of the initial agent based on the difference information so that the agent tends to choose actions that can obtain higher rewards when faced with similar states.

[0011] According to an embodiment of this application, the reward function is a weighted combination of multiple evaluation metrics.

[0012] According to a second aspect of this application, a beam parameter determination device based on an intelligent agent is provided, comprising: an input module for inputting currently acquired clinical data into a target intelligent agent, so that the target intelligent agent constructs a state space based on the clinical data and selects a target beam adjustment action in the state space, and outputs an optimal beam weight combination; wherein the target beam adjustment action is obtained after evaluation based on a pre-set reward function, and the reward function is constructed based on clinical evaluation indicators; the beam weight combination is a parameter setting that the target intelligent agent believes achieves the best dose distribution in the current state; and a generation module for generating a beam parameter adjustment command based on the beam weight combination output by the target intelligent agent, the beam parameter adjustment command being used to adjust the beam parameters of the treatment device.

[0013] According to a third aspect of this application, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described method.

[0014] According to a fourth aspect of this application, a computer-readable storage medium is provided that stores a computer program or instructions thereon, characterized in that the computer program or instructions, when executed by a processor, implement the steps of the method described above.

[0015] According to a fifth aspect of this application, a computer program product is also provided, comprising a computer program or instructions that, when executed by a processor, implement the steps of the described method. Attached Figure Description

[0016] The above and other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0017] Figure 1 This diagram schematically illustrates a system architecture of an agent-based beam parameter determination method according to an embodiment of this application.

[0018] Figure 2 A flowchart illustrating an agent-based beam parameter determination method according to an embodiment of this application is shown schematically.

[0019] Figure 3 This illustration schematically shows a flowchart of training an initial agent based on training data to obtain a target agent according to an embodiment of this application;

[0020] Figure 4 This schematic diagram illustrates the structural block diagram of a particle beam relative biological effect determination device according to an embodiment of this application;

[0021] Figure 5 A block diagram of an electronic device for determining the relative biological effects of a particle beam according to an embodiment of this application is shown schematically. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of a feature, step, operation, and / or component, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0024] In this application, unless otherwise expressly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection, an electrical connection, or a connection that allows communication between them; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0025] In the description of this application, it should be understood that the terms "longitudinal", "length", "circumferential", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the subsystem or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0026] Throughout the accompanying drawings, identical elements are represented by the same or similar reference numerals. Conventional structures or configurations have been omitted where they may cause confusion in understanding this application. Furthermore, the shapes, dimensions, and positional relationships of the components in the drawings do not reflect actual size, scale, or actual positional relationships. Additionally, any reference symbols placed within parentheses should not be construed as limiting.

[0027] Similarly, to simplify this application and aid in understanding one or more of the various disclosed aspects, in the above description of exemplary embodiments of this application, various features of this application are sometimes grouped together into a single embodiment, figure, or description thereof. The use of terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicates that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0028] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0029] In the technical solution of this application, the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information all comply with relevant laws and regulations, necessary confidentiality measures have been taken, and they do not violate public order and good morals. In the technical solution of this application, user authorization or consent has been obtained before acquiring or collecting user personal information.

[0030] This application provides a method, apparatus, device, and medium for determining the relative biological effects of particle beams. Before introducing the technical solutions provided in this application, the relevant technologies involved in this application will be explained.

[0031] Currently, the mainstream method for adjusting beam parameters in heavy ion radiotherapy mainly relies on a combination of manual intervention and gradient descent iterative optimization. The specific implementation process is as follows:

[0032] 1. Technicians manually adjust and set the loss function, weights, and related parameter values ​​based on clinical experience;

[0033] 2. The gradient descent algorithm is used to perform N iterations to optimize the loss function (expected dose distribution and actual calculated dose distribution), and finally adjust the parameters of each beam.

[0034] 3. Use independent evaluation indicators to evaluate the optimized dose distribution results and determine whether they meet the clinical treatment requirements;

[0035] 4. If the clinical treatment requirements are met, exit the cycle; otherwise, continue with steps 1, 2, and 3 above.

[0036] Current techniques primarily quantify dose distribution deviations using loss functions and approximate optimal beam parameters through the iterative properties of gradient descent. However, as clinical demands for treatment precision and efficiency continue to rise, the inherent limitations of this method are becoming increasingly apparent, making it difficult to meet the evolving needs of modern heavy ion radiotherapy. The following problems exist:

[0037] 1. The weights and parameter values ​​of the loss function need to be manually adjusted by technicians, and this adjustment process lacks a unified objective standard, relying entirely on individual clinical experience. Differences in experience among different technicians may lead to deviations in parameter settings, and multiple trials are required to determine an optimal combination of loss functions. This is not only cumbersome but also time-consuming, seriously affecting the efficiency of treatment plan development.

[0038] 2. In clinical treatment, the desired dose distribution is that the dose is completely deposited within the tumor target area, completely avoiding normal organs at risk, thereby reducing the occurrence of complications. In reality, the distribution is affected by a variety of factors such as the patient's individual anatomical structure, tumor location and size, making it almost impossible to achieve the desired dose distribution. This leads to a deviation between the dose calculation distribution results and the ideal clinical requirements, affecting the treatment effect.

[0039] 3. The optimization process for beam parameter adjustment is independent of the final evaluation process. The optimization process is guided by the loss function, while the evaluation process relies on independent clinical evaluation indicators. Although there is some correlation between the two, they are not completely related. This means that the optimization process cannot directly respond to the needs of the evaluation indicators. Technicians can only manually adjust the optimization loss function parameters multiple times, repeatedly performing the "optimization-evaluation" cycle. This not only further increases the operational complexity but also cannot guarantee that the optimization results will necessarily meet the evaluation requirements, reducing the reliability of beam parameter adjustment.

[0040] To address the problems existing in the prior art, this disclosure provides a beam parameter determination method based on an intelligent agent, comprising: inputting currently acquired clinical data into a target intelligent agent, so that the target intelligent agent constructs a state space based on the clinical data and selects a target beam adjustment action in the state space, and outputs the optimal beam weight combination; wherein, the target beam adjustment action is obtained after evaluation based on a pre-set reward function, and the reward function is constructed based on clinical evaluation indicators; the beam weight combination is a parameter setting that the target intelligent agent believes achieves the best dose distribution in the current state; and generating a beam parameter adjustment command based on the beam weight combination output by the target intelligent agent, the beam parameter adjustment command being used to adjust the beam parameters of the treatment device.

[0041] The beam parameter determination method provided in this disclosure automatically adjusts and optimizes beam weights using an intelligent agent, eliminating the need for experienced technical personnel. This significantly reduces the time cost of dose prediction, improves the efficiency of treatment planning, and avoids subjective biases inherent in manual operations, ensuring the consistency of prediction results. Through global guidance from a reward function and multi-round interactive learning, the intelligent agent can explore the global beam weight combination space, effectively avoiding the tendency of gradient descent algorithms to get trapped in local optima. Its optimization objective directly targets clinical evaluation indicators, making dose prediction results closer to ideal clinical needs, improving the accuracy of dose distribution, thereby ensuring treatment efficacy and reducing damage to normal tissues.

[0042] By directly using the final clinical evaluation indicators as the value function (reward function) for reinforcement learning, the optimization process is directly linked to the evaluation requirements, forming an end-to-end closed loop integrating "optimization-evaluation". The model's learning process always focuses on meeting the evaluation indicators as the core objective, eliminating the need for subsequent independent evaluation verification and repeated adjustments. This significantly improves the reliability of dose prediction results and reduces treatment risks.

[0043] Figure 1 The diagram illustrates a system architecture of a beam parameter determination method based on an agent according to an embodiment of this application.

[0044] like Figure 1 As shown, application scenario 100 according to this embodiment may include terminal devices 101, 102, and 103, network 104, and server 105. Network 104 is used as a medium to provide a communication link between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0045] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as physics simulation applications, shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0046] Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0047] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0048] It should be noted that the agent-based beam parameter determination method provided in this application embodiment can generally be executed by server 105. Correspondingly, the agent-based beam parameter determination device provided in this application embodiment can generally be located in server 105. The agent-based beam parameter determination method provided in this application embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the agent-based beam parameter determination device provided in this application embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.

[0049] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0050] The following will be based on Figure 1 The described scene, through Figures 2-3 The beam parameter determination method based on intelligent agents according to embodiments of this disclosure will be described in detail.

[0051] Figure 2 A flowchart illustrating an agent-based beam parameter determination method according to an embodiment of this application is shown.

[0052] like Figure 2 As shown, the beam parameter determination method based on the intelligent agent in this embodiment includes operations S210 to S220.

[0053] In operation S210, the currently collected clinical data is input into the target agent so that the target agent can construct a state space based on the clinical data and select a target beam adjustment action in the state space, and output the optimal beam weight combination. The target beam adjustment action is obtained after evaluation based on a pre-set reward function, which is constructed based on clinical evaluation indicators. The beam weight combination is the parameter setting that the target agent believes achieves the best dose distribution in the current state.

[0054] In operation S220, a beam parameter adjustment command is generated based on the beam weight combination output by the target agent. The beam parameter adjustment command is used to adjust the beam parameters of the treatment device.

[0055] In some embodiments, clinical data may include, for example, images of the patient's anatomical structure, tumor target area contours, initial treatment plans, etc. By performing data preprocessing on the clinical data, the clinical data is converted into a format that the agent can recognize (such as dose distribution matrix, numericalization of clinical constraints, etc.), and the preprocessed data is input into the pre-trained target agent.

[0056] In response to received clinical data, the agent constructs a current state space based on the input patient anatomical imaging data, tumor target area, and organ at-risk contour information, combined with pre-defined clinical constraints. For example, it calculates the dose distribution of the tumor target area and normal tissues under the current beam weight combination based on the imaging data and contour information, and uses this distribution, along with the clinical constraints, as a feature representation of the state space. This state space reflects the current status and limitations of the treatment plan, providing a basis for the agent to make decisions.

[0057] Based on the constructed state space, the agent selects a beam weight adjustment action within its action space. The action space defines the adjustable range and adjustment step size of the beam weights. The agent selects what it considers the optimal beam weight adjustment action from these possible actions based on its own strategy (learned during training). For example, if the current state indicates that the dose to a certain part of the tumor target area is too low, the agent might choose to increase the weight of the corresponding beam.

[0058] The agent may also include a dose calculation module, which can simulate and calculate the actual dose distribution in the tumor target area and normal tissue after adjusting the beam weights based on the interaction principle between heavy ions and human tissues. A pre-defined reward function is used to evaluate the dose distribution after adjusting the beam weights, and the reward value corresponding to the action is calculated. The better the evaluation index, the higher the reward value. For example, if the dose uniformity index of the tumor target area increases and the probability of complications in normal tissues decreases after adjustment, then the agent will be given a higher reward value; conversely, the reward value will be lower. Beam adjustment actions that satisfy the preset conditions are identified as target beam adjustment actions, and the optimal beam weight combination is output.

[0059] In some embodiments, the target agent outputs the optimal beam weight combination under the current clinical data based on the learned optimal strategy. This beam weight combination is then converted into beam parameter adjustment instructions that the treatment device can execute, allowing the device to adjust the beam parameters accordingly. Beam weight is a key parameter controlling the intensity and distribution of the heavy ion beam; different beam weight combinations result in different radiation dose distributions received by the tumor target area and normal tissues. The optimal beam weight combination refers to the beam weight setting that, while meeting clinical constraints, achieves an ideal dose distribution in the tumor target area (e.g., good dose uniformity and high coverage) while minimizing damage to normal tissues.

[0060] This disclosure discloses embodiments that achieve automated optimization of beam parameters through a target intelligent agent, improving the efficiency and reliability of treatment planning. The target intelligent agent constructs a state space based on rich clinical data and selects beam adjustment actions and output beam weight combinations by comprehensively considering a reward function built from multiple clinical evaluation indicators. This allows for a more comprehensive and accurate analysis of the current treatment status, enabling precise adjustment of beam parameters. This results in more accurate dose distribution covering the tumor area while minimizing damage to surrounding normal tissues, thus improving the precision and effectiveness of treatment.

[0061] Figure 3 The illustration shows a flowchart of training an initial agent based on training data to obtain a target agent according to an embodiment of this application.

[0062] like Figure 3 As shown, this embodiment trains an initial agent based on training data to obtain a target agent, including operations S310~S350.

[0063] In operating S310, an initial agent is constructed with the objective of adjusting beam parameters. The interaction environment of the initial agent is set as the heavy ion radiotherapy and treatment planning software environment.

[0064] In some embodiments, a reinforcement learning agent (i.e., the initial agent) with beam parameter adjustment as its core task is constructed, and the interaction environment, state space, action space and reward function are set for the initial agent.

[0065] The interactive environment can be set up as a heavy ion radiotherapy and treatment planning software (TPS) environment for a specific patient. It includes key elements such as the patient's anatomical imaging data, tumor target area and organ at risk contour information, and TPS software settings. This provides a rich source of information for the initial agent's learning and decision-making, enabling it to learn how to adjust beam parameters in real-world scenarios and improve the applicability and reliability of the agent's decisions.

[0066] The state space can include the dose distribution and clinical constraints corresponding to the current beam weight combination (such as normal tissue dose thresholds, tumor target coverage requirements, etc.). This allows the agent to clearly understand the gap between the current treatment and the expected goal, facilitating the agent's learning and decision-making.

[0067] The action space can be the adjustable range and adjustment step size of the beam weight. The adjustable range and adjustment step size can be determined according to the operational limitations of the equipment and the rationality of the treatment plan design in actual treatment, so that the agent can explore and learn within the actual action space, generate an operable beam parameter adjustment scheme, and avoid proposing actions that exceed the equipment capabilities or do not conform to the treatment specifications.

[0068] The reward function can directly adopt clinically recognized final evaluation indicators, such as the tumor target area dose homogeneity index and the probability of complications in normal tissues, to directly link to the treatment goal, ensuring that the agent's learning objective is highly consistent with the clinical treatment objective. The better the evaluation indicator, the higher the reward value.

[0069] For example, the combination of reinforcement learning algorithms, reward functions, and state space representation methods can be flexibly adjusted according to the clinical data characteristics, treatment needs, and computational resource conditions of different patients to adapt to the treatment needs of patients with various tumor types and different anatomical structures, and have broad clinical application prospects.

[0070] In some embodiments, algorithms such as Q-learning, Deep Q-Network (DQN), and Proximal Policy Optimization (PPO) can be used to construct the intelligent agent. DQN is suitable for high-dimensional state-space scenarios and can improve dose prediction accuracy for patients with complex anatomical structures; the PPO algorithm is more stable and suitable for scenarios with limited clinical data; the A2C algorithm has high training efficiency and can meet the needs of rapid planning in emergency treatment cases. Those skilled in the art can choose appropriate algorithms to construct the intelligent agent according to the actual situation, and this application does not impose any limitations.

[0071] When operating the S320, the training data is input into the initial agent, which then selects the beam to adjust the action based on the current state space.

[0072] In some embodiments, in response to the agent receiving input training data, the state space corresponding to the training data is parsed; the state space includes structural information, current dose distribution, and clinical constraints; the agent determines the beam weight adjustment action based on the current state information.

[0073] In some embodiments, the training data includes a large amount of raw treatment data, covering heavy ion radiotherapy cases under different conditions, including treatment data for different types of tumor patients, data for different treatment stages, etc. The raw treatment data is input into an initial agent, which can transform the received raw treatment data into a current state space representation. The state space is an abstract description of the treatment scenario, and can represent information such as the patient's anatomical structure and current dose distribution in the form of vectors or matrices.

[0074] The initial agent selects beam adjustment actions based on the current state space using its internal algorithms and strategies, which may include adjusting the beam intensity, direction, energy, etc.

[0075] In operation S330, the reward function is used to calculate the reward value corresponding to the beam adjustment action; the reward value is used to quantitatively evaluate the effect of the beam adjustment action on the treatment target.

[0076] In some embodiments, a pre-defined reward function is used to calculate the reward value corresponding to the selected beam adjustment action. The reward value is used to quantify the effect of the beam adjustment action on the therapeutic objective. If an action can enable the tumor to receive more precise irradiation while reducing the damage to normal tissues, it will receive a higher reward value; conversely, if the action results in poor treatment effect or causes significant damage to normal tissues, the reward value will be lower.

[0077] During operation of S340, the parameters of the initial agent are updated according to the reward value so that the agent can learn the correlation between different beam adjustment actions and the reward.

[0078] In some embodiments, by adjusting the internal parameters of the agent, it can learn the correlation between different beam-adjusting actions and reward values. For example, if an action yields a high reward value, the agent will increase the probability of selecting that action; conversely, if the action's reward value is low, it will decrease the selection probability. In this way, the agent gradually optimizes its decision-making strategy.

[0079] In operation S350, repeat the above operation N times until the initial agent converges, and then determine the converged initial agent as the target agent.

[0080] In some embodiments, the operations of inputting data, selecting an action, calculating the reward value, and updating parameters are repeated N times. During each repetition, the initial agent continuously learns and improves. As the number of training iterations increases, the agent's understanding of the optimal beam adjustment actions for different treatment scenarios deepens, and its decision-making performance continuously improves. When the agent's performance no longer significantly improves, i.e., it reaches convergence, the converged initial agent is determined as the target agent. At this point, the target agent possesses a relatively accurate and stable beam parameter adjustment capability and can be applied to actual heavy ion radiotherapy.

[0081] This embodiment of the disclosure uses a large amount of raw treatment data for repeated training, enabling the target agent to fully learn the complex relationship between beam adjustment actions and treatment effects in different treatment scenarios. This allows the agent to more accurately select the optimal beam adjustment action when facing new treatment cases, thereby improving the accuracy and reliability of decision-making and providing patients with more precise treatment plans.

[0082] Because the training data encompasses a wide variety of treatment cases and scenarios, the target agent is exposed to diverse treatment environments during training. This allows the agent to adapt to different patient anatomy, tumor types, and treatment stages, demonstrating strong generalization ability and enhancing its adaptability across various treatment scenarios. Compared to traditional methods that rely on human experience and repeated trial-and-error adjustments of beam parameters, the target agent, through training, can make rapid decisions. It can analyze large amounts of treatment data in a short time and select the optimal beam adjustment action based on learned knowledge, significantly shortening the time required to determine treatment parameters, improving the efficiency of the entire treatment process, and enabling more patients to receive timely and effective treatment.

[0083] According to one embodiment of this disclosure, the reward value corresponding to the beam weight adjustment action is calculated using a reward function, including: calculating the actual dose distribution corresponding to the beam weight adjustment action; and determining the reward value corresponding to the beam weight adjustment action based on the actual dose distribution using the reward function.

[0084] In some embodiments, the agent may include a dose calculation module for performing dose distribution calculations. This module can perform dose distribution calculations using algorithms such as Monte Carlo simulation or pencil beam algorithms. In practical applications, the appropriate module can be selected or combined based on the requirements for computational accuracy and speed. Taking the Monte Carlo simulation algorithm as an example, the dose calculation module simulates the motion trajectories of a large number of heavy ion particles in the tissue based on beam adjustment actions and raw treatment data, and statistically analyzes the dose received by each voxel to obtain the actual dose distribution of the entire treatment area.

[0085] The reward function can be a weighted combination of multiple evaluation indicators (such as tumor target volume dose homogeneity, normal tissue complication probability NPV, target volume coverage, etc.). For example, by combining multiple key clinical indicators such as tumor target volume coverage, normal tissue dose limitation, and dose homogeneity index, the weight of each indicator can be determined through the analytic hierarchy process (AHP) to construct a comprehensive reward function, thereby further improving the clinical applicability of dose prediction results.

[0086] Evaluation indicators can be extracted from the actual dose distribution and calculated based on the reward function to obtain the reward value corresponding to the current beam weight combination. The merits of the beam adjustment action can be evaluated based on the calculated reward value, providing feedback information for subsequent reinforcement learning algorithms to guide the agent in selecting a better beam adjustment strategy.

[0087] Reward values, serving as feedback signals between the agent and the environment, provide clear direction for the agent's learning. By setting appropriate reward mechanisms, the agent can learn effective beam adjustment strategies with fewer attempts, shortening training time and improving learning efficiency. Calculating reward values ​​using a reward function allows for rapid and quantitative evaluation of the impact of each beam weight adjustment action on the treatment plan. The reward function used in this embodiment comprehensively considers multiple treatment objectives, such as dose coverage and homogeneity of the tumor target area, and protection of organs at risk, in order to find a globally optimal beam adjustment scheme and improve the overall quality of the treatment plan.

[0088] According to one embodiment of this disclosure, updating the parameters of the initial agent based on the reward value includes: comparing the difference information between the expected reward of the current action and the actual reward obtained; and adjusting the parameters of the initial agent based on the difference information so that the agent tends to choose actions that can obtain higher rewards when faced with similar states.

[0089] In some embodiments, the expected reward is an estimate of the reward value that can be obtained after performing a beam adjustment action, based on current knowledge, models, and understanding of the environment, before taking such an action. It reflects the agent's expectation of possible future outcomes and is an important basis for the agent's decision-making. The actual reward is the reward value calculated by the environment based on the actual result of the action after the agent performs it, using a predefined reward function. It is an objective feedback on the actual effect of the agent's action and is used to measure whether the action has achieved the expected goal.

[0090] The agent can update model parameters based on reward feedback using temporal difference learning (TD learning) or Monte Carlo methods, achieving the association learning of "action-reward". Temporal difference learning evaluates the value of the current state-action pair and adjusts the current value estimate using the actual and expected rewards of subsequent states. It does not require waiting for the treatment process to end; updates can be made after each interaction, making it suitable for real-time adjustments. For example, the agent predicts the potential cumulative reward based on the current state and action (e.g., "the current action may lead to increased target coverage"). It compares the difference between the actual reward and the expected reward (e.g., "the actual reward is 5 points higher than expected"). If the actual reward is higher than expected, the probability of that action being selected is increased; otherwise, the probability is decreased.

[0091] This disclosure updates agent parameters by dynamically comparing the difference between expected and actual rewards. Parameters can be updated without waiting for the entire session to end, significantly reducing training time. TD error calculation relies only on the current state-action pair, eliminating the need to store historical data, making it suitable for dynamic beam adjustment in real-time treatment (such as adaptive radiotherapy).

[0092] Figure 4A schematic block diagram of a beam parameter determination device based on an agent according to an embodiment of this application is shown.

[0093] like Figure 4 As shown, the beam parameter determination device 400 based on intelligent agents in this embodiment includes an input module 410 and a generation module 420.

[0094] The input module 410 is used to input the currently collected clinical data into the target agent, so that the target agent can construct a state space based on the clinical data, select a target beam adjustment action in the state space, and output the optimal beam weight combination. The target beam adjustment action is obtained after evaluation based on a pre-set reward function, which is constructed based on clinical evaluation indicators. The beam weight combination is a parameter setting that the target agent believes achieves the optimal dose distribution in the current state. In one embodiment, the input module 410 can be used to execute the operation S210 described above, which will not be repeated here.

[0095] The generation module 420 is used to generate a beam parameter adjustment instruction based on the beam weight combination output by the target agent. The beam parameter adjustment instruction is used to adjust the beam parameters of the treatment device. In one embodiment, the generation module 420 can be used to perform the operation S220 described above, which will not be repeated here.

[0096] According to embodiments of this application, any plurality of modules in the input module 410 and the generation module 420 may be combined into one module, or any one of these modules may be split into multiple modules. Alternatively, at least a portion of the functionality of one or more of these modules may be combined with at least a portion of the functionality of other modules and implemented in one module. According to embodiments of this application, at least one of the input module 410 and the generation module 420 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any appropriate combination of any of these three implementation methods. Alternatively, at least one of the input module 410 and the generation module 420 may be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.

[0097] Figure 5 A block diagram of an electronic device according to an embodiment of the agent-based beam parameter determination method of this application is illustrated.

[0098] like Figure 5As shown, an electronic device 500 according to an embodiment of this application includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage portion 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.

[0099] RAM 503 stores various programs and data required for the operation of electronic device 1100. Processor 501, ROM 502, and RAM 503 are interconnected via bus 504. Processor 501 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 502 and / or RAM 503. It should be noted that the program may also be stored in one or more memories other than ROM 502 and RAM 503. Processor 501 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in one or more memories.

[0100] According to embodiments of this application, the electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to a bus 504. The electronic device 500 may also include one or more of the following components connected to the input / output (I / O) interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 1107 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output (I / O) interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 510 as needed so that computer programs read from it can be installed into the storage section 508 as needed.

[0101] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0102] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0103] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should all be included within the protection scope of this application.

Claims

1. A method for determining beam parameters based on an intelligent agent, characterized in that, include: The currently collected clinical data is input into the target agent, so that the target agent can construct a state space based on the clinical data, select a target beam adjustment action in the state space, and output the optimal beam weight combination; wherein, the target beam adjustment action is obtained after evaluation based on a pre-set reward function, and the reward function is constructed based on clinical evaluation indicators; the beam weight combination is the parameter setting that the target agent believes achieves the best dose distribution in the current state; A beam parameter adjustment instruction is generated based on the beam weight combination output by the target intelligent agent. The beam parameter adjustment instruction is used to adjust the beam parameters of the treatment device.

2. The method according to claim 1, characterized in that, The target intelligent agent is obtained through training data, and the training process includes: An initial intelligent agent is constructed with the objective of adjusting beam parameters. The interaction environment of the initial intelligent agent is set as a heavy ion radiotherapy and treatment planning software environment. The training data is input into the initial agent, which selects the beam and adjusts the action based on the current state space; the training data includes heavy ion radiotherapy cases under different conditions; A reward function is used to calculate the reward value corresponding to the beam adjustment action; the reward value is used to quantitatively evaluate the effect of the beam adjustment action on the treatment target. The parameters of the initial agent are updated according to the reward value so that the agent learns the correlation between different beam adjustment actions and reward values; Repeat the above operation N times until the initial agent converges, and then determine the converged initial agent as the target agent.

3. The method according to claim 2, characterized in that, The process of inputting raw treatment data into the initial agent, which then selects beam adjustment actions based on the current state space, includes: In response to the intelligent agent receiving input training data, the state space corresponding to the training data is parsed; the state space includes structural information, current dose distribution, and clinical constraints. The agent determines the beam weight adjustment action based on the current state information.

4. The method according to claim 2, characterized in that, The calculation of the reward value corresponding to the beam adjustment action using the reward function includes: Calculate the actual dose distribution corresponding to the beam adjustment action; The reward value corresponding to the beam weight adjustment action is determined by the reward function based on the actual dose distribution.

5. The method according to claim 2, characterized in that, The step of updating the parameters of the initial agent based on the reward value includes: Compare the expected reward of the current action with the actual reward received. The parameters of the initial agent are adjusted based on the difference information so that the agent tends to choose actions that yield higher rewards when faced with similar states.

6. The method according to claim 2, characterized in that, The reward function is a weighted combination of multiple evaluation indicators.

7. A beam parameter determination device based on an intelligent agent, characterized in that, include: The input module is used to input the currently collected clinical data into the target agent, so that the target agent can construct a state space based on the clinical data, select a target beam adjustment action in the state space, and output the optimal beam weight combination; wherein, the target beam adjustment action is obtained after evaluation based on a pre-set reward function, and the reward function is constructed based on clinical evaluation indicators; the beam weight combination is the parameter setting that the target agent believes achieves the best dose distribution in the current state; The generation module is used to generate a beam parameter adjustment instruction based on the beam weight combination output by the target intelligent agent. The beam parameter adjustment instruction is used to adjust the beam parameters of the treatment device.

8. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 6.