Method and device for adjusting position of particle beam
By building and training the KAN network model, combining reinforcement learning and DBSCAN clustering, the automatic adjustment of particle beam position is achieved, which solves the problem of poor flexibility in the existing technology and improves the adjustment efficiency and accuracy.
Patent Information
- Application Number
- CN202510552047.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-02-10
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-12
AI Technical Summary
The existing particle beam adjustment methods are poor in flexibility and rely on manual adjustment, which is time-consuming and inefficient, making it difficult to respond quickly to environmental and equipment performance changes.
The KAN network model is used for automatic adjustment. By building the initial KAN network model, the particle beam training set is used for preliminary training, combined with reinforcement learning strategies and DBSCAN clustering, the target KAN network model is formed to realize automatic adjustment of the particle beam position.
It significantly reduces the need for manual intervention, reduces operating costs, improves equipment operation efficiency, and improves the flexibility and accuracy of particle beam adjustment.
Smart Images

Figure CN120473111A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of medical radiation dose calculation, and specifically to a method and device for adjusting the position of a particle beam. Background Art
[0002] In heavy ion and proton radiotherapy, precise particle beam adjustment is crucial to treatment effectiveness. For both heavy ion and proton radiotherapy, beam adjustment typically occurs in two phases: the accelerator beam diagnosis phase and the treatment room terminal commissioning phase. During the project installation and commissioning phase, as well as the subsequent operation and maintenance phase, particle beam adjustment becomes particularly complex due to environmental and equipment performance factors.
[0003] Existing particle beam adjustment methods mainly rely on manual adjustment. Manual beam adjustment requires the operator to have rich experience and professional knowledge. The adjustment process is time-consuming and difficult to quickly respond to changes in the environment and equipment performance. It has poor flexibility, low precision, and low efficiency, which increases treatment time and cost.
[0004] Therefore, it is urgent to provide a particle beam position adjustment method to at least solve the problem of poor flexibility in existing particle beam adjustment methods. Summary of the Invention
[0005] In view of this, the embodiments of the present application provide a method and device for adjusting the position of a particle beam, which can realize automatic adjustment of the position of the particle beam, solve the problem of poor flexibility in existing particle beam adjustment methods, significantly reduce the need for manual intervention, reduce operating costs and improve equipment operation efficiency.
[0006] To achieve the above objectives, in a first aspect, an embodiment of the present application provides a method for adjusting the position of a particle beam, comprising:
[0007] Determining effective data of the particle beam, wherein the effective data includes an initial position of the particle beam and an actual scanning magnet current value;
[0008] Constructing an initial KAN network model, wherein the initial KAN network model includes at least one selector and multiple adaptive network modules;
[0009] Performing preliminary training on the initial KAN network model using a particle beam training set to form a target KAN network model, wherein the particle beam training set includes a correspondence between a particle beam position and a scanning magnet current value, wherein the scanning magnet current value is a current value flowing through the scanning magnet;
[0010] The target KAN network model is used to predict the initial position of the particle beam, so as to adjust the initial position of the particle beam to the target position based on the prediction result.
[0011] Optionally, after the initial KAN network model is preliminarily trained using the particle beam training set and before the target KAN network model is formed, the method further includes:
[0012] The policy is re-adjusted using a reinforcement learning strategy that limits the magnitude of policy updates.
[0013] Optionally, the further adjustment using the reinforcement learning strategy includes:
[0014] The initial KAN network model after preliminary training is adjusted using the benefit ratio between gradients and the advantage function, and the network model with higher than the preset benefit is used as the target KAN network model. The benefit ratio between the gradients is the benefit of the current network model adopting the previous strategy position and the current strategy position. The advantage function is used to represent the benefit loss value corresponding to the current adjustment step under the previous strategy and the current strategy. The smaller the benefit loss value, the larger the advantage function.
[0015] Optionally, the preset benefit is determined based on the position benefit and the reward benefit, the position benefit is determined based on the benefits of the future position and the target position, and the reward benefit is the benefit corresponding to the Kullbeck-Leibler divergence of the selector.
[0016] Optionally, before predicting the initial position of the particle beam using the target KAN network model, the method further includes:
[0017] Evaluate the prediction quality of the target KAN network model based on some actual point pairs;
[0018] If the prediction quality of the target KAN network model does not meet the preset value, the reinforcement learning strategy is used to update it again.
[0019] Optionally, also include:
[0020] A target KAN network model is set in any treatment room, and multiple target KAN network models in multiple treatment rooms are transmitted to a central server;
[0021] Iterative training is performed based on the received global optimization model to optimize the current target KAN network model. The global optimization model is implemented by aggregating model weights of multiple target KAN network models by a central server.
[0022] Optionally, determining valid data of the particle beam includes:
[0023] Collect and process real-time data of particle beams;
[0024] Use DBSCAN clustering to divide the beam position and the target position according to the Euclidean distance to obtain the target cluster;
[0025] Data filtering is performed on the data in the target cluster to obtain valid data, wherein the data filtering is used to remove scanning points of the particle beam that are outside the deviation between the scanning point and the target beam position.
[0026] Optionally, the adaptive network module includes: a single-layer KAN unit and a double-layer KAN unit, the single-layer KAN unit is a one-dimensional function matrix, and the corresponding weight is 0; the double-layer KAN unit is used to fit white noise, and the corresponding weight is obtained by calculation.
[0027] In a second aspect, an embodiment of the present application provides a device for adjusting the position of a particle beam, comprising the following steps:
[0028] A data generation module is used to determine valid data of the particle beam, wherein the valid data includes an initial position of the particle beam and an actual scanning magnet current value;
[0029] An initial KAN network model construction module, configured to construct an initial KAN network model, wherein the initial KAN network model includes at least one selector and a plurality of adaptive network modules;
[0030] a training module for performing preliminary training on the initial KAN network model using a particle beam training set to form a target KAN network model, wherein the particle beam training set includes a correspondence between the position of the particle beam and the current value of the scanning magnet, wherein the scanning magnet current value is the current value flowing through the scanning magnet;
[0031] The target KAN network model is used to make predictions based on the initial position of the particle beam, so as to adjust the initial position of the particle beam to the target position based on the prediction results.
[0032] It can be seen that the particle beam position adjustment method in the present application constructs an initial KAN network model, performs preliminary training on the initial KAN network model using a particle beam training set to form a target KAN network model, and uses the target KAN network model to predict based on the initial position of the particle beam, so as to adjust the initial position of the particle beam to the target position based on the prediction result. In the embodiment of the present application, a trained KAN network (Kolmogorov-Arnold Networks) is adopted, and the ability of the KAN network to efficiently fit complex high-dimensional functions is utilized to automatically realize the alignment of the actual position of the particle beam and the target position of the particle beam, thereby avoiding manual adjustment by the operator, reducing the need for manual intervention, reducing operating costs and improving equipment operation efficiency.
[0033] In addition, the target KAN network model is a model with interpretability and visualization capabilities. The model is relatively simple, and the working mechanism can be understood intuitively. Each step of the adjustment process has a clear physical meaning and explanation basis, which makes it easier for operators to understand the adjustment mechanism and effectively control the operation of the target KAN network model. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0035] Figure 1 Schematic diagram of the particle beam formation principle.
[0036] Figure 2 A schematic diagram of the adjustment of the actual position point and the target position point of the particle beam provided in an embodiment of the present application.
[0037] Figure 3 A schematic diagram of the steps of the particle beam position adjustment method provided in an embodiment of the present application.
[0038] Figure 4 This is an optional block diagram of the target KAN network model provided in an embodiment of the present application.
[0039] Figure 5 Schematic diagram of the relationship between the target KAN network model and the treatment room provided in the embodiment of the present application.
[0040] Figure 6 This is a schematic diagram of the interaction of the optimization target KAN network model provided in an embodiment of the present application.
[0041] Figure 7 This is an optional block diagram of the particle beam position adjustment device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0043] As described in the background technology, the principle of particle beam position adjustment and position monitoring is as follows: Figure 1As described above, the monitoring system is specifically implemented by a deflection magnet, a treatment head position monitoring device, and an isocenter position monitoring device. Among them, the deflection magnet mainly uses two sets of scanning deflection magnets in the X and Y directions to adjust the beam position. The treatment head position monitoring device is generally fixedly installed on the treatment head, and can obtain the position of the particle beam at the treatment head for beam position calibration, monitoring, and interlocking; the isocenter position monitoring device can obtain the beam position at the isocenter, and is generally placed during position calibration and verification. In addition, there is a conversion relationship between the position of the particle beam at the treatment head position monitoring device and the isocenter position monitoring device and its distance from the deflection magnet. In the particle radiotherapy scenario, the current value of the deflection magnet and the deflection surface of the particle beam at the isocenter monitoring device position theoretically have the following relationship:
[0044] △I=k*△r+△b Formula 1
[0045] Where: x: coordinate value of a certain direction of the particle beam position (X+, X-, Y+, Y-);
[0046] I: Scanning magnet current value at a certain direction coordinate value;
[0047] k: the slope of the scanning magnet current value and the beam position coordinate value;
[0048] b: The intercept of the scanning magnet current value and the beam position coordinate value.
[0049] In order to align the actual position of the particle beam with the target position of the particle beam, the operator manually adjusts the slope value and intercept value to adjust the final position of the beam, such as Figure 2 As shown in the figure, ● is the actual position of the particle beam, and ★ is the target position of the particle beam. By uniformly shifting the particle beam position in a certain direction, that is, by manually adjusting the slope value, the actual position of the particle beam and the target position of the particle beam are finally made to coincide.
[0050] The above-mentioned manual adjustment of the slope and intercept values by the operator to adjust the final position of the beam can achieve particle calibration to a certain extent, but it takes a lot of time, making the use efficiency of the particle accelerator that forms the particle beam low. If there are multiple treatment rooms, they need to be calibrated one by one, which takes even longer. Based on this, the inventors have discovered that a trained Kolmogorov-Arnold Network (KAN Network) can be used. By leveraging the KAN network's ability to efficiently fit complex high-dimensional functions, the actual position of the particle beam can be automatically aligned with the target position of the particle beam, avoiding manual adjustments by the operator, reducing the need for human intervention, lowering operating costs, and improving equipment operation efficiency.
[0051] Specifically, an embodiment of the present invention provides a method for adjusting the position of a particle beam, such as Figure 3 As shown, the following steps are included:
[0052] Step S31, determining valid data of the particle beam, wherein the valid data includes an initial position of the particle beam and an actual scanning magnet current value;
[0053] Step S32: constructing an initial KAN network model, wherein the initial KAN network model includes at least one selector and multiple adaptive network modules;
[0054] In an optional embodiment, the adaptive network module may be an AdaptNet module.
[0055] Step S33: Preliminary training is performed on the initial KAN network model using a particle beam training set to form a target KAN network model, wherein the particle beam training set includes a correspondence between the position of the particle beam and the scanning magnet current value, and the scanning magnet current value is the current value flowing through the scanning magnet;
[0056] Step S34 : using the target KAN network model to predict based on the initial position of the particle beam, so as to adjust the initial position of the particle beam to the target position based on the prediction result.
[0057] The target KAN network model is used in the embodiment of the present invention to predict the initial position of the particle beam and adjust it to the target position. The target KAN network model in this application has an intelligent self-adjustment module and a reinforcement learning algorithm. The target KAN network model can autonomously optimize the adjustment process under different environments. Therefore, the particle beam position adjustment method in this application can realize automatic adjustment of the particle beam position, significantly reducing the need for manual intervention, reducing operating costs and improving equipment operation efficiency.
[0058] Further, the target KAN network model reference Figure 4 As shown, the system includes at least one selector and multiple AdaptNet modules. Based on the effective data of the particle beam, the selector selects the corresponding AdaptNet module, and uses the KAN unit in the AdaptNet module to predict the effective data and the target position, thereby providing a basis for position adjustment.
[0059] refer to Figure 4As mentioned above, the AdaptNet module consists of two branches, one of which is a single-layer KAN unit (the single black square in the figure) and the other is a double-layer KAN unit (the two side-by-side black squares in the figure). The single-layer KAN unit is used to represent the relationship between the K value and the current and position under ideal conditions, expressed as y = K x + b. The formula represented by this layer of the network is known information. Therefore, there is no need to update the weights during training of this layer of the network; the double-layer KAN unit is used to fit white noise, and the current network weights need to be updated each time training is performed.
[0060] Optionally, the steps for updating the current network weights include:
[0061] Step S41: Obtain the current value I corresponding to the current AdaptNet quadrant x I y As model input. x is the x-axis scanning magnet current value, I y is the y-axis scanning magnet current value;
[0062] Step S42: define a loss function L;
[0063]
[0064] In formula 2, x targe t and y target is the coordinate of the target position, x pred and y pred is the predicted position coordinate calculated using AdaptNet. The loss function L minimizes the sum of the squared deviations between the x and y coordinates to optimize the network output.
[0065] Step S43: Use the selected AdaptNet network to calculate the result (x pred ,y pred );
[0066] Step S44: Calculate the partial derivative of the loss function with respect to the network output using the loss function L on x pred and y pred Compute partial derivatives:
[0067]
[0068] Step S45: Use back propagation to calculate the weight gradient. Specifically, according to the chain rule, calculate the partial derivative of the loss function with respect to the weight W:
[0069]
[0070] Step S46: Update the network weight W based on the partial derivative of the loss function L with respect to the weight W. old To the latest network weight Wnew .
[0071]
[0072] Among them, η is the learning rate, which controls the step size of each update.
[0073] In an optional embodiment, the selector can select multiple AdaptNet modules to be updated each time, activating one or more AdaptNet modules for simultaneous updates. This update involves the KAN unit within the AdaptNet module simultaneously predicting and self-learning the available particle beam data based on the available data. Specifically, the KAN unit continuously optimizes its prediction and control capabilities through continuous feedback and self-adjustment, achieving self-updates.
[0074] In addition, the target KAN network model is a model with interpretability and visualization capabilities. The model is relatively simple, and the working mechanism can be understood intuitively. Each step of the adjustment process has a clear physical meaning and explanation basis, which makes it easier for operators to understand the adjustment mechanism and effectively control the system.
[0075] Furthermore, in order to obtain effective data of the particle beam, the specific steps may include:
[0076] Step S41, collecting and processing real-time data of the particle beam;
[0077] The real-time data usually includes the position coordinates of the particle beam in the X and Y directions and the corresponding scanning magnet current values. Each scanning point contains the actual position of the beam (position_x, position_y) and the scanning current (current_x, current_y).
[0078] Furthermore, to ensure the standardization and consistency of real-time data, the system will format the raw data, convert the units, and transform it into a unified coordinate system.
[0079] Step S42: using DBSCAN clustering to divide the beam position and the target position according to the Euclidean distance to obtain a target cluster;
[0080] The purpose of using DBSCAN clustering is to group similar scanning points into one category based on the proximity of their positions and current values.
[0081] Specifically, DBSCAN clustering is used to divide the beam position into groups based on the Euclidean distance between the beam position and the target position. This clustering process can aggregate the scan point data based on the beam deviation value, thus providing a clearer input for the KAN model.
[0082] The cluster distance calculation formula is as follows:
[0083]
[0084] Among them, d is the Euclidean distance between two points, position x1 is the horizontal coordinate of cluster point 1, position x2 is the horizontal coordinate of cluster point 2, position y1 is the ordinate of cluster point 1, position y2 is the vertical coordinate of cluster point 2.
[0085] Step S43: performing data filtering on the data in the target cluster to obtain valid data. The data filtering is used to remove the scanning points of the particle beam that are outside the deviation between the scanning point and the target beam position.
[0086] The data filtering criteria are based on the deviation between the scanning point and the target beam position. If the error of a scanning point exceeds the preset reasonable range, or the current value is not within the predetermined range, the data will be considered invalid and will be removed. The specific data filtering conditions are:
[0087] |△position x |<△and|△position y |<△ Formula 8
[0088] Among them, Δ represents the maximum allowable error value, Δposition x is the horizontal coordinate deviation between the scanning point and the target beam position, Δposition y is the vertical coordinate deviation between the scanning point and the target beam position.
[0089] In order to obtain the target KAN network model in this application, after the initial KAN network model is preliminarily trained using the particle beam training set, before forming the target KAN network model, the following steps are also included:
[0090] The policy is re-adjusted using a reinforcement learning strategy that limits the magnitude of policy updates.
[0091] In an optional embodiment, the reinforcement learning strategy may be a Proximal Policy Optimization (PPO) algorithm, which is used to limit the amplitude of the policy update to ensure smooth regulation.
[0092] Specifically, the readjustment using the reinforcement learning strategy includes:
[0093] The initial KAN network model after preliminary training is adjusted using the benefit ratio between gradients and the advantage function, and the network model with higher than the preset benefit is used as the target KAN network model. The benefit ratio between the gradients is the benefit of the current network model adopting the previous strategy position and the current strategy position. The advantage function is used to represent the benefit loss value corresponding to the current adjustment step under the previous strategy and the current strategy. The smaller the benefit loss value, the larger the advantage function.
[0094] In an optional embodiment, the calculation formula of the reinforcement learning strategy is:
[0095]
[0096] Among them, L PPO (θ) represents the loss function of the PPO algorithm, r t (θ) represents the ratio of the previous strategy position to the current strategy position, is the advantage function, ∈ controls the update amplitude.
[0097] Among them, the smaller the benefit loss value, the larger the advantage function, and the benefit loss value and the advantage function are negatively correlated.
[0098] Furthermore, the preset benefit of the target KAN network model is determined based on the position benefit and the reward benefit, wherein the position benefit is determined based on the benefits of the future position and the target position, and the reward benefit is the benefit corresponding to the Kulbeck-Leibler divergence of the selector.
[0099] The formula for calculating positional gain is:
[0100] R position =(|y future -y target |+|x futur ex target |) Formula 10
[0101] Among them, (x_future, y_future) are the coordinates of the future position, (x_target, y_target) are the coordinates of the target position, y_future is the profit corresponding to the vertical coordinate of the future position, y_target is the profit corresponding to the vertical coordinate of the target position, x_future is the profit corresponding to the horizontal coordinate of the future position, x_target is the profit corresponding to the horizontal coordinate of the target position, and Rpostion represents the total profit of the position.
[0102] The reward is calculated by the selector based on the Kullback-Leibler Divergence (KL Divergence). KL Divergence is used to measure the change between the new and old strategies. A small divergence indicates high selection stability. The formula for calculating the reward is:
[0103] |Rselection=-DKL(πcurrent||πprevious)-λ·a∈A∑I(a) Formula 11
[0104] Among them, |Rselection is the selection reward (SelectionReward), D KL (π current ||π previous ) is the current strategy (π current ) and the previous strategy (π previous ) between the Kulbeck-Lebler (KL) divergence; λ is the penalty weight factor; a∈A is the action in the decision space; I(a) is the indicator function; ∑I(a) is the total number of actions selected under the current strategy.
[0105] Furthermore, |Rselection is used to measure the adjustment benefit of the current strategy relative to the previous strategy, which contains dual constraints on strategy changes and the number of actions; D KL (π current |π previus ) is used to measure the magnitude of the strategy change; the penalty weight factor λ is used to control the impact of the number of selected actions on the benefit.
[0106] It should be noted that when the current strategy is significantly different from the previous strategy, the KL divergence increases, resulting in reduced returns. Therefore, D KL (π current |π previos ) can ensure that the policy update does not deviate too far from the previous policy to maintain stability.
[0107] In the process of calculating reward benefits in this application, we try to maintain the similarity with the previous strategy during adjustment (optimization using KL divergence) and reduce unnecessary action selection (optimization achieved using the penalty weight factor λ), thereby improving the stability and efficiency of the strategy.
[0108] Furthermore, the preset return R of the target KAN network model can be expressed as follows:
[0109] R=R position +R selection Formula 12
[0110] It can be seen that the preset benefit of the target KAN network model is the sum of position benefit and reward benefit.
[0111] In other embodiments, before predicting the initial position of the particle beam using the target KAN network model, the method further includes:
[0112] Evaluate the prediction quality of the target KAN network model based on some actual point pairs;
[0113] If the prediction quality of the target KAN network model does not meet the preset value, the reinforcement learning strategy is used to update it again.
[0114] The reinforcement learning strategy can be specifically the PPO algorithm, a policy gradient-based reinforcement learning algorithm that aims to stabilize the learning process by limiting the difference between the new policy and the old policy. Its core idea is to calculate the ratio between the new policy and the old policy at each update and adjust the gradient based on this ratio to ensure that the policy update amplitude is not too large.
[0115] In practical applications, particle accelerators are typically located in multiple treatment rooms, each of which may have different environmental conditions—such as temperature, humidity, particle beam stability, equipment layout, and personnel operating procedures. Therefore, to ensure that the particle beam regulation system in each treatment room operates optimally, a separate target KAN network model can be deployed for each treatment room. These models need to be optimized for the specific characteristics of their respective environments to maximize treatment effectiveness under varying conditions. Furthermore, the deployment and updating of these models must be able to rapidly respond to environmental changes and self-adjust through continuous learning to ensure treatment accuracy and safety.
[0116] In an alternative embodiment, reference Figure 5 As shown, a target KAN network model is set in any treatment room (as indicated by the white blocks in each treatment room in the figure), and multiple target KAN network models in multiple treatment rooms are transmitted to the synchronization ring 51. The synchronization ring 51 is also connected to the ion source 52. The ion source 52 is transmitted to the central server (not shown in the figure) through the synchronization ring 51. The central server can aggregate the model weights of the multiple target KAN network models to generate a global optimization model.
[0117] Iterative training is performed based on the received global optimization model to optimize the current target KAN network model. The global optimization model is implemented by aggregating model weights of multiple target KAN network models by a central server.
[0118] Because each treatment room has different environmental conditions, training and deploying a personalized model individually can achieve good results in the short term. However, over time, the model will gradually overfit to its specific environment, resulting in poor performance in other treatment rooms. To avoid this localized overfitting and leverage the common information across multiple treatment rooms, federated learning is used for iterative training and model adjustment.
[0119] Specifically, the process of iterative training using federated learning to achieve model adjustment can be referred to Figure 6 , the equipment in each treatment room uploads the target KAN network model in the local environment to the central server. The central server aggregates the weights of multiple target KAN network models to generate a global optimization model. The treatment room performs iterative training based on the received global optimization model to optimize the current target KAN network model and form the final target KAN network model.
[0120] In the current step, the global model not only retains the personalized adjustments of each treatment room, but also can share the experience of other treatment rooms, improving the overall performance and adaptability of the treatment system corresponding to multiple treatment rooms.
[0121] As an optional implementation of the disclosure of the embodiments of the present invention, refer to Figure 7 , an embodiment of the present invention further provides a particle beam position adjustment device, which may specifically include:
[0122] A data generation module 71 is used to determine valid data of the particle beam, wherein the valid data includes an initial position of the particle beam and an actual scanning magnet current value;
[0123] An initial KAN network model construction module 72 is used to construct an initial KAN network model, wherein the initial KAN network model includes at least one selector and multiple adaptive network modules;
[0124] a training module 73 for performing preliminary training on the initial KAN network model using a particle beam training set to form a target KAN network model, wherein the particle beam training set includes a correspondence between the position of the particle beam and the scanning magnet current value, wherein the scanning magnet current value is the current value flowing through the scanning magnet;
[0125] The target KAN network model 74 is used to make predictions based on the initial position of the particle beam, so as to adjust the initial position of the particle beam to the target position based on the prediction result.
[0126] In the particle beam position adjustment device provided in an embodiment of the present application, after the initial KAN network model is preliminarily trained using the particle beam training set and before the target KAN network model is formed, the device further includes:
[0127] The reinforcement learning module is used for readjustment using a reinforcement learning strategy, wherein the reinforcement learning strategy is used to limit the magnitude of the strategy update.
[0128] In the particle beam position adjustment device provided in an embodiment of the present application, the reinforcement learning module, used for the readjustment using the reinforcement learning strategy, includes:
[0129] The initial KAN network model after preliminary training is adjusted using the benefit ratio between gradients and the advantage function, and the network model with higher than the preset benefit is used as the target KAN network model. The benefit ratio between the gradients is the benefit of the current network model adopting the previous strategy position and the current strategy position. The advantage function is used to represent the benefit loss value corresponding to the current adjustment step under the previous strategy and the current strategy. The smaller the benefit loss value, the larger the advantage function.
[0130] In the particle beam position adjustment device provided in an embodiment of the present application, the preset benefit is determined based on the position benefit and the reward benefit, the position benefit is determined based on the benefits of the future position and the target position, and the reward benefit is the benefit corresponding to the Kullbeck-Leibler divergence of the selector.
[0131] In the particle beam position adjustment device provided in an embodiment of the present application, before predicting the initial position of the particle beam using the target KAN network model, the method further includes:
[0132] Evaluate the prediction quality of the target KAN network model based on some actual point pairs;
[0133] If the prediction quality of the target KAN network model does not meet the preset value, the reinforcement learning strategy is used to update it again.
[0134] The particle beam position adjustment device provided in the embodiment of the present application further includes:
[0135] A target KAN network model is set in any treatment room, and multiple target KAN network models in multiple treatment rooms are transmitted to a central server;
[0136] Iterative training is performed based on the received global optimization model to optimize the current target KAN network model. The global optimization model is implemented by aggregating model weights of multiple target KAN network models by a central server.
[0137] In the particle beam position adjustment device provided in the embodiment of the present application, determining valid data of the particle beam includes:
[0138] Collect and process real-time data of particle beams;
[0139] Use DBSCAN clustering to divide the beam position and the target position according to the Euclidean distance to obtain the target cluster;
[0140] Data filtering is performed on the data in the target cluster to obtain valid data, wherein the data filtering is used to remove scanning points of the particle beam that are outside the deviation between the scanning point and the target beam position.
[0141] In the particle beam position adjustment device provided in an embodiment of the present application, the adaptive network module includes: a single-layer KAN unit and a double-layer KAN unit. The single-layer KAN unit is a one-dimensional function matrix, and the corresponding weight is 0; the double-layer KAN unit is used to fit white noise, and the corresponding weight is obtained by calculation.
[0142] The above describes multiple embodiment schemes provided by the embodiments of the present application. The various optional methods introduced in each embodiment scheme can be combined and cross-referenced with each other without conflict, thereby extending a variety of possible embodiment schemes, which can all be considered as embodiment schemes disclosed and open in the embodiments of the present application.
[0143] Although the embodiments of the present application are disclosed above, the present application is not limited thereto. Any person skilled in the art may make various changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims.
Claims
1. A method for adjusting the position of a particle beam, characterized in that: The following steps are involved: Determining effective data of the particle beam, wherein the effective data includes an initial position of the particle beam and an actual scanning magnet current value; Constructing an initial KAN network model, wherein the initial KAN network model includes at least one selector and multiple adaptive network modules; Performing preliminary training on the initial KAN network model using a particle beam training set to form a target KAN network model, wherein the particle beam training set includes a correspondence between a particle beam position and a scanning magnet current value, wherein the scanning magnet current value is a current value flowing through the scanning magnet; The target KAN network model is used to predict the initial position of the particle beam, so as to adjust the initial position of the particle beam to the target position based on the prediction result.
2. The particle beam position adjustment method according to claim 1, wherein: After the initial KAN network model is preliminarily trained using the particle beam training set, and before the target KAN network model is formed, the method further includes: The policy is re-adjusted using a reinforcement learning strategy that limits the magnitude of policy updates.
3. The particle beam position adjustment method according to claim 2, characterized in that: The method of adjusting the reinforcement learning strategy again includes: The initial KAN network model after preliminary training is adjusted using the benefit ratio between gradients and the advantage function, and the network model with higher than the preset benefit is used as the target KAN network model. The benefit ratio between the gradients is the benefit of the current network model adopting the previous strategy position and the current strategy position. The advantage function is used to represent the benefit loss value corresponding to the current adjustment step under the previous strategy and the current strategy. The smaller the benefit loss value, the larger the advantage function.
4. The particle beam position adjustment method according to claim 3, wherein: The preset benefit is determined based on the position benefit and the reward benefit. The position benefit is determined based on the benefits of the future position and the target position. The reward benefit is the benefit corresponding to the Kullback-Leibler divergence of the selector.
5. The particle beam position adjustment method according to claim 1, wherein: Before predicting the initial position of the particle beam using the target KAN network model, the method further includes: Evaluate the prediction quality of the target KAN network model based on some actual point pairs; If the prediction quality of the target KAN network model does not meet the preset value, the reinforcement learning strategy is used to update it again.
6. The particle beam position adjustment method according to claim 1, wherein: Also includes: A target KAN network model is set in any treatment room, and multiple target KAN network models in multiple treatment rooms are transmitted to a central server; Iterative training is performed based on the received global optimization model to optimize the current target KAN network model. The global optimization model is implemented by aggregating model weights of multiple target KAN network models by a central server.
7. The particle beam position adjustment method according to claim 1, wherein: The effective data of the particle beam is determined, including: Collect and process real-time data of particle beams; Use DBSCAN clustering to divide the beam position and the target position according to the Euclidean distance to obtain the target cluster; Data filtering is performed on the data in the target cluster to obtain valid data, wherein the data filtering is used to remove scanning points of the particle beam that are outside the deviation between the scanning point and the target beam position.
8. The particle beam position adjustment method according to claim 1, wherein: The adaptive network module includes: a single-layer KAN unit and a double-layer KAN unit. The single-layer KAN unit is a one-dimensional function matrix, and the corresponding weight is 0; the double-layer KAN unit is used to fit white noise, and the corresponding weight is obtained by calculation.
9. A particle beam position adjustment device, characterized in that: The following steps are involved: A data generation module is used to determine valid data of the particle beam, wherein the valid data includes an initial position of the particle beam and an actual scanning magnet current value; An initial KAN network model construction module, configured to construct an initial KAN network model, wherein the initial KAN network model includes at least one selector and a plurality of adaptive network modules; a training module for performing preliminary training on the initial KAN network model using a particle beam training set to form a target KAN network model, wherein the particle beam training set includes a correspondence between the position of the particle beam and the current value of the scanning magnet, wherein the scanning magnet current value is the current value flowing through the scanning magnet; The target KAN network model is used to make predictions based on the initial position of the particle beam, so as to adjust the initial position of the particle beam to the target position based on the prediction results.