Monte Carlo rapid dose calculation method
By leveraging the structured sparseness and MIG characteristics of the NVIDIA Ampere architecture in Monte Carlo simulation, dynamically compressing data and optimizing GPU resources, the problem of slow calculation speed of Monte Carlo simulation is solved, and efficient particle transport calculation is achieved.
Patent Information
- Application Number
- CN202510530176.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The existing Monte Carlo simulation software has bottlenecks in computing speed, making it difficult to achieve near-real-time computing requirements in clinical applications.
A Monte Carlo fast dose calculation method adapted to NVIDIA Ampere and the newer generation architecture was designed to dynamically compress input data and optimize GPU resource utilization by leveraging structured sparseness and MIG characteristics.
It significantly improves the computing speed of Monte Carlo simulation, improves GPU resource utilization, and supports efficient computing needs for particle transportation.
Smart Images

Figure CN120216202A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of radiotherapy dose calculation, and particularly to a Monte Carlo fast dose calculation method. Background Art
[0002] The Monte Carlo dose calculation technology is a numerical simulation method relying on random sampling. When dealing with particle transport problems, this technology uses pseudo-random numbers generated by a computer to randomly sample relevant parameters in the physical transport process of particles according to a specific probability distribution, so as to simulate the motion trajectory of a single particle in a medium. As a recognized high-precision calculation method in the field of radiotherapy dose calculation, the Monte Carlo technology occupies an important position in medical physics and radiation dosimetry due to its accuracy in particle transport simulation and geometric modeling.
[0003] However, although general Monte Carlo simulation software, such as EGS4, EGSnrc, GEANT, MCNP, etc., uses high-precision transport algorithms and reaction cross-section data to simulate the transport process of particles with different energies in various media, the problem of its excessive calculation time is still the main bottleneck in clinical applications. Over the years, in order to improve the calculation efficiency and shorten the calculation time, researchers have proposed a variety of fast Monte Carlo calculation models based on different principles, which have significantly improved the calculation speed of the Monte Carlo method. However, to meet the requirement of near real-time calculation in clinical applications, further technological breakthroughs are still needed.
[0004] Since NVIDIA Corporation launched the GPU-based parallel computing architecture CUDA in 2006, using GPUs for Monte Carlo simulation has gradually become an important direction in dose calculation research. Compared with traditional CPU computing, GPUs have significant advantages in the cost and energy consumption of unit computing power, and these characteristics make them show great application potential in the field of medical physics. During this period, many research teams around the world have successively developed GPU-based Monte Carlo simulation tools, such as gDPM[2], GPUMCD, GMC, SMC, etc. [3-5]. These tools have significantly improved the calculation speed while maintaining the high precision of Monte Carlo simulation.
[0005] NVIDIA's Ampere architecture is a revolutionary GPU architecture launched in 2020 [6], representing a major breakthrough in GPU technology in terms of performance, energy efficiency, and functionality. This architecture is based on TSMC's 7-nanometer or Samsung's 8-nanometer manufacturing process, significantly enhancing transistor density and energy efficiency ratio, and providing powerful hardware support for fields such as high-performance computing, artificial intelligence, and graphics rendering. The core innovations of the Ampere architecture include the third-generation Tensor Core and the second-generation RT Core. The third-generation Tensor Core supports multiple data types (such as FP64, TF32, FP16, etc.), greatly accelerating deep learning training and inference tasks. At the same time, it introduces sparse computing technology, which further improves AI computing efficiency by leveraging the sparsity in neural networks. The second-generation RT Core significantly enhances real-time ray tracing performance through improved ray tracing algorithms and hardware acceleration, bringing more realistic visual effects to game development, film and television rendering, and industrial design.
[0006] The structural sparsity feature and the MIG feature of NVIDIA Ampere architecture are important innovations for enhancing GPU performance and utilization. The structural sparsity feature mainly uses pruning technology to set non-important parameters in the neural network to zero, thus generating a sparse network. In the NVIDIA Ampere architecture, this sparsity is presented in a 2:4 pattern, that is, at least two out of every four elements must be zero. This pattern reduces the data occupancy space and bandwidth of matrix multiplication operations by half. At the same time, by skipping the calculation of zero values through sparse Tensor Core, the computing throughput is theoretically doubled. This feature is particularly useful in deep learning inference because it can significantly accelerate matrix multiplication operations, improve the inference speed of the model, and maintain the accuracy of the model. The MIG (Multi-Instance GPU) feature allows a physical GPU to be safely partitioned into up to 7 independent GPU instances, each with its own independent memory path, cache, and computing core. This partitioning method ensures resource isolation for each instance, enabling multiple users or applications to use GPU resources simultaneously without interfering with each other. The MIG feature is especially suitable for workloads that are not fully saturated, maximizing GPU utilization by running different workloads in parallel. In addition, MIG supports multiple deployment configurations, including bare metal, virtual machine passthrough, and MIG-based vGPUs, providing flexibility for different application scenarios.
[0007] gDPM, GPUMCD, GMC and other Monte Carlo programs based on GPU memory have not fully utilized the advantages of NVIDIA's new Ampere architecture. There is still room for further improvement in the computational speed of parallelized Monte Carlo simulations. Based on the original Monte Carlo simulation, this method designs a set of computational methods adapted to the Ampere and later generations of architectures, thereby further improving the computational speed of GPU Monte Carlo simulations. Summary of the Invention
[0008] The object of the present invention is to provide a Monte Carlo fast dose calculation method that, based on the original Monte Carlo simulation, designs a set of computational methods adapted to NVIDIA Ampere and later generations of architecture graphics cards, thereby further improving the computational speed of GPU Monte Carlo simulations.
[0009] To achieve the above object, the present invention provides the following technical solutions: including the following steps: S1: Load simulation parameters; S2: Obtain the phase space data of the particle source of the radiotherapy accelerator; S3: Obtain the human anatomical structure image data and reconstruct it into a human three-dimensional matrix; S4: Mesh the three-dimensional human matrix; S5: Input the Monte Carlo dose calculation parameters and construct a refined Monte Carlo dose calculation physical model; S6: Utilize the structured sparsity of the Ampere architecture to dynamically compress the low-contribution voxels in the input matrix, reduce the video memory occupancy, data transmission overhead, and calculations. Adopt an "index-value" sparse storage mode for voxels. Assume that M is the length of the matrix, N is the width of the matrix, and L is the height of the matrix. Establish an index array to record its spatial coordinates (i, j, k), convert the original voxel matrix V ∈ R(M×N×L) into a sparse matrix, and perform dynamic compression on the low-contribution voxels in the input matrix to reduce the video memory occupancy, data transmission overhead, and computational amount; S7: Utilize the MIG feature of the Ampere architecture to input the data of each dose calculation unit for multi-instance MIG calculations and allocate them to different video memories, and store the corresponding particle simulation data in different MIG instance video memories in the same way; S8: In different MIG instances, start the transport calculation of particles and obtain the initial energy E and direction V of the particles by sampling; S9: After determining the energy E and direction V of the particles, adopt a sampling method, combine the particle cross-section data of different materials, judge the collision reactions of the particles in different cross-sections and the mean free path of the particle travel, calculate the dose deposition on the travel route, and save the results to the memory of the corresponding MIG instance; S10: Superimpose the normalized grid dose calculation results in each MIG instance to obtain the total radiation dose; S11: Copy the result to the CPU buffer and save the data in the buffer to disk.
[0010] Preferably, in S5 for particle movement step size sampling, according to the linear attenuation coefficient of the medium, random sampling of the step size is performed through a probability density function, which is closely related to the medium material, particle type, and energy, and the value of the corresponding medium is obtained in real time through a look-up table method.
[0011] Preferably, in the interaction type determination step of S5, based on the current energy E of the particle, combined with the atomic number, density, and other characteristics of the medium material, the pre-established cross-section database is called, and the occurrence probabilities of interactions such as collision, Compton scattering, photoelectric effect, and pair production are determined through probability calculation.
[0012] Preferably, when calculating the energy deposition in S5, according to the specific interaction type, the corresponding physical formula is used to calculate the energy deposition amount ΔE. In the photoelectric effect, the energy deposition is related to the electron binding energy and photon energy of the atom, and is calculated through the formula (Bi is the electron binding energy), and the remaining energy of the particle is updated to .
[0013] Preferably, in S6, a voxel contribution calculation model is constructed. For the voxels in the three-dimensional dose calculation space, their dose contribution C is determined by accumulating the energy deposition of particle transport. Among them, N is the number of particles passing through the voxel, ΔEn is the energy deposition of particle n, δ is the position determination function, a contribution threshold Dth is set, and the voxels below this threshold are determined as low contribution voxels. Among them, is a coefficient between 0 and 1. It is set to remove the voxels below the contribution threshold and use the voxels above the contribution threshold, considering the voxels with less influence on the particle collision process, reducing the video memory occupancy and calculation amount.
[0014] Preferably, in S7, the graphics card hardware resources are detected, including the number of SM units, video memory capacity, video memory bandwidth, and register file size. Check whether the graphics card supports the MIG feature. In the instance initialization stage, load the customized calculation kernel program, initialize the video memory space, establish the thread scheduling queue inside the instance to ensure the efficient execution of the calculation task. According to the size of the video memory occupied by the input data, configure different video memory occupancy units and reserve some shared resources for data interaction. The MIG instance division needs to meet the video memory capacity constraint conditions and reserve a dynamic adjustment space at the same time.
[0015] Preferably, after the particle simulation in S10, a calculation result will be obtained in each MIG instance. Superimpose the results of multiple dose calculations in the calculation area to obtain the total radiation dose distribution.
[0016] In summary, due to the adoption of the above technology, the beneficial effects of the present invention are: Compared with the prior art, the embodiment of the present application provides a GPU parallel Monte Carlo fast dose calculation method based on the NVIDIA architecture, which optimizes the Monte Carlo simulation particle transport process through two main methods. The present invention has significant advantages: using the structured sparse characteristics of the NVIDIA graphics card to dynamically compress input data, improve the effective utilization of video memory, and ensure the continuous operation of computing tasks; MIG instance management technology combines multi-dimensional scheduling with cross-instance collaboration to improve GPU resource utilization, especially in multi-task parallel scenarios. The two work together to give full play to the hardware performance of NVIDIA Ampere and newer generation architecture graphics cards, improve the efficiency of Monte Carlo dose calculation to a new level, reduce calculation time, and strongly support the GPU Monte Carlo calculation requirements for particle transport. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention, so that other features, purposes and advantages of the present invention become more obvious. The accompanying drawings of the exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings: Figure 1 It is a schematic diagram of the calculation process of the present invention. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.
[0019] In the description of the present invention, it is necessary to understand that the terms indicating orientation or positional relationship are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0020] In the present invention, unless otherwise clearly specified and defined, terms such as "installation", "connection", "linkage", "fixation" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific situation of the specification.
[0021] The present invention provides a Monte Carlo fast dose calculation method, a parallel Monte Carlo fast dose calculation method based on NVIDIA Ampere graphics cards, including the following steps S1: Load simulation parameters S2: Obtain the particle source phase space data of the radiotherapy accelerator S3: Obtain the human anatomical structure image data and reconstruct it into a human three-dimensional matrix S4: Mesh the three-dimensional human matrix S5: Input Monte Carlo dose calculation parameters and construct a refined Monte Carlo dose calculation physical model: For the particle motion step size sampling, according to the linear attenuation coefficient of the medium , perform random sampling of the step size through the probability density function , where is closely related to the medium material, particle type and energy, and the value of the corresponding medium is obtained in real time through the look-up table method; In the interaction type determination link, based on the current energy E of the particle, combined with the atomic number, density and other characteristics of the medium material, call the pre-established cross-section database, and determine the occurrence probabilities of interactions such as collision, Compton scattering, photoelectric effect, electron pair production, etc. through probability calculation; When calculating the energy deposition, according to the specific interaction type, use the corresponding physical formula to calculate the energy deposition amount ΔE. For example, in the photoelectric effect, the energy deposition is related to the electron binding energy of the atom and the photon energy, and is calculated by the formula (Bi is the electron binding energy), and the remaining energy of the particle is updated to .
[0022] S6: Utilize the structured sparsity of the Ampere architecture to dynamically compress the low - contribution voxels in the input matrix, reducing the video memory occupancy, data transfer overhead, and computation. Adopt the "index - value" sparse storage mode for voxels. Assume M is the length of the matrix, N is the width of the matrix, and L is the height of the matrix. Establish an index array to record their spatial coordinates (i, j, k), convert the original voxel matrix V ∈ R(M×N×L) into a sparse matrix, and implement dynamic compression on the low - contribution voxels in the input matrix, effectively reducing the video memory occupancy, data transfer overhead, and computational amount.
[0023] First, construct a voxel contribution degree calculation model. For the voxels in the three - dimensional dose calculation space , its dose contribution degree C is determined by the cumulative energy deposition of particle transport:
[0024] where N is the number of particles passing through the voxel, ΔEn is the energy deposition of particle n, and δ is the position determination function (δ = 1 when the particle is at (i, j, k), otherwise 0). Set a contribution degree threshold Dth, and the voxels below this threshold are determined as low - contribution voxels.
[0025]
[0026] where is a coefficient between 0 and 1. Set the ones below the contribution degree threshold to be removed, and only use the voxels above the contribution degree threshold, thus only considering the voxels with less impact on the particle collision process, reducing the video memory occupancy and computational amount.
[0027] S7: Utilize the MIG feature of the Ampere architecture to perform multi - instance MIG calculations on the data input of each dose calculation unit and allocate them to different video memories, and store the corresponding particle simulation data in the video memories of different MIG instances in the same way First, detect the graphics card hardware resources, including the number of SM units, video memory capacity, video memory bandwidth, register file size, etc. Check whether the graphics card supports the MIG feature. In the instance initialization stage, load the customized calculation kernel program, initialize the video memory space, and establish the thread scheduling queue inside the instance to ensure the efficient execution of the calculation task.
[0028] According to the size of the video memory occupancy of the input data, configure different video memory occupancy units and reserve some shared resources for data interaction. Therefore, the MIG instance division needs to meet the video memory capacity constraint conditions and reserve a dynamic adjustment space at the same time.
[0029] We first set the following core variables: Video memory of input data Vinput: Occupied by static data such as input data and medium parameters
[0030] Particle trajectory dynamic video memory Vdynamic: Particle states (position / energy / history) generated in real time
[0031] Computing reserved space Vcompute: Intermediate result cache, thread stack, burst memory growth buffer, storing intermediate data such as scattering cross-section temporary table, energy sink accumulation array, etc. Dynamic adjustment factor: Safety factor α (1.2 - 1.5), to cope with local video memory peak caused by particle clustering, dynamically adjusted based on historical task variance Video memory occupancy constraint for a single MIG
[0032] The number of MIGs is:
[0033] Rounding down to ensure that the number of MIG instances is an integer and can meet the requirements of all capacities, which is the maximum number of instances supported by the graphics card.
[0034] In a separate MIG instance, the number of SM units is allocated according to the set number of particles.
[0035] S8: In different MIG instances, start the transport calculation of particles. First, obtain the initial energy E and direction V of the particles by sampling.
[0036] S9: After determining the energy E and direction V of the particles, use sampling method, combined with particle cross-section data of different materials, to judge the collision reaction of particles in different cross-sections and the mean free path of particle travel, calculate the dose deposition on the travel route, and save the results to the memory of the corresponding MIG instance.
[0037] S10: Superimpose the normalized grid dose calculation results in each MIG instance to obtain the total radiation dose.
[0038] After the particle simulation is completed, a calculation result will be obtained in each MIG instance. Finally, we superimpose the results of multiple dose calculations in the calculation area to obtain the total radiation dose distribution
[0039] where n is the number of MIGs, Wn is the weight of the result of the nth MIG, is the result obtained from the nth MIG instance, For the total calculation result.
[0040] S11: Copy the result to the CPU buffer and then save the data in the buffer to the disk.
[0041] Example: Based on the existing GPU Monte Carlo calculation, the transport process of Monte Carlo simulated particles is optimized through two main methods to improve the calculation efficiency of the GPU Monte Carlo method and reduce the calculation time required. One is to perform dynamic compression optimization on the low-contribution voxels in the input data, only considering the voxels that have a greater impact on the particles during the calculation of particle transport, reducing the video memory occupancy and calculation of invalid voxels; the second is to use the MIG technology of the NVIDIA Ampere architecture for fine-grained configuration of instances, dividing the Monte Carlo program originally running on a single GPU into multiple independent entities for separate calculations, thereby improving the calculation speed.
[0042] This method specifically includes the following steps (as Figure 1 shown) Load the required data, including simulation data, particle cross-section data, etc. The loaded simulation data can be human body CT data, or measured phantom data, or entity data of other different materials. The loaded phase space data can also be the spatial distribution formula of particles, including energy spectrum, flux, etc., or the specific structural materials and shapes.
[0043] Input the Monte Carlo dose calculation parameters, construct a fine-grained Monte Carlo dose calculation physical model, and based on the basic principle of the Monte Carlo method, construct a highly refined particle transport model. In model construction, accurately define the key physical processes of particle movement in the medium: In Monte Carlo simulation, the particle movement step length is achieved by exponential distribution sampling: according to the linear attenuation coefficient of the medium (obtain the μ value of the current medium by looking up the table), generate a random step length that satisfies the probability density function to simulate the random collision behavior of particles in the medium. This process uses the inverse transform sampling instruction accelerated by hardware to ensure that the calculation time per million particle step lengths is < 5ms, providing a spatio-temporal distribution basis for dose deposition.
[0044] The energy deposition calculation is processed differently based on the reaction type: in the photoelectric effect, the energy deposition (Bi is the electron binding energy), and in Compton scattering, the energy loss is calculated according to the scattering angle formula, and the remaining energy of the particle is updated synchronously . This process uses the Tensor Core of Ampere to accelerate floating-point operations to ensure that the single-particle energy update delay is < 20ns, supporting the real-time tracking of millions of particles.
[0045] In this embodiment, the reaction of photons is taken as an example. In addition, the reaction principles of related particles such as protons and carbon ions are similar. Those skilled in the art should know the processes of collision and reaction of different particle types in different materials. The specific reaction particles and transport scenarios are not limited here.
[0046] Based on the structural sparsity characteristics of the NVIDIA Ampere architecture, Monte Carlo simulation is commonly used in radiotherapy for dose calculation and particle transport simulation, and these processes involve a large number of matrix operations. By using the structural sparsity compression technology of NVIDIA graphics cards, these matrices can be pruned and compressed, reducing the computational amount and storage requirements, thereby improving the efficiency and speed of simulation.
[0047] In this embodiment, dynamic compression optimization is performed on low-contribution voxels in dose calculation. The Ampere architecture supports structured sparse computing. Based on this, a "index-value" sparse storage mode is adopted for voxels. Assuming that M is the length of the matrix, N is the width of the matrix, and L is the height of the matrix, an index array is established to record its spatial coordinates (i, j, k), and the original voxel matrix V∈R(M×N×L) is converted into a sparse matrix, and dynamic compression is implemented on the low-contribution voxels in the input matrix, effectively reducing the video memory occupancy, data transmission overhead, and computational amount. The voxels adopt the "index-value" sparse storage mode, and its dilution storage mode can be one-dimensional index values such as D(i)=Value, or three-dimensional index values D(i, j, k)=Value, and different memory spaces are allocated according to different settings.
[0048] In an exemplary embodiment of the present invention, the determination of low-contribution voxels is determined by the formula in the invention content. Similarly, the determination of low-contribution voxels can also be obtained by the method of the region of interest of the user through calculation using one or a combination of physical factors and biomedical factors; where the physical factors reflect the material composition of the patient or the phantom and the irradiation physical conditions; the material composition of the patient or the phantom includes the density of the phantom, CT value, mass number, and atomic number; the irradiation physical conditions include: beam distribution, source distribution. The biomedical factors include: organ tissue irradiation threshold, biological sensitivity damage probability, etc.
[0049] In addition, for different applicable scenarios, the corresponding input phantom data will have obvious differences, and the voxel values of different phantoms will also have obvious differences in the collision reaction of particles. Different threshold settings can be selected to remove voxels. Taking CT as an example, if the gray value of CT is less than 1000, we can understand it as air, and its collision, blocking, absorption, etc. of photons or protons are very small. The voxels with a CT gray value less than 1000HU can be set to be removed, and only the voxels greater than or equal to 1000HU are retained.
[0050] Utilizing the MIG feature of the Ampere architecture, the data input of each dose calculation unit is calculated using multi-instance MIG and distributed to different video memories, and the corresponding particle simulation data is stored in different MIG instance memories in the same way.
[0051] First, check the graphics card hardware resources, including the number of SM units, video memory capacity, video memory bandwidth, register file size, whether the MIG feature is supported, etc. In the instance initialization phase, load the customized computing kernel program, initialize the video memory space, and establish the thread scheduling queue inside the instance to ensure efficient execution of computing tasks.
[0052] According to the size of the input data memory, different memory occupation units are configured, and some shared resources are reserved for data interaction. Therefore, the MIG instance division must meet the memory capacity constraint conditions and reserve dynamic adjustment space. For relevant calculation formulas, refer to the invention content section.
[0053] In different MIG instances, the particle transport calculation is started. First, the initial energy E and direction V of the particle are obtained by sampling. After the energy E and direction V of the particle are determined, the particle cross-section data of different materials are combined by sampling to determine the collision reaction of the particle in different cross-sections and the mean free path of the particle. The dose deposition on the travel path is calculated and the results are saved in the memory of the corresponding MIG instance. The normalized grid dose calculation results in each MIG instance are superimposed to obtain the total radiation dose. After the particle simulation is completed, a calculation result will be obtained in each MIG instance. Finally, we superimpose the results of multiple dose calculations in the calculation area to obtain the total radiation dose distribution. Finally, the results are copied to the CPU buffer, and then the data in the buffer is saved to the disk.
[0054] In Monte Carlo simulation, the MIG (multi-instance GPU) technology of the NVIDIA Ampere architecture is used to split computing tasks, which can achieve hardware-level resource isolation and parallel optimization: by dividing a single card into independent instances (such as 10GB video memory / 10SM or 40GB video memory / 40SM), each instance exclusively occupies the video memory and computing unit, avoiding resource competition among multiple tasks and improving stability; at the same time, it supports parallel processing of multiple cases (such as 3 MIG instances synchronously calculating 3 fields), combined with NVLink high-speed communication to achieve cross-instance particle data interaction, and throughput is increased by 3-5 times; dynamic video memory sharding (such as each instance only stores 20% of the dose matrix) combined with structural sparse compression, video memory usage is reduced by 60%, supporting real-time calculation of complex cases; efficiency is improved by 2 to 3 times compared with the traditional single instance mode.
[0055] In addition, the new generation of accelerated computing platform, Hopper architecture, launched by NVIDIA in 2022, is designed for large-scale AI, HPC, and supercomputing workloads, and also has the characteristics of structural sparsification and MIG. Therefore, the present invention is also applicable to Hopper architecture graphics cards, as well as all graphics cards released by NVIDIA in the future that support these two characteristics.
[0056] Compared with the prior art, the present invention has significant advantages: by utilizing the structural sparsification characteristics of NVIDIA graphics cards, the input data is dynamically compressed, improving the effective utilization rate of video memory and ensuring the continuous operation of computing tasks; the MIG instance management technology combines multi-dimensional scheduling and cross-instance collaboration, improving the utilization rate of GPU resources, especially in the multi-task parallel scenario. The two work together to give full play to the hardware performance of NVIDIA Ampere and newer generation architecture graphics cards, boosting the Monte Carlo dose calculation efficiency to a new level and strongly supporting the Monte Carlo calculation requirements for particle transport.
Claims
1. A Monte Carlo rapid dose calculation method, characterized in that: The following steps are involved: S1: Load simulation parameters; S2: Acquire the phase space data of the particle source of the radiotherapy accelerator; S3: Acquire human anatomical structure image data and reconstruct the human body three-dimensional matrix; S4: Mesh the human body three-dimensional matrix; S5: Input Monte Carlo dose calculation parameters and build a refined Monte Carlo dose calculation physical model; S6: Using the structured sparsity of the Ampere architecture, the low-contribution voxels in the input matrix are dynamically compressed to reduce video memory usage, data transmission overhead and calculation. The "index-value" sparse storage mode is used for voxels. Assuming that M is the length of the matrix, N is the width of the matrix, and L is the height of the matrix, an index array is established to record the spatial coordinates (i, j, k), and the original voxel matrix V∈R (M×N×L) is converted into a sparse matrix, and the low-contribution voxels in the input matrix are dynamically compressed; S7: Using the MIG feature of the Ampere architecture, a multi-instance MIG calculation method is used to process the input data of each dose calculation unit and distribute it to different video memories, and the corresponding particle simulation data is stored in different MIG instance video memories in the same way; S8: In different MIG instances, the particle transport calculation is started, and the initial energy E and direction V of the particles are obtained by sampling; S9: After determining the energy E and direction V of the particle, a sampling method is used to combine the particle cross-section data of different materials to determine the collision reaction of the particle in different cross-sections and the mean free path of the particle, calculate the dose deposition on the travel path, and save the result to the corresponding MIG instance memory; S10: superimpose the normalized grid dose calculation results in each MIG instance to obtain the total radiation dose; S11: copy the total radiation dose to the CPU buffer, and save the data in the buffer to the disk.
2. The Monte Carlo rapid dose calculation method according to claim 1, characterized in that: In S5, the step length of particle motion is sampled by randomly sampling the step length through a probability density function according to the linear attenuation coefficient of the medium. The sampling value is closely related to the medium material, particle type and energy, and the value of the corresponding medium is obtained in real time through a table lookup method.
3. A Monte Carlo rapid dose calculation method according to claim 2, characterized in that: In the interaction type determination stage in S5, based on the current energy E of the particle, combined with the atomic number and density characteristics of the dielectric material, a pre-established cross-section database is called to determine the probability of collision, Compton scattering, photoelectric effect, and electron pair interaction through probability calculation.
4. The Monte Carlo rapid dose calculation method according to claim 3, characterized in that: When calculating energy deposition in S5, the energy deposition amount ΔE is calculated using the corresponding physical formula according to the specific interaction type. In the photoelectric effect, energy deposition is related to the electron binding energy and photon energy of the atom. (Bi is the electron binding energy) calculation, update the particle residual energy to .
5. The Monte Carlo rapid dose calculation method according to claim 1, characterized in that: In the S6, a voxel contribution calculation model is constructed. For voxels in the three-dimensional dose calculation space, the dose contribution C is determined by accumulating the energy deposition of particle transport, wherein N is the number of particles passing through the voxel, ΔEn is the energy deposition of particle n, δ is the position determination function, and a contribution threshold Dth is set. Voxels below the threshold are determined as low-contribution voxels, wherein δ is a coefficient between 0 and 1. Voxels below the contribution threshold are removed, and voxels above the contribution threshold are used. Voxels that are affected by particle collisions are considered to reduce video memory usage and computational complexity.
6. The Monte Carlo rapid dose calculation method according to claim 1, characterized in that: The S7 detects the graphics card hardware resources, including the number of SM units, video memory capacity, video memory bandwidth, and register file size, and checks whether the graphics card supports the MIG feature. In the instance initialization phase, a customized computing kernel program is loaded, the video memory space is initialized, and a thread scheduling queue is established within the instance to ensure efficient execution of computing tasks. Different video memory occupation units are configured according to the size of the video memory occupied by the input data, and some shared resources are reserved for data interaction. The MIG instance division must meet the video memory capacity constraint and reserve dynamic adjustment space.
7. The Monte Carlo rapid dose calculation method according to claim 1, characterized in that: After the particle simulation in S10 is completed, a calculation result is obtained in each MIG instance, and the results of multiple dose calculations in the calculation area are superimposed to obtain the total radiation dose distribution.
Citation Information
Patent Citations
Method for accelerating voxel human body model dose evaluation based on GPU acceleration in radiation protection
CN104298871A
Monte Carlo grid parallel dose calculation method and device and storage medium
CN110504016A
Dose computation for radiation therapy using heterogeneity compensated superposition
US20140235923A1
Calculation of radiation therapy dose using all particle Monte Carlo transport
US5870697A
Cited By
Proton dose information processing device and method and electronic equipment
CN120493674A
Proton dose information processing apparatus, method, and electronic device
CN120493674B