Software and hardware collaborative design method for three-dimensional imaging internal structure reconstruction
By designing a three-nested loop general programming model and a Waffle hardware accelerator, the algorithm for reconstructing the internal structure of 3D imaging was optimized, solving the problems of positioning and synchronization, and achieving efficient 3D reconstruction.
Patent Information
- Application Number
- CN202410597111.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-14
- Publication Date
- 2025-11-14
AI Technical Summary
Existing 3D imaging internal structure reconstruction technology suffers from poor positioning and high synchronization overhead, resulting in slow reconstruction speed and difficulty in achieving both quality and accuracy.
A three-nested loop general programming model is adopted, and the data flow is optimized by the voxel-driven method and a Waffle hardware accelerator is built to improve locality and synchronization, enhance data locality, and reduce redundant computation.
It achieves efficient execution of the 3D reconstruction algorithm, improves reconstruction speed and quality, and meets the requirements of data dependency and synchronization.
Smart Images

Figure CN120953474A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of three-dimensional imaging internal structure reconstruction technology, specifically to a hardware and software co-design method for three-dimensional imaging internal structure reconstruction, which is used to improve the performance and efficiency of three-dimensional image reconstruction algorithms. Background Technology
[0002] Three-dimensional volumetric imaging has been widely used in medical diagnostics (such as CT and PET), microscopic analysis, non-destructive testing (such as customs inspection), archaeology, and geological exploration. The demand for lung CT scans has surged, especially during the COVID-19 pandemic. Low-dose CT, which reduces cumulative radiation dose and is beneficial to human health, is highly favored. However, the process of reconstructing three-dimensional models using low-dose CT is time-consuming and slow, making it difficult to meet the practical needs of the rapidly increasing number of patients. FPGA and GPU implementations specifically for CT reconstruction cannot solve the locality problem and are not universally applicable to the reconstruction of three-dimensional internal structures.
[0003] Given the widespread applications of 3D imaging and recent practical problems, existing 3D reconstruction technologies mainly suffer from the following shortcomings:
[0004] (1) Poor localization: Due to the irregular access of rays and voxels in the reconstruction algorithm, the cache utilization efficiency is low.
[0005] (2) High synchronization overhead: There are a lot of synchronization operations during the reconstruction algorithm iteration process, which limits the performance improvement.
[0006] (2) Reconstruction task partitioning is required: Due to localization and synchronization issues, traditional methods have to divide the reconstruction task into multiple parts for local gradient descent, resulting in slow algorithm convergence.
[0007] In summary, existing 3D image internal structure reconstruction techniques struggle to balance reconstruction speed and reconstruction quality. Summary of the Invention
[0008] To overcome the shortcomings of the existing technologies, this invention proposes a hardware-software co-design solution for the reconstruction of internal structures in 3D imaging. This co-design fully leverages the advantages of near-memory computing, including establishing a three-nested loop general programming model, redesigning a hardware-friendly algorithm, proposing a novel near-memory computing architecture, Waffle. Experimental results verify the overall solution's effectiveness in improving speed and energy efficiency for the reconstruction of internal structures in 3D imaging.
[0009] The technical solution of this invention comprises three parts: establishing a three-nested loop general programming model covering different 3D reconstruction tasks; redesigning the data flow of the algorithm to improve data locality and reducing redundant computation through pre-computation and simplified synchronization schemes; and constructing a Waffle hardware accelerator with a near-memory computing architecture to accelerate algorithm execution and perform related operations under customized computational logic. Specifically, it includes the following method steps:
[0010] (1) For the three-dimensional ray-driven iterative reconstruction model, the Mumford-Shah functional, which is used to measure the attenuation coefficient and edge smoothness of the three-dimensional image, is introduced to optimize the attenuation coefficient and edge tensor of the three-dimensional reconstruction result.
[0011] This paper summarizes the problems existing in 3D ray-driven iterative reconstruction, which involves reconstructing the attenuation coefficient of an object through forward and backward projection. In forward projection, the light beam passes through the object, and the intersection length and absorption coefficient of each pixel with the beam are calculated. In backward projection, the absorption coefficient is updated to detect the attenuation. However, in practical applications, simple backward projection may lead to unsmooth reconstruction results and edge oscillations, which do not conform to the smooth characteristics of natural objects and reduce the reconstruction quality. Therefore, this invention introduces the Mumford-Shah functional, which measures the attenuation coefficient and edge smoothness, and optimizes the results of the attenuation coefficient and edge tensor by minimizing this functional.
[0012] (2) Establish a general programming model for three nested loops, where the nested loops are an outer loop, a parallelized loop, and an inner loop;
[0013] To address the locality and synchronization issues in 3D reconstruction, a general programming model is proposed. This model implements 3D reconstruction task allocation, parallelization, and internal processing through three nested loops. The outer loop divides the 3D reconstruction task into multiple parts based on voxels or rays to solve localization and synchronization problems. The parallelization loop performs forward and backward projection on each core (processor kernel) and synchronizes the projections. The inner loop processes the forward and backward projections of beams on a single core.
[0014] (3) The voxel-driven method was adopted, and the synchronization operation before updating the decay coefficient f and the edge tensor υ in the three nested loop general programming model established in step (2) was deleted. The three-dimensional reconstruction algorithm was optimized and the optimized data dependency (i.e. data flow) was obtained.
[0015] The voxel-driven method reduces redundant computation and improves data locality by storing the voxel's attenuation coefficient f, intersection length L, and beam ID, and by using pre-computation and simplified synchronization schemes.
[0016] Taking forward projection as an example, the voxel-driven method stores the attenuation coefficient f and the IDs of all beams passing through it for each voxel. The intersection length L between the beam and the voxel is pre-calculated and stored in the voxel data structure. When traversing voxels, the attenuation is calculated based on the stored attenuation coefficient f and intersection length L, and the attenuation of the relevant beams is updated.
[0017] In the voxel-driven method, the decay coefficient f and edge tensor υ of each voxel are updated directly without affecting other voxels. Therefore, the voxel-driven method eliminates the synchronization operation before updating the decay coefficient f and edge tensor υ in the original general programming model.
[0018] The calculation process of updating voxel features using 2(∑fL-g)L for forward projection and backward projection shares the multiplication logic, reusing hardware resources in the two stages of the inner loop to reduce computation and storage overhead.
[0019] The outer loop in the three nested loop general programming model established in step (2) is removed, making the algorithm converge faster.
[0020] The 3D reconstruction algorithm, after the above optimizations, achieves a new data dependency (data flow).
[0021] (4) Build a hardware accelerator.
[0022] The Waffle hardware accelerator is constructed using logic chips and multiple memory chips forming a cubic structure connected via through-silicon vias (TSVs). Data structures are allocated for voxels and beams within each cube to store voxel and beam information. The storage layout employs a serpentine approach to traverse voxels. The Waffle hardware accelerator includes: an offset decoder, a getter, a Mumford-Shah functional solver, a communicator, an exclusive updater, and processor elements (PEs).
[0023] The first stage of the data stream performs forward projection and related calculations; the offset decoder and acquirer decode the address offset of each voxel based on the number of beams passing through the voxel and acquire the corresponding data; the Mumford-Shah functional solver is used for calculation; the communicator is used for communication; and the exclusive updater implements conflict-free updates of the beam buffer.
[0024] The second phase of the data stream performs back projection and update operations; the offset decoder and getter are used to traverse voxels; the Mumford-Shah functional solver is used for computation; and the exclusive updater implements conflict-free updates to the voxel buffer.
[0025] The Mumford-Shah functional solver is used to compute all expressions containing (Mumford-Shah) MS functionals. That is, the Mumford-Shah functional solver part is used to compute... and (Where f represents the decay coefficient, υ represents the edge tensor, and MS represents the Mumford-Shah functional).
[0026] Exclusive updaters are used to implement conflict-free updates of beam buffers and voxel buffers.
[0027] Through the above steps, the reconstruction of the internal structure of 3D imaging through hardware and software collaboration is achieved.
[0028] Compared with the prior art, the present invention has the following technical effects:
[0029] This invention defines a three-nested loop general programming model to improve the locality and synchronization problems of 3D reconstruction; optimizes the algorithm by adopting a voxel-driven method instead of the original ray-driven method to improve locality and reduce synchronization overhead; and designs a near-memory architecture Waffle hardware accelerator to address the data dependencies and synchronization problems in the reconstruction algorithm. Through optimization of data structures and hardware design, the Waffle hardware accelerator further improves performance and energy efficiency.
[0030] The three-nested loop general programming model of this invention overcomes the locality and synchronization problems of traditional methods, improving the performance and efficiency of 3D reconstruction. By redesigning the algorithm's data flow, high efficiency and optimized performance of the reconstruction algorithm are achieved. High-efficiency execution of the reconstruction algorithm is realized through a Waffle hardware accelerator employing a specific near-memory architecture and data flow scheme. This hardware-software co-design solution for 3D internal structure reconstruction improves performance and accelerates the reconstruction process while meeting the requirements of data dependency and synchronization. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the three-dimensional reconstruction process of an object proposed in this invention;
[0032] Figure 1 Chinese f k,i L represents the attenuation coefficient of the i-th pixel through which beam k passes. k,i The length of the intersection point between beam k and the i-th pixel is represented by g, and the attenuation is represented by the subscript (the subscript indicates the beam number, e.g., g). k G represents the attenuation of beam k. n (This represents the attenuation of beam n).
[0033] Figure 2 This is the programming model proposed in this invention;
[0034] Figure 3 This is a schematic diagram illustrating the data dependency relationship proposed in this invention;
[0035] Figure 3In this context, f represents the attenuation coefficient, υ represents the edge tensor, L represents the intersection length of the beam and the voxel, g represents the attenuation degree, and MS represents the Mumford-Shah functional.
[0036] Figure 4 This is a schematic diagram of the data flow after the algorithm proposed in this invention is optimized;
[0037] Figure 5 This is an overview diagram of the Waffle hardware accelerator architecture proposed in this invention;
[0038] Figure 5 In the context of kernel processing, PE represents the processor element; in voxel data structures, f represents the attenuation coefficient, υ represents the edge tensor, and info... b Information about the beam passing through the voxel, id b Indicates a unique ID for the beam, end b Indicates whether it is the last voxel of the nucleus, L b Indicates the length of the intersection point between the beam and the voxel; beam data structure: id b Indicates a unique ID, at b Indicates decay rate, lastCore b Indicates the last magnetic core the light beam passed through, nextCore b Indicates the next magnetic core the light beam passes through, sync b Indicates whether the previous core has passed part of the result to the current core, finish b This indicates whether the current nucleus has completed partial forward projection of the beam; TSV (through-silicon-via) represents through-silicon via technology. In this invention, the symbol g is used to represent attenuation during calculation and inference; in the structure of the hardware accelerator, att is used. b Represents the attenuation degree, such as Figure 5 As shown.
[0039] Figure 6 This is a schematic diagram illustrating the division and storage of the cube reconstruction object proposed in this invention;
[0040] Where m and n represent the division of the cube into m*n parts, and p represents the parallelism of the Mumford-Shah functional solver. Detailed Implementation
[0041] The present invention will be further described below with reference to the accompanying drawings and embodiments, but the scope of the invention is not limited in any way.
[0042] This invention provides a hardware-software co-design solution for 3D internal structure reconstruction. It introduces a unified programming model, redesigns and optimizes the algorithm, and constructs a Waffle hardware accelerator with a near-memory computing architecture. Through hardware-software co-design, effective task allocation, data storage, and synchronization mechanisms, the speed, quality, and efficiency of 3D reconstruction are improved.
[0043] The solution of this invention mainly includes proposing a general programming model to address the locality and synchronization problems in 3D reconstruction. This model implements task allocation, parallelization, and internal processing through three nested loops. To improve the locality problem in 3D internal reconstruction, a voxel-driven method is used instead of the traditional ray-driven method. Synchronization overhead is reduced by accurately analyzing data dependencies and eliminating unnecessary synchronization. A new data flow is implemented through algorithm optimization. A Waffle hardware accelerator with a specific near-memory architecture and data flow scheme is designed to achieve efficient execution of the reconstruction algorithm. Specifically, it includes:
[0044] 1. Construct a three-dimensional ray-driven iterative reconstruction model.
[0045] 1.1 Obtain the attenuation coefficient f for each pixel (or voxel), i.e., the 3D reconstruction target;
[0046] The goal of the reconstruction is to obtain the attenuation coefficient f for each pixel (or voxel). The forward projection process of a light beam passing through an object: when the light beam k penetrates the object (e.g., ... Figure 1 The arrow in the image represents the attenuation, which can be expressed as ∑ i f k,i L k,i (f k,i L represents the attenuation coefficient of the i-th pixel through which beam k passes. k,i This represents the length of the intersection point between beam k and the i-th pixel. After beam k leaves the object, its attenuation g will be detected. k Back projection process: implemented by the gradient descent algorithm, i.e., using 2(∑ i f k,i L k,i -g k )L k,i Update f k,i To minimize (∑ i f k,i L k,i -g k ) 2 .
[0047] 1.2 The Mumford-Shah functional is introduced to measure the attenuation coefficient and edge smoothness of the 3D reconstruction results;
[0048] Since simple backprojection can lead to unsmoothness and edge oscillations, the Mumford-Shah (MS) functional is introduced to measure the attenuation coefficient and edge smoothness. The Mumford-Shah functional introduces an additional tensor υ representing the probability that each voxel belongs to the edge of the reconstructed object. The gradient descent algorithm is used to minimize (∑... i f k,i L k,i -g k ) 2 +MS. Where, f k,i L represents the attenuation coefficient of the i-th pixel through which beam k passes. k,i denoted by g, which represents the length of the intersection point between beam k and the i-th pixel. k Let k represent the attenuation of the beam. MS represents the Mumford-Shah functional. The gradient is calculated using the following formula: (α, β, ε are hyperparameters) Δ represents the two-dimensional operators for divergence, gradient, and Laplacian, respectively.
[0049] 1.3 To achieve good reconstruction results, multiple beams from multiple directions are needed to cover the entire object, such as parallel beams, fan beams, and conical beams, which are projection geometries.
[0050] 2. Establish a programming model, such as Figure 2 The model has three nested loops for task partitioning, parallelization, and processing of beams on each core.
[0051] 2.1 The outer loop divides the task and solves the positioning and synchronization problems. The outer loop divides the task into several parts based on voxels or rays.
[0052] 2.2 Parallelized Loops Assign Tasks and Synchronize Each Core. Parallelized loops are assigned to each core, processing the forward and backward projections of the beam in parallel on each core. Cores in all sections are synchronized before updating the attenuation coefficient f and the edge tensor υ.
[0053] 2.3 The inner loop is the forward and backward projection of the beam processed by the kernel.
[0054] 3. Algorithm optimization.
[0055] 3.1 The voxel-driven method (traversing voxels and finding all beams passing through them) reduces redundant computation, improves data locality, and alleviates synchronization overhead, making it superior to the previous ray-driven method (traversing beams and finding all voxels passed through). Taking forward projection as an example, for each voxel, the attenuation coefficient f and the IDs of all beams passing through it are stored. The intersection length L between the beam and the voxel is pre-calculated and stored in the voxel data structure. When traversing voxels, the attenuation is calculated based on the stored attenuation coefficient f and intersection length L, and the attenuation of the relevant beams is updated.
[0056] 3.2 Establishing data dependencies, such as Figure 3 .
[0057] Back projection can be divided into three parts: calculating 2(∑fL-g)L, (f represents the attenuation coefficient, υ represents the edge tensor, L represents the intersection length of the beam and the voxel, g represents the attenuation degree, and MS represents the Mumford-Shah functional). The last two terms are independent of the forward projection and can be started simultaneously with the forward projection.
[0058] In the voxel-driven method, the processing of each voxel is independent. The voxel feature parameters (decay coefficient) f and (edge tensor) υ are updated directly in the voxel-driven method without affecting other voxels. The voxel-driven method does not require separating the back projection and update stages. Therefore, the synchronization operation before updating the decay coefficient f and edge tensor υ in the original general programming model is removed, eliminating the synchronization that was the most expensive part in previous work.
[0059] 3.3 After algorithm optimization, the data stream becomes Figure 4 .and Figure 2 In comparison, the main differences and advantages include:
[0060] The internal circulation becomes voxel-driven to improve locality.
[0061] The decay coefficient f and the edge tensor υ are updated directly during the calculation process.
[0062] The voxel-driven method eliminates the synchronization operation before updating the decay coefficient f and the edge tensor υ in the original general programming model.
[0063] Hardware resources are reused in two phases of the inner loop: the forward projection process and the computation of 2(∑fL-g)L share the multiplication logic; computation and Shared gradient logic.
[0064] Removing the outer loop makes the algorithm converge faster.
[0065] 4. The Waffle hardware accelerator is designed to enable efficient reconstruction algorithm execution.
[0066] 4.1 Data Structure and Layout.
[0067] 4.1.1 Allocate data structures for storing voxels and beams for each cube. The data structure for a voxel includes basic features (attenuation coefficient f, edge tensor υ) and beam information info passing through the voxel. b (unique ID of the beam) b Is it the last voxel of the nucleus? b The length L of the intersection point between the beam and the voxel b The beam's data structure includes a unique IID. b attenuation b The last core through which the beam of light passed. b The next core that the beam passes through b Has the previous core already passed part of the result to the current core (sync)? b Has the current core completed a partial forward projection of the beam? b sync b and finish b The main idea derived from synchronization is to process a beam of light passing through multiple cores, where each core receives a partial result from the previous core and performs an addition operation.
[0068] 4.1.2 The storage layout uses a serpentine traversal method for voxels, such as... Figure 6 As shown, this reduces the required size of the buffer.
[0069] 4.2 Data Flow and Logic Chip Design.
[0070] 4.2.1 The logic module is responsible for Figure 4 The voxel-driven internal loop and synchronization.
[0071] 4.2.2 The data flow is divided into two stages, and hardware resources can be reused.
[0072] The first stage involves forward projection and calculation. And update υ. An offset decoder and acquirer are needed based on the number of beams n passing through the voxel. b Decode the address offset of each voxel and obtain the corresponding data.
[0073] The Mumford-Shah functional solver calculates gradients according to the gradient calculation formula in section 1.2. And update υ. Each processor element PE calculates fL. b (f represents the attenuation coefficient, L) bThis represents the length of the intersection point between the beam and the voxel, and the relevant information is passed to the exclusive updater. Updates occurring in each step are independent and conflict-free. The exclusive updater updates the att value in the beam buffer without potential conflicts. b (att b (Indicates the attenuation level). If end b (end b If the last voxel is true, the PE will retrieve the relevant beam from the beam buffer and finish. b (finish b Change the value to true (indicating whether the current kernel has completed partial orthogonal projection of the beam) and check sync. b (sync b This indicates whether the previous core has passed part of the result to the current core (whether it is true). b and sync b A value of true for both means that the PE will notify the communicator to continue communication until the forward projection of the core has ended.
[0074] After receiving the PE notification, the communicator waits three cycles to achieve synchronization. (1) If it is the last nucleus the beam passes through, the communicator calculates att. b -g forward projection results and pass them to the communicator of the previous core. (2) If it is not the last core, the communicator passes att b The next core will then send the att message to the communicator in that core. b The data is passed to the exclusive updater, which updates the corresponding att in the beam buffer. b .
[0075] The second stage calculates 2(∑fL-g)L and To update f (where f represents the attenuation coefficient, L represents the intersection length of the beam and the voxel, g represents the attenuation degree, and MS represents the Mumford-Shah functional), the hardware resources from the first stage can be reused. The offset decoder and acquirer are used to traverse the voxels. The Mumford-Shah functional solver calculates... The exclusive updater is still responsible for implementing conflict-free updates to the voxel buffer.
[0076] PE can also be reused, used in the first stage for fL b Multiplication is performed in the second stage to calculate 2(∑fL-g)L. Both stages primarily involve multiplication, with inputs acquired via voxel-driven methods and outputs sent to the exclusive updater.
[0077] 4.2.3 The Mumford-Shah functional solver is used for computation. and Three rows of data are needed to calculate the second derivative, therefore, we use a method that... Figure 6 The method of storing and traversing voxels is shown to improve data reuse.
[0078] 4.2.4 The Exclusive Updater section is used to implement conflict-free updates of the beam buffer and voxel buffer. Content-Addressable Memory (CAM) is used to implement conflict-free updates of the buffers, guaranteeing the atomicity of the updates. When a new request arrives, if its address matches a request in the CAM, the value of that request is merged. If the Exclusive Updater issues a request, the relevant CAM entry is marked as "issued". New matching requests cannot be merged; instead, their values are written to the entry, and the entry's state changes from "issued" to "locked". Entries in the "locked" state still support merge requests but cannot issue requests. Only when the request ends does the entry return to the "normal" state.
[0079] The above embodiments are only used to illustrate the method of the present invention and are not intended to limit it. Those skilled in the art can modify or make equivalent substitutions to the method of the present invention without departing from the scope of the present invention. The scope of protection of the present invention should be determined by the claims.
Claims
1. A hardware and software co-design method for reconstructing the internal structure of three-dimensional imaging, characterized in that, Includes the following steps: 1) Introduce the Mumford-Shah functional to measure the attenuation coefficient and edge smoothness of the 3D image, optimize the attenuation coefficient and edge tensor of the 3D reconstruction result, and construct a 3D ray-driven iterative reconstruction model. 2) Establish a three-nested loop general programming model; the programming model includes three nested loops; the nested loops are an outer loop, a parallelization loop, and an inner loop, used for the segmentation and allocation of 3D reconstruction tasks, to achieve parallelization, and to process beams on each kernel; wherein, the outer loop divides the 3D reconstruction task into multiple parts according to the voxels or rays in the 3D reconstruction; the parallelization loop performs forward and backward projection on each kernel and synchronizes them; the inner loop processes the forward and backward projection of beams on a single kernel; 3) A voxel-driven method is adopted. By storing the attenuation coefficient, intersection length, and beam of voxels, and using a pre-calculation and simplified synchronization scheme, the synchronization operation before updating the attenuation coefficient and edge tensor of voxels in the three nested loop general programming model established in step 2) is removed. This optimizes the 3D reconstruction algorithm and obtains the optimized data dependency, i.e., data flow, reducing redundant calculations and improving data locality. This includes: 3a) Traversing voxels and finding all beams passing through them; for each voxel in the forward projection, storing the attenuation coefficient and all beams passing through it; pre-calculating the intersection length of the beam and the voxel and storing it in the voxel data structure; calculating the attenuation degree based on the stored attenuation coefficient and intersection length when traversing voxels, and updating the attenuation degree of the beam. 3b) Establish data dependencies; including: In reverse projection, items unrelated to forward projection are started simultaneously with forward projection; Voxel feature parameters are updated directly in the voxel-driven method without synchronization; voxel feature parameters include decay coefficients and edge tensors; the synchronization operation before updating decay coefficients and edge tensors in the original general programming model is removed. 4) Build a hardware accelerator; A near-memory computing architecture hardware accelerator is constructed, using logic chips and multiple memory chips to form a cubic structure and connected through silicon vias; including: offset decoder, acquirer, Mumford-Shah functional solver, communicator, exclusive updater and processor element PE; 4a) Construct the data structure and storage layout; including: Each cube stores information about voxels and beams. The voxel data structure includes attenuation coefficient, intersection length, and beam ID field. The beam data structure includes unique ID, attenuation, the last kernel the beam passed through, the next kernel the beam passed through, whether the previous kernel has passed some results to the current kernel, and whether the current kernel has completed a partial forward projection of the beam. The storage layout uses a serpentine approach to traverse voxels; 4b) Design the data flow and logic chip; including: performing forward projection and calculation in the first stage of the data flow, and using an exclusive updater to buffer, integrate and schedule the corresponding data; The second phase of the data stream involves back projection and update operations, using an exclusive updater to resolve write conflicts, and a communicator to achieve synchronization and communication. In the constructed near-memory computing architecture hardware accelerator, the Mumford-Shah functional solver is used to compute all expressions containing Mumford-Shah functionals: and Where MS represents the Mumford-Shah functional; f represents the decay coefficient; and υ represents the edge tensor. The exclusive updater section is used to implement conflict-free updates of the beam buffer and voxel buffer; Through the above steps, the internal structure reconstruction of 3D imaging can be achieved through hardware and software collaboration.
2. The hardware and software co-design method for three-dimensional imaging internal structure reconstruction as described in claim 1, characterized in that, Step 1) Construct a 3D ray-driven iterative 3D reconstruction model, including: 1.1 Obtain the attenuation coefficient of each pixel or voxel, i.e., the 3D reconstruction target; 1.2 The Mumford-Shah functional is introduced to measure the attenuation coefficient and edge smoothness of the 3D reconstruction results; 1.3 Acquire multiple beams from multiple directions to cover the reconstructed object.
3. The hardware and software co-design method for reconstructing the internal structure of three-dimensional imaging as described in claim 2, characterized in that, Multiple beams in multiple directions include parallel beams, fan-shaped beams, and conical beams.
4. The hardware and software co-design method for reconstructing the internal structure of three-dimensional imaging as described in claim 1, characterized in that step 4) constructing the hardware accelerator includes the following process: 4.1 Design the data structure and layout; including: 4.1.1 Allocate data structures for storing voxels and beams for each cube; 4.1.2 The storage layout uses a serpentine approach to traverse voxels, reducing the required buffer size; 4.2 Design the data flow and logic chips; including: 4.2.1 The logic module is used for voxel-driven internal loops and synchronization; 4.2.2 The data flow is divided into two stages for reusing hardware resources; The first phase includes: Perform forward projection and calculation And update the edge tensor υ; the offset decoder and acquirer are based on the number of beams n passing through the voxels. b Decode the address offset of each voxel and obtain the corresponding data; Mumford-Shah functional solver calculations And update υ; calculate fL for each processor element PE. b And pass the relevant information to the exclusive updater; where L b This represents the length of the intersection point between the beam and the voxel; the exclusive updater will update the attenuation at in the beam buffer without potential conflicts. b When the current voxel is the last voxel, the processor element PE obtains the relevant beam from the beam buffer, ends the forward projection of the core, and the communicator continues to communicate. The communicator synchronizes after waiting for three cycles, including: if it is the last nucleus the beam has passed through, the communicator calculates the forward projection result and transmits it to the communicator of the previous nucleus; if it is not the last nucleus, the communicator transmits the attenuation att. b The next core is given the information, and the communicator in the next core passes the information to the exclusive updater; the exclusive updater adds the corresponding att to the beam buffer. b middle; The second phase includes: Calculate and update the voxel decay coefficient, and reuse the hardware resources from the first stage; The offset decoder and acquirer are used to traverse voxels; the Mumford-Shah functional solver is used for computation. The exclusive updater is responsible for conflict-free updates of the voxel buffer; 4.2.3 Calculation using the Mumford-Shah functional solver and 4.2.4 The exclusive updater uses content-addressable memory to achieve conflict-free updates of the buffer and ensure the atomicity of the update; When a new request arrives, if the address matches a request in the Content Addressable Memory (CAM), the value of that request is integrated. If the exclusive updater issues a request, the relevant entry in the Content Addressable Memory (CAM) is marked as "issued"; new matching requests cannot request consolidation; its value is written to the entry, and the entry status changes from "issued" to "locked"; entries in the "locked" state still support consolidation requests, but cannot issue them; when the request ends, the entry returns to the "normal" state.
5. The hardware and software co-design method for three-dimensional imaging internal structure reconstruction as described in claim 4, characterized in that, In 4.1.1, the data structure of a voxel includes features such as attenuation coefficient, edge tensor, and beam information passing through the voxel.
6. The hardware and software co-design method for three-dimensional imaging internal structure reconstruction as described in claim 5, characterized in that, Information about a beam passing through a voxel includes: a unique beam ID, whether it is the last voxel of the nucleus, and the length of the intersection point between the beam and the voxel.
7. The hardware and software co-design method for three-dimensional imaging internal structure reconstruction as described in claim 4, characterized in that, In 4.1.1, the data structure of the beam includes a unique ID, attenuation, the last kernel the beam passed through, the next kernel the beam passed through, whether the previous kernel has passed some results to the current kernel, and whether the current kernel has completed a partial forward projection of the beam.