Acceleration calculation method for three-dimensional image processing

Through hardware acceleration technology, parallelization and pipeline processing are used to accelerate octree construction and traversal, solving the problems of low computing efficiency, large memory usage and insufficient real-time performance in three-dimensional image processing, and achieving efficient three-dimensional image processing, suitable for fields such as autonomous driving and virtual reality.

CN120107479APending Publication Date: 2025-06-06XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510184243.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When processing high-resolution and massive three-dimensional point cloud data, existing three-dimensional image processing technology faces the problems of low computing efficiency, large memory usage and insufficient real-time processing capabilities. Especially in areas such as autonomous driving and virtual reality, real-time requirements are high.

Method used

Using hardware acceleration means, octree construction and traversal are accelerated through parallelization and pipeline processing, storage management is optimized to reduce memory usage, and point cloud data processing with extremely high requirements for real-time performance. Specific measures include building a hardware architecture for accelerated octree algorithm, using acceleration units to process key steps, and improving processing efficiency through parallel computing and pipeline design.

Benefits of technology

It significantly improves the efficiency of three-dimensional image processing, reduces memory bandwidth consumption, and realizes point cloud data processing with extremely high requirements for real-time performance, which is suitable for autonomous driving, virtual reality and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107479A_ABST
    Figure CN120107479A_ABST
Patent Text Reader

Abstract

The invention discloses an acceleration calculation method for three-dimensional image processing. The method comprises the following steps: S1, constructing a hardware architecture of an acceleration octree algorithm; s2, an acceleration unit is specially used for processing key steps in the octree algorithm construction process, and computing-intensive tasks are migrated to hardware from the software level so as to improve the processing efficiency and reduce the overall computing complexity of the system; s3, data of multiple dimensions are processed at the same time through parallel calculation, and the calculation time is shortened; s4, tasks in each processing stage are executed in parallel through assembly line design, so that the processing efficiency of the system is improved; according to the method, a hardware acceleration means is adopted, octree construction and traversal are accelerated through parallelization and assembly line processing, storage management is optimized to reduce memory occupation, point cloud data processing with extremely high real-time performance requirements is achieved, and therefore wide application prospects are shown in the fields of automatic driving, virtual reality and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of three-dimensional image processing, and in particular relates to an accelerated calculation method for three-dimensional image processing. Background Art

[0002] 3D image processing is a technology that generates 3D spatial information from 2D image data. It is widely used in fields such as autonomous driving, robot navigation, virtual reality, and 3D modeling. It extracts depth information in space and expands pixels on a 2D plane into points or surfaces in 3D space, providing machines with stereoscopic perception of the real world. This technology can directly measure depth through active sensors (such as lidar, infrared light, ultrasound, etc.), or use passive sensors (such as binocular stereo vision, optical flow method) to calculate the 3D structure of the scene from images from multiple perspectives. The depth images generated by these methods are the basis for 3D image generation. The depth information can be further used to derive point cloud data containing 3D coordinates (x, y, z) to form an accurate geometric description of the scene.

[0003] However, with the improvement of sensor accuracy and resolution, the resolution of the generated depth images is getting higher and higher, and the amount of 3D point cloud data is also growing exponentially. After a high-resolution depth image is converted by the camera's intrinsic parameters, it will generate a 3D point cloud with millions or even hundreds of millions of points. These point clouds not only take up a lot of storage space, but also place extremely high demands on data transmission bandwidth and computing processing capabilities. For example, a single frame of high-resolution point cloud may be hindered by bandwidth limitations when it is transmitted to edge computing devices or the cloud for processing. At the same time, the massive amount of point cloud data will also greatly reduce the processing efficiency of subsequent algorithms (such as point cloud segmentation, registration, object recognition, etc.). Therefore, when acquiring, storing and processing these 3D point clouds, efficient optimization and compression methods need to be introduced.

[0004] In addition, the processing of 3D images faces other technical challenges, such as how to effectively compress point clouds to reduce storage requirements without losing key geometric information, how to dynamically adjust the calculation accuracy to adapt to the performance limitations of different hardware, and how to use hardware acceleration (such as GPU, FPGA, etc.) to process massive point cloud data in real time. Especially in scenarios with high real-time requirements, such as LIDAR point clouds used to perceive obstacles in autonomous driving or depth data used to determine the shape of objects in robot grasping tasks, how to quickly and accurately process 3D images becomes the key to the technology.

[0005] Octrees recursively divide the three-dimensional space into eight equal-volume subspaces. They can efficiently organize and store sparse point cloud data and support fast spatial query, storage optimization, and hierarchical processing. However, in traditional computers and embedded systems, there are many challenges in implementing point cloud processing based on octrees. The implementation principle of octrees can be divided into the following steps: (1) Setting the maximum tree depth and root node cube size; (2) Creating the first cube from the space as the root node; (3) Inserting the point cloud data containing x, y, and z coordinates into the octree in sequence, and finding the corresponding child node according to the spatial position index; (4) If the maximum tree depth is not reached, divide it into 8 sub-cubes as child nodes until the recursive depth is equal to the maximum tree depth; (5) Repeat step (3) until all point clouds are inserted into the octree. From this process, it can be seen that, first, the computational complexity of constructing an octree is high, and it usually requires multiple spatial judgments for each point. As the scale of the point cloud increases, the processing efficiency is heavily dependent on hardware performance. At the same time, the storage requirements of the octree are large, and the number of nodes increases exponentially with the layer depth, which places high demands on memory resources. In addition, the single-thread efficiency of traditional CPUs is limited. Even with multi-thread optimization, bottlenecks may still be encountered when processing high-density point cloud data. These problems are more prominent in embedded systems. Embedded devices usually have limited resources, and the processors are mostly ARM architectures with small computing power and memory capacity, but they need to meet real-time requirements. For example, the processing of LIDAR point clouds in autonomous driving must be completed in a very short time. In order to achieve efficient octree point cloud processing in a resource-constrained environment, common optimization strategies include using fixed-point calculations to reduce floating-point operation overhead and hierarchical storage to accelerate access. However, even so, traditional octree implementations still find it difficult to meet the real-time processing requirements of ultra-large-scale point clouds. Its bottlenecks are mainly concentrated on the high memory overhead and time-consuming splitting operations when dynamically splitting child nodes. Summary of the invention

[0006] To solve the above problems, the present invention proposes an accelerated computing method for three-dimensional image processing. The method adopts hardware acceleration to accelerate octree construction and traversal through parallelization and pipeline processing, optimizes storage management to reduce memory usage, and realizes point cloud data processing with extremely high real-time requirements, thereby showing broad application prospects in the fields of autonomous driving, virtual reality, etc.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] An accelerated computing method for three-dimensional image processing comprises the following steps:

[0009] S1. Build a hardware architecture to accelerate the octree algorithm;

[0010] S2, using acceleration units specifically for processing key steps in the octree algorithm construction process, migrating computationally intensive tasks from the software level to the hardware, to improve processing efficiency and reduce the overall computational complexity of the system;

[0011] S3, through parallel computing, processing data of multiple dimensions at the same time, to shorten the computing time;

[0012] S4. Through pipeline design, tasks in each processing stage are executed in parallel to improve the processing efficiency of the system.

[0013] Preferably, the hardware architecture in step S1 includes a main controller, an on-chip memory, a DMA controller, a peripheral controller, an AXI system bus, multiple cascaded processing units, a configuration data memory, multiple arrays of shift register units, an off-chip memory and an off-chip memory controller; the main controller, on-chip memory, DMA controller and peripheral controller are respectively connected to the AXI system bus; the processing unit is respectively connected to the AXI system bus interface and the shift register unit; the shift register unit is connected to the AXI system bus interface through the configuration data memory, and the output data of the shift register unit is returned to the AXI system bus interface; the off-chip memory is connected to the AXI system bus through the off-chip memory controller.

[0014] Preferably, in step S1, each of the processing units completes one layer of traversal of the octree and is connected in a cascade manner to achieve real-time traversal of the entire octree; wherein the cascade structure is used to improve the computing power of the hardware and enable the system to dynamically adjust the processing power as needed; through the AXI bus control module, the internal registers are configured to achieve cascade structures of different lengths, which are used to adapt the hardware architecture to task requirements of different scales and complexities.

[0015] Preferably, in step S1, the calculation results of each level of processing units are cached by shift registers in the form of an array, so that the data can be efficiently transmitted and output synchronously, and the processed point cloud data will be stored as an index sequence and stored in the system memory; wherein, the index sequence of the point cloud contains the spatial position of each point and its corresponding leaf node path information, which is used when constructing an octree through software. Since the index sequence already contains the position and hierarchical information of each point in the point cloud, the main controller performs simple path search and tree structure update to reduce the calculation complexity.

[0016] Preferably, the acceleration unit in step S2 includes unit A and unit B;

[0017] The unit A is a subnode space size update module, which sets the root node as a cube, and the divided subnodes are all cubes, and the side length of the subnode is half of the side length of the parent node; the size signal input by the subnode space size update module describes the cube space size with the side length, and implements the space size update of new_size=size*0.5 through a floating-point multiplier;

[0018] The unit B is a subnode index update module, which realizes the synchronous update of three-dimensional coordinates through parallel design; the input of the subnode index update module is the point cloud coordinates (point_x, pont_y, point_z), the index value (comp_x, comp_y, comp_z) and the center point coordinates (center_x, center_y, center_z) of the cube corresponding to the current node; among them, new_point_x = point_x, new_point_y = point_y, new_point_z = point_z, indicating that the three-dimensional point coordinates remain unchanged and are passed to the next level; the formula for updating the index value is:

[0019]

[0020] That is, when the point cloud coordinates are greater than the cube center coordinates, the corresponding sub-cube is located on the positive semi-axis and is encoded as 1. Otherwise, the corresponding sub-cube is located on the negative semi-axis and is encoded as 0. The xyz three-dimensional coordinates are concatenated into 3-bit values ​​as the indexes of the eight child nodes, and the center coordinates are finally updated as follows:

[0021] new_center_x=center_x+f(comp_x)*new_center,where,

[0022]

[0023] new_center_y=center_y+f(comp_y)*new_center, where,

[0024]

[0025] new_center_z=center-_+f(comp_z)*new_center, where,

[0026]

[0027] It means that if the 3D point is located on the positive semi-axis with the center of the current cube as the midpoint, the corresponding sub-cube center coordinate is translated by new_size distance toward the positive semi-axis direction, otherwise, it is translated by new_size distance toward the negative semi-axis direction; all floating-point calculations are performed using DSP resources, where comparison operations are implemented using numerical comparators, conditional functions are implemented through multiplexers, and pipeline operations between calculation modules are implemented by inserting registers.

[0028] Preferably, in step S3, multiple cascaded processing units are used to accelerate the recursive operation of the octree algorithm, by decomposing the tasks in the recursive body into multiple parallel tasks, using multiple processing units to process them in parallel at the same time, and then merging the calculation results of the multiple processing units to obtain the final result; wherein, the traversal operation of each layer of the octree is implemented by a processing unit, the position of the node of the corresponding layer is obtained by the processing unit, and the position output of the node of the corresponding layer is used as the input of the next level, the position of the node of each layer is obtained in turn, and finally the position of the node of each layer is combined to obtain a path information from the root node to the leaf node; the main controller uses the path information to construct the structure of the octree to reduce the amount of calculation.

[0029] Preferably, each level of processing unit in step S3 is implemented by an octree algorithm accelerator, the input point of the first level is the point cloud coordinate, and the size is the initial cube size, that is, the center of the first level is the center point coordinate of the cube corresponding to the virtual parent node, then the initial value of (comp_x, comp_y, comp_z) is (0, 0, 0), and each level outputs (comp_x, comp_y, comp_z) as a position index output to the corresponding Buffer buffer for cache, and finally all Buffer buffers are output synchronously and the octree index sequence corresponding to the point cloud is obtained through bit splicing.

[0030] Preferably, in step S4, the fixed delay characteristic of each processing unit is used to define the data delay period T PE , using the corresponding depth D = T PE / T clkThe shift register cascade realizes data delay control; the shift register unit consists of a shift register of fixed depth and an enable control register; all shift register units are connected in array form, and the array controller configures their interconnection network structure. When the tree depth is set to a maximum of 9 layers, a total of 9 processing units are required to be cascaded, then the first-level processing unit requires 8 shift register units to be cascaded, the second-level processing unit requires 7 shift register units to be cascaded, and so on. The eighth-level processing unit requires 1 shift register unit, and the ninth-level processing unit directly outputs without the need for a shift register unit, requiring a total of 36 shift register units; by configuring the register, the number of enabled processing units is selected to achieve acceleration of octree algorithms of different depths, and at the same time, the shift register unit array controller is modified to reduce the number of shift register units corresponding to each processing unit, and the unused processing units and shift register units are turned off by the enable controller to reduce system power consumption; an address space is allocated for all configurable registers, and the main controller performs dynamic configuration through the AXI system bus.

[0031] Preferably, in step S4, since the three-dimensional image data after hardware acceleration is no longer in the conventional three-dimensional coordinate form, the octree algorithm is modified accordingly at the software level to adapt to the new data representation method. The specific modification process is:

[0032] S41, initialize an empty octree, and prepare a structure for storing nodes, the structure contains relevant information of the nodes, and the relevant information includes a position index, a parent node pointer, and a child node pointer;

[0033] S42, read the index sequence from the memory in sequence, divide the index sequence into n 3-bit values, and store them in an array of length n; loop through the entire array, and the value of each array is the child node position index;

[0034] S43, starting from the root node of the octree, determine the position where each element of the array should be inserted according to its value; if the current node is empty or has not yet been assigned a child node, directly insert the node; if the current node already has a child node, skip it and continue to go down along the hierarchy of the octree until the last level of the octree is reached and a leaf node is created; until each node of the octree is gradually filled, forming a complete link from the root node to the leaf node;

[0035] S44. After traversing the index sequences corresponding to all three-dimensional images in the memory, a complete octree structure is finally obtained.

[0036] After adopting the above technical solution, the present invention has the following beneficial effects: the present invention adopts hardware acceleration means, accelerates octree construction and traversal through parallelization and pipeline processing, optimizes storage management to reduce memory usage, and realizes point cloud data processing with extremely high real-time requirements, thereby showing broad application prospects in the fields of autonomous driving, virtual reality, etc. This processing method based on hardware acceleration and octree structure not only improves the efficiency of data processing and reduces memory bandwidth consumption, but also provides a more efficient data access method for subsequent three-dimensional image algorithms such as collision detection, object recognition, scene reconstruction, etc. More importantly, due to the adaptability of the octree structure, it can automatically optimize the storage and processing strategies according to the spatial distribution of the data, thereby greatly improving the processing performance and system scalability in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a flow chart of the present invention;

[0038] Figure 2 It is a structural schematic diagram of the hardware architecture of the present invention;

[0039] Figure 3 It is a structural schematic diagram of the acceleration unit of the present invention;

[0040] Figure 4 It is a schematic diagram of the cascade structure of multiple cascaded processing units of the present invention. DETAILED DESCRIPTION

[0041] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0042] like Figure 1 and Figure 4 As shown, a three-dimensional image processing acceleration calculation method includes the following steps:

[0043] S1. Build a hardware architecture to accelerate the octree algorithm;

[0044] The hardware architecture in step S1 includes a main controller (CPU), an on-chip memory, a DMA controller, a peripheral controller, an AXI system bus, a plurality of cascaded processing units (PEs), a configuration data memory, a plurality of arrays of shift register units (SRs), an off-chip memory, and an off-chip memory controller; the main controller, the on-chip memory, the DMA controller, and the peripheral controller are respectively connected to the AXI system bus; the processing unit is respectively connected to the AXI system bus interface and the shift register unit; the shift register unit is connected to the AXI system bus interface through the configuration data memory, and the output data of the shift register unit is returned to the AXI system bus interface; the off-chip memory is connected to the AXI system bus through the off-chip memory controller;

[0045] In step S1, each of the processing units completes a layer of traversal of the octree, and is connected in a cascade manner to achieve real-time traversal of the entire octree; wherein the cascade structure is used to improve the computing power of the hardware and enable the system to dynamically adjust the processing power as needed; through the AXI bus control module, the internal registers are configured to achieve cascade structures of different lengths, so as to adapt the hardware architecture to task requirements of different scales and complexities;

[0046] In step S1, the calculation results of each level of processing unit are cached by shift registers in the form of an array, so that the data can be efficiently transmitted and synchronously output. The processed point cloud data will be stored as an index sequence and stored in the system memory; wherein, the index sequence of the point cloud contains the spatial position of each point and its corresponding leaf node path information, which is used when the octree is constructed by software. Since the index sequence already contains the position and hierarchical information of each point in the point cloud, the main controller performs simple path search and tree structure update to reduce the calculation complexity;

[0047] S2, using acceleration units specifically for processing key steps in the octree algorithm construction process, migrating computationally intensive tasks from the software level to the hardware, to improve processing efficiency and reduce the overall computational complexity of the system;

[0048] In step S2, the acceleration unit includes unit A and unit B;

[0049] The unit A is a subnode space size update module, which sets the root node as a cube, and the divided subnodes are all cubes, and the side length of the subnode is half of the side length of the parent node; the size signal input by the subnode space size update module describes the cube space size with the side length, and implements the space size update of new_size=size*0.5 through a floating-point multiplier;

[0050] The unit B is a subnode index update module, which realizes the synchronous update of three-dimensional coordinates through parallel design; the input of the subnode index update module is the point cloud coordinates (point_x, pont_y, point_z), the index value (comp_x, comp_y, comp_z) and the center point coordinates (center_x, center_y, center_z) of the cube corresponding to the current node; among them, new_point_x = point_x, new_point_y = point_y, new_point_z = point_z, indicating that the three-dimensional point coordinates remain unchanged and are passed to the next level; the formula for updating the index value is:

[0051]

[0052]

[0053] That is, when the point cloud coordinates are greater than the cube center coordinates, the corresponding sub-cube is located on the positive semi-axis and is encoded as 1. Otherwise, the corresponding sub-cube is located on the negative semi-axis and is encoded as 0. The xyz three-dimensional coordinates are concatenated into 3-bit values ​​as the indexes of the eight child nodes, and the center coordinates are finally updated as follows:

[0054] new_center_x=center_x+f(comp_x)*new_center,where,

[0055]

[0056] new_center_y=center_y+f(comp_y)*new_center, where,

[0057]

[0058] new_center_z=center_z+f(comp_z)*new_center, where,

[0059]

[0060] Indicates that if the 3D point is located on the positive semi-axis with the center of the current cube as the midpoint, the corresponding sub-cube center coordinate is translated to the positive semi-axis direction by a distance of new_size, otherwise, it is translated to the negative semi-axis direction by a distance of new_size; all floating-point calculations are performed using DSP resources, where comparison operations are implemented using numerical comparators, conditional functions are implemented through multiplexers, and pipeline operations between calculation modules are implemented by inserting registers;

[0061] S3, through parallel computing, processing data of multiple dimensions at the same time, to shorten the computing time;

[0062] In step S3, multiple cascaded processing units are used to accelerate the recursive operation of the octree algorithm, by decomposing the tasks in the recursive body into multiple parallel tasks, using multiple processing units to process them in parallel at the same time, and then merging the calculation results of the multiple processing units to obtain the final result; wherein, the traversal operation of each layer of the octree is implemented by a processing unit, the position of the node of the corresponding layer is obtained by the processing of the processing unit, and the position output of the node of the corresponding layer is used as the input of the next level, the position of the node of each layer is obtained in turn, and finally the position of the node of each layer is combined to obtain a path information from the root node to the leaf node; the main controller uses the path information to construct the structure of the octree to reduce the amount of calculation;

[0063] In step S3, each level of processing unit is implemented by an octree algorithm accelerator. The input point of the first level is the point cloud coordinate, and the size is the initial cube size, that is, the center of the first level is the center point coordinate of the cube corresponding to the virtual parent node, and the initial value of (comp_x, comp_y, comp_z) is (0, 0, 0). The output (comp_x, comp_y, comp_z) of each level is output as a position index to the corresponding Buffer buffer for cache. Finally, all Buffer buffers are output synchronously and the octree index sequence corresponding to the point cloud is obtained through bit splicing.

[0064] S4. Through pipeline design, tasks in each processing stage are executed in parallel to improve the processing efficiency of the system;

[0065] In step S4, the fixed delay characteristic of each processing unit is used to define the data delay period T PE , using the corresponding depth D = T PE / T clkThe shift register cascade realizes data delay control; the shift register unit consists of a shift register of fixed depth and an enable control register; all shift register units are connected in array form, and the array controller configures its interconnection network structure. When the tree depth is set to a maximum of 9 layers, a total of 9 processing units are required to be cascaded, then the first-level processing unit requires 8 shift register units to be cascaded, and the second-level processing unit requires 7 shift register units to be cascaded. In sequence, the eighth-level processing unit requires 1 shift register unit, and the ninth-level processing unit directly outputs without the need for a shift register unit, requiring a total of 36 shift register units; through the configuration register, the number of enabled processing units is selected to achieve acceleration of the octree algorithm of different depths, and at the same time, the shift register unit array controller is modified to reduce the number of shift register units corresponding to each processing unit, and the unused processing units and shift register units are turned off by the enable controller to reduce system power consumption; an address space is allocated for all configurable registers, and the main controller is dynamically configured through the AXI system bus;

[0066] In step S4, since the three-dimensional image data after hardware acceleration is no longer in the conventional three-dimensional coordinate form, the octree algorithm is modified accordingly at the software level to adapt to the new data representation method. The specific modification process is as follows:

[0067] S41, initialize an empty octree, and prepare a structure for storing nodes, the structure contains relevant information of the nodes, and the relevant information includes a position index, a parent node pointer, and a child node pointer;

[0068] S42, read the index sequence from the memory in sequence, divide the index sequence into n 3-bit values, and store them in an array of length n; loop through the entire array, and the value of each array is the child node position index;

[0069] S43, starting from the root node of the octree, determine the position where each element of the array should be inserted according to its value; if the current node is empty or has not yet been assigned a child node, directly insert the node; if the current node already has a child node, skip it and continue to go down along the hierarchy of the octree until the last level of the octree is reached and a leaf node is created; until each node of the octree is gradually filled, forming a complete link from the root node to the leaf node;

[0070] S44. After traversing the index sequences corresponding to all three-dimensional images in the memory, a complete octree structure is finally obtained.

[0071] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. An accelerated computing method for three-dimensional image processing, characterized in that: The following steps are involved: S1. Build a hardware architecture to accelerate the octree algorithm; S2, using acceleration units specifically for processing key steps in the octree algorithm construction process, migrating computationally intensive tasks from the software level to the hardware, to improve processing efficiency and reduce the overall computational complexity of the system; S3, through parallel computing, processing data of multiple dimensions at the same time, to shorten the computing time; S4. Through pipeline design, tasks in each processing stage are executed in parallel to improve the processing efficiency of the system.

2. The accelerated computing method for three-dimensional image processing according to claim 1, characterized in that: The hardware architecture described in step S1 includes a main controller, an on-chip memory, a DMA controller, a peripheral controller, an AXI system bus, multiple cascaded processing units, a configuration data memory, multiple arrays of shift register units, an off-chip memory and an off-chip memory controller; the main controller, on-chip memory, DMA controller, and peripheral controller are respectively connected to the AXI system bus; the processing unit is respectively connected to the AXI system bus interface and the shift register unit; the shift register unit is connected to the AXI system bus interface through the configuration data memory, and the output data of the shift register unit is returned to the AXI system bus interface; the off-chip memory is connected to the AXI system bus through the off-chip memory controller.

3. The accelerated computing method for three-dimensional image processing according to claim 2, characterized in that: In step S1, each of the processing units completes a layer of traversal of the octree and is connected in a cascade manner to achieve real-time traversal of the entire octree; wherein the cascade structure is used to improve the computing power of the hardware and enable the system to dynamically adjust the processing power as needed; through the AXI bus control module, the internal registers are configured to achieve cascade structures of different lengths, which are used to adapt the hardware architecture to task requirements of different scales and complexities.

4. The accelerated computing method for three-dimensional image processing according to claim 2, characterized in that: In step S1, the calculation results of each level of processing unit are cached by shift registers in the form of an array, so that the data can be efficiently transmitted and output synchronously. The processed point cloud data will be stored as an index sequence and stored in the system memory; wherein, the index sequence of the point cloud contains the spatial position of each point and its corresponding leaf node path information, which is used when constructing an octree through software. Since the index sequence already contains the position and hierarchical information of each point in the point cloud, the main controller performs simple path search and tree structure update to reduce the calculation complexity.

5. The accelerated computing method for three-dimensional image processing according to claim 2, characterized in that: In step S2, the acceleration unit includes unit A and unit B; The unit A is a subnode space size update module, which sets the root node as a cube, and the divided subnodes are all cubes, and the side length of the subnode is half of the side length of the parent node; the size signal input by the subnode space size update module describes the cube space size with the side length, and implements the space size update of new_size=size*0.5 through a floating-point multiplier; The unit B is a subnode index update module, which realizes the synchronous update of three-dimensional coordinates through parallel design; the input of the subnode index update module is the point cloud coordinates (point_x, pont_y, point_z), the index value (comp_x, comp_y, comp_z) and the center point coordinates (center_x, center_y, center_z) of the cube corresponding to the current node; among them, new_point_x = point_x, new_point_y = point_y, new_point_z = point_z, indicating that the three-dimensional point coordinates remain unchanged and are passed to the next level; the formula for updating the index value is: That is, when the point cloud coordinates are greater than the cube center coordinates, the corresponding sub-cube is located on the positive semi-axis and is encoded as 1. Otherwise, the corresponding sub-cube is located on the negative semi-axis and is encoded as 0. The xyz three-dimensional coordinates are concatenated into 3-bit values ​​as the indexes of the eight child nodes, and the center coordinates are finally updated as follows: new_center_x=center_x+f(comp_x)*new_center,where, new_center_y=center_y+f(comp_y)*new_center, where, new_center_z=center_z+f(comp_z)*new_center, where, It means that if the 3D point is located on the positive semi-axis with the center of the current cube as the midpoint, the corresponding sub-cube center coordinate is translated by new_size distance toward the positive semi-axis direction, otherwise, it is translated by new_size distance toward the negative semi-axis direction; all floating-point calculations are performed using DSP resources, where comparison operations are implemented using numerical comparators, conditional functions are implemented through multiplexers, and pipeline operations between calculation modules are implemented by inserting registers.

6. The accelerated computing method for three-dimensional image processing according to claim 2, characterized in that: In step S3, multiple cascaded processing units are used to accelerate the recursive operation of the octree algorithm. The tasks in the recursive body are decomposed into multiple parallel tasks, and multiple processing units are used to process them in parallel at the same time. Then, the calculation results of the multiple processing units are merged to obtain the final result. Among them, the traversal operation of each layer of the octree is implemented by a processing unit, and the position of the node of the corresponding layer is obtained by the processing of the processing unit, and the position of the node of the corresponding layer is output as the input of the next level, and the position of each layer of nodes is obtained in turn. Finally, the position of each layer of nodes is combined to obtain a path information from the root node to the leaf node; the main controller uses the path information to construct the structure of the octree to reduce the amount of calculation.

7. The accelerated computing method for three-dimensional image processing according to claim 6, characterized in that: In step S3, each level of processing unit is implemented by an octree algorithm accelerator. The input point of the first level is the point cloud coordinate, and the size is the initial cube size, that is, the center of the first level is the center point coordinate of the cube corresponding to the virtual parent node, then the initial value of (comp_x, comp_y, comp_z) is (0, 0, 0), and each level outputs (comp_x, comp_y, comp_z) as a position index output to the corresponding Buffer buffer for cache. Finally, all Buffer buffers are output synchronously and the octree index sequence corresponding to the point cloud is obtained through bit splicing.

8. The accelerated computing method for three-dimensional image processing according to claim 2, characterized in that: In step S4, the fixed delay characteristic of each processing unit is used to define the data delay period T PE , using the corresponding depth D = T PE / T clk The shift register cascade realizes data delay control; the shift register unit consists of a shift register of fixed depth and an enable control register; all shift register units are connected in array form, and the array controller configures their interconnection network structure. When the tree depth is set to a maximum of 9 layers, a total of 9 processing units are required to be cascaded, then the first-level processing unit requires 8 shift register units to be cascaded, the second-level processing unit requires 7 shift register units to be cascaded, and so on. The eighth-level processing unit requires 1 shift register unit, and the ninth-level processing unit directly outputs without the need for a shift register unit, requiring a total of 36 shift register units; by configuring the register, the number of enabled processing units is selected to achieve acceleration of octree algorithms of different depths, and at the same time, the shift register unit array controller is modified to reduce the number of shift register units corresponding to each processing unit, and the unused processing units and shift register units are turned off by the enable controller to reduce system power consumption; an address space is allocated for all configurable registers, and the main controller performs dynamic configuration through the AXI system bus.

9. The accelerated computing method for three-dimensional image processing according to claim 8, characterized in that: In step S4, since the three-dimensional image data after hardware acceleration is no longer in the conventional three-dimensional coordinate form, the octree algorithm is modified accordingly at the software level to adapt to the new data representation method. The specific modification process is as follows: S41, initialize an empty octree, and prepare a structure for storing nodes, the structure contains relevant information of the nodes, and the relevant information includes a position index, a parent node pointer, and a child node pointer; S42, read the index sequence from the memory in sequence, divide the index sequence into n 3-bit values, and store them in an array of length n; loop through the entire array, and the value of each array is the child node position index; S43, starting from the root node of the octree, determine the position where each element of the array should be inserted according to its value; if the current node is empty or has not yet been assigned a child node, directly insert the node; if the current node already has a child node, skip it and continue to go down along the hierarchy of the octree until the last level of the octree is reached and a leaf node is created; until each node of the octree is gradually filled, forming a complete link from the root node to the leaf node; S44. After traversing the index sequences corresponding to all three-dimensional images in the memory, a complete octree structure is finally obtained.