A Multi-GPU Parallel Pyramid Construction Method under CUDA Mode

By using a multi-GPU parallel architecture to build pyramids in CUDA mode, the problem of inefficient processing of traditional CPUs is solved, and the rapid pyramid construction and real-time display of super-large image data is realized.

CN120031706BActive Publication Date: 2025-07-01JILIN GAOFEN REMOTE SENSING APPL RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510508188.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-01
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Traditional CPU processing methods lack processing and storage efficiency in ultra-large resolution images, resulting in slow and inefficient creation of image pyramids, affecting the efficiency of data processing and analysis.

Method used

The multi-GPU parallel pyramid construction method is adopted in the CUDA mode, and the powerful parallel computing power of the GPU is used to compress images layer by layer through convolutional downsampling operations, generate multi-resolution pyramid images, and real-time communication and collaborative control between GPUs are realized through NVIDIA NVLink technology.

Benefits of technology

The pyramid construction speed of super-large image data is significantly improved, and the processing speed is 11.95 times that of traditional CPU methods, which improves the utilization rate of computing resources, and ensures the real-time display capability of pyramid images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031706B_ABST
    Figure CN120031706B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of space remote sensing technology, and specifically provides a method for constructing a multi-GPU parallel pyramid in CUDA mode. The original input image is pre-blocked at multiple levels according to a predefined block size. Under the CUDA parallel computing architecture, the NVIDIA NVLink technology is used to coordinate the multi-GPU parallel processing of block tasks, dynamically allocate loads and synchronize data. Using the CUDA parallel computing architecture, a convolution downsampling operation is performed on each block to obtain the target resolution image corresponding to each level of the pyramid image. The images output by each GPU are geospatially aligned to generate a multi-resolution pyramid file conforming to the OVR format. The present invention first proposes to use a multi-GPU parallel architecture in CUDA mode for pyramid construction, and utilizes the powerful graphics parallel computing ability of the GPU to accelerate the pyramid construction process of ultra-large image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of space remote sensing, and particularly relates to a method for constructing a multi-GPU parallel pyramid in CUDA mode. Background Art

[0002] The pyramid model is the key to the visualization of massive remote sensing images in virtual earth. It is a multi-resolution hierarchical structure. On the premise of ensuring the display accuracy, data with different resolutions are adopted for different regions, thereby improving the rendering efficiency. The process of constructing the pyramid is essentially to perform block and layer processing on the image. Although this will increase the overall data volume, it can significantly shorten the time required for rendering. The pyramid model has been widely used in two-dimensional map visualization, and the rendering speed can be effectively accelerated by resampling the original data.

[0003] However, with the development of remote sensing technology and geographic information systems, the scale of image data has been continuously expanding, and the processing and storage efficiency of traditional CPU processing methods for ultra-high-resolution images is still insufficient. Especially in traditional GIS software, when creating an image pyramid, the processing speed is slow and the efficiency is low, which seriously affects the efficiency of data processing and analysis.

[0004] Therefore, new technical means are needed to solve the problem of too low efficiency in creating a pyramid for large images. Summary of the Invention

[0005] In view of this, the present invention aims to provide a method for constructing a multi-GPU parallel pyramid in CUDA mode, which utilizes the powerful graphics parallel computing ability of the GPU to accelerate the process of constructing the pyramid of ultra-large image data. After creation, it can be recognized by mainstream geographic information and remote sensing software, making the display of ultra-large images real-time.

[0006] To achieve the above object, the technical solution of the present invention is realized as follows:

[0007] The present invention provides a method for constructing a multi-GPU parallel pyramid in CUDA mode, including: presetting the image block size, calculating the number of levels of the pyramid image to be generated according to the block size, and pre-blocking the target resolution images of all levels of the pyramid image according to the block size;

[0008] Under the CUDA parallel computing architecture, the blocks of the target resolution images of each level are assigned to multiple GPUs, and all GPUs are controlled by NVIDIA NVLink technology to perform parallel convolution downsampling operations on the blocks layer by layer to obtain the target resolution images corresponding to each level of the pyramid image; the calculation formula for a single convolution downsampling operation is:

[0009] ;

[0010] Among them, represents the target resolution image of the current level and the convolution kernel at the coordinate The convolution result at this point, represents the radius of the convolution kernel, represents the convolution kernel at the coordinate The value at this point, represents the target resolution image of the current level at the coordinate The pixel value at this point; represents the target resolution image of the next level after downsampling, represents the sampling scaling factor, that is, sampling once every pixels;

[0011] For each level of the target resolution image after GPU downsampling, perform geospatial alignment to generate a multi-resolution pyramid image.

[0012] Preferably, the image block size is , among which, is preset according to the video memory size of the GPU and the number of computing units of the GPU.

[0013] Preferably, the calculation formula for the number of levels of the pyramid image to be generated is:

[0014] ;

[0015] Among them, represents the number of levels of the pyramid image to be generated (excluding the initial input image), that is, the number of times of convolutional downsampling, respectively represent the width and height of the initial input image.

[0016] Preferably, use the CPU to block the target resolution image of each level.

[0017] Preferably, the number of blocks of the target resolution image of the nth level is:

[0018] .

[0019] Preferably, the value of the sampling scaling factor is 2.

[0020] Preferably, the CPU uses a dynamic load balancing algorithm to distribute the blocks of the target resolution image of each level to multiple GPUs.

[0021] Preferably, determine the longitude and latitude information of each block through affine transformation.

[0022] Preferably, it further includes: generating a binary general OVR file for target resolution images of different levels, where the OVR file includes a file header and a data body.

[0023] Preferably, the file header includes the width, height, longitude and latitude range, block size, number of levels of the pyramid image, and affine transformation data of the target resolution image of each level of the initial input image; the data body includes the initial input image and the pixel values of the target resolution images of each level in the pyramid image.

[0024] Compared with the prior art, the present invention can achieve the following beneficial effects:

[0025] The present invention first proposes to use a multi-GPU parallel architecture in the CUDA mode for pyramid construction. Compared with the traditional CPU processing method, it has obvious advantages in the processing of ultra-large image data. When the image to be processed is 124581×158742, the image processing speed of the method of the present invention is 11.95 times that of the traditional CPU method.

[0026] The present invention pre-chunks the target resolution images of all levels of the pyramid images to be generated, which can effectively avoid memory overflow and lay a foundation for subsequent parallel computing. And the NVIDIA NVLink technology is used to achieve real-time communication and cooperative control between GPUs, allowing multiple GPUs to simultaneously and parallelly process different image chunks. Each GPU independently executes the convolution downsampling task, greatly improving the processing speed. And a combined calculation formula of convolution and downsampling is proposed, which is convenient for parallelization through CUDA programming.

[0027] The chunking, multi-GPU task allocation, and geospatial alignment operations of the present invention are all implemented by the CPU, reducing the computing pressure through the combined application of the CPU and GPU. In addition, in order to avoid overloading of some GPUs, the CPU dynamically adjusts the workload of each GPU through a load balancing mechanism, effectively improving the utilization rate of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0029] Figure 1 is a schematic diagram of a multi-GPU parallel pyramid construction method in the CUDA mode according to an embodiment of the present invention;

[0030] Figure 2 is a data structure diagram of a pyramid file according to an embodiment of the present invention;

[0031] Figure 3 It is a statistical chart of the time taken to construct pyramids from images of different sizes by the traditional method and the method of the present invention according to an embodiment of the present invention. Detailed implementation manners

[0032] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention. Similar elements in different embodiments are labeled with related similar element numbers. In the following embodiments, many detailed descriptions are provided to enable a better understanding of the present invention. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other elements, materials, or methods. In some cases, some operations related to the present invention are not shown or described in the specification to avoid obscuring the core part of the present invention with excessive descriptions. For those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations based on the descriptions in the specification and general technical knowledge in the art.

[0033] It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other to form various implementation manners. At the same time, the steps or actions in the method description can also be adjusted in the order that is obvious to those skilled in the art. Therefore, the various orders in the specification and drawings are only for clearly describing a certain embodiment and do not mean that they are the necessary orders, unless it is stated that a certain order must be followed.

[0034] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. can explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0035] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be a direct connection or an indirect connection through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0036] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0037] Please refer to Figure 1 , in an embodiment of the present invention, a method for constructing a multi-GPU parallel pyramid in CUDA mode is provided to solve the problem of low efficiency in generating pyramid images for ultra-large images using a CPU conventionally. The specific construction process is as follows:

[0038] First, for the problem that the size of the original input ultra-large remote sensing image is too large for a single processor to process efficiently, the embodiment of the present invention performs pre-block processing on the original input image, and this block process is carried out in a CPU (Graphic Processing Unit, central processing unit) environment. Set the image block size to , and the design basis of this size mainly depends on the video memory size of the GPU (Central Processing Unit, graphics processing unit) and the number of computing units of the GPU, and the value of can be flexibly adjusted according to the video memory size and the number of computing units of the GPU to ensure that each GPU can independently undertake the subsequent processing tasks of one or more blocks. In addition, according to the block size and the scaling ratio of the adjacent-level resolution images of the pyramid image to be generated, the number of levels of the generated pyramid image is calculated. In the embodiment of the present invention, assuming that the width of the original input image is , and the height is , the number of levels of the pyramid image to be generated (excluding the initial input image) is calculated using the following formula:

[0039] ;

[0040] where, represents the number of levels of the pyramid image to be generated (excluding the initial input image, and the initial input image is the target resolution image of level 0, simply referred to as the pyramid of level 0), that is, the number of times of convolutional downsampling.

[0041] In the above formula, the scaling ratio is 2 each time, that is, the target resolution image resolution of the next-level pyramid is half of the target resolution image resolution of the current-level pyramid. The stopping condition for scaling is that the smaller side of the width and height of the scaled image is less than or equal to the side length of the block. .

[0042] After determining the block size of the pyramid image and the number of pyramid image levels, pre-block the target resolution images of all levels of the pyramid image according to the block size to form an index block list, that is, the number of blocks and block positions of each level of the pyramid. Among them, the blocks of the initial input image are the original resolution tiles. Here, the target resolution images of all levels refer to the different resolution images of different levels in the pyramid image, which can be simply referred to as the 0-level pyramid, 1-level pyramid... n-level pyramid. The design of pre-blocking not only effectively avoids the risk of memory overflow but also creates favorable conditions for subsequent parallel computing.

[0043] When processing ultra-large image data, the traditional single CPU processing mode often faces the bottleneck of computing resources. The computing power of a single processor is not enough to efficiently process large-scale data, resulting in slow processing speed and low efficiency. In the embodiments of the present invention, with the help of the powerful parallel computing power of the GPU released by CUDA (Compute Unified Devices Architecture), in the CUDA parallel computing architecture, for the ultra-large image pyramid construction task, this task is to compress the image layer by layer to a low resolution through convolutional downsampling operations. The calculation formula for a single convolutional downsampling operation task is:

[0044] ;

[0045] Among them, represents the target resolution image of the current level and the convolution kernel at the coordinate The convolution result at the location. At the first iteration, the image is the original input image. represents the radius of the convolution kernel, represents the convolution kernel at the coordinate The value at the location, represents the target resolution image of the current level at the coordinate The pixel value at the location; represents the target resolution image of the next level after downsampling, represents the sampling scaling factor, that is, sample once every pixels.

[0046] After the task is created, the CPU is used to read the index block list, and the blocks of the target resolution image at the current level are allocated to the multi-GPU resource pool according to the index block list, and the available GPUs in the current resource pool are called to execute the above convolution downsampling task. Specifically, the CPU divides the target resolution image at the current level according to the index block list. The number of blocks of the target resolution image at any nth level is:

[0047] .

[0048] The total number of blocks of the target resolution images at all levels of the pyramid image is:

[0049] .

[0050] According to the number of GPU threads q in the CUDA architecture and the number of blocks of the target resolution image at the current level, calculate the number of blocks that each CUDA thread should be responsible for .

[0051] To achieve parallel processing, the embodiment of the present invention controls all GPUs to perform parallel convolution downsampling operations on the blocks layer by layer through NVIDIA NVLink (NVIDIA Scalable Link Interface), that is, multiple GPUs simultaneously perform the following tasks in parallel for each block:

[0052] .

[0053] Each GPU independently executes the convolution downsampling task, which greatly improves the processing speed. Although each GPU performs calculations independently, the image data generated by each GPU finally needs to be merged so that the data at all levels obtained by the construction of the image pyramid are geospatially aligned. Therefore, NVIDIA NVLink is also required to perform data transmission and synchronization between GPUs during the convolution downsampling task.

[0054] As a preferred embodiment, in order to avoid overloading of some GPUs, a load balancing mechanism is adopted. During parallel computing, the CPU will monitor the load conditions of each GPU in real time, and the CPU will dynamically adjust the workload of the GPUs or reallocate tasks as needed. Through dynamic adjustment, it is ensured that the resources of all GPUs are fully utilized, avoiding overloading of some GPUs while other GPUs are idle, thereby improving the utilization rate of the overall computing resources.

[0055] For each level of the target resolution image output by GPU downsampling, a geospatial alignment process is required, so that the same geographical location should correspond to the same pixel position in images of different levels. The specific process is as follows: This alignment process is coordinated and integrated by the CPU, and the results processed by each GPU are integrated into a complete multi-resolution pyramid image structure. Through affine transformation, the east-west longitude resolution and the north-south latitude resolution of each level of the image are ensured, so that the pixels at the same spatial position remain consistent in images of different levels. Among them, the longitude and latitude resolutions of the next-level image are twice that of the current level, and for each level down, the resolutions in the east-west direction and the north-south direction are halved.

[0056] As Figure 2 shown, after the pyramid images are aligned, the image data of different resolutions at each level are integrated into a binary general OVR file to obtain a binary set of pyramid levels. The OVR format is the standard format used by ArcGIS to store pyramid images. It can accelerate the display and zoom operations of maps and support the loading of mainstream geographic information and remote sensing software such as GIS. The OVR file has a specific structure, including a file header and a data body. Among them, the file header stores the metadata information of the pyramid and image tiles, that is, it includes the width, height, longitude and latitude range, block size, number of levels of the pyramid image, and the affine transformation data of the target resolution image at each level. The data body stores the color information of the image, that is, it includes the pixel values of the initial input image and the target resolution image at each level of the pyramid image.

[0057] To verify the effect of the present invention, please refer to Figure 3 . In the embodiments of the present invention, pyramids are also constructed for multiple images of different sizes by using the traditional CPU method and the method of the present invention respectively, and the beneficial effects of the method of the present invention are verified according to the construction time consumption. Specifically, it can be found from Figure 3 that when the processed image is relatively small, the time consumption of the traditional CPU method and the present invention is relatively close. However, as the size of the original input image gradually increases, the time consumption of the method of the present invention is significantly shorter, and the gap between the two gradually increases. When the image size reaches 124581×158742, it can be found that the processing efficiency of the method of the present invention is significantly higher, and the processing speed is 11.95 times that of the traditional method.

[0058] In summary, the above are only the preferred embodiments of this specification and are not intended to limit the protection scope of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included in the protection scope of this specification.

[0059] The systems, devices, modules or units described in one or more of the above embodiments may be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0060] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0061] Each embodiment in this specification is described in a progressive manner, and the same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

[0062] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A multi-GPU parallel pyramid construction method in CUDA mode, characterized in that: include: Preset an image block size, calculate the number of levels of a pyramid image to be generated according to the block size, and pre-block the target resolution images of all levels of the pyramid image according to the block size; Under the CUDA parallel computing architecture, the blocks of the target resolution image at each level are assigned to multiple GPUs. Through NVIDIA NVLink technology, all GPUs are controlled to perform parallel convolution downsampling operations on the blocks layer by layer to obtain the target resolution image corresponding to each level of the pyramid image. The calculation formula for a single convolution downsampling operation is: ; in, Indicates the target resolution image of the current level and convolution kernel In coordinates The convolution result at , represents the radius of the convolution kernel, Represents the convolution kernel In coordinates The value at Indicates the target resolution image of the current level In coordinates The pixel value at ; Represents the target resolution image of the next level after downsampling, Represents the sampling scaling factor, that is, every Each pixel is sampled once; The geospatial alignment is performed on each level of target resolution images after GPU downsampling to generate multi-resolution pyramid images.

2. The multi-GPU parallel pyramid construction method under the CUDA mode according to claim 1, characterized in that: The image block size is ,in, The preset value is based on the GPU memory size and the number of GPU computing units.

3. The multi-GPU parallel pyramid construction method under the CUDA mode according to claim 2, characterized in that: The calculation formula for the number of pyramid image levels to be generated is: ; in, Represents the number of levels of the pyramid image to be generated, excluding the initial input image, that is, the number of convolution downsampling. Represent the width and height of the initial input image respectively.

4. The multi-GPU parallel pyramid construction method under CUDA mode according to claim 1, characterized in that: The CPU is used to divide the target resolution image of each level into blocks.

5. The multi-GPU parallel pyramid construction method under CUDA mode according to claim 3, characterized in that: The number of blocks of the target resolution image at level n is: 。 6. The multi-GPU parallel pyramid construction method under CUDA mode according to claim 1, characterized in that: The sampling scaling factor is set to 2.

7. The multi-GPU parallel pyramid construction method under CUDA mode according to claim 1, characterized in that: The CPU uses a dynamic load balancing algorithm to distribute the blocks of the target resolution image at each level to multiple GPUs.

8. The multi-GPU parallel pyramid construction method under CUDA mode according to claim 1, characterized in that: The longitude and latitude information of each block is determined through affine transformation.

9. The multi-GPU parallel pyramid construction method under CUDA mode according to claim 8, characterized in that: Also includes: Generate binary universal OVR files from images of different target resolutions. The OVR files include file headers and data bodies.

10. The multi-GPU parallel pyramid construction method under CUDA mode according to claim 9, characterized in that: The file header includes the width, height, latitude and longitude range, block size, number of levels of the pyramid image, and affine transformation data of each level target resolution image of the initial input image; the data body includes the pixel values ​​of the initial input image and each level target resolution image in the pyramid image.

Citation Information

Patent Citations

  • Three-line-array stereo aerial survey camera parallel spectrum band registration method based on GPU technology

    CN105894494A

  • Pyramid mutual information image registration method based on parallel programming model on GPU cluster

    CN111445503A