Multi-GPU parallel pyramid construction method in CUDA mode

By using multiple GPUs to construct parallel pyramids in CUDA mode, the problem of inefficient processing of traditional CPUs with ultra-large resolution remote sensing images is solved, and the effect of significantly improving processing speed and efficiency is achieved.

CN120031706AActive Publication Date: 2025-05-23JILIN GAOFEN REMOTE SENSING APPL RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510508188.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-23
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Traditional CPU processing methods are inefficient when processing and storing ultra-large resolution remote sensing images, resulting in slow pyramid construction speed, affecting data processing and analysis efficiency.

Method used

The multi-GPU parallel pyramid construction method in CUDA mode is adopted, and the GPU's parallel computing power is used to compress images layer by layer through convolutional downsampling operations, and real-time communication and collaborative control between GPUs are achieved in combination with NVIDIA NVLink technology.

Benefits of technology

The pyramid construction speed of super-large image data is significantly improved, and the processing speed is 11.95 times that of traditional CPU methods, improving the efficiency of data processing and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031706A_ABST
    Figure CN120031706A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of space remote sensing, in particular to a multi-GPU parallel pyramid construction method in a CUDA mode, which comprises the following steps of: performing multi-level pre-blocking on an original input image according to a predefined blocking size, coordinating a multi-GPU parallel processing blocking task, dynamically distributing load and synchronizing data through an NVIDIA NVLink technology under a CUDA parallel computing architecture, and constructing a multi-GPU parallel pyramid by utilizing the CUDA parallel computing architecture. And carrying out convolution downsampling operation on each block, carrying out geographic space alignment on images output by each GPU according to a target resolution image corresponding to each level of the pyramid image, and generating a multi-resolution pyramid file conforming to an OVR format. According to the method, a multi-GPU parallel architecture in a CUDA mode is put forward for pyramid construction for the first time, and the pyramid construction process of super-large image data is accelerated by utilizing the powerful graph parallel computing capability of the GPUs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of space remote sensing technology, and in particular relates to a multi-GPU parallel pyramid construction method in a CUDA mode. Background Art

[0002] The pyramid model is the key to the visualization of massive remote sensing images in the virtual earth. It is a multi-resolution hierarchical structure. Under the premise of ensuring display accuracy, different resolution data are used for different regions to improve rendering efficiency. The construction process of the pyramid is essentially to divide and layer the image. Although this will increase the overall data volume, it can significantly shorten the time required for drawing. The pyramid model has been widely used in two-dimensional map visualization. By resampling the original data, it can effectively speed up the rendering speed.

[0003] However, with the development of remote sensing technology and geographic information systems, the scale of image data continues to expand, and the traditional CPU processing method is still inefficient in processing and storing ultra-large resolution images. In particular, in traditional GIS software, the processing speed is slow and inefficient when creating image pyramids, which seriously affects the efficiency of data processing and analysis.

[0004] Therefore, new technical means are needed to solve the problem of low efficiency in creating pyramids from large images. Summary of the invention

[0005] In view of this, the present invention aims to provide a multi-GPU parallel pyramid construction method under CUDA mode, which utilizes the powerful graphics parallel computing capability of GPU to accelerate the pyramid construction process of ultra-large image data. After the pyramid is created, it can be recognized by mainstream geographic information and remote sensing software, making the ultra-large image display real-time.

[0006] To achieve the above object, the technical solution created by the present invention is implemented as follows: The invention provides a multi-GPU parallel pyramid construction method under CUDA mode, comprising: presetting an image block size, calculating the number of levels of a pyramid image to be generated according to the block size, and pre-blocking target resolution images of all levels of the pyramid image according to the block size; Under the CUDA parallel computing architecture, the blocks of the target resolution image at each level are assigned to multiple GPUs. Through NVIDIA NVLink technology, all GPUs are controlled to perform parallel convolution downsampling operations on the blocks layer by layer to obtain the target resolution image corresponding to each level of the pyramid image. The calculation formula for a single convolution downsampling operation is: ; in, Indicates the target resolution image of the current level and convolution kernel In coordinates The convolution result at , represents the radius of the convolution kernel, Represents the convolution kernel In coordinates The value at Indicates the target resolution image of the current level In coordinates The pixel value at ; Represents the target resolution image of the next level after downsampling, Represents the sampling scaling factor, that is, every Each pixel is sampled once; The geospatial alignment is performed on each level of target resolution images after GPU downsampling to generate multi-resolution pyramid images.

[0007] Preferably, the image block size is ,in, The preset value is based on the GPU memory size and the number of GPU computing units.

[0008] Preferably, the calculation formula for the number of pyramid image levels to be generated is: ; in, Represents the number of levels of the pyramid image to be generated (excluding the initial input image), that is, the number of convolution downsampling. Represent the width and height of the initial input image respectively.

[0009] Preferably, the target resolution image of each level is divided into blocks using a CPU.

[0010] Preferably, the number of blocks of the target resolution image at level n is: .

[0011] Preferably, the sampling scaling factor is 2.

[0012] Preferably, the CPU uses a dynamic load balancing algorithm to distribute the blocks of each level of the target resolution image to multiple GPUs.

[0013] Preferably, the longitude and latitude information of each block is determined by affine transformation.

[0014] Preferably, the method further comprises: generating a binary universal OVR file from images of different levels of target resolution, wherein the OVR file comprises a file header and a data body.

[0015] Preferably, the file header includes the width, height, latitude and longitude range, block size, number of levels of the pyramid image, and affine transformation data of each level target resolution image of the initial input image; the data body includes the pixel values ​​of the initial input image and each level target resolution image in the pyramid image.

[0016] Compared with the prior art, the invention can achieve the following beneficial effects: The present invention proposes for the first time to use a multi-GPU parallel architecture in CUDA mode to construct pyramids. Compared with the traditional CPU processing method, it has obvious advantages in processing ultra-large image data. After testing, when the image to be processed is 124581×158742, the image processing speed of the method of the present invention is 11.95 times that of the traditional CPU method.

[0017] The present invention pre-blocks the target resolution images of all levels of the pyramid image to be generated, which can effectively avoid memory overflow and lay the foundation for subsequent parallel computing. In addition, NVIDIA NVLink technology is used to achieve real-time communication and collaborative control between GPUs, allowing multiple GPUs to process different image blocks in parallel at the same time. Each GPU independently performs convolution and downsampling tasks, which greatly improves the processing speed. A calculation formula combining convolution and downsampling is proposed, which facilitates parallelization through CUDA programming.

[0018] The block division, multi-GPU task allocation and geospatial alignment operations of the present invention are all implemented by the CPU, and the combined application of the CPU and GPU reduces the computing pressure. In addition, in order to avoid overloading of some GPUs, the CPU is used to dynamically adjust the workload of each GPU through a load balancing mechanism, effectively improving the utilization rate of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings constituting part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation on the present invention. In the drawings: Figure 1 is a schematic diagram of a multi-GPU parallel pyramid construction method in CUDA mode provided according to an embodiment of the present invention; Figure 2 A pyramid file data structure diagram provided according to an embodiment of the present invention; Figure 3 It is a statistical diagram of the time consumption for constructing pyramids from images of different sizes using the traditional method provided by the embodiment of the present invention and the method of the present invention. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical scheme and advantages of the invention clearer, the invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the invention and do not constitute a limitation to the invention. Similar components in different embodiments use associated similar component numbers. In the following embodiments, many detailed descriptions are to enable the invention to be better understood. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other components, materials, and methods. In some cases, some operations related to the invention are not shown or described in the specification, in order to avoid the core part of the invention being overwhelmed by too much description, and for those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations according to the description in the specification and the general technical knowledge in the art.

[0021] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to form various implementation methods. At the same time, the steps or actions in the method description can also be interchanged or adjusted in a manner that is obvious to those skilled in the art. Therefore, the various sequences in the specification and the drawings are only for the purpose of clearly describing a certain embodiment and are not meant to be a necessary sequence, unless otherwise specified that a certain sequence must be followed.

[0022] In the description of the invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first", "second", etc. may explicitly or implicitly include one or more of the features. In the description of the invention, unless otherwise specified, the meaning of "multiple" is two or more.

[0023] In the description of the invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installation", "connection" and "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the invention can be understood according to specific circumstances.

[0024] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments.

[0025] See also Figure 1 In one embodiment of the present invention, a multi-GPU parallel pyramid construction method in CUDA mode is provided to solve the problem of low efficiency in traditional pyramid image generation using CPU for large images. The specific construction process is as follows: First, in order to solve the problem that the original input super-large remote sensing image is too large to be efficiently processed by a single processor, the embodiment of the present invention pre-blocks the original input image. The block process is performed in the CPU (Graphic Processing Unit, central processing unit) environment. Set the image block size to The design basis of this size is mainly based on the GPU (Central Processing Unit) video memory size and the number of GPU computing units. It can be flexibly adjusted according to the GPU video memory size and the number of computing units. The value of ensures that each GPU can independently undertake the subsequent processing tasks of one or more blocks. In addition, the number of levels of the generated pyramid image is calculated according to the block size and the scaling ratio of the adjacent level resolution images of the pyramid image to be generated. In this embodiment of the present invention, it is assumed that the width of the original input image is , the height is , use the following formula to calculate the number of pyramid image levels to be generated (excluding the initial input image): ; in, Indicates the number of levels of the pyramid image to be generated (excluding the initial input image, which is the target resolution image of level 0, referred to as the level 0 pyramid), that is, the number of convolution downsampling.

[0026] In the above formula, the scaling ratio is 2 each time, that is, the target resolution image resolution of the next pyramid layer is half of the target resolution image resolution of the current pyramid layer. The stopping condition for scaling is that the smaller side of the width and height of the scaled image is less than or equal to the side length of the block. .

[0027] After determining the block size and the number of pyramid image levels, the target resolution images of all levels of the pyramid image are pre-blocked according to the block size to form an index block list, that is, the number of blocks and block positions of each level of the pyramid, where the blocks of the initial input image are the original resolution tiles. The target resolution images of all levels here refer to the different resolution images of different levels in the pyramid image, which can be referred to as 0-level pyramid, 1-level pyramid...n-level pyramid. The pre-block design not only effectively avoids the risk of memory overflow, but also creates favorable conditions for subsequent parallel computing.

[0028] When processing large image data, the traditional single CPU processing mode often faces the bottleneck of computing resources. The computing power of a single processor is not enough to efficiently process large-scale data, resulting in slow processing speed and low efficiency. In an embodiment of the present invention, with the help of the powerful parallel computing power of the GPU released by the CUDA (Compute Unified Devices Architecture) technology, under the CUDA parallel computing architecture, a task is constructed for a large image pyramid, which is to compress the image to a low resolution layer by layer through convolution downsampling operations. The calculation formula for a single convolution downsampling operation task is: ; in, Indicates the target resolution image of the current level and convolution kernel In coordinates The convolution result at , at the first iteration, the image is the original input image. represents the radius of the convolution kernel, Represents the convolution kernel In coordinates The value at Indicates the target resolution image of the current level In coordinates The pixel value at ; Represents the target resolution image of the next level after downsampling, Represents the sampling scaling factor, that is, every Each pixel is sampled once.

[0029] After the task is created, the CPU is used to read the index block list, and the blocks of the target resolution image of the current level are allocated to the multi-GPU resource pool according to the index block list, and the available GPUs in the current resource pool are called to perform the above convolution downsampling task. Specifically, the CPU divides the target resolution image of the current level according to the index block list. The number of blocks of the target resolution image of any level n is: .

[0030] The total number of blocks of the target resolution image at all levels of the pyramid image is: .

[0031] According to the number of GPU threads q in the CUDA architecture and the number of blocks of the target resolution image at the current level, calculate the number of blocks that each CUDA thread should be responsible for .

[0032] To achieve parallel processing, the embodiment of the present invention controls all GPUs to perform parallel convolution downsampling operations on blocks layer by layer through NVIDIA NVLink (NVIDIA Scalable Link Interface) technology, that is, multiple GPUs perform the following tasks on each block in parallel at the same time: .

[0033] Each GPU performs convolution downsampling tasks independently, greatly improving processing speed. Although each GPU performs calculations independently, the image data generated by each GPU ultimately needs to be merged so that all levels of data obtained by building the image pyramid are aligned in geographic space. Therefore, NVIDIA NVLink is also required to transfer and synchronize data between GPUs during the convolution downsampling task.

[0034] As a preferred embodiment, in order to avoid overloading of certain GPUs, a load balancing mechanism is adopted. During the parallel computing process, the CPU will monitor the load of each GPU in real time, and dynamically adjust the workload of the GPU or reallocate tasks as needed. Through dynamic adjustment, it is ensured that the resources of all GPUs are fully utilized, avoiding overloading of certain GPUs while other GPUs are idle, thereby improving the utilization of overall computing resources.

[0035] Each target resolution image of the GPU downsampling output needs to be geospatial aligned so that the same geographic location should correspond to the same pixel position in different levels of images. The specific process is as follows: The alignment process is coordinated and integrated by the CPU, integrating the results of each GPU processing into a complete multi-resolution pyramid image structure. Affine transformation is used to ensure that the longitude and latitude resolution of each level of image in the east-west direction is consistent. and the latitudinal resolution in the north-south direction , so that the pixels of the same spatial position in different levels of images remain consistent. Among them, the latitude and longitude resolution of the next level of images is twice that of the current level, and the resolution in the east-west and north-south directions is halved for each layer down.

[0036] like Figure 2 As shown in the figure, after the pyramid image is aligned, the image data of each level with different resolutions are integrated into a binary universal OVR file to obtain a binary set of pyramid hierarchical structure. The OVR format is the standard format used by ArcGIS to store pyramid images. It can accelerate the display and zooming operations of maps and support the loading of mainstream geographic information and remote sensing software such as GIS. The OVR file has a specific structure, which includes a file header and a data body. Among them, the file header stores metadata information of the pyramid and image tiles, including the width, height, longitude and latitude range, block size, number of levels of the pyramid image, and affine transformation data of the target resolution image of each level. The data body stores the color information of the image, including the pixel values ​​of the initial input image and the target resolution image of each level in the pyramid image.

[0037] To verify the effect of the present invention, please refer to Figure 3 The embodiment of the present invention also uses the traditional CPU method and the method of the present invention to construct pyramids for multiple images of different sizes, and verifies the beneficial effect of the method of the present invention based on the construction time. Figure 3 It can be found that when the processed image is relatively small, the time consumption of the traditional CPU method and the present invention is relatively close, but as the size of the original input image gradually increases, the time consumption of the present invention method is significantly shorter, and the difference between the two gradually increases. When the image size reaches 124581×158742, it can be found that the processing efficiency of the present invention method is significantly higher, and the processing speed is 11.95 times that of the traditional method.

[0038] In short, the above description is only a preferred embodiment of this specification and is not intended to limit the protection scope of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included in the protection scope of this specification.

[0039] The systems, devices, modules or units described in one or more of the above embodiments may be implemented by a computer chip or entity, or by a product having a certain function. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0040] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0041] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0042] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A multi-GPU parallel pyramid construction method in CUDA mode, characterized in that: include: Preset an image block size, calculate the number of levels of a pyramid image to be generated according to the block size, and pre-block the target resolution images of all levels of the pyramid image according to the block size; Under the CUDA parallel computing architecture, the blocks of the target resolution image at each level are assigned to multiple GPUs. Through NVIDIA NVLink technology, all GPUs are controlled to perform parallel convolution downsampling operations on the blocks layer by layer to obtain the target resolution image corresponding to each level of the pyramid image. The calculation formula for a single convolution downsampling operation is: ; in, Indicates the target resolution image of the current level and convolution kernel In coordinates The convolution result at , represents the radius of the convolution kernel, Represents the convolution kernel In coordinates The value at Indicates the target resolution image of the current level In coordinates The pixel value at ; Represents the target resolution image of the next level after downsampling, Represents the sampling scaling factor, that is, every Each pixel is sampled once; The geospatial alignment is performed on each level of target resolution images after GPU downsampling to generate multi-resolution pyramid images.

2. The multi-GPU parallel pyramid construction method under the CUDA mode according to claim 1, characterized in that: The image block size is ,in, The preset value is based on the GPU memory size and the number of GPU computing units.

3. The multi-GPU parallel pyramid construction method under the CUDA mode according to claim 2, characterized in that: The calculation formula for the number of pyramid image levels to be generated is: ; in, Represents the number of levels of the pyramid image to be generated (excluding the initial input image), that is, the number of convolution downsampling. Represent the width and height of the initial input image respectively.

4. The multi-GPU parallel pyramid construction method under CUDA mode according to claim 1, characterized in that: The CPU is used to divide the target resolution image of each level into blocks.

5. The multi-GPU parallel pyramid construction method under CUDA mode according to claim 3, characterized in that: The number of blocks of the target resolution image at level n is: 。 6. The multi-GPU parallel pyramid construction method under CUDA mode according to claim 1, characterized in that: The sampling scaling factor is set to 2.

7. The multi-GPU parallel pyramid construction method under CUDA mode according to claim 1, characterized in that: The CPU uses a dynamic load balancing algorithm to distribute the blocks of the target resolution image at each level to multiple GPUs.

8. The multi-GPU parallel pyramid construction method under CUDA mode according to claim 1, characterized in that: The longitude and latitude information of each block is determined through affine transformation.

9. The multi-GPU parallel pyramid construction method under CUDA mode according to claim 8, characterized in that: Also includes: Generate binary universal OVR files from images of different target resolutions. The OVR files include file headers and data bodies.

10. The multi-GPU parallel pyramid construction method under CUDA mode according to claim 9, characterized in that: The file header includes the width, height, latitude and longitude range, block size, number of levels of the pyramid image, and affine transformation data of each level target resolution image of the initial input image; the data body includes the pixel values ​​of the initial input image and each level target resolution image in the pyramid image.

Citation Information

Patent Citations

  • Three-line-array stereo aerial survey camera parallel spectrum band registration method based on GPU technology

    CN105894494A

  • Pyramid mutual information image registration method based on parallel programming model on GPU cluster

    CN111445503A