Pixel extraction and reconstruction method for large-size SAR image simulation of unmanned aerial vehicle
By combining image size alignment and automatic edge-padding mechanisms with array slicing and the CycleGAN model, efficient segmentation and seamless reconstruction of UAV-borne SAR images were achieved, solving the problems of high computational complexity and insufficient memory in large-size image processing, and improving image quality and processing efficiency.
Patent Information
- Application Number
- CN202511429495.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-10-09
AI Technical Summary
When processing SAR images on UAVs, existing technologies suffer from problems such as high computational complexity, excessive memory consumption, inaccurate image reconstruction, and edge misalignment. In particular, in the processing of large-size images, traditional segmentation methods are time-consuming and produce poor-quality reconstructed images, while models such as CycleGAN are prone to memory overflow and information loss when large-size images are input.
By employing image size alignment strategies, automatic edge-padding mechanisms, and block numbering logic, image segmentation and stitching are performed by replacing multiple loops with array slicing. Combined with deep learning models such as CycleGAN, efficient segmentation and seamless reconstruction of large-size images are achieved.
It significantly reduces image processing time complexity and memory usage, improves spatial consistency and visual quality of image reconstruction, solves the boundary demarcation problem of large-size images, and is highly adaptable to various types of image processing.
Smart Images

Figure CN120892586A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of remote sensing image processing and computer vision, and particularly relates to a pixel extraction and reconstruction method for large-size SAR image simulation of an unmanned aerial vehicle. BACKGROUND
[0002] The unmanned aerial vehicle-borne SAR system has become an important means for obtaining high-resolution ground observation data due to its characteristics of being flexible and maneuverable, not being restricted by weather and illumination, and being capable of realizing all-weather and all-day imaging. Compared with the traditional satellite SAR, the unmanned aerial vehicle-borne SAR is closer to the observation target and can obtain image data with centimeter-level to meter-level resolution, but it can produce large-size images of dozens to hundreds of GB in a single flight, such as scene images with a width of several kilometers and a resolution of 0.5 meters, and there are significant speckle noise, geometric distortion, and radiation non-uniformity in the images.
[0003] When processing such images, direct processing often faces problems of high memory occupation and low computing efficiency due to large data volume and high resolution. In the prior art, the large image is usually divided into multiple sub-blocks for step-by-step processing, and the processed sub-blocks are then reconstructed into a complete image. However, the traditional segmentation method usually uses multiple loops for pixel-level processing, which has the problems of high computational complexity and slow processing speed. Meanwhile, in the reconstruction process, due to inaccurate sub-block splicing, edge misalignment and image distortion may occur, affecting the quality of the reconstructed image.
[0004] With the development of generative adversarial networks, unsupervised image translation models such as CycleGAN have been widely used in style transfer, remote sensing image pseudo-color conversion, SAR image reconstruction, etc. CycleGAN can realize high-quality mapping between different image domains without paired samples by introducing a cycle consistency constraint, and has good cross-domain generalization ability. However, there are certain limitations in processing large-size images such as 2048x2048 and 4096x4096 using CycleGAN and other image generation networks. Inputting a large-size image into a deep neural network will cause a sharp increase in memory consumption, and the training process may be interrupted due to memory overflow, so it is usually only possible to process small-size images such as 256x256 or 512x512. Large-size images often have sizes that cannot be evenly divided, resulting in inconsistent sizes of image blocks in the last row or column, affecting the accuracy of model training and subsequent image splicing. If cropping or padding is used, it may cause loss or distortion of the original image information. In practical applications, in order to adapt to network input, large images are often divided into several fixed-size small blocks for separate processing. However, due to the inability of the block cutting operation to completely preserve the context information, CycleGAN may generate areas with inconsistent styles or obvious boundaries at the boundaries of each image block, resulting in grid-like or splicing marks in the reconstructed image.
[0005] In the unmanned aerial SAR disaster monitoring, after the occurrence of earthquake or flood, the obtained large-size SAR image of disaster area needs to be segmented to identify the house damage area, and then the local result is reconstructed into a complete disaster map. If the traditional method is used, not only the segmentation is time-consuming and affects the rescue decision, but also the reconstructed image may misjudge the road connectivity due to the misalignment between blocks, which may cause serious consequences.
[0006] Therefore, in view of the characteristics of large size, strong noise and high precision reconstruction of unmanned aerial SAR (synthetic aperture radar) image, a technical solution is needed to efficiently segment, low memory occupation, accurate reconstruction and adapt to deep learning model, in order to promote the landing of unmanned aerial SAR in high-precision remote sensing application. SUMMARY
[0007] The present application provides a large-size image segmentation and reconstruction method based on extracted pixels, which realizes information preservation and edge continuity in the image segmentation process by introducing image size alignment strategy, automatic edge filling mechanism, block numbering logic and lossless splicing method in reconstruction, improves the spatial consistency and visual quality of the final image translation result, and is suitable for high-resolution image processing scenes such as SAR image.
[0008] To achieve the above purpose, the present application adopts the following technical scheme: A pixel extraction and reconstruction method for large-size SAR image simulation of unmanned aerial vehicle, comprising the following steps: S1. Obtain the path and image type of the image to be processed, and create a storage directory; S2. Read the image to be processed to obtain its original width and original height; S3. According to the preset sub-block size, calculate the width and height of the filled image, and the width and height of the filled image are both integer multiples of the sub-block size; S4. Adjust the image to be processed to the width and height after filling to obtain the filled image, and store the filled image under the storage directory; S5. According to the size of the filled image and the sub-block size, determine the number of rows and columns of the segmented sub-blocks; S6. According to the number of rows and columns of the sub-blocks, the filled image is segmented into a plurality of sub-block images by using array slicing method, and each sub-block image is stored under the storage directory and its storage path is recorded; S7. Obtain the storage directory of the sub-block image, the width and height of the filled image, and the width and height of the original image; S8. Create an empty array with the same size as the filled image as the reconstruction canvas; S9. Based on the number of rows and columns of the sub-block, read the image of each sub-block in sequence, and write the sub-block image to the corresponding position on the reconstruction canvas using array slicing; S10. Convert the reconstructed canvas into an image, adjust its width and height to match the original image, obtain the reconstructed image, and save it; S11. Delete the storage directory of the sub-block image.
[0009] To optimize the above technical solution, the specific measures also include: Furthermore, in S1, the path of the image to be processed is a local storage path or a network path, and the image type includes jpg format and png format; The creation of the storage directory is specifically done recursively to ensure that parent directories at all levels are automatically created when the parent directory does not exist, and that the output directory is not created again when it already exists. The naming rule for the storage directory is set to "original image name + _split + _timestamp".
[0010] Furthermore, S2 specifically refers to: Open image files using functions from a library that supports multiple image formats; The original width and original height can be obtained through the size property of the image object. The first element of the tuple returned by the size property is the width, and the second element is the height. If an error occurs during the reading process, the error message will be logged, and empty image data and zero-value size information will be returned.
[0011] Furthermore, S3 specifically refers to: The preset sub-block size is used to calculate the width and height of the image after filling, based on the preset sub-block size. The calculation method is as follows: The formula for calculating the width after padding is: padded_width = max (split_width, math.ceil(original_width / split_width) * split_width), where padded_width is the width of the image to be processed after padding, split_width is the width of the sub-block, and original_width is the original width of the image to be processed; The formula for calculating the height after filling is: padded_height = max (split_height, math.ceil(original_height / split_height) * split_height), where padded_height is the height of the image to be processed after filling, split_height is the height of the sub-block, and original_height is the original height of the image to be processed; The `math.ceil` function is used to round up; the `max` function is used to find the maximum value.
[0012] Furthermore, S4 specifically refers to: An image scaling algorithm is used to adjust the width and height of the image to be processed to the padded width and height. The storage path of the padded image is set to "padded_image + image type" in the output directory, and the original color mode of the image to be processed is preserved when storing.
[0013] Furthermore, S5 specifically refers to: The formula for calculating the number of sub-block rows is: blocks_row_num = padded_height / split_height, where blocks_row_num represents the number of sub-block rows, padded_height represents the height of the image to be processed after filling, and split_height represents the height of the sub-block; The formula for calculating the number of sub-block columns is: blocks_col_num = padded_width / split_width, where blocks_col_num is the number of sub-block columns, padded_width is the width of the image to be processed after filling, and split_width is the width of the sub-block.
[0014] Furthermore, S6 specifically refers to: The padded image is converted into a NumPy array with dimensions padded_height * padded_width * channels. Here, padded_height represents the height of the padded image, padded_width represents the width of the padded image, and channels represents the number of channels in the image. Images in RGB mode have 3 channels, images in RGBA mode have 4 channels, and images in grayscale mode have 1 channel. For the sub-block in row 'row' and column 'col': The starting row coordinates are start_row = row * split_height; where start_row is the starting row coordinates and split_height represents the height of the sub-block. The coordinates of the end row are given by end_row = start_row + split_height; where end_row is the coordinates of the end row, start_row is the coordinates of the start row, and split_height represents the height of the sub-block. The starting column coordinates are start_col = col * split_width; where start_col is the starting column coordinates and split_width is the width of the sub-block. The coordinates of the end column are given by end_col = start_col + split_width; where end_col is the coordinate of the end column. Extract the array data corresponding to the sub-block using the array slice `img_array [start_row:end_row, start_col:end_col]`; where `img_array` is a NumPy array. The extracted sub-block array data is converted into image objects using the Image.fromarray function and stored according to the naming rule of "padded_image + row index + column index + image type"; The storage paths of the sub-block images are all recorded in a list, which corresponds one-to-one with the row and column indices of the sub-blocks.
[0015] Furthermore, S7 specifically refers to: The process involves traversing the storage directory to obtain all sub-block images, the width of the image to be processed after filling, the height of the image to be processed after filling, the original width of the image to be processed, and the original height of the image to be processed. If any abnormalities are found in the obtained size information, the reconstruction process is terminated and an error message is returned. S8 specifically refers to: The dimensions of the empty array are set to padded_height*padded_width*channels, where padded_height represents the height of the image to be processed after being padded, padded_width is the width of the image to be processed after being padded, and channels is the number of channels of the image, which is consistent with the number of channels of the image before segmentation; An empty array has its data type set to an 8-bit unsigned integer, and its initial value is set to 0.
[0016] Furthermore, S9 specifically refers to: For the sub-block in row 'row' and column 'col': The filename of the sub-padded image is "padded_image + row + col + image type", and the corresponding sub-padded image file can be found in the storage directory accordingly; If the sub-block image file does not exist, record a warning message and skip the sub-block, then continue processing other sub-blocks; Read the sub-block image and convert it to a NumPy array; The position of this sub-block image in the reconstruction canvas is determined by the following formula: The starting row coordinates are start_row = row * split_height; where start_row is the starting row coordinates and split_height represents the height of the sub-block. The coordinates of the end row are given by end_row = start_row + split_height; where end_row is the coordinates of the end row, start_row is the coordinates of the start row, and split_height represents the height of the sub-block. The starting column coordinates are start_col = col * split_width; where start_col is the starting column coordinates and split_width is the width of the sub-block. The coordinates of the end column are given by end_col = start_col + split_width; where end_col is the coordinate of the end column. The sub-block data is written to the reconstructed canvas array `rebuilt_array` by slicing the array `rebuilt_array [start_row:end_row, start_col:end_col] = block_array`, where `block_array` represents the NumPy array of the sub-block image.
[0017] Furthermore, S10 specifically refers to: Convert the reconstructed canvas array into an image object; A scaling algorithm is used to adjust the image object to its original width and height; The naming convention for reconstructed images is set to "merge_ + original image name + image type"; The storage path for the reconstructed image is set to the parent directory of the sub-block storage directory.
[0018] The beneficial effects of this invention are: Replacing traditional multiple loops with array slicing for image segmentation and stitching reduces the time complexity of image segmentation and reconstruction from O(n^2) to O(n^2). 2 m 2The processing efficiency is reduced to O(nm) (where n and m are the number of pixels in the height and width directions of the image, respectively), which significantly improves processing efficiency.
[0019] By segmenting large images into sub-blocks for processing, memory consumption is significantly reduced, resolving the issue of program crashes or slow performance caused by insufficient memory when CycleGAN directly processes large images. This makes style transfer and other processing of large images possible. It can adapt to the step-by-step processing requirements of models like CycleGAN for large images, shortening the overall processing time.
[0020] This invention can significantly reduce boundary line problems caused by image stitching. Traditional methods often encounter problems such as color abruptness and edge misalignment when restoring image blocks, resulting in obvious grid-like boundary lines or discontinuous areas in the generated image. This invention adopts a unified edge patching, non-overlapping block cutting, and in-situ alignment restoration mechanism to ensure seamless stitching of image blocks during reconstruction, minimizing boundary errors in image reconstruction.
[0021] It is highly adaptable and suitable for various types of images and different sub-block size settings. It has good versatility and flexibility, especially in the field of large-size image processing such as SAR images. It can also be seamlessly integrated with deep learning models such as CycleGAN, expanding the application scope of the model. Attached Figure Description
[0022] Figure 1 This is an overall flowchart of the method of the present invention; Figure 2 Here is the flowchart for the image segmentation module; Figure 3 A schematic diagram illustrating pixel extraction and segmentation into sub-image blocks for a large image; Figure 4 Here is a flowchart of the image reconstruction module; Figure 5 This is a flowchart illustrating the combination of this invention with CycleGAN for large-scale image block processing. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0024] Example 1 This invention proposes a pixel extraction and reconstruction method for simulating large-size SAR images from unmanned aerial vehicles (UAVs). The overall process of this method is as follows: Figure 1As shown, it includes two parts: image segmentation and image reconstruction. Image segmentation is as follows: Figure 2 As shown, image reconstruction is as follows Figure 4 As shown, the overall method includes the following steps: S1. Obtain the path and image type of the image to be processed, and create a storage directory; the path of the image to be processed is a local storage path or a network path, and the image type includes jpg and png formats; the creation of the storage directory is specifically done recursively to ensure that parent directories at all levels are automatically created when the parent directory does not exist, and that the output directory is not created repeatedly when it already exists, thus avoiding redundant operations; the naming rule of the storage directory is set to "original image name + _split + _timestamp" to ensure the uniqueness of the directory and prevent the segmented sub-blocks of different images from being confused with each other.
[0025] S2. Read the image to be processed and obtain its original width and original height; specifically: Use functions from image reading libraries that support multiple image formats to open image files; such as the Image.open function in the PIL library, to ensure that images of different formats can be read effectively. The original width and original height can be obtained through the size property of the image object. The first element of the tuple returned by the size property is the width, and the second element is the height. If an error occurs during the reading process, such as the file not existing or the format not being supported, the error message will be logged and empty image data and zero-value size information will be returned.
[0026] S3. Based on the preset sub-block size, calculate the width and height of the image to be processed after filling, where the width and height after filling are both integer multiples of the sub-block size; specifically: The preset sub-block size can be adjusted according to the actual application scenario. Typical values are 128×128, 256×256, 512×512, etc., and the width and height of the sub-blocks can be set to the same or different values. Same values result in square sub-blocks, while different values result in rectangular sub-blocks. Based on the preset sub-block size, the width and height of the image to be processed after filling are calculated. When the original width is less than the sub-block width, the filled width is the sub-block width; when the original width is an integer multiple of the sub-block width, the filled width is equal to the original width; otherwise, the filled width is the minimum value that is greater than the original width and an integer multiple of the sub-block width. The same applies to the height. The calculation method is expressed by the following formula: The formula for calculating the padded width is: padded_width = max (split_width, math.ceil(original_width / split_width) * split_width), where padded_width is the padded width of the image to be processed, split_width is the width of the sub-block, original_width is the original width of the image to be processed, the math.ceil function is used to round up, and the max function is used to get the maximum value. The formula for calculating the height after filling is: padded_height = max (split_height, math.ceil(original_height / split_height) * split_height), where padded_height is the height of the image to be processed after filling, split_height is the height of the sub-block, and original_height is the original height of the image to be processed.
[0027] S4. Adjust the image to be processed to the width and height after padding, obtain the padded image, and save the padded image in the storage directory; specifically: An image scaling algorithm is used to adjust the image to the width and height after padding, ensuring the clarity and detail retention of the padded image. The storage path for the padded image is set to "padded_image + image type" in the output directory. For example, if the output directory is "test_split_202310011200" and the image type is ".png", then the storage path is "test_split_202310011200 / padded_image.png". The original color mode of the image is preserved during storage, such as RGB, RGBA, and grayscale, ensuring complete preservation of image information.
[0028] S5. Based on the dimensions of the filled image and the sub-block dimensions, determine the number of rows and columns of the segmented sub-blocks; S5 specifically involves: The formula for calculating the number of sub-block rows is: blocks_row_num = padded_height / split_height, where blocks_row_num represents the number of sub-block rows, padded_height represents the height of the image to be processed after filling, and split_height represents the height of the sub-block; that is, the height after filling is divided by the height of the sub-block, and the quotient obtained by integer division is the number of sub-block rows.
[0029] The formula for calculating the number of sub-block columns is: blocks_col_num = padded_width / split_width, where blocks_col_num is the number of sub-block columns, padded_width is the width of the image after filling, and split_width is the width of the sub-block. That is, the padded width is divided by the sub-block width using integer division, and the quotient is the number of sub-block columns.
[0030] For example, if the height after filling is 1024 and the height of the child block is 256, then the number of rows in the child block is 1024 / 256=4; if the width after filling is 768 and the width of the child block is 256, then the number of columns in the child block is 768 / 256=3.
[0031] S6. Based on the number of rows and columns of each sub-block, divide the filled image into multiple sub-block images using array slicing. Store each sub-block image in a storage directory and record its storage path; for example... Figure 3 As shown, S6 specifically refers to: The padded image is converted into a NumPy array (img_array) with dimensions padded_height * padded_width * channels. Here, padded_height represents the height of the padded image, padded_width represents the width of the padded image, and channels represents the number of channels in the image. Images in RGB mode have 3 channels, images in RGBA mode have 4 channels, and images in grayscale mode have 1 channel. For the sub-block in row 'row' (where row ranges from 0 to blocks_row_num - 1) and column 'col' (where col ranges from 0 to blocks_col_num - 1): The starting row coordinates are start_row = row * split_height; where start_row is the starting row coordinates and split_height represents the height of the sub-block. The coordinates of the end row are given by end_row = start_row + split_height; where end_row is the coordinates of the end row, start_row is the coordinates of the start row, and split_height represents the height of the sub-block. The starting column coordinates are start_col = col * split_width; where start_col is the starting column coordinates and split_width is the width of the sub-block. The coordinates of the end column are given by end_col = start_col + split_width; where end_col is the coordinate of the end column. Extract the array data corresponding to the sub-block using the array slice `img_array [start_row:end_row, start_col:end_col]`; where `img_array` is a NumPy array. The extracted sub-block array data is converted into image objects using the Image.fromarray function and stored according to the naming rule of "padded_image + row index + column index + image type"; for example, the sub-block in row 0 and column 1 is stored as "padded_image_0_1.png".
[0032] The storage paths of the sub-block images are all recorded in a list, which corresponds one-to-one with the row and column indices of the sub-blocks.
[0033] S7. Obtain the storage directory of the sub-block image, the width and height of the padded image, and the width and height of the original image; S7 specifically includes: The process involves traversing the storage directory to obtain all sub-block images, the width of the image to be processed after filling, the height of the image to be processed after filling, the original width of the image to be processed, and the original height of the image to be processed. If the obtained size information is abnormal, such as being negative or zero, the reconstruction process is terminated and an error message is returned. S8. Create an empty array with the same dimensions as the filled image as the reconstruction canvas; specifically: The dimensions of the empty array are set to padded_height*padded_width*channels, where padded_height represents the height of the image to be processed after being padded, padded_width is the width of the image to be processed after being padded, and channels is the number of channels of the image, which is consistent with the number of channels of the image before segmentation; The data type of the empty array is set to an 8-bit unsigned integer (np.uint8) to match the range of image pixel values from 0 to 255; the initial value of the array is set to 0 to ensure that the uncovered areas do not affect the quality of the reconstructed image.
[0034] S9. Based on the number of rows and columns of each sub-block, read the image of each sub-block sequentially, and write the sub-block image to the corresponding position on the reconstructed canvas using an array slicing method; specifically: For the sub-block in row 'row' and column 'col': The filename of the sub-padded image is "padded_image + row + col + image type", and the corresponding sub-padded image file can be found in the storage directory accordingly; If the sub-block image file does not exist, record a warning message and skip the sub-block, then continue processing other sub-blocks; Read the sub-block image and convert it to a NumPy array; The position of this sub-block image in the reconstruction canvas is determined by the following formula: The starting row coordinates are start_row = row * split_height; where start_row is the starting row coordinates and split_height represents the height of the sub-block. The coordinates of the end row are given by end_row = start_row + split_height; where end_row is the coordinates of the end row, start_row is the coordinates of the start row, and split_height represents the height of the sub-block. The starting column coordinates are start_col = col * split_width; where start_col is the starting column coordinates and split_width is the width of the sub-block. The coordinates of the end column are given by end_col = start_col + split_width; where end_col is the coordinate of the end column. The sub-block data is written to the reconstructed canvas array `rebuilt_array` by slicing the array `rebuilt_array [start_row:end_row, start_col:end_col] = block_array`, where `block_array` represents the NumPy array of the sub-block image.
[0035] S10. Convert the reconstructed canvas into an image, adjust its width and height to match the original image, obtain the reconstructed image, and save it; specifically: Convert the reconstructed canvas array into an image object; Scaling algorithms, such as bicubic interpolation, are used to adjust the image objects to their original width and height, ensuring that the image size is consistent with the original image before segmentation. The naming convention for reconstructed images is set to "merge_ + original image name + image type"; for example, if the original image is "SARMap.jpg", then the reconstructed image is stored as "merge_SARMap.jpg".
[0036] The storage path for the reconstructed image is set to the parent directory of the sub-block storage directory to distinguish it from the original image and the segmentation directory.
[0037] S11. Delete the storage directory of sub-block images. After the reconstructed image is successfully stored, the sub-block storage directory and all sub-block image files contained therein are removed using a recursive deletion method. If errors such as insufficient permissions occur during the deletion process, the error information is logged, but this does not affect the availability of the reconstructed image.
[0038] This invention is highly adaptable and can accommodate the step-by-step processing requirements of models such as CycleGAN for large images. For example, if style transfer of a large image is required using CycleGAN, only the style transfer of the sub-block images obtained in step S6 needs to be performed on each sub-block image to obtain style-transferred sub-block images; then, by performing steps S7-S11 on these style-transferred sub-block images, the style-transferred large image can be reconstructed. Figure 5 As shown, this solves the problem of program crashes or slow operation caused by insufficient memory when CycleGAN directly processes large images.
[0039] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0040] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A pixel extraction and reconstruction method for simulating large-size SAR images from unmanned aerial vehicles, characterized in that, Includes the following steps: S1. Obtain the path and image type of the image to be processed, and create a storage directory; S2. Read the image to be processed and obtain its original width and original height; S3. Calculate the width and height of the image to be processed after filling according to the preset sub-block size. The width and height after filling are both integer multiples of the sub-block size. S4. Adjust the image to be processed to the width and height after filling, obtain the filled image, and save the filled image in the storage directory; S5. Determine the number of rows and columns of the segmented sub-blocks based on the size of the filled image and the size of the sub-blocks; S6. Based on the number of rows and columns of the sub-blocks, the filled image is divided into multiple sub-block images using array slicing. Each sub-block image is stored in the storage directory and its storage path is recorded. S7. Obtain the storage directory of the sub-block image, the width and height of the padded image, and the width and height of the original image; S8. Create an empty array with the same size as the filled image as the reconstruction canvas; S9. Based on the number of rows and columns of the sub-block, read the image of each sub-block in sequence, and write the sub-block image to the corresponding position on the reconstruction canvas using array slicing; S10. Convert the reconstructed canvas into an image, adjust its width and height to match the original image, obtain the reconstructed image, and save it.
2. The pixel extraction and reconstruction method for simulating large-size SAR images of UAVs as described in claim 1, characterized in that, In S1, the path of the image to be processed is a local storage path or a network path, and the image type includes jpg format and png format; The creation of the storage directory is specifically done recursively to ensure that parent directories at all levels are automatically created when the parent directory does not exist, and that the output directory is not created again when it already exists. The naming rule for the storage directory is set to "original image name + _split + _timestamp".
3. The pixel extraction and reconstruction method for simulating large-size SAR images of UAVs as described in claim 1, characterized in that, S2 specifically refers to: Open image files using functions from a library that supports multiple image formats; The original width and original height can be obtained through the size property of the image object. The tuple returned by the size property contains two elements, one for the width and the other for the height. If an error occurs during the reading process, the error message will be logged, and empty image data and zero-value size information will be returned.
4. The pixel extraction and reconstruction method for simulating large-size SAR images of UAVs as described in claim 1, characterized in that, S3 specifically refers to: The preset sub-block size is used to calculate the width and height of the image after filling, based on the preset sub-block size. The calculation method is as follows: The formula for calculating the width after padding is: padded_width = max (split_width, math.ceil(original_width / split_width) * split_width), where padded_width is the width of the image to be processed after padding, split_width is the width of the sub-block, and original_width is the original width of the image to be processed; The formula for calculating the height after filling is: padded_height = max (split_height, math.ceil(original_height / split_height) * split_height), where padded_height is the height of the image to be processed after filling, split_height is the height of the sub-block, and original_height is the original height of the image to be processed; The `math.ceil` function is used to round up; the `max` function is used to find the maximum value.
5. The pixel extraction and reconstruction method for simulating large-size SAR images of UAVs as described in claim 1, characterized in that, S4 specifically refers to: An image scaling algorithm is used to adjust the width and height of the image to be processed to the padded width and height. The storage path of the padded image is set to "padded_image + image type" in the output directory, and the original color mode of the image to be processed is preserved when storing.
6. The pixel extraction and reconstruction method for simulating large-size SAR images of UAVs as described in claim 1, characterized in that, S5 specifically refers to: The formula for calculating the number of sub-block rows is: blocks_row_num = padded_height / split_height, where blocks_row_num represents the number of sub-block rows, padded_height represents the height of the image to be processed after filling, and split_height represents the height of the sub-block; The formula for calculating the number of sub-block columns is: blocks_col_num = padded_width / split_width, where blocks_col_num represents the number of sub-block columns, padded_width represents the width of the image to be processed after filling, and split_width represents the width of the sub-block.
7. The pixel extraction and reconstruction method for simulating large-size SAR images of UAVs as described in claim 1, characterized in that, S6 specifically refers to: The padded image is converted into a NumPy array with dimensions padded_height * padded_width * channels. Here, padded_height represents the height of the padded image, padded_width represents the width of the padded image, and channels represents the number of channels in the image. Images in RGB mode have 3 channels, images in RGBA mode have 4 channels, and images in grayscale mode have 1 channel. For the sub-block in row 'row' and column 'col': The starting row coordinates are start_row = row * split_height; where start_row is the starting row coordinates and split_height represents the height of the sub-block. The coordinates of the end row are given by end_row = start_row + split_height; where end_row is the coordinates of the end row, start_row is the coordinates of the start row, and split_height represents the height of the sub-block. The starting column coordinates are start_col = col * split_width; where start_col is the starting column coordinates and split_width is the width of the sub-block. The coordinates of the end column are given by end_col = start_col + split_width; where end_col is the coordinate of the end column. Extract the array data corresponding to the sub-block using the array slice `img_array [start_row:end_row, start_col:end_col]`; where `img_array` is a NumPy array. The extracted sub-block array data is converted into image objects using the Image.fromarray function and stored according to the naming rule of "padded_image + row index + column index + image type"; The storage paths of the sub-block images are all recorded in a list, which corresponds one-to-one with the row and column indices of the sub-blocks.
8. The pixel extraction and reconstruction method for simulating large-size SAR images of UAVs as described in claim 1, characterized in that, S7 specifically refers to: The process involves traversing the storage directory to obtain all sub-block images, the width of the image to be processed after filling, the height of the image to be processed after filling, the original width of the image to be processed, and the original height of the image to be processed. If any abnormalities are found in the obtained size information, the reconstruction process is terminated and an error message is returned. S8 specifically refers to: The dimensions of the empty array are set to padded_height*padded_width*channels, where padded_height represents the height of the image to be processed after being padded, padded_width is the width of the image to be processed after being padded, and channels is the number of channels of the image, which is consistent with the number of channels of the image before segmentation; An empty array has its data type set to an 8-bit unsigned integer, and its initial value is set to 0.
9. The pixel extraction and reconstruction method for simulating large-size SAR images of UAVs as described in claim 1, characterized in that, S9 specifically refers to: For the sub-block in row 'row' and column 'col': The filename of the sub-padded image is "padded_image + row + col + image type", which is used to locate the corresponding sub-padded image file in the storage directory; If the sub-block image file does not exist, record a warning message and skip the sub-block, then continue processing other sub-blocks; Read the sub-block image and convert it to a NumPy array; The position of this sub-block image in the reconstruction canvas is determined by the following formula: The starting row coordinates are start_row = row * split_height; where start_row is the starting row coordinates and split_height represents the height of the sub-block. The coordinates of the end row are given by end_row = start_row + split_height; where end_row is the coordinates of the end row, start_row is the coordinates of the start row, and split_height represents the height of the sub-block. The starting column coordinates are start_col = col * split_width; where start_col is the starting column coordinates and split_width is the width of the sub-block. The coordinates of the end column are given by end_col = start_col + split_width; where end_col is the coordinate of the end column. The sub-block data is written to the reconstructed canvas array `rebuilt_array` by slicing the array `rebuilt_array [start_row:end_row, start_col:end_col] = block_array`, where `block_array` represents the NumPy array of the sub-block image.
10. The pixel extraction and reconstruction method for simulating large-size SAR images of UAVs as described in claim 1, characterized in that, S10 specifically refers to: Convert the reconstructed canvas array into an image object; A scaling algorithm is used to adjust the image object to its original width and height; The naming convention for reconstructed images is set to "merge_ + original image name + image type"; The storage path for the reconstructed image is set to the parent directory of the sub-block storage directory.
Citation Information
Patent Citations
Edge priority guide single-frame remote sensing image super-resolution processing method
CN104732491A
Infrared image volume cloud detection method based on boundary fractal dimensions
CN109658429A
Field rice image automatic segmentation method based on generative adversarial network
CN118628743A
Apparatus and method for recovering image based on blocks
KR100803045B1