Self-adaptive real-time panoramic splicing system and method thereof

Through the adaptive real-time panoramic stitching system, combined with the dynamic blocking algorithm and the improved MobileNet-V3 network, the problems of poor real-time performance and weak adaptability of drone image stitching are solved, and efficient and accurate image stitching is achieved, which adapts to different scenes and lighting conditions and meets the real-time stitching needs of drones in different scenarios.

CN120672568APending Publication Date: 2025-09-19GUANGXI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510824417.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The existing image stitching technology has poor real-time performance, insufficient stitching accuracy, and weak adaptability, and cannot meet the needs of efficient and accurate stitching of drones in different scenes and lighting conditions.

Method used

An adaptive real-time panoramic stitching system is adopted, including a UAV flight control system and a ground computing end. Through a dynamic block image acquisition unit, an image matching unit and an image stitching unit, an improved MobileNet-V3 network and an adaptive stitching seam search strategy are used to realize an adaptive stitching system for image stitching. The adaptive stitching method is combined with the method of adaptive stitching, and is applied to the field of UAV technology, especially the field of UAV technology, especially the field of UAV technology, especially the field of UAV technology, especially the field of UAV technology, especially the field of UAV technology, especially the field of UAV technology, especially the field of UAV technology, especially the field of UAV technology, specifically involving the method of adaptive stitching, and is applied to UAV image processing.

Benefits of technology

It realizes real-time, efficient and accurate image stitching on drones, improves stitching quality and adaptability, reduces image transmission delay, and improves processing speed and stitching effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672568A_ABST
    Figure CN120672568A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a self-adaptive real-time panoramic stitching system and method, and the system comprises an unmanned aerial vehicle flight control system which is used for controlling the flight attitude and image collection of an unmanned aerial vehicle; the ground computing end is in communication connection with the unmanned aerial vehicle flight control system and is used for processing the acquired image and outputting a spliced image, and the ground computing end comprises an image acquisition unit which is used for acquiring an original image and compressing the acquired image into two paths of video streams; the image matching unit is used for dynamically partitioning the image and extracting feature points for feature matching; the image splicing unit is used for performing splicing processing after screening the matched feature points; and the image transmission unit is used for outputting the spliced image, and adaptively adjusting the partitioning size according to the image texture complexity by adopting a dynamic partitioning algorithm, so that the pertinence and efficiency of feature extraction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an adaptive real-time panoramic stitching system and method thereof, which are used for real-time stitching processing of images collected by unmanned aerial vehicles. Background Art

[0002] With the rapid development of drone technology, drone aerial photography has been widely used in surveying and mapping, surveillance, search and rescue, and other fields. In practical applications, single-view images often cannot meet the needs of large-scale scene monitoring and analysis, so multiple images need to be stitched together to obtain a panoramic view.

[0003] Traditional image stitching methods usually involve transmitting images back to a ground computing center for processing. This approach has many problems: first, the image transmission process occupies a large amount of bandwidth, resulting in poor real-time performance; second, the ground processing center has a heavy computational burden, making it difficult to meet the demand for rapid response; third, existing stitching algorithms usually use fixed parameters and cannot adapt to changes in different scenes and lighting conditions, affecting the stitching quality.

[0004] Therefore, how to achieve efficient and accurate image stitching has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The purpose of the invention is to provide an adaptive real-time panoramic stitching system and method thereof, aiming to solve the problems of poor real-time performance, insufficient stitching accuracy, and weak adaptability of image stitching in the existing technology.

[0006] The present invention proposes an adaptive real-time panoramic stitching system, comprising:

[0007] UAV flight control system, used to control the UAV flight attitude and image acquisition;

[0008] A ground computing terminal is connected to the UAV flight control system for processing the collected images and outputting a stitched image. The ground computing terminal includes:

[0009] An image acquisition unit, configured to acquire original images and compress the acquired images into two video streams;

[0010] Image matching unit, used to dynamically divide the image into blocks and extract feature points for feature matching;

[0011] An image stitching unit, used for screening the matched feature points and then performing stitching processing;

[0012] The image transmission unit is used to output the spliced ​​image.

[0013] Preferably, the image acquisition unit comprises:

[0014] Dual CMOS image sensors for capturing images at 1080p resolution;

[0015] An image encoder, connected to the dual-channel CMOS image sensor, for converting the collected images into H265 format;

[0016] A dual-channel encoder is connected to the image encoder and is used to output two H265 encoded streams to the image matching unit and the image splicing unit respectively.

[0017] Preferably, the image matching unit includes:

[0018] An image segmentation module is used to dynamically segment each frame of an image and adjust the size of the image blocks according to the texture complexity of the image blocks;

[0019] An ORB feature extraction module is connected to the image segmentation module and is used to extract features from the segmented images using an improved MobileNet-V3 network;

[0020] The ORB feature matching module is connected to the ORB feature extraction module and is used to perform feature matching on the extracted feature points in combination with the epipolar geometry principle.

[0021] Preferably, the calculation formula of the image block size S in the image segmentation module is:

[0022] ,

[0023] Among them, S is the current image block size, is the initial image block size, K is the number of points in the image whose grayscale value is greater than the threshold, is the number of points whose initial grayscale value is greater than the threshold, R is the maximum extreme value coordinate of the point whose grayscale value is greater than the threshold, The maximum extreme value coordinates of the point whose initial grayscale value is greater than the threshold.

[0024] Preferably, the ORB feature extraction module adopts an improved MobileNet-V3 network, including:

[0025] Two convolutional layers for preliminary feature extraction;

[0026] Seven bneck blocks, connected to the two convolutional layers, each bneck block includes a 3×3 convolution, a 1×1 convolution and an adaptive enhancement module;

[0027] A global average pooling layer, connected to the seven bneck blocks, is used to output the feature map as feature points;

[0028] Among them, the input resolution is 512×512, and the final output feature map size is 112×112 32-channel feature maps.

[0029] Preferably, the image stitching unit includes:

[0030] The splicing point screening module is used to screen matching points using an adaptive splicing seam search strategy and eliminate points with excessive matching errors;

[0031] The fusion module is connected to the splicing point screening module and is used to detect the illumination difference of the splicing result by using a phase consistency algorithm and process the result.

[0032] Preferably, the processing of the fusion module includes:

[0033] Calculate the image transformation matrix H for the matching point coordinates;

[0034] Performing coordinate transformation on the matching points according to the transformation matrix H;

[0035] Calculate the pixel difference between the mapping image and the target image. When the pixel difference is within [-T max ,T max ] is determined as a normal matching point;

[0036] After removing all mismatched points, the final splicing result is generated;

[0037] Among them, T max is the maximum pixel difference, which is determined by image quality, lighting conditions, camera model, target height and image block size.

[0038] Preferably, the system is implemented based on an FPGA development board and connected to an industrial camera via a USB3.0 interface. The UAV flight control system and the ground computing end exchange data through wireless transmission. The ground computing end also includes an FPGA timer module for detecting UAV status information and dynamically adjusting processing parameters.

[0039] Preferably, the improved MobileNet-V3 network introduces a channel separation residual mechanism in the concatenation of the Bottleneck residual module and the depthwise separable convolution, including:

[0040] Separate the two inputs;

[0041] Each input is normalized, activated, and convolved.

[0042] One input passes through a 3×1 convolution layer to obtain 2-channel features, and the other input passes through a 1×3 convolution layer to obtain 2-channel features;

[0043] The two features are spliced ​​together to form a multi-scale feature expression.

[0044] The adaptive panoramic stitching method based on edge computing adopts the system described above and includes the following steps:

[0045] Image acquisition step: the drone collects images and encodes and compresses the two-way data before transmitting them to the image matching unit and the image stitching unit respectively;

[0046] In the image matching step, each frame is dynamically divided into blocks, the block size is adaptively adjusted based on the texture complexity of the image blocks, and the image blocks are input into the improved MobileNet-V3 network to perform feature extraction and feature matching based on the principle of epipolar geometry;

[0047] In the image stitching step, an adaptive stitching seam search strategy is used to screen the matching results, calculate the transformation matrix, and apply the phase consistency algorithm to process the illumination differences in the stitching area to generate the stitched image;

[0048] In the image transmission step, the panoramic image stitched by the ground computing end is sent to the terminal via wireless transmission for display and post-processing.

[0049] This invention achieves real-time image processing and precise stitching by combining a dynamic block algorithm, an improved MobileNet-V3 network, and an adaptive seam search strategy with ground-based computing units. The invention has the following beneficial effects:

[0050] 1. Using a dynamic block algorithm, the block size is adaptively adjusted according to the complexity of the image texture, which improves the pertinence and efficiency of feature extraction;

[0051] 2. Adopting an improved MobileNet-V3 network structure, the lightweight design meets the resource constraints of edge computing while maintaining feature extraction capabilities;

[0052] 3. Through the adaptive seam search strategy and phase consistency algorithm, the stitching traces and illumination differences are effectively eliminated, and the stitching quality is improved;

[0053] 4. The various modules of the system work together to form a complete processing chain, adapt to the needs of different scenarios, and improve the adaptability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 Schematic diagram of the overall structure of the adaptive real-time panoramic stitching system of the present invention;

[0055] Figure 2 This is a schematic diagram of the functional module structure of the ground computing terminal of the present invention;

[0056] Figure 3Schematic diagram of the structure of the image matching unit of the present invention;

[0057] Figure 4 Schematic diagram of the structure of the image splicing unit of the present invention;

[0058] Figure 5 This is a schematic diagram of the improved MobileNet-V3 network structure of the present invention;

[0059] Figure 6 This is a flow chart of the adaptive seam search strategy of the present invention;

[0060] Figure 7 Schematic diagram of the dynamic block algorithm of the present invention;

[0061] Figure 8 This is a flow chart of the adaptive panoramic stitching method based on edge computing of the present invention;

[0062] Figure 9 This is a schematic diagram of the feature matching optimization process based on epipolar geometry of the present invention;

[0063] Figure 10 Schematic diagram of the phase consistency fusion algorithm principle of the present invention. DETAILED DESCRIPTION

[0064] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0065] like Figure 1 As shown, the adaptive real-time panoramic stitching system provided by the present invention includes a UAV flight control system 10 and a ground computing terminal 20. The UAV flight control system 10 is responsible for controlling the UAV's flight attitude and image acquisition, while the ground computing terminal 20 is responsible for processing the acquired images and outputting a stitched image.

[0066] Preferably, if Figure 2 As shown, the ground computing terminal 20 of the present invention includes an image acquisition unit 21 , an image matching unit 22 , an image stitching unit 23 and an image transmission unit 24 .

[0067] The image acquisition unit 21 is used to capture the original image and compress the captured image into two video streams. Specifically, the image acquisition unit 21 uses a CMOS image sensor to capture images at a resolution of 1080p, and compresses the captured image into two H265 video streams, which are then output to the image matching unit 22 and the image stitching unit 23 respectively.

[0068] like Figure 3As shown, the image matching unit 22 is used to dynamically block the image and extract feature points for feature matching. Specifically, the image matching unit 22 includes an image blocking module 221, an ORB feature extraction module 222, and an ORB feature matching module 223. The image blocking module 221 dynamically blocks each frame of the image and adjusts the image block size based on the texture complexity of the image block; the ORB feature extraction module 222 uses an improved MobileNet-V3 network to extract features from the blocked image; and the ORB feature matching module 223 combines the principles of epipolar geometry to perform feature matching on the extracted feature points.

[0069] In a specific embodiment of the present invention, the calculation formula of the image block size S in the image segmentation module 221 is:

[0070] ,

[0071] in, is the current image block size, is the initial image block size, usually set to Pixels; The number of points in the image whose grayscale value is greater than the threshold, the threshold is preferably set to 128 (the median of the grayscale range 0-255); is the number of points whose initial grayscale value is greater than the threshold, which is set to 30% of the total pixels of the image according to the empirical value; is the maximum extreme value coordinate of the point where the gray value is greater than the threshold value, The maximum extreme value coordinates of the point whose initial grayscale value is greater than the threshold.

[0072] This dynamic blocking algorithm, employed in this paper, increases the block size appropriately in areas with complex image textures (e.g., when the K value is large), thereby capturing more texture information. In areas with simpler textures, the block size is reduced accordingly, improving processing efficiency. Experiments have shown that this adaptive adjustment mechanism can reduce processing time by approximately 35% while maintaining feature extraction quality.

[0073] like Figure 5 As shown, the ORB feature extraction module 222 of the present invention uses an improved MobileNet-V3 network, including two convolutional layers, seven bneck blocks, a global average pooling layer, and three activation functions. The input resolution is 512×512, and the final output feature map size is 112×112 32-channel feature maps.

[0074] Specifically, the improved MobileNet-V3 network first performs preliminary feature extraction through two convolutional layers with a kernel size of 3×3×3, a stride of 2, and a Leaky_ReLU activation function. The input feature map then passes through seven bneck blocks, each consisting of a 3×3 convolution, a 1×1 convolution, and an adaptive enhancement module. Within each bneck block, a 3×3 convolution first reduces the number of input channels by half, and then the adaptive enhancement module quadruples the number of channels. In the adaptive enhancement module, shallow feature extraction is performed on the input feature map using a 1×1 convolution branch. This feature map is then decomposed into multiple channel branches. Each channel is superimposed with a 3×3 convolution branch and a 1×1 convolution branch. Finally, a channel-wise attention mechanism is employed, with a 1×1 convolution branch outputting dynamic weights for each channel.

[0075] Preferably, the present invention introduces a channel-separated residual mechanism in the combination of a Bottleneck residual module and a depthwise separable convolution. This involves separating the two inputs; normalizing each input, performing activation function processing, and performing depthwise convolution; passing one input through a 3×1 convolutional layer to obtain two-channel features, and the other input through a 1×3 convolutional layer to obtain two-channel features; and concatenating the two features to form a multi-scale feature representation. This design significantly improves the robustness of feature extraction and has greater adaptability to changes in illumination and viewing angle.

[0076] In particular, the ORB feature matching module 223 of the present invention adopts an optimized matching strategy based on epipolar geometry constraints, such as Figure 9 As shown. In traditional feature matching, the matching relationship is judged only by the Euclidean distance between descriptors, which is easily affected by noise and produces mismatches. The present invention significantly improves the matching accuracy by introducing epipolar geometry constraints. The specific implementation steps are as follows:

[0077] 1. First, use the Euclidean distance based preliminary matching on the feature points extracted from the two images to obtain a set of potential matching point pairs ,in and are the corresponding feature points in the two images respectively;

[0078] 2. Estimate the fundamental matrix using the RANSAC algorithm , which satisfies the epipolar geometry constraint equation:

[0079] ,

[0080] in, and Expressed in homogeneous coordinate form.

[0081] 3. Calculate the epipolar geometric consistency error of each pair of matching points:

[0082] ,

[0083] in, Represents a vector No. A portion.

[0084] 4. Set the epipolar geometry consistency threshold (Preferably set to 1.5 pixels in the experiment), remove the four matching points whose error is greater than the threshold:

[0085] ,

[0086] 5. To further improve the matching stability, the present invention introduces local geometric consistency verification. For any two pairs of matching points and , compute the affine consistency between them:

[0087] ,

[0088] in, represents the Euclidean distance between two points, Represents the absolute value of the angle between two vectors (3 radians), is the weight coefficient, and the preferred value is .

[0089] 6. If (Preferably set to 0.25 in the experiment), then the two pairs of matching points are considered to be consistent in local geometric relationship. For each pair of matching points, the number of matching points that are consistent with their local geometry is counted and recorded as .

[0090] 7. The final set of matching points retained is:

[0091] ,

[0092] in, is the support threshold, which is preferably set to 15% of the total number of matching points.

[0093] This optimization strategy significantly improves matching accuracy and robustness. Experiments show that compared to traditional methods, the false matching rate is reduced by approximately 70%, and a high matching success rate is maintained in the presence of large viewpoint changes, lighting differences, and partial occlusion.

[0094] like Figure 4As shown, the image stitching unit 23 is used to screen the matched feature points and then perform stitching. Specifically, the image stitching unit 23 includes a stitching point screening module 231 and a fusion module 232. The stitching point screening module 231 uses an adaptive stitching seam search strategy to screen matching points and eliminate points with excessive matching errors. The fusion module 232 uses a phase congruency algorithm to detect illumination differences in the stitching results and process the results.

[0095] In a preferred embodiment of the present invention, the processing process of the fusion module 232 includes: calculating the image transformation matrix H for the coordinates of the matching points; performing coordinate transformation on the matching points according to the transformation matrix H; calculating the pixel difference between the mapped image and the target image, and when the pixel difference is within [-T max ,T max ] is determined as a normal matching point; after eliminating all incorrect matching points, the final splicing result is generated. max is the maximum pixel difference, which is determined by image quality, lighting conditions, camera model, target height and image block size.

[0096] Specifically, T max The calculation formula is:

[0097] ,

[0098] in, is the maximum pixel difference, For image quality, , usually based on image clarity and contrast evaluation, preferably set to 0.8; is the lighting condition level, with a sunny day value of 1.0, a cloudy day value of 0.9, and a night value of 0.7; is the standard height of the target, is the actual height of the target; is the image block size, is the distance between the feature point and the center point, is the maximum block size.

[0099] In particular, the fusion module 232 of the present invention adopts an image fusion algorithm based on phase consistency, such as Figure 10 As shown in Figure 2, this algorithm provides a high-quality solution to the problem of inconsistent illumination in the stitching area. Traditional image fusion often uses simple weighted averaging or gradient domain fusion, which cannot effectively handle illumination differences in complex scenes. The phase consistency fusion algorithm proposed in this paper achieves seamless and natural fusion by analyzing the phase information of local images. The specific implementation is as follows:

[0100] 1. First, locate the splicing area and determine the overlapping area , whose boundaries are

[0101] 2. For the two images to be fused and , calculate the phase consistency map in the overlapping area

[0102] ,

[0103] in, Is in the direction and scale The amplitude under , is calculated as:

[0104] ,

[0105] is an image In the direction and scale The phase under , is calculated as:

[0106] ,

[0107] here, and They are images Through the scale of , direction is In this embodiment, it is preferred to use a Gabor filter bank with 6 directions (0°, 30°, 60°, 90°, 120°, 150°) and 4 scales.

[0108] 3. Based on phase consistency diagram , construct the fusion weight graph

[0109] ,

[0110] in, is the average value of phase consistency, Controls the smoothness of the transition, preferably set to 0.2.

[0111] 4. To ensure a smooth transition of the fusion boundary, the distance field modulation function is introduced

[0112] ,

[0113] in, Yes To the overlapping area boundary The shortest distance, Control the width of the transition zone, preferably set to 20% of the width of the overlapping area.

[0114] The final fusion weight is:

[0115] ,

[0116] in, is the weight at the boundary. To ensure a smooth transition at the boundary, linear interpolation is usually used.

[0117] 6. Final fused image for:

[0118] ,

[0119] This algorithm demonstrates significant advantages when stitching scenes with significant lighting variations. Experimental results show that, compared to the traditional weighted averaging method, the subjective visual effect is improved by approximately 40%, and the objective evaluation metrics of PSNR and SSIM are improved by 3.2dB and 15%, respectively. This algorithm is particularly effective in stitching together images captured by drones at high altitudes at different times and under different lighting conditions, effectively eliminating stitching artifacts and achieving a natural transition.

[0120] The present invention uses an adaptive threshold setting mechanism to dynamically adjust the matching criteria according to the actual scene conditions to improve the accuracy of stitching. For example, when the lighting conditions are poor (K value is small), the system will lower the matching criteria to ensure a sufficient number of matching points; when processing high-definition images ( If the matching accuracy is higher, the system will increase the matching accuracy requirements to ensure the stitching quality.

[0121] The image transmission unit 24 is used to output the stitched image and transmit the stitched panoramic image from the ground computing end to the terminal via wireless transmission for display and post-processing.

[0122] The system of the present invention is preferably implemented on an FPGA development board, connected to an industrial camera via a USB 3.0 interface. The drone flight control system 10 and the ground computing terminal 20 exchange data via wired or wireless transmission. Furthermore, the ground computing terminal 20 includes an FPGA timer module for detecting drone status information and dynamically adjusting processing parameters.

[0123] By using FPGA implementation, the present invention can fully utilize hardware acceleration capabilities to improve the real-time performance of the system. Experiments have shown that the FPGA-based implementation is about five times faster than the pure software implementation, and can meet the real-time processing requirements of more than 25 frames per second.

[0124] The following combination Figure 8 , the adaptive panoramic stitching method based on edge computing provided by the present invention is described in detail.

[0125] The method comprises the following steps:

[0126] Step S1: Image acquisition,

[0127] The drone captures images, and the ground computing end encodes and compresses the two data streams before transmitting them to the image matching unit and image stitching unit, respectively. Specifically, dual CMOS image sensors capture images at 1080p resolution and compress them using the H.265 encoding format. One stream of data is used for feature extraction and matching, while the other is used for the final stitching process.

[0128] Step S2: Image matching: Each frame image is dynamically divided into blocks, the block size is adaptively adjusted according to the texture complexity of the image block, and the image block is input into the improved MobileNet-V3 network to perform feature extraction and feature matching based on the epipolar geometry principle.

[0129] Specifically, the image is first segmented using the aforementioned dynamic segmentation algorithm. For these segmented images, ORB feature points are extracted using a modified MobileNet-V3 network. Next, based on epipolar geometry, the RANSAC algorithm is used for feature matching, removing anomalous matching points. Furthermore, the aforementioned local geometric consistency verification method is employed to further improve matching stability and accuracy.

[0130] Step S3: Image stitching, using an adaptive stitching seam search strategy to screen the matching results, calculate the transformation matrix, apply the phase consistency algorithm to process the illumination difference in the stitching area, and generate a stitched image.

[0131] Specifically, the image transformation matrix H is first calculated for the coordinates of the matching points, and the coordinates of the matching points are transformed. The pixel difference between the mapped image and the target image is calculated, and the validity of the matching points is determined based on the aforementioned Tmax threshold. Then, the aforementioned phase congruence fusion algorithm is applied to analyze the phase information of the local images and calculate the fusion weights to achieve a smooth transition between the stitching areas, ultimately forming a seamless panoramic image.

[0132] Step S4: Image transmission: The panoramic image stitched by the ground computing end is sent to the terminal via wireless transmission for display and post-processing.

[0133] The method of the present invention achieves real-time image processing and precise stitching through an edge computing architecture. Compared to traditional methods, this method offers the following advantages: First, by performing processing on the drone end, it avoids the transmission of large amounts of raw images and reduces transmission latency. Second, a dynamic block segmentation algorithm and an improved MobileNet-V3 network improve feature extraction efficiency and accuracy. Finally, an optimized matching strategy based on epipolar geometry constraints and a phase consistency fusion algorithm address key challenges in feature matching and image fusion, respectively, effectively improving stitching quality.

[0134] Experimental results show that, under the same hardware conditions, the proposed method achieves approximately 40% faster processing speed than traditional methods, improves stitching quality scores (assessed using PSNR and SSIM metrics) by approximately 25%, and significantly enhances adaptability to diverse scenes and lighting conditions. In particular, the proposed method maintains stable stitching results under conditions of significant lighting variations and viewing angle differences, thanks to the synergistic effect of an optimized matching strategy based on epipolar geometry constraints and a phase congruency fusion algorithm.

[0135] Through the above detailed description, those skilled in the art can clearly understand the implementation mode and technical effects of the present invention. It should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. The scope of protection of the present invention is subject to the claims.

[0136] It should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. Adaptive real-time panoramic stitching system, characterized by: include: UAV flight control system, used to control the UAV flight attitude and image acquisition; A ground computing terminal is connected to the UAV flight control system for processing the collected images and outputting a stitched image. The ground computing terminal includes: An image acquisition unit, configured to acquire original images and compress the acquired images into two video streams; Image matching unit, used to dynamically divide the image into blocks and extract feature points for feature matching; An image stitching unit, used for screening the matched feature points and then performing stitching processing; The image transmission unit is used to output the spliced ​​image.

2. The adaptive real-time panoramic stitching system according to claim 1, characterized in that: The image acquisition unit includes: Dual CMOS image sensors for capturing images at 1080p resolution; An image encoder, connected to the dual-channel CMOS image sensor, for converting the collected images into H265 format; A dual-channel encoder is connected to the image encoder and is used to output two H265 encoded streams to the image matching unit and the image splicing unit respectively.

3. The adaptive real-time panoramic stitching system according to claim 1, characterized in that: The image matching unit includes: An image segmentation module is used to dynamically segment each frame of an image and adjust the size of the image blocks according to the texture complexity of the image blocks; An ORB feature extraction module is connected to the image segmentation module and is used to extract features from the segmented images using an improved MobileNet-V3 network; The ORB feature matching module is connected to the ORB feature extraction module and is used to perform feature matching on the extracted feature points in combination with the epipolar geometry principle.

4. The adaptive real-time panoramic stitching system according to claim 3, characterized in that: The calculation formula of the image block size S in the image segmentation module is: , Among them, S is the current image block size, is the initial image block size, K is the number of points in the image whose grayscale value is greater than the threshold, is the number of points whose initial grayscale value is greater than the threshold, R is the maximum extreme value coordinate of the point whose grayscale value is greater than the threshold, The maximum extreme value coordinates of the point whose initial grayscale value is greater than the threshold.

5. The adaptive real-time panoramic stitching system according to claim 3, characterized in that: The ORB feature extraction module uses an improved MobileNet-V3 network, including: Two convolutional layers for preliminary feature extraction; Seven bneck blocks, connected to the two convolutional layers, each bneck block includes a 3×3 convolution, a 1×1 convolution and an adaptive enhancement module; A global average pooling layer, connected to the seven bneck blocks, is used to output the feature map as feature points; Among them, the input resolution is 512×512, and the final output feature map size is 112×112 32-channel feature maps.

6. The adaptive real-time panoramic stitching system according to claim 1, characterized in that: The image stitching unit includes: The splicing point screening module is used to screen matching points using an adaptive splicing seam search strategy and eliminate points with excessive matching errors; The fusion module is connected to the splicing point screening module and is used to detect the illumination difference of the splicing result by using a phase consistency algorithm and process the result.

7. The adaptive real-time panoramic stitching system according to claim 6, characterized in that: The processing process of the fusion module includes: Calculate the image transformation matrix H for the matching point coordinates; Performing coordinate transformation on the matching points according to the transformation matrix H; Calculate the pixel difference between the mapping image and the target image. When the pixel difference is within [-T max ,T max ] is determined as a normal matching point; After removing all mismatched points, the final splicing result is generated; Among them, T max is the maximum pixel difference, which is determined by image quality, lighting conditions, camera model, target height and image block size.

8. The adaptive real-time panoramic stitching system according to claim 1, characterized in that: The system is implemented based on an FPGA development board and connected to an industrial camera via a USB3.0 interface. The UAV flight control system and the ground computing end exchange data through wireless transmission. The ground computing end also includes an FPGA timer module for detecting UAV status information and dynamically adjusting processing parameters.

9. The adaptive real-time panoramic stitching system according to claim 3, characterized in that: The improved MobileNet-V3 network introduces a channel-separated residual mechanism in the concatenation of the Bottleneck residual module and the depthwise separable convolution, including: Separately process the two inputs; Each input is normalized, activated, and convolved. One input passes through a 3×1 convolution layer to obtain 2-channel features, and the other input passes through a 1×3 convolution layer to obtain 2-channel features; The two features are spliced ​​together to form a multi-scale feature expression.

10. An adaptive panoramic stitching method based on edge computing, using the system according to any one of claims 1 to 9, characterized in that: The following steps are involved: Image acquisition step: the drone collects images and encodes and compresses the two-way data before transmitting them to the image matching unit and the image stitching unit respectively; In the image matching step, each frame is dynamically divided into blocks, the block size is adaptively adjusted based on the texture complexity of the image blocks, and the image blocks are input into the improved MobileNet-V3 network to perform feature extraction and feature matching based on the principle of epipolar geometry; In the image stitching step, an adaptive stitching seam search strategy is used to screen the matching results, calculate the transformation matrix, and apply the phase consistency algorithm to process the illumination differences in the stitching area to generate the stitched image; In the image transmission step, the panoramic image stitched by the ground computing end is sent to the terminal via wireless transmission for display and post-processing.