Distortion correction method and device for fisheye image
By using controllable convolutional modulation blocks and controllable attention modulation blocks in fisheye image distortion correction, control conditions are dynamically generated to adapt to different distortion degrees, solving the problem of insufficient generalization ability for different distortion conditions in the prior art, and achieving high-precision and flexibility distortion correction effects.
Patent Information
- Application Number
- CN202510250291.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art lacks the ability to generalize different distortion conditions in fisheye image distortion correction, resulting in low distortion correction accuracy.
The controllable convolutional modulation block and the controllable attention modulation block are used to construct a hierarchical correction network. By querying vectors, the control conditions are dynamically generated, different distortion degrees are adapted, and multi-level feature modulation is performed in the network.
It improves the accuracy and flexibility of fisheye image distortion correction, enhances the modeling ability of complex distortion modes, and can show good adaptability under different distortion conditions.
Smart Images

Figure CN120182145A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to a method and device for correcting distortion of fish-eye images. Background Art
[0002] With the rapid development of computer vision and image processing technologies, fish-eye lenses have been widely used in many fields due to their wide-angle coverage characteristics, such as autonomous driving, video surveillance, virtual reality, and drone navigation. However, while fish-eye lenses capture large-range images, due to their non-linear optical structure, significant geometric distortion problems are often introduced. This distortion not only affects the visual effect of the image but may also have an adverse impact on the performance of subsequent tasks (such as object detection and scene understanding).
[0003] In recent years, with the rapid development of deep learning technologies, data-driven image processing methods have begun to be applied to the task of correcting distortion of fish-eye images. Such methods usually learn the mapping relationship from a large number of fish-eye images and their corresponding corrected images through an end-to-end learning method. However, although such methods perform well under fixed distortion degrees, they generally lack the generalization ability for different distortion conditions, resulting in low accuracy in correcting the distortion of fish-eye images. Summary of the Invention
[0004] Based on this, it is necessary to provide a method and device for correcting distortion of fish-eye images for the above technical problems. This method has a high generalization ability for different distortion conditions and can improve the accuracy of correcting the distortion of fish-eye images.
[0005] The present invention adopts the following technical solutions:
[0006] The present invention provides a method for correcting distortion of fish-eye images, including:
[0007] Constructing a hierarchical correction network by combining a controllable convolution modulation block and a controllable attention modulation block; the controllable convolution modulation block is used to extract local texture details of features, and the controllable attention modulation block can capture global spatial dependence relationships using the attention mechanism;
[0008] Obtaining a query vector of the fish-eye image to be corrected from a query vector set according to the distortion degree of the fish-eye image to be corrected; the query vector set includes multiple query vectors, and each query vector corresponds to a distortion degree of an image;
[0009] Determining control conditions of the fish-eye image to be corrected at multiple levels according to the query vector of the fish-eye image to be corrected; the control conditions at each level correspond to each modulation block in the hierarchical correction network in sequence;
[0010] Input the original features of the fish-eye image to be corrected and the control conditions at each level of the fish-eye image to be corrected into the hierarchical correction network. Use the controllable convolution modulation block to fuse the original features and the corresponding control conditions, and in the controllable attention modulation block, globally enhance the features output by the controllable convolution modulation block through the control conditions to obtain the distortion-corrected image of the fish-eye image to be corrected.
[0011] Optionally, determine the control conditions of the fish-eye image to be corrected at multiple levels according to the query vector, including:
[0012] Input the query vector into the pre-trained learnable query network for distortion perception to extract the control conditions of the fish-eye image to be corrected at multiple levels.
[0013] Optionally, the learnable query network includes a feature extraction module and a multi-layer perceptron; extracting the control conditions of the fish-eye image to be corrected at multiple levels includes:
[0014] Convert the query vector into a low-dimensional latent feature through the feature extraction module;
[0015] In the first level of the learnable query network, process the low-dimensional latent feature through the multi-layer perceptron to obtain the control condition of the first level;
[0016] In the second level to the last level of the learnable query network, process the control condition generated by the previous level through the multi-layer perceptron to obtain the control condition of the corresponding level.
[0017] Optionally, the hierarchical correction network includes a plurality of first controllable convolution modulation blocks, a plurality of controllable attention modulation blocks, and a plurality of second controllable convolution modulation blocks connected in series in sequence. The control condition of each level corresponds to a modulation block in sequence; both the first controllable convolution modulation block and the second controllable convolution modulation block are controllable convolution modulation blocks; using the controllable convolution modulation block to fuse the original features and the corresponding control conditions, and in the controllable attention modulation block, globally enhance the features output by the controllable convolution modulation block through the control conditions to obtain the distortion-corrected image of the fish-eye image to be corrected, including:
[0018] Pass the original features and the corresponding control conditions through a plurality of first controllable convolution modulation blocks, a plurality of controllable attention modulation blocks, and a plurality of second controllable convolution modulation blocks in sequence to obtain the distortion-corrected image of the fish-eye image to be corrected; starting from the second modulation block, the features input to the subsequent modulation block are all generated according to the output features of the previous modulation block.
[0019] Optionally, there is a symmetric association between the input and output of the modulation block in the hierarchical correction network. The method for obtaining the features input to the last second controllable convolution modulation block includes:
[0020] Optical flow estimation is performed on the original features to obtain the optical flow field of the original features;
[0021] The output features of the first first controllable convolutional modulation block are corrected through the optical flow field;
[0022] The corrected features and the output features of the previous modulation block of the last second controllable convolutional modulation block are concatenated in channels to obtain fused features, and the fused features are used as the input of the last second controllable convolutional modulation block;
[0023] Correspondingly, the features input to the penultimate second controllable convolutional modulation block are generated based on the corrected features, the output features of the second first controllable convolutional modulation block, and the output features of the previous modulation block of the penultimate second controllable convolutional modulation block.
[0024] Optionally, the controllable convolutional modulation block includes a first channel connection unit, a predicted fusion ratio predictor, a first control unit, a feature modulation unit, and a first feed-forward neural network; the implementation process of the controllable convolutional modulation block includes:
[0025] The input features and the corresponding control conditions are concatenated in channels through the first channel connection unit to obtain concatenated features;
[0026] The concatenated features are analyzed by the predicted fusion ratio predictor to obtain the predicted fusion ratio;
[0027] The input features and the control conditions are multiplied by the first control unit to obtain control features;
[0028] The predicted fusion ratio, the input features, and the control features are input to the feature modulation unit to obtain modulated features;
[0029] The modulated features are input to the first feed-forward neural network to obtain the output features of the controllable convolutional modulation block. Optionally, the implementation manner of the feature modulation unit is:
[0030] F out = θF c +(1 - θ)F in ;
[0031] Among them, F out represents the output features, θ represents the predicted fusion ratio, F c represents the control features, and F in represents the input features.
[0032] Optionally, the controllable attention modulation block includes a second control unit, a first projection unit, a second projection unit, an attention unit, a second channel connection unit, and a second feed-forward neural network. The implementation process of the controllable attention modulation block includes:
[0033] Multiply the input feature by the corresponding control condition through the second control unit to obtain a control feature;
[0034] Multiply the first projection matrix by the control feature through the first projection unit to obtain a query matrix;
[0035] Multiply the second projection matrix by the input feature through the second projection unit to obtain a key matrix, multiply the third projection matrix by the input feature to obtain a value matrix, and multiply the fourth projection matrix by the input feature to obtain the input feature after projection transformation;
[0036] Input the query matrix, key matrix, and value matrix into the attention unit to obtain control attention;
[0037] Connect the control attention with the input feature after projection transformation through the second channel connection unit to obtain a connection feature;
[0038] Perform residual connection on the connection feature through the second feed-forward neural network to obtain the output feature of the controllable attention modulation block.
[0039] Optionally, the implementation manner of the attention unit is:
[0040]
[0041] where CTRL-ATTN(Q, K, V) represents control attention, Q represents the query matrix, K represents the key matrix, m represents the sequence length of the matrix, and V represents the value matrix.
[0042] The present invention provides a distortion correction device for a fish-eye image, including:
[0043] A construction module for constructing a hierarchical correction network by combining a controllable convolution modulation block and a controllable attention modulation block; the controllable convolution modulation block is used to extract local texture details of features, and the controllable attention modulation block can capture global spatial dependencies using the attention mechanism;
[0044] An acquisition module for obtaining a query vector of the fish-eye image to be corrected from a query vector set according to the distortion degree of the fish-eye image to be corrected; the query vector set includes multiple query vectors, and each query vector corresponds to a distortion degree of an image;
[0045] A first determination module for determining control conditions of the fish-eye image to be corrected at multiple levels according to the query vector of the fish-eye image to be corrected; the control conditions at each level correspond to each modulation block in the hierarchical correction network in sequence;
[0046] A second determination module, configured to input the original features of the fish-eye image to be corrected and the control conditions at each level of the fish-eye image to be corrected into a hierarchical correction network, fuse the original features and the corresponding control conditions by using a controllable convolution modulation block, and in the controllable attention modulation block, globally enhance the features output by the controllable convolution modulation block through the control conditions to obtain a distortion-corrected image of the fish-eye image to be corrected.
[0047] The present invention provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-mentioned method for correcting the distortion of a fish-eye image.
[0048] The present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the above-mentioned method for correcting the distortion of a fish-eye image when executing the program.
[0049] The above at least one technical solution adopted by the present invention can achieve the following beneficial effects:
[0050] In the present invention, the query vector represents the distortion degree of the fish-eye image to be corrected, the query vector of the fish-eye image to be corrected is determined from a preset query vector set, and the control conditions of the fish-eye image to be corrected at multiple levels are determined according to the query vector. The control conditions at each level correspond to each modulation block in the hierarchical correction network in sequence; this is equivalent to dynamically adjusting the working modes of the modulation blocks at each level according to the actual distortion situation of the fish-eye image to be corrected. Different distortion degrees correspond to different control conditions, and each level of modulation block can make corresponding adaptive changes. This mechanism of flexible adjustment according to distortion enables the hierarchical correction network to well handle various distortion situations and improves the distortion correction accuracy; further, the hierarchical correction network includes a controllable convolution modulation block and a controllable attention modulation block. The controllable convolution modulation block extracts the local texture details of the features, can dynamically fuse the original features and the control conditions, and provides a good foundation for subsequent correction work. The controllable attention modulation block captures the global spatial dependence relationship, analyzes the connection between different regions from a global perspective, improves the modeling ability for complex distortions, and further improves the adaptability to different distortion situations, thereby improving the distortion correction accuracy of the fish-eye image to be corrected. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The illustrative embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0052] Figure 1 It is a schematic flowchart of a method for correcting the distortion of a fish-eye image provided by the present invention;
[0053] Figure 2 Schematic diagram of the structure of a hierarchical correction network provided by the present invention;
[0054] Figure 3 Schematic diagram of the implementation principle of a controllable convolution modulation block provided by the present invention;
[0055] Figure 4 Schematic diagram of the implementation principle of a controllable attention modulation block provided by the present invention;
[0056] Figure 5 Schematic diagram of an existing method and the method of the present invention provided by the present invention;
[0057] Figure 6 Schematic diagram of the correction result in a real scenario provided by the present invention;
[0058] Figure 7 Schematic diagram of the correction result on a public dataset provided by the present invention;
[0059] Figure 8 Schematic diagram of a fisheye image distortion correction device provided by the present invention;
[0060] Figure 9 Schematic diagram of a computer device for implementing a method for correcting fisheye image distortion provided by the present invention. Detailed implementation manners
[0061] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0062] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that an article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed. Without further limitations, the element defined by the statement "including one..." does not exclude the existence of another identical element in the article or device including the said element.
[0063] In recent years, with the rapid development of deep learning technology, data-driven image processing methods have begun to be applied to the task of fisheye image distortion correction. These methods usually learn the mapping relationship from a large number of fisheye images and their corresponding corrected images through an end-to-end learning method. For example, the regression method based on convolutional neural network directly predicts pixel transformation parameters or the corrected image, avoiding the cumbersome calibration process in traditional methods; the method based on Generative Adversarial Network (GAN) improves the image quality while ensuring the correction effect through the adversarial training of the generator and discriminator. These methods perform well under fixed distortion degrees, but generally lack the generalization ability for different distortion conditions, and users cannot flexibly adjust the correction effect. In addition, the computational cost of deep learning models is relatively high, especially in generative adversarial networks, which require high-quality training data and are prone to problems such as overfitting or blurred correction results. Therefore, how to design a correction model with efficient control ability, strong generalization ability and suitable for resource-constrained environments has become the key research difficulty.
[0064] The rise of controllable image processing tasks provides a new idea for solving the above problems. By introducing additional control conditions, users can dynamically adjust the specific effects of image processing. For example, in controllable image denoising, users can adjust the denoising intensity according to the noise level; in controllable image enhancement, users can adjust the brightness or contrast according to the enhancement target. These methods improve the flexibility and interpretability of the algorithm through the control mechanism. However, introducing the controllable mechanism into the fisheye distortion correction task still faces significant challenges, including how to design control conditions that accurately characterize the degree of fisheye distortion, how to adapt the control conditions to complex fisheye distortion patterns, and how to meet the real-time requirements in dynamic scenarios. Solving these problems requires breakthroughs in control condition design, network architecture optimization, and real-time computational efficiency.
[0065] To sum up, traditional methods and existing deep learning methods both have deficiencies in terms of flexibility, generalization ability, and real-time performance in dealing with fisheye image distortion correction. Therefore, there is an urgent need to design a new type of fisheye image distortion correction method that can adapt to different distortion degrees by adjusting control conditions, show good generalization ability in complex scenarios, and achieve real-time and efficient correction effects in resource-constrained environments.
[0066] Based on this, the present invention proposes a controllable distortion correction method for fisheye images based on a query vector set, aiming to solve the problems in traditional fisheye image correction methods, such as poor adaptability to a single distortion degree, insufficient control flexibility, and limited generalization ability. The innovation of this method lies in the introduction of a learnable query vector mechanism and the construction of a hierarchical correction network by combining two types of controllable modulation blocks, achieving flexible and efficient correction for various distortion degrees. Specifically, the method of the present invention learns the mapping relationship between the query vector and the control condition, which can not only achieve accurate distortion correction but also generate continuous correction effects through interpolation, thus flexibly coping with the fisheye image correction tasks in multiple scenarios and with multiple requirements.
[0067] The core idea of the present invention is to represent different distortion control conditions through a query vector set and transform the correction problem of non-linear optical distortion in fisheye images into a controllable learning problem of a deep neural network. First, by inputting a query vector representing the distortion condition, the user can dynamically generate a control condition to guide the correction process. Second, the designed Controllable Convolutional Modulation Block (CCMB) and Controllable Attention Modulation Block (CAMB) model the local details and global spatial relationships at different scales of the network respectively, ensuring the quality and robustness of the correction result. In addition, the two-stage training strategy (coarse-grained pre-training and fine-grained fine-tuning) proposed by the present invention further improves the network's adaptability to data with multiple distortion degrees and significantly reduces the dependence on calibration data.
[0068] Through the detailed description of the following specific implementation manners, a deeper understanding of the actual effects of the present invention in different application scenarios can be obtained. For example, the cameras in the autonomous driving scenario often face the distortion problem caused by an overly wide viewing angle, and the method proposed by the present invention can not only efficiently correct the distortion but also achieve a balance between real-time performance and accuracy. In addition, by combining the drawings to illustrate the network structure, query control mechanism, and specific implementation of the controllable modulation module of the present invention, the technical advantages of the present invention can be clearly presented.
[0069] The distortion correction method for fisheye images provided by the present invention can be applied to a server, which can be a server set up in a business platform or a device such as a desktop computer or a laptop computer that can execute the solution of the present invention. As Figure 1 shown, Figure 1 is a schematic flowchart of a distortion correction method for fisheye images in the present invention, which specifically includes the following steps:
[0070] S101. Construct a hierarchical correction network by combining a controllable convolution modulation block and a controllable attention modulation block. The controllable convolution modulation block is used to extract local texture details of features, and the controllable attention modulation block can capture global spatial dependencies using the attention mechanism.
[0071] The present invention first proposes a Distortion-Aware Learnable Query Mechanism (DLQM). DLQM represents different distortion control conditions of fish-eye images by constructing a series of learnable query vectors. These query vectors can flexibly capture changes in the distortion degree. Users can generate corresponding correction results by selecting or adjusting the query vectors, avoiding the cumbersome process of re-calibrating or re-training the model in traditional methods. Through the query feature extraction module \(CE(\cdot)\), DLQM maps high-dimensional query vectors to a low-dimensional latent space and generates control conditions for guiding the correction operations of each layer of the network. Users can also generate continuous control conditions through interpolation between query vectors, thereby achieving dynamic adjustment and precise control of the image correction effect. Secondly, to enable the control conditions to effectively guide the correction process, the present invention designs two types of controllable modulation blocks: a controllable convolution modulation block CCMB based on a convolutional neural network and a controllable attention modulation block CAMB based on the deep learning architecture Transformer. CCMB is mainly used to extract local texture details and can dynamically fuse the original features and control features to retain the local texture. CAMB uses the attention mechanism to capture global spatial dependencies and improve the modeling ability for complex distortion patterns. By combining CCMB and CAMB, a hierarchical correction network is constructed, which can achieve fine-grained modulation of features at different scales.
[0072] Secondly, the CCMB module realizes the retention of local texture details by dynamically fusing the input features and control features, and at the same time adaptively adjusts the local correction intensity through the control conditions, ensuring precise correction of the local area. In contrast, the CAMB module constructs a global attention mechanism using the control conditions to capture the complex global spatial dependencies in the fish-eye image to be corrected, thereby improving the modeling ability of the hierarchical correction network for complex distortion patterns. These two types of modules play roles at different scales in the network: CCMB is suitable for feature maps with a larger resolution and is mainly used to retain local image details; CAMB is suitable for feature maps with a smaller resolution and is mainly used to model global dependencies. By combining these two types of modules, a hierarchical correction network is constructed, which can achieve multi-level fine-grained modulation of image features at different scales.
[0073] In addition, the present invention also proposes a two-stage training strategy, which further enhances the generalization ability and robustness of the hierarchical correction network. In the first-stage coarse-grained pre-training, the hierarchical correction network is trained on a dataset with a single distortion degree to learn the basic image correction ability. In the second-stage fine-grained fine-tuning, the hierarchical correction network is trained on a dataset containing various distortion degrees, and combined with the control conditions generated by DLQM, the applicability of the network in different scenarios and distortion modes is extended. This training strategy not only reduces the dependence on diverse training data, but also significantly accelerates the convergence of the model and improves the stability of the correction effect.
[0074] Coarse-grained pre-training: The training strategy is an important means to improve the generalization ability of the model in this method. First, in the coarse-grained pre-training stage, the network is trained on a dataset with a single distortion degree, and by optimizing the reconstruction loss L_r and the multi-scale loss L_m, it learns the basic correction ability:
[0075] Objective: Learn the basic correction ability of the model.
[0076] Loss function:
[0077] L pre = L r + L m , (1)
[0078] where, L pre is the value of the loss function, L r is the reconstruction loss, and L m is the multi-scale loss.
[0079] Fine-grained fine-tuning: In the fine-grained fine-tuning stage, the network is trained through a dataset containing various distortion degrees, and at the same time, diverse control conditions are generated by using query vectors to further improve the adaptability to unseen scenarios. This two-stage strategy reduces the dependence of the model on the amount of data and accelerates the convergence of the model. Objective: Improve the generalization ability of the model to different distortion degrees. Method: Fine-tune the model on a dataset containing various distortion degrees, and use the query vector set Q s to generate corresponding control conditions.
[0080] S102, according to the distortion degree of the fish-eye image to be corrected, obtain the query vector of the fish-eye image to be corrected from the query vector set; the query vector set includes multiple query vectors, and each query vector corresponds to a distortion degree of the image.
[0081] Each query vector in the query vector set represents a specific distortion control condition. Constructing the query vector set is one of the core innovations of this method, which is used to achieve flexible control of different distortion degrees. Traditional image correction methods usually train a fixed model for specific distortion parameters. When encountering new distortion conditions, the lens needs to be recalibrated or the model needs to be retrained. By constructing the query vector set, the degree and characteristics of fisheye distortion can be described in a data-driven manner. Specifically, the query vector set can represent a set of control conditions, and each query vector Q i corresponds to a specific distortion degree. Query vectors can not only represent discrete distortion levels, but also generate continuous control conditions through interpolation, so as to achieve fine-grained adjustment of the correction results. This mechanism not only improves the flexibility of the model, but also reduces the need for a large amount of training data. The initialization of query vectors during training can be carried out in various ways, for example:
[0082] (1) Random initialization: Generate initial values through Gaussian distribution or uniform distribution;
[0083] (2) Pre-training initialization: Generate query vectors using pre-calibrated distortion data;
[0084] (3) Manual design: Set manually according to prior knowledge of distortion patterns.
[0085] Optionally, the query vector of the fisheye image to be corrected can be set manually according to the prior knowledge of the distortion pattern of the fisheye image, or can be obtained from the query vector set; optionally, the user selects the query vector corresponding to the distortion degree from the query vector set according to the distortion degree of the fisheye image to be corrected, so as to achieve the correction of the corresponding degree.
[0086] In an exemplary embodiment, the query vector of the fisheye image to be corrected is determined by interpolating the vectors in the query vector set.
[0087] For example, in practical applications, the distortion degree of the fisheye lens may not always match the query vectors during training exactly. For example, the query vector set may only contain two discrete distortion conditions, Q8 and Q9. And the distortion degree of the fisheye image input by the user may be between the two, such as 8.5. DLQM generates continuous control conditions through the interpolation mechanism, for example: Q 8.5 = 0.5Q8 + 0.5Q9.
[0088] This interpolation method can not only generate new control conditions, but also maintain the continuity and interpretability of the control conditions, so as to flexibly meet the diverse correction requirements in actual scenarios.
[0089] The advantages of the interpolation mechanism include: flexible adaptation to unseen data: users can process new distortion conditions without recalibrating or retraining the model; and smooth transition effect: continuous control conditions can ensure the coherence of the correction results and avoid obvious switching traces.
[0090] Through the dynamic generation of query vectors, the invention can quickly adapt to various distortion scenarios without retraining, significantly improving the flexibility of the system. Users can adjust the query vectors or interpolate to generate new control conditions to achieve precise control of the correction process. For example, various correction results of different degrees can be generated according to actual needs to facilitate the selection of the optimal effect. Traditional methods can often only handle the distortion conditions seen during training, while the query interpolation mechanism expands the applicable range of the network, enabling it to adapt to unseen distortion scenarios. Through the position encoding mechanism, the invention can capture complex spatial distortion patterns in the image, thereby more accurately restoring the true structure of the image. After the query vectors are dimension-reduced by the feature extraction network, the computational amount is significantly reduced, and at the same time, the resource utilization of the network is optimized through the dynamic generation of hierarchical control conditions.
[0091] S103. Determine the control conditions of the fish-eye image to be corrected at multiple levels according to the query vector of the fish-eye image to be corrected; the control conditions at each level correspond to each modulation block in the hierarchical correction network in sequence.
[0092] Determine the control conditions of the fish-eye image to be corrected at multiple levels according to the query vector, including: inputting the query vector into a pre-trained learnable query network for distortion perception to extract the control conditions of the fish-eye image to be corrected at multiple levels.
[0093] The learnable query network includes a feature extraction module and a multi-layer perceptron; extracting the control conditions of the fish-eye image to be corrected at multiple levels includes: converting the query vector into a low-dimensional latent feature through the feature extraction module; in the first level of the learnable query network, processing the low-dimensional latent feature through the multi-layer perceptron to obtain the control conditions of the first level; in the second level to the last level of the learnable query network, processing the control conditions generated in the previous level through the multi-layer perceptron to obtain the control conditions of the corresponding level.
[0094] Specifically, DLQM is the core component of the invention, mainly dynamically generating the control conditions of the network through the query vector input by the user. The control conditions not only constrain the network behavior globally but also can be dynamically adjusted according to the spatial position, thereby ensuring the accuracy and flexibility of the correction results.
[0095] After the query vector is input into the learnable query network, it is first converted into a low-dimensional latent feature through the feature extraction module CE(·), expressed as:
[0096] Qex = CE(Q i ), Q ex ∈ R C (2)
[0097] where Q ex is used to generate the control conditions for each layer of the network. CE(·) is a feature extraction module (such as a fully connected network or a convolutional network), and C represents the number of channels of the low-dimensional features. Through this mapping process, the query vector is embedded into the latent space for efficient processing by the subsequent network.
[0098] This low-dimensional feature representation has the following advantages: reducing computational cost: the high-dimensional query vector reduces the number of parameters through the dimensionality reduction operation; enhancing robustness: the mapped features show stronger generalization ability in the latent space.
[0099] Generate the control conditions for each layer. At each layer i of the learnable query network, Q ex is processed through a Multilayer Perceptron (MLP) to generate the layer-by-layer control conditions The formula is:
[0100]
[0101] where the initial condition is
[0102] Specifically, for each layer i in the learnable query network, the query vector is further processed using fully connected layers FC1 and FC2 to generate the layer control conditions These control conditions dynamically guide the processing of features in each layer, enabling the network to flexibly adjust the correction strategy for different distortion conditions.
[0103]
[0104] where is the control condition of the previous layer. Initially
[0105] The parameters of the MLP are optimized through backpropagation during training, enabling the generated control conditions to adapt to the feature requirements of different layers. The dynamic generation of layer control conditions ensures that the network can adjust its behavior according to the input features and query vectors at each layer, thereby achieving refined distortion correction.
[0106] In an exemplary embodiment, to more precisely handle the non-linear distortion in the fish-eye image, DLQM also introduces a control mechanism related to the spatial position. Through the position encoding PosEnc(x, y), the control condition is associated with the spatial position of the fish-eye image to be corrected:
[0107]
[0108] Among them, PosEnc(x, y) is a position encoding based on the pixel position (x, y) of the fish-eye image to be corrected, which can be generated by a sine function or other mapping methods. The introduction of position encoding enables the network to capture spatial distribution information, thereby solving the local distortion problem that is difficult to handle by traditional scalar control methods.
[0109] It is also possible to use the output in formula (5) as the control condition output by the learnable query network.
[0110] The present invention proposes a learnable query network based on query vectors. The core innovation lies in the introduction of a distortion-aware learnable query mechanism to achieve flexible, efficient, and highly generalizable fish-eye image correction. Traditional fish-eye image correction methods usually rely on fixed parameters or preset calibration processes, lacking flexibility when adapting to various complex scenarios, and it is difficult to perform efficient correction for unseen distortion conditions. By introducing a query mechanism, the network of the present invention can dynamically generate control conditions, adapt to various distortion scenarios, and eliminate the need for recalibration, greatly reducing the deployment cost and complexity.
[0111] S104: Input the original features of the fish-eye image to be corrected and the control conditions at each level of the fish-eye image to be corrected into the hierarchical correction network. Use the controllable convolution modulation block to fuse the original features and the corresponding control conditions, and in the controllable attention modulation block, globally enhance the features output by the controllable convolution modulation block through the control conditions to obtain the distortion-corrected image of the fish-eye image to be corrected.
[0112] In one embodiment, as Figure 2 described, the hierarchical correction network includes a plurality of first controllable convolution modulation blocks, a plurality of controllable attention modulation blocks, and a plurality of second controllable convolution modulation blocks connected in series in sequence. The control conditions at each level correspond to a modulation block in sequence; both the first controllable convolution modulation block and the second controllable convolution modulation block are controllable convolution modulation blocks; using the controllable convolution modulation block to fuse the original features and the corresponding control conditions, and in the controllable attention modulation block, globally enhance the features output by the controllable convolution modulation block through the control conditions to obtain the distortion-corrected image of the fish-eye image to be corrected, including: sequentially passing the original features and the corresponding control conditions through a plurality of first controllable convolution modulation blocks, a plurality of controllable attention modulation blocks, and a plurality of second controllable convolution modulation blocks to obtain the distortion-corrected image of the fish-eye image to be corrected; starting from the second modulation block, the features input to the subsequent modulation block are all generated based on the output features of the previous modulation block.
[0113] Optionally, please continue to refer to Figure 2, there is a symmetric correlation between the input and output of the modulation blocks in the hierarchical correction network. Taking the last second controllable convolutional modulation block as an example, the method for obtaining the features input to the last second controllable convolutional modulation block includes: performing optical flow estimation on the original features to obtain the optical flow field of the original features, and correcting the output features of the first first controllable convolutional modulation block through the optical flow field; concatenating the corrected features and the output features of the previous modulation block of the last second controllable convolutional modulation block in the channel dimension to obtain the fused features, and using the fused features as the input to the last second controllable convolutional modulation block.
[0114] Correspondingly, the features input to the penultimate second controllable convolutional modulation block are generated based on the corrected features, the output features of the second first controllable convolutional modulation block, and the output features of the previous modulation block of the penultimate second controllable convolutional modulation block, and so on until the input-output connections of the modulation blocks in the hierarchical correction network are completed.
[0115] The controllable modulation block is a key module for processing feature maps in the present invention. The core of its design lies in combining the input features and control conditions to achieve precise modeling of local details and global relationships. The controllable convolutional modulation block (CCMB) is a module for processing local texture details. Local details are an important feature dimension in fish-eye image correction. Traditional methods are prone to ignoring small-scale texture deformations, while the CCMB dynamically fuses the input feature F in and the control feature Q c by predicting the fusion ratio θ, thereby achieving adaptive retention of local textures.
[0116] In one embodiment, the controllable convolutional modulation block includes a first channel connection unit, a predicted fusion ratio predictor, a first control unit, a feature modulation unit, and a first feedforward neural network; the implementation process of the controllable convolutional modulation block includes: concatenating the input features and the corresponding control conditions in the channel dimension through the first channel connection unit to obtain the concatenated features; analyzing the concatenated features through the predicted fusion ratio predictor to obtain the predicted fusion ratio; multiplying the input features and the control conditions through the first control unit to obtain the control feature; inputting the predicted fusion ratio, the input features, and the control feature into the feature modulation unit to obtain the modulated features; and inputting the modulated features into the first feedforward neural network to obtain the output features of the controllable convolutional modulation block.
[0117] As Figure 3 shown, Figure 3 is the schematic diagram of the implementation of the controllable convolutional modulation block, with the input original feature F in and the control condition Q c .
[0118] The predicted fusion ratio θ:
[0119] θ = CP(F in , Qc ) (6)
[0120] Among them, CP(·) is a prediction fusion ratio predictor, which consists of two fully connected layers.
[0121] Generate control features:
[0122]
[0123] Among them, represents element-wise multiplication.
[0124] The implementation method of the feature modulation unit is:
[0125] F out = θF c +(1 - θ)F in (8)
[0126] Among them, F out represents the output feature, θ represents the predicted fusion ratio, F c represents the control feature, and F in represents the input feature.
[0127] In one embodiment, the controllable attention modulation block includes a second control unit, a first projection unit, a second projection unit, an attention unit, a second channel connection unit, and a second feed-forward neural network. The implementation process of the controllable attention modulation block includes: multiplying the input feature by the corresponding control condition through the second control unit to obtain a control feature; multiplying the first projection matrix by the control feature through the first projection unit to obtain a query matrix; multiplying the second projection matrix by the input feature through the second projection unit to obtain a key matrix, multiplying the third projection matrix by the input feature to obtain a value matrix, and multiplying the fourth projection matrix by the input feature to obtain the input feature after projection transformation; inputting the query matrix, key matrix, and value matrix into the attention unit to obtain control attention; connecting the control attention with the input feature after projection transformation through the second channel connection unit to obtain a connection feature; performing a residual connection on the connection feature through the second feed-forward neural network to obtain the output feature of the controllable attention modulation block.
[0128] As Figure 4 shown, Figure 4 is the implementation schematic diagram of the controllable attention modulation block. The query matrix Q, key matrix K, and value matrix V are constructed as:
[0129] Q = W Q Fc, K = W K F in , V = W V F in (9)
[0130] Among them, WQ , W K and W V are the first projection matrix, the second projection matrix, and the third projection matrix, respectively.
[0131] Calculate control attention:
[0132]
[0133] Among them, CTRL-ATTN(Q, K, V) represents control attention, Q represents the query matrix, K represents the key matrix, m represents the sequence length of the matrix, and V represents the value matrix. It should be noted that the sequence lengths of the query matrix, the key matrix, and the value matrix are the same, and m represents the sequence length of the query matrix, the key matrix, or the value matrix.
[0134] Connect the control attention with the input features through the second channel connection unit to obtain the connected feature Fa.
[0135] Output modulation feature: Generate the final output feature F through residual connection and feed-forward neural network out . In contrast, the Controllable Attention Modulation Block (CAMB) focuses on modeling global dependencies. The global distortion of fisheye images often contains complex spatial distribution characteristics, and traditional convolutions cannot effectively capture such global relationships. By embedding control conditions into the generation process of the query matrix Q, the key matrix K, and the value matrix V, CAMB can capture correlations across the entire image range.
[0136] The hierarchical correction network is of U-shaped structure, realizing the fusion and correction of multi-scale features. CCMB is used at larger feature map scales (the first several layers) to retain local details; CAMB is used at smaller feature map scales (the last several layers) to capture global dependencies.
[0137] This global modeling ability significantly improves the correction effect under complex distortion patterns. By combining CCMB and CAMB, a hierarchical correction network is constructed. In the first several layers of the network, CCMB is used to extract local details; in the last several layers, CAMB is used to model global relationships. Through this division of labor, the network can take into account both local and global correction effects.
[0138] The present invention proposes a controllable fisheye image distortion correction method based on a query vector set, which solves the problems of insufficient generalization ability of the model to different distortion degrees and inability to flexibly control the correction process in the prior art. By introducing a distortion-aware learnable query mechanism (DLQM), users can flexibly control the correction process by adjusting the query vector without retraining the model. The design of the controllable convolutional modulation block (CCMB) and the controllable attention modulation block (CAMB) enables the control conditions to effectively guide the correction process while retaining rich image details and capturing global distortion patterns. The two-stage training strategy reduces the dependence on diverse training data, accelerates the convergence of the model, and improves the robustness and generalization ability of the model. The present invention has broad application prospects in scenarios such as autonomous driving and security monitoring that require processing wide-angle fisheye images.
[0139] In one embodiment, as Figure 5 shown, Figure 5 is a schematic diagram of an existing method and the method of the present invention provided by the present invention. The existing methods include correction based on a regression model and correction based on a generative model.
[0140] To verify the effectiveness of the controllable fisheye image distortion correction method based on the query vector set proposed by the present invention, we conducted a comparative experiment on the publicly available COCO fisheye image dataset. The following evaluation metrics were used in the experiment: Peak Signal-to-Noise Ratio (PSNR): Measures the restoration accuracy between the corrected image and the real image. The larger the value, the higher the image quality. As shown in Table 1, Table 1 shows the test results based on PSNR on the COCO fisheye image dataset; Structural Similarity Index (SSIM): Used to evaluate the similarity of images at the structural level. The closer the value is to 1, the more accurate the structural restoration.
[0141] To comprehensively verify the performance of the learnable query network (QueryCDR), the method provided by the present invention was compared with the following several classic and up-to-date calibration methods: Automatic Lens Distortion Correction Using Two Parameter Polynomial and Division Models with iterative optimization (SC), A Deep Learning Approach for Automatic Intrinsic Calibration of Wide Field-of-View Cameras (DeepCalib), Blind geometric distortion correction on images through deep learning (Blind), Automatic radial distortion rectification using conditional gan in real-time (DR-GAN), Progressively complementary network for fisheye image rectification using appearance flow (PCN), Dual diffusion architecture for fisheye image rectification: Synthetic-to-real generalization (DDA), and A simple framework for fisheye image rectification with self-supervised representation learning (SimFIR), where d i represents the corresponding distortion degree, as Figure 6 and Figure 7 shown, Figure 6 is a schematic diagram of the correction result in the real scene in the embodiment of the present invention, Figure 7 is a schematic diagram of the correction result on the public dataset in the embodiment of the present invention.
[0142] Table 1
[0143] Method d1 d2 d3 d4 d5 d6 d7 d8 d9 Average SC 10.05 11.59 11.81 11.03 11.50 10.19 11.11 10.19 9.14 10.73 DeepCalib 10.00 10.69 11.01 11.19 11.28 11.46 11.45 11.13 11.11 11.04 Blind 12.98 11.24 10.30 9.62 9.99 11.75 12.69 12.75 12.80 11.57 DR-GAN 15.68 17.40 17.97 18.34 18.50 18.44 17.95 17.94 17.47 17.74 PCN 14.93 17.43 18.43 18.86 18.86 18.88 18.74 17.35 18.26 17.97 DDA 16.39 17.41 17.43 19.48 20.12 18.90 18.85 18.17 18.22 18.33 SimFIR 16.57 17.88 18.43 18.97 19.31 19.28 19.19 18.65 18.48 18.53 QueryCDR 20.01 20.29 20.39 20.41 20.72 20.81 20.58 19.11 20.53 20.32
[0144] Table 2
[0145] Method d1 d2 d3 d4 d5 d6 d7 d8 d9 Average SC 0.101 0.113 0.149 0.182 0.283 0.175 0.141 0.126 0.093 0.151 eepCalib 0.184 0.210 0.223 0.230 0.234 0.246 0.250 0.245 0.246 0.229 Blind 0.308 0.244 0.199 0.176 0.194 0.296 0.367 0.395 0.420 0.289 DR-GAN 0.295 0.330 0.339 0.344 0.344 0.332 0.314 0.312 0.299 0.323 PCN 0.420 0.547 0.589 0.607 0.608 0.610 0.615 0.576 0.603 0.575 DDA 0.455 0.589 0.592 0.620 0.675 0.626 0.619 0.564 0.581 0.591 SimFIR 0.492 0.581 0.626 0.635 0.640 0.628 0.622 0.591 0.595 0.601 QueryCDR 0.643 0.665 0.668 0.677 0.688 0.699 0.692 0.656 0.693 0.676
[0146] The experimental results show that in the task of fisheye image distortion correction, the method provided by the present invention outperforms the existing traditional methods on multiple benchmark datasets. Especially in complex scenarios, the method shows higher robustness and stability, and can effectively process fisheye images with various degrees of distortion. It is worth noting that the method provided by the present invention has high adaptability and usability under various environmental conditions. Especially in practical applications with large lighting changes or complex scenarios, it can maintain a good correction effect.
[0147] The advantage of the method provided by the present invention is that by using a custom query vector-based learnable query network, it can achieve efficient and accurate correction in different degrees of distortion and application scenarios. First, the present invention designs a query mechanism based on a deep neural network (DLQM), which dynamically transforms the query vector provided by the user into control conditions in the correction process, and realizes precise correction of different degrees of distortion by adjusting these control conditions. Secondly, during the processing, the network models the local details and global spatial dependencies respectively through two types of controllable modulation modules (CCMB and CAMB), so as to take into account the detail recovery and global distortion adjustment in the correction process. In order to enhance the accuracy and flexibility of the correction, we design an adaptive query interpolation method, which allows the network to generate new control conditions through the interpolation of query vectors when facing unseen distortion situations, thus ensuring the continuity and flexibility of the correction process. Finally, by introducing a two-stage training strategy, our network can not only converge quickly on training data with multiple degrees of distortion, but also achieve efficient real-time correction in complex scenarios.
[0148] Through these innovative designs, the fisheye image distortion correction method provided by the present invention can correct fisheye images in real time and accurately according to different degrees of distortion and scene requirements, and has broad application prospects. Especially in fields such as autonomous driving and security monitoring that require high precision and need to process large-view images, it provides an efficient and accurate fisheye image distortion correction solution.
[0149] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
[0150] When applying the fisheye image distortion correction method provided by the present invention, it is not necessary to execute according to the Figure 1 sequence of the steps shown. The specific execution sequence of each step can be determined as needed, and the present invention does not limit this.
[0151] The above is the fisheye image distortion correction method provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding fisheye image distortion correction device, as Figure 8 shown.
[0152] Figure 8 It is a schematic diagram of a fisheye image distortion correction device provided by the present invention. The device 800 includes:
[0153] A construction module 801, configured to construct a hierarchical correction network by combining a controllable convolution modulation block and a controllable attention modulation block; the controllable convolution modulation block is used to extract local texture details of features, and the controllable attention modulation block can capture global spatial dependence relationships by using an attention mechanism;
[0154] An acquisition module 802, configured to obtain a query vector of the fisheye image to be corrected from a query vector set according to the distortion degree of the fisheye image to be corrected; the query vector set includes multiple query vectors, and each query vector corresponds to a distortion degree of an image;
[0155] A first determination module 803, configured to determine control conditions of the fisheye image to be corrected at multiple levels according to the query vector of the fisheye image to be corrected; the control condition of each level corresponds to each modulation block in the hierarchical correction network in sequence;
[0156] A second determination module 804, configured to input the original features of the fisheye image to be corrected and the control conditions of the fisheye image to be corrected at each level into the hierarchical correction network, use the controllable convolution modulation block to fuse the original features and the corresponding control conditions, and in the controllable attention modulation block, globally enhance the features output by the controllable convolution modulation block through the control conditions to obtain a distortion correction image of the fisheye image to be corrected.
[0157] For the specific limitations on the fisheye image distortion correction device, reference can be made to the limitations on the fisheye image distortion correction method in the above text, which will not be elaborated here. Each module in the above fisheye image distortion correction device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0158] The present invention also provides a computer-readable storage medium storing a computer program, which can be used to execute the above-mentioned Figure 1 fisheye image distortion correction method provided.
[0159] The present invention also provides Figure 9 a schematic structural diagram of the computer device shown, as Figure 9 shown. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned Figure 1 fisheye image distortion correction method provided.
[0160] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include at least one of non-volatile and volatile memories. The non-volatile memory can include a read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. The volatile memory can include a random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0161] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded by the present invention.
Claims
1. A method for fisheye image distortion correction, characterized in that: include: A hierarchical correction network is constructed by combining a controllable convolution modulation block and a controllable attention modulation block; the controllable convolution modulation block is used to extract local texture details of features, and the controllable attention modulation block can capture global spatial dependencies using an attention mechanism; According to the distortion degree of the fisheye image to be corrected, a query vector of the fisheye image to be corrected is obtained from a query vector set; the query vector set includes a plurality of query vectors, each query vector corresponding to a degree of distortion of the image; Determining control conditions of the fisheye image to be corrected at multiple levels according to the query vector of the fisheye image to be corrected; the control conditions of each level correspond to each modulation block in the hierarchical correction network in turn; The original features of the fisheye image to be corrected and the control conditions of the fisheye image to be corrected at each level are input into the hierarchical correction network, the original features and the corresponding control conditions are fused using the controllable convolution modulation block, and in the controllable attention modulation block, the features output by the controllable convolution modulation block are globally enhanced by the control conditions to obtain the distortion-corrected image of the fisheye image to be corrected.
2. The method according to claim 1, characterized in that The step of determining control conditions of the fisheye image to be corrected at multiple levels according to the query vector of the fisheye image to be corrected includes: The query vector is input into a pre-trained distortion-aware learnable query network to extract control conditions of the fisheye image to be corrected at multiple levels.
3. The method according to claim 2, characterized in that The learnable query network includes a feature extraction module and a multi-layer perceptron; the control conditions for extracting the fisheye image to be corrected at multiple levels include: Converting the query vector into low-dimensional potential features by the feature extraction module; In the first level of the learnable query network, the low-dimensional latent features are processed by the multi-layer perceptron to obtain a control condition of the first level; In the second level to the last level of the learnable query network, the control conditions generated in the previous level are processed by the multi-layer perceptron to obtain the control conditions of the corresponding level.
4. The method according to claim 1, characterized in that: The hierarchical correction network includes a plurality of first controllable convolution modulation blocks, a plurality of controllable attention modulation blocks and a plurality of second controllable convolution modulation blocks connected in series in sequence, and the control conditions of each level correspond to a modulation block in sequence; the first controllable convolution modulation block and the second controllable convolution modulation block are both controllable convolution modulation blocks; the controllable convolution modulation block is used to fuse the original features and the corresponding control conditions, and in the controllable attention modulation block, the features output by the controllable convolution modulation block are globally enhanced by the control conditions to obtain the distortion-corrected image of the fisheye image to be corrected, including: The original features and corresponding control conditions are sequentially passed through a plurality of first controllable convolution modulation blocks, a plurality of controllable attention modulation blocks and a plurality of second controllable convolution modulation blocks to obtain the distortion-corrected image of the fisheye image to be corrected; starting from the second modulation block, the features input to the latter modulation block are generated according to the output features of the previous modulation block.
5. The method according to claim 4, characterized in that The input and output of the modulation block in the hierarchical correction network are symmetrically related, and the method for obtaining the input feature of the last second controllable convolution modulation block includes: Performing optical flow estimation on the original feature to obtain an optical flow field of the original feature; Correcting the output features of the first controllable convolution modulation block through the optical flow field; Perform channel connection on the corrected feature and the output feature of the previous modulation block of the last second controllable convolution modulation block to obtain a fused feature, and use the fused feature as the input of the last second controllable convolution modulation block; Correspondingly, the input feature of the second-to-last second controllable convolution modulation block is generated based on the correction feature, the output feature of the second first controllable convolution modulation block and the output feature of the previous modulation block of the second-to-last second controllable convolution modulation block.
6. The method according to claim 1, characterized in that The controllable convolution modulation block includes a first channel connection unit, a prediction fusion ratio predictor, a first control unit, a feature modulation unit and a first feedforward neural network; the implementation process of the controllable convolution modulation block includes: Performing channel connection on the input feature and the corresponding control condition through the first channel connection unit to obtain a connection feature; Analyzing the connection features by the predicted fusion ratio predictor to obtain a predicted fusion ratio; Multiplying the input feature and the control condition by the first control unit to obtain a control feature; Inputting the predicted fusion ratio, the input feature and the control feature into the feature modulation unit to obtain a modulation feature; The modulation feature is input into the first feedforward neural network to obtain the output feature of the controllable convolution modulation block.
7. The method according to claim 6, characterized in that The characteristic modulation unit is implemented as follows: F out =θF c +(1-θ)F in ; Among them, F out represents the output feature, θ represents the predicted fusion ratio, F c represents the control feature, F in Represents the input features.
8. The method according to claim 1, characterized in that The controllable attention modulation block includes a second control unit, a first projection unit, a second projection unit, an attention unit, a second channel connection unit and a second feedforward neural network. The implementation process of the controllable attention modulation block includes: The second control unit multiplies the input feature and the corresponding control condition to obtain a control feature; Multiplying the first projection matrix by the control feature through the first projection unit to obtain a query matrix; Multiplying the second projection matrix by the input feature through the second projection unit to obtain a key matrix, multiplying the third projection matrix by the input feature to obtain a value matrix, and multiplying the fourth projection matrix by the input feature to obtain the input feature after projection transformation; Inputting the query matrix, the key matrix and the value matrix into the attention unit to obtain control attention; Connecting the control attention with the input feature after the projection transformation through the second channel connection unit to obtain a connection feature; The connection features are residually connected through the second feedforward neural network to obtain the output features of the controllable attention modulation block.
9. The method according to claim 8, characterized in that The attention unit is implemented as follows: Among them, CTRL-ATTN(Q,K,V) represents controlled attention, Q represents query matrix, K represents key matrix, m represents sequence length of matrix, and V represents value matrix.
10. A fisheye image distortion correction device, characterized in that: include: A construction module, for constructing a hierarchical correction network based on a controllable convolution modulation block and a controllable attention modulation block; the controllable convolution modulation block is used to extract local texture details of features, and the controllable attention modulation block can capture global spatial dependencies using an attention mechanism; An acquisition module, used to acquire a query vector of the fisheye image to be corrected from a query vector set according to the distortion degree of the fisheye image to be corrected; the query vector set includes a plurality of query vectors, each query vector corresponding to a degree of distortion of the image; A first determination module is used to determine control conditions of the fisheye image to be corrected at multiple levels according to a query vector of the fisheye image to be corrected; the control conditions of each level correspond to each modulation block in the hierarchical correction network in turn; The second determination module is used to input the original features of the fisheye image to be corrected and the control conditions of the fisheye image to be corrected at each level into the hierarchical correction network, use the controllable convolution modulation block to fuse the original features and the corresponding control conditions, and in the controllable attention modulation block, globally enhance the features output by the controllable convolution modulation block through the control conditions to obtain the distortion-corrected image of the fisheye image to be corrected.
Citation Information
Cited By
Fisheye image correction method and system based on synthetic distortion enhancement
CN120807371A
A method and system for fisheye image correction based on synthetic distortion enhancement
CN120807371B