Fisheye image correction method and system based on synthetic distortion enhancement

By combining a dual midpoint circle mapping model and a distortion-aware Transformer module, strong distortion image samples are generated and the network is trained. This solves the problems of insufficient data coverage and limited modeling capabilities of fisheye images in strong distortion scenarios, and achieves high-quality image correction results.

CN120807371AActive Publication Date: 2025-10-17SHANDONG UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511241529.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-10-17
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Existing fisheye image correction methods suffer from insufficient data coverage, limited modeling capabilities, and poor structural restoration in scenes with strong distortion, especially in addressing image blurring and geometric distortion in edge regions.

Method used

A dual midpoint circle mapping model is used to generate synthetic fisheye image samples with strong distortion features. An encoder-decoder network is constructed through distortion perception and adaptive Transformer modules and trained with multiple loss functions to achieve accurate restoration of the distortion region.

Benefits of technology

It improves image restoration capabilities and model robustness in complex scenes, effectively corrects distorted areas in fisheye images, and outputs clear, structurally consistent, distortion-free images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807371A_ABST
    Figure CN120807371A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of fisheye image correction, and provides a fisheye image correction method and system based on synthetic distortion enhancement in order to solve the problem that an existing method is prone to causing image blurring or geometric structure distortion. The synthetic distortion enhancement-based fisheye image correction method comprises the steps of performing nonlinear geometric transformation on an original image by using a dual midpoint circle mapping model, generating a synthetic fisheye image sample with a strong distortion feature and a clear structure, and forming a synthetic fisheye image sample set; and training a distortion correction network model based on the synthesized fisheye image sample set in combination with a multiple loss function, and correcting the to-be-processed fisheye image by using the trained distortion correction network model so as to realize restoration of a distortion region and output a distortionless fisheye image. According to the method, the image restoration capability and the model robustness under the complex distortion condition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of fisheye image rectification, and particularly relates to a fisheye image rectification method and system based on synthetic distortion enhancement. BACKGROUND

[0002] The statements in this section merely provide background information related to the application and do not necessarily constitute prior art.

[0003] Fisheye cameras have been widely used in automatic driving, intelligent monitoring, virtual reality and many other scenes due to their characteristics of super large field of view and compact volume. However, due to its special optical structure, fisheye lens is prone to produce strong non-linear distortion in the imaging process, especially in the edge area of the image, which seriously affects the geometric structure expression of the image and the accuracy of subsequent visual understanding tasks. To solve the distortion problem of fisheye images, existing image rectification methods include geometric modeling-based methods and deep learning-based methods. The geometric modeling-based method usually needs to calibrate the camera accurately, and estimate and correct the distortion by establishing the imaging model parameters. Although this method is theoretically perfect, it has high requirements for calibration accuracy in practical application, and is difficult to adapt to unknown or dynamic imaging environment. The fisheye image rectification method based on deep learning can realize end-to-end mapping from distorted image to normal image with the powerful feature extraction and fitting ability of convolutional neural network. This method has strong flexibility and adaptability, but its performance depends largely on the diversity and representativeness of the training data. Due to the high cost of obtaining real fisheye images and their corresponding distortion parameters or labeled images, the current available training data sets generally have the problems of insufficient samples and incomplete distortion type coverage, which limits the generalization ability of the model in complex scenes.

[0004] In addition, the existing image synthesis method mainly uses simple projection transformation or low-order polynomial model when generating the simulated fisheye images required for training, which cannot fully simulate the non-linear and high-intensity distortion produced by fisheye lens in reality, resulting in insufficient recovery ability of the trained model for extreme distortion structures. At the same time, the current network structure has limited modeling ability for distorted regions, especially in edge structure restoration and local detail preservation, which is prone to cause image blur or geometric structure distortion. SUMMARY

[0005] To solve the above technical problems, the application provides a fisheye image rectification method and system based on synthetic distortion enhancement, which can realistically simulate high-intensity distortion and introduce distortion perception and structure guidance mechanisms in the network structure to improve the image restoration ability and robustness of the model under complex distortion conditions.

[0006] To achieve the above purpose, the application adopts the following technical solutions: The first aspect of the present application provides a fish-eye image rectification method based on synthetic distortion enhancement.

[0007] In one or more embodiments, a fish-eye image rectification method based on synthetic distortion enhancement is provided, comprising: A nonlinear geometric transformation is performed on the original image using a double midpoint circle mapping model to generate a synthetic fish-eye image sample with strong distortion characteristics and clear structure, forming a synthetic fish-eye image sample set; A distortion rectification network model is trained based on the synthetic fish-eye image sample set and combined with multiple loss functions, and the trained distortion rectification network model is used to rectify the fish-eye image to be processed to restore the distortion area and output a non-distorted fish-eye image; Wherein, the distortion rectification network model adopts an encoder-decoder structure; distortion perception modules and distortion adaptive Transformer modules are introduced in the encoder and the decoder; the distortion perception module in the encoder is used to extract key distortion area features in the synthetic fish-eye image sample / processed fish-eye image, and the distortion adaptive Transformer module in the encoder is used to capture global structural relationship features in the synthetic fish-eye image sample / processed fish-eye image based on the key distortion area features; the distortion perception module in the decoder is used to refine and enhance the key distortion area features, and the distortion adaptive Transformer module in the decoder is used to optimize the global structural relationship features.

[0008] As an implementation, the process of performing nonlinear geometric transformation on the original image using a double midpoint circle mapping model is: A geometric circular path tangent to the boundary is constructed in a two-dimensional image coordinate system, and the intersection of the two circles is used as the mapping position of the pixel in the distortion image to realize distortion transformation of the original image pixels.

[0009] As an implementation, the distortion perception module includes a three-path feature extraction path, one of which is used to extract the texture background of the synthetic fish-eye image sample / processed fish-eye image, and the other two are used to extract the high-frequency features of the high-frequency image and the contour features of the edge enhanced image, respectively.

[0010] As an implementation, the distortion perception module further includes a feature fusion sub-module, which introduces a spatial attention mechanism and a channel attention mechanism, wherein the spatial attention mechanism generates a position-sensitive attention map based on the differences of each spatial position in the input feature map, guiding the model to focus on the area where the structure deformation is prominent; the channel attention mechanism weights according to the importance of different channels in semantic expression.

[0011] As an implementation form, the distortion adaptive Transformer module adopts a deep separable convolution instead of a standard linear mapping layer to generate the query, key and value matrices, and introduces local spatial information to enhance the expression capability for edge structure and texture details.

[0012] As an implementation form, a learnable temperature factor is introduced in the process of multiplying the query and the key to form the attention weight matrix, for adjusting the numerical range and sensitivity of the inner product result.

[0013] As an implementation form, the multiple loss functions include a structure reconstruction loss, an adversarial loss, a perception and style enhancement loss, and a multi-scale supervision loss, and each type of loss function is combined according to a weight to form a total loss function of the distortion correction network model.

[0014] The second aspect of the present application provides a fisheye image correction system based on synthetic distortion enhancement.

[0015] In one or more embodiments, a fisheye image correction system based on synthetic distortion enhancement includes: A sample set construction module is configured to perform a non-linear geometric transformation on an original image by using a double midpoint circle mapping model to generate a synthetic fisheye image sample with strong distortion characteristics and clear structure, thereby forming a synthetic fisheye image sample set. A model training module is configured to train a distortion correction network model based on the synthetic fisheye image sample set and in combination with multiple loss functions, and to correct a to-be-processed fisheye image by using the trained distortion correction network model, so as to restore the distortion region and output a fisheye image without distortion. The distortion correction network model adopts an encoder-decoder structure, and a distortion perception module and a distortion adaptive Transformer module are introduced in the encoder and the decoder. The distortion perception module in the encoder is configured to extract key distortion region features in the synthetic fisheye image sample / to-be-processed fisheye image, and the distortion adaptive Transformer module in the encoder is configured to capture global structure relationship features in the synthetic fisheye image sample / to-be-processed fisheye image based on the key distortion region features. The distortion perception module in the decoder is configured to refine and enhance the key distortion region features, and the distortion adaptive Transformer module in the decoder is configured to optimize the global structure relationship features.

[0016] The third aspect of the present application provides a computer readable storage medium.

[0017] A computer readable storage medium has a computer program stored thereon, and the program is executed by a processor to implement the steps in the fisheye image correction method based on synthetic distortion enhancement as described above.

[0018] A fourth aspect of the present invention provides an electronic device.

[0019] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the fisheye image correction method based on synthetic distortion enhancement as described above are implemented.

[0020] Compared with the prior art, the present invention has the following beneficial effects: The present invention's fisheye image correction method and system based on synthetic distortion enhancement aims to address existing issues such as insufficient data coverage, limited modeling capabilities, and poor structural restoration in highly distorted scenes. This method uses a dual midpoint circle mapping model to perform a nonlinear transformation on a standard image, generating a synthetic image with diverse, high-intensity distortion features. By introducing a distortion-aware dual attention mechanism and a distortion-adaptive Transformer module, key distortion features are extracted and global structural relationships are modeled. Finally, a multi-scale deep correction network is constructed that integrates local perception and global modeling capabilities. End-to-end training is performed using multiple loss functions to achieve accurate restoration of distorted areas and high-quality image output. This method is suitable for geometric correction and visual quality improvement of fisheye images in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0022] Figure 1 is a principle diagram of a fisheye image correction method based on synthetic distortion enhancement according to an embodiment of the present invention; Figure 2 Schematic diagram of a dual midpoint circle mapping model in a fisheye image correction method according to an embodiment of the present invention; Figure 3 Schematic diagram of the structure of the distortion-aware dual attention mechanism in the fisheye image correction method according to an embodiment of the present invention; Figure 4 Schematic diagram of the structure of the distortion adaptive Transformer module in the fisheye image correction method according to an embodiment of the present invention; Figure 5 1 is a schematic structural diagram of a fisheye image correction system based on synthetic distortion enhancement according to an embodiment of the present invention; Figure 6 This is a synthetic comparison between the fisheye image correction method according to an embodiment of the present invention and other methods; Figure 7 This is the correction result of a real fisheye image captured by an OV5640 220-degree fisheye camera and a web search according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The application will be further described below in connection with the drawings and examples.

[0024] It should be noted that the following detailed description is illustrative only and is intended to provide further description of the application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0025] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments in accordance with the present application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, elements, components, and / or groups thereof, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.

[0026] Figure 1 is a schematic diagram of a fish-eye image rectification method based on synthetic distortion enhancement according to an embodiment of the application. In combination with Figure 1 , the fish-eye image rectification method based on synthetic distortion enhancement can include: S101, performing non-linear geometric transformation on the original image by using a double midpoint circle mapping model to generate a synthetic fish-eye image sample with strong distortion characteristics and clear structure, thereby forming a synthetic fish-eye image sample set.

[0027] The process of performing non-linear geometric transformation on the original image by using the double midpoint circle mapping model is as follows: A geometric circular path tangent to the boundary is constructed in a two-dimensional image coordinate system, and the intersection of the two circles is taken as the mapping position of the pixel point in the distortion image, so as to realize distortion transformation of the original image pixel.

[0028] In the embodiment of the application, a standard view image from a public data set is selected as the original input image, and the image is subjected to non-linear geometric transformation by the double midpoint circle mapping model to simulate the strong distortion effect of the fish-eye lens in the actual imaging process, so as to construct a synthetic fish-eye image with high non-linear characteristics and complex geometric distortion.

[0029] Specifically, as shown in Figure 2 , let the single image coordinate system be XOY, and let the pixel point A on the vertical line and the pixel point on the horizontal line in the image be taken at random, and construct the circles 、 tangent to , , whose centers are , , and the intersection is denoted as q.

[0030] In the vertical direction, let the coordinates of point A be , and let the foot of the perpendicular from A to the bottom edge (y = 0) of the image be point C. A circle is constructed that passes through point C and is tangent to the vertical line . The center of the circle is denoted as , and the coordinates are obtained by the following formula: ; The radius of the circle is obtained by the following formula: ; Since point q ∈ , and satisfies the same vertical height as point , i.e. , the distortion mapping point of point A in the vertical direction is calculated.

[0031] Similarly, in the horizontal direction, let the coordinates of point B be , and let the foot of the perpendicular from B to the left boundary (x = 0) of the image be point D. A circle is constructed that passes through point D and is tangent to the horizontal line . The center of the circle is denoted as , and the coordinates are obtained by the following formula: ; The radius of the circle is obtained by the following formula: ; Similarly, point q ∈ , and , so the distortion mapping point of point B in the horizontal direction is obtained.

[0032] To maintain the geometric consistency and structural symmetry of the distortion mapping, the midpoint of and is selected: as the input, and the mapping result in the distorted image is calculated. The final distorted pixel point is defined by the joint mapping function: ; where represents the distortion function based on the double midpoint circle mapping, represents the intersection of circles , , and the point in the mapping direction is the final output.

[0033] ​​​​On the basis of the above geometrical configuration, a batch image processing strategy is adopted to uniformly distort a large-scale standard image dataset to obtain a pair of distorted images and corresponding standard images, and by adjusting the parameter r, the type and intensity of distortion are controlled to form a training image set containing different distortion degrees.

[0034] S102, based on the synthetic fisheye image sample set and combined with a multiple loss function, a distortion correction network model is trained, and the trained distortion correction network model is used to correct the fisheye image to be processed to realize the restoration of the distortion area and output the fisheye image without distortion.

[0035] Wherein, the distortion correction network model adopts an encoder-decoder structure; distortion perception modules and distortion adaptive Transformer modules are introduced in the encoder and the decoder; the distortion perception module in the encoder is used to extract the key distortion area features in the synthetic fisheye image sample / processed fisheye image, and the distortion adaptive Transformer module in the encoder is used to capture the global structure relationship features in the synthetic fisheye image sample / processed fisheye image based on the key distortion area features; the distortion perception module in the decoder is used to refine and enhance the key distortion area features, and the distortion adaptive Transformer module in the decoder is used to optimize the global structure relationship features.

[0036] In view of the problem that the distortion area in the fisheye image usually presents complex visual features such as non-uniform deformation, edge blur and local structure distortion, the embodiment introduces a double attention mechanism of distortion perception to realize the significant feature extraction and representation enhancement of the key distortion area.

[0037] As shown in Figure 3 , first, the input distorted image is respectively subjected to high-pass filtering and edge detection processing to obtain a high-frequency image and an edge image ; wherein the high-frequency image retains the image texture change, and the edge image highlights the structure contour information. The three images are jointly input into a multi-branch feature extraction network, and multi-scale primary features are extracted through a lightweight convolution module. Each branch network extracts deep semantic features through shared parameters and local adaptive path adjustment, and outputs to a feature fusion unit.

[0038] The distortion perception module of the embodiment includes three-path feature extraction paths, one of which is used to extract the texture background of the synthetic fisheye image sample / original image, and the other two are used to extract the high-frequency features of the high-frequency image and the contour features of the edge enhanced image respectively, and the calculation formula of the three paths is as follows: ; 、 and denote three pathway characteristics; wherein, denotes element-wise multiplication. denotes edge enhancement function.

[0039] The above three groups of characteristics are spliced in the channel dimension to form a multi-dimensional feature expression tensor that fuses spatial details and edge responses, and compression and normalization are realized through a fusion convolution module: ; wherein, is an activation function; is a fusion function; is a convolution function.

[0040] In the embodiment of the application, the distortion perception module further comprises a feature fusion sub-module, which introduces a spatial attention mechanism and a channel attention mechanism, wherein the spatial attention mechanism generates a position-sensitive attention map based on the differences of each spatial position in the input feature map, guiding the model to focus on the area where the structure deformation is prominent; the channel attention mechanism weights according to the importance of different channels in semantic expression.

[0041] In the fusion stage, the spatial attention mechanism is first applied to emphasize the distortion significant area in an adaptive manner. This module allocates weights according to the importance of each pixel position, preferentially preserving the area features of the edge distortion and the texture fracture, and filtering the background of the redundant or invalid area in the image. Then the channel attention mechanism is introduced to weight and reconstruct the feature map in the channel dimension, guiding the model to focus on the feature channel with strong expression ability and sensitive response change. The spatial and channel attention mechanisms are used in series to form a structure-texture joint enhancement path, which realizes robust modeling of complex distortion areas without introducing additional structure bias.

[0042] In specific implementation, the attention module adopts residual connection and Sigmoid activation function design to ensure the stability of gradient and the smoothness of feature enhancement. The attention weight and the original feature map are fused pixel by pixel to form the enhanced distortion perception feature map . The specific calculation is as follows: ; wherein, denotes the attention module function.

[0043] Finally, the feature map will be input into the subsequent global structure modeling module as input, realizing further modeling of long-distance structure dependence and global semantics.

[0044] To achieve global modeling of fisheye image structure information and geometric consistency constraints, the embodiment introduces a distortion adaptive Transformer module to realize structural analysis of long-distance dependencies in images and multi-scale feature fusion. The module is improved on the basis of the traditional Transformer encoder structure, which has the ability of global semantic modeling while enhancing the ability to analyze and respond to local structural distortion areas.

[0045] As shown in Figure 4 , in the attention mechanism construction process, the Transformer encoder uses a depth separable convolution to replace the traditional linear mapping layer to generate query , key and value matrices, and introduces local spatial information to enhance the expression ability of edge structure and texture details. The specific calculation is as follows: ; In the process of calculating the attention weight, a learnable temperature factor is further introduced to adjust the scale and sensitivity of the query and key matrix product results. First, normalize and : ; Then, calculate the normalized attention weight matrix , which is defined as follows: ; The temperature factor can be dynamically optimized according to the training process, so that the attention mechanism can more accurately respond to cross-regional distortion correlation.

[0046] Finally, based on the calculated attention weight and value matrix , feature aggregation is completed: ; In the feedforward network part, to overcome the limitations of traditional fully connected structure in structure information preservation, the embodiment uses a gated convolution feedforward network module instead of a standard feedforward subnetwork. The gated convolution feedforward network module is composed of two parallel paths, one of which is processed by an activation function, and the other is used as a gate control branch to generate a weight map; the two features are multiplied element by element before output to realize dynamic feature adjustment, and the specific calculation is as follows: ; ; To effectively restore and reconstruct severely distorted areas in fisheye images at the pixel level, this embodiment constructs a distortion correction model with multi-scale collaborative modeling capabilities. This model employs an encoder-decoder architecture and integrates the aforementioned distortion-aware module and the distortion-adaptive Transformer module at different resolution levels, forming a deep neural network with both structural guidance and distortion modeling capabilities.

[0047] Specifically, the encoder part of the model consists of multiple downsampling convolutional layers, and each encoding layer is responsible for extracting image features at different scales. In the high-resolution 6-layer shallow feature, the model introduces a distortion perception module to enhance the ability to capture local edge deformation areas, distorted textures, and detail changes; while in the low-resolution 5-layer deep feature, a distortion-adaptive Transformer module is embedded to model global structural dependencies and restore cross-regional geometric relationships, ensuring that the network has the ability to restore the overall structure and globally consistent expression. In the decoder part of the model, a layer-by-layer upsampling strategy is adopted, combined with a jump connection mechanism, to fuse the features extracted from each layer of the encoder back to the corresponding layer of the decoder, realizing the cascade integration of multi-scale semantic information and original structural features.

[0048] To achieve high-fidelity restoration and structural consistency maintenance of the distorted areas of the fisheye image, this embodiment adopts an end-to-end training strategy driven by multiple loss functions to globally optimize the aforementioned distortion correction network model.

[0049] Multiple loss functions include structure reconstruction loss, adversarial loss, perception and style enhancement loss, and multi-scale supervision loss; Various loss functions are combined according to weighted combinations to form the total loss function of the distortion correction network model. : .

[0050] Structural reconstruction losses ; Fighting Losses ; Perceptual and style enhancement losses: ; Multi-scale supervision loss ; in, Indicates the calibration results is the corresponding reference image, and Respectively represent The feature map after layer correction and the original feature map, is the downsampling operation, represents the convolutional map, is the feature splicing operation, and is a loss weight coefficient; represents the feature map of the i-th j layer, and C, H and W represent the channel number, height and width of the feature map of the i-th layer, respectively. j represents the feature map of the i-th layer of the loss network. j represents the Gram matrix of the i-th measures the ability of the discriminator to correctly identify real data. measures the ability of the generator to deceive the discriminator. is a discriminator function. is a generator function.

[0051] During the training process, the model uses the Adam optimizer to update the parameters, and the initial learning rate is set to 1e-4, and a learning rate decay strategy is combined to adapt to the training requirements at different stages. To avoid gradient explosion and overfitting, a gradient clipping mechanism and an early stopping strategy are introduced, and dynamic adjustment is performed based on the performance of the validation set after each round of training.

[0052] After training is completed, the model performs fast forward inference on any input fisheye image in the test phase, completes the structure restoration of the image distortion area and generates a non-distortion image.

[0053] The correction process of the embodiment of the application has end-to-end processing capability, and only a single fisheye image is required as input, and the network can automatically complete distortion area identification, feature extraction, structure modeling and image reconstruction and other operations, and the output is a non-distortion image with excellent visual quality, clear edges and consistent structure. Relying on the pre-synthetic distortion enhancement and multi-task training strategy, the correction network still has good generalization ability and robustness when facing fisheye distortion images of different degrees and different types, and can be widely used in image preprocessing tasks in complex application scenarios such as autonomous driving, intelligent security and virtual reality.

[0054] The embodiment of the application adopts SMIA-TV distortion as an objective indicator to measure the degree of image distortion, and introduces a modulation transfer function (MTF, Modulation Transfer Function) as an evaluation standard for image sharpness. The greater the absolute value of SMIA-TV distortion, the more significant the distortion; the greater the MTF value, the higher the image sharpness.

[0055] Wherein, SMIA TV distortion is defined by the Standard Mobile Imaging Architecture (SMIA) organization, which is a quantitative standard specially used to evaluate image distortion in the field of mobile device cameras. Its analysis method is performed by comparing the relative distortion of the edge and the middle area.

[0056] Table 1 compares the performance of the double midpoint circle mapping model and other mainstream synthesis methods in the above indicators. The results show that the double midpoint circle mapping model performs best in SMIA-TV distortion, and is better than other methods in the clarity evaluation, proving that it can effectively maintain high image clarity and structural fidelity while generating strong distortion images. To visually display the effects, Figure 6 The synthesis comparison of the method and other methods is shown, and the double midpoint circle mapping model proposed in the application exhibits more significant distortion characteristics while maintaining the coherence of the image structure.

[0057] Table 1 quantitative comparison of different distortion models;

[0058] Among them, ↑ represents that the larger the numerical value, the better the performance of the index, and ↓ represents that the smaller the numerical value, the better the performance of the index.

[0059] To evaluate the performance of the method in the fish-eye image correction, six representative comparison methods are selected, including the traditional geometric model SC, the regression-based DeepCalib, the generation adversarial network-based DR-GAN, SimFIR and PCN, covering geometric modeling, generation modeling, end-to-end regression and other technical paths. The peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) are used for quantitative evaluation to measure the pixel and structural restoration ability, the multi-scale structural similarity (MS-SSIM) and the composite wavelet structural similarity (CW-SSIM) are used to reflect the image multi-scale and edge structure consistency, the generated image perceptual quality (FID) and the perceptual image similarity (LPIPS) are used to evaluate the perceptual authenticity and detail quality of the image. Table 2 summarizes the evaluation results of each method on the synthesis dataset. The results show that the application performs best in all indicators, better than existing methods, and has significant comprehensive advantages. Figure 7 Real fish-eye image correction results from network search and OV5640 220-degree fish-eye camera shooting are shown, unlike other methods which have problems such as being unable to correct, structural distortion or artifacts when dealing with complex distortion, the application can effectively repair structural distortion and maintain the geometric consistency and content integrity of the image without camera parameters.

[0060] Table 2 quantitative evaluation of different methods on the synthesis dataset;

[0061] Among them, ↑ represents that the larger the numerical value, the better the performance of the index, and ↓ represents that the smaller the numerical value, the better the performance of the index.

[0062] Figure 5 is a fish-eye image correction system structure schematic diagram based on synthetic distortion enhancement in an embodiment of the application, asFigure 5 As shown, the fish-eye image correction system based on synthetic distortion enhancement in the embodiment can include: A sample set construction module 501, configured to perform non-linear geometric transformation on an original image by using a double midpoint circle mapping model, to generate a synthetic fish-eye image sample with strong distortion characteristics and clear structure, and form a synthetic fish-eye image sample set; A model training module 502, configured to train a distortion correction network model based on the synthetic fish-eye image sample set and in combination with a multiple loss function, and use the trained distortion correction network model to correct a fish-eye image to be processed, so as to restore the distortion region and output a non-distortion fish-eye image; The distortion correction network model adopts an encoder-decoder structure, and the distortion perception module and the distortion adaptive Transformer module are introduced in the encoder and the decoder; the distortion perception module in the encoder is configured to extract key distortion region features in the synthetic fish-eye image sample / the fish-eye image to be processed, and the distortion adaptive Transformer module in the encoder is configured to capture global structure relationship features in the synthetic fish-eye image sample / the fish-eye image to be processed based on the key distortion region features; the distortion perception module in the decoder is configured to refine and enhance the key distortion region features, and the distortion adaptive Transformer module in the decoder is configured to optimize the global structure relationship features.

[0063] It should be noted here that, Figure 5 The modules in the fish-eye image correction system based on synthetic distortion enhancement correspond one-to-one to the steps in the fish-eye image correction method based on synthetic distortion enhancement, and the specific implementation process is the same, which will not be described here.

[0064] The fish-eye image correction method and system based on synthetic distortion enhancement aims to solve the problems of insufficient data coverage, limited modeling capability and poor structure restoration effect in the prior art in a strong distortion scene. The method performs non-linear transformation on a standard image based on a double midpoint circle mapping model to generate synthetic images with diversified high-intensity distortion characteristics; by introducing a distortion perception double attention mechanism and a distortion adaptive Transformer module, key distortion features are extracted and global structure relationships are modeled; finally, a multi-scale deep correction network that integrates local perception and global modeling capability is constructed, and end-to-end training is performed in combination with a multiple loss function, to realize accurate restoration of the distortion region and output of high-quality images. The fish-eye image correction system based on synthetic distortion enhancement includes image generation, feature extraction, structure modeling, model training and reasoning, image correction and reconstruction, and other functional modules, has good robustness and generalization performance, and is suitable for fish-eye image geometric correction and visual quality improvement in complex scenes.

[0065] In one or more embodiments, an electronic device is provided that includes a central processing unit (CPU) that can perform various appropriate actions and processes in accordance with a program stored in a read only memory (ROM) or a program loaded from a storage section into a random access memory (RAM). In the RAM, various programs and data required for system operation are also stored. The central processing unit, the ROM, and the RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0066] Connected to the I / O interface are an input section including a keyboard, a mouse, etc.; an output section including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section including a hard disk, etc.; and a communication section including a network interface card such as a local area network (LAN) card, a modem, etc. The communication section performs communication processing via a network such as the Internet. A drive is also connected to the I / O interface as necessary. A removable medium such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive as necessary, so that a computer program read out therefrom is installed in the storage section as necessary.

[0067] The central processing unit in the electronic device of the present embodiment implements the steps in the above-described fisheye image rectification method based on synthetic distortion enhancement when executing the program.

[0068] In particular, according to embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the fisheye image rectification method based on synthetic distortion enhancement. In such embodiments, the computer program can be downloaded and installed from a network via the communication section, and / or installed from a removable medium. When the computer program is executed by the central processing unit, various functions defined in the apparatus of the present application are performed.

[0069] The computer program instructions corresponding to the fisheye image rectification method based on synthetic distortion enhancement can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instruction means that implement the function specified in the flowchart Figure 1 one or more flows and / or blocks Figure 1 one or more flows and / or blocks

[0070] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing relevant hardware, and the program can be stored in a computer readable storage medium. When the program is executed, the program can include the processes of the above-mentioned embodiment methods. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM), a random access memory (RAM), or the like.

[0071] The above only describes the preferred embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A fisheye image correction method based on synthetic distortion enhancement, characterized in that: include: The double midpoint circle mapping model is used to perform nonlinear geometric transformation on the original image to generate synthetic fisheye image samples with strong distortion characteristics and clear structure, forming a synthetic fisheye image sample set; Training a distortion correction network model based on the synthetic fisheye image sample set and in combination with multiple loss functions, and using the trained distortion correction network model to correct the fisheye image to be processed, so as to restore the distorted area and output an undistorted fisheye image; The distortion correction network model adopts an encoder-decoder structure; a distortion perception module and a distortion adaptive Transformer module are introduced in both the encoder and the decoder; the distortion perception module in the encoder is used to extract key distortion area features in the synthesized fisheye image samples / the fisheye image to be processed, and the distortion adaptive Transformer module in the encoder is used to capture the global structural relationship features in the synthesized fisheye image samples / the fisheye image to be processed based on the key distortion area features; The distortion perception module in the decoder is used to refine and enhance the features of the key distortion area, and the distortion adaptive Transformer module in the decoder is used to optimize the global structural relationship features.

2. The fisheye image correction method based on synthetic distortion enhancement according to claim 1, characterized in that: The process of performing nonlinear geometric transformation on the original image using the dual midpoint circle mapping model is as follows: A geometric circular path tangent to the boundary is constructed in the two-dimensional image coordinate system, and the intersection of the two circles is used as the mapping position of the pixel point in the distorted image to achieve the distortion transformation of the original image pixels.

3. The fisheye image correction method based on synthetic distortion enhancement according to claim 1, wherein: The distortion perception module includes a three-channel feature extraction path, one of which is used to extract the texture background of the synthetic fisheye image sample / original image, and the other two are used to extract the high-frequency features of the high-frequency image and the contour features of the edge-enhanced image, respectively.

4. The fisheye image correction method based on synthetic distortion enhancement according to claim 3, characterized in that: The distortion perception module also includes a feature fusion submodule, which introduces a spatial attention mechanism and a channel attention mechanism. The spatial attention mechanism generates a position-sensitive attention map based on the differences in spatial positions in the input feature map, guiding the model to focus on areas with prominent structural deformation; the channel attention mechanism weights different channels according to their importance in semantic expression.

5. The fisheye image correction method based on synthetic distortion enhancement according to claim 1, characterized in that: The distortion-adaptive Transformer module uses depthwise separable convolution instead of standard linear mapping layers to generate query, key, and value matrices, and introduces local spatial information to enhance the ability to express edge structures and texture details.

6. The fisheye image correction method based on synthetic distortion enhancement according to claim 5, characterized in that: A learnable temperature factor is introduced in the process of multiplying the query and key to form the attention weight matrix to adjust the numerical range and sensitivity of the inner product result.

7. The fisheye image correction method based on synthetic distortion enhancement according to claim 1, characterized in that: Multiple loss functions include structure reconstruction loss, adversarial loss, perception and style enhancement loss, and multi-scale supervision loss. Various loss functions are combined according to weighted combinations to form the total loss function of the distortion correction network model.

8. A fisheye image correction system based on synthetic distortion enhancement, characterized in that: include: A sample set construction module is used to perform nonlinear geometric transformation on the original image using a dual midpoint circle mapping model to generate synthetic fisheye image samples with strong distortion characteristics and clear structure, thereby forming a synthetic fisheye image sample set; A model training module, which is used to train a distortion correction network model based on the synthetic fisheye image sample set and in combination with multiple loss functions, and use the trained distortion correction network model to correct the fisheye image to be processed, so as to restore the distorted area and output an undistorted fisheye image; The distortion correction network model adopts an encoder-decoder structure; a distortion perception module and a distortion adaptive Transformer module are introduced in both the encoder and the decoder; the distortion perception module in the encoder is used to extract key distortion area features in the synthesized fisheye image samples / the fisheye image to be processed, and the distortion adaptive Transformer module in the encoder is used to capture the global structural relationship features in the synthesized fisheye image samples / the fisheye image to be processed based on the key distortion area features; The distortion perception module in the decoder is used to refine and enhance the features of the key distortion area, and the distortion adaptive Transformer module in the decoder is used to optimize the global structural relationship features.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the fisheye image correction method based on synthetic distortion enhancement as described in any one of claims 1 to 7 are implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the fisheye image correction method based on synthetic distortion enhancement according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Distortion image correction method and system based on distortion distribution diagram

    CN111260565A

  • Fisheye image correction method and system based on SKNet attention mechanism

    CN117522748A

  • Fisheye image correction method and system based on refraction law

    CN118608434A

  • Distortion correction method and device for fisheye image

    CN120182145A

  • Pipe wall fisheye distortion correction method based on multi-scale discriminator and center self-attention

    CN120339141A