A fisheye lens image pixel-level edge enhancement and denoising method

By generating a spatial density weight map and an adaptive surface convolution kernel, and dynamically adjusting the denoising and enhancement processes, the problem of scale mismatch in traditional fisheye lens image processing is solved, achieving high-precision edge feature preservation and spatial awareness.

CN121998860BActive Publication Date: 2026-06-23XIAMEN ALAUD OPTICAL CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAMEN ALAUD OPTICAL CO LTD
Filing Date
2026-04-08
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

In traditional fisheye lens image processing methods, the mismatch between mathematical processing scale and physical stretching scale leads to the irreversible erasure of key semantic features at the edge of the field of view, and insufficient spatial perception accuracy and feature preservation under extreme fields of view.

Method used

By acquiring the optical distortion parameters of the fisheye lens, a spatial density weight map and an adaptive surface convolution kernel are generated. The denoising threshold and edge enhancement processing are dynamically adjusted, and the distortion weighted structural similarity evaluation is combined to ensure that the mathematical processing matches the physical stretching scale.

Benefits of technology

It effectively avoids the irreversible erasure of key semantic features at the edges, improves spatial perception accuracy and feature preservation under extreme field of view, and ensures high reliability and accuracy of image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998860B_ABST
    Figure CN121998860B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing, in particular to a fisheye lens image pixel-level edge enhancement and denoising method, comprising: a data acquisition step: acquiring original fisheye image data and preset fisheye lens optical distortion parameters; a weight generation step: determining the distortion stretching rate of a pixel point according to the optical distortion parameters; and generating a spatial density weight map; a threshold determination and denoising step: determining a target denoising threshold according to the spatial density weight map; denoising the original fisheye image data to generate an intermediate denoising image; a convolution kernel generation step: generating an adaptive curved surface convolution kernel according to the optical distortion parameters; an edge enhancement step: using the adaptive curved surface convolution kernel to perform edge enhancement processing on the intermediate denoising image; and generating an enhanced image; the present application effectively avoids the risk of irreversible erasure of key semantic features in the edge field, and ensures high-reliability feature preservation in the extreme field.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a fish-eye lens image pixel-level edge enhancement and denoising method. BACKGROUND

[0002] With the rapid iteration of automatic driving and security monitoring systems, various image processing devices equipped with fish-eye lenses appear on the market. The non-linear optical distortion of such fish-eye lenses has very high requirements for image processing accuracy. How to achieve extreme edge enhancement and denoising is also a current technical challenge. Under the premise of pursuing target detection accuracy in the field of machine vision, whether to preserve edge high-frequency details has also become a research focus in the field of image processing.

[0003] Traditional fish-eye image processing currently mainly relies on the following methods: global uniform denoising, fixed square convolution kernel edge enhancement, and global uniform structural similarity evaluation.

[0004] However, the global uniform denoising, fixed square convolution kernel edge enhancement, and global uniform structural similarity evaluation methods all have certain defects. For example, the global uniform denoising method considers the weak high-frequency details of the edge as noise and removes them, resulting in the loss of key semantic features. The fixed square convolution kernel edge enhancement method has a mismatch between the mathematical processing scale and the physical stretching scale, which easily amplifies the artifacts caused by optical dispersion in the edge region. The global uniform structural similarity evaluation method excessively pursues global smoothness and sets a uniform index, but the uniform index is difficult to adapt to the edge region that is forcibly stretched in the physical space. SUMMARY

[0005] The present application aims to provide a fish-eye lens image pixel-level edge enhancement and denoising method to solve the following technical problems:

[0006] Avoiding the irreversible erasure or reconstruction deviation of key semantic features in the edge field caused by the mismatch between the mathematical processing scale and the physical stretching scale, and greatly improving the spatial perception accuracy in the extreme field and ensuring high-reliable feature preservation.

[0007] The purpose of the present application can be achieved by the following technical solutions:

[0008] A fish-eye lens image pixel-level edge enhancement and denoising method, the method is applied to an image processing device, and the method comprises the following steps: an image processing device acquires original fish-eye image data and a preset fish-eye lens optical distortion parameter;

[0009] The image processing device determines the distortion stretching rate of a pixel point in the original fish-eye image data according to the fish-eye lens optical distortion parameter;

[0010] The image processing device generates a spatial density weight map based on the distortion stretching rate; the image processing device determines the target denoising threshold based on the spatial density weight map;

[0011] The image processing device performs denoising processing on the original fisheye image data according to the target denoising threshold to generate an intermediate denoised image;

[0012] The image processing device generates an adaptive surface convolution kernel based on the optical distortion parameters of the fisheye lens;

[0013] The image processing device uses the adaptive surface convolution kernel to perform edge enhancement processing on the intermediate denoised image to generate an enhanced image.

[0014] Furthermore, determining the target denoising threshold based on the spatial density weight map includes: the image processing device determining whether the weight value in the spatial density weight map is greater than or equal to a preset weight threshold.

[0015] If so, the image processing device determines the target denoising threshold as a preset first denoising threshold;

[0016] If not, the image processing device determines the target denoising threshold to be a preset second denoising threshold; wherein the first denoising threshold is greater than the second denoising threshold.

[0017] Furthermore, generating an adaptive surface convolution kernel based on the fisheye lens optical distortion parameters includes: the image processing device determining the optical distortion center based on the fisheye lens optical distortion parameters; and the image processing device constructing a radial tangential orthogonal coordinate system based on the optical distortion center.

[0018] The image processing device determines the target deformation direction based on the radial vector in the radial-tangential orthogonal coordinate system;

[0019] The image processing device deforms the preset initial convolution kernel according to the target deformation direction to generate the adaptive surface convolution kernel.

[0020] Furthermore, the intermediate denoised image is subjected to edge enhancement processing using the adaptive surface convolution kernel to generate an enhanced image, including: the image processing device determining the tangent direction according to the radial tangential orthogonal coordinate system;

[0021] The image processing device performs edge enhancement processing on the intermediate denoised image by applying the adaptive surface convolution kernel along the tangent direction, thereby generating the enhanced image.

[0022] Furthermore, the method also includes: the image processing device acquiring preset reference image data; the image processing device calculating an initial structural similarity based on the enhanced image and the reference image data;

[0023] The image processing device performs weighted processing on the initial structural similarity according to the spatial density weight map to generate a distortion-weighted structural similarity.

[0024] The image processing device adjusts the preset parameters in the denoising process or the edge enhancement process based on the distortion weighted structural similarity.

[0025] Furthermore, the initial structural similarity is weighted according to the spatial density weight map to generate a distortion weighted structural similarity, including: the image processing device determining whether the weight value in the spatial density weight map is less than a preset center determination threshold;

[0026] If yes, the image processing device assigns a preset first evaluation weight to the initial structural similarity; if no, the image processing device assigns a preset second evaluation weight to the initial structural similarity.

[0027] The image processing device multiplies the initial structural similarity by the assigned evaluation weight to generate the distortion-weighted structural similarity; wherein the first evaluation weight is greater than the second evaluation weight.

[0028] Furthermore, determining the distortion stretching rate of pixels in the original fisheye image data based on the fisheye lens optical distortion parameters includes: the image processing device constructing a physical projection mapping model based on the fisheye lens optical distortion parameters;

[0029] The image processing device inputs the coordinates of the pixels in the original fisheye image data into the physical projection mapping model to calculate the actual physical area corresponding to the pixel.

[0030] The image processing device determines the distortion stretching rate based on the ratio of the actual physical area to the preset standard pixel area.

[0031] Furthermore, the method also includes: the image processing device sending the enhanced image to a preset machine vision recognition system, so that the machine vision recognition system performs target detection processing based on the enhanced image.

[0032] Furthermore, the machine vision recognition system is deployed in pre-installed autonomous driving equipment;

[0033] The machine vision recognition system performs target detection processing based on the enhanced image, including: the machine vision recognition system extracts semantic features of edge regions in the enhanced image;

[0034] The machine vision recognition system generates target detection results based on the semantic features; the machine vision recognition system generates control command data to indicate the operating status of the autonomous driving equipment based on the target detection results.

[0035] The beneficial effects of this invention are:

[0036] 1. This invention addresses the problem that traditional global uniform denoising easily erases high-frequency details at the edges. This method determines the distortion stretching rate of pixels based on the optical distortion parameters of the fisheye lens, and then generates a spatial density weight map. Based on this weight map, the target denoising threshold is dynamically determined for different regions of the image. Stronger denoising is applied to the high-density area in the center, and the denoising threshold is reduced in the low-density area at the edge. This adaptive mechanism effectively avoids the irreversible erasure of key semantic features at the edges, and maximizes the integrity of the true semantics of the edge image.

[0037] 2. This invention addresses the problems of scale mismatch and magnified artifacts caused by traditional fixed square convolution kernel edge enhancement. This method constructs a radial-tangential orthogonal coordinate system based on optical distortion parameters to determine the target deformation direction, thereby generating an adaptive surface convolution kernel. This surface convolution kernel is used to enhance the image edges along the tangential direction. This enables precise matching between mathematical processing and physical stretching scale, effectively avoiding non-physical artifacts, and accurately capturing real structural features even in extremely stretched areas.

[0038] 3. This invention addresses the shortcomings of traditional global unified evaluation methods that overemphasize smoothness while neglecting edge features. This method uses a spatial density weight map to weight the initial structural similarity, generating a distortion-weighted structural similarity. By determining the weight values, higher evaluation weights are assigned to low-density edge regions. This evaluation mechanism forces the algorithm to focus on edge semantic fidelity during optimization, avoiding feature loss and greatly improving the target recognition accuracy of machine vision under extreme field of view. Attached Figure Description

[0039] The invention will now be further described with reference to the accompanying drawings.

[0040] Figure 1 This is a flowchart illustrating a pixel-level edge enhancement and denoising method for fisheye lens images provided in an embodiment of this application. Detailed Implementation

[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0042] Please see Figure 1 A pixel-level edge enhancement and denoising method for fisheye lens images is disclosed. The method is applied to an image processing device and includes: the image processing device acquiring original fisheye image data and preset fisheye lens optical distortion parameters; the image processing device determining the distortion stretching rate of pixels in the original fisheye image data based on the fisheye lens optical distortion parameters; and the image processing device generating a spatial density weight map based on the distortion stretching rate.

[0043] The image processing device determines the target denoising threshold based on the spatial density weight map; the image processing device performs denoising processing on the original fisheye image data based on the target denoising threshold to generate an intermediate denoised image; the image processing device generates an adaptive surface convolution kernel based on the optical distortion parameters of the fisheye lens; the image processing device uses the adaptive surface convolution kernel to perform edge enhancement processing on the intermediate denoised image to generate an enhanced image.

[0044] Specifically, the method provided in this embodiment is applied to image processing devices equipped with fisheye lenses, such as image signal processors (ISPs) or related computing modules in systems like autonomous driving surround view, security panoramic monitoring, or VR devices; the image processing device acquires raw fisheye image data from sensor data in formats such as RAW or RGB, as well as preset fisheye lens optical distortion parameters;

[0045] Due to the inherent nonlinear optical distortion field and spatial information density gradient characteristics of fisheye lenses, the center resolution of the image is extremely high, while the edge field of view is extremely stretched.

[0046] Therefore, the image processing device calculates the distortion stretching rate corresponding to each pixel in the original fisheye image data based on the obtained fisheye lens optical distortion parameters such as the camera's intrinsic parameter matrix or the physical projection model. This distortion stretching rate quantifies the physical stretching scale of the image in different regions. A spatial density weight map of the same size as the original image is generated based on the distortion stretching rate.

[0047] The specific mapping relationship can use a normalized inverse proportional mapping function, that is, the spatial density weight of the current pixel is equal to a preset constant divided by the corresponding distortion stretching rate, and normalized to the range of 0 to 1; in the specific calculation, let the distortion stretching rate of the pixel be... The default constant is Typically, a stretching rate benchmark of 1.0 is taken from the distortion-free region at the center of the image. To avoid overflow from dividing by zero, a spatial density weight is set. The calculation formula is:

[0048]

[0049] in, To presuppose a minimum constant, such as taking Using the maximum and minimum weight values ​​obtained by traversing the entire graph, all weights are normalized using the minima-maximum method. Linear mapping to the range of 0 to 1; for example, in the spatial density weight map, the central region has a very high information density, so its corresponding weight is close to 1; while the closer to the edge region, the lower the information density is due to the forced stretching of the real physical space, the smaller its corresponding weight.

[0050] Subsequently, the image processing device changes the traditional global unified processing mode. Based on the generated spatial density weight map, it dynamically determines the target denoising threshold for different regions in the image, and performs adaptive denoising processing on the original fisheye image data according to the target denoising threshold to generate an intermediate denoised image.

[0051] For example, the target denoising threshold can be used as the brightness similarity weight attenuation parameter in a bilateral filtering algorithm or a nonlocal mean filtering (NLM) algorithm. When the target denoising threshold is large, the tolerance for differences between pixels can be relaxed, allowing more neighboring pixels to participate in the mean calculation to fully filter out noise.

[0052] When the target noise reduction threshold is small, the number of pixels involved in the calculation is strictly limited, so that only highly similar areas are smoothed, thereby preserving edge details.

[0053] To avoid the amplification of non-physical artifacts caused by optical dispersion in the edge region of the fisheye lens by traditional fixed square convolution kernels such as conventional 3x3 convolution kernels or Sobel operators, the image processing device further generates an adaptive surface convolution kernel based on the optical distortion parameters of the fisheye lens. This convolution kernel dynamically deforms according to the distortion vector at its location. Finally, the image processing device uses this adaptive surface convolution kernel to perform edge enhancement processing on the intermediate denoised image, outputting an enhanced image with uniform noise distribution and high-fidelity edge semantics.

[0054] This embodiment generates a spatial density weight map and an adaptive surface convolution kernel by introducing the nonlinear optical distortion parameters of a fisheye lens. Compared with conventional processing algorithms in the prior art that do not incorporate physical characteristics, this embodiment constructs an image processing engine that is aware of optical distortion.

[0055] This method enables the denoising intensity and enhancement direction of each pixel to be deeply coupled with the physical laws of the fisheye lens, effectively avoiding the risk of irreversible erasure of key semantic features at the edge of the field of view due to the mismatch between the mathematical processing scale and the physical stretching scale, and ensuring highly reliable feature preservation even in extreme fields of view.

[0056] Furthermore, this embodiment provides a step for determining a target denoising threshold based on a spatial density weight map, including: an image processing device determining whether the weight value in the spatial density weight map is greater than or equal to a preset weight threshold; if yes, the image processing device determines the target denoising threshold as a preset first denoising threshold; if no, the image processing device determines the target denoising threshold as a preset second denoising threshold; wherein, the first denoising threshold is greater than the second denoising threshold.

[0057] Specifically, in the process of determining the target denoising threshold, the image processing device will determine whether the weight value in the spatial density weight map is greater than or equal to the preset weight threshold pixel by pixel or region by region. The preset weight threshold is used to distinguish high-density areas such as the central region and low-density areas such as the edge of the extremely distorted region in the image. This embodiment does not limit the specific value of the weight threshold. Those skilled in the art can calibrate it according to the specific field of view (FOV) of the lens.

[0058] Regarding the method for obtaining the preset weight threshold, it can be obtained through the offline calibration stage: extract multiple fisheye images containing standard targets, calculate the average distortion stretching rate under different field of view, find the inflection point where the real physical area stretching rate suddenly increases and the image resolution begins to decrease significantly, and calibrate the spatial density weight value corresponding to the inflection point as the preset weight threshold; if the weight value is greater than or equal to the preset weight threshold, it means that the pixel is located in the central region with high information density and small optical distortion. At this time, the image processing device determines the target denoising threshold as the preset first denoising threshold to apply a stronger denoising constraint.

[0059] If the weight value is less than the preset weight threshold, it means that the pixel is located in the low-density edge area where the physical space is severely stretched. At this time, the image processing device determines the target denoising threshold to be the preset second denoising threshold, and the first denoising threshold is greater than the second denoising threshold.

[0060] For example, assuming the preset weight threshold is 0.6, for a pixel at the center of the image, its weight value may be 0.9, which is greater than 0.6. In this case, the image processing device uses a larger first denoising threshold to perform strong denoising to obtain highly clean image quality. However, for edge pixels with a field of view between 150° and 180°, their weight value may be only 0.2, which is less than 0.6. At this time, the algorithm determines that the pixel information density in this area is low and automatically reduces the denoising threshold, using a smaller second denoising threshold. In this case, reducing the denoising threshold acts as a physical buffer, allowing some noise to be retained in the edge area.

[0061] This embodiment dynamically allocates denoising thresholds based on spatial density weighting maps. In the high-density central area, a stronger denoising effort is used to ensure overall image quality, while in the low-density edge area, the denoising threshold is reduced to further ensure that faint, real high-frequency details such as the outlines of distant pedestrians or the edges of obstacles in blind spots are not erased as noise. This adaptive denoising mechanism, which uses the physical information density gradient as a constraint, maximizes the integrity of the true semantics of the edge image and lays a solid data foundation for avoiding missed reporting of critical events.

[0062] In a preferred embodiment of the present invention, this embodiment further provides a step of generating an adaptive surface convolution kernel based on the optical distortion parameters of a fisheye lens, comprising: an image processing device determining the optical distortion center based on the optical distortion parameters of the fisheye lens;

[0063] The image processing device constructs a radial-tangential orthogonal coordinate system based on the optical distortion center; the image processing device determines the target deformation direction according to the radial vector in the radial-tangential orthogonal coordinate system; the image processing device performs deformation processing on the preset initial convolution kernel according to the target deformation direction to generate an adaptive surface convolution kernel.

[0064] Specifically, after the denoising process is completed to obtain the intermediate denoised image, in order to perform accurate edge enhancement, the image processing device needs to generate a convolution kernel that matches the physical characteristics of the fisheye image. Since the straight lines in the real world in the fisheye image are stretched and bent into curves in the edge region, if the traditional fixed square convolution kernel such as the horizontal or vertical edge detection operator Sobel is used directly for processing, it will often fail in the edge region or amplify the non-physical artifacts caused by optical dispersion.

[0065] Therefore, in this embodiment, the optical distortion center of the lens on the image plane is determined based on the pre-acquired optical distortion parameters of the fisheye lens. Then, with the optical distortion center as the origin, a corresponding radial-tangential orthogonal coordinate system is constructed for each processed pixel in the image. In this coordinate system, the direction from the optical distortion center to the current pixel is the radial direction, and the direction perpendicular to the radial direction is the tangential direction.

[0066] The image processing device determines the target deformation direction of the pixel's location based on the radial vector in the radial-tangential orthogonal coordinate system. Since the physical distortion of the image increases non-linearly outward along the radial direction, the target deformation direction is closely related to the physical direction of this non-linear stretching. Based on the determined target deformation direction, the image processing device performs deformation processing on a preset initial convolution kernel, such as a standard 3x3 or 5x5 square convolution kernel.

[0067] For example, assuming the currently processed pixel is located at the edge of the image and its radial stretching is extremely severe, the image processing device will remap and interpolate the sampling point coordinates of the initial square convolution kernel along the vector direction affected by distortion, so that it bends and fits the actual physical stretching curvature of the region, thereby generating a dynamically deformed adaptive surface convolution kernel.

[0068] Specifically, the image processing device acquires the local coordinates of each sampling point of the initial square convolution kernel relative to the center point; calculates the physical offset of each local coordinate under nonlinear projection based on the curvature of the target deformation direction determined by the derivative of the local radial distance; and adds the physical offset to the local coordinates to obtain the deformed non-integer sampling coordinates.

[0069] The specific logic is as follows: Assume the local coordinates of a sampling point in the initial square convolution kernel relative to the center point are... The target deformation direction of the current pixel in the radial tangential orthogonal coordinate system corresponds to a radial deformation coefficient, which is obtained by differentiating the fisheye lens distortion polynomial with respect to the current radial distance. The fisheye lens distortion polynomial can adopt the conventional odd-order Taylor expansion model in this field, and its calculation formula is as follows:

[0070]

[0071] in This is the distorted radial distance. For the ideal, distortion-free radial distance, These are the preset optical distortion parameters for the fisheye lens;

[0072] The image processing device multiplies the radial component of the local coordinates by the radial deformation coefficient to determine the radial physical offset, while the tangential component remains unchanged, thereby accurately calculating the physical offset under extreme stretching, and superimposing it on the original local coordinates to obtain non-integer sampled coordinates;

[0073] This embodiment does not limit the specific size of the initial convolution kernel or the interpolation algorithm. Those skilled in the art can choose conventional methods such as bilinear interpolation or bicubic interpolation to resample the pixels of the original image using non-integer sampling coordinates, thereby completing the deformation processing.

[0074] This embodiment constructs an orthogonal coordinate system based on the optical distortion center and guides the convolution kernel to deform according to the radial vector, thus changing the defect of the fixed convolution kernel in traditional image processing;

[0075] By introducing the nonlinear optical distortion parameters of the fisheye lens, the generated adaptive surface convolution kernel can effectively match the optical deformation laws at different locations of the image, ensuring that the real structural features can be accurately captured even in areas with extremely stretched edges, and avoiding reconstruction deviations caused by the mismatch between mathematical processing scale and physical stretching scale.

[0076] Furthermore, this embodiment provides a step of performing edge enhancement processing on an intermediate denoised image using an adaptive surface convolution kernel to generate an enhanced image, including: an image processing device determining the tangent direction according to a radial-tangential orthogonal coordinate system; and the image processing device performing edge enhancement processing on the intermediate denoised image along the tangent direction using an adaptive surface convolution kernel to generate an enhanced image.

[0077] Specifically, after obtaining the adaptive surface convolution kernel customized for the current pixel, the image processing device needs to use the convolution kernel to perform a convolution operation on the intermediate denoised image; the image processing device accurately determines the tangent direction perpendicular to the radial vector based on the radial-tangential orthogonal coordinate system constructed above; in the physical laws of fisheye imaging, the edges of objects, especially the originally straight edges, are stretched after nonlinear projection, and the direction of the stretched edges is often highly consistent with the tangent direction.

[0078] Subsequently, the image processing device applies the generated adaptive surface convolution kernel strictly along the tangent direction to perform pixel-level edge enhancement processing on the intermediate denoised image;

[0079] It should be noted that standard two-dimensional image convolution operations require traversing pixels within the two-dimensional coordinate system of the image, while "along the tangent direction" here refers to the main sliding tracking trajectory of the convolution kernel when extracting continuous edge features along the direction of the edge, i.e., the tangent direction.

[0080] Meanwhile, based on the basic principles of image processing, if the physical orientation of an object's edge is parallel to the tangent direction, in order to extract and sharpen the high-frequency contour details of the edge, the weight distribution inside the adaptive surface convolution kernel is designed to perform gradient differentiation or contrast enhancement calculation in the direction that crosses the edge.

[0081] For example, when processing a pixel block containing a pedestrian outline located in the edge region of the field of view of 150° to 180°, the image processing device extracts the tangent direction at that location and makes the adaptive surface convolution kernel slide convolution tracking along the tangent direction, while using the differential weights of the convolution kernel in the radial direction to perform pixel-level sharpening across the edge.

[0082] Because the shape of the convolution kernel has been bent and deformed according to the distortion vector, and the sliding tracking trajectory of the enhancement operation strictly follows the tangent direction of the physical distortion, while the sharpening force is precisely applied in the radial direction across the edge, it can accurately act on the real object contour when extracting and amplifying high-frequency details, without mistakenly amplifying noise or specks caused by dispersion; after a comprehensive traversal and processing of the intermediate denoised image, the final output is an enhanced image with high-fidelity edge semantics and clear structure.

[0083] This embodiment uses an adaptive surface convolution kernel guided by the tangent direction to enhance the edge, thus deeply binding the edge sharpening operation with the physical projection law of the fisheye lens;

[0084] Compared to traditional algorithms that easily destroy weak, critical high-frequency details at the edge of the field of view, the nonlinear transformation and adaptive enhancement mechanism in this embodiment can effectively restore and enhance the extremely stretched real edge features, significantly improving the spatial perception accuracy under extreme fields of view, and providing high-quality image input for downstream machine vision systems to achieve highly reliable target detection and recognition in blind areas.

[0085] In a preferred embodiment of the present invention, this embodiment further provides a pixel-level edge enhancement and denoising method for fisheye lens images. The method further includes: an image processing device acquiring preset reference image data; the image processing device calculating an initial structural similarity based on the enhanced image and the reference image data; the image processing device weighting the initial structural similarity according to a spatial density weight map to generate a distortion-weighted structural similarity; and the image processing device adjusting preset parameters in the denoising or edge enhancement processing according to the distortion-weighted structural similarity.

[0086] Specifically, during the algorithm training or parameter tuning phase, in order to objectively evaluate the effect of image processing and optimize the model, the image processing device acquires preset reference image data, which is usually a high-quality calibration image with no noise and clear edges.

[0087] Subsequently, the image processing device calculates the initial structural similarity between the enhanced image output in the above embodiment and the preset reference image data, such as the conventional SSIM index. However, the traditional structural similarity calculation is globally uniform, which can easily lead to the algorithm removing weak high-frequency details with extremely stretched edges as noise in order to pursue the smoothness of the overall image.

[0088] Therefore, this embodiment no longer uses the global initial structural similarity alone, but performs pixel-level or region-level weighting on the initial structural similarity based on the spatial density weight map generated in the aforementioned steps, generating a new evaluation index, namely Distortion Weighted Structural Similarity DW-SSIM.

[0089] The image processing device multiplies the calculated initial structural similarity matrix element-wise with the assigned evaluation weight matrix to generate a distortion-weighted structural similarity matrix. Subsequently, the image processing device calculates the global average or sums all elements in the distortion-weighted structural similarity matrix to generate the final distortion-weighted structural similarity as a single scalar. To highlight the importance of structural features in edge regions, the first evaluation weight is greater than the second evaluation weight.

[0090] The image processing device uses the distortion-weighted structural similarity, which is a single scalar, as the objective function for optimization. It employs a gradient descent algorithm or a grid search strategy to calculate the feedback error of the parameters relative to the objective function. Based on this feedback error, it calculates the corresponding update step size and iteratively adjusts preset parameters such as the preset denoising threshold and the deformation coefficient of the initial convolution kernel in the denoising or edge enhancement processing until the rate of change of the objective function is less than the preset convergence threshold, so as to guide the algorithm to converge in the optimal direction.

[0091] This embodiment abandons the traditional evaluation method that pursues global peak signal-to-noise ratio (PSNR) or global smoothness, and innovatively constructs a distortion-weighted structural similarity evaluation function.

[0092] By using the spatial density weight map as the constraint and incentive term for evaluation, the algorithm is set to focus on areas with severe physical stretching during the optimization process. This ensures that the optimization direction serves high-fidelity edge semantics, avoids the damage to the spatial perception accuracy of the system caused by excessive pursuit of appearance indicators, and enables the final output image to effectively serve downstream machine vision recognition tasks.

[0093] Furthermore, this embodiment provides a step of weighting the initial structural similarity based on the spatial density weight map to generate a distortion-weighted structural similarity, including: the image processing device determining whether the weight value in the spatial density weight map is less than a preset center determination threshold;

[0094] If yes, the image processing device assigns a preset first evaluation weight to the initial structural similarity; if no, the image processing device assigns a preset second evaluation weight to the initial structural similarity; the image processing device multiplies the initial structural similarity by the assigned evaluation weight to generate a distortion-weighted structural similarity; wherein, the first evaluation weight is greater than the second evaluation weight.

[0095] Specifically, in the process of generating distortion-weighted structural similarity, the image processing device will determine whether the weight value in the spatial density weight map is less than the preset center determination threshold pixel by pixel or region by region.

[0096] The preset center determination threshold is used to divide the region of severe edge field distortion and the region of slight center field distortion in the image. This embodiment does not limit the specific value of the threshold, and those skilled in the art can set it according to the actual optical projection model.

[0097] Specifically, the center determination threshold can be obtained by offline analysis of the lens's MTF (Modulation Transfer Function) curve, finding the field of view position corresponding to when the MTF value decays to 50% of the peak value, and using the spatial density weight value corresponding to this position as the preset center determination threshold.

[0098] If the weight value is less than the preset center determination threshold, it means that the pixel is located in the edge region where the physical space is forcibly stretched and the information density is extremely low. In order to increase the weight of the region in the algorithm optimization, the image processing device assigns the region the preset first evaluation weight of the initial structural similarity.

[0099] If the weight value is greater than or equal to the preset center determination threshold, it means that the pixel is located in the central region with extremely high information density, and the image processing device assigns it a preset second evaluation weight.

[0100] The image processing device multiplies the calculated initial structural similarity with the assigned evaluation weights element by element to generate the final distortion-weighted structural similarity; where, in order to highlight the importance of the structural features of the edge region, the first evaluation weight is greater than the second evaluation weight.

[0101] Furthermore, the specific numerical acquisition method for the first evaluation weight and the second evaluation weight can adopt an allocation criterion based on information entropy: calculate the ratio of the local information entropy of the distortion-free reference image in the corresponding edge region and the center region, and use this information entropy ratio as the proportional coefficient of the first evaluation weight and the second evaluation weight to ensure that the setting of the constraint weight strictly corresponds to the amount of high-frequency information actually carried by the physical region.

[0102] For example, assuming the preset center determination threshold is 0.4, for edge pixels with a field of view between 150° and 180°, their weight value in the spatial density weight map may be only 0.2, less than 0.4. In this case, the image processing device assigns them a larger first evaluation weight, such as 1.5. For pixels at the center of the image, their weight value may be 0.8, not less than 0.4. In this case, the image processing device assigns them a smaller second evaluation weight, such as 0.8.

[0103] The image processing device multiplies the initial structural similarity of edge pixels by 1.5 and the initial structural similarity of center pixels by 0.8 to obtain the distortion-weighted structural similarity. When the algorithm over-smooths the edge region, causing feature loss, the weight multiplied by 1.5 will significantly amplify this error, forcing the image processing device to adjust the denoising intensity or enhancement direction in subsequent iterations.

[0104] This embodiment sets a center determination threshold and assigns differentiated evaluation weights, enabling the evaluation function to have physical perception capabilities. This mechanism forces the evaluation model to assign higher weights to the structural information of the image edge region during evaluation, so that the optimized algorithm can effectively retain the edge blind zone small target features such as pedestrians and traffic cones, greatly improving the target recognition accuracy under extreme field of view, and solving the technical pain point that the improvement of global image indicators in traditional algorithms leads to a significant decrease in the recognition rate of key edge targets.

[0105] In a preferred embodiment of the present invention, this embodiment further provides a step for determining the distortion stretching rate of pixels in the original fisheye image data based on the optical distortion parameters of the fisheye lens, including: an image processing device constructing a physical projection mapping model based on the optical distortion parameters of the fisheye lens; the image processing device inputting the coordinates of pixels in the original fisheye image data into the physical projection mapping model to calculate the actual physical area corresponding to the pixel; and the image processing device determining the distortion stretching rate based on the ratio of the actual physical area to the preset standard pixel area.

[0106] Specifically, when determining the distortion stretching rate of pixels in the original fisheye image data, the image processing device constructs a corresponding physical projection mapping model based on the acquired optical distortion parameters of the fisheye lens. In ordinary planar images, the actual physical area represented by one pixel is uniform, but in a fisheye lens, there exists a non-linear projection mapping, such as equidistant projection, the calculation formula of which is:

[0107]

[0108] If an equal solid angle projection is used, the calculation formula is:

[0109]

[0110] in, This represents the radial distance from a pixel on the image plane to the optical center. This indicates the focal length of the fisheye lens. This represents the angle of incidence between the incident ray and the optical axis; this physical projection mapping model achieves a nonlinear transformation.

[0111] The image processing device inputs the coordinates of each pixel in the original fisheye image data into the physical projection mapping model, and calculates the actual physical area of ​​the pixel in the real world through reverse calculation.

[0112] In practice, the image processing device constructs a pixel micro-element, such as a 2x2 pixel block, containing a preset number of pixels, centered on the current pixel point; obtains the image coordinates of each vertex of the pixel micro-element, and back-projects the image coordinates of each vertex to a preset three-dimensional world coordinate system or a unit sphere through the inverse function of the physical projection mapping model.

[0113] Calculate the area of ​​the polygon or solid angle enclosed by each vertex after back projection, and use it as the actual physical area corresponding to the pixel; the image processing device compares the actual physical area with the preset standard pixel area, calculates the ratio between the two, and accurately determines the distortion stretching rate of the pixel.

[0114] For example, for pixels in the central region of an image, the physical spatial deformation is minimal. After inputting into the physical projection mapping model, 10x10 pixels may only represent 0.1 square meters of actual physical area in the real world. The ratio of this area to the standard pixel area is close to 1, and the corresponding distortion stretching rate is small.

[0115] For areas near the edge of the image, the same 10x10 pixels may be forcibly stretched to represent the actual physical area of ​​10 square meters in the real world after non-linear projection. The calculated actual physical area is much larger than the standard pixel area, and the corresponding distortion stretching rate is extremely large.

[0116] This embodiment accurately quantifies the impact of nonlinear optical distortion of fisheye lens on pixel spatial information density by constructing a physical projection mapping model and solving the actual physical area.

[0117] Compared to conventional algorithms that perform uniform mathematical operations on images in existing technologies, this embodiment uses lens optical distortion parameters as the basic physical basis to reveal the physical stretching scale of different regions in the image, providing a precise quantitative indicator for avoiding the risk of irreversible erasure of key semantic features at the edge of the field of view.

[0118] Furthermore, this embodiment provides a pixel-level edge enhancement and denoising method for fisheye lens images. The method further includes: an image processing device sending the enhanced image to a preset machine vision recognition system so that the machine vision recognition system performs target detection processing based on the enhanced image.

[0119] Specifically, after the denoising processing based on the spatial density weight map and the edge enhancement processing based on the adaptive surface convolution kernel in the above embodiments, the image processing device generates an enhanced image with high-fidelity edge semantics and uniform noise distribution; the image processing device sends the enhanced image to a preset downstream machine vision recognition system;

[0120] After receiving the enhanced image, the machine vision recognition system uses its internally deployed detection algorithm model based on convolutional neural networks to perform subsequent target detection processing based on the high-frequency details and real structural features preserved in the enhanced image.

[0121] Specifically, the machine vision recognition system in this embodiment is deployed in a preset autonomous driving device, such as the surround view system of an autonomous vehicle or an ADAS advanced driver assistance system. When the machine vision recognition system receives the enhanced image, it extracts the feature map of the edge region in the enhanced image through the feature extraction backbone network in the detection algorithm model and parses out the semantic features.

[0122] Since the front-end image processing device has already performed physical buffering and edge enhancement in the tangential direction for the areas where the edges are extremely stretched, the original faint high-frequency details of the edge area are preserved. The machine vision recognition system can extract these key semantic features. The machine vision recognition system inputs the extracted semantic features into the detection head network, and through bounding box regression and classification calculation, generates target detection results containing information such as target category and location.

[0123] Ultimately, based on the target detection results, the machine vision recognition system generates control command data to indicate the operating status of the autonomous driving equipment and sends it to the vehicle's underlying control module.

[0124] For example, when an autonomous driving device is in motion, its fisheye lens can detect a pedestrian or a traffic cone that suddenly appears in the edge blind spot with a field of view of 150° to 180°.

[0125] The machine vision recognition system accurately extracts the contour semantic features of pedestrians in the edge region from the enhanced image, thereby generating a target detection result of pedestrians in the side and rear blind spot; based on the target detection result, the system quickly generates control command data for emergency braking or fine-tuning the steering wheel in the opposite direction to instruct the autonomous driving equipment to take evasive action.

[0126] This embodiment applies image processing methods to autonomous driving scenarios. By extracting high-fidelity edge semantic features from enhanced images, the accuracy of small target recognition in edge blind spots can be significantly improved.

[0127] This solution not only effectively reduces the missed detection rate of critical events and lowers the risk of traffic accidents, but also avoids ineffective over-computation by introducing spatial information density weights, saving computing power consumption of edge devices and maximizing the application of fisheye lens systems.

[0128] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A pixel-level edge enhancement and denoising method for fisheye lens images, characterized in that, The method is applied to an image processing device, and the method includes: the image processing device acquiring raw fisheye image data and preset fisheye lens optical distortion parameters; The image processing device determines the distortion stretching rate of pixels in the original fisheye image data based on the optical distortion parameters of the fisheye lens. The image processing device generates a spatial density weight map based on the distortion stretching rate; the image processing device determines the target denoising threshold based on the spatial density weight map; The image processing device performs denoising processing on the original fisheye image data according to the target denoising threshold to generate an intermediate denoised image; The image processing device generates an adaptive surface convolution kernel based on the optical distortion parameters of the fisheye lens; The image processing device uses the adaptive surface convolution kernel to perform edge enhancement processing on the intermediate denoised image to generate an enhanced image; Determining the target denoising threshold based on the spatial density weight map includes: the image processing device determining whether the weight value in the spatial density weight map is greater than or equal to a preset weight threshold. If so, the image processing device determines the target denoising threshold as a preset first denoising threshold; If not, the image processing device determines the target denoising threshold to be a preset second denoising threshold; wherein the first denoising threshold is greater than the second denoising threshold; Let the distortion stretching rate of a pixel be... The default constant is The stretching ratio baseline value of 1.0 is taken from the distortion-free region at the center of the image, and the spatial density weight is... The calculation formula is: ; in, To predetermine a minimum constant; using the maximum and minimum weight values ​​obtained by traversing the entire graph, all weights are normalized using a maximum-minimum normalization method. Linearly mapped to the interval between 0 and 1; The process of generating an adaptive surface convolution kernel based on the optical distortion parameters of the fisheye lens includes: the image processing device determining the optical distortion center based on the optical distortion parameters of the fisheye lens; and the image processing device constructing a radial tangential orthogonal coordinate system based on the optical distortion center. The image processing device determines the target deformation direction based on the radial vector in the radial-tangential orthogonal coordinate system; The image processing device deforms the preset initial convolution kernel according to the target deformation direction to generate the adaptive surface convolution kernel; The step of using the adaptive surface convolution kernel to perform edge enhancement processing on the intermediate denoised image to generate an enhanced image includes: the image processing device determining the tangent direction according to the radial tangential orthogonal coordinate system; The image processing device performs edge enhancement processing on the intermediate denoised image by applying the adaptive surface convolution kernel along the tangent direction, thereby generating the enhanced image.

2. The pixel-level edge enhancement and denoising method for fisheye lens images according to claim 1, characterized in that, The method further includes: the image processing device acquiring preset reference image data; the image processing device calculating an initial structural similarity based on the enhanced image and the reference image data; The image processing device performs weighted processing on the initial structural similarity according to the spatial density weight map to generate a distortion-weighted structural similarity. The image processing device adjusts the preset parameters in the denoising process or the edge enhancement process based on the distortion weighted structural similarity.

3. The pixel-level edge enhancement and denoising method for fisheye lens images according to claim 2, characterized in that, The step of weighting the initial structural similarity according to the spatial density weight map to generate a distortion-weighted structural similarity includes: the image processing device determining whether the weight value in the spatial density weight map is less than a preset center determination threshold. If yes, the image processing device assigns a preset first evaluation weight to the initial structural similarity; if no, the image processing device assigns a preset second evaluation weight to the initial structural similarity. The image processing device multiplies the initial structural similarity by the assigned evaluation weight to generate the distortion-weighted structural similarity; wherein the first evaluation weight is greater than the second evaluation weight.

4. The pixel-level edge enhancement and denoising method for fisheye lens images according to claim 1, characterized in that, The step of determining the distortion stretching rate of pixels in the original fisheye image data based on the optical distortion parameters of the fisheye lens includes: the image processing device constructing a physical projection mapping model based on the optical distortion parameters of the fisheye lens; The image processing device inputs the coordinates of the pixels in the original fisheye image data into the physical projection mapping model to calculate the actual physical area corresponding to the pixel. The image processing device determines the distortion stretching rate based on the ratio of the actual physical area to the preset standard pixel area.

5. The pixel-level edge enhancement and denoising method for fisheye lens images according to claim 1, characterized in that, The method further includes: the image processing device sending the enhanced image to a preset machine vision recognition system, so that the machine vision recognition system performs target detection processing based on the enhanced image.

6. The pixel-level edge enhancement and denoising method for fisheye lens images according to claim 5, characterized in that, The machine vision recognition system is deployed in a pre-set autonomous driving device; The machine vision recognition system performs target detection processing based on the enhanced image, including: the machine vision recognition system extracts semantic features of edge regions in the enhanced image; The machine vision recognition system generates target detection results based on the semantic features; the machine vision recognition system generates control command data to indicate the operating status of the autonomous driving equipment based on the target detection results.