Image enhancement method based on ISP image signal processing visual sensor
Through the multi-frame image acquisition and processing mechanism, combined with analog-to-digital conversion and timestamp recording, dark current correction, lens shadow correction, multi-scale feature extraction and LSTM network are used to solve the distortion and noise problems of ISP vision sensors during image acquisition, improve the clarity and detail fidelity of the image, and achieve high-quality image enhancement effect.
Patent Information
- Application Number
- CN202510501483.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional ISP vision sensors face problems such as dark current noise, lens distortion and uneven light during image acquisition, which makes it difficult to meet the actual application requirements. The existing enhancement methods lack adaptability and timing-related information utilization, affecting the visual quality and practicality of the image.
A multi-frame image acquisition and processing mechanism is adopted, combined with analog-to-digital conversion and timestamp recording, through dark current correction and lens shadow correction, SIFT feature matching and affine transformation, a multi-scale feature extraction network and LSTM recursive neural network based on CFP-ISP are designed, combined with the attention mechanism of channel and spatial dimensions, feature fusion and deconvolution processing are performed, structural layer edge enhancement and texture layer denoising are performed, and the dynamic range and color restoration of the image are optimized.
It significantly reduces distortion and noise interference during image acquisition, improves image clarity and detail fidelity, enhances image timing correlation and visual quality, and meets application needs in fields such as intelligent manufacturing, security monitoring and medical diagnosis.
Smart Images

Figure CN120339100A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image enhancement, and in particular, to an image enhancement method based on an ISP image signal processing vision sensor. Background Art
[0002] The image enhancement technology based on the ISP image signal processing vision sensor has been widely applied in the fields of intelligent manufacturing, security monitoring, medical diagnosis, etc. However, traditional ISP vision sensors often face problems such as dark current noise, lens distortion, uneven illumination, etc. during the image acquisition process, resulting in the quality of the output image being difficult to meet the actual application requirements.
[0003] Currently, most mainstream image enhancement methods only focus on the processing of single-frame images, ignoring the temporal correlation information between multiple frames of images, and unable to make full use of the redundant information in the continuous acquisition process to improve the image quality. At the same time, existing enhancement algorithms often adopt a fixed processing flow, lacking the adaptive ability to different scenes and lighting conditions, and the enhancement effect in complex environments is not ideal. In addition, in the process of feature extraction and fusion of traditional ISP processing methods, the importance of different scales and types of features cannot be effectively distinguished, easily resulting in the loss or over-enhancement of detail information. In the image reconstruction stage, the balanced processing of structural information and texture details also faces challenges, affecting the visual quality and practicality of the final enhanced image. Summary of the Invention
[0004] The present invention provides an image enhancement method based on an ISP image signal processing vision sensor, which reduces various distortions and noise interferences in the image acquisition process and ensures the clarity and detail fidelity of the enhanced image.
[0005] In a first aspect, the present invention provides an image enhancement method based on an ISP image signal processing vision sensor, and the image enhancement method based on the ISP image signal processing vision sensor includes: Collect multiple frames of images of the same scene, perform analog-to-digital conversion on the collected multiple frames of images and generate an RGB three-channel digital matrix, record the timestamp of each frame of image, and obtain an N-frame original image sequence, where N is an integer greater than or equal to 3; Perform dark current correction and lens shadow correction on the N-frame original image sequence, calculate the affine transformation matrix from the N-1th frame to the first frame of image based on SIFT feature point extraction and matching, and obtain an image sequence after spatial alignment; Input the image sequence after spatial alignment into the CFP-ISP feature extraction network for spatial feature extraction to obtain a multi-scale feature map sequence, and input the multi-scale feature map sequence into the LSTM recurrent neural network for residual feature connection operation to obtain a temporal correlation feature map sequence; Perform attention calculations on the sequence of time-series correlation feature maps in the channel dimension and the spatial dimension to generate a two-dimensional fusion feature map, and input the two-dimensional fusion feature map into a three-layer transposed convolutional network for upsampling. At the same time, perform structural layer edge enhancement and texture layer denoising processing to obtain the target enhanced image.
[0006] In a second aspect, the present invention provides an image enhancement device based on an ISP image signal processing vision sensor. The image enhancement device based on an ISP image signal processing vision sensor includes: An acquisition module, configured to acquire multiple frames of images of the same scene, perform analog-to-digital conversion on the acquired multiple frames of images and generate an RGB three-channel digital matrix, record the timestamp of each frame of image, and obtain an N-frame original image sequence, where N is an integer greater than or equal to 3; A correction module, configured to perform dark current correction and lens shading correction on the N-frame original image sequence, and calculate the affine transformation matrix from the N-1th frame to the first frame image based on SIFT feature point extraction and matching to obtain an image sequence after spatial alignment; An extraction module, configured to input the spatially aligned image sequence into a CFP-ISP feature extraction network for spatial feature extraction to obtain a multi-scale feature map sequence, and input the multi-scale feature map sequence into an LSTM recurrent neural network for residual feature connection operations to obtain a time-series correlation feature map sequence; A generation module, configured to perform attention calculations on the time-series correlation feature map sequence in the channel dimension and the spatial dimension to generate a two-dimensional fusion feature map, and input the two-dimensional fusion feature map into a three-layer transposed convolutional network for upsampling. At the same time, perform structural layer edge enhancement and texture layer denoising processing to obtain the target enhanced image.
[0007] In a third aspect, the present invention provides an image enhancement device based on an ISP image signal processing vision sensor, including: a memory and at least one processor, where instructions are stored in the memory; the at least one processor calls the instructions in the memory so that the image enhancement device based on an ISP image signal processing vision sensor executes the above-mentioned image enhancement method based on an ISP image signal processing vision sensor.
[0008] In a fourth aspect of the present invention, there is provided a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a computer, the computer is caused to execute the above-mentioned image enhancement method based on an ISP image signal processing vision sensor.
[0009] In the technical solution provided by the present invention, by designing a multi-frame image acquisition and processing mechanism, combined with analog-to-digital conversion and timestamp recording, the integrity and temporal consistency of image information are effectively improved, providing high-quality input data for subsequent enhancement processing. A dual correction mechanism of dark current correction and lens shading correction, combined with SIFT feature matching and affine transformation, significantly reduces various distortions and noise interferences during image acquisition. A multi-scale feature extraction network based on CFP-ISP is designed, and through convolutional kernels of different sizes and ReLU activation functions, multi-level and multi-scale extraction of image features is achieved. An LSTM recurrent neural network is innovatively introduced for temporal feature processing, and through a three-gate control mechanism and residual connection, the temporal correlation between image sequences is enhanced. Combining a dual attention mechanism in the channel dimension and spatial dimension, adaptive weighting and fusion of features are realized, improving the expression ability of important features. A three-layer transposed convolution network is used for feature upsampling, combined with a hierarchical processing strategy of edge enhancement in the structure layer and denoising in the texture layer, ensuring the clarity and detail fidelity of the enhanced image. Through HDR tone mapping and automatic white balance technology, the dynamic range and color restoration effect of the image are optimized, improving the visual quality of the enhancement result.
[0010] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification, or be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims and drawings.
[0011] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a schematic diagram of an embodiment of an image enhancement method based on an ISP image signal processing vision sensor in an embodiment of the present invention; Figure 2 It is a schematic diagram of an embodiment of an image enhancement device based on an ISP image signal processing vision sensor in an embodiment of the present invention; Figure 3 It is a schematic diagram of an embodiment of an image enhancement device based on an ISP image signal processing vision sensor in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0014] As used in the embodiments of the present invention, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes other steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0015] To facilitate the understanding of this embodiment, first, a detailed introduction will be given to an image enhancement method based on an ISP image signal processing vision sensor disclosed in the embodiments of the present invention. As Figure 1 shown, this method includes the following steps: 101. Collect multiple frames of images of the same scene, perform analog-to-digital conversion on the collected multiple frames of images and generate an RGB three-channel digital matrix, record the timestamp of each frame of image, and obtain an N-frame original image sequence, where N is an integer greater than or equal to 3; It can be understood that the execution subject of the present invention can be an image enhancement device based on an ISP image signal processing vision sensor, or a terminal or a server. Specifically, no limitation is made here. In the embodiments of the present invention, the server is taken as an example of the execution subject for illustration.
[0016] Specifically, continuous image sampling is performed on the same scene through an ISP vision sensor. The sampling frequency is set to 30 frames per second, and the sampling duration is set to T seconds, obtaining T×30 frames of raw image data. During this process, the sensor captures image frames at a fixed frequency to ensure continuity and stability. While image sampling is taking place, a real-time light sensor monitors the illuminance of the surrounding environment. By measuring the intensity of the light, the corresponding environmental illuminance detection result is obtained, and this result falls within the range of 0 to 1000 lux. Based on this illuminance result, the LED fill light coefficient is dynamically calculated to generate light compensation data for improving the image quality in low-light or high-light environments. An analog-to-digital conversion operation is performed on the light compensation data to convert the analog signal into a digital signal so that it can be accepted by subsequent digital processing modules. After completing the analog-to-digital conversion, the digital image data is separated into channels according to the standard RGB color space, and the red channel (R), green channel (G), and blue channel (B) are respectively extracted. To ensure computational stability and result consistency in subsequent processing, normalization operations are performed on the data of these three channels respectively, normalizing the data values of each channel to the range of [0, 255]. The normalized R channel, G channel, and B channel are respectively combined to perform a channel merging operation. By rearranging the three-channel digital matrix according to the arrangement rules of the RGB color space, a complete RGB three-channel digital matrix is generated. This matrix contains the complete information of each frame of the image on different color channels. To record the time information of each frame of the image, based on the high-precision system clock of the ISP processor, the generated RGB three-channel digital matrix is marked with a microsecond-level timestamp. The addition of the timestamp gives each frame of the image clear time series information, thus ensuring that the frames can be arranged and processed in chronological order during subsequent processing. Marking the timestamp generates an image matrix sequence with serial numbers. Based on this image matrix sequence with serial numbers, the quality of each frame of the image is evaluated, and the image is scored using various indicators such as image sharpness, brightness distribution, and contrast. The N frames of the original image sequence with the highest scores are selected.
[0017] 102. Perform dark current correction and lens shading correction on the N-frame original image sequence. Based on SIFT feature point extraction and matching, calculate the affine transformation matrix from the N - 1th frame to the first frame of the image to obtain an image sequence after spatial alignment; Specifically, in order to eliminate the influence of the dark current noise signal, perform a no-light illumination scan on the N-frame original image sequence. By turning off the light source or shielding the sensor, the corresponding dark current noise signal is obtained. This dark current noise signal is a numerical representation of the fixed pattern noise generated by the sensor in a low-light or no-light environment. Perform a pixel-by-pixel subtraction operation on the dark current noise signal and each frame of the original image data to eliminate the influence of these noises. At the same time, combined with the linear stretching compensation technique, adjust the dynamic range of the image data after the subtraction operation to restore it to an ideal brightness and contrast distribution state, generating an image sequence after dark current correction. To solve the problem of optical distortion introduced by the lens, calculate the brightness mean and variance of each frame of the image in the image sequence after dark current correction respectively, and construct a lens shadow distribution map based on these statistical data. The lens shadow distribution map describes the uneven distribution of image brightness caused by the optical characteristics of the lens, especially the phenomenon of reduced brightness in the edge part. Using this lens shadow distribution map, perform a pixel-by-pixel compensation operation on each frame of the image to restore the brightness uniformity of the image. At the same time, for the problems of radial distortion and tangential distortion introduced by the lens, adopt a projective transformation method to eliminate these distortions through precise geometric correction, generating an image sequence after lens correction. Based on the first frame of the image sequence after lens correction, calculate the difference between adjacent scale images. Through the calculation of the difference between adjacent scale images, construct a multi-scale pyramid, and detect extreme points in the scale space to obtain preliminary scale space feature points. Input the scale space feature points into the SIFT feature extractor, calculate the main direction and 128-dimensional feature descriptor of the local area for each feature point, generating a set of feature points for the first frame of the image. Similarly, perform feature extraction operations on the remaining N - 1 frames of the image sequence after lens correction to obtain N - 1 sets of feature points to be matched. Based on the nearest neighbor principle of Euclidean distance, perform feature matching between the set of feature points of the first frame of the image and the sets of feature points to be matched of the remaining N - 1 frames of the image. Use the bidirectional nearest neighbor method or the ratio test method to screen out a reliable sequence of matching point pairs. Based on the reliable sequence of matching point pairs, by establishing a least squares equation, minimize the error of the matching point pairs, and use the singular value decomposition method to accurately solve the affine transformation matrix. The affine transformation matrix can describe the spatial relationships such as rotation, scaling, and translation between images. Using this matrix, transform the N - 1 frames of the image from their original coordinate system to the coordinate system of the first frame of the image to ensure the spatial consistency of all images. Finally, generate an image sequence after spatial alignment.
[0018] 103. Input the image sequence after spatial alignment into the CFP-ISP feature extraction network for spatial feature extraction to obtain a sequence of multi-scale feature maps, and input the sequence of multi-scale feature maps into the LSTM recurrent neural network for residual feature connection operations to obtain a sequence of temporal correlation feature maps; Specifically, the spatially aligned image sequence is input into the first convolutional network of the CFP-ISP feature extraction network. This convolutional network uses a convolutional kernel of size 3×3 for feature extraction operations. The stride of the convolution is set to 1, and the number of padding pixels is 1. After the convolution operation, the feature map undergoes a non-linear transformation through the ReLU activation function to effectively capture the low-level spatial features of the image while maintaining computational efficiency, resulting in the first feature map. The first feature map is input into the second convolutional network of the CFP-ISP feature extraction network for processing. The size of the convolutional kernel of the second convolutional network is set to 5×5, the convolutional stride is set to 2, and the number of padding pixels is 2, reducing the spatial resolution of the feature map to half of that of the first feature map, thereby improving the compactness and computational efficiency of the features. After convolution, the data is processed through the ReLU activation function and important features are extracted through a 2×2 max-pooling operation while reducing redundant information, generating the second feature map. The second feature map is input into the 7×7 convolutional layer of the third convolutional network. A convolutional kernel of size 7×7 is used to extract features from the input data. The convolutional stride is set to 2, and the number of padding pixels is 3 to ensure that a larger receptive field can capture the middle and high-level spatial features of the image. After the convolution operation, the data is processed through the ReLU activation function, and a 2×2 max-pooling operation is performed to compress the size of the feature map while retaining key information, resulting in the third feature map. The third feature map is input into the fourth convolutional network of the CFP-ISP feature extraction network for deeper feature extraction. This convolutional network uses a convolutional kernel of 9×9, the convolutional stride is set to 2, and the number of padding pixels is 4, effectively extracting the high-level semantic features of the image while expanding the receptive field. After the convolution process, the data undergoes a non-linear transformation through the ReLU activation function and a 2×2 max-pooling operation, generating the fourth feature map, which represents the deepest feature information in the input image sequence. The above four feature maps are used to perform a feature alignment operation on the data to ensure the spatial consistency of features with different resolutions and scales, generating an aligned feature sequence. The aligned feature sequence is processed through a channel attention module. In this module, the channel weights of the feature sequence are calculated, and weights are dynamically assigned according to the importance of each channel for feature expression; then the feature sequence is weighted and calibrated to highlight the expression ability of significant features, resulting in the final multi-scale feature map sequence. The multi-scale feature map sequence is input into an LSTM recurrent neural network to capture the temporal dependencies between different frames in the image sequence. The LSTM network selectively retains and updates features through its unique input gate, forget gate, and output gate mechanisms, retaining both important temporal information and filtering out redundant or irrelevant features. At the same time, the network performs a residual feature connection operation to strengthen the temporal correlation of the features, generating the final sequence of temporally correlated feature maps.
[0019] Input the sequence of multi-scale feature maps into the input gate of the LSTM recurrent neural network. In the input gate, the current input features and the input weight matrix are operated on through the sigmoid activation function to calculate the weight values of the features. These weight values are used to selectively filter the current input features, only retaining the important features while suppressing irrelevant or redundant information, and generating input gated features. The input gated features are subjected to a tanh activation operation, and the features are restricted to the range of [-1, 1] through a non-linear mapping operation to enhance the network's ability to model non-linear relationships. The activated features are element-wise multiplied by the input weight matrix to generate candidate memory units. The candidate memory units are input into the forget gate module of the LSTM network. The role of the forget gate is to selectively forget the memory state of the previous moment of the network, ensuring that the network can adapt to dynamically changing input features. In the forget gate, the sigmoid activation function is used to calculate the forget weight matrix, and the memory state of the previous moment is weighted to generate forget gated features. The design of the forget gate enables the LSTM network to dynamically retain or discard past information according to the requirements of the current task, thus avoiding the ineffective accumulation of memory states or the interference of redundant features. The forget gated features and the candidate memory units are subjected to a weighted summation operation to update the memory unit state at the current moment and generate a new memory state. By combining the valid information from the past and the input features at the current moment, the content of the memory state is dynamically adjusted, enabling the network to retain key features while promptly adapting to changes in input features, thereby enhancing the model's temporal modeling ability and memory ability. After updating the memory unit, it is input into the output gate of the LSTM network. The output gate calculates the output weight matrix through the sigmoid activation function, selectively outputs the features at the current moment, and generates output gated features. The role of the output gate is to determine which information needs to be passed to the next layer of the network or the output state, and which information can be ignored, thus effectively controlling the information flow of the network. The updated memory unit is subjected to a tanh activation operation to adjust the memory state through a non-linear transformation. The activated memory unit is element-wise multiplied by the output gated features to generate the output state at the current moment. By combining the roles of the memory unit and the output gate, the state information with temporal correlation and key feature expression is output. To enhance the temporal correlation of the feature map sequence and the transmission efficiency of the information flow, a residual connection path is established based on the multi-scale feature map sequence. In the residual path, the input multi-scale feature map and the output state at the current moment are subjected to an element-wise addition operation to retain the original input features while combining the output state of the LSTM network to form the final sequence of temporally correlated feature maps.
[0020] 104. Perform attention calculations on the sequence of temporal correlation feature maps in the channel dimension and the spatial dimension to generate a two-dimensional fusion feature map, and input the two-dimensional fusion feature map into a three-layer transposed convolutional network for upsampling. At the same time, perform structural layer edge enhancement and texture layer denoising processing to obtain the target enhanced image.
[0021] Specifically, a global average pooling operation is performed on each temporal correlation feature map of the temporal correlation feature map sequence to obtain an initial channel feature vector. The initial channel feature vector is input into the first fully connected layer for dimensionality reduction transformation to reduce the dimension of the feature vector and reduce the computational complexity to obtain an intermediate feature vector. The intermediate feature vector is input into the second fully connected layer for dimensionality increase transformation to restore the original number of channels and generate a channel attention weight vector. By gradually reducing and increasing the dimensionality, the network can better learn the relationship between channels, thereby dynamically adjusting the weight of each channel. The channel attention weight vector is multiplied by the temporal correlation feature map sequence channel by channel, and each channel is weighted to obtain a channel enhanced feature sequence. By strengthening the expression ability of key channels and suppressing the information of irrelevant or redundant channels, the channel expression ability of the feature map is effectively improved. The spatial dimension attention calculation is performed on the channel enhanced feature sequence. A 3×3 convolution operation is performed on the channel enhanced feature sequence to extract local spatial features. The convolved feature maps are processed by two branches, namely, maximum pooling and average pooling, to extract spatial saliency information from different angles and generate the first spatial feature map and the second spatial feature map. These two spatial feature maps represent the maximum value distribution and average value distribution of the input features in the local area, respectively, and can effectively capture important information at the spatial level. The first spatial feature map and the second spatial feature map are spliced in the channel dimension to form a feature map containing richer spatial information. The spliced feature map is fused through a 1×1 convolution layer to integrate different spatial features, and the fused feature map is normalized using the sigmoid function to generate a spatial attention weight map. This weight map is used to highlight the important areas with spatial significance in the feature map while reducing the influence of unimportant areas. The channel enhancement feature sequence is multiplied element by element with the spatial attention weight map, and the channel and space attention mechanisms are combined to generate a preliminary spatiotemporal attention feature sequence. On this basis, in order to enhance the transitivity of the original features, the residual connection mechanism is adopted to add the spatiotemporal attention feature sequence to the original temporal association feature map sequence element by element to generate a complete spatiotemporal attention feature sequence. Based on the spatiotemporal attention feature sequence, the attention coefficient of each temporal position is calculated, and a weighted average fusion operation is performed to fuse the N-frame feature sequence into a single-frame expression to obtain a two-dimensional fused feature map. The two-dimensional fusion feature map is input into the three-layer deconvolution network for upsampling. Through layer-by-layer deconvolution and feature decoding, the feature map is gradually restored to the target resolution to generate the upsampled feature map. The structural layer edge enhancement processing is performed on the upsampled feature map to enhance the structural information and contour clarity in the image. At the same time, the texture layer is denoised to eliminate noise interference and improve the image's detail expression ability. The target enhanced image is generated.
[0022] The two-dimensional fusion feature map is input into the first deconvolution network of the three-layer deconvolution network. A 4×4 deconvolution kernel with a stride of 2 is used to perform 2x upsampling on the input feature map. The spatial resolution of the feature map is expanded to twice the original to restore finer spatial details. After the deconvolution calculation is completed, the ReLU activation function is used to perform a non-linear transformation on the upsampling result to enhance the expression ability of the network. At the same time, batch normalization is performed on the result to reduce the distribution difference between different channels in the feature map, obtaining the first-layer feature map at a 2x scale. The first-layer feature map is passed as input to the second deconvolution network, and the same 4×4 deconvolution kernel with a stride of 2 is used to perform 2x upsampling again. The resolution of the feature map is increased to 4 times the original, generating a higher-resolution feature map. Similarly, the ReLU activation function and batch normalization are applied to the upsampling result of the second layer to ensure the non-linear expression ability and numerical stability of the features, obtaining the second-layer feature map at a 4x scale. The second-layer feature map is passed to the third layer of the three-layer deconvolution network, and a 2x upsampling operation is performed through a 4×4 deconvolution kernel with a stride of 2, increasing the resolution to 8 times the original feature map, and completely restoring the feature map at the target resolution. After the deconvolution operation is completed, the ReLU activation function and batch normalization are applied to the upsampling result of the third layer to obtain the final upsampling feature map at an 8x scale. After generating the upsampling feature map at an 8x scale, in order to enhance the structural layer information and texture layer detail information of the image respectively, multi-scale decomposition is performed on the upsampling feature map. Through multi-scale decomposition, the feature map is separated into a structural layer feature map and a texture layer feature map. The structural layer feature map mainly contains the contour and edge information of the target, while the texture layer feature map contains finer image details. For the structural layer feature map, a 3×3 gradient operator in the horizontal and vertical directions is used for gradient calculation to extract edge features. Gradient calculation can effectively enhance the edge structure in the image, making the contour clearer, especially more obvious in the high-frequency region. At the same time, for the texture layer feature map, a special denoising algorithm is used to process it, removing unnecessary noise interference while retaining important detail information. The purpose of texture denoising is to optimize the detail expression and eliminate random noise introduced by low-quality inputs or interference from the sensor itself. Through this process, a denoised texture feature map is generated. The enhanced structural edge features and the denoised texture features are weighted and superimposed. The structural information and texture details are fused to generate a high-quality enhanced fusion feature. HDR (High Dynamic Range) tone mapping and color distribution correction are performed on the enhanced fusion feature. HDR tone mapping enhances the detail expression ability of the image in the dark and bright regions by readjusting the brightness range, making the image closer to the visual perception effect of the human eye; while color distribution correction corrects the color deviation to ensure that the colors of the image are more real and natural. Through the above steps, the target enhanced image is generated.
[0023] In the embodiments of the present invention, by designing a multi-frame image acquisition and processing mechanism, combined with analog-to-digital conversion and timestamp recording, the integrity and temporal consistency of image information are effectively improved, providing high-quality input data for subsequent enhancement processing. A dual correction mechanism of dark current correction and lens shading correction is adopted, combined with SIFT feature matching and affine transformation, significantly reducing various distortions and noise interferences in the image acquisition process. A multi-scale feature extraction network based on CFP-ISP is designed, and through convolutional kernels of different sizes and ReLU activation functions, multi-level and multi-scale extraction of image features is achieved. An LSTM recurrent neural network is innovatively introduced for temporal feature processing, and through a three-gate control mechanism and residual connection, the temporal correlation between image sequences is enhanced. Combining the dual attention mechanisms of channel dimension and spatial dimension, adaptive weighted sum and fusion of features are realized, improving the expression ability of important features. A three-layer deconvolution network is used for feature upsampling, combined with a hierarchical processing strategy of edge enhancement in the structure layer and denoising in the texture layer, ensuring the clarity and detail fidelity of the enhanced image. Through HDR tone mapping and automatic white balance technology, the dynamic range and color restoration effect of the image are optimized, improving the visual quality of the enhancement result.
[0024] In a specific embodiment, the process of executing step 101 may specifically include the following steps: Continuously sample images of the same scene through an ISP vision sensor, with the sampling frequency set to 30 frames per second and the sampling duration set to T seconds, obtaining T×30 frames of image data; Detect the ambient illuminance of the T×30 frames of image data through a real-time light sensor, obtain the ambient illuminance detection result, and calculate the LED fill light coefficient within the range of 0 - 1000 lux based on the ambient illuminance detection result to obtain the light compensation data; Perform analog-to-digital conversion on the light compensation data to obtain quantized digital image data, and separate the channels of the quantized digital image data according to the RGB color space, and perform normalization operations in the range of [0, 255] on the R channel, G channel, and B channel respectively to obtain three normalized digital matrices; Perform a channel merging operation on the three normalized digital matrices, arrange and combine the digital matrices of the R, G, and B channels according to the color space to obtain an RGB three-channel digital matrix; Based on the system clock of the ISP processor, mark microsecond-level timestamp information for the RGB three-channel digital matrix to obtain an image matrix sequence with serial numbers, and perform image quality evaluation based on the image matrix sequence with serial numbers to screen out the N-frame original image sequence with the highest score.
[0025] Specifically, continuously sample images of the same scene through an ISP vision sensor, and set the sampling frequency as , and the sampling duration as The total number of sampled frames is expressed as: ; The image frames obtained by sampling are denoted as , where represents the th frame, represents the spatial coordinates of the pixel represents the RGB color channels. The ambient illuminance of each frame of image data is detected by a real-time light sensor. Let the illuminance be (unit: Iux), and the range is [0, 1000]. According to the illuminance , calculate the corresponding LED fill light coefficient The formula for is: ; Among them, represents the fill light intensity ratio for the th frame. When , , it means that no fill light is required; when , , it means that full fill light is required. The image data after fill light is subjected to analog-to-digital conversion, and the quantization range is [0, 255]. The quantized pixel value is , and its value is expressed as: ; Among them, ADC represents the analog-to-digital conversion operation. The quantized image data is separated into channels according to the RGB color space, and the red, green, and blue channels are respectively extracted to obtain three matrices , and . The data of each channel is normalized, and the pixel value is mapped to the range of [0, 1]. The normalization formula is: ; The normalized matrices are respectively , and . After normalization, the three-channel matrix is recombined into an RGB three-channel matrix : ; Use the system clock of the ISP processor to mark the timestamp in microseconds for each frame of image, denoted as , and generate a sequence of image matrices with timestamps . Perform image quality assessment on each frame of image matrix with timestamp to screen out the one with the highest score Frame image. Quality score is calculated based on multiple metrics, including sharpness ( ), contrast ( ), and brightness uniformity ( ). The comprehensive scoring formula is: ; where are weight parameters that satisfy . The specific calculation for each metric is as follows: Sharpness : Calculated by summing the gradient magnitudes, with the formula: ; Contrast : Calculated by the range of brightness distribution, with the formula: ; Brightness uniformity : Calculated by the mean square error of pixel brightness, with the formula: ; where is the average brightness of the image, and are the height and width of the image respectively. By calculating the for each frame, the frame image with the highest score is selected to generate the final high-quality original image sequence . .
[0026] Among them, continuous image sampling is performed on the same scene through an ISP vision sensor. The sampling frequency is set to 30 frames per second, and the sampling duration is set to T seconds, obtaining T×30 frames of image data, including: initializing the circuit of the imaging unit of the ISP vision sensor, setting the pixel array to 1920×1080, the quantum efficiency of each pixel unit to 0.6, and the photosensitive area to 3.45μm×3.45μm to obtain the initial sensor parameters; inputting the initial sensor parameters into the exposure control unit, dynamically adjusting the exposure time within the range of 1 / 30 second to 1 / 60000 second according to the scene brightness, and performing automatic gain control on the ISO within the range of 100 - 3200 to obtain the exposure control matrix; generating a 30Hz sampling timing for the exposure control matrix through the on-chip clock circuit, setting the horizontal and vertical blanking time to 2μs and the pixel readout time to 10ns to obtain the sampling timing signal; controlling the CMOS photodiode array to perform photoelectric conversion based on the sampling timing signal, reading out the pixel voltage through the source follower and row-column gating circuit to obtain the analog image signal; performing CDS correlated double sampling processing on the analog image signal, eliminating the fixed pattern noise and dark current noise through the differential amplifier to obtain the denoised analog signal; inputting the denoised analog signal into a 12-bit column-parallel ADC, performing quantization sampling at a clock frequency of 80MHz using the successive approximation algorithm to obtain the digital image signal; performing Bayer interpolation reconstruction on the digital image signal, calculating the missing color information using the adaptive direction interpolation algorithm, and automatically adjusting the weight coefficient according to the edge gradient to obtain the panchromatic Raw data; continuously collecting and temporarily storing the panchromatic Raw data in the on-chip SRAM cache according to the set sampling duration T, with the cache depth of 12 bits×(1920×1080)×T×30, and transmitting it to the system memory through the DMA channel to obtain T×30 frames of image data.
[0027] In a specific embodiment, the process of executing step 102 may specifically include the following steps: Performing non-illuminated scanning on the N-frame original image sequence, obtaining the dark current noise signal, and performing pixel-level subtraction operation and linear stretching compensation on the dark current noise signal and the N-frame original image sequence to obtain the dark current corrected image sequence; Calculating the brightness mean and variance of each frame of the image in the dark current corrected image sequence respectively, constructing the lens shadow distribution map, and performing per-pixel compensation operation on the lens shadow distribution map and the dark current corrected image sequence, and eliminating the radial and tangential distortion through projective transformation to obtain the lens corrected image sequence; Calculating the adjacent scale image difference of the first frame of the image in the lens corrected image sequence to obtain the scale space feature points; Input the scale-space feature points into the SIFT feature extractor, calculate the main direction and 128-dimensional feature descriptor for each feature point, obtain the set of feature points of the first frame image, and repeat the feature extraction process for the remaining N - 1 frame images in the image sequence after lens correction to obtain N - 1 sets of feature points to be matched; Based on the principle of the nearest neighbor in Euclidean distance, perform feature matching on the set of feature points of the first frame image and the N - 1 sets of feature points to be matched to obtain a sequence of reliable matching point pairs; Establish a least-squares equation according to the sequence of reliable matching point pairs, calculate the affine transformation matrix using singular value decomposition, and transform the N - 1 frame images into the coordinate system of the first frame image to obtain the spatially aligned image sequence.
[0028] Specifically, perform non-illuminated scanning on the original image sequence of frames to obtain the sensor dark current noise signal. The dark current noise signal is a fixed bias generated by the sensor for each pixel in a lightless environment, and its value remains stable throughout the image frame. The result of the non-illuminated scanning is stored as a two-dimensional matrix consistent with the image resolution, denoted as . Subtract this dark current noise signal from each pixel in each frame of the original image sequence, and perform pixel-level subtraction operations to eliminate the influence of the dark current. Assume that the th frame of the original image is , and the subtraction operation formula is: ; where, represents the pixel value after removing the dark current. Since the subtraction operation may cause the pixel value to exceed the normal range (such as becoming negative), perform linear stretching compensation to remap the pixel value to the effective range [0, 255]. The linear stretching formula is: ; where, and respectively represent the minimum and maximum values in, is the pixel value after linear stretching compensation. After this processing, obtain the dark current corrected image sequence . Calculate the brightness mean and variance for each frame of the dark current corrected image, which are used to construct the lens shadow distribution map. The formula for calculating the brightness mean is: ; The formula for calculating the variance is: ; where, and are the height and width of the image. Based on the brightness mean and variance of all frames, a lens shadow distribution map is constructed , representing the uneven brightness distribution caused by the optical characteristics of the lens. Apply this distribution map to each frame of the image, and adjust the brightness non-uniformity through per-pixel compensation operations. The formula is: ; where is the brightness difference to be compensated. Correct the radial distortion and tangential distortion through projective transformation. The formula for projective transformation is: ; where is 's projective transformation matrix, is the corrected pixel coordinate. After applying this transformation matrix, a sequence of lens-corrected images is generated. Calculate the image difference at adjacent scales for the first frame of the lens-corrected image sequence to construct a scale space to detect scale space feature points. Let the first frame of the image be , and the Gaussian blur function be , then the scale space is expressed as: ; The difference image at adjacent scales is: ; where is the scale parameter, is the multiple of the scale interval. By detecting the local extreme points of , scale space feature points are obtained. Input these feature points into the SIFT feature extractor to calculate the principal direction and 128-dimensional feature descriptor for each feature point, and obtain the feature point set of the first frame of the image. For the remaining frames in the lens-corrected image sequence, repeat this process to generate the feature point set for each frame. Based on the nearest neighbor principle of Euclidean distance, match the feature point set of the first frame of the image with the feature point sets of the remaining frames. The calculation formula for the matching distance is: ; where and are two 128-dimensional feature vectors, and are the feature values of the th dimension respectively. By screening the nearest neighbor matching point pairs, a reliable matching point pair sequence is obtained. Based on the matching point pair sequence, establish a least squares equation to calculate the affine transformation matrix . Let the matching point pair be , then the affine transformation matrix satisfies: ; Solve the optimal using Singular Value Decomposition (SVD) and transform to the coordinate system of the first frame image, and finally obtain the spatially aligned image sequence .
[0029] Among them, for the dark current corrected image sequence, calculate the brightness mean and variance of each frame image respectively, construct a lens shadow distribution map, and perform pixel-by-pixel compensation operation on the lens shadow distribution map and the dark current corrected image sequence. Eliminate radial and tangential distortions through projective transformation to obtain the lens corrected image sequence, including: converting the dark current corrected image sequence to the YUV color space, calculating the local brightness mean of the Y channel using a 5×5 pixel sliding window, and determining the brightness abnormal area through a double-threshold segmentation method to obtain the brightness distribution feature map; performing adaptive weighted Gaussian filtering on the brightness distribution feature map, the size of the filter kernel is dynamically adjusted according to the image resolution, and the weight coefficient is jointly determined by the Euclidean distance and brightness difference between pixels to obtain the smoothed brightness distribution map; rearrange the smoothed brightness distribution map according to polar coordinate mapping, calculate the radial brightness attenuation curve, and use polynomial fitting to obtain the brightness attenuation model parameters to obtain the lens shadow distribution map; construct a compensation illumination mapping table based on the lens shadow distribution map, perform block brightness compensation on the dark current corrected image sequence, and the compensation coefficient changes non-linearly with the image area position to obtain the uniformly illuminated image sequence; apply the improved Zhang calibration method to the uniformly illuminated image sequence, calculate the camera internal parameters and lens distortion parameters through multi-view checkerboard images, construct a 15th-order radial distortion model to obtain high-precision distortion correction parameters; establish a pixel remapping matrix according to the high-precision distortion correction parameters, and perform spatial transformation on the uniformly illuminated image sequence in combination with the bicubic interpolation algorithm to obtain the geometrically corrected image sequence; apply an image restoration network based on deep learning to the edge area of the geometrically corrected image sequence, and generate reasonable filling content through multi-scale feature extraction and attention mechanism to obtain the completely restored image sequence; input the completely restored image sequence into an edge enhancement module containing a spatial adaptive threshold, and dynamically adjust the sharpening intensity according to the local gradient distribution to obtain the lens corrected image sequence.
[0030] In a specific embodiment, the process of executing step 103 may specifically include the following steps: Input the spatially aligned image sequence into the first convolutional network of the CFP-ISP feature extraction network. The first convolutional network uses a 3×3 convolutional kernel for feature extraction, the convolutional stride is 1, the padding pixels are 1, and non-linear transformation is performed through the ReLU activation function to obtain the first feature map; Perform a 5×5 convolution operation on the first-layer feature map with a convolution stride of 2 and a padding pixel of 2, reducing the feature map size to half of the first-layer feature map. After processing through the ReLU activation function, perform a max pooling operation with a pooling kernel size of 2×2 to obtain the second-layer feature map; Input the second-layer feature map into the 7×7 convolutional layer of the third-layer convolutional network with a convolution stride of 2 and a padding pixel of 3. After passing through the ReLU activation function and a 2×2 max pooling operation, obtain the third-layer feature map; Extract features from the third-layer feature map using a 9×9 convolutional kernel in the fourth-layer convolutional network with a convolution stride of 2 and a padding pixel of 4. After passing through the ReLU activation function and a 2×2 max pooling operation, obtain the fourth-layer feature map; Align the features based on the four-layer feature maps to obtain an aligned feature sequence. Input the aligned feature sequence into the channel attention module for channel weight calculation and weighted calibration to obtain a multi-scale feature map sequence; Input the multi-scale feature map sequence into the LSTM recurrent neural network. Selectively retain and update the features through the input gate, forget gate, and output gate, and perform a residual feature connection operation to obtain a sequence of temporally correlated feature maps.
[0031] Specifically, input the spatially aligned image sequence into the first-layer convolutional network of the CFP-ISP feature extraction network. The first convolutional operation of this network uses convolutional kernel, with a stride , and a padding pixel to extract local features without changing the input image size. The formula for the convolutional operation is: ; where, is the output of the first-layer convolutional feature map, is the weight of the convolutional kernel, is the bias term, represents the ReLU activation function . After activation by ReLU, obtain the first-layer feature map , whose dimension is the same as the input image. Input the first-layer feature map into the second-layer convolutional network, which uses convolutional kernel, with a stride , and a padding pixel . The formula for the convolutional operation is: ; Since the stride is 2, the spatial size of the feature map will be reduced to half of the first-layer feature map. Apply a max pooling operation to the convolutional result with a pooling kernel size of , the pooling formula is: ; After ReLU activation and pooling, the second-layer feature map is obtained . Taking the second-layer feature map as the input, it is passed to the third-layer convolutional network, which uses convolution kernels, with a stride of , and padding pixels . The formula for the convolution operation is: ; Similarly, the main features are extracted through max-pooling operation, and the formula is: ; After ReLU activation and pooling, the third-layer feature map is obtained . The third-layer feature map is input into the fourth-layer convolutional network, which uses convolution kernels, with a stride of , and padding pixels . The convolution formula is: ; Similarly, features are extracted through the max-pooling operation, and the formula is: ; After processing, the fourth-layer feature map is obtained . Using these four-layer feature maps for feature alignment, an aligned feature sequence is generated. The aligned feature sequence is input into the channel attention module, and the channel attention module calculates the importance of each channel through global average pooling. The formula is: ; Then, a fully connected layer and an activation function are applied to the weight vector to calculate the channel weights. The formula is: ; The channel weights are multiplied with the feature sequence channel by channel to obtain the channel-enhanced feature map: ; Perform LSTM recurrent neural network processing on the channel-enhanced feature map. LSTM dynamically models the features through the input gate, forget gate, and output gate. Its state update formula is as follows: Input gate: ; Forget gate: ; Memory cell: ; Output gate: ; Hidden state update: ; Through the above steps, the LSTM network performs temporal modeling on the input sequence of multi-scale feature maps and generates a sequence of hidden states. A residual connection is performed between the hidden state at each time step and the input features, and the formula is: ; Finally, a sequence of temporally correlated feature maps is obtained .
[0032] In a specific embodiment, the process of performing the steps of inputting the multi-scale feature map sequence into the LSTM recurrent neural network, selectively retaining and updating the features through the input gate, forget gate, and output gate, and performing the residual feature connection operation to obtain the sequence of temporally correlated feature maps may specifically include the following steps: Input the multi-scale feature map sequence into the input gate of the LSTM recurrent neural network, calculate the input weight matrix through the sigmoid function, selectively filter the input features at the current time step, and obtain the input gated features; Perform a tanh activation operation on the input gated features and perform an element-wise multiplication operation with the input weight matrix to obtain the candidate memory units; Input the candidate memory units into the forget gate of the LSTM recurrent neural network, calculate the forget gate weight matrix through the sigmoid function, selectively forget the memory state at the previous time step, and obtain the forget gated features; Perform a weighted sum operation on the forget gated features and the candidate memory units to update the memory unit state at the current time step and obtain the updated memory units; Input the updated memory units into the output gate of the LSTM recurrent neural network, calculate the output weight matrix through the sigmoid function, selectively output the output features at the current time step, and obtain the output gated features; Perform a tanh activation operation on the updated memory units and perform an element-wise multiplication operation with the output gated features to obtain the output state at the current time step; Based on the multi-scale feature map sequence, establish a residual connection path, perform an element-wise addition operation on the input features and the output state at the current time step, and obtain the sequence of temporally correlated feature maps.
[0033] Specifically, define the input multi-scale feature map sequence as , where represents the time step, represents the spatial coordinates, Represents the channel index. The LSTM network dynamically models features through input gates, forget gates, and output gates, and generates a sequence of feature maps with temporal correlation through residual connections. For the input features at time step of , it is processed by the input gate to selectively filter the valid part of the current input features. The calculation formula of the input gate is: ; where is the output of the input gate, represents the sigmoid function, which is used to compress the value to the range of [0, 1]; is the weight matrix of the input gate, is the bias vector, is the hidden state at the previous moment, represents the concatenation of the hidden state and the current input features as the input. The output of the input gate selectively filters the current features, indicating the retention ratio of each feature. The output of the input gate is subjected to a tanh activation operation to generate a candidate memory unit. The calculation formula of the candidate memory unit is: ; where is the candidate memory unit, is the weight matrix of the candidate unit, is the bias vector. The candidate memory unit captures the potential memory information at the current moment. On this basis, the output of the input gate is multiplied element-wise with the candidate memory unit to obtain the input-gated feature: ; where represents the element-wise multiplication operation, is the input-gated feature. The candidate memory unit is input into the forget gate to determine the proportion of forgetting the memory state at the previous moment. The calculation formula of the forget gate is: ; where is the output of the forget gate, is the weight matrix of the forget gate, is the bias vector. The forget gate controls the degree of forgetting of the memory state at the previous moment. The forget-gated feature is calculated as: ; The memory unit state at the current moment is updated through the weighted sum of the forget-gated feature and the input-gated feature, and the formula is: ; where is the updated memory cell state, which integrates the memory of the previous moment and the input information of the current moment. The updated memory cell state is input into the output gate to selectively generate the output feature of the current moment. The calculation formula of the output gate is: ; where is the output of the output gate, is the weight matrix of the output gate, is the bias vector. The output gate controls which memory cell information will be used as the output feature of the current moment. The updated memory cell state is subjected to the tanh activation operation and multiplied element-wise with the output gate control feature to obtain the output state of the current moment: ; where is the hidden state of the current moment, representing the output feature after the LSTM network models the input feature. To enhance the feature transfer ability and stability of the network, a residual connection path is established based on the hidden state of the current moment, and the input feature is directly added to the hidden state to obtain the output of the current moment of the sequence of temporal correlation feature maps: ; Through this residual connection, the network effectively retains the original information of the input feature and at the same time uses the temporal correlation features in the hidden state to achieve enhanced expression. This process is carried out for each time step of the sequence of multi-scale feature maps in turn to obtain the sequence of temporal correlation feature maps .
[0034] In a specific embodiment, the process of executing step 104 may specifically include the following steps: Perform global average pooling operation on each temporal correlation feature map of the sequence of temporal correlation feature maps to obtain an initial channel feature vector, and input the initial channel feature vector into the first fully connected layer for dimensionality reduction transformation to obtain an intermediate feature vector; Input the intermediate feature vector into the second fully connected layer for dimensionality increase transformation to obtain a channel attention weight vector, and perform element-wise multiplication of the channel attention weight vector with the sequence of temporal correlation feature maps to obtain a channel enhanced feature sequence; Perform a 3×3 convolution operation on the channel enhanced feature sequence, and extract spatial saliency information through the max pooling and average pooling branches respectively to obtain a first spatial feature map and a second spatial feature map; Concatenate the first spatial feature map and the second spatial feature map in the channel dimension, fuse the spatial information through a 1×1 convolutional layer, and perform normalization processing using the sigmoid function to obtain a spatial attention weight map; Perform an element-wise multiplication operation on the channel-enhanced feature sequence and the spatial attention weight map, and perform a residual connection with the temporal correlation feature map sequence to obtain a spatio-temporal attention feature sequence; Calculate the attention coefficient at each temporal position based on the spatio-temporal attention feature sequence, perform a weighted average fusion operation, fuse the N-frame feature sequence into a single-frame representation, and obtain a two-dimensional fusion feature map; Input the two-dimensional fusion feature map into a three-layer transposed convolutional network for upsampling to obtain an upsampled feature map, and perform structural layer edge enhancement and texture layer denoising processing on the upsampled feature map to obtain the target enhanced image.
[0035] Specifically, define the input temporal correlation feature map sequence as , where 𝑡 represents the time frame, is the spatial coordinate, is the channel index. For each temporal correlation feature map , perform global average pooling to extract the global features of each channel. The formula is: ; where, is the average feature value of channel , and are the height and width of the feature map respectively, to obtain the initial channel feature vector . Input the initial channel feature vector into the first fully connected layer for dimensionality reduction transformation. The formula for the intermediate feature vector after dimensionality reduction is: ; where, and are the weights and biases of the first fully connected layer respectively, and ReLU is the non-linear activation function. Input the intermediate feature vector into the second fully connected layer for dimensionality increase transformation to generate the channel attention weight vector : ; where, and are the weights and biases of the second fully connected layer respectively, is the sigmoid function, which is used to limit the output value to [0,1]. The channel attention weight vector is used to adjust the importance of each channel. Multiply it with the original temporal features channel by channel to obtain the channel-enhanced feature sequence : ; Perform convolution operation on the channel-enhanced feature sequence to extract local spatial features. The formula is: ; Where is the convolution result, and are the weights and biases of the convolution kernel. Extract spatial saliency information from the convolution result through the max pooling and average pooling branches respectively. The formula is: Max pooling: ; Average pooling: ; Where and represent the results of max pooling and average pooling respectively, obtaining the first spatial feature map and the second spatial feature map. Concatenate the two spatial feature maps in the channel dimension to form a fused feature map . Then, further fuse the information through the convolutional layer. The formula is: ; Normalize the fused feature map using the sigmoid function to generate a spatial attention weight map : ; Multiply the channel-enhanced feature sequence element-wise with the spatial attention weight map, and add it to the original temporal correlation feature map through residual connection to obtain the spatio-temporal attention feature sequence: ; Calculate the attention coefficient at each time position based on the spatio-temporal attention feature sequence, and fuse the frame feature sequence into a single-frame representation through weighted average. The formula is: ; Where is the time attention coefficient, generated by the network or a specific function. Input the two-dimensional fused feature map into a three-layer transposed convolution network for progressive upsampling. Each layer uses a transposed convolution kernel with a stride of 2. After upsampling, a higher-resolution feature map is obtained. Perform structural layer edge enhancement and texture layer denoising on the upsampled feature map, strengthen the edge details by extracting gradient information, and at the same time apply a denoising algorithm to eliminate random noise, finally generating the target enhanced image .
[0036] In a specific embodiment, the process of inputting the two-dimensional fusion feature map into a three-layer transposed convolutional network for upsampling to obtain the upsampled feature map, and performing structural layer edge enhancement and texture layer denoising processing on the upsampled feature map to obtain the target enhanced image may specifically include the following steps: Input the two-dimensional fusion feature map into the first transposed convolutional network layer of the three-layer transposed convolutional network, and perform 2-fold upsampling on the two-dimensional fusion feature map through a 4×4 transposed convolutional kernel with a stride of 2 to obtain the first upsampling result, and perform the ReLU activation function and batch normalization processing on the first upsampling result to obtain the first layer feature map at a 2-fold scale; Input the first layer feature map into the second transposed convolutional network layer of the three-layer transposed convolutional network, and perform 2-fold upsampling through a 4×4 transposed convolutional kernel with a stride of 2 to obtain the second upsampling result, and perform the ReLU activation function and batch normalization processing on the second upsampling result to obtain the second layer feature map at a 4-fold scale; Input the second layer feature map into the third transposed convolutional network layer of the three-layer transposed convolutional network, and perform 2-fold upsampling through a 4×4 transposed convolutional kernel with a stride of 2 to obtain the third upsampling result, and perform the ReLU activation function and batch normalization processing on the third upsampling result to obtain the upsampled feature map at an 8-fold scale; Perform multi-scale decomposition on the upsampled feature map at an 8-fold scale to obtain the structural layer feature map and the texture layer feature map. The structural layer feature map contains the target contour information, and the texture layer feature map contains the detail information; For the structural layer feature map, perform gradient calculation using 3×3 operators in the horizontal and vertical directions to obtain the enhanced structural edge features, and perform denoising processing on the texture layer feature map to obtain the denoised texture features; Perform weighted superposition on the enhanced structural edge features and the denoised texture features to obtain the enhanced fusion features, and perform HDR tone mapping and color distribution correction on the enhanced fusion features to obtain the target enhanced image.
[0037] Specifically, represent the input two-dimensional fusion feature map as , where are the spatial coordinates, is the channel index. This feature map is input into the first transposed convolutional network layer for upsampling operation, and a transposed convolutional kernel with a stride of 2 is used to perform 2-fold upsampling on the input feature map. The calculation formula for the transposed convolution is: ; where, is the output feature map of the first transposed convolution, is the weight of the first layer of transposed convolution kernels, and the stride of 2 ensures that the spatial size of the output is twice that of the input. To enhance the non-linear expression ability, the transposed convolution result passes through the ReLU activation function: ; Next, batch normalization is performed on the activated feature map to reduce the data distribution difference between channels. The formula is: ; Among them, and are learnable normalization parameters, and are the mean and variance of the current channel respectively, is a small constant to prevent division by zero. After processing, a feature map with a scale twice that of the first layer is obtained. The first-layer feature map is input into the second-layer transposed convolution network layer, and the above process is repeated to perform 2x upsampling again to generate a feature map with a scale of 4 times. The formula for the transposed convolution operation is similar: ; After ReLU activation and batch normalization, the second-layer feature map is obtained. The second-layer feature map is input into the third-layer transposed convolution network layer, and the third 2x upsampling is performed through the transposed convolution kernel with a stride of 2 to generate a feature map with a scale of 8 times. The formula is as follows: ; Similarly, after ReLU activation and batch normalization, the final feature map with an 8x upsampling scale is obtained . The 8x upsampled feature map is decomposed into multi-scales, and it is decomposed into a structural layer feature map and a texture layer feature map . The structural layer feature map mainly represents the contour and edge information of the target, and the texture layer feature map contains details and high-frequency information. For the structural layer feature map, the gradient operators in the horizontal and vertical directions are used to calculate the edge features respectively. The formula is: Horizontal direction: ; Vertical direction: ; Among them, and are the gradient operators in the horizontal and vertical directions respectively, such as the Sobel operator. The gradient results in these two directions are combined into an edge intensity feature: ; For the texture layer feature map, a denoising algorithm is used to eliminate random noise while retaining high-frequency details. The denoising process uses bilateral filtering. The formula is: ; Among them, is the spatial distance weight, is the pixel intensity difference weight. The enhanced structural edge features and the denoised texture features are superimposed according to the weight ratio to generate an enhanced fusion feature map: ; Among them, is the fusion weight, which controls the relative importance of the structure and the texture. Perform HDR tone mapping and color distribution correction on the enhanced fusion features. The HDR tone mapping compresses the dynamic range through a non-linear function, and the formula is: ; The color distribution correction uses the mapping matrix in the target color space to complete the generation of the corrected target enhanced image: ; Through the above steps, the conversion from the two-dimensional fusion features to the target enhanced image is realized, and the output image has better edge sharpness and texture details, meeting the requirements of high-quality image enhancement.
[0038] In this embodiment, an adaptive spectral compensation process is further included: performing wavelet transform decomposition on the target enhanced image, obtaining 7 high-frequency subbands and 1 low-frequency subband through three-layer wavelet decomposition, extracting frequency features and energy distribution features for each subband to obtain a multi-scale spectral feature matrix; inputting the multi-scale spectral feature matrix into a spectral analysis network, calculating attention weights in the frequency domain and energy domain respectively through a dual-branch attention module, performing importance weighting on features in different frequency bands to obtain a spectral weight mapping graph; constructing an adaptive spectral compensation filter bank based on the spectral weight mapping graph, where the filter bank includes a high-frequency enhancement filter, an intermediate-frequency modulation filter, and a low-frequency preservation filter, performing adaptive enhancement on different frequency bands to obtain compensated spectral components; inputting the compensated spectral components into a non-linear enhancement module, using a multi-scale decomposition algorithm based on the Retinex theory to separate the image into an illumination component and a reflection component to obtain illumination-invariant features; performing local contrast adaptive enhancement on the illumination-invariant features, dynamically adjusting the enhancement intensity according to the illumination conditions and texture complexity of the local area to obtain an enhanced reflection component; adaptively fusing the enhanced reflection component with the original illumination component, where the fusion weight is determined by the local statistical characteristics of the image to obtain an enhanced image after spectral compensation; performing color restoration and correction on the enhanced image after spectral compensation, using a color mapping algorithm based on color consistency constraints to maintain the naturalness and authenticity of the color to obtain a color-corrected image; constructing an HDR tone mapping model based on the color-corrected image, and obtaining the final high-dynamic-range enhanced image through adaptive brightness adjustment and local detail enhancement.
[0039] The image enhancement method based on an ISP image signal processing visual sensor in the embodiments of the present invention has been described above. Next, the image enhancement device based on an ISP image signal processing visual sensor in the embodiments of the present invention will be described. Please refer to Figure 2 , an embodiment of the image enhancement device based on an ISP image signal processing visual sensor in the embodiments of the present invention includes: An acquisition module 201, configured to acquire multiple frames of images of the same scene, perform analog-to-digital conversion on the acquired multiple frames of images and generate an RGB three-channel digital matrix, record the timestamp of each frame of image, and obtain an N-frame original image sequence, where N is an integer greater than or equal to 3; A correction module 202, configured to perform dark current correction and lens shading correction on the N-frame original image sequence, calculate the affine transformation matrix from the N-1 frame to the first frame image based on SIFT feature point extraction and matching, and obtain an image sequence after spatial alignment; An extraction module 203, configured to input the image sequence after spatial alignment into a CFP-ISP feature extraction network for spatial feature extraction to obtain a multi-scale feature map sequence, and input the multi-scale feature map sequence into an LSTM recurrent neural network for residual feature connection operation to obtain a temporal correlation feature map sequence; A generation module 204 is configured to perform attention calculations on the sequence of temporal correlation feature maps in both the channel dimension and the spatial dimension to generate a two-dimensional fusion feature map, and input the two-dimensional fusion feature map into a three-layer transposed convolution network for upsampling. Meanwhile, edge enhancement of the structure layer and denoising of the texture layer are performed to obtain the target enhanced image.
[0040] Through the collaborative cooperation of the above-mentioned various components, by designing a multi-frame image acquisition and processing mechanism, combined with analog-to-digital conversion and timestamp recording, the integrity and temporal consistency of image information are effectively improved, providing high-quality input data for subsequent enhancement processing. A dual correction mechanism of dark current correction and lens shading correction, combined with SIFT feature matching and affine transformation, significantly reduces various distortions and noise interferences during the image acquisition process. A multi-scale feature extraction network based on CFP-ISP is designed, and through convolutional kernels of different sizes and ReLU activation functions, multi-level and multi-scale extraction of image features is achieved. An LSTM recurrent neural network is innovatively introduced for temporal feature processing, and through a three-gate control mechanism and residual connections, the temporal correlation between image sequences is enhanced. Combining the dual attention mechanism in the channel dimension and the spatial dimension realizes the adaptive weighting and fusion of features, enhancing the expression ability of important features. A three-layer transposed convolution network is used for feature upsampling, combined with a hierarchical processing strategy of edge enhancement of the structure layer and denoising of the texture layer, ensuring the clarity and detail fidelity of the enhanced image. Through HDR tone mapping and automatic white balance technology, the dynamic range and color restoration effect of the image are optimized, improving the visual quality of the enhancement result.
[0041] Above Figure 2 From the perspective of modular functional entities, the image enhancement device based on an ISP image signal processing visual sensor in the embodiments of the present invention is described in detail. Next, the image enhancement device based on an ISP image signal processing visual sensor in the embodiments of the present invention is described in detail from the perspective of hardware processing.
[0042] Figure 3FIG. 0 is a schematic structural diagram of an image enhancement device based on an ISP image signal processing vision sensor provided by an embodiment of the present invention. The image enhancement device 300 based on the ISP image signal processing vision sensor may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 333 or data 332 (for example, one or more mass storage device ends). Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the image enhancement device 300 based on the ISP image signal processing vision sensor. Further, the processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the image enhancement device 300 based on the ISP image signal processing vision sensor to implement the steps of the above-mentioned image enhancement method based on the ISP image signal processing vision sensor.
[0043] The image enhancement device 300 based on the ISP image signal processing vision sensor may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that Figure 3 The shown structural diagram of the image enhancement device based on the ISP image signal processing vision sensor does not limit the image enhancement device based on the ISP image signal processing vision sensor provided by the present invention, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0044] The present invention also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium, or may also be a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are run on a computer, the computer is caused to execute the steps of the image enhancement method based on the ISP image signal processing vision sensor.
[0045] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0046] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0047] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An image enhancement method based on an ISP image signal processing vision sensor, characterized in that, The method includes: Performing multi-frame image acquisition on the same scene, performing analog-to-digital conversion on the acquired multi-frame images and generating an RGB three-channel digital matrix, recording the timestamp of each frame of image, and obtaining an N-frame original image sequence, where N is an integer greater than or equal to 3; Performing dark current correction and lens shading correction on the N-frame original image sequence, calculating the affine transformation matrix from the N-1th frame to the first frame of image based on SIFT feature point extraction and matching, and obtaining the spatially aligned image sequence; Inputting the spatially aligned image sequence into the CFP-ISP feature extraction network for spatial feature extraction to obtain a multi-scale feature map sequence, and inputting the multi-scale feature map sequence into the LSTM recurrent neural network for residual feature connection operation to obtain a temporal correlation feature map sequence; Performing attention calculation on the temporal correlation feature map sequence in the channel dimension and the spatial dimension to generate a two-dimensional fusion feature map, and inputting the two-dimensional fusion feature map into a three-layer deconvolution network for upsampling. At the same time, performing structural layer edge enhancement and texture layer denoising processing to obtain the target enhanced image.
2. The image enhancement method based on an ISP image signal processing vision sensor according to claim 1, characterized in that The performing multi-frame image acquisition on the same scene, performing analog-to-digital conversion on the acquired multi-frame images and generating an RGB three-channel digital matrix, recording the timestamp of each frame of image, and obtaining an N-frame original image sequence, where N is an integer greater than or equal to 3, includes: Performing continuous image sampling on the same scene through an ISP vision sensor, setting the sampling frequency to 30 frames per second, and setting the sampling duration to T seconds to obtain T×30 frame image data; Performing ambient light intensity detection on the T×30 frame image data through a real-time light sensor to obtain the ambient light intensity detection result, and calculating the LED fill light coefficient within the range of 0-1000 lux based on the ambient light intensity detection result to obtain the light compensation data; Performing analog-to-digital conversion on the light compensation data to obtain the quantized digital image data, separating the channels of the quantized digital image data according to the RGB color space, and performing normalization operations in the [0,255] interval on the R channel, G channel, and B channel respectively to obtain three normalized digital matrices; Performing channel merging operation on the three normalized digital matrices, arranging and combining the digital matrices of the R, G, and B channels according to the color space to obtain an RGB three-channel digital matrix; Based on the system clock of the ISP processor, marking microsecond-level timestamp information on the RGB three-channel digital matrix to obtain an image matrix sequence with serial numbers, and performing image quality evaluation based on the image matrix sequence with serial numbers to screen out the N-frame original image sequence with the highest score.
3. The image enhancement method based on an ISP image signal processing vision sensor according to claim 2, characterized in that, The performing dark current correction and lens shading correction on the N-frame original image sequence, calculating the affine transformation matrix from the N-1th frame to the first frame of image based on SIFT feature point extraction and matching, and obtaining the spatially aligned image sequence, includes: Performing lightless scanning on the N-frame original image sequence to obtain the dark current noise signal, and performing pixel-level subtraction operation and linear stretching compensation on the dark current noise signal and the N-frame original image sequence to obtain the dark current correction image sequence; Calculate the brightness mean and variance of each frame image in the dark current corrected image sequence respectively, construct a lens shadow distribution map, and perform pixel-by-pixel compensation operation on the lens shadow distribution map and the dark current corrected image sequence. Eliminate radial and tangential distortion through projective transformation to obtain the lens corrected image sequence; Calculate the adjacent scale image difference based on the first frame image in the lens corrected image sequence to obtain scale space feature points; Input the scale space feature points into the SIFT feature extractor, calculate the main direction and 128-dimensional feature descriptor of each feature point to obtain the feature point set of the first frame image, and repeat the feature extraction process for the remaining N - 1 frame images in the lens corrected image sequence to obtain N - 1 sets of feature points to be matched; Perform feature matching on the feature point set of the first frame image and the N - 1 sets of feature points to be matched based on the principle of nearest neighbor in Euclidean distance to obtain a reliable matching point pair sequence; Establish a least squares equation according to the reliable matching point pair sequence, calculate the affine transformation matrix using singular value decomposition, and transform the N - 1 frame images into the coordinate system of the first frame image to obtain the spatially aligned image sequence.
4. The image enhancement method of the ISP image signal processing vision sensor according to claim 3, characterized in that Input the spatially aligned image sequence into the CFP-ISP feature extraction network for spatial feature extraction to obtain a multi-scale feature map sequence, and input the multi-scale feature map sequence into the LSTM recurrent neural network for residual feature connection operation to obtain a temporal correlation feature map sequence, including: Input the spatially aligned image sequence into the first convolutional network of the CFP-ISP feature extraction network. The first convolutional network uses a 3×3 convolutional kernel for feature extraction, the convolutional stride is 1, the padding pixels are 1, and non-linear transformation is performed through the ReLU activation function to obtain the first layer feature map; Perform a 5×5 convolutional operation of the second convolutional network on the first layer feature map, the convolutional stride is 2, the padding pixels are 2, reduce the feature map size to half of the first layer feature map, perform maximum pooling operation after being processed by the ReLU activation function, and the pooling kernel size is 2×2 to obtain the second layer feature map; Input the second layer feature map into the 7×7 convolutional layer of the third convolutional network, the convolutional stride is 2, the padding pixels are 3, and after passing through the ReLU activation function and 2×2 maximum pooling operation, obtain the third layer feature map; Perform feature extraction on the third layer feature map through the 9×9 convolutional kernel in the fourth convolutional network, the convolutional stride is 2, the padding pixels are 4, and after passing through the ReLU activation function and 2×2 maximum pooling operation, obtain the fourth layer feature map; Perform feature alignment according to the four-layer feature map to obtain an aligned feature sequence, input the aligned feature sequence into the channel attention module for channel weight calculation and weighted calibration to obtain a multi-scale feature map sequence; Input the multi-scale feature map sequence into the LSTM recurrent neural network, selectively retain and update the features through the input gate, forget gate and output gate, and perform residual feature connection operation to obtain a temporal correlation feature map sequence.
5. The image enhancement method based on an ISP image signal processing vision sensor according to claim 4, characterized in that, Inputting the multi-scale feature map sequence into an LSTM recurrent neural network, selectively retaining and updating features through an input gate, a forget gate, and an output gate, and performing a residual feature connection operation to obtain a sequence of time-series related feature maps, including: Inputting the multi-scale feature map sequence into the input gate of the LSTM recurrent neural network, calculating an input weight matrix through a sigmoid function, and selectively filtering the input features at the current moment to obtain input gated features; Performing a tanh activation operation on the input gated features and performing an element-wise multiplication operation with the input weight matrix to obtain candidate memory units; Inputting the candidate memory units into the forget gate of the LSTM recurrent neural network, calculating a forget gate weight matrix through a sigmoid function, and selectively forgetting the memory state at the previous moment to obtain forget gated features; Performing a weighted sum operation on the forget gated features and the candidate memory units to update the memory unit state at the current moment to obtain updated memory units; Inputting the updated memory units into the output gate of the LSTM recurrent neural network, calculating an output weight matrix through a sigmoid function, and selectively outputting the output features at the current moment to obtain output gated features; Performing a tanh activation operation on the updated memory units and performing an element-wise multiplication operation with the output gated features to obtain the output state at the current moment; Based on the multi-scale feature map sequence, establishing a residual connection path, and performing an element-wise addition operation on the input features and the output state at the current moment to obtain a sequence of time-series related feature maps.
6. The image enhancement method of the ISP image signal processing vision sensor according to claim 5, wherein Performing attention calculations on the sequence of time-series related feature maps in the channel dimension and the spatial dimension to generate a two-dimensional fusion feature map, and inputting the two-dimensional fusion feature map into a three-layer transposed convolution network for upsampling. At the same time, performing structural layer edge enhancement and texture layer denoising processing to obtain a target enhanced image, including: Performing global average pooling operations on each time-series related feature map in the sequence of time-series related feature maps to obtain initial channel feature vectors, and inputting the initial channel feature vectors into a first fully connected layer for dimensionality reduction transformation to obtain intermediate feature vectors; Inputting the intermediate feature vectors into a second fully connected layer for dimensionality increase transformation to obtain channel attention weight vectors, and performing an element-wise multiplication operation on the channel attention weight vectors and the sequence of time-series related feature maps to obtain a sequence of channel enhanced features; Performing 3×3 convolution operations on the sequence of channel enhanced features, and respectively extracting spatial saliency information through max pooling and average pooling branches to obtain a first spatial feature map and a second spatial feature map; Concatenating the first spatial feature map and the second spatial feature map in the channel dimension, fusing spatial information through a 1×1 convolutional layer, and performing normalization processing using a sigmoid function to obtain a spatial attention weight map; Performing an element-wise multiplication operation on the sequence of channel enhanced features and the spatial attention weight map, and performing a residual connection with the sequence of time-series related feature maps to obtain a sequence of spatio-temporal attention features; Calculate the attention coefficients for each temporal position based on the spatio-temporal attention feature sequence, perform a weighted average fusion operation to fuse the N-frame feature sequence into a single-frame representation, and obtain a two-dimensional fusion feature map; Input the two-dimensional fusion feature map into a three-layer transposed convolutional network for upsampling to obtain an upsampled feature map, and perform structural layer edge enhancement and texture layer denoising on the upsampled feature map to obtain a target enhanced image.
7. The image enhancement method of the ISP image signal processing vision sensor according to claim 6, characterized in that, The step of inputting the two-dimensional fusion feature map into a three-layer transposed convolutional network for upsampling to obtain an upsampled feature map, and performing structural layer edge enhancement and texture layer denoising on the upsampled feature map to obtain a target enhanced image includes: Input the two-dimensional fusion feature map into the first transposed convolutional network layer of the three-layer transposed convolutional network, perform 2-fold upsampling on the two-dimensional fusion feature map using a 4×4 transposed convolutional kernel with a stride of 2 to obtain a first upsampling result, and perform a ReLU activation function and batch normalization on the first upsampling result to obtain a first-layer feature map at a 2-fold scale; Input the first-layer feature map into the second transposed convolutional network layer of the three-layer transposed convolutional network, perform 2-fold upsampling using a 4×4 transposed convolutional kernel with a stride of 2 to obtain a second upsampling result, and perform a ReLU activation function and batch normalization on the second upsampling result to obtain a second-layer feature map at a 4-fold scale; Input the second-layer feature map into the third transposed convolutional network layer of the three-layer transposed convolutional network, perform 2-fold upsampling using a 4×4 transposed convolutional kernel with a stride of 2 to obtain a third upsampling result, and perform a ReLU activation function and batch normalization on the third upsampling result to obtain an upsampled feature map at an 8-fold scale; Perform multi-scale decomposition on the upsampled feature map at an 8-fold scale to obtain a structural layer feature map and a texture layer feature map, where the structural layer feature map contains target contour information and the texture layer feature map contains detail information; For the structural layer feature map, perform gradient calculation using 3×3 operators in the horizontal and vertical directions to obtain enhanced structural edge features, and perform denoising on the texture layer feature map to obtain denoised texture features; Perform weighted superposition on the enhanced structural edge features and the denoised texture features to obtain enhanced fusion features, and perform HDR tone mapping and color distribution correction on the enhanced fusion features to obtain a target enhanced image.
8. An image enhancement device based on an ISP image signal processing vision sensor, characterized in that, A device for performing the image enhancement method of an ISP image signal processing vision sensor according to any one of claims 1-7, the device includes: An acquisition module for collecting multiple frames of images of the same scene, performing analog-to-digital conversion on the collected multiple frames of images and generating an RGB three-channel digital matrix, recording the timestamp of each frame of image, and obtaining an N-frame original image sequence, where N is an integer greater than or equal to 3; A correction module for performing dark current correction and lens shading correction on the N-frame original image sequence, calculating an affine transformation matrix from the N-1 frame to the first frame image based on SIFT feature point extraction and matching, and obtaining a spatially aligned image sequence; An extraction module for inputting the spatially aligned image sequence into a CFP-ISP feature extraction network for spatial feature extraction to obtain a multi-scale feature map sequence, and inputting the multi-scale feature map sequence into an LSTM recurrent neural network for residual feature connection operations to obtain a temporal correlation feature map sequence; A generation module for performing attention calculations on the temporal correlation feature map sequence in the channel dimension and the spatial dimension to generate a two-dimensional fusion feature map, and inputting the two-dimensional fusion feature map into a three-layer transposed convolution network for upsampling, and at the same time, performing structural layer edge enhancement and texture layer denoising processing to obtain a target enhanced image.
9. An image enhancement device based on an ISP image signal processing vision sensor, characterized in that, The image enhancement device based on an ISP image signal processing vision sensor includes: a memory and at least one processor, and instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the image enhancement device based on an ISP image signal processing vision sensor executes the image enhancement method based on an ISP image signal processing vision sensor according to any one of claims 1-7.
10. A computer-readable storage medium, on which instructions are stored, characterized in that, When the instructions are executed by the processor, the image enhancement method based on an ISP image signal processing vision sensor according to any one of claims 1-7 is implemented.
Citation Information
Cited By
Defect monitoring method based on visual large model
CN120726029A
Hydraulic engineering concrete crack detection method and system
CN121074034A
Submarine cable image splicing method based on double-domain decoupling enhancement and multistage fault-tolerant matching
CN121998820A
Paper electrocardiogram digitization method and electronic equipment
CN122066610A
Low-illumination image enhancement method, system, equipment and medium
CN122265057A