Gradient descent scene consistency adjusting method for fixed scene
Through a dual-branch multi-input neural network model and a color difference-texture similarity joint loss function, the color mapping relationship between the reference image and the image to be adjusted is directly fitted, which solves the accuracy and adaptability problems of image color consistency adjustment in fixed scenes, and realizes color consistency adjustment in remote sensing images and fixed camera scenes.
Patent Information
- Application Number
- CN202511308174.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-15
AI Technical Summary
The existing technology for adjusting image color consistency in complex fixed scenes has problems such as insufficient color matching accuracy and limited scene adaptability.
A dual-branch multi-input neural network model is adopted to perform feature fusion through the stepped downsampling of the encoder and the upsampling of the decoder. Combined with the color difference-texture similarity joint loss function, the gradient descent algorithm is used to optimize the neural network parameters, directly fit the color mapping relationship between the reference image and the image to be adjusted, and output the final adjusted image.
It achieves image color consistency adjustment in multiple types of scenes, avoiding the problem of poor adjustment effect caused by the pre-trained model not having seen images of specific scenes. It has wide applicability and can be applied to color difference adjustment of remote sensing images and fixed cameras across scene types.
Smart Images

Figure CN120807333A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image color processing and computer vision, in particular to a gradient descent scene consistency adjustment method for fixed scenes. BACKGROUND
[0002] In the field of image acquisition and processing, color consistency adjustment of images in fixed scenes is a key technical requirement. Whether it is remote sensing monitoring or security monitoring, due to differences in shooting equipment (such as different types of sensors, cameras), shooting conditions (such as light intensity, weather conditions, shooting time), images of the same scene will have inconsistent color distribution, blurred or distorted details, etc., which seriously affects the accuracy and reliability of subsequent image analysis.
[0003] At present, the existing technologies for image scene color consistency adjustment can be mainly divided into two categories: one is the traditional correction method, including algorithms based on gray world assumption, white balance adjustment, histogram matching, etc., which has the limitation of relying on manually set empirical parameters, lacking adaptive learning ability, and low color matching accuracy; the other is the color adjustment method based on neural network, which is characterized by training based on a large amount of data, and the model weight is fixed after training, which has the limitation of good adaptability to specific scenes and poor adaptability to general scenes, and if the application scene is greatly different from the training data scene, the adjustment effect will be poor (such as remote sensing color adjustment model cannot be used for day and night color adjustment). In summary, the existing technologies still face the problems of insufficient color matching accuracy and limited scene adaptability in color consistency adjustment in complex fixed scenes. SUMMARY
[0004] Therefore, the present application provides a gradient descent scene consistency adjustment method for fixed scenes to solve the problems of insufficient color matching accuracy and limited scene adaptability in color consistency adjustment in complex fixed scenes in the prior art.
[0005] The present application provides a gradient descent scene consistency adjustment method for fixed scenes, comprising:
[0006] acquiring reference images and images to be adjusted by collecting image data;
[0007] concatenating the reference images and the images to be adjusted in the channel dimension to form a double-input joint feature, inputting the double-input joint feature into a double-branch multi-input neural network model, performing feature fusion through stepwise down-sampling of the encoder and up-sampling of the decoder, and outputting a preliminary adjustment image;
[0008] The color difference-texture similarity joint loss function is used to calculate a color difference loss between the preliminary adjusted image and the reference image and a structure similarity loss between the preliminary adjusted image and the image to be adjusted, and the color difference loss and the structure similarity loss are weighted and fused to obtain a total loss value.
[0009] Based on the total loss value, a gradient descent algorithm is used to iteratively optimize the neural network parameters, and when the optimization termination condition is met, a final adjusted image is output.
[0010] The reference image is obtained by,
[0011] In a fixed scene, a single high-quality image collected is used as the reference image, and at the same time, in the same scene, images obtained by changing the collection parameters are used as the images to be adjusted.
[0012] The double-branch multi-input neural network model comprises an encoder and a decoder.
[0013] The encoder performs stepwise down-sampling coding on the reference image and the image to be adjusted to extract local features and global features; in the down-sampling process, the global features and the local features of the reference image and the image to be adjusted are abstractly coded, and intermediate features carrying different spatial scales are output through different down-sampling levels.
[0014] The decoder learns the color mapping relationship from the image to be adjusted to the reference image step by step based on the feature difference between the reference image and the image to be adjusted extracted by coding, through up-sampling and cross-stage feature fusion, and restores the features abstracted by the encoder stage to image data after color adjustment using the learned color mapping relationship.
[0015] The color difference-texture similarity joint loss function promotes the network to adjust the color mapping relationship through the gradient feedback of the color difference loss, so that the overall color distribution of the image to be adjusted approaches the reference image, and at the same time, the color difference-texture similarity joint loss function forces the network to retain the structural information of the edge contour and texture details of the original image during color conversion through the constraint of the texture similarity loss.
[0016] Let the reference image be , the image to be adjusted be , and the adjusted image output after processing by the neural network model be ;
[0017] The color difference loss function comprises,
[0018] ;
[0019] Wherein, H represents the height of the image, W represents the width of the image, 3 represents the number of RGB channels, i represents the number of image height direction, j represents the number of image width direction, and c represents the number of RGB three color channels;
[0020] The calculation of the texture similarity loss function includes,
[0021] The to-be-adjusted image and the preliminary adjustment image output after the neural network model processing are respectively converted into gray images And The Sobel operator is used to extract the texture of the gray image.
[0022] The definition of the Sobel operator x direction is:
[0023] ;
[0024] The definition of the Sobel operator y direction is:
[0025] ;
[0026] The texture map of the gray image of the to-be-adjusted image includes,
[0027] ;
[0028] The texture map of the gray image of the preliminary adjustment image includes,
[0029] .
[0030] The texture map of the gray image of the to-be-adjusted image And the texture map of the gray image of the preliminary adjustment image Global Softmax normalization is performed, and the texture value range is compressed to [0, 1];
[0031] ;
[0032] ;
[0033] Wherein, The normalized texture probability map of the gray image of the to-be-adjusted image, The normalized texture probability map of the gray image of the preliminary adjustment image, exp represents the exponential function with the natural constant e as the base, i represents the number of image height direction, and j represents the number of image width direction.
[0034] The difference of the normalized texture map is measured by MSE, the texture structure consistency of the preliminary adjustment image and the to-be-adjusted image is constrained, and the texture similarity loss is obtained ,
[0035] .
[0036] The total loss value The calculation includes
[0037] The color difference loss and the texture similarity loss are weighted and fused to obtain the total loss value ,
[0038] ;
[0039] Among them, , Indicate the weighting coefficient.
[0040] The optimization termination condition includes
[0041] After each iteration, it is determined whether the optimization termination condition is met, if the number of iterations reaches the maximum set value, the iteration is stopped, otherwise, the neural network parameters are iteratively optimized.
[0042] Beneficial effects: the present application directly uses the reference image and the image to be adjusted in the fixed scene to iteratively optimize the model parameters, directly fits the specific color mapping relationship between the actual data, does not need to rely on the pre-trained model, effectively avoids the problem that the pre-trained model has poor adjustment effect due to not seeing the specific scene image, has wide applicability of multiple types of scenes, and can be applied to remote sensing image color difference adjustment, fixed camera color difference adjustment across scene types.
[0043] It should be understood that the contents described in this part are not intended to identify the key or important features of the embodiments of the present application, nor are they used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0044] The accompanying drawings are used to better understand the present scheme and do not constitute a limitation on the present application. Among them:
[0045] Figure 1 is a whole process schematic diagram according to the present application;
[0046] Figure 2 is a model overall structure schematic diagram according to the present application;
[0047] Figure 3 is an encoder structure schematic diagram according to the present application;
[0048] Figure 4 is a decoder structure schematic diagram according to the present application;
[0049] Figure 5is a gradient descent optimization loss function curve diagram of multi-period remote sensing images according to the present application;
[0050] Figure 6 is a reference image schematic diagram of a multi-period remote sensing image scene according to the present application;
[0051] Figure 7 is a to-be-adjusted image schematic diagram of a multi-period remote sensing image scene according to the present application;
[0052] Figure 8 is an adjusted image schematic diagram of a multi-period remote sensing image scene according to the present application;
[0053] Figure 9 is a gradient descent optimization loss function curve diagram of fixed camera day and night scenes according to the present application;
[0054] Figure 10 is a reference image schematic diagram of a fixed camera day and night scene according to the present application;
[0055] Figure 11 is a to-be-adjusted image schematic diagram of a fixed camera day and night scene according to the present application;
[0056] Figure 12 is an adjusted image schematic diagram of a fixed camera day and night scene according to the present application. DETAILED DESCRIPTION
[0057] Exemplary embodiments of the present application are described below with reference to the accompanying drawings, which include various details of the embodiments of the present application to help in understanding them. These should be considered in their context only as illustrative. Thus, those of ordinary skill in the art will recognize various changes and modifications of the embodiments described herein, without departing from the scope and spirit of the present application. Also, descriptions of well-known functions and constructions are omitted in the following description for clarity and conciseness.
[0058] As shown in Figure 1 , the present application provides a gradient descent scene consistency adjustment method for fixed scenes, including:
[0059] S1: Collect image data to obtain a reference image and a to-be-adjusted image. It should be noted that:
[0060] The acquisition of the reference image includes,
[0061] In a fixed scene, a single high-quality image collected is taken as a reference image; at the same time, in the same scene, images obtained by changing the collection parameters (such as different shooting periods, different sensors, and different light weather) are taken as images to be adjusted. For example, in a remote sensing scene, multi-period images of the same region are collected, one of which is selected as a reference image, and the others are taken as images to be adjusted; in a fixed camera scene, a clear daytime image is collected as a reference image, and a low-illumination nighttime image is taken as an image to be adjusted.
[0062] S2: The reference image and the image to be adjusted are spliced in the channel dimension to form a double-input joint feature, the double-input joint feature is input into a double-branch multi-input neural network model, feature fusion is performed through stepwise down-sampling of an encoder and up-sampling of a decoder, and a preliminary adjusted image is output. It should be noted that:
[0063] The double-branch multi-input neural network model includes an encoder and a decoder.
[0064] The multi-input neural network model for scene image color consistency adjustment proposed in the application adopts an encoding-decoding integrated structure, and deeply integrates feature encoding and color adjustment decoding of double-input images in the same calculation process (the overall structure of the model is shown in Figure 2 The multi-input neural network model for scene image color consistency adjustment proposed in the application adopts an encoding-decoding integrated structure, and deeply integrates feature encoding and color adjustment decoding of double-input images in the same calculation process (the overall structure of the model is shown in
[0065] In the application, the color distribution conversion relationship between the reference image and the image to be adjusted can be abstracted as a continuous function, the encoder performs multi-scale extraction and abstract encoding on the joint features of the double-input images through multiple layers of convolution and nonlinear activation (linear rectifier function), and mines the potential correlation of the color distribution; the decoder gradually restores the abstract features after encoding to the image after color adjustment by means of deconvolution, feature splicing and other operations. The entire encoding-decoding neural network structure can be used as a multi-layer feedforward network, can simulate the color distribution conversion function between the reference image and the image to be adjusted, can learn and implement scene color consistency adjustment by using the approximation ability of the network, and can conform to the theoretical logic supported by the universal approximation theorem.
[0066] The encoder performs stepwise down-sampling coding on the reference image and the image to be adjusted to extract local features and global features; in the down-sampling process, the global features and the local features of the reference image and the image to be adjusted are abstractly coded, intermediate features carrying different spatial scales are output through different down-sampling levels, and the intermediate features include intermediate features 1 retaining relatively complete structure contours and intermediate features 2 mining more abstract semantic and color correlations.
[0067] The encoder structure (as shown in Figure 3 is described as follows:
[0068] Feature dimension splicing ([H, W, 6]): the reference image ([H, W, 3]) and the image to be adjusted ([H, W, 3]) are spliced in the channel dimension, the initial color and structure information of the two input images are fused, and a joint feature basis containing the differences and commonalities of the two images is provided for subsequent coding.
[0069] Convolution (kernel width 3, stride 2, feature 128) -> linear rectification function: the [H, W, 6] features after splicing are processed, a 3×3 convolution kernel is used to down-sample the features with a stride of 2 (the spatial dimension is compressed from [H, W] to [H / 2, W / 2]), local features are extracted and the amount of calculation is reduced, and the feature dimension is expanded to 128 channels; the linear rectification function introduces nonlinearity, enhances the expression ability of the model to complex color and structure features, and outputs the intermediate features 1 ([H / 2, W / 2, 128]).
[0070] Convolution (kernel width 3, stride 1, feature 128) -> linear rectification function: based on the [H / 2, W / 2, 128] features, a 3×3 convolution kernel is used to further extract local detail features with a stride of 1, the spatial dimension is kept unchanged, and the feature expression is deepened; the linear rectification function continues to inject nonlinearity, so that the features contain more rich color correlation patterns, and the intermediate features 1 ([H / 2, W / 2, 128]) are output.
[0071] Intermediate features 1 ([H / 2, W / 2, 128]): the features after two convolutions and activations are stored as an internal feature transmission node of the encoder, carrying the fusion features of the two images at the [H / 2, W / 2] scale, and providing key information for subsequent coding and decoder cross-stage fusion.
[0072] Convolution (kernel width 3, stride 2, feature 128) -> linear rectification function: the intermediate features 1 are down-sampled again, a 3×3 convolution kernel is used to compress the spatial dimension to [H / 4, W / 4] with a stride of 2, focus on more abstract global features, and reduce redundant information; the linear rectification function maintains nonlinearity, helping the model to mine deep color mapping rules of the two images.
[0073] Convolution(kernel width 3, stride 1, feature 128) -> Linear activation: 3x3 convolution kernel with stride 1 to extract local features, complementing the details lost in downsampling; linear activation to enhance non-linear representation, outputting intermediate feature 2 ([H / 4, W / 4, 128]), completing the encoding process from raw to abstract features for both input images.
[0074] Intermediate feature 2 ([H / 4, W / 4, 128]): stores features after multiple convolutions and activations, providing key information for subsequent decoder cross-stage fusion.
[0075] The decoder learns the color mapping relationship from the to-be-adjusted image to the reference image by upsampling and cross-stage feature fusion based on the feature difference between the reference image and the to-be-adjusted image extracted by the encoder, and restores the features abstracted by the encoder to image data after color adjustment using the learned color mapping relationship.
[0076] The decoder structure (as shown in Figure 4 ) is described as follows:
[0077] Deconvolution(kernel width 2, stride 2, feature 128): receives the intermediate feature 2 ([H / 4, W / 4, 128]) output by the encoder, and restores the feature space dimension by upsampling (from [H / 4, W / 4] to [H / 2, W / 2]) through deconvolution with kernel width 2 and stride 2, while keeping the channel number at 128, achieving preliminary spatial restoration of low-resolution encoded features and preparing for subsequent fusion of more rich scale information.
[0078] Intermediate feature 3 ([H / 2, W / 2, 128]): stores the feature data after deconvolution, serving as an internal feature transfer node in the decoder, bridging the pre- and post-processing procedures, and providing a carrier for fusion with intermediate feature 1.
[0079] Feature dimension concatenation ([H / 2, W / 2, 256]): concatenates the intermediate feature 3 ([H / 2, W / 2, 128]) with the intermediate feature 1 ([H / 2, W / 2, 128]) output by the encoder in the channel dimension, fusing feature information of different encoding stages and different spatial scales (intermediate feature 1 retains relatively more complete structural outlines, and intermediate feature 3 restores details after deconvolution), enabling the decoder to optimize color mapping learning using multi-scale features.
[0080] Deconvolution(kernel width 2, stride 2, feature 3): performs deconvolution with kernel width 2 and stride 2 on the concatenated [H / 2, W / 2, 256] feature, further restoring the spatial dimension to [H, W] and compressing the channel number to 3, preliminarily constructing a color feature mapping matching the output image size, and laying the foundation for restoring image color.
[0081] Intermediate Feature 4 ([H, W, 3]): Temporary storage of the deconvolved feature, at this time the feature dimension has been equal to the output image ([H, W, 3]), as a transition node, to prepare for the subsequent fusion of the original image to be adjusted information.
[0082] Feature Dimension Splicing ([H, W, 6]): Channel splicing of intermediate feature 4 ([H, W, 3]) and the original image to be adjusted ([H, W, 3]) is performed, introducing the color, structure details of the original image, making up for the possible loss of original information in the encoding-decoding process, assisting the decoder to learn the color adjustment strategy more accurately, and ensuring the relevance of the output image and the original scene.
[0083] Convolution (kernel width 1, stride 1, feature 3): 1x1 convolution is performed on the spliced [H, W, 6] feature, the channel number is compressed back to 3, the fused feature is integrated through linear transformation, and the key color adjustment information is focused, preparing for the final output.
[0084] Normalized exponential function: the Softmax function is used to normalize the features output by the convolution, and the adjusted value range is [0, 1].
[0085] Output image ([H, W, 3]): the result processed by the decoder full process is output, the output image color value range is [0, 255], and the output image is color adapted to the reference image and the details are retained intact.
[0086] S3: using the color difference-texture similarity joint loss function, the color difference loss between the preliminary adjustment image and the reference image and the structural similarity loss between the preliminary adjustment image and the image to be adjusted are calculated, and the color difference loss and the structural similarity loss are weighted and fused to obtain the total loss value. It should be noted that:
[0087] Wherein, the color difference loss between the preliminary adjustment image and the reference image is calculated to ensure color consistency, and the structural similarity loss between the preliminary adjustment image and the original image to be adjusted is to ensure the retention of texture details.
[0088] The color difference-texture similarity joint loss function promotes the network to adjust the color mapping relationship through the gradient feedback of the color difference loss, so that the overall color distribution of the image to be adjusted tends to the reference image, and at the same time, the color difference-texture similarity joint loss function forces the network to retain the edge contour and texture detail structure information of the original image when color conversion.
[0089] Let the reference image be , the image to be adjusted is , and the adjusted image output after processing by the neural network model is ;
[0090] Color difference loss function The calculation of the color difference loss function comprises,
[0091] ;
[0092] Wherein, H represents the height of the image, W represents the width of the image, 3 represents the number of RGB channels, i represents the number of image height direction, j represents the number of image width direction, and c represents the number of RGB three color channels.
[0093] The texture similarity loss function realizes detail preservation through edge texture feature constraint and avoids color interference.
[0094] The calculation of the texture similarity loss function comprises,
[0095] The gray image of the to-be-adjusted image and the preliminary adjustment image output after the neural network model processing are converted into gray images respectively And The Sobel operator is used to extract the texture of the gray image.
[0096] The definition of the Sobel operator x direction is:
[0097] ;
[0098] The definition of the Sobel operator y direction is:
[0099] ;
[0100] The texture map of the gray image of the to-be-adjusted image The extraction comprises,
[0101] ;
[0102] The texture map of the gray image of the preliminary adjustment image The extraction comprises,
[0103] .
[0104] The texture map of the gray image of the to-be-adjusted image And the texture map of the gray image of the preliminary adjustment image Global Softmax normalization is performed, and the texture value range is compressed to [0, 1];
[0105] ;
[0106] ;
[0107] Wherein, a texture probability map representing a grayscale image of the normalized image to be adjusted, a texture probability map representing a grayscale image of the normalized preliminary adjusted image, exp represents an exponential function with the natural constant e as the base, i represents the number in the image height direction, and j represents the number in the image width direction;
[0108] The difference of the normalized texture map is measured by the MSE, the texture structure consistency of the preliminary adjusted image and the image to be adjusted is constrained, and the texture similarity loss is obtained ,
[0109] .
[0110] S4: Based on the total loss value, the neural network parameters are iteratively optimized by using a gradient descent algorithm; when the optimization termination condition is met, the final adjusted image is output. It should be noted that:
[0111] The calculation of the total loss value includes,
[0112] The color difference loss and the texture similarity loss are weighted and fused to obtain the total loss value ,
[0113] ;
[0114] wherein, , both represent weighting coefficients, , which can be empirically taken as 0.5.
[0115] The judgment of the optimization termination condition includes,
[0116] After each iteration, it is judged whether the optimization termination condition is met. If the number of iterations reaches the maximum set value, the iteration is stopped; otherwise, the neural network parameters are iteratively optimized.
[0117] When the optimization termination condition is met, the final adjusted image is output. The final adjusted image is color consistent with the reference image, and retains the edge and texture key details of the image to be adjusted, realizing the color consistency adjustment of the image in a fixed scene.
[0118] Embodiment 1: Scene color consistency adjustment of multi-period remote sensing images
[0119] In this embodiment, remote sensing images of a certain place are selected as experimental data, and the specific data information is as follows:
[0120] Reference image: The image taken on May 28, 2024, with a size of 45591x39388 pixels, a spatial resolution of 1.8 meters, and a spectral range covering RGB three channels. The weather was sunny and the light was uniform during imaging, and the color features of the ground objects were clear, serving as a color anchor.
[0121] Image to be adjusted: The image taken on December 28, 2024, in the same area, with a size of 23844x20884 pixels and a spatial resolution of 1.8 meters. The weather was overcast during imaging, and the overall tone was dark, with a slight greenish deviation in the vegetation area.
[0122] The two sets of images were standardized by normalizing the pixel values to the [0, 255] interval, aligning the center point coordinates, and cropping the same range to scale the image size to 2000x2000 pixels. The image to be adjusted and the reference image were respectively concatenated in the channel dimension to form two sets of double-input features (size 2000x2000x6).
[0123] The double-branch multi-input neural network model proposed in this application was used, where the encoder contains 4 layers of convolution (kernel width 3, stride 2 / 1 alternation, feature number 128), and the decoder contains 2 layers of deconvolution (kernel width 2, stride 2, feature number 128). The gradient descent algorithm uses the Adam optimizer with an initial learning rate of 0.001 and a maximum of 1000 iterations.
[0124] The input features are input into the neural network model, and the "color difference-texture similarity" joint loss function proposed in this application is used as the optimization objective to start the parameter iteration optimization process and update the network weights in real time to minimize the loss value.
[0125] During the optimization process, the total loss function value curve is recorded as shown in Figure 5 The curve shows that the total loss value converges from the initial 0.15 to 0.04 at the 40th iteration, and remains stable thereafter. The curve is smoother than that of a conventional neural network model, with no obvious oscillation, showing good convergence characteristics.
[0126] As shown in Figure 6 to Figure 8 Before adjustment, the reference image presents a blue-green color overall, and the image to be adjusted has a color bias towards light green, with a significant difference in color distribution from the reference image. After adjustment, the colors of the two remote sensing images tend to be blue-green, and the color tone is uniformly improved. The details of the adjusted image are consistent with the image to be adjusted, and the visual consistency of the same ground object in different phases is significantly improved.
[0127] Example 2: Day and night scene color consistency adjustment of fixed camera
[0128] This embodiment selects the day and night video collected by a fixed camera in a certain seawall scene as experimental data. The scene contains typical elements such as vegetation and mountains. The specific data information is as follows:
[0129] Reference image: The image taken at 11:49:27 on June 19, 2024 is selected as the color reference benchmark. The image size is 1920x1080 pixels, and the color mode is RGB.
[0130] Image to be adjusted: The same scene image taken by the same camera at 19:34:03 on July 2, 2024 (no light at night) is selected. The image size is consistent with the reference image. Due to insufficient light, the overall color tone is dim, the scene recognition degree is reduced, and there are a small amount of noise.
[0131] The two images have completed spatial registration, ensuring that the positions of all objects in the scene correspond completely, and only the color and brightness deviations caused by the difference in day and night lighting exist.
[0132] The reference image and the image to be adjusted are preprocessed, the pixel values are normalized to the [0, 255] interval, the image size is kept at 1920x1080 pixels, and the two are spliced in the channel dimension to form a double-input feature (size 1920x1080x6).
[0133] The double-branch multi-input neural network model proposed in this application is adopted, wherein the encoder includes 4 layers of convolution (kernel width 3, stride 2 / 1 alternately, feature number 128), and the decoder includes 2 layers of deconvolution (kernel width 2, stride 2, feature number 128); the gradient descent algorithm selects the Adam optimizer, the initial learning rate is set to 0.001, and the upper limit of the number of iterations is set to 1000 times.
[0134] The double-input feature is input into the neural network model, the joint loss function is used as the optimization objective to start iterative optimization, and the network parameters are adjusted in real time through gradient back propagation, so that the output image gradually approaches the color features of the reference image.
[0135] During the optimization process, the change curves of the total loss value, the color difference loss value, and the structural similarity loss value are recorded (as shown in Figure 9 In the figure, the initial total loss value is 0.25; when the iteration reaches 50 times, the loss value slightly fluctuates; when the iteration reaches 100 times, the total loss value converges to 0.19, and the loss value basically does not change in the subsequent optimization process, meeting the convergence condition. The curve is smoother than the curve of the conventional neural network model, with no obvious fluctuation, showing good convergence characteristics.
[0136] As shown in Figure 10 to Figure 12As shown, before adjustment, the night image is dark and dull as a whole, and the difference in bright color tone with the daytime image is obvious; after adjustment, the overall brightness and color tone of the night image tend to be close to the daytime image, the sky presents bright color, the vegetation reappears green, and the overall scene is consistent. The adjusted night image is highly consistent with the daytime image in color, while retaining the original structure details of the night scene.
[0137] The above merely illustrates the specific embodiments of the present application, but the protection scope of the present application is not limited thereto, any change or replacement within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A gradient descent scene consistency adjustment method for a fixed scene, characterized in that: include: Collect image data and obtain a reference image and an image to be adjusted; The reference image and the image to be adjusted are spliced in the channel dimension to form a dual-input joint feature, which is input into a dual-branch multi-input neural network model. Feature fusion is performed through step-wise downsampling of the encoder and upsampling of the decoder to output a preliminary adjusted image. Using a color difference-texture similarity joint loss function, the color difference loss between the preliminary adjusted image and the reference image and the structural similarity loss between the preliminary adjusted image and the image to be adjusted are calculated, and the color difference loss and the structural similarity loss are weightedly fused to obtain a total loss value; Based on the total loss value, the neural network parameters are iteratively optimized using a gradient descent algorithm; when the optimization termination condition is met, the final adjusted image is output.
2. The method for adjusting the consistency of a gradient descent scene for a fixed scene according to claim 1, characterized in that: The acquisition of the reference image includes: In a fixed scene, a single high-quality image is collected as a reference image; At the same time, in the same scene, the image obtained by changing the acquisition parameter conditions is used as the image to be adjusted.
3. A method for adjusting scene consistency of gradient descent for a fixed scene according to claim 1 or 2, characterized in that: The dual-branch multi-input neural network model includes an encoder and a decoder.
4. The method for adjusting the consistency of a gradient descent scene for a fixed scene according to claim 3, characterized in that: The encoder performs step-by-step downsampling encoding on the reference image and the image to be adjusted to extract local features and global features; During the downsampling process, the global and local features of the reference image and the image to be adjusted are abstractly encoded, and intermediate features carrying different spatial scales are output through different downsampling levels.
5. The method for adjusting the consistency of a gradient descent scene for a fixed scene according to claim 4, characterized in that: Based on the feature differences between the reference image extracted by encoding and the image to be adjusted, the decoder gradually learns the color mapping relationship from the image to be adjusted to the reference image through upsampling and cross-stage feature fusion, and uses the learned color mapping relationship to restore the features abstracted by the encoder stage to the color-adjusted image data.
6. The method for adjusting the scene consistency of a fixed scene using gradient descent according to claim 5, characterized in that: The color difference-texture similarity joint loss function uses the gradient feedback of the color difference loss to drive the network to adjust the color mapping relationship, so that the overall color distribution of the image to be adjusted approaches that of the reference image. At the same time, the color difference-texture similarity joint loss function forces the network to retain the structural information of the edge contours and texture details of the original image during color conversion through the constraint of texture similarity loss.
7. The method for adjusting the scene consistency of a gradient descent method for a fixed scene according to claim 6, characterized in that: Set the base image to , the image to be adjusted is , the adjusted image output after processing by the neural network model is ; Color difference loss function The calculation includes, ; Where H represents the height of the image, W represents the width of the image, 3 represents the number of RGB channels, i represents the number of the image in the height direction, j represents the number of the image in the width direction, and c represents the number of the three RGB color channels; The calculation of the texture similarity loss function includes, Convert the image to be adjusted and the preliminary adjustment image output after processing by the neural network model into grayscale images respectively and , Sobel operator is used to extract texture from grayscale image; The Sobel operator in the x direction is defined as: ; The Sobel operator in the y direction is defined as: ; Texture map of the grayscale image to be adjusted The extraction includes, ; Initially adjust the texture map of the grayscale image The extraction includes, 。 8. The method for adjusting the scene consistency of a gradient descent method for a fixed scene according to claim 7, characterized in that: The texture map of the grayscale image to be adjusted and the texture map of the grayscale image of the preliminary adjusted image Perform global Softmax normalization to compress the texture value range to [0,1]; ; ; in, Represents the texture probability map of the normalized grayscale image to be adjusted, represents the texture probability map of the grayscale image of the normalized preliminary adjustment image, exp represents the exponential function with the natural constant e as the base, i represents the number in the height direction of the image, and j represents the number in the width direction of the image; The difference of the normalized texture map is measured by MSE, and the consistency of the texture structure between the initial adjustment image and the image to be adjusted is constrained to obtain the texture similarity loss , 。 9. The method for adjusting the scene consistency of a fixed scene using gradient descent according to claim 8, characterized in that: Total loss value The calculation includes, The color difference loss and texture similarity loss are weightedly fused to obtain the total loss value , ; in, 、 Both represent weighting coefficients.
10. A method for adjusting scene consistency of gradient descent for a fixed scene according to claim 1 or 9, characterized in that: The optimization termination conditions include: After each iteration, determine whether the optimization termination condition is met. If the number of iterations reaches the maximum set value, stop the iteration; Otherwise, continue to iteratively optimize the neural network parameters.
Citation Information
Patent Citations
Cross-domain image conversion method and device based on unsupervised neural network, computer equipment and storage medium
CN112819687A
Unsupervised image fusion method and system based on structure texture decomposition
CN115631428A
Mural image color restoration method and device based on migration refinement network
CN116993579A
Infrared image colorization method fusing Res-ASPP-Unet network
CN117576233A
End-to-end trace detection method
CN118015056A