Image graying processing method, system and device and storage medium
The grayscale processing neural network selects the channel with the largest gradient information in the RGB channel as the label information, which solves the problem of information loss in image grayscale processing, and improves the detailed performance and recognition accuracy of grayscale processing.
Patent Information
- Application Number
- CN202510305806.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art has problems with information loss in image graying processing, especially when the image shooting environment changes, which affects the recognition accuracy.
The grayscale processing neural network is used to calculate the color conversion matrix of the image, and the channel with the largest gradient information among the three RGB channels is selected as the label information, so that more image details are retained by training the neural network.
The detailed performance capability of grayscale processing is improved, and the generated grayscale map is closer to the target grayscale map in details such as edges and textures, improving the recognition accuracy.
Smart Images

Figure CN120387958A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of image processing technology, and in particular to an image grayscale processing method, system, device and storage medium. Background Art
[0002] Compared with grayscale images, color images contain more information. However, limited by the actual industrial inspection requirements, in many cases, it is necessary to convert the image into a grayscale image before subsequent processing.
[0003] Image grayscaling refers to converting an RGB image into a single-channel grayscale image. The currently commonly used grayscale processing algorithm is the weighted average algorithm that is close to the human visual system. When the shooting environment of the image to be recognized, such as the venue, light, weather, etc., changes, a single grayscale processing algorithm is uniformly used for grayscale processing, resulting in the loss of some image features, thus affecting the recognition accuracy. Image grayscaling is a process of information suppression and will lose a large amount of image information.
[0004] Therefore, how to retain as much image information as possible while performing image grayscale processing and improve the performance of subsequent industrial inspections is an important problem to be solved. Summary of the Invention
[0005] To this end, the embodiments of the present application provide an image grayscale processing method, system, device and storage medium, which use a grayscale processing neural network to calculate the color conversion matrix of the image and convert the color information into grayscale information, retaining more image gradient information than traditional methods. Select the channel with the largest gradient information among the RGB three channels in the target grayscale image as the label information to ensure that more detailed information can be retained during the grayscaling process, so that the generated grayscale image is closer to the target grayscale image in terms of details such as edges and textures, thereby improving the detail performance of the grayscaling process.
[0006] To achieve the above object, the embodiments of the present application provide the following technical solutions:
[0007] According to the first aspect of the embodiments of the present application, an image grayscale processing method is provided, and the method includes:
[0008] Collect a number of image sample pairs, each image sample pair including an original image and the target grayscale image corresponding to the original image; input the original image into the initial model to generate a predicted grayscale image;
[0009] At each pixel position in the target grayscale image, determine the channel with the largest gradient information among the RGB three channels as the label information of the pixel position;
[0010] Determine the gradient information of the predicted grayscale image according to the label information, and calculate the total image gradient loss between the predicted grayscale image and the target grayscale image; train the grayscale processing neural network according to the total image gradient loss;
[0011] Perform RGB image normalization processing on the image to be processed; input the processed image to be processed into the grayscale processing neural network, and output the weights of the image to be processed in the three RGB channels respectively;
[0012] Obtain the grayscale image corresponding to the image to be processed based on the weights of the image to be processed in the three RGB channels respectively.
[0013] Optionally, the calculating the total image gradient loss between the predicted grayscale image and the target grayscale image includes:
[0014] At each pixel position, obtain the gradient information of the predicted grayscale image and the gradient information of the target grayscale image;
[0015] Calculate the image gradient loss between the predicted grayscale image and the target grayscale image according to the gradient information of the predicted grayscale image and the gradient information of the target grayscale image;
[0016] Calculate the total image gradient loss according to the image gradient losses at all pixel positions.
[0017] Optionally, the training the grayscale processing neural network according to the total image gradient loss includes:
[0018] If the total image gradient loss does not meet the convergence condition, adjust the parameters of the grayscale processing neural network according to the total image gradient loss until the total image gradient loss meets the convergence condition; the convergence condition includes that the total image gradient loss no longer decreases during the iteration process, or reaches a preset maximum number of iterations, or the total image gradient loss is less than a set loss threshold; the grayscale processing neural network includes a convolutional layer and a fully connected layer; the parameters of the convolutional layer include the number of input channels, the number of output channels, the convolutional kernel size, and the convolutional stride; the parameters of the fully connected layer include the number of input channels and the number of output channels.
[0019] Optionally, inputting the image to be processed into the grayscale processing neural network and outputting the weights of the image to be processed in the three RGB channels respectively includes:
[0020] Input the image to be processed into the grayscale processing neural network so that the grayscale processing neural network extracts image features through forward propagation calculation;
[0021] Calculate the weights of the three RGB channels based on the image features;
[0022] Normalize the weights of the three RGB channels so that the sum of the weights of the three RGB channels is 1.
[0023] Optionally, the RGB image standardization processing of the image to be processed includes:
[0024] Scale the image to be processed to the input size set during the training of the grayscale processing neural network by an interpolation method;
[0025] Determine whether the image to be processed is a grayscale image or an RGB image;
[0026] If the image to be processed is a grayscale image, adjust the channel order so as to expand the grayscale image into an RGB image;
[0027] Normalize the pixel values of the image to be processed to a set range.
[0028] Optionally, before normalizing the pixel values of the scaled image to be processed to a set range, the method further includes:
[0029] Perform enhancement processing on the image to be processed and randomly perform rotation and / or flipping operations.
[0030] Optionally, obtaining a grayscale image corresponding to the image to be processed based on the weights of the image to be processed in the three RGB channels respectively includes:
[0031] Calculate the grayscale value of each pixel according to the weights of the three RGB channels and the corresponding pixel values;
[0032] Combine the grayscale values of all pixels into a single-channel grayscale image.
[0033] According to the second aspect of the embodiments of the present application, an image grayscale processing system is provided, and the system includes:
[0034] A predicted grayscale map generation module, configured to collect a plurality of image sample pairs, each image sample pair including an original image and a target grayscale map corresponding to the original image; input the original image into the initial model to generate a predicted grayscale map;
[0035] A training module, configured to determine, at each pixel position in the target grayscale map, the channel with the largest gradient information among the three RGB channels as the label information of the pixel position; determine the gradient information of the predicted grayscale map according to the label information, and calculate the total image gradient loss between the predicted grayscale map and the target grayscale map; train the grayscale processing neural network according to the total image gradient loss;
[0036] A channel weight calculation module is used to perform RGB image normalization on the image to be processed; input the processed image to be processed into the grayscale processing neural network, and output the weights of the image to be processed in the three RGB channels respectively.
[0037] A grayscale image generation module is used to obtain the grayscale image corresponding to the image to be processed based on the weights of the image to be processed in the three RGB channels respectively.
[0038] According to the third aspect of the embodiments of the present application, an electronic device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor runs the computer program, it is configured to implement the method described in the first aspect above.
[0039] According to the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored. The computer-readable instructions can be executed by a processor to implement the method described in the first aspect above.
[0040] In summary, the embodiments of the present application provide an image grayscale processing method, system, device, and storage medium. By collecting a number of image sample pairs, each image sample pair includes an original image and the target grayscale image corresponding to the original image; inputting the original image into the initial model to generate a predicted grayscale image; at each pixel position in the target grayscale image, determining the channel with the largest gradient information among the three RGB channels as the label information of the pixel position; determining the gradient information of the predicted grayscale image according to the label information, and calculating the total image gradient loss between the predicted grayscale image and the target grayscale image; training the grayscale processing neural network according to the total image gradient loss; performing RGB image normalization on the image to be processed; inputting the processed image to be processed into the grayscale processing neural network, and outputting the weights of the image to be processed in the three RGB channels respectively; obtaining the grayscale image corresponding to the image to be processed based on the weights of the image to be processed in the three RGB channels respectively. Using a grayscale processing neural network to calculate the color conversion matrix of the image and convert the color information into grayscale information, more image gradient information is retained compared to traditional methods. Selecting the channel with the largest gradient information among the three RGB channels in the target grayscale image as the label information ensures that more detail information can be retained during the grayscale process, making the generated grayscale image closer to the target grayscale image in terms of details such as edges and textures, thereby improving the detail performance of the grayscale processing. Description of the Drawings
[0041] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only exemplary. For those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.
[0042] The structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they do not have substantial technical significance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0043] Figure 1 Schematic flow chart of an image grayscale processing method provided by an embodiment of the present application;
[0044] Figure 2 Structural diagram of a grayscale processing neural network provided by an embodiment of the present application;
[0045] Figure 3 Block diagram of an image grayscale processing system provided by an embodiment of the present application;
[0046] Figure 4 Schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0047] Figure 5 Schematic diagram of a computer-readable storage medium provided by an embodiment of the present application. Specific embodiments
[0048] The following specific embodiments illustrate the embodiments of the present invention. Those who are familiar with this technology can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0049] The embodiment of the present application provides an image grayscale processing method, which uses a grayscale processing neural network to calculate the color conversion matrix of the image, converts the color information into grayscale information, and retains more image gradient information compared with the traditional method. In the target grayscale image, the channel with the largest gradient information among the RGB three channels is selected as the label information, ensuring that more detailed information can be retained during the grayscale process, so that the generated grayscale image is closer to the target grayscale image in terms of details such as edges and textures, thereby improving the detail performance ability of the grayscale processing.
[0050] Figure 1 The image grayscale processing method provided by the embodiment of the present application is shown, including the following steps:
[0051] Step S101: Collect a number of image sample pairs, each image sample pair including the original image and the target grayscale image corresponding to the original image; input the original image into the initial model to generate a predicted grayscale image;
[0052] Step S102: At each pixel position in the target grayscale image, determine the channel with the largest gradient information among the RGB three channels as the label information of the pixel position;
[0053] Step S103: Determine the gradient information of the predicted grayscale image according to the label information, and calculate the total image gradient loss between the predicted grayscale image and the target grayscale image; train the grayscale processing neural network according to the total image gradient loss;
[0054] Step S104: Perform RGB image normalization processing on the image to be processed; input the processed image to be processed into the grayscale processing neural network, and output the weights of the image to be processed in the RGB three channels respectively;
[0055] Step S105: Obtain the grayscale image corresponding to the image to be processed based on the weights of the image to be processed in the RGB three channels respectively.
[0056] First, collect image sample pairs (original image and target grayscale image), and use a deep learning method (grayscale processing neural network) to perform grayscale processing on the original image to generate a predicted grayscale image. By comparing the gradient information of the predicted grayscale image and the target grayscale image, calculate the total image gradient loss, and train the neural network with this, so that the model can more accurately convert the color image into a grayscale image, reducing information loss and error during the grayscale process. By selecting the channel with the largest gradient information among the RGB three channels in the target grayscale image as the label information, it is ensured that more detailed information can be retained during the grayscale process.
[0057] In a possible implementation manner, in step 103, the calculation of the total image gradient loss between the predicted grayscale image and the target grayscale image includes:
[0058] At each pixel position, obtain the gradient information of the predicted grayscale image and the gradient information of the target grayscale image; calculate the image gradient loss between the predicted grayscale image and the target grayscale image according to the gradient information of the predicted grayscale image and the gradient information of the target grayscale image; calculate the total image gradient loss according to the image gradient losses at all pixel positions.
[0059] Traditional grayscale methods usually convert the RGB channels to grayscale values using simple formulas (such as weighted average method). This method may lose details, especially in the edge and texture parts. The method provided in the embodiments of this application calculates the gradient loss between the predicted grayscale image and the target grayscale image pixel by pixel, and the model can pay special attention to the details of the contour and texture parts. By separately calculating the gradient information of the predicted grayscale image and the target grayscale image at each pixel position and calculating the image gradient loss between them, the error of the grayscale process can be measured more precisely. This pixel-by-pixel gradient loss calculation method enables the model to more accurately capture the detailed differences in the grayscale process, thereby optimizing the grayscale effect.
[0060] In a possible implementation manner, in step 103, the training of the grayscale processing neural network according to the total image gradient loss includes:
[0061] If the total image gradient loss does not meet the convergence condition, adjust the parameters of the grayscale processing neural network according to the total image gradient loss until the total image gradient loss meets the convergence condition; the convergence condition includes that the total image gradient loss no longer decreases during the iteration process, or reaches a preset maximum number of iterations, or the total image gradient loss is less than a set loss threshold; the grayscale processing neural network includes a convolutional layer and a fully connected layer; the parameters of the convolutional layer include the number of input channels, the number of output channels, the size of the convolutional kernel, and the convolutional stride; the parameters of the fully connected layer include the number of input channels and the number of output channels.
[0062] During the training process, dynamically adjust the parameters of the neural network (such as the size and stride of the convolutional kernel in the convolutional layer, and the number of input and output channels in the fully connected layer) according to the total image gradient loss. This dynamic adjustment mechanism enables the model to adaptively optimize its own structure to better adapt to different image grayscale tasks. By adjusting the parameters of the convolutional layer and the fully connected layer, the model can more efficiently extract the feature information of the image and convert it into a high-quality grayscale image.
[0063] The model dynamically adjusts the parameters of the convolutional layer and the fully connected layer according to the total gradient loss of the image. For example: the convolutional layer adjusts the size and stride of the convolutional kernel to better capture the edge and texture features of the image. If the total gradient loss of the image indicates that there is a significant loss of edge details, the model may increase the size of the convolutional kernel to extract richer local features. The fully connected layer adjusts the number of input and output channels to optimize the feature integration process. If the model finds that certain feature channels contribute less to the grayscale quality, it may reduce the weights of these channels or adjust their connection methods. If the total gradient loss of the image does not decrease in consecutive iterations, it means that the model has converged and the training stops. If the preset maximum number of iterations (such as 1000 times) is reached, the training will stop even if the loss has not fully converged to avoid overtraining. If the total gradient loss of the image is less than the set loss threshold (such as 0.01), it means that the model has reached a high accuracy and the training stops. The convolutional layer can capture the local features of the image (such as edges and textures), while the fully connected layer is responsible for integrating these features to generate the final grayscale image. By dynamically adjusting the network parameters, the model can be adaptively optimized according to different image samples (the original image and the target grayscale image).
[0064] In a possible implementation manner, in step 104, the image to be processed is input into the grayscale processing neural network, and the weights of the image to be processed in the three RGB channels are output, including:
[0065] The image to be processed is input into the grayscale processing neural network, so that the grayscale processing neural network extracts image features through forward propagation calculation; calculates the weights of the three RGB channels based on the image features; and normalizes the weights of the three RGB channels so that the sum of the weights of the three RGB channels is 1.
[0066] The method provided by the embodiments of the present application does not rely on a fixed grayscale formula, but dynamically calculates weights through a neural network. By dynamically calculating the weights of the image to be processed in the three RGB channels through the grayscale processing neural network, the grayscale process can be adaptively adjusted according to the features of the input image. In this way, different types of images (such as landscape photos, portraits, or medical images) can obtain the grayscale results most suitable for their content, rather than relying on fixed weight allocation. The grayscale processing neural network extracts image features through forward propagation and calculates the RGB channel weights based on these features. This feature-based weight allocation method can better retain the detailed information of the image, such as edges, textures, and contrast, thereby improving the quality of the grayscale process.
[0067] Suppose there are the following different types of images that need to be grayscaled:
[0068] Category 1: Landscape photos (dominated by blue sky and green trees); Traditional grayscale methods use a fixed RGB weight formula (such as 0.299R + 0.587G + 0.114B) for grayscaling, which may cause loss of details in the blue sky and green trees because the fixed weights cannot fully consider the characteristics of the image. In this application, after the grayscale processing neural network extracts the image features through forward propagation, it determines that the green (G channel) is more important in the green tree part, and the blue (B channel) is more important in the blue sky part. Therefore, the network dynamically calculates that the weight of the G channel is higher, the weight of the B channel is the second, and the weight of the R channel is lower. After normalization, the sum of the weights is 1. This adaptive weight assignment can better retain the details of the green trees and blue sky and generate high-quality grayscale images.
[0069] Category 2: Portrait photos (dominated by skin color); Using a fixed weight formula in traditional grayscale methods may cause the grayscale value of the skin color part to be too dark or too bright, unable to accurately reflect the facial features of the person. After the neural network in this application extracts the image features, it determines that the R channel is more important in the skin color part. Therefore, the network dynamically calculates that the weight of the R channel is higher, and the weights of the G and B channels are lower. After normalization, the sum of the weights is 1. This weight assignment can better retain the details of the skin color and make the grayscale result of the portrait photo more natural and realistic.
[0070] Category 3: Medical images (such as X-ray images); Using a fixed weight formula in traditional grayscale methods may not be able to highlight the key details in medical images, such as the contrast of bones or soft tissues. After the neural network in this application extracts the image features, it dynamically calculates the weights according to the characteristics of the medical image. For example, if the contrast of bones in the image mainly depends on the R channel, the network will assign a higher weight to the R channel and adjust the weights of the G and B channels to optimize the overall grayscaling effect. After normalization, the sum of the weights is 1. This adaptive weight assignment can better highlight the key information in medical images and facilitate doctors' diagnosis.
[0071] By dynamically calculating and normalizing the RGB channel weights, high-quality and adaptive image grayscaling processing is achieved. This feature-based weight assignment method can flexibly adjust the grayscaling process according to the characteristics of the input image, while maintaining the consistency and stability of the grayscaling result, enhancing the generality and flexibility of the model, and enabling it to adapt to various complex scenarios and different types of images.
[0072] In a possible implementation manner, in step 104, the RGB image standardization processing of the image to be processed includes:
[0073] The to-be-processed image is scaled to the input size set during the training of the grayscale processing neural network through an interpolation method; it is determined whether the to-be-processed image is a grayscale image or an RGB image; if the to-be-processed image is a grayscale image, the channel order is adjusted so that the grayscale image is expanded into an RGB image; the pixel values of the to-be-processed image are normalized to a set range.
[0074] In a possible implementation manner, in step 104, before normalizing the pixel values of the scaled to-be-processed image to the set range, the method further includes: performing enhancement processing on the to-be-processed image and randomly performing rotation and / or flipping operations.
[0075] By scaling the to-be-processed image to the input size set during the neural network training, it is ensured that the image can adapt to the input requirements of the model. This normalization processing can avoid processing errors caused by inconsistent image sizes and improve the compatibility and stability of the model. Further, it is determined whether the to-be-processed image is a grayscale image or an RGB image, and the grayscale image is expanded into an RGB image (by adjusting the channel order). This processing method enables the model to uniformly process input images of different formats and avoid processing problems caused by differences in image types.
[0076] The image can also be enhanced (such as randomly rotating and / or flipping), which can increase the diversity of the input image and simulate more possible image scenarios. This data augmentation method can improve the robustness of the model, enabling it to maintain good performance when facing input images at different angles and in different directions. Normalizing the pixel values of the image to a set range (such as [0, 1] or [-1, 1]) can reduce the pixel value differences between different images and avoid model training biases caused by different numerical ranges.
[0077] In a possible implementation manner, in step 105, obtaining the grayscale image corresponding to the to-be-processed image based on the weights of the to-be-processed image in the three RGB channels includes:
[0078] Calculating the grayscale value of each pixel according to the weights of the three RGB channels and the corresponding pixel values; combining the grayscale values of all pixels into a single-channel grayscale image.
[0079] By using dynamically calculated RGB channel weights (instead of a fixed weight formula) to calculate the grayscale value of each pixel, the characteristics and details of the input image can be more accurately reflected. This method avoids the problem of detail loss or distortion caused by fixed weights in traditional grayscale conversion methods, thereby generating a higher-quality grayscale image. Finally, the grayscale values of all pixels are combined into a single-channel grayscale image, and this single-channel output conforms to the standard format of grayscale images, facilitating subsequent processing and applications (such as image analysis, feature extraction, display, etc.).
[0080] The following will provide a detailed description of the image grayscale processing method provided by the embodiments of the present application in conjunction with the accompanying drawings.
[0081] In the first aspect, a neural network architecture and configuration parameters are built.
[0082] Figure 2 The overall framework of the grayscale processing neural network model provided by the embodiments of the present application is shown, including:
[0083] 1. Convolutional layer: The model contains four convolutional layers, and the parameters of each layer are as follows:
[0084] The first layer: The number of input channels is 3 (RGB image), the number of output channels is 32, the convolutional kernel size is 3x3, and the stride is 1. The second layer: The number of input channels is 32, the number of output channels is 32, the convolutional kernel size is 3x3, and the stride is 2. The third layer: The number of input channels is 32, the number of output channels is 64, the convolutional kernel size is 3x3, and the stride is 2. The fourth layer: The number of input channels is 64, the number of output channels is 128, the convolutional kernel size is 3x3, and the stride is 1.
[0085] 2. Pooling layer: The model contains an average pooling layer and a max pooling layer, which are used to reduce the dimension of the feature map while retaining important features.
[0086] 3. Fully connected layer: The model contains three fully connected layers, and the parameters are as follows: The first layer: The number of input channels is 256, and the number of output channels is 512. The second layer: The number of input channels is 512, and the number of output channels is 512. The third layer: The number of input channels is 512, and the number of output channels is 3 (corresponding to the weights of the RGB three channels).
[0087] 4. Activation function: The ReLU activation function is used after each convolutional layer to increase the non-linear expression ability of the network.
[0088] 5. Output layer: The output layer does not use an activation function and directly outputs three weight values for calculating the grayscale image.
[0089] In a possible implementation manner, the parameters of the grayscale processing neural network structure are configured: Optimizer: Use the Adam optimizer, and the learning rate is set to 1e-4. Loss function: Use the L1 loss function to calculate the difference between the predicted grayscale image gradient information and the maximum gradient information in the RGB three channels. Gradient calculation: For each pixel, calculate the gradients in the horizontal and vertical directions, and then take the sum of squares to obtain the gradient information of each channel. Take the maximum of the gradient information in the RGB three channels as the label. Number of training epochs: Conduct 1000 epochs of training. Batch size: Each batch contains 2 image samples.
[0090] The second stage: Train the grayscale processing neural network. The specific training process includes the following steps:
[0091] Step 1: Collect a number of image sample pairs, each image sample pair including an original image and the target grayscale image corresponding to the original image; input the original image into the initial model to generate a predicted grayscale image; in this step, the original color image (RGB image) is input into a convolutional neural network. The convolutional layer of the network is responsible for extracting features of the image, such as edges, textures, etc. These features are passed to the fully connected layer, and the fully connected layer further processes these features to generate the predicted grayscale image.
[0092] Step 2: At each pixel position in the target grayscale image, determine the channel with the largest gradient information among the three RGB channels as the label information for the pixel position; during the training process, it is necessary to determine the label information for each pixel position, that is, which channel among the three RGB channels has the largest gradient information. The gradient information is used for the training of the network, enabling the network to learn how to generate a more accurate grayscale image based on the gradient information.
[0093] Step 3: Determine the gradient information of the predicted grayscale image according to the label information, and calculate the total image gradient loss between the predicted grayscale image and the target grayscale image; train the grayscale processing neural network according to the total image gradient loss; the fully connected layer of the network outputs the predicted grayscale image, and then calculates the gradient loss between the predicted grayscale image and the target grayscale image. The gradient loss is used in the training process of the network, and by adjusting network parameters (such as the convolution kernel size and stride of the convolutional layer, and the number of input and output channels of the fully connected layer) to minimize the loss.
[0094] Step 4: Perform RGB image normalization processing on the image to be processed; input the processed image to be processed into the grayscale processing neural network. The convolutional layer of the network extracts image features, and the fully connected layer calculates the weights of each pixel in the three RGB channels and outputs the weights of the image to be processed in the three RGB channels respectively. Among them, the normalization processing may include data preprocessing and enhancement processes, such as scaling the image to 128x128 pixels, and / or randomly performing rotation and flipping operations, and / or normalizing using the mean and standard deviation of the ImageNet dataset.
[0095] Step 5: Obtain the grayscale image corresponding to the image to be processed based on the weights of the image to be processed in the three RGB channels respectively. The fully connected layer generates the final grayscale image according to the calculated weights. These weights reflect the contribution of each pixel in the three RGB channels, so that the generated grayscale image can more accurately reflect the features of the original image.
[0096] In the loss function provided in the embodiment of the present application, the largest gradient information in the three RGB channels is used as the label information, and the L1 loss of the gradient information of the generated grayscale image. The largest gradient information is calculated as follows:
[0097] Grad label = Max(Grad R , Grad G , Grad B )
[0098] GradR, GradG, and GradB represent the gradient information calculated on the red, green, and blue channels respectively. Gradient information is usually used to describe the intensity of brightness or color changes in an image, and it can be obtained by calculating the differences in pixel values in the horizontal and vertical directions.
[0099] The pixel gradient information of RGB is calculated as follows:
[0100] Grad K = (K[1:, 1:] - K[1:, :-1]) 2 + (K[1:, :-1] - K[:-1, :-1]) 2
[0101] where K takes the value of the RGB grayscale image, representing different color channels. It is the sum of the square of the difference between the pixel below the current pixel and the current pixel and the square of the difference between the current pixel and the pixel to the left of the current pixel, that is, the sum of the squares of the horizontal direction gradient and the vertical direction gradient. This formula is only for experimental implementation, and the same effect can be achieved using different implementation schemes such as subtracting the right side from the left side, subtracting the lower side from the upper side, and taking the absolute value.
[0102] Through the loss function, the model can learn how to generate high-quality grayscale images based on the features of the input image. This structure and training method enable the network to adaptively process different types of images and generate grayscale images with rich details and high quality.
[0103] In summary, the embodiment of the present application provides an image grayscale processing method. By collecting a number of image sample pairs, each image sample pair includes an original image and the target grayscale image corresponding to the original image; inputting the original image into the initial model to generate a predicted grayscale image; at each pixel position in the target grayscale image, determining the channel with the largest gradient information among the three RGB channels as the label information of the pixel position; determining the gradient information of the predicted grayscale image according to the label information, and calculating the total image gradient loss between the predicted grayscale image and the target grayscale image; training the grayscale processing neural network according to the total image gradient loss; performing RGB image normalization processing on the image to be processed; inputting the processed image to be processed into the grayscale processing neural network, and outputting the weights of the image to be processed in the three RGB channels respectively; obtaining the grayscale image corresponding to the image to be processed based on the weights of the image to be processed in the three RGB channels respectively. Using the grayscale processing neural network to calculate the color conversion matrix of the image and convert the color information into grayscale information, more image gradient information is retained compared with the traditional method. Selecting the channel with the largest gradient information among the three RGB channels in the target grayscale image as the label information ensures that more detailed information can be retained during the grayscale process, making the generated grayscale image closer to the target grayscale image in terms of details such as edges and textures, thereby improving the detail performance ability of the grayscale processing.
[0104] Based on the same technical concept, the embodiment of the present application further provides an image grayscale processing system, as Figure 3 shown. The system includes:
[0105] A predicted grayscale image generation module 301, configured to collect a number of image sample pairs, each image sample pair includes an original image and the target grayscale image corresponding to the original image; input the original image into the initial model to generate a predicted grayscale image;
[0106] A training module 302, configured to, at each pixel position in the target grayscale image, determine the channel with the largest gradient information among the three RGB channels as the label information of the pixel position; determine the gradient information of the predicted grayscale image according to the label information, and calculate the total image gradient loss between the predicted grayscale image and the target grayscale image; train the grayscale processing neural network according to the total image gradient loss;
[0107] A channel weight calculation module 303, configured to perform RGB image normalization processing on the image to be processed; input the processed image to be processed into the grayscale processing neural network, and output the weights of the image to be processed in the three RGB channels respectively;
[0108] A grayscale image generation module 304, configured to obtain the grayscale image corresponding to the image to be processed based on the weights of the image to be processed in the three RGB channels respectively.
[0109] Embodiments of the present application also provide an electronic device corresponding to the method provided in the foregoing embodiments. Please refer to Figure 4 , which shows a schematic diagram of an electronic device provided in some embodiments of the present application. The electronic device 20 may include: a processor 200, a memory 201, a bus 202, and a communication interface 203. The processor 200, the communication interface 203, and the memory 201 are connected through the bus 202. A computer program that can run on the processor 200 is stored in the memory 201. When the processor 200 runs the computer program, it executes the method provided in any of the foregoing embodiments of the present application.
[0110] Among them, the memory 201 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one physical port (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0111] The bus 202 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 201 is used to store a program. After receiving an execution instruction, the processor 200 executes the program. The method disclosed in any of the foregoing embodiments of the present application can be applied to or implemented by the processor 200.
[0112] The processor 200 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 200 or the instructions in the form of software. The above-mentioned processor 200 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware decoding processor, or executed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 201, and the processor 200 reads the information in the memory 201 and combines its hardware to complete the steps of the above method.
[0113] The electronic device provided by the embodiment of the present application and the method provided by the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by it.
[0114] The embodiment of the present application also provides a computer-readable storage medium corresponding to the method provided by the foregoing embodiment. Please refer to Figure 5 ., the shown computer-readable storage medium is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the method provided by any of the foregoing embodiments.
[0115] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here one by one.
[0116] The computer-readable storage medium provided by the above embodiment of the present application and the method provided by the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by the application program stored therein.
[0117] It should be noted that:
[0118] The algorithms and displays provided herein are not inherently related to any particular computer, virtual apparatus, or other device. Various general-purpose apparatuses may also be used in conjunction with the teachings presented herein. The structure required to construct such apparatuses will be apparent from the above description. In addition, the present application is not directed to any particular programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the description of the specific language above is for the purpose of disclosing the best mode of the present application.
[0119] In the specification provided herein, a number of specific details are set forth. However, it can be understood that embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0120] Similarly, it should be understood that in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all of the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.
[0121] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.
[0122] In addition, those skilled in the art will understand that although some of the embodiments described herein include certain features included in other embodiments but not others, the combination of features of different embodiments is meant to be within the scope of this application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0123] Each component embodiment of this application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that in practice, a microprocessor or a digital signal processor (DSP) can be used to implement some or all of the functions of some or all of the components in the virtual machine creation device according to the embodiments of this application. This application can also be implemented as a device or device program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0124] It should be noted that the above embodiments illustrate rather than limit this application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
[0125] As described above, only the preferred specific embodiments of this application are provided, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by this application should be covered by the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claimed claims.
Claims
1. An image grayscale processing method, characterized in that, The method includes: Collecting a number of image sample pairs, where each image sample pair includes an original image and a target grayscale image corresponding to the original image; inputting the original image into the initial model to generate a predicted grayscale image; At each pixel position in the target grayscale image, determining the channel with the largest gradient information among the three RGB channels as the label information of the pixel position; Determining the gradient information of the predicted grayscale image according to the label information, and calculating the total image gradient loss between the predicted grayscale image and the target grayscale image; training the grayscale processing neural network according to the total image gradient loss; Performing RGB image normalization processing on the image to be processed; inputting the processed image to be processed into the grayscale processing neural network, and outputting the weights of the image to be processed in the three RGB channels respectively; Obtaining the grayscale image corresponding to the image to be processed based on the weights of the image to be processed in the three RGB channels respectively.
2. The method according to claim 1, characterized in that The calculating of the total image gradient loss between the predicted grayscale image and the target grayscale image includes: At each pixel position, obtaining the gradient information of the predicted grayscale image and the gradient information of the target grayscale image; Calculating the image gradient loss between the predicted grayscale image and the target grayscale image according to the gradient information of the predicted grayscale image and the gradient information of the target grayscale image; Calculating the total image gradient loss according to the image gradient losses at all pixel positions.
3. The method according to claim 2, characterized in that, The training of the grayscale processing neural network according to the total image gradient loss includes: If the total image gradient loss does not meet the convergence condition, adjusting the parameters of the grayscale processing neural network according to the total image gradient loss until the total image gradient loss meets the convergence condition; the convergence condition includes that the total image gradient loss no longer decreases during the iteration process, or reaches a preset maximum number of iterations, or the total image gradient loss is less than a set loss threshold; the grayscale processing neural network includes a convolutional layer and a fully connected layer; the parameters of the convolutional layer include the number of input channels, the number of output channels, the convolutional kernel size, and the convolutional stride; the parameters of the fully connected layer include the number of input channels and the number of output channels.
4. The method according to claim 1, wherein Inputting the image to be processed into the grayscale processing neural network and outputting the weights of the image to be processed in the three RGB channels respectively includes: Inputting the image to be processed into the grayscale processing neural network so that the grayscale processing neural network extracts image features through forward propagation calculation; Calculating the weights of the three RGB channels based on the image features; Performing normalization processing on the weights of the three RGB channels so that the sum of the weights of the three RGB channels is 1.
5. The method according to claim 4, characterized in that, The performing of RGB image normalization processing on the image to be processed includes: Scaling the image to be processed to the input size set during the training of the grayscale processing neural network by an interpolation method; Judging whether the image to be processed is a grayscale image or an RGB image; If the image to be processed is a grayscale image, performing channel order adjustment to expand the grayscale image into an RGB image; Normalizing the pixel values of the image to be processed to a set range.
6. The method according to claim 5, wherein Before normalizing the pixel values of the processed image after scaling to a set range, the method further includes: Performing enhancement processing on the processed image and randomly performing rotation and / or flipping operations.
7. The method according to claim 1, characterized in that, Obtaining a grayscale image corresponding to the processed image based on the weights of the processed image in the three RGB channels, including: Calculating the grayscale value of each pixel according to the weights of the three RGB channels and the corresponding pixel values; Combining the grayscale values of all pixels into a single-channel grayscale image.
8. An image grayscale processing system, characterized in that, The system includes: A predicted grayscale image generation module, configured to collect a plurality of image sample pairs, each image sample pair including an original image and a target grayscale image corresponding to the original image; inputting the original image into the initial model to generate a predicted grayscale image; A training module, configured to determine, at each pixel position in the target grayscale image, the channel with the largest gradient information among the three RGB channels as the label information of the pixel position; determining the gradient information of the predicted grayscale image according to the label information, and calculating the total image gradient loss between the predicted grayscale image and the target grayscale image; training the grayscale processing neural network according to the total image gradient loss; A channel weight calculation module, configured to perform RGB image normalization processing on the processed image; inputting the processed image into the grayscale processing neural network, and outputting the weights of the processed image in the three RGB channels respectively; A grayscale image generation module, configured to obtain a grayscale image corresponding to the processed image based on the weights of the processed image in the three RGB channels respectively.
9. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A computer-readable instruction is stored thereon, and the computer-readable instruction can be executed by the processor to implement the method according to any one of claims 1-7.