Image processing method and device, electronic equipment and readable storage medium

By acquiring the brightness and color difference maps of an image and adjusting the text clarity, this method solves the quality degradation problem caused by image processing in existing technologies, achieving the effect of improving clarity while maintaining color.

CN115294055BActive Publication Date: 2026-05-29VIVO MOBILE COMM CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2022-08-03
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing methods for improving the clarity of text content in images can lead to a decrease in image quality, resulting in a reduction in the image's overall quality.

Method used

By acquiring the brightness and color difference maps of an image, feature information of the text content is extracted, the text clarity of the brightness map is adjusted, and a new image is generated by combining the color difference map, while preserving the original color difference without changing the color.

Benefits of technology

While improving text clarity, the image colors remain unchanged, ensuring that the image quality is not compromised.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115294055B_ABST
    Figure CN115294055B_ABST
Patent Text Reader

Abstract

The application discloses an image processing method and device, electronic equipment and a readable storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring a first luminance graph corresponding to a first image and a first color difference graph corresponding to the first image; acquiring first feature information of text content in the first luminance graph, and adjusting the definition of the text content in the first luminance graph according to the first feature information to obtain an adjusted first luminance graph; and generating a second image according to the first color difference graph and the adjusted first luminance graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to an image processing method, apparatus, electronic device, and readable storage medium. Background Technology

[0002] Currently, images captured by electronic devices may contain text content, such as product information. However, due to factors such as the shooting environment and angle, the text in the image may not be clear enough. Therefore, users need to process the image to make the text clearer.

[0003] In existing technologies, common processing methods include user-adjusted image contrast and color to make text in images clearer. However, these methods alter the inherent properties of the image, thereby reducing image quality.

[0004] It is evident that existing methods for improving the clarity of text content in images reduce image quality by altering the image's inherent properties. Summary of the Invention

[0005] The purpose of this application is to provide an image processing method that can solve the problem that existing methods for improving the clarity of text content in images reduce image quality by altering the image's inherent properties.

[0006] In a first aspect, embodiments of this application provide an image processing method, the method comprising: acquiring a first brightness map corresponding to a first image and a first color difference map corresponding to the first image; acquiring first feature information of text content in the first brightness map, and adjusting the clarity of the text content in the first brightness map according to the first feature information to obtain an adjusted first brightness map; and generating a second image according to the first color difference map and the adjusted first brightness map.

[0007] Secondly, embodiments of this application provide an image processing apparatus, which includes: an acquisition module, configured to acquire a first brightness map corresponding to a first image and a first color difference map corresponding to the first image; an adjustment module, configured to acquire first feature information of text content in the first brightness map and adjust the clarity of the text content in the first brightness map according to the first feature information to obtain an adjusted first brightness map; and a generation module, configured to generate a second image according to the first color difference map and the adjusted first brightness map.

[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions, when executed by the processor, implementing the steps of the method described in the first aspect.

[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0011] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0012] Thus, in the embodiments of this application, for a first image including text content, a first luminance map and a first color difference map corresponding to the first image are obtained respectively. First, first feature information of the text content is extracted from the first luminance map to adjust the clarity of the text content based on the first feature information, so as to make the text content clear. Then, based on the adjusted first luminance map and combined with the first color difference map, a second image is generated. The second image is generated based on the original color difference of the first image, which can preserve the color of the original image. It can be seen that, based on the embodiments of this application, only the luminance of the image is processed, and the color difference of the image is not processed. While ensuring the clarity of the text content in the image, the color of the image is not changed, thereby ensuring the high quality of the image. Attached Figure Description

[0013] Figure 1 This is a flowchart of an image processing method according to an embodiment of this application;

[0014] Figures 2 to 4 This is a schematic diagram of a model according to an embodiment of this application;

[0015] Figures 5 to 8 This is a schematic diagram illustrating the image processing method according to an embodiment of this application;

[0016] Figure 9 This is a block diagram of an image processing apparatus according to an embodiment of this application;

[0017] Figure 10 This is one of the hardware structure diagrams of the electronic device according to an embodiment of this application;

[0018] Figure 11 This is the second schematic diagram of the hardware structure of the electronic device according to an embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0020] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0021] The image processing method provided in this application embodiment can be executed by the image processing device provided in this application embodiment, or by an electronic device integrating the image processing device, wherein the image processing device can be implemented in hardware or software.

[0022] The image processing method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0023] Figure 1 A flowchart of an image processing method according to an embodiment of this application is shown, exemplified by its application to an electronic device, including:

[0024] Step 110: Obtain the first brightness map corresponding to the first image and the first color difference map corresponding to the first image.

[0025] Optionally, the electronic device captures an image of the target scene as the first image.

[0026] In this embodiment, the image processing method includes text content in the target scene, such as a large screen displaying text content in a conference room, and the user taking a picture of the large screen; or a paper product manual containing text content, and the user taking a picture of the product manual.

[0027] In this embodiment, due to factors such as the shooting environment and shooting angle, such as shooting from a rear position, the text content in the first image has clarity issues, such as blurring around the edges, color distortion, and noise.

[0028] In this step, based on the captured first image, corresponding image processing is performed to obtain the first brightness map and the first color difference map of the first image.

[0029] The first brightness map is used to represent the brightness information of each pixel in the first image, and the first color difference map is used to represent the color difference information of each pixel in the first image.

[0030] The data format of the first image is Red (R), Green (G), and Blue (B).

[0031] For reference, the formula for calculating the brightness information Y is:

[0032] Y = 0.299 * R + 0.587 * G + 0.114 * B

[0033] The formula for calculating color difference information U is:

[0034] U = -0.169 * R - 0.331 * G + 0.500 * B

[0035] The formula for calculating color difference information V is:

[0036] V = 0.500 * R - 0.439 * G - 0.081 * B

[0037] In the above formula, R, G, and B represent the R color channel value, G color channel value, and B color channel value of a pixel, respectively.

[0038] Step 120: Obtain the first feature information of the text content in the first brightness image, and adjust the clarity of the text content in the first brightness image according to the first feature information to obtain the adjusted first brightness image.

[0039] In this step, the first feature information of the text content is obtained based on the first brightness map.

[0040] Optionally, the first feature information includes stroke information of the text content, outline information of the text content, overall information of the text content, etc.

[0041] Optionally, the first feature information is a feature map.

[0042] Furthermore, based on the extracted first feature information, the clarity of the text content is adjusted to improve the clarity of the text content.

[0043] Optionally, this embodiment adjusts the clarity of the text content to the clarity of the text content under ideal conditions. For example, based on professional equipment, the clarity of the text content in an image captured under ideal conditions.

[0044] Step 130: Generate a second image based on the first color difference image and the adjusted first brightness image.

[0045] In this step, based on the adjusted first luminance map and combined with the first color difference map, a second image can be obtained. The second image has the same color difference information as the first image, that is, the colors are the same. The difference is that the text content in the second image is clearer than the text content in the first image.

[0046] Optionally, the first luminance map and the first chromaticity map are converted into RGB data, and the conversion formula includes:

[0047] R = Y + 1.402 * V;

[0048] G = Y - 0.344 * U - 0.792 * V;

[0049] B = Y + 1.772 * U.

[0050] Furthermore, by combining the conversion formula, a second image in RGB data format is generated.

[0051] Thus, in the embodiments of this application, for a first image including text content, a first luminance map and a first color difference map corresponding to the first image are obtained respectively. First, first feature information of the text content is extracted from the first luminance map to adjust the clarity of the text content based on the first feature information, so as to make the text content clear. Then, based on the adjusted first luminance map and combined with the first color difference map, a second image is generated. The second image is generated based on the original color difference of the first image, which can preserve the color of the original image. It can be seen that, based on the embodiments of this application, only the luminance of the image is processed, and the color difference of the image is not processed. While ensuring the clarity of the text content in the image, the color of the image is not changed, thereby ensuring the high quality of the image.

[0052] In another embodiment of the image processing method of this application, step 120 includes:

[0053] Sub-step A1: Divide the first brightness map into N1 sub-brightness maps, where N1 is a positive integer.

[0054] In this step, the first brightness map is cropped into N1 sub-brightness maps according to a preset order.

[0055] The image dimensions are: image width (W) x image height (H).

[0056] Optionally, the size of a typical captured image is "3000x4000", and the size of the sub-luminance map is "512x512". This image size is more conducive to capturing text content. The more sub-luminance maps are cropped, the more refined the processing, but this will also consume more computing resources and limit computing efficiency.

[0057] Sub-step A2: Obtain feature information of the text content in N1 sub-brightness maps.

[0058] In this step, feature information of the text content is extracted for each sub-brightness map.

[0059] Sub-step A3: Based on the feature information of the text content in the N1 sub-brightness maps, adjust the clarity of the text content in the N1 sub-brightness maps accordingly.

[0060] In this step, for each sub-brightness map, the clarity of its text content is adjusted based on the extracted feature information to improve the clarity of the text content in each sub-brightness map.

[0061] Sub-step A4: Based on the adjusted N1 sub-brightness maps, synthesize the adjusted first brightness map.

[0062] In this step, based on the adjusted sub-brightness maps, the first brightness map is restored according to a preset order.

[0063] In this embodiment, a large image is broken down into multiple smaller images, and the clarity of the text content in each smaller image is adjusted. This adjustment method is more meticulous, ensuring the clarity of the text content in the overall image after adjustment. Furthermore, processing large images consumes a significant amount of memory and computing power; this embodiment can also greatly reduce the use of memory and computing resources.

[0064] In another embodiment of the image processing method of this application, step 120 includes:

[0065] Sub-step B1: Based on N2 size information, adjust the size of the first brightness map N2 times to obtain N2 first brightness maps, where N2 is a positive integer.

[0066] In this embodiment, in order to obtain more comprehensive first feature information, the size of the first brightness map can be adjusted to obtain feature information of first brightness maps of different sizes, and then the feature information is fused to use the finally fused feature information as the first feature information, thereby ensuring that the obtained feature information is more comprehensive.

[0067] Referring to the previous embodiment, a sub-brightness map is used as a processing object.

[0068] Optionally, the text feature extraction module extracts the feature information of the text content.

[0069] See Figure 2 The text feature extraction module is shown in the figure. It employs a 3x3 convolutional layer with activation operations (indicated by arrows in the figure). This avoids excessively amplifying the receptive field, allowing the network to focus on the strokes of the text. Since typical text sizes are twelve pixels or larger, the 3x3 convolutional operation can compute adjacent 3x3 pixel regions during calculation. Figure 2 The four-layer 3x3 convolution operation can effectively extract feature information from text content. Specifically, in... Figure 2 In the diagram, feature map 201 is used to represent the feature map corresponding to the input sub-brightness map, and feature map 202 is used to represent the feature map corresponding to the output feature information.

[0070] Among them, see Figure 2 The input and output feature maps of the text feature extraction module are of the same size, both being [W, H, C]. W represents the width of the feature map, H represents the height of the feature map, and the values ​​of W and H are the same. C represents the number of channels of the feature map, which is usually "8, 16, 32".

[0071] Optionally, taking into account the size of the text, with the input feature map size being "512x512", the adjusted values ​​of W for the multiple sizes are "512, 256, 128". Feature maps of these sizes can better capture text content.

[0072] Optionally, feature maps can be scaled to different sizes after undergoing different types of convolution operations to obtain feature maps of different sizes. These feature maps contain different amounts of feature information. For example, a convolution with a stride of "2" will output a feature map containing more global information.

[0073] For example, see Figure 3 The feature map 301 with size [W, H, C] is processed by convolutional layers with activation stride = "1", convolutional layers with activation stride = "2", and convolutional layers with activation stride = "3" to obtain three sizes: [W, H, C], [W / 2, H / 2, C], and [W / 4, H / 4, C].

[0074] Sub-step B2: Obtain feature information of the text content in N2 first brightness images.

[0075] Optionally, in order to extract feature information corresponding to feature maps of different sizes, multiple corresponding text feature extraction modules can be used, with different text feature extraction modules processing feature maps of different sizes.

[0076] For example, see Figure 3 The feature map sizes processed by the three text feature extraction modules 302 are [W, H, C], [W / 2, H / 2, C], and [W / 4, H / 4, C], respectively.

[0077] Sub-step B3: Fuse the feature information of the text content in the N2 first brightness images to obtain the first feature information.

[0078] In this step, feature maps of different sizes need to be fused.

[0079] Optionally, the output feature maps can be upsampled to their original size. This upsampling is done through interpolation, which is computationally efficient and requires less computation.

[0080] For example, see Figure 3 First, the two feature maps 303 with dimensions [256, 256, C] and [128, 128, C] are directly concatenated to obtain feature map 304 with dimensions [512, 512, 2C]. Then, feature map 304 is fused with the output feature map with dimensions [512, 512, C] through the concat operation, which can effectively preserve the feature information of different dimensions of the text content. Finally, feature map 305 with dimensions [512, 512, 3C] is output.

[0081] In this embodiment, text content feature information can be extracted based on feature maps of multiple sizes. This can effectively handle the details of text strokes and the entire text content, which is beneficial for adjusting the clarity of the text content.

[0082] In another embodiment of the image processing method of this application, the multi-size feature map fusion step provided in the previous embodiment can be repeated to connect multiple fusion steps in series, thereby increasing the depth of the network and improving the expressive power of the model.

[0083] For example, see Figure 4 Taking a sub-brightness map as an example, sub-brightness map 401 corresponds to the state before the sharpness adjustment, and sub-brightness map 402 corresponds to the state after the sharpness adjustment.

[0084] exist Figure 4 In the process, based on the sub-brightness map 401, it is converted into the corresponding feature map 403. After the first multi-size feature fusion module 404 performs the first fusion, it is then fused a second time by the second multi-size feature fusion module 405. The output feature map 406 is concatted with the original feature map 403 to obtain the final feature map 407, which is then converted into the sub-brightness map 402.

[0085] Among them, Figure 4 In this process, the feature map output by each multi-size feature fusion module has a size of [512, 512, 3C]. After convolution and activation, the size is transformed to [512, 512, C]. Here, the length and width of the feature map are both "512", which is consistent with the size of the input feature map.

[0086] In this embodiment, the output feature map contains features of both text details and the text as a whole. Due to the increased depth of the network, the final output feature map has a larger receptive field and its size is consistent with the input size. Therefore, the final output feature map can also well take into account global brightness, contrast and other information, and its size is [512, 512, C].

[0087] It is worth mentioning that the concat operation is used in this application to facilitate the fast convergence of the model and to create a skip link between the input and output. Specifically, the two feature maps of size [512, 512, C] are directly concatenated along the third dimension to form a thicker feature map of size [512, 512, 2C].

[0088] For example, see Figure 4 Two feature maps of size [512, 512, C] are used. One is the input feature map 403, and the other is the output feature map 407. A concat operation is performed on the two feature maps.

[0089] For feature map 403, since the size of sub-brightness map 401 is [512, 512, 1], the input image needs to go through a convolutional layer first to modify its feature size to [512, 512, C].

[0090] In this embodiment, on the one hand, by extracting features from multiple multi-size feature maps, more detailed features of the text content can be extracted. On the other hand, by combining the characteristic information in the input original feature map, the final feature information of the text content becomes more comprehensive. Therefore, this embodiment utilizes the features of the text image to extract detailed features of the text content, effectively reconstructing text details and improving text clarity.

[0091] In another embodiment of the image processing method of this application, step 120 includes:

[0092] Sub-step C1: Obtain the second feature information of the text content in the first brightness map.

[0093] Sub-step C2: Based on the second feature information, adjust the clarity of the text content in the first brightness map to obtain the second brightness map.

[0094] Optionally, this embodiment may employ Document Ultra-High Definition Network (DocSR) for image processing to output a second image.

[0095] Optionally, the model corresponding to the document ultra-high definition network can be trained before outputting a clearer second image.

[0096] For reference, during the training process, the model outputs a predicted graph P1.

[0097] Among them, the prediction image P1 is obtained after adjusting the clarity of the text content based on the second feature information.

[0098] Optionally, the prediction map P1 corresponds to the sub-brightness map.

[0099] Sub-step C3: Obtain the third brightness map corresponding to the first brightness map. The first brightness map corresponds to the pixels of the third brightness map, and the clarity of the text content in the third brightness map meets the first preset condition.

[0100] In this step, images are taken based on the same scene to obtain the corresponding third brightness map from the captured images.

[0101] Alternatively, you can use professional equipment for shooting, such as a device with better optical properties, like a DSLR camera. Such devices are generally larger and have better optical imaging effects.

[0102] Therefore, in this step, the clarity of the text content in the third brightness map corresponding to the captured image can meet the first preset condition.

[0103] Optionally, the first preset condition is the clarity of the text content in the captured image under ideal shooting conditions.

[0104] Sub-step C4: Obtain the difference information between the second brightness map and the third brightness map.

[0105] Optionally, taking the predicted image P1 as an example, in the third brightness image, find the corresponding region y2 and calculate the difference between the two.

[0106] Furthermore, the difference values ​​of all regions in a brightness map are used as the difference information in this step.

[0107] Sub-step C5: If the difference information meets the second preset condition, the obtained second feature information is determined as the first feature information.

[0108] Optionally, the second preset condition is a minimum value.

[0109] Optionally, the difference between P1 and y2 is calculated and backpropagated; then, the difference between the two is calculated iteratively, causing the model to train in the direction where the difference is minimized. After one hundred epochs of iteration, the similarity between P1 and y2 becomes extremely high.

[0110] Furthermore, each sub-brightness map in a brightness map is used as a training sample to complete model training.

[0111] In the process of training the model, when cropping large images, there can be overlap between the cropped small images. Moreover, while ensuring computational efficiency, the number of small images can be increased to ensure sample diversity.

[0112] Therefore, after training the model, the second feature information output by the model is more accurate and can be used as the first feature information, which is then used in actual image processing.

[0113] In this embodiment, the model is trained based on the difference between the predicted image and the ideal image, so that in subsequent use scenarios, the clarity of the text content in the adjusted first brightness image output by the model is close to the clarity of the text content in the ideal image.

[0114] In the flow of the image processing method according to another embodiment of this application, step C4 includes:

[0115] Sub-step D1: Obtain the first absolute value of the difference between two corresponding pixel values ​​in the second brightness map and the third brightness map, obtain the corresponding first mean value based on the first absolute value, and use the first mean value as the first difference value.

[0116] In this embodiment, the difference between p1 and y2 is described in detail.

[0117] In this step, please refer to Formula 1:

[0118]

[0119] In Formula 1, L1 represents the first difference between p1 and y2; cnt represents the number of pixels in p1 or y2. For example, cnt = 512 x 512.

[0120] In this step, the first difference is applied to the entire p1 image to remove noise and improve overall detail.

[0121] Sub-step D2: Obtain the first stroke gradient map corresponding to the second luminance map, the second stroke gradient map corresponding to the third luminance map, and obtain the second absolute value of the difference between two corresponding pixel values in the first stroke gradient map and the second stroke gradient map. Obtain the corresponding second mean value according to the second absolute value, and use the second mean value as the second difference.

[0122] In this step, for p1, values are filled in the four directions of up, down, left, and right, and the filled area is cropped at the same time.

[0123] For example, refer to Figure 5 , for p1, the size is "512x512", the corresponding original area is 501, values are filled in the right direction (the thin rectangular strip on the right in the figure), and the newly cropped filled area 502 (shown by the dashed line) is obtained.

[0124] Furthermore, the filled areas cropped on the left and right are subtracted from each other, and the filled areas cropped on the up and down are subtracted from each other. The obtained results are superimposed, and finally the stroke gradient map is obtained.

[0125] For example, refer to Figure 6 the stroke gradient maps corresponding to the characters "one", "two", and "three" shown.

[0126] The same applies to y2.

[0127] In this step, refer to Formula 2:

[0128]

[0129] Among them, in Formula 2, L2 is used to represent the second difference between p1 and y2; grad1 and grad2 are respectively used to represent the stroke gradient maps of p1 and y2; cnt is used to represent the number of pixels of p1 or y2.

[0130] In this step, taking advantage of the fact that the background area of the image is relatively flat and the text area has obvious gradients, the edge area of the text is found by the method of shifting and subtracting, and the gradient difference between the predicted image and the ideal image is calculated.

[0131] Sub-step D3: According to the second stroke gradient map, expand the first area corresponding to the text content to the second area, identify the third area corresponding to the second area in the second luminance map, and obtain the third absolute value of the difference between two corresponding pixel values in the second area and the third area. Obtain the corresponding third mean value according to the third absolute value, and use the third mean value as the third difference.

[0132] In this step, based on the stroke gradient map of y2 obtained in the previous step, the first area where the text content is located can be found, and thus the first area can be expanded to obtain the expanded second area.

[0133] Optionally, perform dilation processing on the first region.

[0134] For example, refer to Figure 7 , take the stroke gradient map as the original image, binaryize the original image 701 into "0" and "1" values, where "0" is the black region and "1" is the white region. Use a "3x3" template 702 to sequentially search for all regions with a value of "1" and set the surrounding eight values to "1" to obtain the dilated image 703. For example, the comparison between the original text "一" 801 and the dilated text "一" 802 in the original image is as Figure 8 shown.

[0135] Furthermore, calculate the mask of the text region (i.e., the second region) of y2, and then calculate the average value of the absolute value between p1 and y2 in the same region.

[0136] In this step, refer to Formula 3:

[0137]

[0138] where, in Formula 3, L3 is used to represent the third difference between p1 and y2; mask_cnt is used to represent the number of pixels in the second region of y2.

[0139] In this step, calculate the difference between the prediction map and the ideal map according to the text region, so as to constrain the convergence of the entire model to focus on the text region, which is conducive to improving the clarity of the text.

[0140] Sub-step D4: Perform weighted processing on the first difference, the second difference and the third difference to obtain the difference information.

[0141] Since the above differences respectively focus on different dimensional information, weighted processing is performed on multiple differences to ensure that a single difference will not overly affect the convergence direction of the model.

[0142] Refer to Formula 4:

[0143] sum_loss = a * L1 + b * L2 + c * L3

[0144] In Formula 4, sum_loss is used to represent the difference information between p1 and y2, and a, b, c are the weight values of each difference. The values of a, b, c are ensured to make L1, L2, L3 in the same order of magnitude.

[0145] Furthermore, perform backpropagation on sum_loss and calculate sum_loss iteratively, so that the model is trained in the direction of minimizing sum_loss. After 100 epochs of iteration, the output p1 and the ideal y2 are relatively similar.

[0146] In this embodiment, model training considers not only the global information of the entire text content but also information such as the edges of the strokes and the overall shape of the characters. Therefore, during model training, the constraints outlined in this embodiment can effectively improve the overall quality of the text content. Ultimately, the trained model can effectively enhance the details of the character strokes, thereby improving text clarity and ultimately enhancing image quality.

[0147] In summary, this application can improve the clarity of captured text content in scenarios where text content is captured, without affecting the color information of the overall image. Specifically, this application comprehensively considers image characteristics such as: the strokes of characters are independent of each other; text blurring often manifests as unclear strokes of individual characters; the proportion of individual characters in the overall image is relatively small; and non-text areas are generally relatively flat and uniform. Based on these characteristics, a suitable network model is designed to process the captured images.

[0148] The image processing method provided in this application can be executed by an image processing device. This application uses an image processing device executing the image processing method as an example to illustrate the image processing device provided in this application.

[0149] Figure 9 A block diagram of an image processing apparatus according to another embodiment of this application is shown, the apparatus comprising:

[0150] The acquisition module 10 is used to acquire the first brightness map corresponding to the first image and the first color difference map corresponding to the first image;

[0151] The adjustment module 20 is used to obtain the first feature information of the text content in the first brightness image, and adjust the clarity of the text content in the first brightness image according to the first feature information to obtain the adjusted first brightness image.

[0152] The generation module 30 is used to generate a second image based on the first color difference map and the adjusted first brightness map.

[0153] Thus, in the embodiments of this application, for a first image including text content, a first luminance map and a first color difference map corresponding to the first image are obtained respectively. First, first feature information of the text content is extracted from the first luminance map to adjust the clarity of the text content based on the first feature information, so as to make the text content clear. Then, based on the adjusted first luminance map and combined with the first color difference map, a second image is generated. The second image is generated based on the original color difference of the first image, which can preserve the color of the original image. It can be seen that, based on the embodiments of this application, only the luminance of the image is processed, and the color difference of the image is not processed. While ensuring the clarity of the text content in the image, the color of the image is not changed, thereby ensuring the high quality of the image.

[0154] Optionally, the adjustment module 20 includes:

[0155] The splitting unit is used to split the first brightness map into N1 sub-brightness maps, where N1 is a positive integer;

[0156] The first acquisition unit is used to acquire feature information of the text content in N1 sub-brightness maps;

[0157] The first adjustment unit is used to adjust the clarity of the text content in the N1 sub-brightness maps according to the feature information of the text content in the N1 sub-brightness maps.

[0158] The synthesis unit is used to synthesize the adjusted first brightness map based on the adjusted N1 sub-brightness maps.

[0159] Optionally, the adjustment module 20 includes:

[0160] The second adjustment unit is used to adjust the size of the first brightness map N2 times based on N2 size information to obtain N2 first brightness maps, where N2 is a positive integer.

[0161] The second acquisition unit is used to acquire feature information of the text content in N2 first brightness images;

[0162] The fusion unit is used to fuse the feature information of the text content in N2 first brightness images to obtain the first feature information.

[0163] Optionally, the adjustment module 20 includes:

[0164] The third acquisition unit is used to acquire the second feature information of the text content in the first brightness image;

[0165] The third adjustment unit is used to adjust the clarity of the text content in the first brightness image according to the second feature information to obtain the second brightness image;

[0166] The fourth acquisition unit is used to acquire a third brightness map corresponding to the first brightness map, wherein the first brightness map corresponds to the pixels of the third brightness map, and the clarity of the text content in the third brightness map meets the first preset condition.

[0167] The fifth acquisition unit is used to acquire the difference information between the second brightness map and the third brightness map;

[0168] The determining unit is used to determine the acquired second feature information as the first feature information when the difference information meets the second preset condition.

[0169] Optionally, the fifth acquisition unit includes:

[0170] The first acquisition subunit is used to acquire the first absolute value of the difference between two corresponding pixel values ​​in the second brightness map and the third brightness map, acquire the corresponding first mean value based on the first absolute value, and use the first mean value as the first difference value.

[0171] The second acquisition subunit is used to acquire the first stroke gradient map corresponding to the second brightness map, the second stroke gradient map corresponding to the third brightness map, and to acquire the second absolute value of the difference between two corresponding pixel values ​​in the first stroke gradient map and the second stroke gradient map, acquire the corresponding second mean value based on the second absolute value, and use the second mean value as the second difference value.

[0172] The third acquisition subunit is used to expand the first region corresponding to the text content into a second region according to the second stroke gradient map, identify the third region corresponding to the second region in the second brightness map, and obtain the third absolute value of the difference between the two pixel values ​​corresponding to the second region and the third region, obtain the corresponding third mean value according to the third absolute value, and use the third mean value as the third difference value.

[0173] The weighted sub-unit is used to weight the first difference, the second difference, and the third difference to obtain the difference information.

[0174] The image processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0175] The image processing apparatus in this application embodiment can be a device with a motion system. This motion system can be an Android motion system, an iOS motion system, or other possible motion systems; this application embodiment does not specifically limit it.

[0176] The image processing apparatus provided in this application embodiment can implement the various processes implemented in the above method embodiments, and will not be described again here to avoid repetition.

[0177] Optionally, such as Figure 10 As shown, this application embodiment also provides an electronic device 100, including a processor 101, a memory 102, and a program or instructions stored in the memory 102 and executable on the processor 101. When the program or instructions are executed by the processor 101, they implement the various steps of any of the above-described image processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0178] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0179] Figure 11 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0180] The electronic device 1000 includes, but is not limited to, components such as: radio frequency unit 1001, network module 1002, audio output unit 1003, input unit 1004, sensor 1005, display unit 1006, user input unit 1007, interface unit 1008, memory 1009, and processor 1010.

[0181] Those skilled in the art will understand that the electronic device 1000 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1010 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 11 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0182] The processor 1010 is configured to acquire a first brightness map corresponding to a first image and a first color difference map corresponding to the first image; acquire first feature information of the text content in the first brightness map, and adjust the clarity of the text content in the first brightness map according to the first feature information to obtain an adjusted first brightness map; and generate a second image according to the first color difference map and the adjusted first brightness map.

[0183] Thus, in the embodiments of this application, for a first image including text content, a first luminance map and a first color difference map corresponding to the first image are obtained respectively. First, first feature information of the text content is extracted from the first luminance map to adjust the clarity of the text content based on the first feature information, so as to make the text content clear. Then, based on the adjusted first luminance map and combined with the first color difference map, a second image is generated. The second image is generated based on the original color difference of the first image, which can preserve the color of the original image. It can be seen that, based on the embodiments of this application, only the luminance of the image is processed, and the color difference of the image is not processed. While ensuring the clarity of the text content in the image, the color of the image is not changed, thereby ensuring the high quality of the image.

[0184] Optionally, the processor 1010 is further configured to split the first brightness image into N1 sub-brightness images, where N1 is a positive integer; obtain feature information of the text content in the N1 sub-brightness images; adjust the clarity of the text content in the N1 sub-brightness images according to the feature information of the text content in the N1 sub-brightness images; and synthesize the adjusted first brightness image according to the adjusted N1 sub-brightness images.

[0185] Optionally, the processor 1010 is further configured to adjust the size of the first brightness image N2 times based on N2 size information to obtain N2 first brightness images, where N2 is a positive integer; obtain feature information of the text content in the N2 first brightness images; and fuse the feature information of the text content in the N2 first brightness images to obtain the first feature information.

[0186] Optionally, the processor 1010 is further configured to: acquire second feature information of the text content in the first brightness image; adjust the clarity of the text content in the first brightness image according to the second feature information to obtain a second brightness image; acquire a third brightness image corresponding to the first brightness image, wherein the first brightness image corresponds to the pixels of the third brightness image, and the clarity of the text content in the third brightness image satisfies a first preset condition; acquire difference information between the second brightness image and the third brightness image; and, if the difference information satisfies the second preset condition, determine the acquired second feature information as the first feature information.

[0187] Optionally, the processor 1010 is further configured to: obtain a first absolute value of the difference between two corresponding pixel values ​​in the second brightness map and the third brightness map; obtain a corresponding first mean value based on the first absolute value; and use the first mean value as the first difference value; obtain a first stroke gradient map corresponding to the second brightness map and a second stroke gradient map corresponding to the third brightness map; obtain a second absolute value of the difference between two corresponding pixel values ​​in the first stroke gradient map and the second stroke gradient map; obtain a corresponding second mean value based on the second absolute value; and use the second mean value as the second difference value; expand the first region corresponding to the text content into a second region based on the second stroke gradient map; identify a third region corresponding to the second region in the second brightness map; obtain a third absolute value of the difference between two corresponding pixel values ​​in the second region and the third region; obtain a corresponding third mean value based on the third absolute value; and use the third mean value as the third difference value; and perform weighted processing on the first difference value, the second difference value, and the third difference value to obtain the difference information.

[0188] In summary, this application can improve the clarity of captured text content in scenarios where text content is captured, without affecting the color information of the overall image. Specifically, this application comprehensively considers image characteristics such as: the strokes of characters are independent of each other; text blurring often manifests as unclear strokes of individual characters; the proportion of individual characters in the overall image is relatively small; and non-text areas are generally relatively flat and uniform. Based on these characteristics, a suitable network model is designed to process the captured images.

[0189] It should be understood that, in this embodiment, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042. The GPU 10041 processes image data of still images or video images obtained by an image capture device (such as a camera) in video image capture mode or image capture mode. The display unit 1006 may include a display panel 10061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1007 includes a touch panel 10071 and at least one of other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include a touch detection device and a touch controller. Other input devices 10072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here. The memory 1009 can be used to store software programs and various data, including but not limited to applications and motion systems. Processor 1010 may integrate an application processor and a modem processor. The application processor mainly handles the action system, user page, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 1010.

[0190] The memory 1009 can be used to store software programs and various data. The memory 1009 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1009 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1009 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0191] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor 1010.

[0192] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image processing method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0193] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0194] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described image processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0195] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0196] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described image processing method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0197] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0198] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0199] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An image processing method, characterized in that, The method includes: Obtain the first brightness map corresponding to the first image, and the first color difference map corresponding to the first image; Obtain the first feature information of the text content in the first brightness image, and adjust the clarity of the text content in the first brightness image according to the first feature information to obtain the adjusted first brightness image. A second image is generated based on the first color difference map and the adjusted first brightness map; The step of obtaining the first feature information of the text content in the first brightness image includes: Obtain the second feature information of the text content in the first brightness image; Based on the second feature information, the clarity of the text content in the first brightness image is adjusted to obtain the second brightness image; Obtain a third brightness map corresponding to the first brightness map, wherein the first brightness map corresponds to the pixels of the third brightness map, and the clarity of the text content in the third brightness map meets a first preset condition. Obtain the difference information between the second brightness map and the third brightness map; If the difference information meets the second preset condition, the obtained second feature information will be determined as the first feature information; The step of obtaining the difference information between the second brightness map and the third brightness map includes: Obtain the first absolute value of the difference between two corresponding pixel values ​​in the second brightness map and the third brightness map, obtain the corresponding first mean value based on the first absolute value, and use the first mean value as the first difference; Obtain the first stroke gradient map corresponding to the second brightness map, the second stroke gradient map corresponding to the third brightness map, and obtain the second absolute value of the difference between two corresponding pixel values ​​in the first stroke gradient map and the second stroke gradient map. Obtain the corresponding second mean value based on the second absolute value and use the second mean value as the second difference value. Based on the second stroke gradient map, the first region corresponding to the text content is expanded into the second region, the third region corresponding to the second region in the second brightness map is identified, and the third absolute value of the difference between the two pixel values ​​corresponding to the second region and the third region is obtained. Based on the third absolute value, the corresponding third mean value is obtained, and the third mean value is used as the third difference value. The first difference, the second difference, and the third difference are weighted to obtain the difference information.

2. The method according to claim 1, characterized in that, The step of obtaining first feature information of the text content in the first brightness image and adjusting the clarity of the text content in the first brightness image based on the first feature information to obtain the adjusted first brightness image includes: The first brightness map is divided into N1 sub-brightness maps, where N1 is a positive integer; Obtain feature information of the text content in the N1 sub-brightness images; Based on the feature information of the text content in the N1 sub-brightness images, the clarity of the text content in the N1 sub-brightness images is adjusted accordingly. Based on the adjusted N1 sub-brightness maps, the adjusted first brightness map is synthesized.

3. The method according to claim 1, characterized in that, The step of obtaining the first feature information of the text content in the first brightness image includes: Based on N2 size information, the size of the first brightness image is adjusted N2 times to obtain N2 first brightness images, where N2 is a positive integer; Obtain feature information of the text content in the N2 first brightness images; The feature information of the text content in the N2 first brightness images is fused to obtain the first feature information.

4. An image processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire a first brightness map corresponding to the first image and a first color difference map corresponding to the first image; The adjustment module is used to obtain the first feature information of the text content in the first brightness image, and adjust the clarity of the text content in the first brightness image according to the first feature information to obtain the adjusted first brightness image. The generation module is used to generate a second image based on the first color difference image and the adjusted first brightness image; The adjustment module includes: The third acquisition unit is used to acquire the second feature information of the text content in the first brightness image; The third adjustment unit is used to adjust the clarity of the text content in the first brightness image according to the second feature information to obtain the second brightness image. The fourth acquisition unit is used to acquire a third brightness map corresponding to the first brightness map, wherein the first brightness map corresponds to the pixels of the third brightness map, and the clarity of the text content in the third brightness map meets the first preset condition. The fifth acquisition unit is used to acquire the difference information between the second brightness map and the third brightness map; The determining unit is configured to determine the acquired second feature information as the first feature information when the difference information satisfies the second preset condition; The fifth acquisition unit includes: The first acquisition subunit is used to acquire the first absolute value of the difference between two corresponding pixel values ​​in the second brightness map and the third brightness map, acquire the corresponding first mean value based on the first absolute value, and use the first mean value as the first difference value. The second acquisition subunit is used to acquire the first stroke gradient map corresponding to the second brightness map, the second stroke gradient map corresponding to the third brightness map, and to acquire the second absolute value of the difference between two corresponding pixel values ​​in the first stroke gradient map and the second stroke gradient map, acquire the corresponding second mean value based on the second absolute value, and use the second mean value as the second difference value. The third acquisition subunit is used to expand the first region corresponding to the text content into a second region according to the second stroke gradient map, identify the third region corresponding to the second region in the second brightness map, and obtain the third absolute value of the difference between the two pixel values ​​corresponding to the second region and the third region, obtain the corresponding third mean value according to the third absolute value, and use the third mean value as the third difference value. The weighting subunit is used to perform weighted processing on the first difference, the second difference, and the third difference to obtain the difference information.

5. The apparatus according to claim 4, characterized in that, The adjustment module includes: The splitting unit is used to split the first brightness map into N1 sub-brightness maps, where N1 is a positive integer; The first acquisition unit is used to acquire feature information of the text content in the N1 sub-brightness images; The first adjustment unit is used to adjust the clarity of the text content in the N1 sub-brightness images according to the feature information of the text content in the N1 sub-brightness images. The synthesis unit is used to synthesize the adjusted first brightness map based on the adjusted N1 sub-brightness maps.

6. The apparatus according to claim 4, characterized in that, The adjustment module includes: The second adjustment unit is used to adjust the size of the first brightness map N2 times based on N2 size information to obtain N2 first brightness maps, where N2 is a positive integer. The second acquisition unit is used to acquire feature information of the text content in the N2 first brightness images; The fusion unit is used to fuse the feature information of the text content in the N2 first brightness images to obtain the first feature information.

7. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the image processing method as described in any one of claims 1 to 3.

8. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the image processing method as described in any one of claims 1 to 3.