Method and device for detecting media font materials
By converting RGB images to the LAB color space using a deep learning model, and combining a feature pyramid network and a self-attention mechanism, the problem of color extraction in complex and colorful text and background scenes is solved. This enables accurate identification of background color, fill color, and stroke color in various complex scenes, improving the accuracy and stability of color extraction.
Patent Information
- Application Number
- CN202510834079.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-17
AI Technical Summary
Existing text color extraction techniques are not accurate enough in complex and colorful text and background scenes, especially when the foreground and background colors are similar, they are prone to recognition errors. Traditional methods such as clustering based on the RGB color space and Log-Gabor filters have limited generalization ability in colorful text and complex backgrounds.
Using a deep learning model, RGB images are converted to the LAB color space. The multi-scale feature maps are then processed by combining the Feature Pyramid Network (FPN), the Channel Attention Block (SE Block), and the Self-Attention mechanism. This generates multi-scale feature maps and adjusts the channel weights to enhance attention to key color channels, capture the correlation of global colors in the image, and output LAB color prediction values for background color, fill color, and stroke color.
This greatly enhances the model's generalization ability and adaptability in various complex scenarios, enabling it to accurately identify color information in colorful text and complex backgrounds, thus improving the accuracy and stability of color extraction.
Smart Images

Figure CN120808368A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a method and device for detecting media font materials. Background Art
[0002] In the field of text recognition, traditional optical character recognition (OCR) technology primarily focuses on identifying character shapes, typically using binarization to separate text areas from the background in an image. However, these methods are limited in extracting text color and are unable to cope with complex, colorful text scenes. Existing color extraction technologies, mostly based on binarization, masking, and pixel clustering, can achieve basic color recognition in simple, monochrome text and backgrounds, but their accuracy and generalization capabilities are poor for scenes with rich colors and similar foreground and background colors.
[0003] Existing technical solutions, such as the text color extraction method in the Rewire system, use a technology based on Delta-E color difference measurement. First, the color histogram of the text box edge pixels is calculated to identify the background color. Then, the foreground pixels are clustered and the weighted average is calculated to finally determine the foreground color (font color) of the text.
[0004] This method relies on pixel clustering and color difference calculation, and can work effectively in simple foreground and background scenes. However, when faced with colorful text or complex backgrounds, its method is prone to color recognition errors due to inaccurate clustering, especially when the foreground and background colors are close, the accuracy is poor.
[0005] Another example is a clustering method based on selective metrics, which extracts text color. By using multiple metrics in the RGB color space, similar colors are merged to achieve text segmentation. However, relying solely on color information is not sufficient to solve all problems in natural scenes. Therefore, this method also combines intensity and spatial information obtained through Log-Gabor filters to further segment characters. This method is designed to cope with the complex lighting changes, color variations, and background interference in natural scenes to improve the final character recognition rate.
[0006] This method extracts text color by clustering colors based on physical light reflectance. However, its core relies on clustering in the RGB color space, grouping and merging similar colors. While this method is somewhat effective when dealing with complex color variations, its adaptability is limited when dealing with colorful text and complex backgrounds. In particular, color segmentation accuracy is insufficient when the foreground and background colors are similar. Summary of the Invention
[0007] The application aims to provide a method and device for detecting media font materials, which have stronger generalization ability when dealing with complex multi-color text scenes and different background conditions.
[0008] To solve the above problems, the technical scheme of the application is: A method for detecting media font materials, which uses a deep learning model to extract background color, fill color and stroke color from a multi-color text image. The method includes: Converting the RGB image to be tested into LAB color space and generating corresponding color labels; the color labels include background color, fill color and stroke color; Feature extraction is performed on the color-converted image to obtain different levels of feature maps; Using a feature pyramid network to fuse different levels of feature maps to generate multi-scale feature maps; Using a channel attention mechanism to adjust the channel weight of the multi-scale feature maps to enhance the attention to different color channels; Using a self-attention mechanism to calculate the correlation between different pixels in the channel-enhanced feature map to generate new weighted features; Fully connecting or convolving the new weighted features to output LAB color prediction values of the background color, fill color and stroke color.
[0009] According to an embodiment of the application, the RGB image to be tested is converted into LAB color space to generate corresponding color labels, which further includes: Using the LABTransform class to resize the RGB image and convert it into LAB color space; at the same time, generating color labels including background color, fill color and stroke color, and converting the color labels into LAB color to ensure that the color space of the image and the label is consistent.
[0010] According to an embodiment of the application, the feature extraction of the color-converted image to obtain different levels of feature maps further includes: Using a ResNet34 neural network to extract features from the color-converted image to obtain different levels of feature maps including edges and textures, as well as shapes and semantics.
[0011] According to an embodiment of the application, the feature pyramid network is used to fuse different levels of feature maps to generate multi-scale feature maps, which further includes: Different levels of feature maps are processed through the lateral convs module, and different scale feature maps are fused layer by layer upwards to generate multi-scale features.
[0012] According to an embodiment of the present application, the channel attention mechanism is used to adjust the channel weights of the multi-scale feature maps, which further includes: Based on the channel attention mechanism, global average pooling is performed on each channel to generate a channel-level description; and two fully connected layers are used to adjust the weight of each channel to highlight the key color channels.
[0013] According to an embodiment of the present application, the self-attention mechanism is used to calculate the correlation between different pixels in the feature map after channel enhancement to generate new weighted features, which further includes: Based on the self-attention mechanism, the correlation weights between different pixels in the feature map after channel enhancement are calculated through query, key and value matrices, which are used to weight the feature map to generate adaptive feature representation, so as to enhance the understanding of global correlation in the image and facilitate color extraction in a multi-color background.
[0014] According to an embodiment of the present application, after obtaining the LAB color prediction value, the error between the predicted color and the true label is calculated by the ColorLoss loss function, and the AdamW optimizer is used to update the weights of the model according to the loss value to minimize the loss.
[0015] According to an embodiment of the present application, the deep learning model is trained, and the CosineAnnealingWarmRestarts scheduler is used to dynamically adjust the learning rate according to the training progress of each epoch to avoid falling into a local optimal solution. And through the early stopping mechanism, when the loss of the deep learning model on the validation set does not improve within the specified patience, the training is ended early to avoid overfitting.
[0016] According to an embodiment of the present application, the generalization performance of the deep learning model is evaluated by the validation set, and the LAB color values of the background color, the fill color and the stroke color are output according to the finally trained model; at the same time, an attention map is generated to show the area in the image that the deep learning model focuses on.
[0017] A device for detecting media font materials, which uses a deep learning model to extract background color, fill color and stroke color from a multi-color text image; the device comprises: A data processing module is configured to convert an RGB image to be tested into a LAB color space and generate a corresponding color label; the color label includes background color, fill color and stroke color. A feature extraction module is configured to extract features from the image after color conversion to obtain feature maps at different levels. An FPN module is configured to use a feature pyramid network to fuse feature maps at different levels to generate multi-scale feature maps. The SE Block module uses the channel attention mechanism to adjust the channel weights of multi-scale feature maps to enhance attention to different color channels; The Self-Attention module uses the self-attention mechanism to calculate the correlation between different pixels in the feature map after channel enhancement and generate new weighted features; The color prediction module is used to fully connect or convolve the new weighted features and output the LAB color prediction values of the background color, fill color, and stroke color.
[0018] Due to the adoption of the above technical solution, the present invention has the following advantages and positive effects compared with the prior art: The method for detecting media fonts in one embodiment of the present invention combines a feature pyramid network (FPN), a channel-wise attention mechanism (SE Block), and a self-attention mechanism (Self-Attention), significantly enhancing the model's generalization and adaptability in a variety of complex scenarios. The FPN provides multi-scale feature fusion, enabling the model to process text color information at varying resolutions. The SE Block enhances the model's perception of key color channels by adjusting the weights of specific channels. Self-Attention further captures global color correlations in the image, ensuring the model can consistently extract color information despite varying color distributions and complex backgrounds. This multi-module collaboration effectively improves the model's overall performance, enabling it to handle a wide range of text and background color combinations and accurately identify background, fill, and stroke colors in a variety of complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 1 is a flow chart of a method for detecting media font materials in one embodiment of the present invention; Figure 2 A schematic diagram of a portion of the principles of a deep learning model in one embodiment of the present invention; Figure 3 Schematic diagram of another part of the principle of the deep learning model in one embodiment of the present invention. DETAILED DESCRIPTION
[0020] The following is a detailed description of a method and device for detecting media font materials proposed by the present invention in conjunction with the accompanying drawings and specific embodiments. The advantages and features of the present invention will become more apparent from the following description and claims.
[0021] With the rapid growth of digital content, especially the wide application of image processing and computer vision technology, there is an urgent need for automated text color extraction technology in many fields. Traditional color extraction methods, such as techniques based on RGB color space, are difficult to accurately distinguish similar colors or color gradient text and background, and are easily affected by color confusion. Methods based on color clustering or selective metric clustering lack generalization ability in complex multi-color text and background scenes, especially when the foreground and background colors are close, and are prone to recognition errors.
[0022] To solve the above problems, the embodiment provides a method for detecting media font materials, which can accurately extract color information (background color, fill color and stroke color) of the text region. This provides technical support for many industries, especially in scenarios that require complex image processing.
[0023] Please refer to Figure 1 The method for detecting media font materials uses a deep learning model to extract background color, fill color and stroke color from multi-color text images. The method includes the following steps: Convert the RGB image to be tested to LAB color space to generate corresponding color labels; the color labels include background color, fill color and stroke color; Feature extraction is performed on the color-converted image to obtain feature maps at different levels; Using a feature pyramid network, the feature maps at different levels are fused to generate multi-scale feature maps; Using a channel attention mechanism, the channel weight of the multi-scale feature map is adjusted to enhance the attention to different color channels; Using a self-attention mechanism, the correlation between different pixels in the channel-enhanced feature map is calculated to generate new weighted features; The new weighted features are fully connected or convolved to output the LAB color prediction values of the background color, fill color and stroke color.
[0024] This method combines a feature pyramid network (FPN), a channel-wise attention mechanism (SE Block), and a self-attention mechanism (Self-Attention), significantly enhancing the model's generalization and adaptability in a variety of complex scenarios. The FPN provides multi-scale feature fusion, enabling the model to process text color information at varying resolutions. The SE Block enhances the model's perception of key color channels by adjusting the weights of specific channels. Self-Attention further captures global color correlations in the image, ensuring the model can consistently extract color information despite varying color distributions and complex backgrounds. This multi-module collaboration effectively improves the model's overall performance, enabling it to handle a wide range of text and background color combinations, ensuring accurate recognition of background, fill, and stroke colors in a variety of complex scenarios.
[0025] For details, please see Figure 2 and Figure 3 The deep learning model includes input layer, feature extraction layer, FPN layer, SEBlock layer, Self-Attention layer and output layer. Due to space constraints, the figure only shows depth 1. In actual applications, there may be multiple depths.
[0026] This example uses a deep learning model to extract background, fill, and stroke colors from multicolored text images. At the input layer, the RGB image to be tested must be converted to the LAB color space to generate corresponding color labels. The RGB image to be tested can be a captured image or the output of a text box after a text detection phase using tools like PaddleOCR or YOLO.
[0027] The color distribution in the RGB color space performs poorly when dealing with lighting changes and color gradients, while the brightness (L), red-green axis (a), and yellow-blue axis (b) components in the LAB color space are more consistent with human visual perception. This application converts images from RGB to the LAB color space, enabling the model to more effectively capture color differences, thereby improving color extraction accuracy.
[0028] Specifically, the LABTransform class is used to resize the RGB image (for example, to 224×224) and convert it to the LAB color space for subsequent neural network processing; at the same time, color labels including background color, fill color, and stroke color are generated, and the LABTargetTransform class is used to convert the color labels to LAB colors to ensure that the color space of the image and the label are consistent.
[0029] The present embodiment effectively reduces the confusion between background color, fill color and stroke color using the LAB color space, and the model can maintain high color extraction accuracy even in scenes with similar or gradient colors. Moreover, by training the network in the LAB space, the model can maintain high accuracy when processing color similar text and background, enhancing the accuracy and stability of the color prediction task.
[0030] In the feature extraction layer, feature extraction needs to be performed on the color-converted image to obtain feature maps of different levels. The present embodiment uses a ResNet34 neural network to perform feature extraction on the color-converted image to obtain feature maps of different levels. The neural network includes multiple convolutional layers and pooling layers, which gradually extract feature maps of different scales from the input image, including low-level (edge and texture) and high-level (shape and semantic) features. The ResNet34 neural network is a relatively mature network, and its structure is not described here.
[0031] In the FPN layer, a feature pyramid network is used to fuse feature maps of different levels to generate multi-scale feature maps. Specifically, the feature maps extracted from the ResNet34 are processed by lateral convs, and different scale feature maps are fused layer by layer to generate comprehensive multi-scale feature representations. This feature fusion method enables the model to obtain detailed information of the image from different resolutions, especially in the color extraction task to help capture the details and structure of the image.
[0032] Feature Pyramid Network (FPN) is a deep learning architecture for solving multi-scale object detection problems, which enhances the robustness of the model to scale changes by fusing feature maps of different levels. Traditional methods rely on image pyramids (multi-scale input) which are computationally intensive and memory-intensive, while deep convolutional networks naturally have a pyramid feature hierarchy, but high-level features have strong semantics and low resolution, while low-level features are rich in details but weak in semantics. FPN fuses features of different levels to build a feature pyramid with high semantics and high resolution. It includes three core paths: Bottom-up Path: Feedforward computation of backbone networks (such as ResNet) generates multi-scale feature maps, with deeper levels having lower resolution but stronger semantics.
[0033] Top-down Path: Upsampling (such as bilinear interpolation) of high-level features gradually restores high resolution.
[0034] Lateral Connections: Fuse the upsampled feature with the same scale bottom layer feature (usually through 1x1 convolution to adjust the channel number), combine the high-level semantic and bottom-level detail.
[0035] The fused feature map has the same number of channels, and each layer can be used independently for target detection or segmentation tasks.
[0036] In the SE Block layer, the channel weight of the feature map after FPN processing is adjusted to enhance the model's attention to important color channels. Specifically, based on the channel attention mechanism, the global average pooling is performed on each channel to generate a channel-level description, and then two fully connected layers are used to adjust the weight of each channel. The feature map after the SE Block can highlight the key color channels, thereby improving the model's attention to color and performance in the color extraction task.
[0037] Channel Attention Mechanism is a core component in deep learning target detection, which dynamically adjusts the weight of feature channels to enhance key features and suppress redundant information, thereby improving model performance.
[0038] The channels of the feature map represent different semantic information, but traditional convolution treats all channels equally, resulting in key features being submerged. Channel attention mechanism learns channel weights to achieve adaptive enhancement of feature channels. Its essential principle is as follows: where u c is the c-th channel of the input feature map, and s c is the learned channel weight.
[0039] In the Self-Attention layer, the self-attention mechanism is used to calculate the correlation between different pixels in the channel-enhanced feature map to generate new weighted features. Specifically, based on the self-attention mechanism, the correlation weights (attention distribution) between different pixels in the channel-enhanced feature map are calculated through the query, key, and value matrices, and then these weights are used to weight the input feature map to generate a more adaptive feature representation. This process enhances the model's understanding of global correlations in the image (such as color changes and complex backgrounds), which helps to handle color extraction tasks in multi-color backgrounds.
[0040] The self-attention mechanism captures global dependency between features through dynamic weight distribution, and its calculation process can be decomposed as follows: where Q, K, V are Query, Key, Value vectors generated by linear transformation of input features; d k is a scaling factor to prevent gradient vanishing caused by large dot product; Z i is the weighted combination of input features, and the weight a ij represents the correlation between features i and j.
[0041] At the output layer, the feature maps processed by FPN, SE Block, and Self-Attention are passed into a fully connected layer or a convolutional layer, and each channel outputs the LAB color values of background color, fill color, and stroke color, respectively. The final output is a tensor of shape (3, 3), representing the LAB prediction values of these three colors.
[0042] Before being put into application, the above deep learning model needs to be trained. After obtaining the LAB color prediction value each time, the error between the predicted color and the true label is calculated through the ColorLoss loss function, and the AdamW optimizer is used to update the weights of the model according to the loss value to minimize the loss.
[0043] The ColorLoss loss function is a key component in computer vision tasks for optimizing the color reconstruction quality of images. Its core idea is to quantify the difference between the generated image and the real image in the color space, driving the model to learn more accurate color distribution. Its function expression is as follows: where c pred is the color prediction value, and c gt is the true label value.
[0044] The AdamW optimizer is a widely used adaptive optimization algorithm in the field of deep learning. As an improved version of Adam, its core innovation lies in decoupling weight decay and gradient update, significantly improving the model's generalization ability and training stability.
[0045] AdamW applies weight decay as an independent term directly to parameter updates, avoiding the distortion of adaptive learning rate on regularization strength, allowing weight decay to truly control model complexity. AdamW can be considered as an approximation to the Proximal Gradient, where weight decay corresponds to a closed-form proximal mapping (such as the soft threshold function of L2 regularization). Under the condition that the random gradient satisfies the bounded variance condition, the convergence rate of AdamW is: where d is the dimension of the parameter, K is the number of iterations, and the optimal convergence rate of SGD is matched.
[0046] The deep learning model is trained, and the CosineAnnealingWarmRestarts scheduler is used to dynamically adjust the learning rate according to the training progress of each epoch to avoid falling into a local optimal solution. And through the early stopping mechanism, when the loss of the deep learning model on the validation set does not improve within the specified patience, the training is ended in advance to avoid overfitting.
[0047] The CosineAnnealingWarmRestarts scheduler periodically resets the learning rate to make the model jump out of the local minimum in training, improving the generalization ability and convergence efficiency. Its expression is: Where η max is the initial learning rate of the optimizer, η min is the scheduler parameter, T i is the length of the i-th cycle, and T cur is the step number in the current cycle (from 0).
[0048] At the end of each epoch, the validation is performed, the generalization performance of the deep learning model is evaluated through the validation set, and the LAB color values of the background color, fill color and stroke color are output according to the finally trained model; at the same time, the attention map is generated to show the area in the image that the deep learning model focuses on.
[0049] The above method can still maintain high efficient color extraction ability in diversified color backgrounds and complex scenes through the design of network structure (FPN, SE Block and Self-Attention). Compared with traditional color extraction methods, the introduction of this innovative network structure enables the present technology to adapt to different market demands, especially in complex text and image environments.
[0050] Based on the same idea, the present embodiment also provides a device for detecting media font materials, which extracts background color, fill color and stroke color from multi-color text images by using a deep learning model. The device comprises: A data processing module is used to convert the RGB image to be tested into LAB color space and generate corresponding color labels; the color labels include background color, fill color and stroke color; A feature extraction module is used to extract features from the color-converted image to obtain feature maps at different levels; An FPN module is used to adopt a feature pyramid network to fuse feature maps at different levels to generate multi-scale feature maps; The SE Block module is configured to adjust the channel weights of the multi-scale feature maps by using a channel attention mechanism, so as to enhance the attention to different color channels. The Self-Attention module is configured to calculate the correlation between different pixels in the feature maps after the channel enhancement by using a self-attention mechanism, and generate new weighted features. The color prediction module is configured to perform full connection or convolution on the new weighted features, and output LAB color prediction values of the background color, the fill color and the stroke color.
[0051] The device is used to implement the method for detecting media font materials, and the embodiments are similar, which will not be described here again.
[0052] The device can be easily integrated into existing OCR, image processing and video analysis systems as a complementary module of the existing systems, helping enterprises to improve the functions and effects of the existing technologies. For example, it can improve the brand identification accuracy in advertisements and video content, and can also enhance the automated analysis capability of media content.
[0053] The embodiments of the present application are described in detail above in combination with the drawings, but the present application is not limited to the above-described embodiments. Even if various changes are made to the present application, if the changes belong to the scope of the claims of the present application and equivalent technologies, they still fall within the protection scope of the present application.
Claims
1. A method for detecting media font materials, using a deep learning model to extract background color, fill color, and stroke color from a colorful text image; characterized in that: include: Convert the RGB image to be tested into LAB color space and generate corresponding color labels; the color labels include background color, fill color and stroke color; Perform feature extraction on the color-converted image to obtain feature maps at different levels; Using feature pyramid network, feature maps at different levels are fused to generate multi-scale feature maps; Adopting the channel attention mechanism, the channel weights of the multi-scale feature maps are adjusted to enhance the attention to different color channels; Using the self-attention mechanism, the correlation between different pixels in the feature map after channel enhancement is calculated to generate new weighted features; The new weighted features are fully connected or convolved to output the LAB color prediction values of the background color, fill color, and stroke color.
2. The method for detecting media font material according to claim 1, wherein: The converting of the RGB image to be measured into the LAB color space and generating the corresponding color label further comprises: Use the LABTransform class to resize the RGB image and convert it to the LAB color space. At the same time, generate color labels including background color, fill color, and stroke color, and convert the color labels to LAB colors to ensure that the color space of the image and the label are consistent.
3. The method for detecting media font material according to claim 1, wherein: The feature extraction of the color-converted image to obtain feature maps at different levels further includes: The ResNet34 neural network is used to extract features from the color-converted image to obtain feature maps at different levels, including edges and textures, shapes, and semantics.
4. The method for detecting media font material according to claim 1, wherein: The method of using a feature pyramid network to fuse feature maps at different levels to generate a multi-scale feature map further includes: The feature maps of different levels are processed by the lateral convs module, and the feature maps of different scales are fused layer by layer to generate multi-scale features.
5. The method for detecting media font material according to claim 1, wherein: The channel attention mechanism is used to adjust the channel weights of the multi-scale feature map further comprising: Based on the channel attention mechanism, global average pooling is performed on each channel to generate channel-level description; the weight of each channel is adjusted through two fully connected layers to highlight the key color channel.
6. The method for detecting media font material according to claim 1, wherein: The self-attention mechanism is used to calculate the correlation between different pixels in the feature map after channel enhancement to generate new weighted features, which further includes: Based on the self-attention mechanism, the correlation weights between different pixels in the channel-enhanced feature map are calculated through the query, key and value matrices. These weighted feature maps are used to generate adaptive feature representations to enhance the understanding of global correlation in the image and facilitate color extraction in colorful backgrounds.
7. The method for detecting media font material according to claim 1, wherein: After obtaining the LAB color prediction value, the error between the predicted color and the true label is calculated through the ColorLoss loss function, and the AdamW optimizer is used to update the model weight according to the loss value to minimize the loss.
8. The method for detecting media font material according to claim 1, wherein: For deep learning model training, the CosineAnnealingWarmRestarts scheduler is used to dynamically adjust the learning rate based on the training progress of each epoch to avoid falling into local optimal solutions. And through the early stopping mechanism, when the loss of the deep learning model on the validation set does not improve within the specified patience, the training ends early to avoid overfitting.
9. The method for detecting media font material according to claim 1, wherein: The generalization performance of the deep learning model is evaluated through the validation set. Based on the final trained model, the LAB color values of the background color, fill color, and stroke color are output. At the same time, an attention map is generated to show the areas in the image that the deep learning model focuses on.
10. A device for detecting media font material, using a deep learning model to extract background color, fill color, and stroke color from a colorful text image; characterized in that: include: A data processing module is used to convert the RGB image to be tested into the LAB color space and generate corresponding color labels; the color labels include background color, fill color and stroke color; Feature extraction module, used to extract features from the color-converted image and obtain feature maps at different levels; The FPN module is used to fuse feature maps at different levels using a feature pyramid network to generate multi-scale feature maps; The SE Block module uses the channel attention mechanism to adjust the channel weights of multi-scale feature maps to enhance attention to different color channels; The Self-Attention module uses the self-attention mechanism to calculate the correlation between different pixels in the feature map after channel enhancement and generate new weighted features; The color prediction module is used to fully connect or convolve the new weighted features and output the LAB color prediction values of the background color, fill color, and stroke color.