Neural Network-Based Text Matting Method, Device, Equipment, and Storage Medium
Through the text cutout method based on neural network, the problem of loss of text style and artistic font features in text recognition technology is solved, and the accurate extraction and retention of text features is achieved, and the scope of application is expanded.
Patent Information
- Application Number
- CN202280005241.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-23
AI Technical Summary
Existing text recognition technologies are difficult to preserve the text style and artistic font features in images, resulting in the loss of personalized features.
The text cutout method based on neural network is used to process the image through feature extraction network, intermediate processing network and feature fusion network, and extract and retain the font and color features of the text.
It realizes the personalized features of the text during text cutout, expands the application areas and scope of text features, and improves the accuracy of text feature extraction.
Smart Images

Figure CN118541736B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and more particularly, to a text matting method, apparatus, device, and storage medium based on a neural network. Background Art
[0002] Text recognition technology can detect and recognize text information in images, such as applications for determining text positions, recognizing the meanings of words, etc. However, pure text recognition is difficult to retain personalized features such as artistic fonts and text styles of specific unique designs in images such as posters and advertisements.
[0003] With the development of artificial intelligence technology, it has been widely applied to the field of image processing. Image processing generally can include processing tasks such as image retouching and color adjustment, image beautification, image denoising, image super-resolution conversion, and feature extraction. Summary of the Invention
[0004] Some embodiments of the present disclosure provide a text matting method, apparatus, device, and storage medium based on a neural network, which use the neural network to perform text matting on the text included in the image to obtain an output image including complete text features, where personalized features such as the font and color of the text are retained.
[0005] According to one aspect of the present disclosure, there is provided a neural network-based text matting method, including: processing a first image by using a feature extraction network to obtain a feature map, where the feature extraction network includes n extraction convolutional network modules connected in sequence. The first extraction convolutional network module processes the first image and outputs a first feature map. The i-th extraction convolutional network module processes the (i - 1)-th feature map and outputs an i-th feature map; processing the feature map by using an intermediate processing network to obtain an intermediate feature map, where the intermediate processing network includes n spatial convolutional network modules. The first spatial convolutional network module processes the first feature map and outputs a first intermediate feature map. The i-th spatial convolutional network module processes the i-th feature map and outputs an i-th intermediate feature map; and processing the intermediate feature map by using a feature fusion network to obtain a second image, where the feature fusion network includes n fusion convolutional network modules connected in sequence. The first fusion convolutional network module processes the n-th intermediate feature map output by the n-th spatial convolutional network module to obtain a first fusion map. The i-th fusion convolutional network module processes the (n - i + 1)-th intermediate feature map output by the (n - i + 1)-th spatial convolutional network module and the (i - 1)-th fusion map output by the (i - 1)-th fusion convolutional network module to obtain an i-th fusion map. The n-th fusion map output by the n-th fusion convolutional network module is used as the second image, and the second image includes text features extracted from the first image, where n is an integer greater than 1, and i is an integer greater than 1 and less than or equal to n.
[0006] According to some embodiments of the present disclosure, the n extraction convolutional network modules in the feature extraction network are respectively composed of one or more of a convolutional layer, a pooling layer, and a residual convolutional module; the n fusion convolutional network modules in the feature fusion network are respectively composed of one or more of a convolutional layer, a residual convolutional module, and a first connection layer, where the first connection layer includes a feature merging layer and a super-resolution layer.
[0007] According to some embodiments of the present disclosure, n is equal to 4. In the feature extraction network, the first extraction convolutional network module is composed of a convolutional layer and a max pooling layer. The second extraction convolutional network module is composed of a residual convolutional module, a convolutional layer, and a max pooling layer. The third extraction convolutional network module is composed of a residual convolutional module, a convolutional layer, and a max pooling layer. The fourth extraction convolutional network module is composed of a residual convolutional module and a convolutional layer. In the feature fusion network, the first fusion convolutional network module is composed of a convolutional layer and a residual convolutional module. The second fusion convolutional network module is composed of a first connection layer, a convolutional layer, and a residual convolutional module. The third fusion convolutional network module is composed of a first connection layer, a convolutional layer, and a residual convolutional module. The fourth fusion convolutional network module is composed of a first connection layer and a convolutional layer, where the first connection layer includes a feature merging layer and a super-resolution layer.
[0008] According to some embodiments of the present disclosure, the feature merging layer is used to merge a plurality of input images by increasing the number of image channels, and the super-resolution layer is used to increase the image size by reducing the number of image channels; wherein, the residual convolution module is composed of a convolutional layer, a batch normalization layer, and an activation function.
[0009] According to some embodiments of the present disclosure, the structures of n spatial convolution network modules are the same. Each spatial convolution network module includes m parallel processing units and a second connection layer. The m processing units are used to process the input image respectively, and the second connection layer is used to merge the processing results of the m processing units. Among them, the second connection layer includes a feature merging layer, a convolutional layer, and a batch normalization layer. The feature merging layer is used to merge the m images respectively output by the m processing units by increasing the number of image channels, where m is an integer greater than 1.
[0010] According to some embodiments of the present disclosure, m is equal to 5. In each spatial convolution network module, the first processing unit is composed of a first convolutional layer and a batch normalization layer connected in sequence, the second processing unit is composed of a second convolutional layer, a batch normalization layer, a third convolutional layer, and a batch normalization layer connected in sequence, the third processing unit is composed of a fourth convolutional layer, a batch normalization layer, a third convolutional layer, and a batch normalization layer connected in sequence, the fourth processing unit is composed of a fifth convolutional layer, a batch normalization layer, a third convolutional layer, and a batch normalization layer connected in sequence, and the fifth processing unit is composed of a batch normalization layer, an adaptive average pooling layer, a third convolutional layer, and an upsampling layer connected in sequence. Among them, the network parameters of the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer are different.
[0011] According to some embodiments of the present disclosure, the pixel points of the nth fusion graph are represented as text probability values in the range of 0-1 and quantized into grayscale images in the range of 0-255.
[0012] According to some embodiments of the present disclosure, the text matting method further includes: generating a training image set and annotating the training images in the training image set to generate corresponding matte images; and training the feature extraction network, the intermediate processing network, and the feature fusion network by using the training image set and based on a training function.
[0013] According to some embodiments of the present disclosure, generating a training image set includes: constructing a corpus composed of text characters, a font library composed of text fonts, a texture feature library composed of text colors, and a background library composed of background images; and randomly extracting a set of text characters, text fonts, text colors, and background images from the corpus, font library, texture feature library, and background library, and forming a training image, wherein the text characters include Chinese characters and English characters, the text fonts include Chinese fonts and English fonts, and the text colors include solid colors, textures, and patterns.
[0014] According to some embodiments of the present disclosure, the text matting method further includes: performing text detection on the original image to extract a detection image including text from the original image, wherein the detection image serves as the first image.
[0015] According to some embodiments of the present disclosure, the text matting method further includes: performing image processing on the text features in the second image to generate a third image after text style conversion; and combining the third image with the target background image to generate a fourth image after image style conversion.
[0016] According to some embodiments of the present disclosure, the text matting method further includes: performing image processing on the second image to generate a mask image having the same image size as the original image; and using the mask image to perform image completion on the original image to generate the original image after eliminating text features.
[0017] According to another aspect of the present disclosure, there is also provided a text matting device based on a neural network. The text matting device includes: a feature extraction unit configured to process a first image using a feature extraction network to obtain a feature map, where the feature extraction network includes n extraction convolutional network modules connected in sequence. Among them, the first extraction convolutional network module processes the first image and outputs the first feature map, and the i-th extraction convolutional network module processes the (i - 1)-th feature map and outputs the i-th feature map; an intermediate processing unit configured to process the feature map using an intermediate processing network to obtain an intermediate feature map, where the intermediate processing network includes n spatial convolutional network modules. Among them, the first spatial convolutional network module processes the first feature map and outputs the first intermediate feature map, and the i-th spatial convolutional network module processes the i-th feature map and outputs the i-th intermediate feature map; and a feature fusion unit configured to process the intermediate feature map using a feature fusion network to obtain a second image, where the feature fusion network includes n fusion convolutional network modules connected in sequence. The first fusion convolutional network module processes the n-th intermediate feature map output by the n-th spatial convolutional network module to obtain the first fusion map, and the i-th fusion convolutional network module processes the (n - i + 1)-th intermediate feature map output by the (n - i + 1)-th spatial convolutional network module and the (i - 1)-th fusion map output by the (i - 1)-th fusion convolutional network module to obtain the i-th fusion map. Among them, the n-th fusion map output by the n-th fusion convolutional network module is used as the second image, and the second image includes text features extracted from the first image, where n is an integer greater than 1, and i is an integer greater than 1 and less than or equal to n.
[0018] According to yet another aspect of the present disclosure, there is also provided an image processing device, including: a processor; and a memory, where computer-readable code is stored in the memory, and when the computer-readable code is run by the processor, it executes the text matting method as described above.
[0019] According to yet another aspect of the present disclosure, there is also provided a computer-readable storage medium having instructions stored thereon, and when the instructions are executed by a processor, the processor is caused to execute the text matting method as described above.
[0020] Using a neural network-based text matting method, apparatus, device, and storage medium according to some embodiments of the present disclosure, it is possible to perform text matting on a first image including text information by using a neural network composed of a feature extraction network, an intermediate processing network, and a feature fusion network, so as to extract text features in the first image and obtain a second image including the text features. The personalized information such as the font and color of the text is retained in the output second image, which is beneficial to expanding the application fields and scope of the text features. In addition, the intermediate processing network can serve as a connection layer between the feature extraction network and the feature fusion network, extract multi-scale intermediate features, and improve the accuracy of text feature extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 FIG. shows a schematic flowchart of a neural network-based text matting method according to an embodiment of the present disclosure;
[0023] Figure 2 FIG. shows a schematic structural diagram of a neural network for implementing text matting according to an embodiment of the present disclosure;
[0024] Figure 3 FIG. shows an overall structural diagram of a neural network for implementing text matting according to an embodiment of the present disclosure;
[0025] Figure 4 FIG. shows a network structural diagram of a residual convolution module according to an embodiment of the present disclosure;
[0026] Figure 5 FIG. shows a network structural diagram of a spatial convolution network module according to an embodiment of the present disclosure;
[0027] Figure 6 FIG. shows a schematic diagram of data elements for making a data set according to an embodiment of the present disclosure;
[0028] Figure 7 FIG. shows a schematic flowchart of making a data set according to an embodiment of the present disclosure;
[0029] Figure 8 FIG. shows an application flowchart of using the text matting method according to an embodiment of the present disclosure;
[0030] Figure 9Shows another application flowchart of the text matting method according to an embodiment of the present disclosure;
[0031] Figure 10 Shows a schematic block diagram of a neural network-based text matting device according to an embodiment of the present disclosure;
[0032] Figure 11 Shows a schematic block diagram of an image processing device according to an embodiment of the present disclosure;
[0033] Figure 12 Shows a schematic diagram of the architecture of an exemplary computing device according to an embodiment of the present disclosure;
[0034] Figure 13 Shows a schematic diagram of a computer storage medium according to an embodiment of the present disclosure. Detailed implementation manners
[0035] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0036] The "first", "second" and similar terms used in the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "include" or "comprise" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. "Connection" or "coupling" and similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.
[0037] Flowcharts are used in the present disclosure to illustrate the steps of the methods according to the embodiments of the present disclosure. It should be understood that the steps before or after do not necessarily need to be performed precisely in order. On the contrary, they can be performed in reverse order or various steps can be processed simultaneously. At the same time, other operations can also be added to these processes.
[0038] It can be understood that the professional terms and nouns involved herein have the meanings well-known to those skilled in the art.
[0039] The text recognition technology can detect and recognize the text information in the image, such as the position, content, etc. of the text. For example, the extracted text can be used to implement applications such as semantic recognition. However, the extraction of text content cannot retain the personalized features such as the artistic fonts and text styles of the specific unique designs in images such as posters and advertisements.
[0040] With the development of artificial intelligence, it has been widely applied in the field of image processing. Artificial Intelligence (AI) technology uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0041] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies, mainly including several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. By training a neural network based on training samples, image processing tasks such as image feature extraction and segmentation can be achieved, for example.
[0042] Some embodiments of the present disclosure provide a text matting method, apparatus, device, and storage medium based on a neural network. By constructing a neural network architecture capable of implementing the text matting task, an input image including text is processed, and an output image with text features is output to achieve text matting. The output image includes complete text features and retains personalized information such as the font and color of the text, so that it can be flexibly applied to application fields such as creative font extraction and subtitle elimination.
[0043] Specifically, in the text matting method, apparatus, device, and storage medium based on a neural network provided by some embodiments of the present disclosure, a neural network composed of a feature extraction network, an intermediate processing network, and a feature fusion network is used to perform text matting on a first image including text information to extract the text features in the input image, and an output image including the text features is obtained. The output image retains personalized information such as the font and color of the text, which is beneficial to expanding the application fields and scope of the text features. In addition, the intermediate processing network can serve as a connection layer between the feature extraction network and the feature fusion network to extract multi-scale intermediate features, thereby improving the accuracy of text feature extraction.
[0044] Figure 1 Shows a schematic flowchart of a text matting method based on a neural network according to an embodiment of the present disclosure. Next, the implementation process of the text matting method according to an embodiment of the present disclosure will be described first in combination with Figure 1 The overall implementation process of the text matting method according to an embodiment of the present disclosure will be described, and then the structure of the neural network for implementing the text matting method will be described in detail.
[0045] First, as Figure 1 shown, in step S101, a feature extraction network is used to process the first image to obtain a feature map. Among them, the feature extraction network includes n extraction convolutional network modules connected in sequence. Among them, the first extraction convolutional network module processes the first image and outputs the first feature map. The i-th extraction convolutional network module processes the (i - 1)-th feature map and outputs the i-th feature map, where n is an integer greater than 1, and i is an integer greater than 1 and less than or equal to n.
[0046] According to some embodiments of the present disclosure, the network structure design and specific parameters of each extraction convolutional network module can be the same or different, which can be set according to factors such as specific application scenarios and the size of the first image. In addition, the value of n can also be set according to factors such as specific application scenarios and the size of the first image. It can be understood that although some parameters or values are described for specific examples or application scenarios in the following description of this article, the implementation manner according to the embodiments of the present disclosure is not limited to this.
[0047] According to some embodiments of the present disclosure, the feature extraction network is used to extract features from the input first image. For example, the n extraction convolutional network modules connected in sequence included therein are used to extract feature information of different scales. According to some embodiments of the present disclosure, the n extraction convolutional network modules in the feature extraction network can be respectively composed of one or more of a convolutional layer, a pooling layer, and a residual convolutional module. The specific network implementation manner will be described in combination with Figure 3 the following. Generally speaking, in the feature extraction network, the convolutional layer is used for feature extraction, and the pooling layer is used to reduce the spatial size of the image features extracted by the convolutional layer and provide a larger receptive field. According to some embodiments of the present disclosure, the residual convolutional module can be composed of a convolutional layer, a batch normalization layer, and an activation function. The structure of the residual convolutional module will be described in combination with Figure 4 the following.
[0048] Next, in step S102, an intermediate processing network is used to process the feature map to obtain an intermediate feature map. Among them, the intermediate processing network includes n spatial convolutional network modules. Among them, the first spatial convolutional network module processes the first feature map and outputs the first intermediate feature map. The i-th spatial convolutional network module processes the i-th feature map and outputs the i-th intermediate feature map. Similarly, according to some embodiments of the present disclosure, the network structure design and specific parameters of each spatial convolutional network module can be the same or different, which can be set according to factors such as specific application scenarios and the size of the first image. In some implementation manners, the network structures and network parameters of each spatial convolutional network module are the same.
[0049] According to some embodiments of the present disclosure, the intermediate processing network serves as a connection layer between the feature extraction network and the feature fusion network, and is used to further extract multi-scale intermediate features, and output the intermediate feature maps at different scales to the corresponding modules of the feature fusion network, so as to improve the accuracy and integrity of the text features in the output image. For example, the n spatial convolution network modules included in the intermediate processing network are used to process the feature maps with different sizes output by the n extraction convolution network modules in the feature extraction network respectively, so as to further obtain intermediate feature maps.
[0050] Next, referring to Figure 1 , in step S103, the feature fusion network is used to process the intermediate feature map to obtain a second image. The feature fusion network includes n fusion convolution network modules connected in sequence. The first fusion convolution network module processes the nth intermediate feature map output by the nth spatial convolution network module to obtain the first fusion map. The ith fusion convolution network module processes the (n - i + 1)th intermediate feature map output by the (n - i + 1)th spatial convolution network module and the (i - 1)th fusion map output by the (i - 1)th fusion convolution network module to obtain the ith fusion map. According to some embodiments of the present disclosure, the nth fusion map output by the nth fusion convolution network module serves as the second image, and the second image includes the text features extracted from the first image. Similarly, according to some embodiments of the present disclosure, the network structure design and specific parameters of each fusion convolution network module may be the same or different, which can be set according to factors such as specific application scenarios and the size of the first image. The specific network implementation manner will be described in detail below in conjunction with Figure 3 Expand the description.
[0051] According to some embodiments of the present disclosure, the feature fusion network is used to perform feature fusion on the intermediate feature map output by the intermediate processing network to output a final output image, that is, the second image, and the image size of the second image may be the same as that of the first image. Specifically, the second image includes the text features extracted from the first image, and the text features include not only text information but also personalized design features such as text fonts.
[0052] According to some embodiments of the present disclosure, the n fusion convolutional network modules in the feature fusion network may be respectively composed of one or more of a convolutional layer, a residual convolutional module, and a first connection layer, wherein the first connection layer includes a feature merging layer and a super-resolution layer. According to some embodiments of the present disclosure, the feature merging layer is used to merge multiple input images by increasing the number of image channels, and the super-resolution layer is used to increase the image size by reducing the number of image channels. According to some embodiments of the present disclosure, the residual convolutional module may be composed of a convolutional layer, a batch normalization layer, and an activation function. In the feature fusion network, the first connection layer may be used to perform feature connection on the intermediate feature map output by the spatial convolutional network module and the fusion map output by the fusion convolutional network module at the previous level for joint processing.
[0053] As an example, when n = 4, that is, the feature extraction network includes 4 extraction convolutional network modules connected in sequence, wherein the first extraction convolutional network module processes the first image (i.e., the image input to the feature extraction network) and outputs the first feature map, the second extraction convolutional network module processes the first feature map and outputs the second feature map, the third extraction convolutional network module processes the second feature map and outputs the third feature map, and the fourth extraction convolutional network module processes the third feature map and outputs the fourth feature map.
[0054] As an example, when n = 4, that is, the intermediate processing network includes 4 spatial convolutional network modules, wherein the first spatial convolutional network module is used to process the first feature map output by the first extraction convolutional network module and output the first intermediate feature map, the second spatial convolutional network module is used to process the second feature map output by the second extraction convolutional network module and output the second intermediate feature map, the third spatial convolutional network module is used to process the third feature map output by the third extraction convolutional network module and output the third intermediate feature map, and the fourth spatial convolutional network module is used to process the fourth feature map output by the fourth extraction convolutional network module and output the fourth intermediate feature map. Regarding the network implementation of the intermediate processing network and the spatial convolutional network modules therein, it will be described below in conjunction with Figure 5 will be described.
[0055] As an example, when n = 4, that is, the feature fusion network includes 4 fusion convolutional network modules connected in sequence. Among them, the first fusion convolutional network module processes the fourth intermediate feature map output by the fourth spatial convolutional network module to obtain the first fusion map; the second fusion convolutional network module processes the third intermediate feature map output by the third spatial convolutional network module and the first fusion map output by the first fusion convolutional network module to obtain the second fusion map; the third fusion convolutional network module processes the second intermediate feature map output by the second spatial convolutional network module and the second fusion map output by the second fusion convolutional network module to obtain the third fusion map; and the fourth fusion convolutional network module processes the first intermediate feature map output by the first spatial convolutional network module and the third fusion map output by the third fusion convolutional network module to obtain the fourth fusion map.
[0056] Specifically, Figure 2 FIG. shows a schematic structural diagram of 4 extraction convolutional network modules and 4 fusion convolutional network modules in a specific example where n = 4. It can be understood that the specific structures of the feature extraction network and the feature fusion network according to the embodiments of the present disclosure are not limited thereto.
[0057] As Figure 2 shown, in the feature extraction network, the first extraction convolutional network module consists of a convolutional layer and a max pooling layer, the second extraction convolutional network module consists of a residual convolutional module, a convolutional layer and a max pooling layer, the third extraction convolutional network module consists of a residual convolutional module, a convolutional layer and a max pooling layer, and the fourth extraction convolutional network module consists of a residual convolutional module and a convolutional layer. For example, the residual convolutional module therein can consist of a convolutional layer, a batch normalization layer and an activation function.
[0058] Next, as Figure 2 shown, in the feature fusion network, the first fusion convolutional network module consists of a convolutional layer and a residual convolutional module, the second fusion convolutional network module consists of a first connection layer, a convolutional layer and a residual convolutional module, the third fusion convolutional network module consists of a first connection layer, a convolutional layer and a residual convolutional module, and the fourth fusion convolutional network module consists of a first connection layer and a convolutional layer. Among them, the first connection layer can include, for example, a feature merging layer and a super-resolution layer. Similarly, the feature merging layer is used to merge multiple input images (for example, 2) by increasing the number of image channels, the super-resolution layer is used to increase the image size by reducing the number of image channels, and the residual convolutional module therein can consist of a convolutional layer, a batch normalization layer and an activation function. According to the embodiments of the present disclosure, in Figure 2In the example, the 4th fusion image output by the 4th fusion convolutional network module is used as the second image, the image size of the second image can be the same as the size of the first image, and the second image includes text features extracted from the first image.
[0059] According to some embodiments of the present disclosure, Figure 2 In the feature fusion network, the last layer is the direct output of the convolution layer, and the pixel value of the output image is constrained to be distributed in the range of 0-1, so that the extracted text contour transitions naturally and retains the special forms such as gradient or translucency of the original text. That is to say, the pixel points of the fourth fusion image are represented as text probability values in the range of 0-1, and are quantized and output as a grayscale image in the range of 0-255. The second image output by the text cutout method of the embodiment of the present disclosure is a grayscale image with a color scale in the range of 0-255, which can reflect the characteristics of using a neural network for the cutout task. The final output reflects the probability value. The closer to the black area, the closer the probability value is to 0, and the closer to the white area, the closer the probability value is to 1. Finally, the quantization output is a grayscale image with 256 color scales, that is, the text contour obtained by the cutout transitions naturally and supports the special forms such as gradient or translucency of the original text, so that the output second image can be used as a mask image including personalized text features, which is conducive to further applying the mask image obtained by text cutout to specific application scenarios such as video subtitle removal, which will be described below. In comparison, in the scenario of using neural networks for image segmentation tasks, each pixel in the input image is subjected to binary classification, that is, the output value is either 0 or 1, which is different from the text cutout solution involved in the present disclosure.
[0060] Next, we will combine Figure 3 , Figure 4 and Figure 5 To describe the specific implementation of the feature extraction network, the intermediate processing network and the feature fusion network according to the embodiment of the present disclosure, and give specific parameter values in an example manner.
[0061] Figure 3 FIG. 2 shows the overall structure of a neural network for implementing text cutout according to an embodiment of the present disclosure. Figure 3 As shown, the first image including text information can be first input into the feature extraction network, that is, input into the first extraction convolutional network module in the feature extraction network. Figure 3 The first image includes text features and background. The text features specifically include text characters "太阳升起" and "The sun rises", text fonts (Chinese fonts and English fonts) and text colors. Specifically, in this article, the text color can refer to pure colors, textures, patterns, etc. Figure 3In the grayscale image shown, the text color is the grayscale information included in the text itself. It can be understood that the text matte extraction method according to the embodiments of the present disclosure can be used to process various text colors, etc. In addition, in Figure 3 In the example of, the single text resolution can be uniformly scaled to 150 * 150 pixels for easy processing.
[0062] Figure 3 The design goal of the neural network shown is to extract the text features in the first image, that is, to implement text matte extraction, which extracts both text characters and personalized features such as the above-mentioned text font, text color, and text position. For the specific text matte extraction effect, reference can be made to Figure 3 The example of the output second image. As described above, the second image is a grayscale image with pixel values in the range of 0 - 255, which makes the extracted text contour transition naturally and preserves special forms such as the gradient or semi-transparency of the original text. In addition, in some embodiments according to the present disclosure, the image size of the output second image can be the same as that of the first image, which is beneficial to reflecting the position information of the text in the original input image.
[0063] As Figure 3 shown, in the feature extraction network, there are 4 extraction convolutional network modules. Among them, the first extraction convolutional network module processes the first image and outputs the first feature map, and the first feature map is output both downward to the second extraction convolutional network module and to the spatial convolutional network module for extracting the first intermediate feature map.
[0064] Specifically, the first extraction convolutional network module includes a convolutional layer and a max pooling layer. For ease of description, in this article, the parameters of the convolutional layer are represented as Conv_c1_c2_k_s_d_g, and the parameters of the max pooling layer are represented as Maxpool_k_s, where c1 represents the number of input channels, c2 represents the number of output channels, k represents the filter kernel size, s represents the stride, d represents the spatial convolution dilation coefficient (default value is 1), and g represents the separable convolution group coefficient (default value is 1). Thus, Figure 3 The convolutional layer Conv_3_64_7_1 in the first extraction convolutional network module shown in can be understood as a convolutional function with 3 input channels, 64 output channels, a filter kernel size of 7, and a stride of 1, and the spatial convolution dilation coefficient and the separable convolution group coefficient are the default values of 1. Then, the max pooling layer Maxpool_2_2 in the first extraction convolutional network module can be understood as a max pooling function with a filter kernel size of 2 and a stride of 2.
[0065] Next, as Figure 3As shown, in the feature extraction network, the second extraction convolutional network module processes the first feature map and outputs the second feature map. This second feature map is output both downward to the third extraction convolutional network module and to the spatial convolutional network module for extracting the second intermediate feature map. Specifically, the second extraction convolutional network module includes a residual convolutional module, a convolutional layer, and a max pooling layer. The parameters of the convolutional layer and the max pooling layer can refer to the parameter meanings above and will not be elaborated here.
[0066] Figure 4 FIG. shows the network structure diagram of the residual convolutional module according to an embodiment of the present disclosure. Similarly, for ease of description, the parameters of the residual convolutional module (Resnet Block, RB) are denoted as RB_c1_c2, where c1 represents the number of input channels and c2 represents the number of output channels. Further, referring to Figure 4 , the residual convolutional module consists of a convolutional layer, a batch normalization layer (BatchNorm2d), and an activation function (ReLU). It can be understood that in the residual convolutional module, there are two processing paths. One is the processing path including the convolutional layer, the batch normalization layer (BatchNorm2d), and the activation function (ReLU), and the other is used to implement cross-level connection and add the two. The difference between the residual network and the ordinary network lies in the introduction of the skip connection of the above cross-level connection, which can make the information of the previous residual block enter the next processing layer without hindrance, improving the information flow and also avoiding the vanishing gradient problem and degradation problem caused by the increase of network depth.
[0067] Next, as Figure 3 shown, in the feature extraction network, the third extraction convolutional network module processes the second feature map and outputs the third feature map. This third feature map is output both downward to the fourth extraction convolutional network module and to the spatial convolutional network module for extracting the third intermediate feature map. Specifically, the third extraction convolutional network module includes a residual convolutional module, a convolutional layer, and a max pooling layer. The parameters of the convolutional layer and the max pooling layer can refer to the parameter meanings above and will not be elaborated here. Referring to Figure 3 , the third extraction convolutional network module and the second extraction convolutional network module only differ in the parameters of the convolutional layer. As Figure 3 shown, the fourth extraction convolutional network module processes the third feature map and outputs the fourth feature map. This fourth feature map will be output to the spatial convolutional network module for extracting the fourth intermediate feature map. Specifically, the fourth extraction convolutional network module includes a residual convolutional module and a convolutional layer.
[0068] Figure 3Also schematically shown is the network structure diagram of the feature fusion network. In the feature fusion network, there are 4 fusion convolutional network modules. Among them, the first fusion convolutional network module processes the fourth intermediate feature map output by the fourth spatial convolutional network module to obtain the first fusion map. Specifically, the first fusion convolutional network module consists of a convolutional layer and a residual convolutional module. The second fusion convolutional network module processes the third intermediate feature map output by the third spatial convolutional network module and the first fusion map output by the first fusion convolutional network module to obtain the second fusion map. Specifically, the second fusion convolutional network module consists of a first connection layer, a convolutional layer, and a residual convolutional module. Refer to Figure 3 , the first connection layer includes a feature merging layer and a super-resolution layer. The feature merging layer is used to merge multiple input images by increasing the number of image channels, and the super-resolution layer is used to increase the image size by reducing the number of image channels.
[0069] As an implementation, in Figure 3 , the feature merging layer adopts the merging function (Concat), which can be used to integrate two or more groups of input feature maps and increase the number of channels of the output feature map. For example, when the number of channels of the two input feature maps is x and y respectively, the number of channels of the feature map after Concat merging can be x + y. Then, the feature map after the merging operation will be output to the super-resolution layer. In Figure 3 The implementation shown, the super-resolution layer is implemented as the PixelShuffle function. Compared with the general upsampling function, PixelShuffle combines the information of the channel dimension to fill the pixels, which can improve the resolution and generate a more realistic image. Here, PixelShuffle can be understood as transforming a low-resolution image of H×W into a high-resolution image of rH×rW through sub-pixel operations. However, the process of PixelShuffle to achieve resolution improvement is not directly generated by interpolation or other methods, but by the method of periodic shuffling to obtain a high-resolution image. For example, when the number of feature layers before PixelShuffle is 256, an output image with 64 feature layers but improved image resolution can be obtained after PixelShuffle.
[0070] Then refer to Figure 3, the third fusion convolutional network module processes the second intermediate feature map output by the second spatial convolutional network module and the second fusion map output by the second fusion convolutional network module to obtain the third fusion map. The third fusion convolutional network module is composed of a first connection layer, a convolutional layer, and a residual convolutional module. The fourth fusion convolutional network module processes the first intermediate feature map output by the first spatial convolutional network module and the third fusion map output by the third fusion convolutional network module to obtain the fourth fusion map. The fourth fusion convolutional network module is composed of a first connection layer and a convolutional layer. According to some embodiments of the present disclosure, the fourth fusion map output by the fourth fusion convolutional network module will be output as the second image. Specifically, the second image includes text features extracted from the first image.
[0071] As Figure 3 shown, the output second image is a grayscale image with values in the range of 0 - 255. The area with grayscale close to black represents the background part of the first image except for the text, and the area with grayscale close to white represents the text features extracted from the first image, that is, text matting is achieved. The image obtained through text matting includes not only text characters but also features such as text fonts and text colors.
[0072] According to the embodiments of the present disclosure, each spatial convolutional network module in the intermediate processing network can be designed to have the same network structure and parameters. Each spatial convolutional network module includes m parallel processing units and a second connection layer. The m processing units are used to process the input image (for example, the feature map output by the feature extraction network) respectively, and the second connection layer is used to merge the processing results of the m processing units. Among them, the second connection layer includes a feature merging layer, a convolutional layer, and a batch normalization layer. The feature merging layer is used to merge the m images respectively output by the m processing units by increasing the number of image channels, where m is an integer greater than 1.
[0073] Figure 5 shows the network structure diagram of the spatial convolutional network module according to the embodiments of the present disclosure. In Figure 5In the example, m = 5. As an implementation, when m = 5, in each spatial convolution network module, the first processing unit consists of a first convolutional layer and a batch normalization layer connected in sequence, the second processing unit consists of a second convolutional layer, a batch normalization layer, a third convolutional layer, and a batch normalization layer connected in sequence, the third processing unit consists of a fourth convolutional layer, a batch normalization layer, a third convolutional layer, and a batch normalization layer connected in sequence, the fourth processing unit consists of a fifth convolutional layer, a batch normalization layer, a third convolutional layer, and a batch normalization layer connected in sequence, and the fifth processing unit consists of a batch normalization layer, an adaptive average pooling layer, a third convolutional layer, and an upsampling layer connected in sequence, where the network parameters of the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer are different.
[0074] Thus, in Figure 5 's example, for the input feature map, it is first processed separately by 5 parallel processing units. Compared with the feature extraction network, the spatial convolution network module is used to further extract multi-scale feature information, and this can be achieved by using convolutional layers with different spatial convolution dilation coefficients (d). In Figure 5 's example, the spatial convolution dilation coefficients are designed to be 1, 2, 4, 8, and full respectively. As Figure 5 shows, among the 5 processing units, there are convolutional layers with spatial convolution dilation coefficients of 1, 2, 4, 8, and full respectively, where full corresponds to Figure 5 the adaptive average pooling layer (AdaptiveAvgPool2d) included in the fifth processing unit in Figure 3 This layer is used to calculate the global average value for statistical horizontal feature distribution. Regarding the selection of the above spatial convolution dilation coefficients, in the text matting implementation shown in
[0075] Next, referring to Figure 5 , the second connection layer includes a feature merging layer (Concat), a convolutional layer (Conv_1280_c2_1_1_1_1), and a batch normalization layer (BatchNorm2d). Similarly, this feature merging layer Concat is used to merge the 5 processing results obtained from the above parallel processing to obtain a merged feature map, and after being processed by the convolutional layer and the batch normalization layer, the intermediate feature map is output to the corresponding fusion convolutional network module in the feature fusion network.
[0076] According to the text matting method of some embodiments of the present disclosure, it may further include the training process of the neural network. When using such as Figure 3Before performing the text matting processing task on the neural network structure shown, it is also necessary to train the network parameters of the neural network using a training image set. Specifically, the training process includes: generating a training image set and annotating the training images in the training image set to generate corresponding matte images; and using the training image set and based on a training function to train the feature extraction network, the intermediate processing network, and the feature fusion network. Specifically, the above process of annotating the generated training images can be understood as annotating the ground truth, which serves as the goal of the training task.
[0077] According to some embodiments of the present disclosure, generating a training image set includes: constructing a corpus composed of text characters, a font library composed of text fonts, a texture feature library composed of text colors, and a background library composed of background images; and randomly extracting a set of text characters, text fonts, text colors, and background images from the corpus, font library, texture feature library, and background library, and forming a training image.
[0078] According to an embodiment of the present disclosure, a method for making a data set for a text matting training task is provided. Among them, the data set for the text matting task mainly includes four elements: text characters, text fonts, text colors, and background images. Figure 6 The schematic diagram of the data elements for making a data set according to an embodiment of the present disclosure is shown. As Figure 6 shown, the text characters can include Chinese characters and English characters, the text fonts can include Chinese fonts and English fonts, and the text colors can include solid colors, textures, and patterns, etc. According to the examples of the respective elements as Figure 6 shown, a corpus composed of text characters, a font library composed of text fonts, a texture feature library composed of text colors, and a background library composed of background images can be constructed. As an example, the corpus can use a Chinese-English machine translation corpus, which contains approximately 10 million pairs of Chinese-English parallel sentences as text character data. The text fonts can be obtained from a common font library. For example, there are about 100 Chinese fonts and about 500 English fonts. For the texture feature library, the solid color images therein can be randomly generated through image processing functions, etc. The background images can be obtained from a common image processing database.
[0079] Figure 7 The schematic diagram of the process for making a data set according to an embodiment of the present disclosure is shown. According to some embodiments of the present disclosure, in order to obtain as much training image data as possible and fit more and more complex text contents and formats, during the training process of the neural network model, a set of data is randomly extracted from the database as shown in Figure 6 each iteration, and then a training image is made and the corresponding matte image is annotated. Specifically, as Figure 7As shown, for the text characters extracted from the corpus and the Chinese and English fonts extracted from the font library, a text image can be first generated, and this text image can be labeled as a mask image for training. Next, the generated text image can be processed based on the extracted text color to obtain text features with font colors. Further, the generated text features with colors are inserted into the extracted background image as a training image for training. The process of the neural network processing the training image is similar to the process of processing the first image described above, and will not be repeated here.
[0080] As an example, after inputting one of the training images in the training dataset into the neural network, a training output image corresponding to this training image will be generated. For this training output image, the loss value between it and the labeled mask image (i.e., regarded as the ground truth or the target value) can be calculated according to the training function. Using a large number of training images in the training dataset, with the continuous reduction of this loss value as the training objective, to make the neural network, such as Figure 3 shown, have the text matting processing ability that meets the training requirements.
[0081] According to some embodiments of the present disclosure, the training function may include an Alpha loss function for calculating the absolute difference pixel by pixel between the labeled mask image (α g ) and the training output image (α p ) output by the neural network based on the training image. The loss value calculated based on this Alpha loss function is denoted as the Alpha loss value. As an example, the calculation formula of the Alpha loss function is as follows:
[0082]
[0083] According to some embodiments of the present disclosure, the training function may include a Laplacian loss function for decomposing the training output image (α p ) into 5 levels of Gaussian pyramid levels, and then calculating the Alpha loss value between it and the corresponding mask image (α g ) at each level, and finally making a weighted combination to obtain the Laplacian loss value. As an example, the calculation formula of the Laplacian loss function is as follows:
[0084]
[0085] where represents the pyramid level where it is located, and the range of b is 1 - 5.
[0086] According to some embodiments of the present disclosure, the training function may include a Holistically-nested Edge Detection (HED) loss function. Wherein, the edge map of the training output image (α p ) and the edge map of the mask image (α g ) are respectively extracted using the Hed edge detection network, and the L1 norm between the edge map of the mask image (α g ) and the edge map of the training output image (α p ) output by the network is calculated. The loss value calculated based on this HED loss function is denoted as the HED loss value. As an example, the calculation formula of the HED loss function is as follows:
[0087]
[0088] Where, H j (α) represents the edge image of the j-th layer extracted by the HED edge detection network, where the range of j is 1-5.
[0089] According to some embodiments of the present disclosure, the training function may include one or more of the Alpha loss function, Laplacian loss function, and HED loss function described above. In the case where the training function includes all three of the above, the complete loss value calculation formula is as follows:
[0090] L = L α + L lap + 0.5×L hed_edge (4)
[0091] Where, for the Alpha loss value, Laplacian loss value, and HED loss value, weight coefficients of 1, 1, and 0.5 are respectively multiplied.
[0092] As an example, during the training process of the neural network, the relevant training parameters can be set as follows:
[0093]
[0094] Using the neural network-based text matting method according to some embodiments of the present disclosure, a neural network composed of a feature extraction network, an intermediate processing network, and a feature fusion network can be used to perform text matting on a first image including text information to extract text features in the first image, and obtain a second image including the text features. The personalized information such as the font and color of the text is retained in the output second image. In addition, the intermediate processing network can serve as a connection layer between the feature extraction network and the feature fusion network to extract multi-scale intermediate features, thereby improving the accuracy of text feature extraction.
[0095] Further, the second image (which can be expressed as a text feature image or a text mask) obtained by using the above text matting method can also expand the application fields of personalized text features.
[0096] The text matting method according to some embodiments of the present disclosure may further include: performing text detection on the original image to extract a detection image including text from the original image, where the detection image serves as the first image. That is, for the original image, image preprocessing can be first performed to extract a detection image including text therefrom, and the size of the detection image can be smaller than that of the original image to reduce the data volume of image processing.
[0097] Further, the text matting method according to some embodiments of the present disclosure may further include: performing image processing on the text features in the second image to generate a third image after text style conversion; and combining the third image with a target background image to generate a fourth image after image style conversion.
[0098] As an application scenario, the text mask obtained by using the text matting method according to the embodiments of the present disclosure can be used to implement poster style conversion. Figure 8 The application flow chart of using the text matting method according to the embodiments of the present disclosure is shown, as Figure 8 shown, the text matting method can be used in scenarios such as intelligent poster generation and style switching. For the advertised or promotional pictures to be placed, through image processing algorithms such as text detection and instance segmentation, the text (e.g., "Great gift price cut" and "Buy one get one free for best-selling products") and the product feature map in the picture can be first extracted. Then, the above text matting process can be performed on the detected detection image to obtain a text mask with a gray value in the range of 0-255. Next, intelligent matching of the background and style can be performed in combination with the current activity characteristics. For example, Figure 8 the original advertised picture (original image) is in the style of activity promotion during the Spring Festival, including festive elements such as auspicious clouds and lanterns and mainly in red and gold colors. When the same poster content is displayed during other festivals, such as Women's Day, there will be a sense of incongruity. At this time, a new text color suitable for the application scenario can be selected (such as Figure 8The rose gold texture shown in is used to perform text color conversion on the text mask obtained by text matting, so as to obtain the text mask after style conversion. In addition, a new background image suitable for the application scenario (such as a pink-purple background image) can be selected to replace the background of the commodity feature map obtained by instance segmentation. Finally, the text mask after style conversion and the commodity feature map with the background replaced are combined to obtain a new poster image after style switching, so as to realize the intelligent style switching application process. Moreover, the personalized text features and commodity pictures in the original image are retained in the new poster image. By replacing the background image with a pink-purple tone and the text color with a rose gold texture, the regenerated poster better fits the Women's Day festival atmosphere.
[0099] Furthermore, the text matting method according to some embodiments of the present disclosure may further include: performing image processing on the second image to generate a mask image having the same image size as the original image; and using the mask image to perform image completion on the original image to generate the original image after eliminating text features.
[0100] As another application scenario, the text mask obtained by using the text matting method according to the embodiments of the present disclosure can also be used to eliminate picture or video subtitles. Figure 9 shows another application flow chart of using the text matting method according to the embodiments of the present disclosure. As Figure 9 shown, the text matting method can be used to eliminate subtitles in the original image. For some video images, such as video files that have been repeatedly transcoded and compressed, their subtitles have been superimposed on the video screen, resulting in the subtitle content causing a certain degree of occlusion to the video screen, which will affect the subsequent video image editing and re-production processes. Generally, video editors cover or remove subtitles by cropping the video screen or using a brush to smear, etc., but this will cause the screen to be incomplete and not beautiful. By using the text matting method according to the embodiments of the present disclosure in combination with the image completion technology, video subtitles can be better removed while retaining the complete video screen.
[0101] As Figure 9As shown, first, text detection can be performed on the original image to obtain a detection image including text. Then, for the detection image, a text matting method according to an embodiment of the present disclosure is applied to obtain an output text mask. To facilitate the elimination of subtitles in the original image, a subtitle mask is further generated based on the size of the original image, and the image size of the subtitle mask can be the same as that of the original image. Then, the generated subtitle mask is used to perform image completion processing on the original image to obtain an image after subtitle elimination. Among them, the present disclosure does not limit the process of implementing image completion processing. For example, an image completion technology based on a neural network can be used to implement it. The image after subtitle elimination can have the same size as the original image and retain complete video image information. It can be understood that the text matting method according to an embodiment of the present disclosure can also be applied to other application scenarios, which will not be described one by one here. The text mask obtained by the text matting method contains both the extracted text characters and retains the text features unique to the text in the original image, which is beneficial to expanding the scope of text personalized application scenarios.
[0102] Using the neural network-based text matting method according to some embodiments of the present disclosure, a neural network composed of a feature extraction network, an intermediate processing network, and a feature fusion network can be used to perform text matting on a first image including text information to extract the text features in the first image and obtain a second image including the text features. The output second image retains personalized information such as the font and color of the text, which is beneficial to expanding the application fields and scope of the text features. In addition, the intermediate processing network can serve as a connection layer between the feature extraction network and the feature fusion network to extract multi-scale intermediate features, thereby improving the accuracy of text feature extraction.
[0103] According to another aspect of the present disclosure, some embodiments of the present disclosure also provide a neural network-based text matting device. Specifically, Figure 10 FIG. shows a schematic block diagram of a neural network-based text matting device according to an embodiment of the present disclosure.
[0104] As Figure 10 shown, the device 1000 may include a feature extraction unit 1010, an intermediate processing unit 1020, and a feature fusion unit 1030.
[0105] According to some embodiments of the present disclosure, the feature extraction unit 1010 may be configured to process the first image by using a feature extraction network to obtain a feature map. The feature extraction network includes n extraction convolutional network modules connected in sequence. The first extraction convolutional network module processes the first image and outputs the first feature map. The i-th extraction convolutional network module processes the (i - 1)-th feature map and outputs the i-th feature map, where n is an integer greater than 1, and i is an integer greater than 1 and less than or equal to n.
[0106] According to some embodiments of the present disclosure, the intermediate processing unit 1020 may be configured to process the feature map by using an intermediate processing network to obtain an intermediate feature map. The intermediate processing network includes n spatial convolutional network modules. The first spatial convolutional network module processes the first feature map and outputs the first intermediate feature map. The i-th spatial convolutional network module processes the i-th feature map and outputs the i-th intermediate feature map.
[0107] According to some embodiments of the present disclosure, the feature fusion unit 1030 may be configured to process the intermediate feature map by using a feature fusion network to obtain a second image. The feature fusion network includes n fusion convolutional network modules connected in sequence. The first fusion convolutional network module processes the n-th intermediate feature map output by the n-th spatial convolutional network module to obtain the first fusion map. The i-th fusion convolutional network module processes the (n - i + 1)-th intermediate feature map output by the (n - i + 1)-th spatial convolutional network module and the (i - 1)-th fusion map output by the (i - 1)-th fusion convolutional network module to obtain the i-th fusion map. The n-th fusion map output by the n-th fusion convolutional network module is used as the second image, and the second image includes text features extracted from the first image.
[0108] According to the specific implementation processes of the feature extraction unit 1010, the intermediate processing unit 1020, and the feature fusion unit 1030 in the text matting device 1000 according to the embodiments of the present disclosure, reference may be made to the description made above in conjunction with Figure 1 and details are not described herein again.
[0109] According to some embodiments of the present disclosure, the n extraction convolutional network modules in the feature extraction network are respectively composed of one or more of a convolutional layer, a pooling layer, and a residual convolutional module; the n fusion convolutional network modules in the feature fusion network are respectively composed of one or more of a convolutional layer, a residual convolutional module, and a first connection layer, where the first connection layer includes a feature merging layer and a super-resolution layer.
[0110] According to some embodiments of the present disclosure, n is equal to 4. In the feature extraction network, the first extraction convolutional network module consists of a convolutional layer and a max pooling layer, the second extraction convolutional network module consists of a residual convolutional module, a convolutional layer, and a max pooling layer, the third extraction convolutional network module consists of a residual convolutional module, a convolutional layer, and a max pooling layer, and the fourth extraction convolutional network module consists of a residual convolutional module and a convolutional layer; wherein, in the feature fusion network, the first fusion convolutional network module consists of a convolutional layer and a residual convolutional module, the second fusion convolutional network module consists of a first connection layer, a convolutional layer, and a residual convolutional module, the third fusion convolutional network module consists of a first connection layer, a convolutional layer, and a residual convolutional module, and the fourth fusion convolutional network module consists of a first connection layer and a convolutional layer, wherein the first connection layer includes a feature merging layer and a super-resolution layer.
[0111] According to some embodiments of the present disclosure, the feature merging layer is used to merge multiple input images by increasing the number of image channels, and the super-resolution layer is used to increase the image size by reducing the number of image channels, wherein the residual convolutional module consists of a convolutional layer, a batch normalization layer, and an activation function.
[0112] According to some embodiments of the present disclosure, the pixel points of the nth fusion graph are represented as text probability values in the range of 0-1 and quantized into grayscale images in the range of 0-255.
[0113] According to some embodiments of the present disclosure, the structures of the n spatial convolutional network modules are the same. Each spatial convolutional network module includes m parallel processing units and a second connection layer. The m processing units are used to process the input image respectively, and the second connection layer is used to merge the processing results of the m processing units. Among them, the second connection layer includes a feature merging layer, a convolutional layer, and a batch normalization layer. The feature merging layer is used to merge the m images respectively output by the m processing units by increasing the number of image channels, where m is an integer greater than 1.
[0114] According to some embodiments of the present disclosure, m is equal to 5. In each spatial convolutional network module, the first processing unit consists of a first convolutional layer and a batch normalization layer connected in sequence, the second processing unit consists of a second convolutional layer, a batch normalization layer, a third convolutional layer, and a batch normalization layer connected in sequence, the third processing unit consists of a fourth convolutional layer, a batch normalization layer, a third convolutional layer, and a batch normalization layer connected in sequence, the fourth processing unit consists of a fifth convolutional layer, a batch normalization layer, a third convolutional layer, and a batch normalization layer connected in sequence, and the fifth processing unit consists of a batch normalization layer, an adaptive average pooling layer, a third convolutional layer, and an upsampling layer connected in sequence. Among them, the network parameters of the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer are different.
[0115] Regarding the network structures and network parameter designs of the feature extraction network, intermediate processing network, and feature fusion network for performing text matting tasks in the text matting device 1000 according to the embodiments of the present disclosure, reference may be made to the descriptions made above in conjunction with Figures 2 - 5 and will not be elaborated herein.
[0116] According to some embodiments of the present disclosure, the device 1000 may further include a training unit, which may be configured to: generate a training image set and label the training images in the training image set to generate corresponding matte images; and use the training image set and based on a training function to train the feature extraction network, intermediate processing network, and feature fusion network.
[0117] According to some embodiments of the present disclosure, the training unit generating the training image set includes: constructing a corpus composed of text characters, a font library composed of text fonts, a texture feature library composed of text colors, and a background library composed of background images; and randomly extracting a set of text characters, text fonts, text colors, and background images from the corpus, font library, texture feature library, and background library, and forming a training image, where the text characters include Chinese characters and English characters, the text fonts include Chinese fonts and English fonts, and the text colors include solid colors, textures, and patterns. Regarding the specific implementation process of the training unit in the text matting device 1000 according to the embodiments of the present disclosure, reference may be made to the descriptions made above in conjunction with Figure 6 and Figure 7 and will not be elaborated herein.
[0118] According to some embodiments of the present disclosure, the device 1000 may further include an image processing unit, which may be configured to: perform text detection on the original image to extract a detection image including text from the original image, where the detection image serves as a first image.
[0119] According to some embodiments of the present disclosure, the image processing unit may further be configured to: perform image processing on the text features in the second image to generate a third image after text style conversion; and combine the third image with a target background image to generate a fourth image after image style conversion. The implementation manner and effects of this part of the solution may be referred to the descriptions made above in conjunction with Figure 8 and will not be elaborated herein.
[0120] According to some embodiments of the present disclosure, the image processing unit may further be configured to: perform image processing on the second image to generate a matte image having the same image size as the original image; and use the matte image to perform image completion on the original image to generate an original image after eliminating text features. The implementation manner and effects of this part of the solution may be referred to the descriptions made above in conjunction with Figure 9 and will not be elaborated herein.
[0121] Using a neural network-based text matting device according to some embodiments of the present disclosure, it is possible to perform text matting on a first image including text information by using a neural network composed of a feature extraction network, an intermediate processing network, and a feature fusion network, so as to extract the text features in the first image and obtain a second image including the text features. The personalized information such as the font and color of the text is retained in the output second image, which is beneficial to expanding the application fields and scope of the text features. In addition, the intermediate processing network can serve as a connection layer between the feature extraction network and the feature fusion network to extract multi-scale intermediate features, thereby improving the accuracy of text feature extraction.
[0122] According to another aspect of the present disclosure, there is also provided an image processing device. Figure 11 A schematic block diagram of an image processing device according to an embodiment of the present disclosure is shown.
[0123] As Figure 11 shown, the device 2000 may include a processor 2010 and a memory 2020. According to an embodiment of the present disclosure, computer-readable code is stored in the memory 2020, and when the computer-readable code is run by the processor 2010, it executes the text matting method described above.
[0124] The processor 2010 can perform various actions and processes according to the program stored in the memory 2020. Specifically, the processor 2010 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0125] The memory 2020 stores computer-executable instruction codes, which are used to implement the text matting method according to the embodiments of the present disclosure when executed by the processor 2010. The memory 2020 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memories described herein are intended to include but not be limited to these and any other suitable types of memories.
[0126] As an example, the image processing device can be implemented as a central processing unit (CPU), which can be the operation and control core of a computer system and is the final execution unit for information processing and program running. Alternatively, the image processing device can also be implemented as a graphics processing unit (GPU), which is a microprocessor specialized for performing image and graphics-related operations on personal computers, workstations, game consoles, and some mobile devices.
[0127] The text matting method, text matting device or equipment according to the embodiments of the present disclosure can also be implemented by means of Figure 12 the architecture of the computing device 3000 shown. As Figure 12 shown, the computing device 3000 can include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage devices in the computing device 3000, such as the ROM 3030 or the hard disk 3070, can store various data or files used for the processing and / or communication of the text matting method provided by the present disclosure and the program instructions executed by the CPU. The computing device 3000 can also include a user interface 3080. Of course, Figure 12 the architecture shown is only exemplary, and when implementing different devices, according to actual needs, Figure 12One or more components in the computing device shown.
[0128] According to yet another aspect of the present disclosure, a non-transitory computer-readable storage medium is also provided. Figure 13 A schematic diagram 4000 of a storage medium according to the present disclosure is shown.
[0129] As Figure 13 shown, computer-readable instructions 4010 are stored on computer storage medium 4020. When the computer-readable instructions 4010 are run by a processor, the text matting method described with reference to the above figures can be executed. The computer-readable storage medium includes but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. For example, computer storage medium 4020 may be connected to a computing device such as a computer. Then, in the case where the computing device runs the computer-readable instructions 4010 stored on computer storage medium 4020, the text matting method provided according to the embodiments of the present disclosure as described above can be performed.
[0130] Those skilled in the art can understand that the content disclosed in the present disclosure can have various variations and improvements. For example, the various devices or components described above can be implemented by hardware, or by software, firmware, or some or all combinations of the three.
[0131] In addition, although the present disclosure makes various references to certain units in the system according to the embodiments of the present disclosure, however, any number of different units can be used and run on the client and / or server. The units are only illustrative, and different aspects of the system and method can use different units.
[0132] Those of ordinary skill in the art can understand that all or part of the steps in the above methods can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as read-only memory, magnetic disk or optical disk, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software functional module. The present disclosure is not limited to any specific form of the combination of hardware and software.
[0133] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It should also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and should not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0134] The foregoing is a description of the disclosure and should not be taken as a limitation thereof. Although several exemplary embodiments of the disclosure have been described, those skilled in the art will readily appreciate that many modifications can be made to the exemplary embodiments without departing from the novel teachings and advantages of the disclosure. Accordingly, all such modifications are intended to be included within the scope of the disclosure as defined by the claims. It should be understood that the foregoing is a description of the disclosure and should not be considered limited to the particular embodiments disclosed, and modifications to the disclosed embodiments as well as other embodiments are intended to be included within the scope of the appended claims. The disclosure is defined by the claims and their equivalents.
Claims
1. A text matting method based on a neural network, comprising: Processing a first image by using a feature extraction network to obtain a feature map, wherein the feature extraction network includes n extraction convolutional network modules connected in sequence, wherein the first extraction convolutional network module processes the first image and outputs a first feature map, and the i-th extraction convolutional network module processes the (i - 1)-th feature map and outputs an i-th feature map; Processing the feature map by using an intermediate processing network to obtain an intermediate feature map, wherein the intermediate processing network includes n spatial convolutional network modules, wherein the first spatial convolutional network module processes the first feature map and outputs a first intermediate feature map, and the i-th spatial convolutional network module processes the i-th feature map and outputs an i-th intermediate feature map; and Processing the intermediate feature map by using a feature fusion network to obtain a second image, wherein the feature fusion network includes n fusion convolutional network modules connected in sequence, the first fusion convolutional network module processes the n-th intermediate feature map output by the n-th spatial convolutional network module to obtain a first fusion map, and the i-th fusion convolutional network module processes the (n - i + 1)-th intermediate feature map output by the (n - i + 1)-th spatial convolutional network module and the (i - 1)-th fusion map output by the (i - 1)-th fusion convolutional network module to obtain an i-th fusion map, wherein the n-th fusion map output by the n-th fusion convolutional network module serves as the second image, and the second image includes text features extracted from the first image, wherein n is an integer greater than 1, and i is an integer greater than 1 and less than or equal to n.
2. The method according to claim 1, wherein The n extraction convolutional network modules in the feature extraction network are respectively composed of one or more of a convolutional layer, a pooling layer, and a residual convolutional module; the n fusion convolutional network modules in the feature fusion network are respectively composed of one or more of a convolutional layer, a residual convolutional module, and a first connection layer, wherein the first connection layer includes a feature merging layer and a super-resolution layer.
3. The method according to claim 1, characterized in that, n is equal to 4. In the feature extraction network, the first extraction convolutional network module is composed of a convolutional layer and a max pooling layer, the second extraction convolutional network module is composed of a residual convolutional module, a convolutional layer, and a max pooling layer, the third extraction convolutional network module is composed of a residual convolutional module, a convolutional layer, and a max pooling layer, and the fourth extraction convolutional network module is composed of a residual convolutional module and a convolutional layer; In the feature fusion network, the first fusion convolutional network module is composed of a convolutional layer and a residual convolutional module, the second fusion convolutional network module is composed of a first connection layer, a convolutional layer, and a residual convolutional module, the third fusion convolutional network module is composed of a first connection layer, a convolutional layer, and a residual convolutional module, and the fourth fusion convolutional network module is composed of a first connection layer and a convolutional layer, wherein the first connection layer includes a feature merging layer and a super-resolution layer.
4. The method according to claim 2 or 3, characterized in that, The feature merging layer is used to merge multiple input images by increasing the number of image channels, and the super-resolution layer is used to increase the image size by reducing the number of image channels. Among them, the residual convolution module consists of a convolution layer, a batch normalization layer, and an activation function.
5. The method according to claim 1, wherein The structures of the n spatial convolution network modules are the same. Each spatial convolution network module includes m parallel processing units and a second connection layer. The m processing units are used to process the input image respectively, and the second connection layer is used to merge the processing results of the m processing units. Among them, the second connection layer includes a feature merging layer, a convolution layer, and a batch normalization layer. Among them, the feature merging layer is used to merge the m images respectively output by the m processing units by increasing the number of image channels. Among them, m is an integer greater than 1.
6. The method according to claim 5, wherein m is equal to 5. In each spatial convolution network module, the first processing unit consists of a first convolution layer and a batch normalization layer connected in sequence. The second processing unit consists of a second convolution layer, a batch normalization layer, a third convolution layer, and a batch normalization layer connected in sequence. The third processing unit consists of a fourth convolution layer, a batch normalization layer, a third convolution layer, and a batch normalization layer connected in sequence. The fourth processing unit consists of a fifth convolution layer, a batch normalization layer, a third convolution layer, and a batch normalization layer connected in sequence. The fifth processing unit consists of a batch normalization layer, an adaptive average pooling layer, a third convolution layer, and an upsampling layer connected in sequence. Among them, the network parameters of the first convolution layer, the second convolution layer, the third convolution layer, the fourth convolution layer, and the fifth convolution layer are different.
7. The method according to claim 1, characterized in that The pixel points of the nth fusion graph are represented as text probability values in the 0-1 interval and quantized into grayscale images in the 0-255 interval.
8. The method according to claim 1, further comprising: Generating a training image set and annotating the training images in the training image set to generate corresponding mask images; And Using the training image set and based on a training function to train the feature extraction network, the intermediate processing network, and the feature fusion network.
9. The method according to claim 8, wherein The generating of the training image set includes: Constructing a corpus composed of text characters, a font library composed of text fonts, a texture feature library composed of text colors, and a background library composed of background images; and Randomly extracting a set of text characters, text fonts, text colors, and background images from the corpus, the font library, the texture feature library, and the background library, and forming a training image. Among them, the text characters include Chinese characters and English characters, the text fonts include Chinese fonts and English fonts, and the text colors include solid colors, textures, and patterns.
10. The method according to claim 1, further comprising: Performing text detection on the original image to extract a detection image including text from the original image, where the detection image serves as the first image.
11. The method according to claim 10, further comprising: Performing image processing on the text features in the second image to generate a third image after text style conversion; And Combining the third image with a target background image to generate a fourth image after image style conversion.
12. The method according to claim 10, further comprising: Perform image processing on the second image to generate a mask image with the same image size as the original image; and Use the mask image to perform image completion on the original image to generate an original image after eliminating the text features.
13. A text matting device based on a neural network, comprising: A feature extraction unit configured to process a first image using a feature extraction network to obtain a feature map, wherein the feature extraction network includes n extraction convolutional network modules connected in sequence, wherein the first extraction convolutional network module processes the first image and outputs a first feature map, and the i-th extraction convolutional network module processes the (i - 1)-th feature map and outputs an i-th feature map; An intermediate processing unit configured to process the feature map using an intermediate processing network to obtain an intermediate feature map, wherein the intermediate processing network includes n spatial convolutional network modules, wherein the first spatial convolutional network module processes the first feature map and outputs a first intermediate feature map, and the i-th spatial convolutional network module processes the i-th feature map and outputs an i-th intermediate feature map; and A feature fusion unit configured to process the intermediate feature map using a feature fusion network to obtain a second image, wherein the feature fusion network includes n fusion convolutional network modules connected in sequence, the first fusion convolutional network module processes the n-th intermediate feature map output by the n-th spatial convolutional network module to obtain a first fusion map, and the i-th fusion convolutional network module processes the (n - i + 1)-th intermediate feature map output by the (n - i + 1)-th spatial convolutional network module and the (i - 1)-th fusion map output by the (i - 1)-th fusion convolutional network module to obtain an i-th fusion map, wherein the n-th fusion map output by the n-th fusion convolutional network module is used as the second image, and the second image includes text features extracted from the first image, wherein n is an integer greater than 1, and i is an integer greater than 1 and less than or equal to n.
14. An image processing device, comprising: A processor; and A memory, wherein computer-readable code is stored in the memory, and when the computer-readable code is run by the processor, it executes the text matting method according to any one of claims 1 - 12.
15. A non-transitory computer-readable storage medium, on which instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the text matting method according to any one of claims 1 - 12.
Citation Information
Patent Citations
A natural scene text detection method based on full convolution neural network
CN109299274A
Pet-MRI image denoising method and apparatus based on dual-encoding fusion network model
WO2022126588A1