Text processing method and apparatus, electronic device, and readable storage medium
By obtaining the style characteristics of the user's handwritten text and adjusting it to reference style characteristics, and using the generative adversarial network model processing, the continuous or lack of strokes of the user's handwritten text is solved, and the beautification and recognition experience is improved.
Patent Information
- Application Number
- PCT/CN2023/141539
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, the beautification scheme of handwritten text by users often leads to poor beautification experience due to continuous strokes or lack of strokes, especially when writing at different times or when the same user is writing.
By obtaining the user style features of the original image and adjusting it to reference style features, using the generative adversarial network model for processing, the target text map is generated to solve the problem of missing strokes or continuous strokes and improve the recognition experience.
It effectively solves the problem of missing strokes or continuous strokes of user handwritten text, and improves the beautification and recognition experience of text content.
Smart Images

Figure CN2023141539_03072025_PF_FP_ABST
Abstract
Description
Text processing method, device, electronic device and readable storage medium Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a text processing method, device, electronic device, and readable storage medium. Background Art
[0002] As the number of scenarios in which users can write on electronic devices (such as smartphones, tablet computers, or large-size displays, etc.) increases, there is a demand in related technologies to beautify user handwritten text to facilitate reading.
[0003] In real-world scenarios, each user has a different handwriting style, and even the same character written at different times can vary significantly. For example, when a font has connected strokes or missing strokes, the beautification solutions of related technologies often result in extra or missing strokes, affecting the beautification experience.
[0004] Summary of the Invention
[0005] The present disclosure provides a text processing method, device, electronic device and readable storage medium to address the deficiencies of related technologies.
[0006] According to a first aspect of an embodiment of the present disclosure, a text processing method is provided, the method comprising:
[0007] Obtaining user style features of the text content contained in the original image;
[0008] Acquire an original image for recognition processing to obtain a content font image corresponding to the original image;
[0009] Acquire a reference style feature corresponding to the user style feature of the text content;
[0010] The text content is processed based on the reference style feature to obtain a target text image, and the writing style of the text content of the target text image is the reference style corresponding to the reference style feature.
[0011] Optionally, obtaining user style features of the text content contained in the original image includes:
[0012] Get the style encoding model;
[0013] The original image is input into the style encoding model to obtain the style features output by the style encoding model as the user style features of the text content contained in the original image.
[0014] Optionally, obtaining user style features of the text content contained in the original image includes:
[0015] determining a user style feature of each character in the text content as the user style feature of the text content;
[0016] The obtaining of a reference style feature corresponding to the user style feature of the text content includes:
[0017] Obtaining reference style features of user style features of each word in the text content;
[0018] The step of processing the text content based on the reference style feature to obtain a target text graph includes:
[0019] The characters are processed based on the reference style features of the characters to obtain a target text graph.
[0020] Optionally, obtaining user style features of the text content contained in the original image includes:
[0021] Obtaining user style features of multiple characters in the text content;
[0022] A weighted feature of the user style features of the plurality of characters is obtained as the user style feature of the text content.
[0023] Optionally, obtaining a reference style feature corresponding to the user style feature of the text content includes:
[0024] Obtaining the first candidate style feature of each type of style text in the candidate style feature library;
[0025] Obtaining a plurality of first principal component style features according to the similarity between the user style feature and the first candidate style feature of each type of style text;
[0026] The plurality of first principal component style features are weighted to obtain the reference style feature.
[0027] Optionally, obtaining a reference style feature corresponding to the user style feature of the text content includes:
[0028] Extracting second candidate style features of the same character from a candidate style feature library according to each character in the text content;
[0029] Calculating the similarity between the user style feature of each character and each second candidate style feature of the same character, and obtaining a plurality of second principal component style features with the highest similarity values when sorted in descending order;
[0030] The plurality of second principal component style features are weighted to obtain a first weighted style feature.
[0031] Optionally, the method further includes:
[0032] The first weighted style feature is used as a reference style feature for each character.
[0033] Optionally, the candidate style feature library is implemented by the following steps:
[0034] Get M texts;
[0035] Get N different types of text styles;
[0036] Obtaining a text image of each of the M characters according to N text styles to obtain M*N text images;
[0037] Inputting each text image into a style encoding model to obtain style features output by the style encoding model;
[0038] The identification codes and style features of the various text images are stored in a style feature library to obtain the candidate style feature library.
[0039] Optionally, weighting the plurality of second principal component style features to obtain a first weighted style feature includes:
[0040] respectively obtaining a similarity value corresponding to each second principal component style feature in the plurality of second principal component style features as an initial weight;
[0041] Normalizing the initial weights of the plurality of second principal component style features to obtain a normalized weight of each second principal component style feature;
[0042] A first weighted style feature of the plurality of principal component style features is calculated based on the normalized weights of the respective second principal component style features.
[0043] Optionally, the method further includes:
[0044] Get the specified style features in the configuration data;
[0045] The designated style feature and the first weighted style feature are weighted to obtain a second weighted style feature as a reference style feature for each character.
[0046] Optionally, processing the text content based on the reference style feature to obtain a target text graph includes:
[0047] Obtain content encoding model and content decoding model respectively;
[0048] Obtaining a content font image corresponding to the original image;
[0049] Inputting the content font image into the content coding model to obtain a text feature map output by the content coding model;
[0050] The text feature graph of each character and its reference style feature are input into the content decoding model to obtain a target text graph output by the content decoding model, wherein the writing style of each character in the target text graph is the reference style corresponding to the reference style feature.
[0051] Optionally, obtaining a content font image corresponding to the original image includes:
[0052] Determining that the original image is the content font image;
[0053] or,
[0054] Perform text recognition on the original image to obtain a text recognition result, and obtain a content text image with a preset style according to the text recognition result.
[0055] Optionally, the content encoding model, the content decoding model, and the style encoding model are trained using a generative adversarial approach, including:
[0056] Generate prediction results using a generator consisting of a content encoding model, a content decoding model, and a style encoding model;
[0057] Use the prediction results and original images to train the discriminator;
[0058] Training the generator using the trained discriminator;
[0059] In response to detecting that the number of training times is less than or equal to a preset number threshold, continuing to execute the step of generating a prediction result;
[0060] In response to detecting that the number of training times is greater than a preset number threshold, it is determined that the generator has completed training.
[0061] Optionally, a generator composed of a content encoding model, a content decoding model, and a style encoding model is used to generate prediction results, including:
[0062] Get multiple sample images;
[0063] Input each sample image into the style encoding model in turn to obtain the user style features output by the style encoding model;
[0064] Acquire an original image for recognition processing to obtain a content font image corresponding to the original image;
[0065] Acquire a reference style feature corresponding to the user style feature of the text content;
[0066] Input the sample images of each pair of sample images into the content coding model in turn to obtain a text feature map;
[0067] The text feature map of each character and its reference style feature are input into the content decoding model to obtain the predicted text image output by the content decoding model as the prediction result of the generator.
[0068] Optionally, the content coding model includes 3 deformable convolution layers, 1 convolution layer and 2 residual block layers.
[0069] Optionally, the content decoding model includes 2 residual block layers and 4 upsampling convolution layers.
[0070] According to a second aspect of an embodiment of the present disclosure, a text processing device is provided, the device comprising:
[0071] A user style acquisition module is used to obtain user style features of the text content contained in the original image;
[0072] A reference style acquisition module, configured to acquire reference style features corresponding to user style features of the text content;
[0073] The target text image acquisition module is used to process the text content based on the reference style feature to obtain a target text image, wherein the writing style of the text content of the target text image is the reference style corresponding to the reference style feature.
[0074] Optionally, the user style acquisition module includes:
[0075] The style model acquisition submodule is used to obtain the style encoding model;
[0076] The user style acquisition submodule is used to input the original image into the style encoding model to obtain the style features output by the style encoding model as the user style features of each word in the text content contained in the original image.
[0077] Optionally, the user style acquisition module further includes:
[0078] The first feature determination submodule is configured to determine a user style feature of each character in the text content as the user style feature of the text content.
[0079] Optionally, the user style acquisition module includes:
[0080] The style feature acquisition and determination submodule is further used to obtain user style features of multiple characters in the text content;
[0081] The second feature determination submodule is configured to obtain weighted features of the user style features of the plurality of characters as the user style features of the text content.
[0082] Optionally, the reference style acquisition module includes:
[0083] A candidate style acquisition submodule is used to extract, based on each character in the text content, second candidate style features of the same character from a candidate style feature library;
[0084] The principal component style acquisition submodule is used to calculate the similarity between the user style feature of each character and each second candidate style feature of the same character, and obtain multiple second principal component style features with the highest similarity value when sorted in descending order;
[0085] The first reference style acquisition submodule is configured to perform weighted processing on the plurality of second principal component style features to obtain a first weighted style feature.
[0086] Optionally, the candidate style feature library is implemented by the following steps:
[0087] Get M texts;
[0088] Get N different types of text styles;
[0089] Obtaining a text image of each of the M characters according to N text styles to obtain M*N text images;
[0090] Inputting each text image into a style encoding model to obtain style features output by the style encoding model;
[0091] The identification code and style feature of each character image in the character matrix are stored in a style feature library to obtain the candidate style feature library.
[0092] Optionally, the first reference style acquisition submodule includes:
[0093] an initial weight obtaining unit, configured to obtain a similarity value corresponding to each of the plurality of second principal component style features as an initial weight;
[0094] a normalized weight obtaining unit, configured to perform normalization processing on the initial weights of the plurality of second principal component style features to obtain a normalized weight of each second principal component style feature;
[0095] The first reference style acquisition unit is configured to calculate a first weighted style feature of the plurality of second principal component style features based on the normalized weights of the second principal component style features as a reference style feature of each character.
[0096] Optionally, the device further comprises:
[0097] A specified style acquisition module is used to obtain specified style features in the configuration data;
[0098] The second reference style acquisition module is configured to perform weighted processing on the designated style feature and the first weighted style feature to obtain a second weighted style feature as a reference style feature for each character.
[0099] Optionally, the target text graph acquisition module includes:
[0100] A model acquisition unit, used to respectively acquire a content encoding model and a content decoding model;
[0101] A content image acquisition unit, configured to acquire a content font image corresponding to the original image;
[0102] a text feature acquisition unit, configured to input the content font image into the content coding model to obtain a text feature map output by the content coding model;
[0103] The target text graph acquisition unit is used to input the text feature graph of each character and its reference style feature into the content decoding model to obtain the target text graph output by the content decoding model, wherein the writing style of each character in the target text graph is the reference style corresponding to the reference style feature.
[0104] Optionally, the content image acquisition unit includes:
[0105] a first determining subunit, configured to determine that the original image is the content font image;
[0106] or,
[0107] The second determining subunit is configured to perform recognition processing on the original image to obtain a content text image corresponding to the original image.
[0108] Optionally, the content encoding model, the content decoding model, and the style encoding model are trained using a generative adversarial approach, including:
[0109] Generate prediction results using a generator consisting of a content encoding model, a content decoding model, and a style encoding model;
[0110] Use the prediction results and original images to train the discriminator;
[0111] Training the generator using the trained discriminator;
[0112] In response to detecting that the number of training times is less than or equal to a preset number threshold, continuing to execute the step of generating a prediction result;
[0113] In response to detecting that the number of training times is greater than a preset number threshold, it is determined that the generator has completed training.
[0114] Optionally, a generator composed of a content encoding model, a content decoding model, and a style encoding model is used to generate prediction results, including:
[0115] Get multiple sample images;
[0116] Input each sample image into the style encoding model in turn to obtain the user style features output by the style encoding model;
[0117] Acquire an original image for recognition processing to obtain a content font image corresponding to the original image;
[0118] Acquire a reference style feature corresponding to the user style feature of the text content;
[0119] Input the sample images of each pair of sample images into the content coding model in turn to obtain a text feature map;
[0120] The text feature map of each character and its reference style feature are input into the content decoding model to obtain the predicted text image output by the content decoding model as the prediction result of the generator.
[0121] Optionally, the style encoding model is implemented using an AlexNet network model or a VGG network model.
[0122] Optionally, the content coding model includes 3 deformable convolution layers, 1 convolution layer and 2 residual block layers.
[0123] Optionally, the content decoding model includes 2 residual block layers and 4 upsampling convolution layers.
[0124] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, comprising
[0125] a processor; a memory for storing a computer program executable by the processor;
[0126] The processor is configured to execute the computer program in the memory to implement the method as described in any one of the first aspects.
[0127] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which, when an executable computer program in the storage medium is executed by a processor, can implement the method described in any one of the first aspects.
[0128] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:
[0129] As can be seen from the above embodiments, the solution provided by the embodiments of the present disclosure obtains the user style features of the text content contained in the original image; then, obtains the reference style features corresponding to the user style features of the text content; thereafter, the text content is processed based on the reference style features to obtain a target text image, and the writing style of the text content of the target text image is the reference style corresponding to the reference style features. In this way, in this embodiment, a reference style feature similar to the user style feature is used to replace the user style feature, which can adjust the text content to the reference style and solve problems such as missing strokes, broken strokes, or connected strokes, which is conducive to improving the recognition experience.
[0130] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0131] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0132] Fig. 1 is a flowchart showing a text processing method according to an exemplary embodiment.
[0133] Fig. 2 is a flow chart showing a reference style feature according to an exemplary embodiment.
[0134] Fig. 3 is a schematic diagram showing the style of a certain text in a candidate style feature library according to an exemplary embodiment.
[0135] Fig. 4 is a flow chart showing a method of obtaining weighted style features according to an exemplary embodiment.
[0136] FIG5 is a schematic diagram showing the architecture of a training style encoding model, a content encoding model, and a content decoding model according to an exemplary embodiment.
[0137] Fig. 6 is a schematic diagram showing a generator outputting prediction results according to an exemplary embodiment.
[0138] Fig. 7 is a block diagram showing a text processing apparatus according to an exemplary embodiment. DETAILED DESCRIPTION
[0139] Exemplary embodiments will be described in detail herein, with examples shown in the accompanying drawings. When the following description refers to the drawings, identical numbers in different drawings represent identical or similar elements, unless otherwise indicated. The exemplary embodiments described below do not represent all embodiments consistent with the present disclosure. Rather, they are merely examples of devices consistent with certain aspects of the present disclosure, as detailed in the appended claims. It should be noted that, unless there is a conflict, the features of the following embodiments and implementations may be combined with each other.
[0140] To solve the above technical problems, the embodiments of the present disclosure provide a text processing method, apparatus, electronic device, and readable storage medium. FIG1 is a flowchart of a text processing method according to an exemplary embodiment. Referring to FIG1 , a text processing method includes steps 11 to 13:
[0141] In step 11, the user style features of the text content contained in the original image are obtained.
[0142] In this step, the electronic device may obtain user style features of the text content contained in the original image. The user style features of the text content contained in the original image may include user style features of individual characters in the text content, user style features of individual lines of text, user style features of individual paragraphs of text, or user style features of the entire text content.
[0143] It should be noted that the user style feature is used to characterize the writing style formed by the user's written text. For example, the above-mentioned writing style may include fonts, such as Kaiti, Songti, Fangsongti, and Heiti; the above-mentioned writing style may include pen strokes, such as center-point pen strokes, side-point pen strokes, and a tendency towards various pen strokes; the above-mentioned writing style may include stroke thickness, such as each stroke has the same thickness, thick at the beginning and thin at the end, horizontal lines are thick and thin, and vertical lines are thin; the above-mentioned writing style may include text shape, such as round text, square text, elongated text, and tilted text; the above-mentioned writing style may include text proportions, such as text with a ratio of wide at the top and narrow at the bottom, and text with a ratio of wide at the bottom and thin in the middle; the above-mentioned writing style may include text fluency, such as smooth, slow, and jerky; the above-mentioned writing style may include text artistic conception, such as elegant, heavy, novice, and rounded. It is understandable that this example only illustrates a part of the features included in the user style characteristics. You can choose at least one of the above or add other writing styles according to the specific scenario, such as font size, adjacent font spacing, line spacing and other styles. If each user can be distinguished, the corresponding style characteristics fall within the scope of protection of this disclosure.
[0144] In one embodiment, the electronic device can obtain user style features for each character in the text content. For example, the electronic device can obtain a style encoding model; this style encoding model has been trained. The specific training method will be described in subsequent embodiments and will not be explained here. The electronic device can then input the original image into the style encoding model to obtain the style features output by the style encoding model as the user style features for each character in the text content contained in the original image.
[0145] In one example, when the original image is an 80-pixel*80-pixel image, it can be directly input into the above style encoding model.
[0146] In one example, when the size of the original image is larger than 80 pixels * 80 pixels, the size of the original image can be adjusted first, and finally become an image of 80 pixels * 80 pixels. For example, a user can write text content on the touch screen of an electronic device. The electronic device can collect track point data during the user's writing process, where the track point data includes the current track point x-coordinate, y-coordinate, and whether the pen is lifted, recorded as (x, y, isLeave); then, the track point is converted into an initial image through track point mapping. The electronic device can generate an initial image after the user writes a word. When the initial image is larger than 80 pixels * 80 pixels, the electronic device can calculate the maximum height and maximum width of the text, and select the larger value of the maximum height and maximum width; then, the scaling ratio is calculated based on the maximum side, that is, the length of the maximum side is reduced to 80 pixels, and the other sides other than the maximum side can be filled with additional pixels, thereby obtaining an original image of 80 pixels * 80 pixels.
[0147] In one example, the style encoding model can be implemented using an AlexNet network model or a VGG network model.
[0148] Table 1 Style encoding model structure
[0149] Taking the VGG11 network model as an example, as shown in Table 1, the VGG11 network model includes 3*3 convolutional layers, 5 max pooling layers, 1 average pooling layer, and 1 fully connected layer. The output dimension of the last fully connected layer is 128. In other words, an 80*80 pixel image is converted into 2*2 pixel feature data through the five layers of downsampling and pooling in the VGG11 network model. This 2*2 pixel feature data is then converted into 1*1 pixel feature data with 512 channels through the average pooling layer. The final fully connected layer then reduces the 512-dimensional channels to 128 dimensions, resulting in a 1*1*128-dimensional user style feature. In other words, the style encoding model encodes each original image into a 128-dimensional feature vector.
[0150] In one embodiment, the electronic device may obtain user style features of each line of text. For example, the electronic device may obtain at least one character from each line of text, such as by random selection or selection from a specified location; then, the electronic device may perform weighted processing on the user style features of the at least one character, and use the weighted processing result as the user style features of the line of text.
[0151] It should be noted that the electronic device can obtain the user style features of each text and then calculate the user style features of each line of text, and then calculate the user style features of each line of text; it can also select the user style features of at least one text in each line of text, and then calculate the user style features of each line of text, thereby reducing the amount of calculation.
[0152] In one embodiment, the electronic device can obtain user style features of each paragraph of text content. For example, the electronic device can obtain at least one character from each paragraph, such as by randomly selecting or selecting from a specified location; then, the electronic device can perform weighted processing on the user style features of the at least one character, and use the weighted processing result as the user style features of the paragraph.
[0153] It should be noted that the electronic device can obtain the user style features of each text and then calculate the user style features of each paragraph of text, and then calculate the user style features of each paragraph of text; it can also select the user style features of at least one text in each paragraph of text, and then calculate the user style features of each line of text, thereby reducing the amount of calculation.
[0154] In one embodiment, the electronic device may obtain the user style characteristics of the entire text content. For example, the electronic device may obtain the user style characteristics of at least one character in the text content, perform weighted processing on the user style characteristics of the at least one character, and use the weighted processing result as the user style characteristics of the entire text content.
[0155] It should also be noted that the electronic device can use the user style features of the above-mentioned single text, a line of text, a paragraph of text or the entire text content as the user style features of the text content; or, in the case of multiple texts, multiple lines of text, or multiple paragraphs of text, the user style features can be weighted to obtain weighted features as the user style features of the text content. The weighted processing can be to first obtain the product of the user style features of each parameter (i.e., a single text, a line of text, or a paragraph of text) and the weight and then accumulate them to obtain the final accumulated value as the weighted feature; or to calculate the average value of the above accumulated value and use the average value as the weighted feature. In addition, during the weighted processing, the weights of the various parameters can be equal, thereby reflecting that the various parameters have the same contribution to the user style features; of course, the weights of the various parameters can also be unequal, for example, the style feature weight of a text with more strokes is larger, and the weight of a text with fewer strokes is smaller, or the weight of a line with more characters is larger and the weight of a line with fewer characters is smaller, etc., which can be selected according to the specific scenario, and the corresponding scheme falls within the scope of protection of this disclosure.
[0156] In step 12, a reference style feature corresponding to the user style feature of the text content is obtained.
[0157] In this step, the electronic device may obtain a reference style feature corresponding to the user style feature of the text content, as shown in FIG. 2 , which includes steps 21 to 23 .
[0158] In step 21, second candidate style features of the same character are extracted from a candidate style feature library based on each character in the text content.
[0159] In this step, the electronic device may store a candidate style feature library. The candidate style feature library includes multiple characters and different types of character styles, such as Kaiti, Songti, Fangsongti, Heiti, etc. The candidate style feature library can be implemented by the following steps:
[0160] First, the electronic device can obtain M characters from a specified character encoding library. For example, the specified character encoding library can be the GB18030 character encoding library, and the electronic device can select all first-level Chinese characters and second-level Chinese characters from the GB18030 character encoding library, which is M in number.
[0161] The electronic device can then obtain N different types of text styles. These text styles can be public or private versions of text style data. For example, the electronic device can select N = 1000 different types of text styles from the Chinese character library on the "Fonts World" website (its website is https: / / www.fonts.net.cn) as the "base" for the handwritten text.
[0162] Then, the electronic device can obtain the text images of each of the M characters according to N text styles, obtaining M*N text images, as shown in Figure 3.
[0163] After that, the electronic device can input each character in the text matrix into the style encoding model to obtain the style features input by the style encoding model.
[0164] Finally, the electronic device can store the recognition codes and style features of each text image in the text matrix into the style feature library to obtain a candidate style feature library. For example, the key in this candidate style feature library can be the recognition code of the text image, and this recognition code can be the serial number of the text image, the image encoding feature, or the recognition code, etc., and the value can be the style features of this text image in different text styles.
[0165] In this step, while obtaining the user style features of each character, the character is also recognized. At this time, the electronic device can use each of the above characters to retrieve the above candidate style feature library to obtain multiple candidate style features of the same character, that is, the second candidate style features. For example, after the electronic device obtains the character "遏", it can detect N different types of "遏" in the candidate style feature library, and each type of "遏" corresponds to a candidate style feature.
[0166] In step 22, calculate the similarity values between the user style features of each character and each of the second candidate style features of the same character, and obtain multiple second principal component style features with higher similarity values when sorted in descending order.
[0167] In this step, the electronic device can calculate the similarity values between the user style features of each character and each of the candidate style features of the same character; and sort the multiple similarity values of the same character in descending order, and multiple principal component style features with higher similarity values when sorted in descending order can be obtained, that is, the second principal component style features. Among them, the number of second principal component style features can be selected according to the specific scenario.
[0168] In step 23, perform a weighted processing on the multiple second principal component style features to obtain a first weighted style feature.
[0169] In this step, the electronic device can perform weighted processing on multiple second principal component style features. For example, the electronic device can respectively obtain the similarity values corresponding to each second principal component style feature in the multiple second principal component style features as the initial weights. Then, the electronic device can perform normalization processing on the initial weights of the multiple second principal component style features, such as using the softmax function for normalization processing, to obtain the normalized weights of each second principal component style feature. Finally, the electronic device can calculate the weighted style features of at least one second principal component style feature based on the normalized weights of each second principal component style feature as the reference style feature of the text.
[0170] Continuing with the example of the character "遏", referring to Figure 4, the electronic device can calculate the similarity value between the user style feature and the candidate style features of all "遏" characters in the candidate style feature library. This similarity value can be represented by a cosine value. Then, sort each similarity value and select the top 5 larger similarity values, which are 0.5, 0.4, 0.3, 0.2, and 0.1 respectively. Then, the electronic device can use the softmax function to perform normalization processing on each second principal component style feature to obtain the normalized weights, which are 0.33, 0.26, 0.2, 0.13, and 0.07 in sequence. Since each second principal component style feature is a known quantity and their respective normalized weights are known quantities, at this time, weighted processing can be performed on the 5 second principal component style features, that is, the features at the same position in the 5 second principal component style features are multiplied by their respective normalized weights, and then the 5 products of the same feature are summed to obtain the value of the feature at the same position in the weighted style feature, thereby obtaining the weighted style feature. For the convenience of describing the solution, the weighted style feature here is referred to as the first weighted style feature in this example to distinguish it from the second weighted style feature in subsequent embodiments.
[0171] In one example, the solution of the present disclosure supports the user to configure the text style. For example, the electronic device can display a configuration menu. When the text style option in the configuration menu is selected, multiple text style options can be displayed within the display interface of the electronic device. When it is detected that one or more text styles are selected, such as the italic font, at this time, the style features of the text style selected by the user can be stored as the specified style features in the configuration data.
[0172] The electronic device can obtain the specified style features in the configuration data. Then, the electronic device can perform weighted processing on the specified style features and the first weighted style features to obtain the second weighted style features as the reference style features of each text.
[0173] f = w * f1 + (1 - w) * f2; (1)
[0174] In formula (1), w represents the weight of the specified style feature, such as 0.3, which means retaining 30% of the specified font style feature, f1 represents the specified style feature, f2 represents the first weighted style feature, and f represents the reference style feature.
[0175] In some possible examples, the electronic device can obtain candidate style features of the same type of text in the candidate style feature library, such as Kaiti, Songti, Fangsongti, etc.; it is understandable that the above-mentioned candidate style features are weighted features of the candidate style features of each type of text, for example, the product of the candidate style features of the text of the same type and the weight is accumulated, and the final accumulated value is used as the candidate style feature of the text of the same type; or the above-mentioned accumulated value is calculated as the average value, and the average value is used as the candidate style feature. In addition, during the weighted processing, the weights of the text of the same type can be equal, so that each text has the same contribution to the candidate style feature of this type; of course, the weights of each text can also be different, for example, the weight of the text with more strokes is larger, and the weight of the text with fewer strokes is smaller, using different weights to reflect the different contributions to the candidate style feature of this type. Afterwards, the electronic device can obtain the similarity between the user style feature and the candidate style features of each type of text, obtain multiple first principal component style features, for example, sort the similarities in descending order, select the candidate style features starting from the largest similarity, and obtain multiple first principal component style features. Finally, the electronic device may perform weighted processing on the multiple first principal component style features to obtain a reference style feature.
[0176] In step 13, the text content is processed based on the reference style feature to obtain a target text image, and the writing style of the text content of the target text image is the reference style corresponding to the reference style feature.
[0177] In this step, the electronic device can obtain a content encoding model and a content decoding model respectively. Then, the electronic device can obtain a content font image corresponding to the original image, for example, the original image can be used as the content font image; or the electronic device can perform text recognition on the original image to obtain a text recognition result, and obtain a content text image with a preset style based on the above text recognition result. Then, the above content text image is input into the content encoding model to obtain a text feature map output by the content encoding model; thereafter, the electronic device can input the text feature map of each character and its reference style feature into the content decoding model to obtain a target text map output by the content decoding model, wherein the writing style of each character in the target text map is the reference style corresponding to the reference style feature. In other words, the content encoding model can be used to encode the text content in the original image to obtain a text feature map. The content decoding model can be used to decode the above text feature map to obtain a target text map, wherein the writing style of each character in the target text map is the reference style corresponding to the reference style feature.
[0178] In one example, see Table 2, the content coding model includes three layers of deformable convolutional layers (deform.Conv.), one convolutional layer (Conv.), and two layers of residual block layers (Res Block). The content coding model can extract the structural features of the original image and map them to a spatial feature map, namely a text feature map. It can be understood that for the same text, it can produce style-invariant features. For an original image of 80 pixels * 80 pixels, it will be downsampled twice, and the encoded feature size is 20 pixels * 20 pixels, and the number of channels is 256.
[0179] Table 2 Content Coding Model Structure
[0180] In one example, referring to Table 3, the content decoding model includes 2 layers of residual block layers (Res Block) and 4 layers of upsampling convolution layers (Conv.), that is, the text feature map can be upsampled twice to restore the 20 pixel * 20 pixel text feature map to the size of the original image, that is, 80 pixels * 80 pixels.
[0181] Table 3 Content decoding model structure
[0182] In one example, the content encoding model, the content decoding model, and the style encoding model are trained using a generative adversarial approach, and the training architecture is shown in FIG5 .
[0183] The electronic device can use the content encoding model, content decoding model, and style encoding model as a generator, and then train the generator and the discriminator together, including:
[0184] First, the prediction results are generated using a generator consisting of a content encoding model, a content decoding model, and a style encoding model.
[0185] For example, multiple sample images are obtained; each sample image is input into a style encoding model in turn to obtain the user style features output by the style encoding model; reference style features corresponding to the user style features of the text content are obtained; the sample images of each pair of sample images are input into a content encoding model in turn to obtain a text feature map; the text feature map of each text and its reference style features are input into a content decoding model to obtain a predicted text image output by the content decoding model as the prediction result of the generator.
[0186] Referring to FIG. 6, the style encoding model in the generator can obtain the user style features of the text "Dong" and obtain the reference style features according to the above user style features; the content encoding model can obtain the text feature map of the text "Chen"; then the content decoding model can generate the text "Chen" in the target text map by combining the user network features and the text feature map. It can be seen that the text "Chen" in the target text map is similar in style to the text "Dong".
[0187] Then, use the prediction result and the original image to train the discriminator;
[0188] After that, use the trained discriminator to train the generator;
[0189] Finally, when it is detected that the number of training times is less than or equal to the preset number threshold (such as 400 times, adjustable), continue to execute the step of generating the prediction result, that is, continue to train the discriminator and the generator; when it is detected that the number of training times is greater than the preset number threshold, it is determined that the generator has completed training.
[0190] It can be understood that when the generator completes training, the above content encoding model, content decoding model, and style encoding model are synchronously completed, so that the three models cooperate with each other to generate the target text map.
[0191] So far, in the solution provided by the embodiments of the present disclosure, the user style features of the text content included in the original image are obtained; then, the reference style features corresponding to the user style features of the text content are obtained; after that, the text content is processed based on the reference style features to obtain the target text map, and the writing style of the text content in the target text map is the reference style corresponding to the reference style features. In this way, in this embodiment, the reference style features similar to the user style features are used to replace the user style features, which can adjust the text content to the reference style and can also solve problems such as missing strokes, broken strokes, or connected strokes, which is beneficial to improving the recognition experience.
[0192] Based on the text processing method provided by the embodiments of the present disclosure, this embodiment also provides a text processing device. Referring to FIG. 7, the device includes:
[0193] A user style acquisition module 71 for acquiring user style features of text content included in the original image;
[0194] A reference style acquisition module 72 for acquiring reference style features corresponding to the user style features of the text content;
[0195] A target text map acquisition module 73 for processing the text content based on the reference style features to obtain a target text map, and the writing style of the text content in the target text map is the reference style corresponding to the reference style features.
[0196] In one embodiment, the user style acquisition module includes:
[0197] The style model acquisition submodule is used to obtain the style encoding model;
[0198] The user style acquisition submodule is used to input the original image into the style encoding model to obtain the style features output by the style encoding model as the user style features of the text content contained in the original image.
[0199] In one embodiment, the user style acquisition module further includes:
[0200] A first feature determination submodule is configured to determine a user style feature of each character in the text content as the user style feature of the text content;
[0201] The reference style acquisition module includes:
[0202] A first reference feature acquisition submodule, configured to acquire reference style features of user style features of each character in the text content;
[0203] The target text graph acquisition module includes:
[0204] The first text image acquisition submodule is configured to process the text based on the reference style features of each text to obtain a target text image.
[0205] In one embodiment, the user style acquisition module includes:
[0206] The style feature acquisition and determination submodule is further used to obtain user style features of multiple characters in the text content;
[0207] The second feature determination submodule is configured to obtain weighted features of the user style features of the plurality of characters as the user style features of the text content.
[0208] In one embodiment, the reference style acquisition module includes:
[0209] A candidate feature acquisition submodule is used to acquire the first candidate style feature of each type of style text in the candidate style feature library;
[0210] The first candidate style feature acquisition submodule is used for:
[0211] a principal component style feature acquisition submodule, configured to obtain a plurality of first principal component style features based on the similarity between the user style feature and the candidate style features of each type of style text;
[0212] The reference style feature acquisition submodule is used to perform weighted processing on the multiple first principal component style features to obtain the reference style feature.
[0213] In one embodiment, the reference style acquisition module includes:
[0214] A candidate style acquisition submodule is used to extract, based on each character in the text content, second candidate style features of the same character from a candidate style feature library;
[0215] The principal component style acquisition submodule is used to calculate the similarity between the user style feature of each character and each second candidate style feature of the same character, and obtain multiple second principal component style features with the highest similarity value when sorted in descending order;
[0216] The first weighted style acquisition submodule is configured to perform weighted processing on the plurality of second principal component style features to obtain a first weighted style feature.
[0217] In one embodiment, the reference style acquisition module further includes:
[0218] The first reference style acquisition submodule is configured to use the first weighted style feature as a reference style feature for each character.
[0219] In one embodiment, the candidate style feature library is implemented by the following steps:
[0220] Get M texts;
[0221] Get N different types of text styles;
[0222] Obtaining a text image of each of the M characters according to N text styles to obtain M*N text images;
[0223] Inputting each text image into a style encoding model to obtain style features output by the style encoding model;
[0224] The identification codes and style features of the various text images are stored in a style feature library to obtain the candidate style feature library.
[0225] In one embodiment, the first reference style acquisition submodule includes:
[0226] an initial weight obtaining unit, configured to obtain a similarity value corresponding to each of the plurality of first principal component style features as an initial weight;
[0227] a normalized weight acquisition unit, configured to perform normalization processing on the initial weights of the plurality of first principal component style features to obtain a normalized weight of each first principal component style feature;
[0228] The first reference style acquisition unit is configured to calculate a first weighted style feature of the plurality of first principal component style features based on the normalized weights of the respective first principal component style features.
[0229] In one embodiment, the apparatus further comprises:
[0230] A specified style acquisition module is used to obtain specified style features in the configuration data;
[0231] The second reference style acquisition module is configured to perform weighted processing on the designated style feature and the first weighted style feature to obtain a second weighted style feature as a reference style feature for each character.
[0232] In one embodiment, the target text graph acquisition module includes:
[0233] A model acquisition unit, used to respectively acquire a content encoding model and a content decoding model;
[0234] A content image acquisition unit, configured to acquire a content font image corresponding to the original image;
[0235] a text feature acquisition unit, configured to input the content font image into the content coding model to obtain a text feature map output by the content coding model;
[0236] The target text graph acquisition unit is used to input the text feature graph of each character and its reference style feature into the content decoding model to obtain the target text graph output by the content decoding model, wherein the writing style of each character in the target text graph is the reference style corresponding to the reference style feature.
[0237] In one embodiment, the content image acquisition unit includes:
[0238] a first determining subunit, configured to determine that the original image is the content font image;
[0239] or,
[0240] The second determining subunit is configured to perform text recognition on the original image to obtain a text recognition result, and obtain a text image with a preset style according to the text recognition result.
[0241] In one embodiment, the content encoding model, the content decoding model, and the style encoding model are trained using a generative adversarial approach, including:
[0242] Generate prediction results using a generator consisting of a content encoding model, a content decoding model, and a style encoding model;
[0243] Use the prediction results and original images to train the discriminator;
[0244] Training the generator using the trained discriminator;
[0245] In response to detecting that the number of training times is less than or equal to a preset number threshold, continuing to execute the step of generating a prediction result;
[0246] In response to detecting that the number of training times is greater than a preset number threshold, it is determined that the generator has completed training.
[0247] In one embodiment, generating prediction results using a generator composed of a content encoding model, a content decoding model, and a style encoding model includes:
[0248] Get multiple sample images;
[0249] Input each sample image into the style encoding model in turn to obtain the user style features output by the style encoding model;
[0250] Acquire an original image for recognition processing to obtain a content font image corresponding to the original image;
[0251] Acquire a reference style feature corresponding to the user style feature of the text content;
[0252] Input the sample images of each pair of sample images into the content coding model in turn to obtain a text feature map;
[0253] The text feature map of each character and its reference style feature are input into the content decoding model to obtain the predicted text image output by the content decoding model as the prediction result of the generator.
[0254] In one embodiment, the content coding model includes 3 deformable convolution layers, 1 convolution layer and 2 residual block layers.
[0255] In one embodiment, the content decoding model includes 2 residual block layers and 4 upsampling convolution layers.
[0256] It should be noted that the device embodiment shown in this embodiment matches the content of the above-mentioned method embodiment. You can refer to the content of the above-mentioned method embodiment and will not repeat it here.
[0257] In an exemplary embodiment, an electronic device is also provided, comprising
[0258] processor;
[0259] a memory for storing a computer program executable by the processor;
[0260] The processor is configured to execute the computer program in the memory to implement the above method.
[0261] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including an executable computer program. The executable computer program can be executed by a processor to implement the method of the above embodiment. The computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0262] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the disclosure herein. This disclosure is intended to cover any variations, uses, or adaptations that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0263] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A text processing method, characterized in that, The method includes: Obtain the user style features of the text content included in the original image; Obtain the reference style features corresponding to the user style features of the text content; Process the text content based on the reference style features to obtain a target text image, and the writing style of the text content in the target text image is the reference style corresponding to the reference style features.
2. The method according to claim 1, wherein Obtaining the user style features of the text content included in the original image includes: Obtain a style encoding model; Input the original image into the style encoding model, and obtain the style features output by the style encoding model as the user style features of the text content included in the original image.
3. The method according to claim 1 or 2, characterized in that, Obtaining the user style features of the text content included in the original image includes: Determine the user style features of each character in the text content as the user style features of the text content; The obtaining of the reference style features corresponding to the user style features of the text content includes: Obtain the reference style features of the user style features of each character in the text content; The processing of the text content based on the reference style features to obtain a target text image includes: Process the characters based on the reference style features of each character to obtain a target text image.
4. The method according to claim 2, wherein Obtaining the user style features of the text content included in the original image includes: Obtain the user style features of multiple characters in the text content; Obtain the weighted features of the user style features of the multiple characters as the user style features of the text content.
5. The method according to claim 4, wherein Obtaining the reference style features corresponding to the user style features of the text content includes: Obtain the first candidate style features of each type of style character in the candidate style feature library; According to the similarity between the user style features and the first candidate style features of each type of style character, obtain multiple first principal component style features; Perform weighted processing on the multiple first principal component style features to obtain the reference style features.
6. The method according to claim 3, characterized in that Obtaining the reference style features corresponding to the user style features of the text content includes: Extract the second candidate style features of the same character from the candidate style feature library according to each character in the text content; Calculate the similarity values between the user style features of each character and the second candidate style features of the same character of each character, and obtain multiple second principal component style features with higher similarity values when sorted in descending order; Perform weighted processing on the multiple second principal component style features to obtain a first weighted style feature.
7. The method according to claim 6, characterized in that, The method further includes: Use the first weighted style feature as the reference style feature of each character.
8. The method according to claim 5 or 6, characterized in that, The candidate style feature library is implemented by the following steps: Obtain M characters; Obtain N different types of text styles; According to the N text styles, obtain the text images of each character in the M characters to obtain M*N text images; Input each text image into the style encoding model, and obtain the style features output by the style encoding model; Store the identification codes and style features of each text image in the style feature library to obtain the candidate style feature library.
9. The method according to claim 6, characterized in that, Performing weighted processing on the multiple second principal component style features to obtain a first weighted style feature includes: Obtain the similarity values corresponding to each of the multiple second principal component style features as the initial weights respectively; Perform normalization processing on the initial weights of the multiple second principal component style features to obtain the normalized weights of each second principal component style feature; Calculate the first weighted style feature based on the normalized weights of each second principal component style feature.
10. The method according to claim 6, characterized in that The method further includes: Obtain the specified style feature in the configuration data; Perform weighted processing on the specified style feature and the first weighted style feature to obtain the second weighted style feature as the reference style feature of each character.
11. The method according to claim 1, wherein Processing the text content based on the reference style feature to obtain the target text image, including: Obtain the content encoding model and the content decoding model respectively; Obtain the content font image corresponding to the original image; Input the content font image into the content encoding model to obtain the character Feature map output by the content encoding model; Input the character feature map of each character and its reference style feature into the content decoding model to obtain the target text image output by the content decoding model, and the writing style of each character in the target text image is the reference style corresponding to the reference style feature.
12. The method according to claim 11, characterized in that, Obtaining the content font image corresponding to the original image includes: Determine that the original image is the content font image; Or, Perform character recognition on the original image to obtain the character recognition result, and obtain the content character image with a preset style according to the character recognition result.
13. The method according to claim 11, characterized in that, The content encoding model, the content decoding model, and the style encoding model are trained in a generative adversarial manner, including: Use the generator composed of the content encoding model, the content decoding model, and the style encoding model to generate prediction results; Use the prediction results and the original image to train the discriminator; Use the trained discriminator to train the generator; In response to detecting that the number of training times is less than or equal to the preset number threshold, continue to execute the step of generating prediction results; In response to detecting that the number of training times is greater than the preset number threshold, determine that the generator has completed training.
14. The method according to claim 13, characterized in that, Using the generator composed of the content encoding model, the content decoding model, and the style encoding model to generate prediction results, including: Obtain multiple sample images; Input the multiple sample images into the style encoding model in sequence to obtain the user style features output by the style encoding model; Obtain the reference style features corresponding to the user style features of each character; Input the sample images of each pair of sample images into the content encoding model in sequence to obtain the character feature map; Input the character feature map of each character and its reference style feature into the content decoding model to obtain the predicted text image output by the content decoding model as the prediction result of the generator.
15. The method according to claim 13, wherein The content encoding model includes 3 deformable convolutional layers, 1 convolutional layer, and 2 residual block layers.
16. The method according to claim 13, wherein The content decoding model includes 2 residual block layers and 4 upsampling convolutional layers.
17. A text processing device, characterized in that, The device includes: A user style acquisition module for acquiring the user style features of the text content included in the original image; A reference style acquisition module for acquiring the reference style features corresponding to the user style features of the text content; A target text graph acquisition module, which is used to process the text content based on the reference style feature to obtain a target text graph, and the writing style of the text content of the target text graph is the reference style corresponding to the reference style feature.
18. An electronic device, characterized in that, Comprising A processor; A memory for storing computer programs executable by the processor; Wherein, the processor is configured to execute the computer program in the memory to implement the method according to any one of claims 1 to 16.
19. A computer-readable storage medium, characterized in that, When the executable computer program in the storage medium is executed by the processor, the method according to any one of claims 1 to 16 can be implemented.
Citation Information
Patent Citations
Chinese calligraphy character image style migration method and system, and intelligent terminal
CN113393370A
Handwritten image generation method, model training method, device and equipment
CN113516136A
Handwritten text image generation method and device, electronic equipment and storage medium
CN114255159A
Handwritten text image generation method and device, equipment and storage medium
CN114898380A
Font generation method, device, electronic equipment and system
CN115880701A