Image processing method, device and computer program product

By identifying and generating matching text line image style information, the problem of style mismatch in bitmap and vector image editing is solved, and the flexibility and display effect of image text is improved, and it is suitable for a variety of image word processing scenarios.

CN120495467APending Publication Date: 2025-08-15ZHUHAI KINGSOFT OFFICE SOFTWARE +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510493496.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, in bitmap and vector image editing, the style information of the new text line image does not match the style information of the original text line area, resulting in obvious editing traces and affecting the image display effect.

Method used

By identifying the text content and style information of the image to be processed, a matching second text line image is generated, and the editing traces are weakened in combination with image processing technology to generate a high-quality target image.

Benefits of technology

It improves the flexibility and display effect of image text, and is suitable for a variety of image text processing scenarios, such as poster design, image text proofreading and modification, text translation and generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495467A_ABST
    Figure CN120495467A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing method, an image processing device and a computer program product. The method comprises the following steps: in response to a first operation on a first text line image of an image to be processed, obtaining first text content of the first text line image and style information of the first text line image based on the first text line image, and processing the first text content to obtain second text content; generating a second text line image based on the second text content and the style information; and obtaining the first target image based on the to-be-processed image and the second text line image, thereby not only ensuring the image quality of the first target image, but also being suitable for various image word processing scenes, such as poster design, picture word proofreading and modification, word translation and generation, etc. Therefore, the flexibility, the display effect and the intelligent level of editing the image characters are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to, but is not limited to, the field of information technology, and in particular to an image processing method, apparatus, and computer program product. Background Art

[0002] In the text editing scenario, for bitmap images, a new text line image can be generated based on the user's editing operation, and the new text image can be overlaid on the first text line area corresponding to the editing operation to obtain the edited target image. However, this method is prone to the situation where the style information of the new text line image does not match the style information of the first text line area, making the editing traces of the target image obvious, resulting in poor display effect of the target image.

[0003] For vector images, although the user can directly modify the content information of the first text line area, when the background of the first text line area is relatively complex, the editing traces of the edited target image are still relatively obvious. Summary of the Invention

[0004] In order to overcome the problems existing in the related art, the present disclosure provides an image processing method, device and computer program product, which not only ensures the image quality of the first target image, but also can be applied to a variety of image and text processing scenarios, such as poster design, picture and text proofreading and modification, text translation and generation, etc., thereby significantly improving the flexibility, display effect and intelligence level of image and text editing.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, including:

[0006] In response to a first operation on a first text line image of an image to be processed, obtaining first text content and style information of the first text line image based on the first text line image, and processing the first text content to obtain second text content;

[0007] obtaining a second text line image based on the second text content and the style information;

[0008] A first target image is obtained based on the image to be processed and the second text line image.

[0009] According to a second aspect of an embodiment of the present disclosure, there is provided an image processing apparatus, including:

[0010] a first acquisition module configured to, in response to a first operation on a first text line image of an image to be processed, obtain first text content and style information of the first text line image based on the first text line image, and process the first text content to obtain second text content;

[0011] a generating module configured to generate a second text line image based on the second text content and the style information;

[0012] The synthesis module is configured to obtain a first target image based on the image to be processed and the second text line image.

[0013] According to a third aspect of the present disclosure, an electronic device is provided, including:

[0014] processor;

[0015] a memory for storing processor-executable instructions;

[0016] The processor executes the computer program or instructions to implement the steps of any one of the methods in the first aspect above.

[0017] According to a fourth aspect of an embodiment of the present disclosure, a non-temporary computer-readable storage medium is provided, wherein the storage medium stores a computer program or instructions. When the computer program or instructions in the storage medium are executed by a processor, the steps of the method described in any one of the above-mentioned first aspects are implemented.

[0018] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program or instructions. When the computer program or instructions are executed by a processor, the steps of the method described in any one of the first aspects are implemented. The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:

[0019] The technical solution disclosed in the present invention, in response to a first operation on the first text line image of the image to be processed, first determines the first text content of the first text line image and the style information of the first text line image based on the first text line image, so that the user can edit the first text content to meet the user's personalized needs. At the same time, based on the style information of the first text line image, it lays the foundation for reducing the editing traces of the first text line image; and processes the first text content to obtain the second text content to respond to the user's editing operation; then generates the second text line image based on the second text content and the style information, so that the style information of the second text line image matches the style information of the first text line image; finally, obtains the first target image based on the image to be processed and the second text line image, which not only ensures the image quality of the first target image, but also can be applied to a variety of image and text processing scenarios, such as poster design, picture and text proofreading and modification, text translation and generation, etc., thereby significantly improving the flexibility, display effect and intelligence level of editing image text.

[0020] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0022] Figure 1 The process of an image processing method according to an exemplary embodiment is shown as follows Figure 1 .

[0023] Figure 2 The process of an image processing method according to an exemplary embodiment is shown as follows Figure 2 .

[0024] Figure 3 It is a schematic diagram of a framework of an image processing method according to an exemplary embodiment.

[0025] Figure 4 The process of an image processing method according to an exemplary embodiment is shown as follows Figure 3 .

[0026] Figure 5 FIG. 4 is a schematic diagram showing editing an image to be processed according to an exemplary embodiment.

[0027] Figure 6 is a schematic diagram of an image to be processed according to an exemplary embodiment Figure 1 .

[0028] Figure 7 is a schematic diagram of a first target image according to an exemplary embodiment Figure 1 .

[0029] Figure 8 is a schematic diagram of an image to be processed according to an exemplary embodiment Figure 2 .

[0030] Figure 9 is a schematic diagram of a first target image according to an exemplary embodiment Figure 2 .

[0031] Figure 10 The figure is a block diagram of an image processing apparatus according to an exemplary embodiment.

[0032] Figure 11 The figure is a schematic structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0033] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0034] Figure 1 The process of an image processing method according to an exemplary embodiment is shown as follows Figure 1 .like Figure 1 As shown, the method mainly includes the following steps:

[0035] In step 101, in response to a first operation on a first text line image of an image to be processed, first text content and style information of the first text line image are obtained based on the first text line image, and the first text content is processed to obtain second text content;

[0036] In step 102, a second text line image is generated based on the second text content and style information;

[0037] In step 103, a first target image is obtained based on the image to be processed and the second text line image.

[0038] It should be noted that the image processing method proposed in the present disclosure can be applied to electronic devices as well as servers. Here, electronic devices may include: terminal devices, such as mobile terminals or fixed terminals. Mobile terminals may include: mobile phones, tablet computers, laptop computers, etc. Fixed terminals may include: desktop computers, smart TVs, etc. As a type of computer, a server can provide computing or application services to other clients (such as computers, smartphones, and other terminal devices, or even large devices such as train systems) in a network.

[0039] The image processing method in the embodiment of the present disclosure may be configured in an image processing device, and the image processing device may be provided in a server or in an electronic device, which is not limited in the embodiment of the present disclosure.

[0040] It should be noted that the executing entity of the embodiment of the present disclosure may be, for example, a central processing unit (CPU) in a server or electronic device in terms of hardware, or may be, for example, a related background service in a server or electronic device in terms of software, without limitation.

[0041] In some embodiments, the images to be processed can originate from a variety of scenarios, and different methods can be used to obtain the images to be processed in each scenario. For example, if the images to be processed originate from an advertising design scenario, the images to be processed may include designs for promotional posters, flyers, or product packaging; if the images to be processed originate from a text proofreading scenario, the images to be processed may include PDF screenshots or photographed documents; and if the images to be processed originate from a social media scenario, the images to be processed may include social media images, animated posters, or emoticons.

[0042] In some embodiments, after obtaining the image to be processed, text line detection can be performed on the image to be processed to determine the first text line image in the image to be processed. Exemplarily, text line detection can be performed on the image to be processed using a text line detection model. For example, a text line detection model based on deep learning performs text line detection on the image to be processed, wherein the text line detection model may include: a segmentation-based text detection (Differentiable Binarization, DB) algorithm, an instance segmentation-based text detection model (Mask Text Spotter), a Faster R-CNN-based text line detection algorithm (Connectionist Text Proposal Network, CTPN), etc.

[0043] It is understood that by performing text line detection on the image to be processed, the position information of each text line in the image to be processed, including the bounding box coordinates of the text line, can be quickly and accurately detected, thereby obtaining a first text line image and the orientation information of the first text line image in the image to be processed. The first text line image is any text line image in the image to be processed, and the orientation information includes: direction information and position information.

[0044] Here, the first text line image determined from the image to be processed may be one or more. Regardless of whether it is one first text line image or multiple text line images, the first orientation information of each first text line image may be the same or different.

[0045] Exemplarily, all first text line images are set to a set S, and the set S consists of n elements, where each element is a four-tuple, (X1, Y1) represents the upper left corner coordinate of the first text line image, (X2, Y2) represents the upper right corner coordinate of the first text line image, (X3, Y3) represents the lower left corner coordinate of the first text line image, and (X4, Y4) represents the lower right corner coordinate of the first text line image.

[0046] Specifically, the set S can be expressed as:

[0047] S={((X 1i ,Y1i ),(X 2i Y 2i ),(X 3i Y 3i ),(X 4i Y 4i ))|(X,Y)∈R 2 ,i=1,2,3n} (1);

[0048] In formula (1), (X 1i Y 1i ) is the coordinate of the upper left corner of any first text line image in the image to be processed, (X 2i Y 2i ) is the coordinate of the upper right corner of any first text line image in the image to be processed, (X 3i Y 3i ) is the coordinate of the lower left corner of any first text line image in the image to be processed, (X 4i Y 4i ) is the lower right corner coordinate of any first text line image in the image to be processed.

[0049] In some embodiments, in response to a first operation on a first text line image of an image to be processed, text units of the first text line image can be identified through optical character recognition (OCR) technology, and text information can be extracted from the first text line image and converted into an editable text format through this technology to obtain first text content.

[0050] Here, the OCR method is mainly based on deep learning models, such as Convolutional Recurrent Neural Network or Shape Attention Text Recognizer Network.

[0051] In other embodiments, in addition to using OCR technology to recognize text units and obtain recognition results, the recognition results can also be corrected based on the Bidirectional Encoder Representations from Transformers, BERT language model to enhance the processing capabilities of complex character ambiguity and multi-language mixed scenarios, thereby ensuring the comprehensiveness and accuracy of the first text content obtained.

[0052] Here, the first operation is an editing operation on the first text line image, or a triggering operation on the first text line image. For example, if the first operation is an editing operation, the editing operation is a deletion operation or an insertion operation on the first text content of the first text line image; the editing operation is a modification operation on the first text content of the first text line image, for example, modifying a text unit of the first text content or modifying style information of the second text content. In some embodiments, in response to the editing operation, the first text content is processed based on the editing operation to obtain the second text content.

[0053] Taking the first operation as a trigger operation as an example, the trigger operation includes but is not limited to one or more of a click operation, a long press operation, or a sliding operation on the first text line image. In some embodiments, an association relationship is pre-established between different trigger operations and editing operations of the first text content, for example, a click operation corresponds to a deletion operation of the first text content, a long press operation corresponds to an insertion operation of the first text content, or a sliding operation corresponds to a modification operation of the first text content; in response to the trigger operation, the first text content is processed based on the association relationship and the trigger operation to obtain the second text content.

[0054] It is understandable that, in response to the first operation, in addition to obtaining the first text content, style information of the first text line image may also be determined based on the first text line image, so as to reduce editing traces on the image to be processed.

[0055] Here, the style information includes the style information of the first text content and / or the background image of the first text line image. In some embodiments, the style information of the first text line image can be obtained by model recognition. For example, the first text content is recognized by a first neural network model to obtain the style information of the first text content. In another example, the first text line image is processed by a generative adversarial network (GAN) to obtain a background image that does not contain the first text content, and the background image is recognized based on a second neural network model to obtain the style information of the background image.

[0056] In other embodiments, the style information of the first text content can also be obtained by determining the similarity. For example, the feature information of the text unit of the first text content is extracted, and the similarity between the feature information and the preset style information of the font library is determined, and then the preset style information with the maximum similarity is determined as the style information of the first text content.

[0057] In other embodiments, the style information of the background image can also be obtained through image processing technology. For example, color analysis is performed on the first text line image to obtain features such as the color histogram and color space distribution of the background area, thereby obtaining the style information of the background image. In another example, texture analysis is performed on the first text line image to extract texture features, and then the style information of the background image is obtained through texture analysis functions such as Gray-Level Co-occurrence Matrix (GLCM) and Local Binary Patterns (LBP).

[0058] Here, the style information of the first text content includes but is not limited to the font size, font and / or color of the text unit of the first text content, etc. The style information of the background image includes but is not limited to background color, background pattern and the like.

[0059] It can be understood that when the second text content is obtained, a second text line image is generated based on the second text content and the style information of the first text line image, so that the style information of the text content of the second text line image matches the style information of the first text content, and / or the style information of the background image of the second text line image matches the style information of the background image of the first text line image, thereby improving the text editing effect.

[0060] In some embodiments, the background content of the background image can be adjusted based on the style information of the background image, and / or the second text content can be adjusted based on the style information of the first text content; and then a second text line image is generated based on the adjusted background content and / or the adjusted second text content.

[0061] Exemplarily, when the editing operation is a modification operation, the second text content may be adjusted based on the style information of the first text content; and a second text line image may be generated based on the adjusted second text content and the background image of the first text image.

[0062] In another exemplary embodiment, when the editing operation is a deletion operation, the target position of the deleted text unit is determined; the second text content can be adjusted based on the style information of the first text content; and the background content of the target position is adjusted based on the style information of the first background content; finally, based on the adjusted second text content and the adjusted background content, a second text line image is generated.

[0063] Here, when the second text line image is obtained, a first target image may be generated based on the image to be processed and the second text line image.

[0064] In some embodiments, the image to be processed can be preprocessed, for example, by grayscale conversion, edge detection, etc.; the first text line image is matched with the second text line image so as to maintain the size of the area where the first text line image is located in the preprocessed image to be processed, or the size of the area where the first text line image is located is adjusted, for example, the area where the first text line image is located is expanded; the unadjusted or adjusted area where the first text line image is located is extracted from the preprocessed image to be processed; the second text line image is then fused with the image to be processed, and the fused image is optimized through the Patch GAN model to eliminate edge traces, thereby obtaining a first target image.

[0065] In some embodiments, after the fused image is processed by the Patch GAN model, a denoising algorithm (e.g., non-local means denoising, BM3D, etc.) can be used to remove the noise generated during the fusion process; and a sharpening algorithm (e.g., Laplace sharpening, USM sharpening, etc.) can be used to enhance image details, thereby optimizing the visual effect of the first target image.

[0066] In other embodiments, once the first target image is obtained, the user can adjust the first target image as needed to meet personalized requirements. For example, the user can adjust the size, position, or rotation angle of the text line image in the first target image; or, in another example, adjust the size, transparency, or clarity of the first target image.

[0067] The technical solution disclosed in the present invention, in response to a first operation on the first text line image of the image to be processed, first determines the first text content of the first text line image and the style information of the first text line image based on the first text line image, so that the user can edit the first text content to meet the user's personalized needs. At the same time, based on the style information of the first text line image, it lays the foundation for reducing the editing traces of the first text line image; and processes the first text content to obtain the second text content to respond to the user's editing operation; then generates the second text line image based on the second text content and the style information, so that the style information of the second text line image matches the style information of the first text line image; finally, obtains the first target image based on the image to be processed and the second text line image, which not only ensures the image quality of the first target image, but also can be applied to a variety of image and text processing scenarios, such as poster design, picture and text proofreading and modification, text translation and generation, etc., thereby significantly improving the flexibility, display effect and intelligence level of editing image text.

[0068] In some embodiments, obtaining first text content and style information of the first text line image based on the first text line image includes:

[0069] Performing text recognition on the first text line image to obtain first text content;

[0070] Eliminating the first text content to obtain a background image of the first text line image;

[0071] The text units of the first text content are identified to obtain style information of the first text content.

[0072] In an embodiment of the present disclosure, in response to a first operation on the first text line image of the image to be processed, the text unit of the first text line image can be identified through OCR technology, and this technology can extract text information from the first text line image and convert it into an editable text format to obtain the first text content.

[0073] Here, the background image of the first text line image may be obtained through an image inpainting algorithm, or may be obtained through an image generation model.

[0074] Taking the sample-based image restoration algorithm as an example, a sample area similar to the area covered by the first text content is determined, wherein the sample area can be an area in the first text line image that is not covered by the first text content, or an area in an image similar to the first text line image; for each pixel point that needs to be repaired, based on the color, texture, shape and other characteristics of the pixel point, a sample pixel point that matches the pixel point is determined from the sample area; the matched sample pixel point is directly covered to the area covered by the first text content, or the matched sample pixel point is interpolated with the surrounding pixel points of the area covered by the first text content to obtain the background image of the first text line image.

[0075] Taking the example of a generative adversarial network as an image generation model, a generative adversarial network consists of two neural networks: a generator network and a discriminator network. The generator network is used to generate predicted images that are close to the real dataset, while the discriminator network is used to distinguish whether the input image comes from the real dataset or is generated by the generator. When the first text line image is input into the image generation model, the generator network can generate image content that matches the background of the area outside the area based on the pixel information of the surrounding pixels of the area and the global context features, i.e., the background image of the first text line image.

[0076] In the embodiment of the present disclosure, the style information of the first text content can be determined by at least one of the recognition method of the recognition model and the similarity matching method with the font library. Here, the style information of the first text content includes but is not limited to the font size, font or color.

[0077] Exemplarily, the first text content is input into the recognition model, and the font features of the first text content are extracted using the feature extraction network of the recognition model; the font features are recognized using the recognition network of the recognition model to obtain the font and / or color of the first text content.

[0078] In another exemplary embodiment, the style information of the first text content is obtained by determining the similarity. For example, the feature information of the text unit of the first text content is extracted, and the similarity between the feature information and the preset fonts of the font library is determined, and then the preset font with the maximum similarity is determined as the font of the first text content.

[0079] In some embodiments, when the style information of the first text content and the second text content are obtained, the target text content can be further determined. Here, the target text content is the text content of the second text line image. Different editing operations have different ways of generating the target text content. For example, when the editing operation indicates to modify the style information of the first text content, the target text content can be determined directly based on the second text content; for another example, when the editing operation indicates to insert a text unit in the first text content, the first text content is processed to obtain the second text content and the second text content is processed based on the style information of the first text content to obtain the target text content.

[0080] Therefore, the target text content can be determined based on the style information of the first text content and the second text content; or, the target text content can be determined based on the second text content to flexibly adapt to the user's editing operations, thereby improving the intelligence level of editing image text.

[0081] In some embodiments, when the editing operation is a modification operation, and the modification operation indicates that the first text unit of the first text content is updated to the second text unit (for example, adjusting "weather" to "temperature"), the second text content can be obtained first; and the second text content is processed based on the style information of the first text content to generate the target text content. When the modification operation indicates that the style information of the first text content is adjusted from the first style information to the second style information (for example, adjusting the first text content from non-italic to italic, adjusting the first text content from bold to Songti, or adjusting the first text content from size 4 font to size 6 font), the second text content can be obtained first and determined as the target text content.

[0082] In other embodiments, when the editing operation is a deletion operation, since no text unit is added to the first text content, there is no need to process the second text content based on the style information of the first text content, and the second text content is directly determined as the target text content.

[0083] In other embodiments, when the editing operation is an insert operation, since a text unit is added to the first text content, the second text content can be processed based on the style information of the first text content to generate the target text content.

[0084] Through the technical solution disclosed in the present invention, the background image of the first text line image and the style information of the first text content are obtained, which can meet the user's needs for flexible adjustment of the text content, and is conducive to reducing the editing traces of the image to be processed and improving the image quality of the first target image.

[0085] In some embodiments, processing the first text content to obtain the second text content includes:

[0086] In response to a triggering operation on the first text content, displaying the first text content at the editing location according to the reference style information;

[0087] In response to an editing operation on the first text content at the editing location, second text content is obtained.

[0088] It should be noted that, in order to enhance the user's editing experience, when a trigger operation on the first text content is detected, the first text content can be displayed in the editing area according to the reference style information, so that the user can edit the first text content based on the editing area. The trigger operation includes but is not limited to one or more of a click operation, a slide operation, or a long press operation; the editing area can be understood as the position designated by the cursor on the display screen, or any editing position in the edit box on the display screen.

[0089] Here, the reference style information and the style information of the first text content may be the same or different.

[0090] In some embodiments, considering the user's need to determine the style information of the first text content, the first text content can be displayed at the editing location according to the style information of the first text content, so that the user can determine the style information of the first text content before editing the first text content.

[0091] In other embodiments, considering that the font size indicated by the style information of the first text content may be small, if the first text content is displayed according to the style information of the first text content during editing, the user may not be able to accurately identify the first text content, thereby hindering the user's editing experience. Therefore, reference style information that is easy for the human eye to see can be pre-set, and the first text content can be displayed according to the reference style information.

[0092] In the disclosed embodiment, if the editing operation is an insert operation, the text unit indicated by the insert operation is inserted into the first text content at the location that triggered the insert operation, thereby obtaining the second text content. If the editing operation is a modify operation, the text unit indicated by the modify operation in the first text content is modified, thereby obtaining the second text content. If the editing operation is a delete operation, the text unit indicated by the delete operation in the first text content is deleted, thereby obtaining the second text content.

[0093] In some embodiments, in response to a triggering operation on the second text content, the target text content is determined based on the style information of the first text content and / or the second text content. Here, the triggering operation on the second text content is used to indicate the completion of the editing process of the first text content. Exemplarily, the user can trigger the triggering operation on the second text content by clicking on the first text line image; in another exemplary embodiment, the user can trigger the triggering operation on the second text content by clicking or sliding any area in the image to be processed.

[0094] In some embodiments, when the reference style information is the same as the style information of the first text content and a triggering operation on the second text content is detected, the target text content is determined based on the second text content.

[0095] In other embodiments, when the reference style information is different from the style information of the first text content and a trigger operation on the second content is detected, the target text content is determined according to the type of editing operation based on the style information of the first text content and the second text content.

[0096] Through the technical solution disclosed in the present invention, in response to a triggering operation on the first text content, the first text content can be displayed at the editing location according to the reference style information, so that the user can edit the first text content based on the editing location, thereby improving the user's editing experience; and when the editing operation is detected, the first text content is processed to obtain the second text content, which is conducive to improving the flexibility and intelligence of text editing.

[0097] In some embodiments, when the style information includes style information of the first text content and a background image, obtaining the second text line image based on the second text content and the style information includes:

[0098] When the editing operation is an insert operation or a delete operation, adjusting the background content of the background image using the style information of the background image, and generating a second text line image based on the second text content, the style information of the first text content, and the adjusted background image;

[0099] In a case where the editing operation is a modification operation, a second text line image is generated based on the second text content, the style information of the first text content, and the background image.

[0100] It can be understood that when the editing operation is an insert operation or a delete operation, it is determined that the background content of the first text line image will change during the editing operation. In order to ensure that the style information of the second text line image matches the style information of the first text line image, the background content of the background image can be adjusted based on the style information of the background image, and the second text line image can be generated based on the target text content and the adjusted background image.

[0101] Here, the background content of the background image refers to the background color (e.g., solid color or gradient color) and / or the background elements (e.g., pattern or wood grain). The style information of the background image is used to control the display of the background color and / or the layout of the elements.

[0102] In the case where the editing operation is a modification operation, it is determined that the background content of the first text line image will not change during the editing operation. In order to save processing resources, the second text line image can be directly generated based on the target text content and the background image.

[0103] Through the technical solution disclosed in the present invention, a second text line image can be flexibly generated based on the type of editing operation, which not only matches the style information of the second text line image with the style information of the first text line image, but also reduces invalid processing processes and saves processing resources. Therefore, while reducing the editing traces of the image to be processed, it can also improve the intelligence of the text editing process.

[0104] In some embodiments, when the editing operation is an insert operation or a delete operation, adjusting the background content of the background image using the style information of the background image includes:

[0105] In a case where the insert operation indicates inserting the first text unit into the first text content, moving the text unit of the first text content based on the insertion position of the first text unit;

[0106] Adjust the background content at the insertion location based on the style information of the background image.

[0107] In an embodiment of the present disclosure, when an insertion operation is detected, the insertion position of the first text content can be determined first, and then based on the insertion position, the text unit of the first text content can be moved to facilitate determining the background content that needs to be adjusted, thereby improving the intelligence of the text insertion process.

[0108] Here, the number of text units moved in the first text content matches the number of first text units inserted. The text units after the insertion position in the first text content can be moved, and the text units before the insertion position in the first text content can also be moved, which is not limited in the present embodiment.

[0109] In some embodiments, the background content of the insertion position may be filled with the style information of the background image to ensure the consistency and coherence of the background image, so that the first target image after the text is inserted looks more natural and coordinated.

[0110] Through the technical solution disclosed in the present invention, when the insertion position of the insertion operation is determined, the text unit of the first text content is moved based on the insertion position to determine the background content that needs to be adjusted, thereby reducing the insertion traces of the first text line image and improving the image quality of the first target image.

[0111] In some embodiments, when the editing operation is an insert operation or a delete operation, adjusting the background content of the background image using the style information of the background image includes:

[0112] In a case where the deletion operation indicates deletion of the second text unit in the first text content, moving the text units in the first text content except the second text unit based on the target position of the second text unit;

[0113] Based on the style information of the background image, the background content of the target location is adjusted.

[0114] It can be understood that when a deletion operation is detected, the text units in the first text content except the second text unit can be moved based on the target position, and the background content of the target position of the second text unit indicated by the deletion operation can be determined as the background content that needs to be adjusted, thereby reducing the deletion traces of the first text line image.

[0115] Here, the number of text units moved in the first text content matches the number of deleted second text units. The text units after the second text unit in the first text content can be moved, or the text units before the second text unit in the first text content can be moved, which is not limited in the present embodiment.

[0116] In some embodiments, the background content of the target location can be filled with the style information of the background image to ensure the consistency and coherence of the background image, so that the first target image after deleting the text looks more natural and coordinated.

[0117] Through the technical solution disclosed in the present invention, when the target position of the second text unit indicated by the deletion operation is determined, the text units except the second text unit are moved and adjusted, and the background content of the target position is adjusted, thereby reducing the deletion traces of the first text line image and improving the image quality of the first target image.

[0118] In some embodiments, when the style information includes style information of the first text content and a background image, obtaining the second text line image based on the second text content and the style information includes:

[0119] Determining a degree of matching between style information of the first text content and preset style information of the font library, and determining a target rendering strategy based on the degree of matching;

[0120] Rendering the second text content based on the target rendering strategy to obtain a third text line image;

[0121] Background segmentation processing is performed on the third text line image, and a second text line image is obtained based on the background image and the processed third text line image.

[0122] In some embodiments, a font library is a data set containing multiple font styles and text units, allowing users to select and use different fonts in various applications. Each font in the font library is composed of a set of vector graphics or bitmap images that define the shape and appearance of the text unit.

[0123] Therefore, the target rendering strategy of the second text content can be determined by the matching degree between the style information of the first text content and the preset style information of the font library, thereby improving the accuracy and efficiency of rendering the second text content.

[0124] Taking the font of a text unit as an example, the first feature information of the font of the first text content and the second feature information of a preset font in the font library are extracted, wherein the feature information includes outline shape, stroke thickness, or character spacing, etc.; the extracted first feature information and the second feature information are quantified, and the similarity between the quantized first feature information and the second feature information is calculated, for example, using a measurement method such as cosine similarity and Euclidean distance to obtain a comparison result; and based on the comparison result, the degree of match between the font of the first text content and the preset font in the font library is determined. Here, in order to improve the accuracy of determining the comparison result, a third threshold value can be set in advance. When the similarity is determined, the similarity is compared with the third threshold value to obtain a comparison result.

[0125] Here, the target rendering strategy includes: a rendering strategy based on a font library or a processing strategy based on a preset model.

[0126] When the style information of the first text content has a high degree of match with the preset style information of the font library, the second text content can be rendered based on the rendering strategy of the preset style information of the font library to obtain a third text line image; and when the style information of the first text content has a low degree of match with the preset style information of the font library, the second text content can be processed based on the preset model to obtain a third text line image.

[0127] It should be noted that, considering that the style information of the background content of the third text line image does not match the background image of the first text line image, the editing traces of the first text line image will be obvious. Therefore, the third text line image can be subjected to background segmentation processing to obtain the text content of the third text line image, so that the text content and the background image can be integrated, thereby reducing the editing traces of the first text line image.

[0128] Here, background segmentation can be implemented by a threshold-based segmentation method, an edge-based segmentation method, or a model-based segmentation method, which is not limited in the embodiments of the present disclosure.

[0129] For example, if the contrast between the text content of the third text line image and the background image is high, a second threshold can be set, and pixels with pixel values above or below the second threshold are classified as foreground and background, respectively, to obtain the text content of the third text line image. In another example, a classifier (e.g., a support vector machine, a neural network, etc.) is trained using the labeled foreground and background images, and the trained classifier is then used to perform pixel-level classification on the third text line image to obtain the text content of the third text line image.

[0130] Through the technical solution disclosed in the present invention, the target rendering strategy of the second text content is determined based on the matching degree between the style information of the first text content and the preset style information of the font library, thereby improving the accuracy and efficiency of the target rendering strategy, thereby improving the accuracy of obtaining the third text line image; and the background style processing is performed on the third text line image to obtain the text foreground, so that the text foreground and the background image can be integrated, thereby reducing the editing traces of the first text line image.

[0131] In some embodiments, determining a target rendering strategy based on the matching degree includes:

[0132] If the degree of matching between the target style information in the font library and the style information of the first text content is greater than or equal to a first threshold, determining the rendering strategy of the target style information as the target rendering strategy;

[0133] When the matching degree between the preset style information in the font library and the style information of the first text content is less than a first threshold, the processing strategy of the first preset model is determined as the target rendering strategy.

[0134] It can be understood that in order to determine the efficiency of the target rendering strategy, a first threshold can be preset, and when determining the matching degree between the style information of the first text content and the preset style information of the font library, the matching degree is compared with the first threshold to obtain a comparison result to further determine the target rendering strategy.

[0135] Here, the first threshold can be set arbitrarily according to requirements, such as 90%, which is not limited in the embodiment of the present disclosure.

[0136] In an embodiment of the present disclosure, when the comparison result indicates that the target style information matches the style information of the first text content to a preset style information greater than or equal to a first threshold, the rendering strategy of the target style information can be determined as the target rendering strategy, thereby ensuring the rendering accuracy of the second text content.

[0137] Taking the style information as a font as an example, the matched target font is read from the font library; the graphics library or rendering engine can read the meta information of the target font, such as the font outline, shape, size, and inclination angle; based on the meta information, the graphics library or rendering engine converts the second text content into pixel data, thereby obtaining a third text line image.

[0138] In some embodiments, if at least two target style information in the font library match the style information of the first text content at a degree greater than or equal to a preset style information threshold, the rendering strategy of the target style information with the greatest degree of match among the target style information can be determined as the target rendering strategy. For example, if the degree of match between the font of the first text content and both Songti and Xin Songti in the font library is greater than the first threshold, the font with the greater degree of match between Songti and Xin Songti will be determined as the target font.

[0139] When the comparison result indicates that the matching degree between the preset style information in the font library and the style information of the first text content is less than the first threshold, it is determined that the font library-based rendering method cannot meet the requirements, and the processing strategy of the first preset model is determined as the target rendering strategy. In this way, even if the font of the first text content is a non-standard font, for example, handwriting, multi-texture font or special custom font, the font style of the first text content can be flexibly restored, thereby improving the generation ability of non-standard fonts and reducing the editing traces of the first text line image.

[0140] Through the technical solution disclosed in the present invention, a target rendering strategy based on font library rendering and first preset model processing is pre-set, and based on the matching degree between the style information of the first text content and the preset style information of the font library and the first threshold, the target rendering strategy is flexibly determined, thereby improving the flexibility and intelligence of text generation.

[0141] In some embodiments, when the target rendering strategy is the processing strategy of the first preset model, rendering the second text content based on the target rendering strategy to obtain a third text line image includes:

[0142] The style information of the first text content and the second text content are input into the first preset model, and the style information of the first text content is applied to the second text content using the generator network of the first preset model to obtain a third text line image.

[0143] It should be noted that a first preset model for font style transfer can be pre-trained, and the style information of the first text content and the second text content can be input into the first preset model. The first preset model then processes the second text content to generate the third text line image. In this way, even if the font of the first text content is non-standard, the font style label of the first text content can be applied to the second text content, improving the accuracy of the third text line image.

[0144] Here, the first preset model includes a generator network and a discriminator network.

[0145] The training process of the first preset model is specifically as follows: obtaining training data, wherein the training data includes: a source font image and a real image of the target font, and the size of the input image of the first preset model can be unified as 256×256; a deep neural network can be determined as a generator network, which is used to transfer the font style of the target font to the source font image to obtain a generated image of the target font, wherein the structure of the generator network includes a downsampling layer or an upsampling layer, etc.; a two-classification network is determined as a discriminator network, which is used to distinguish whether the input image (i.e., the image input by the generator) is a real image of the target font or a generated image generated by the generator network; the loss function of the generator network measures the difference between the real image and the generated image, and the loss function of the discriminator network measures the accuracy of the discriminator network classification. During the actual training process, the generator network takes the source font image and the target font as conditional inputs, uses the downsampling layer to extract shallow font features, and restores the feature vector to an image through the upsampling layer to generate a generated image of the target font in order to imitate the font style of the target font to "cheat" the discriminator. At the same time, the discriminator is used for adversarial training. With the help of the idea of adversarial network training, the network parameters of the generator network are optimized until the preset convergence conditions are reached.

[0146] For example, the difference between the real image and the generated image is obtained based on the style loss function, and the formula of the style loss function is:

[0147] L style (G)=||S(G(x))-S(y)||2 (2);

[0148] In formula (2), G(x) is the generated image, S is the style extraction function, y is the font of the first text content, and || ||2 represents the Euclidean distance.

[0149] Through the technical solution disclosed in the present invention, the style information of the first text content can be used in the second text content through the first preset model, so that even if the font of the first text content is a non-standard font, the accuracy of the third text line image can be improved, thereby improving the flexibility and intelligence of text generation.

[0150] In some embodiments, identifying text units of the first text content to obtain font style information of the first text content includes:

[0151] In a case where the style information of the first text content includes a font size, determining the number of text units in the first text content and the size of the first text line image, and determining the font size based on the number of text units and the size;

[0152] In the case where the style information of the first text content includes font and / or color, the first text line image is input into the recognition model, the font features of the first text content are extracted using the feature extraction network of the recognition model, and the font features are recognized using the recognition network of the recognition model to obtain the font and / or color.

[0153] Here, the font size represents the size of the text unit and is used to quantify the size of the text unit.

[0154] In the disclosed embodiment, when the first text content is recognized, each text unit in the first text content is traversed and counted to obtain the number of text units in the first text content. Simultaneously, text line detection is performed on the image to be processed to obtain the bounding box coordinates of the first text line, thereby further determining the size of the first text line image.

[0155] It can be understood that based on the number of text units and the pixel width of the first text line image, the average width of each text unit can be obtained; by comparing the average width with the known reference font size, the font size of the text unit can be determined, thereby improving the efficiency of determining the font size of the text unit.

[0156] For example, Figure 2 The process of an image processing method according to an exemplary embodiment is shown as follows Figure 2 ,like Figure 2 As shown, the image processing method includes:

[0157] In step 201 , text recognition is performed on a first text line image in an image to be processed to obtain first text content of the first text line image.

[0158] In step 202 , text correction is performed on the first text content.

[0159] In step 203 , the number of text units in the first text content is determined.

[0160] In step 204 , the size of the first text line image is determined.

[0161] In step 205 , based on the equal division strategy, the average width of each text unit is determined.

[0162] In step 206, the font size of the text unit is determined.

[0163] In the disclosed embodiments, the feature extraction network of the recognition model can be a convolutional neural network (CNN) or a transformer model, etc.; the feature extraction network includes at least a convolutional layer and a pooling layer. The number of convolutional layers and pooling layers in the feature extraction network can be arbitrarily set according to requirements and is not limited thereto.

[0164] It is understood that the recognition model is used to identify the font and / or color of a text unit. When the recognition model is used to identify the font of a text unit, the font features extracted by the feature extraction network may include: outline shape, stroke thickness, or character spacing; the recognition network of the recognition model is used to identify these font features to determine the font of the text unit.

[0165] When the recognition model is used to identify the color of a text unit, the font features extracted by the feature extraction network may include color features such as color histograms or color moments. The recognition network of the recognition model is used to identify the above color features to determine the color of the text unit.

[0166] Taking the recognition model used to identify fonts as an example, a large number of text images are collected in advance as sample images, where the sample images include text content in different fonts; each sample image is manually annotated and stored in a database to obtain the true label of the font of each sample image; the true label and sample image are input into the recognition model, and the feature extraction network of the recognition model is used to extract the font features of the sample image; and the recognition network of the recognition model is used to identify the font features of the sample image to obtain a predicted label; based on the true label and the predicted label, the model parameters of the recognition model are adjusted until the adjusted recognition model meets the preset convergence conditions. In some embodiments, in order to improve the robustness of the model, the text image can be preprocessed, for example, size normalization, or enhancement processing such as rotation, scaling, and cropping, and the preprocessed text image is then used as the sample image.

[0167] Through the technical solution disclosed in the present invention, on the one hand, based on the number of text units and the size of the first text line image, the accuracy and efficiency of determining the font size are improved, thereby improving the matching accuracy with the preset style information of the font library; on the other hand, based on the recognition model, the first text line image is recognized, thereby improving the accuracy of determining the font and / or color of the text unit of the first text content, thereby improving the matching accuracy with the preset style information of the font library.

[0168] In some embodiments, using a recognition network of a recognition model to recognize font features to obtain a font includes:

[0169] Using the first recognition network, identifying the font features to obtain the target language category of the first text content;

[0170] The second recognition network determined based on the target language category is used to recognize the font features to obtain the font.

[0171] Here, the first recognition network is a language classifier, used to map the extracted font features to language categories. For CNN models, the first recognition network can be a fully connected layer; for Vision Transformer (ViT) models, the first recognition network can be a classification head.

[0172] In the embodiment of the present disclosure, the font features obtained by the feature extraction network are input into the first recognition network, and the first recognition network recognizes the font features to obtain the target language category of the text unit.

[0173] In some embodiments, a language classifier can be trained using an image dataset containing text in multiple languages. Sample images in the image dataset are first labeled to obtain true labels for the sample images. Feature extraction is then performed on the sample images using a feature extraction network to obtain font features for the sample images. These font features are then input into a first recognition network to obtain predicted labels for the sample images. Based on the difference between the predicted labels and the true labels, the network parameters of the first recognition network are adjusted until a predetermined convergence condition is reached.

[0174] Here, different language categories correspond to different second recognition networks. Recognizing font features through the second recognition network of the target language category can improve the accuracy of font recognition.

[0175] Through the technical solution disclosed in the present invention, the target language category of the first text content is first obtained through the first recognition network, and then the font is obtained based on the second recognition network of the target language category, thereby improving the recognition accuracy of the font.

[0176] In some embodiments, performing elimination processing on the first text content to obtain a background image of the first text line image includes:

[0177] The first text line image is input into the second preset model, and the content features of the first text line image are eliminated using the generator network of the second preset model to obtain a background image.

[0178] It can be understood that in order to improve the efficiency of obtaining the background image, a second preset model can be pre-trained, and when the first text line image is obtained, the first text line image is input into the second preset model, and the second preset model processes the first text line image to obtain the background image.

[0179] Here, a second preset model is constructed based on a text elimination algorithm (such as TextFuseNet or EraseNet), and the second preset model includes a generator network and a discriminator network.

[0180] The training process of the second preset model is specifically as follows: obtaining an image dataset containing text and a corresponding text-free background, wherein these images can cover a variety of acquisition scenarios, lighting conditions, or text styles to ensure the generalization ability of the model; a deep neural network can be determined as a generator network to generate text-free background images from random noise, wherein the structure of the generator network includes convolutional layers, deconvolution layers, or upsampling layers, etc.; a binary classification network is determined as a discriminator network to distinguish whether the input image is a real text-free background image or a fake image generated by the generator network; the loss function of the generator network measures the difference between the generated image and the real image, and the loss function of the discriminator network measures the accuracy of the discriminator network classification. In the actual training process, the generator network is first fixed, and the discriminator network is trained to distinguish between real images and generated images; then the discriminator is fixed, and the generator is trained to generate more realistic background images to "fool" the discriminator; the above two steps are performed alternately until the preset convergence condition is reached.

[0181] For example, Figure 3 is a schematic diagram of a framework of an image processing method according to an exemplary embodiment. Figure 3As shown, by inputting the first text line image 30 into the recognition model 31, the language category of the first text content can be obtained, for example, Chinese 33, English 34, or Japanese 35. Each language category corresponds to a second recognition network. For example, Chinese 33 corresponds to CNN network 1 (i.e., Figure 37), English 34 corresponds to CNN network 2 (i.e., Figure 38), and Japanese 35 corresponds to CNN network 3 (i.e., Figure 39). Based on the second recognition network of the target language category, the font of the text unit can be determined. At the same time, by inputting the first text line image 30 into the recognition model 31, the color 36 of the text unit can also be obtained. By inputting the first text line image 30 into the second preset model 32, the background image 40 of the first text line image can be obtained. In this way, through the above structure, the style information of the first text content and the background image of the first text line image can be accurately obtained, which is conducive to reducing the editing traces of the first text line image during the subsequent editing process.

[0182] In the disclosed embodiment, the background image is obtained through the second generation model, which improves the accuracy and efficiency of the background image, thereby laying a foundation for reducing the editing traces of the first text line image.

[0183] In some embodiments, the method further comprises:

[0184] In response to a second operation on the first target image, outputting a reference adjustment template; wherein the reference adjustment template is determined based on historical adjustment data;

[0185] In response to an application instruction for the reference adjustment template, a second target image is obtained based on the first target image and the reference adjustment template.

[0186] It is understandable that in order to improve the user's editing experience, the user's historical adjustment data can be recorded, and a reference adjustment model can be generated based on the historical adjustment data; and when the second target image is obtained and the second operation on the second target image is detected, the reference adjustment template can be output in a timely manner, thereby improving the efficiency of adjusting the first target image while meeting the user's personalized needs.

[0187] Here, the second operation includes but is not limited to one or more of a click operation, a long press operation, or a slide operation.

[0188] Here, the historical adjustment data includes but is not limited to size, transparency, or clarity.

[0189] In the embodiment of the present disclosure, when an application instruction for a reference adjustment template is detected, a second target image can be obtained based on the first target image and the reference adjustment template. In this way, for image and text editing scenarios with large amounts of data, such as poster design or text proofreading, the processing efficiency of the edited image can also be improved, thereby improving the user's editing experience.

[0190] Figure 4 The process of an image processing method according to an exemplary embodiment is shown as follows Figure 3 ,like Figure 4 As shown, the method mainly includes the following steps:

[0191] In step 401 , in response to a first operation on a first text line image of an image to be processed, text recognition is performed on the first text line image to obtain first text content.

[0192] In step 402, text units of the first text content are identified to obtain style information of the first text content.

[0193] In some embodiments, when the style information of the first text content includes font size, the font size is determined based on the number and size of the text units in the first text content and the size of the first text line image.

[0194] In other embodiments, when the style information of the first text content includes font and / or color, the first text line image can be input into a recognition model, and the font features of the first text content can be extracted using the feature extraction network of the recognition model; the font features can be recognized using the recognition network of the recognition model to obtain the color of the font and / or text unit.

[0195] In step 403, the first text content is eliminated to obtain a background image of the first text line image.

[0196] In some embodiments, the first text line image can be input into a second preset model, and the content features of the first text line image can be eliminated using the generator network of the second preset model to obtain a background image.

[0197] In step 404, the first text content is processed to obtain second text content.

[0198] In some embodiments, in response to a triggering operation on the first text content, the first text content is displayed at the editing location of the editing box according to the reference style information; in response to an editing operation on the first text content at the editing location, the second text content is obtained.

[0199] For example, Figure 5 FIG. 1 is a schematic diagram showing editing of an image to be processed according to an exemplary embodiment. Figure 5As shown, image 50 to be processed is a computer advertisement. The first text content of first text line image 51 is "I5 processor," and the first text content of first text line image 52 is "Full HD eye protection screen." When a trigger operation is detected for the first text content of first text line image 51, the first text content of first text line image 51 is displayed in edit box 53 according to reference style information. The font indicated by the reference style information is different from the font indicated by the style information of the first text content, and the font size indicated by the reference style information is also different from the font size indicated by the style information of the first text content.

[0200] In step 405 , the degree of matching between the style information of the first text content and the preset style information of the font library is determined.

[0201] In step 406 , it is determined whether there is a matching degree greater than or equal to a first threshold.

[0202] In some embodiments, if it is determined that there is a matching degree greater than or equal to the first threshold, step 407 is performed.

[0203] In some embodiments, if it is determined that there is no matching degree greater than the first threshold, step 408 is performed.

[0204] In step 407 , the rendering strategy of the target style information is determined as the target rendering strategy.

[0205] In step 408 , the processing strategy of the first preset model is determined as the target rendering strategy.

[0206] In step 409 , the second text content is rendered based on the target rendering strategy to obtain a third text line image.

[0207] In some embodiments, when the rendering strategy of the target style information is determined as the target rendering strategy, the matched target font is read from the font library; the graphics library or rendering engine can read the meta information of the target font, such as the font outline, shape, size, and inclination angle; based on the meta information, the graphics library or rendering engine converts the second text content into pixel data, thereby obtaining a third text line image.

[0208] In other embodiments, when the target rendering strategy is the first preset model, the style information of the first text content and the second text content can be input into the first preset model, and the generator network of the first preset model can be used to apply the style information of the first text content to the second text content to obtain a third text line image.

[0209] In step 410 , background segmentation processing is performed on the third text line image.

[0210] In step 411 , a second text line image is generated based on the processed third text line image and the background image.

[0211] In step 412 , a first target image is obtained based on the image to be processed and the second text line image.

[0212] In some embodiments, in response to a second operation on the first target image, a reference adjustment template is output; wherein, the reference adjustment template is determined based on historical adjustment data; in response to an application instruction for the reference adjustment template, a second target image is obtained based on the first target image and the reference adjustment template.

[0213] For example, Figure 6 is a schematic diagram of an image to be processed according to an exemplary embodiment Figure 1 , Figure 7 is a schematic diagram of a first target image according to an exemplary embodiment Figure 1 ,like Figure 6-7 As shown, the first text content of the first text line image 61 in the image to be processed 60 is "Self-service Hair Dryer Precautions", and the first text content of the first text line image 62 is "Please follow the precautions." In response to the modification operation instruction, the "self-service" in the first text content of the first text line image 61 is modified to "others", and based on the style information of the first text content and "others", the second text line image 71 is obtained. At the same time, in response to the deletion operation instruction, the "note" in the first text content of the first text line image 62 is deleted, and based on the style information of the background image, the background content of the "note" is adjusted; and based on the location of the "note", the text units in the first text content except the "note" are moved.

[0214] In response to an input operation indicating the insertion of the word "whoever" into the first text content of first text line image 62, the text unit of the first text content is moved based on the insertion position, and the background content at the insertion position is adjusted based on the style information of the background image. Based on the style information of the first text content and the word "whoever", a second text line image 72 is obtained. Based on second text line image 71, second text line image 72, and image to be processed 60, a first target image 70 is obtained.

[0215] As another example, Figure 8 is a schematic diagram of an image to be processed according to an exemplary embodiment Figure 2 , Figure 9 is a schematic diagram of a first target image according to an exemplary embodiment Figure 2 ,like Figure 8-9As shown, the first text content of the first text line image 81 in the image to be processed 80 is "-900.00". In response to the modification operation instruction, "9" is modified to "1", and the insertion operation instruction inserts "0000" into the first text content. Based on the insertion position, the text unit of the first text content is moved; and based on the style information of the background image, the background content of the insertion position is adjusted, and then based on the style information of the first text content and "-1000000.00", the second text line image 91 is obtained. Finally, based on the second text line image 91 and the image to be processed 80, the first target image 90 is obtained.

[0216] The technical solution disclosed in the present invention, in response to a first operation on the first text line image of the image to be processed, first determines the first text content of the first text line image and the style information of the first text line image based on the first text line image, so that the user can edit the first text content to meet the user's personalized needs. At the same time, based on the style information of the first text line image, it lays the foundation for reducing the editing traces of the first text line image; and processes the first text content to obtain the second text content to respond to the user's editing operation; then generates the second text line image based on the second text content and the style information, so that the style information of the second text line image matches the style information of the first text line image; finally, obtains the first target image based on the image to be processed and the second text line image, which not only ensures the image quality of the first target image, but also can be applied to a variety of image and text processing scenarios, such as poster design, picture and text proofreading and modification, text translation and generation, etc., thereby significantly improving the flexibility, display effect and intelligence level of editing image text.

[0217] Figure 10 FIG. 1 is a block diagram of an image processing apparatus according to an exemplary embodiment. Figure 10 As shown, the image processing device 1000 mainly includes:

[0218] A first acquisition module 1001 is configured to, in response to a first operation on a first text line image of an image to be processed, obtain first text content and style information of the first text line image based on the first text line image, and process the first text content to obtain second text content;

[0219] A generating module 1002 is configured to generate a second text line image based on the second text content and the style information of the first text line image;

[0220] The synthesis module 1003 is configured to obtain a first target image based on the image to be processed and the second text line image.

[0221] In some embodiments, the first acquisition module 1001 is specifically configured to:

[0222] Performing text recognition on the first text line image to obtain the first text content;

[0223] performing elimination processing on the first text content to obtain a background image of the first text line image;

[0224] The text units of the first text content are identified to obtain style information of the first text content.

[0225] In some embodiments, the first acquisition module 1001 is further configured to:

[0226] In response to a triggering operation on the first text content, displaying the first text content at an editing location according to the reference style information;

[0227] In response to an editing operation on the first text content at the editing location, the second text content is obtained.

[0228] In some embodiments, when the style information includes style information of the first text content and a background image, the generating module 1002 includes:

[0229] a first processing module configured to, when the editing operation is an insert operation or a delete operation, adjust the background content of the background image using the style information of the background image, and generate the second text line image based on the second text content, the style information of the first text content, and the adjusted background image;

[0230] The second processing module is configured to generate the second text line image based on the second text content, the style information of the first text content and the background image when the editing operation is a modification operation.

[0231] In some embodiments, the first processing module is specifically configured to:

[0232] In a case where the insert operation indicates inserting a first text unit into the first text content, moving the text unit of the first text content based on the insertion position of the first text unit;

[0233] Based on the style information of the background image, the background content of the insertion position is adjusted.

[0234] In some embodiments, the first processing module is further configured to:

[0235] In a case where the deletion operation indicates deletion of the second text unit in the first text content, moving the text units in the first text content except the second text unit based on the target position of the second text unit;

[0236] Based on the style information of the background image, the background content of the target location is adjusted.

[0237] In some embodiments, when the style information includes style information of the first text content and a background image, the generating module 1002 is specifically configured to:

[0238] Determining a degree of matching between style information of the first text content and preset style information of a font library, and determining a target rendering strategy based on the degree of matching;

[0239] Rendering the second text content based on the target rendering strategy to obtain a third text line image;

[0240] Background segmentation processing is performed on the third text line image, and the second text line image is obtained based on the background image and the processed third text line image.

[0241] In some embodiments, the generating module 1002 includes:

[0242] a first determining module configured to, if a degree of matching between the target style information in the font library and the style information of the first text content is greater than or equal to a first threshold, determine the rendering strategy of the target style information as the target rendering strategy;

[0243] The second determining module is configured to determine the processing strategy of the first preset model as the target rendering strategy when the matching degree between the preset style information in the font library and the style information of the first text content is less than the first threshold.

[0244] In some embodiments, when the target rendering strategy is the processing strategy of the first preset model, the generating module 1002 is specifically configured to:

[0245] The style information of the first text content and the second text content are input into the first preset model, and the style information of the first text content is applied to the second text content using the generator network of the first preset model to obtain the third text line image.

[0246] In some embodiments, the first acquisition module 1001 is specifically configured to:

[0247] In a case where the style information of the first text content includes a font size, determining the number of text units in the first text content and the size of the first text line image, and determining the font size based on the number of text units and the size;

[0248] In the case where the style information of the first text content includes font and / or color, the first text line image is input into a recognition model, the font features of the first text content are extracted using the feature extraction network of the recognition model, and the font features are recognized using the recognition network of the recognition model to obtain the font and / or the color.

[0249] In some embodiments, the first acquisition module 1001 is further configured to:

[0250] Using a first recognition network, identifying the font features to obtain a target language category of the first text content;

[0251] The font features are recognized using a second recognition network determined based on the target language category to obtain the font.

[0252] In some embodiments, the performing of the elimination process on the first text content to obtain the background image of the first text line image includes:

[0253] The first text line image is input into a second preset model, and the content features of the first text line image are eliminated using the generator network of the second preset model to obtain the background image.

[0254] In some embodiments, the apparatus 1000 further includes:

[0255] an output module configured to output a reference adjustment template in response to a second operation on the first target image; wherein the reference adjustment template is determined based on historical adjustment data;

[0256] The second acquisition module is configured to obtain a second target image based on the first target image and the reference adjustment template in response to an application instruction for the reference adjustment template.

[0257] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0258] Based on the same inventive concept, an embodiment of the present disclosure provides an electronic device, which may be a computer or a terminal in one or more of the above embodiments. Figure 11 FIG. 1 is a schematic diagram showing the structure of an electronic device according to an exemplary embodiment. Figure 11 As shown, the electronic device 1100 adopts general computer hardware and includes a processor 1101 , a memory 1102 , a bus 1103 , an input device 1104 , and an output device 1105 .

[0259] In some possible implementations, the memory 1102 may include computer storage media in the form of volatile and / or non-volatile memory, such as read-only memory and / or random access memory. The memory 1102 may store an operating system, application programs, other program modules, executable code, program data, user data, and the like.

[0260] The input device 1104 can be used to input commands and information to the electronic device. The input device 1104 can be a keyboard or a pointing device such as a mouse, trackball, touchpad, microphone, joystick, game pad, satellite TV antenna, scanner, or similar device. The input device 1104 can be connected to the processor 1101 via the bus 1103.

[0261] The output device 1105 can be used to output information to the electronic device 1100. In addition to the monitor, the output device 1105 can also be other peripheral output devices, such as speakers and / or printing devices. The output device 1105 can also be connected to the processor 1101 via the bus 1103.

[0262] The electronic device 1100 may be connected to a network, such as a local area network (LAN), via an antenna 1106. In a networked environment, executable instructions may be stored in a remote storage device, rather than being limited to local storage.

[0263] When the processor 1101 in the electronic device 1100 executes the executable code or application stored in the memory 1102, the electronic device 1100 can implement the file processing method in the above embodiment. The specific execution process can be found in the above embodiment and will not be repeated here.

[0264] The memory 1102 may store information for implementing Figure 10 Executable instructions for the functions of the first acquisition module 1001, the generation module 1002, and the synthesis module 1003. Figure 10 The functions / implementation processes of the first acquisition module 1001, the generation module 1002, and the synthesis module 1003 can be realized by Figure 11 The processor 1101 in the embodiment calls the executable instructions stored in the memory 1102 to implement the above-mentioned embodiment. For the specific implementation process and functions, please refer to the above-mentioned related embodiments.

[0265] Based on the same inventive concept, an embodiment of the present disclosure further provides a storage medium having instructions stored therein. When the instructions are executed on a computer, the instructions are used to execute the image processing method in one or more of the above embodiments.

[0266] Based on the same inventive concept, the embodiments of the present disclosure further provide a computer program or a computer program product. When the computer program product is executed on a computer, the computer implements the image processing method in one or more of the above embodiments.

[0267] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.

[0268] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An image processing method, characterized in that: The method comprises: In response to a first operation on a first text line image of an image to be processed, obtaining first text content and style information of the first text line image based on the first text line image, and processing the first text content to obtain second text content; obtaining a second text line image based on the second text content and the style information; A first target image is obtained based on the image to be processed and the second text line image.

2. The method according to claim 1, characterized in that The obtaining, based on the first text line image, the first text content of the first text line image and the style information of the first text line image includes: Performing text recognition on the first text line image to obtain the first text content; performing elimination processing on the first text content to obtain a background image of the first text line image; The text units of the first text content are identified to obtain style information of the first text content.

3. The method according to claim 1, characterized in that The processing of the first text content to obtain the second text content includes: In response to a triggering operation on the first text content, displaying the first text content at an editing location according to the reference style information; In response to an editing operation on the first text content at the editing location, the second text content is obtained.

4. The method according to claim 3, characterized in that When the style information includes style information of the first text content and a background image, obtaining the second text line image based on the second text content and the style information includes: When the editing operation is an insert operation or a delete operation, adjusting the background content of the background image using the style information of the background image, and generating the second text line image based on the second text content, the style information of the first text content, and the adjusted background image; In a case where the editing operation is a modification operation, the second text line image is generated based on the second text content, style information of the first text content, and the background image.

5. The method according to claim 4, characterized in that When the editing operation is an insert operation or a delete operation, adjusting the background content of the background image using the style information of the background image includes: In a case where the insert operation indicates inserting a first text unit into the first text content, moving the text unit of the first text content based on the insertion position of the first text unit; Based on the style information of the background image, the background content of the insertion position is adjusted.

6. The method according to claim 4, characterized in that When the editing operation is an insert operation or a delete operation, adjusting the background content of the background image using the style information of the background image includes: In a case where the deletion operation indicates deletion of the second text unit in the first text content, moving the text units in the first text content except the second text unit based on the target position of the second text unit; Based on the style information of the background image, the background content of the target location is adjusted.

7. The method according to claim 1, characterized in that When the style information includes style information of the first text content and a background image, obtaining the second text line image based on the second text content and the style information includes: Determining a degree of matching between style information of the first text content and preset style information of a font library, and determining a target rendering strategy based on the degree of matching; Rendering the second text content based on the target rendering strategy to obtain a third text line image; Background segmentation processing is performed on the third text line image, and the second text line image is obtained based on the background image and the processed third text line image.

8. The method according to claim 7, characterized in that Determining a target rendering strategy based on the matching degree includes: If the degree of matching between the target style information in the font library and the style information of the first text content is greater than or equal to a first threshold, determining the rendering strategy of the target style information as the target rendering strategy; When the matching degree between the preset style information in the font library and the style information of the first text content is less than the first threshold, the processing strategy of the first preset model is determined as the target rendering strategy.

9. The method according to claim 8, characterized in that When the target rendering strategy is the processing strategy of the first preset model, rendering the second text content based on the target rendering strategy to obtain a third text line image includes: The style information of the first text content and the second text content are input into the first preset model, and the style information of the first text content is applied to the second text content using the generator network of the first preset model to obtain the third text line image.

10. The method according to any one of claims 2 to 9, characterized in that The identifying the text unit of the first text content to obtain the font style information of the first text content includes: In a case where the style information of the first text content includes a font size, determining the number of text units in the first text content and the size of the first text line image, and determining the font size based on the number of text units and the size; In the case where the style information of the first text content includes font and / or color, the first text line image is input into a recognition model, the font features of the first text content are extracted using the feature extraction network of the recognition model, and the font features are recognized using the recognition network of the recognition model to obtain the font and / or the color.

11. The method according to claim 10, characterized in that The step of using the recognition network of the recognition model to recognize the font features to obtain the font includes: Using a first recognition network, identifying the font features to obtain a target language category of the first text content; The font features are recognized using a second recognition network determined based on the target language category to obtain the font.

12. The method according to any one of claims 2 to 9, characterized in that The step of performing elimination processing on the first text content to obtain a background image of the first text line image includes: The first text line image is input into a second preset model, and the content features of the first text line image are eliminated using the generator network of the second preset model to obtain the background image.

13. The method according to any one of claims 2 to 9, characterized in that The method further comprises: In response to a second operation on the first target image, outputting a reference adjustment template; wherein the reference adjustment template is determined based on historical adjustment data; In response to an application instruction for the reference adjustment template, a second target image is obtained based on the first target image and the reference adjustment template.

14. An image processing device, characterized in that: include: a first acquisition module configured to, in response to a first operation on a first text line image of an image to be processed, obtain first text content and style information of the first text line image based on the first text line image, and process the first text content to obtain second text content; a generating module configured to generate a second text line image based on the second text content and the style information; The synthesis module is configured to obtain a first target image based on the image to be processed and the second text line image.

15. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.