Image processing method, device and computer program product
By identifying and translating text content and style information in the image, using the neural network model to generate matching text line images, the problems of low accuracy and poor visual consistency in image text translation are solved, and high-quality cross-language translation effect is achieved.
Patent Information
- Application Number
- CN202510493543.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-15
AI Technical Summary
When the existing neural machine translation system processes text translation in images, the language expression of the translation results is stiff, resulting in low accuracy and difficult to maintain visual consistency of the image in complex backgrounds, resulting in poor display of the translated image.
By identifying the text content and style information in the image, using the neural network model for translation, and generating text line images matching the original image based on the translation strategy, combining image processing technology to optimize the final results to ensure the translated image quality and visual effect.
It significantly improves the accuracy and display effect of image text translation, and is suitable for a variety of cross-language image text translation scenarios, such as advertising poster design and public logo translation, weakening the translation traces and improving the level of intelligence.
Smart Images

Figure CN120496104A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to, but is not limited to, the field of information technology, and in particular to an image processing method, apparatus, and computer program product. Background Art
[0002] In text translation scenarios, Neural Machine Translation (NMT) models can be used to translate the recognized text. While current translation systems perform well for standard text translation, when translating text within images, the resulting translations are often linguistically awkward, resulting in lower accuracy.
[0003] Furthermore, the translation results need to be re-embedded into the original image to maintain visual consistency. However, simple font library rendering or direct overlaying methods make it difficult to restore the style information of the text line area against a complex image background, resulting in a visually fragmented translated image and poor display quality. Summary of the Invention
[0004] To overcome the problems existing in the related technologies, the present disclosure provides an image processing method, device, and computer program product, which not only ensure the image quality of the target image, but also can be applied to a variety of cross-language image-to-text translation scenarios, such as advertising poster design, cultural work translation, or public sign translation, thereby significantly improving the accuracy, display effect, and intelligence level of image-to-text translation.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, including:
[0006] In response to a translation instruction for content in an image, recognizing a first text line image in the image to obtain first text content in a first language and style information of the first text line image;
[0007] When the translation instruction indicates that the first text content is to be translated into a second language, translating the first text content based on a translation strategy for the second language to obtain a second text content;
[0008] generating a second text line image based on the second text content and the style information;
[0009] A target image is obtained based on the image and the second text line image.
[0010] According to a second aspect of an embodiment of the present disclosure, there is provided an image processing apparatus, including:
[0011] a first acquisition module configured to, in response to a translation instruction for content in an image, recognize a first text line image in the image and obtain first text content in a first language and style information of the first text line image;
[0012] a translation module configured to, when the translation instruction indicates that the first text content is to be translated into a second language, translate the first text content based on a translation strategy of the second language to obtain a second text content;
[0013] a generating module configured to generate a second text line image based on the second text content and style information of the first text line image;
[0014] The synthesis module is configured to obtain a target image based on the image and the second text line image.
[0015] According to a third aspect of the present disclosure, an electronic device is provided, including:
[0016] processor;
[0017] a memory for storing processor-executable instructions;
[0018] The processor executes the computer program or instructions to implement the steps of any one of the methods in the first aspect above.
[0019] According to a fourth aspect of an embodiment of the present disclosure, a non-temporary computer-readable storage medium is provided, wherein the storage medium stores a computer program or instructions. When the computer program or instructions in the storage medium are executed by a processor, the steps of the method described in any one of the above-mentioned first aspects are implemented.
[0020] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program or instructions. When the computer program or instructions are executed by a processor, the steps of the method described in any one of the first aspects are implemented. The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:
[0021] The technical solution disclosed herein, in response to a translation instruction for content in an image, recognizes a first text line image in the image and obtains first text content in a first language and style information of the first text line image. This facilitates translation of the first text content based on the translation instruction, meets the user's translation needs, helps accurately capture the contextual relevance of the first text content, and, based on the style information, lays the foundation for reducing translation traces in the first text line image. When the translation instruction instructs to translate the first text content into a second language, the first text content is translated based on a translation strategy for the second language to obtain second text content. Because the translation strategy matches the second language, the obtained second text content has natural linguistic expression, thereby improving translation accuracy. A second text line image is generated based on the second text content and the style information, such that the style information of the second text line image matches the style information of the first text line image. A target image is then obtained based on the image and the second text line image. This not only ensures the image quality of the target image but is also applicable to a variety of cross-language image-to-text translation scenarios, such as advertising poster design, cultural work translation, or public sign translation, thereby significantly improving the accuracy, display effect, and intelligence level of image-to-text translation.
[0022] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0024] Figure 1 The process of an image processing method according to an exemplary embodiment is shown as follows Figure 1 .
[0025] Figure 2 The process of an image processing method according to an exemplary embodiment is shown as follows Figure 2 .
[0026] Figure 3 is a schematic diagram of a preset style template according to an exemplary embodiment.
[0027] Figure 4 The figure is a schematic diagram showing how to translate the content of an image to be processed according to an exemplary embodiment.
[0028] Figure 5 The figure is a block diagram of an image processing apparatus according to an exemplary embodiment.
[0029] Figure 6 The figure is a schematic structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0030] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0031] Figure 1 The process of an image processing method according to an exemplary embodiment is shown as follows Figure 1 .like Figure 1 As shown, the method mainly includes the following steps:
[0032] In step 101, in response to a translation instruction for content in an image, a first text line image in the image is recognized to obtain first text content in a first language and style information of the first text line image;
[0033] In step 102, when the translation instruction indicates that the first text content is to be translated into a second language, the first text content is translated based on the translation strategy of the second language to obtain a second text content;
[0034] In step 103, a second text line image is generated based on the second text content and style information;
[0035] In step 104 , a target image is obtained based on the image and the second text line image.
[0036] It should be noted that the image processing method proposed in the present disclosure can be applied to electronic devices as well as servers. Here, electronic devices may include: terminal devices, such as mobile terminals or fixed terminals. Among them, mobile terminals may include: mobile phones, tablet computers, laptops, etc. Fixed terminals may include: desktop computers, smart TVs, etc. As a type of computer, a server can provide computing or application services to other clients (such as computers, smart phones and other terminal devices, and even large devices such as train systems) in the network.
[0037] The image processing method in the embodiment of the present disclosure may be configured in an image processing device, and the image processing device may be provided in a server or in an electronic device, which is not limited in the embodiment of the present disclosure.
[0038] It should be noted that the executing entity of the embodiment of the present disclosure may be, for example, a central processing unit (CPU) in a server or electronic device in terms of hardware, or may be, for example, a related background service in a server or electronic device in terms of software, without limitation.
[0039] Here, the translation instruction is triggered by the user, which can be triggered manually, for example, a selection box is provided on the interface as a translation button, and the names of different languages are listed in the selection box. The translation instruction can be triggered by touching the translation button; it can also be triggered by voice, for example, when the electronic device has a voice recognition function, the user outputs the translation instruction by voice.
[0040] In some embodiments, the image that triggers the translation instruction can be derived from a variety of scenarios, and the image can be obtained in different ways in each scenario. For example, the image that triggers the translation instruction can be derived from an advertising design scenario, and the image may include: a design drawing of a promotional poster, a flyer, or product packaging; the image that triggers the translation instruction can be derived from an educational resource scenario, and the image may include: teaching materials, exhibition images, or paper illustrations; the image that triggers the translation instruction can be derived from a social media scenario, and the image may include: a social media image or a dynamic poster; the image that triggers the translation instruction can be derived from a tourism service scenario, and the image may include: a scenic spot sign or a traffic sign.
[0041] In some embodiments, after obtaining the image, text line detection may be performed on the image to determine the first text line image in the image. Exemplarily, text line detection may be performed on the image using a text line detection model. For example, a text line detection model based on deep learning performs text line detection on the image to be processed, wherein the text line detection model may include: a segmentation-based text detection (Differentiable Binarization, DB) algorithm, an instance segmentation-based text detection model (Mask Text Spotter), a Faster R-CNN-based text line detection algorithm (Connectionist Text Proposal Network, CTPN), etc.
[0042] Here, by performing text line detection on an image, the position information of each text line in the image, including the bounding box coordinates of the text line, can be quickly and accurately detected, thereby obtaining a first text line image and the orientation information of the first text line image in the image. The orientation information includes direction information and position information.
[0043] Here, the number of first text line images determined from the image may be one or more, and the first orientation information of each first text line image may be the same or different.
[0044] Exemplarily, the bounding box of a text line is defined as:
[0045] B i ={(x 1i ,y 1i ),(x 2i ,y 2i )} (1);
[0046] In formula (1), (x 1i ,y 1i ) is the coordinate of the upper left corner of the bounding box of the text line, (x 2i ,y 2i ) is the lower right corner coordinate of the text line's bounding box.
[0047] At the same time, the text lines in the image meet the preset conditions:
[0048] P(B i |I)=softmax(W f ·F(I)) (2);
[0049] Among them, in formula (2), F(I) is the feature extraction of image I, W f is the weight of the classifier.
[0050] In an embodiment of the present disclosure, in response to a translation instruction, text units of the first text line image can be identified through optical character recognition (OCR) technology, and text information can be extracted from the first text line image and converted into an editable text format through this technology to obtain the first text content.
[0051] Here, the OCR method is mainly based on a deep learning model, such as the Convolutional Recurrent Neural Network or the Shape Attention Text Recognizer Network. During the training process, a sample image containing text and the actual character label are input into the deep learning model. The sample image is extracted using the Convolutional Neural Networks (CNN) to obtain an initial feature sequence. The Recurrent Neural Network (RNN) processes the initial feature sequence to capture the temporal dependency between text units. The Connectionist Temporal Classification (CTC) network converts the output of the RNN into a character sequence to obtain a character label. Based on the difference information between the predicted text label and the actual text label, the model parameters of the deep learning model are adjusted until the preset convergence condition is reached. The difference information is obtained based on the CTC loss function. In actual application, the first text line image is input into the deep learning model to obtain the first text content of the first text line image.
[0052] For example, the CTC loss function can be calculated as follows:
[0053]
[0054] Among them, in formula (3), P(y t |x t ) is the character label y output by the deep learning model at time step t t The probability distribution of x t is the feature of the input image.
[0055] In some embodiments, in addition to using OCR technology to identify text units and obtain recognition results of text units, the recognition results can also be corrected based on Bidirectional Encoder Representations from Transformers, BERT's language model to enhance the processing capabilities of complex character ambiguity and multi-language mixed scenarios, thereby ensuring the comprehensiveness and accuracy of the acquired first text content.
[0056] In an embodiment of the present disclosure, the style information of the first text line image includes: style information of the first text content and / or the background image of the first text line image; when the first text content is identified, the first text content in the first text line image can be eliminated to obtain the background image; and the style information of the first text content is obtained by identifying the text units of the first text content. In some embodiments, the background image and the style information of the first text content can be obtained by model recognition. For example, the first text content is recognized by a first neural network model to obtain the style information of the first text content. In another example, the first text line image is processed by a generative adversarial network (GAN) to obtain a background image that does not contain the first text content. In addition, the background image is recognized based on a second neural network model to obtain the style information of the background image.
[0057] In other embodiments, the style information of the first text content can also be obtained by determining similarity. For example, feature information of text units of the first text content is extracted, and similarity between the feature information and preset style information in a font library is determined, and then the preset style information with the greatest similarity is determined as the style information of the first text content.
[0058] In other embodiments, the style information of the background image can also be obtained through image processing technology. For example, color analysis is performed on the first text line image to obtain features such as the color histogram and color space distribution of the background area, thereby obtaining the style information of the background image. In another example, texture analysis is performed on the first text line image to extract texture features, and then the style information of the background image is obtained through texture analysis functions, such as Gray-Level Co-occurrence Matrix (GLCM) and Local Binary Patterns (LBP).
[0059] Here, the style information of the first text content includes but is not limited to the font size, font and / or color of the text unit of the first text content, etc. The style information of the background image includes but is not limited to: background color, background pattern, background content, etc.
[0060] In the disclosed embodiment, when first text content is obtained, the first text content can be input into a language identification model. The feature extraction network of the language identification model extracts features from the first text content to obtain language features, such as vocabulary, grammatical structure, and semantics. The recognition network of the language identification model then determines the similarity between the language features and language features of different languages, thereby obtaining a recognition result that the first text content is in the first language. Here, the language identification model can be Fast Text, Lang ID, or Compact Language Detector 3, among others.
[0061] Here, the first text content being in the first language can be understood as all text units of the first text content being in the first language. For example, when the first text content is an ancient poem, all text units recognized from the poem are in the first language.
[0062] The first text content being in the first language can also be understood as at least two text units in the first text content being in the first language. For example, if the first text content is a biography of a foreign writer, language identification is performed on the biography to determine that some text units in the biography are in the first language, while other text units in the biography are in a third language, where the third language is different from the first language, and the number of text units in the first language is greater than the number of text units in the third language.
[0063] It is understandable that when the translation instruction indicates that the first text content is to be translated into a second language, the first text content can be translated based on the translation strategy of the second language to obtain the second text content. Since the translation strategy matches the second language, it is conducive to improving the accuracy of the translation, thereby ensuring the accuracy of the second text content.
[0064] Here, the second language and the first language may be the same, for example, the first language and the second language are both Chinese.
[0065] Exemplarily, language identification is performed on the first text content to determine that the third text unit in the first text content is in English, and the text units in the first text content other than the third text unit are in Chinese; when the second language is Chinese, the third text unit can be translated from an English text unit into a Chinese text unit; based on the translated third text unit and the text units other than the third text unit, the second text content is determined.
[0066] The second language can also be different from the first language. For example, the first language can be English and the second language can be Japanese, Chinese, Italian, or Thai.
[0067] It is understandable that different languages have different translation strategies. These include, but are not limited to, grammatical rules, language complexity, sentence structure, the cultural context of the language, and terminology. For example, English sentence structures are relatively simple, while Chinese may require more modifiers and idioms. Therefore, when translating an English text unit into a Chinese text unit, the translation can be based on the language complexity of the Chinese language to produce a natural and fluent translation result.
[0068] As another example, the tense expression of Chinese content is relatively vague and depends more on the context and time adverbials, while the tense expression of English content is clear, and each verb has a fixed tense form. Therefore, when translating a Chinese text unit into an English text unit, the translation can be performed according to the grammatical rules of English to obtain a translation result with a more accurate tense expression.
[0069] The translation method in the disclosed embodiments can be integrated into a large language model (LLM), which can then translate the first text content based on the logic of step 102. When a large language model performs translation based on the above process, not only can the contextual translation style be maintained consistent, but the contextual relevance of the translation result can also be improved, resulting in better translation results.
[0070] The large language model is a super-large deep learning model pre-trained on a large amount of data. Its underlying transformer is a set of neural networks consisting of an encoder and decoder with attention. These encoders and decoders extract semantic meaning from a series of text units and understand the relationships between them to complete translation.
[0071] In some embodiments, the translation prompt word and the first text content can be input into a large language model. The large language model can first determine the length of the first text content and determine whether to perform word segmentation on the first text content based on the length of the first text content. If the first text unit is not subjected to word segmentation, the similarity between the first text content and each translation prompt word in the prompt word is directly calculated. If the first text unit is subjected to word segmentation, the text unit after word segmentation is first obtained, and then the similarity between each text unit obtained by word segmentation and each translation prompt word is calculated. If the similarity is greater than a preset threshold, the text unit of the first text content is replaced with the translation prompt word to obtain the replaced first text content, and the replaced first text content is input into the large language model for translation to obtain the second text content. In this way, the accuracy of the text unit replaced by the translation prompt word is improved, thereby not only maintaining the consistency of the contextual translation style, but also improving the contextual relevance of the translation result. Here, the similarity calculation method includes but is not limited to cosine similarity, Euclidean distance, etc.
[0072] For example, if the first text content is a Chinese slogan that reads "The mobile phone has an extra-large display screen and a fingerprint recognition function," and the translation instruction indicates that the first text content is to be translated into English, the first text content can be segmented to obtain various text units, such as "mobile phone," "having," "extra-large," "display screen," "and," "fingerprint recognition," and "function." If the translation prompt indicates that "mobile phone" is to be translated as "mobile phone" and "display screen" is to be translated as "display screen," the text units of the first text content are replaced with the translation prompt, and the replaced first text content becomes "mobile phone has an extra-large display screen and a fingerprint recognition function." The replaced first text content is then input into the large language model to obtain a second text content in English, namely, "The mobile phone has a large display screen and fingerprint recognition function."
[0073] In some embodiments, when the second text content is obtained, in order to facilitate the user to compare the first and second text contents, the style information of the first text content can be applied to the second text content, and the second text content with the style information of the first text content can be displayed in the edit box. For example, based on the position information, font, font color, font size, and font spacing of the first text content in the first text line image, the second text content can be displayed in the edit box with the style information to facilitate user verification.
[0074] It can be understood that since the second text line image is generated based on the second text content and the style information of the first text line image, the style information of the second text line image can be matched with the style information of the first text line image, which is conducive to reducing the translation traces of the first text line image, thereby improving the image quality and display effect of the second text line image.
[0075] In some embodiments, based on the style information of the first text content, the style information of the second text content is adjusted; and based on the adjusted second text content and the background image, a second text line image is generated. For example, if the first text content is in Chinese and the font of the first text content is Songti, the first text content is first translated from Chinese into English to obtain the second text content; then the font of the second English text content is adjusted from Songti to Roman to obtain the adjusted second text content; finally, the adjusted second text content is superimposed on the background image to obtain the second text line image.
[0076] In other embodiments, a target text unit that does not match the style information of the first text content is determined from the second text content; when the target text unit is determined, the style information of the target text unit is adjusted based on the style information of the first text content; and the background content is adjusted based on the style information of the background image; and then, based on the adjusted second text content and the adjusted background image, a second text line image is obtained, thereby improving the efficiency of obtaining the second text unit.
[0077] Here, in the case where the second text line image is obtained, a target image may be generated based on the image and the second text line image.
[0078] In some embodiments, the image can be preprocessed, for example, by grayscale conversion, edge detection, etc.; the first text line image is matched with the second text line image so as to maintain the size of the area where the first text line image is located in the preprocessed image, or the size of the area where the first text line image is located is adjusted, for example, the area where the first text line image is located is expanded; the area where the unadjusted or adjusted first text line image is located is cut out in the preprocessed image to obtain the first image; the second text line image is then fused with the first image, and the fused image is optimized through the Patch GAN model to eliminate edge traces, thereby obtaining the target image.
[0079] In some embodiments, after the fused image is processed by the Patch GAN model, a denoising algorithm (e.g., non-local mean denoising, BM3D, etc.) can be used to remove the noise generated during the fusion process; and a sharpening algorithm (e.g., Laplace sharpening, Unsharp Mask (USM) sharpening, etc.) can be used to enhance image details, thereby optimizing the visual effect of the target image. These effects include that after fusion, the second text line image and the first image will not form different content, that is, the second text line image will be completely integrated, and there will be no blank areas in the image due to the processing of the original text line image. After the second text image is integrated, the first image will be integrated with the background of the second text image, and the background will be intelligently adjusted so that the second text image is integrated without leaving any blank space.
[0080] In other embodiments, once the target image is obtained, the user can adjust the target image as needed to meet personalized requirements. For example, the size, position, or rotation angle of the text line image in the target image can be adjusted; in another example, the size, transparency, or clarity of the target image can be adjusted.
[0081] In some embodiments, in order to improve the user's editing experience, the user's historical adjustment data can be recorded, and a reference adjustment model can be generated based on the historical adjustment data; and when the target image is obtained and the adjustment operation for the target image is detected, the reference adjustment template can be output in a timely manner, thereby improving the efficiency of adjusting the target image while meeting the user's personalized needs.
[0082] The technical solution disclosed herein, in response to a translation instruction for content in an image, recognizes a first text line image in the image and obtains first text content in a first language and style information of the first text line image. This facilitates translation of the first text content based on the translation instruction, meets the user's translation needs, and helps accurately capture the contextual relevance of the first text content. Furthermore, based on the style information, it can lay the foundation for reducing translation traces in the first text line image. When the translation instruction instructs to translate the first text content into a second language, the first text content is translated based on a translation strategy for the second language to obtain second text content. Because the translation strategy matches the second language, the obtained second text content has natural linguistic expression, thereby improving translation accuracy. Furthermore, based on the second text content and the style information, a second text line image is generated, such that the style information of the second text line image matches the style information of the first text line image. A target image is then obtained based on the image and the second text line image. This not only ensures the image quality of the target image but is also applicable to a variety of cross-language image-to-text translation scenarios, such as advertising poster design, cultural work translation, or public sign translation, thereby significantly improving the accuracy, display effect, and intelligence level of image-to-text translation.
[0083] In some embodiments, when the translation instruction indicates that the first text content is to be translated into a second language, translating the first text content based on a translation strategy of the second language to obtain the second text content includes:
[0084] When the translation instruction indicates that the first text content is to be translated into a second language, determining a first prompt word in the second language from a prompt word template library;
[0085] In the case where the first text content includes a key text unit, determining a second prompt word for the key text unit from a prompt word template library; wherein the key text unit is used to indicate the field to which the first text content belongs;
[0086] Based on the first prompt word and / or the second prompt word, the first text content is translated to obtain the second text content.
[0087] It is understandable that different languages have different linguistic features, and a prompt word template library containing translation prompt words in different languages can be pre-established. At the same time, considering that the content to be translated is text content in specific fields with high professionalism, in order to improve the accuracy of translation, the prompt word template library can also include translation prompt words in different fields.
[0088] Here, in the prompt word template library, there is a mapping relationship between each language and the translated prompt word of the language; when the second language is determined, the translated prompt word of the second language can be determined based on the mapping relationship.
[0089] In addition, the translation prompt words in the field are verified terms with high accuracy.
[0090] In the embodiment of the present disclosure, when a translation instruction is detected, a translation prompt word in the second language is selected from the prompt word template library as the first prompt word; and when the first text content includes a key text unit, a translation prompt word in the field matching the key text unit is selected from the prompt word template library as the second prompt word.
[0091] Here, the key text unit is used to indicate the field to which the first text content belongs, including but not limited to the subject, attributive or adverbial of the first text content.
[0092] In some embodiments, when the first text content is obtained, the first text content can be parsed to extract the key text units of the first text content; and the field to which the key text units belong can be determined; and then a second prompt word matching the field can be determined from a prompt word template library.
[0093] Here, when all text units of the first text content belong to the first language and the first text content does not have a key text unit, a first prompt word is determined from a prompt word template library, and based on the first prompt word, the first text content is translated to obtain the second text content.
[0094] When the text unit of the first text content belongs to the first language and there is a key text unit, a first prompt word and a second prompt word are respectively determined from the prompt word template library, and the first text content is translated based on the first prompt word and the second prompt word to obtain the second text content.
[0095] When the key text unit of the first text content belongs to the first language and the text units other than the key text unit belong to the second language, a second prompt word is determined from the prompt word template library, and the first text content is translated based on the second prompt word to obtain the second text content.
[0096] In some embodiments, at least one of the first prompt word and the second prompt word and the first text content can be input into a large language model, and the input prompt word and the first text content can be encoded using the encoder of the large language model to obtain a target feature vector; and the target feature vector can be decoded using the decoder of the large language model to obtain the second text content.
[0097] Through the technical solution of the present disclosure, a prompt word template library is pre-built, and prompt words can be determined from the prompt word template library based on key text units in the second language and / or the first text content; and translation is performed based on the prompt words, which is conducive to improving the accuracy and efficiency of translation.
[0098] In some embodiments, translating the first text content based on the first prompt word and / or the second prompt word to obtain the second text content includes:
[0099] Translating the first text content based on the language feature indicated by the first prompt word to obtain first translated content;
[0100] Based on the terminology information indicated by the second prompt word, the text unit associated with the key text unit in the first text content is translated to obtain second translated content;
[0101] Based on the first translation content and / or the second translation content, a second text content is obtained.
[0102] Here, language features represent the systematic characteristics of a language in terms of vocabulary, syntax, and semantics. Terminology information represents the mapping relationship between language and concepts in a specific field.
[0103] It is understandable that translating the first text content based on the linguistic features of the second language and the terminology information of the field is conducive to generating second text content that is semantically accurate, contextually coherent, and naturally expressed, thereby improving translation accuracy.
[0104] In some embodiments, when the first prompt word and the first text content are input into the large language model, the large language model parses the language features indicated by the first prompt word to obtain the grammatical features, lexical features and related semantics of the second language. At the same time, the large language model can perform word segmentation on the first text content, divide the continuous text into independent words or phrases, and mark punctuation marks, special characters, etc. in the first text content. In the encoding stage, the first text content after word segmentation and the grammatical features and lexical features of the second language are encoded into a series of high-dimensional vectors; in the decoding stage, the encoded high-dimensional vectors are transmitted to the decoder, and the decoder obtains the first translation content based on these vectors.
[0105] In other embodiments, for a text unit associated with a key text unit in the first text unit (hereinafter referred to as a fourth text unit), it may be first determined whether the second prompt word contains a translation result of the fourth text unit. If the second prompt word contains a translation result of the fourth text unit, the prompt word in the second prompt word is determined as the translation result of the fourth text unit. If the second prompt word does not contain a translation result of the fourth text unit, machine translation (MT) or a large language model may be used to translate the fourth text unit.
[0106] Through the technical solution of the present disclosure, semantically accurate, contextually coherent and naturally expressed second text content is generated based on the linguistic characteristics and field requirements of different languages (for example, medical, legal or commercial), thereby improving translation accuracy.
[0107] In some embodiments, translating the first text content based on the language feature indicated by the first prompt word to obtain first translated content includes:
[0108] Determining language similarity between a text unit of the first text content and a reference text unit in the second language in the first prompt word, and determining a text unit having a language similarity greater than or equal to a first threshold as a first text unit;
[0109] Based on the language feature indicated by the first prompt word, the text units in the first text content except the first text unit are translated to obtain first translated content.
[0110] It should be noted that, considering that there may be a situation where the first text content contains text units in the second language, if the text unit in the second language is first identified in the first text content and the text unit is not translated during the translation process of the first text content, it can avoid the situation where the text unit is translated into other translations, and at the same time, it can improve the relevance of the context, thereby improving translation efficiency.
[0111] Therefore, in the disclosed embodiment, a first threshold is pre-set. When the language similarity between a text unit in the first text content and a reference text unit in the second language in the first prompt word is determined, a determination is made based on the language similarity and the first threshold to determine whether the first text content contains the first text unit in the second language. The first threshold can be set arbitrarily as needed, for example, 95%.
[0112] Here, the first prompt word can be understood as the translation template of the second language, and the reference text unit is any text unit in the translation template of the second language, for example, a quantifier or auxiliary word in the translation template.
[0113] Specifically, a text unit whose language similarity is greater than or equal to a first threshold is determined as a first text unit. When the first text unit exists in the first text content, the text units in the first text content except the first text unit are translated to obtain a first translation content; and when the first text unit does not exist in the first text content, all text units in the first text content are translated.
[0114] Here, there are many methods for determining language similarity, for example, methods based on string matching or methods based on neural network models.
[0115] In some embodiments, a string matching algorithm is used to measure the similarity between a text unit of the first text content and a reference text unit to assess the language similarity between the text unit of the first text content and the reference text unit. For example, an edit distance algorithm can calculate the minimum number of edit operations (e.g., insertion, deletion, or replacement) required to convert a text unit of the first text content into a reference text unit, and based on the minimum number of edits, a similarity score is obtained to further determine the language similarity.
[0116] In another example, the Jaccard similarity algorithm can be used to calculate the ratio of the intersection of the text units of the first text content and the reference text units to the union of the text units of the first text content and the reference text units; based on the ratio, a similarity score is obtained to further determine the language similarity.
[0117] Through the technical solution disclosed in the present invention, the language of each text unit in the first text content can be automatically identified. If the first text content contains a first text unit in a second language, there is no need to translate the first text unit, thereby reducing invalid translations and improving context relevance and translation efficiency.
[0118] In some embodiments, the method further comprises:
[0119] generating a third prompt word in response to an application instruction for a target style template in a preset style template, wherein the preset style template is used to indicate an arrangement style of the translated content;
[0120] The first text content is translated based on the first prompt word, the second prompt word, and the third prompt word to obtain the second text content.
[0121] It is understandable that at least one preset style template can be displayed first to facilitate the user to select a target style template that meets the needs; and when an application instruction for the target style template is detected, a third prompt word is generated, and then translation is performed based on the third prompt word to obtain target translation content arranged according to the target style template, so that the user can verify the second text content, thereby improving the user's translation experience.
[0122] Here, different preset style templates indicate different arrangements of translation content. For example, a first preset style template indicates that the first line displays the coordinates of a first text line area, the line below the first line displays the first text content, and the line below the last line displaying the first text content displays the second text content. In another example, a second preset style template indicates that the first line displays the coordinates of the first text line area, and the line below the first line displays the second text content.
[0123] In some embodiments, when there is no key text unit in the first text content, the first text content can be translated based on the language features indicated by the first prompt word to obtain the first translated content; and based on the third prompt word, the arrangement style of the first translated content can be adjusted to obtain the second text content.
[0124] In other embodiments, when the key text unit of the first text content is in a first language and the text units other than the key text unit are in a second language, the text unit associated with the key text unit in the first text content can be translated based on the terminology information indicated by the second prompt word to obtain second translated content; and the second translated content can replace the text unit associated with the key text unit in the first text content to obtain third text content; and then, based on the third prompt word, the arrangement style of the third text content can be adjusted to obtain second text content.
[0125] In other embodiments, when the text unit of the first text content belongs to the first language and there is a key text unit, the first text content can be translated based on the language features indicated by the first prompt word to obtain the first translated content; and based on the terminology information indicated by the second prompt word, the text unit associated with the key text unit in the first text content is translated to obtain the second translated content; then, based on the first translated content and the second translated content, the third text content is obtained; finally, based on the third prompt word, the arrangement style of the third text content is adjusted to obtain the second text content.
[0126] Through the technical solution of the present disclosure, the user can select a target style template that meets the needs based on the preset style template, so that the translated second text content can be output based on the target style template, thereby satisfying the user's personalized experience.
[0127] In some embodiments, recognizing a first text line image in an image to obtain first text content in a first language and style information of the first text line image includes:
[0128] Performing text recognition on the first text line image to obtain first text content;
[0129] Eliminating the first text content to obtain a background image of the first text line image;
[0130] Generating a second text line image based on the second text content and style information includes:
[0131] Determining a target font based on the second language and the font of the first text content;
[0132] determining target text content based on the target font and the second text content;
[0133] A second text line image is generated based on the target text content and the background image.
[0134] In the embodiment of the present disclosure, the text information can be extracted from the first text line image and converted into an editable text format through OCR technology to obtain the first text content.
[0135] Here, the background image of the first text line image may be obtained through an image inpainting algorithm, or may be obtained through an image generation model.
[0136] Taking the sample-based image restoration algorithm as an example, a sample area similar to the area covered by the first text content is determined, wherein the sample area can be an area in the first text line image that is not covered by the first text content, or an area in an image similar to the first text line image; for each pixel point that needs to be repaired, based on the color, texture, shape and other characteristics of the pixel point, a sample pixel point that matches the pixel point is determined from the sample area; the matched sample pixel point is directly covered to the area covered by the first text content, or the matched sample pixel point is interpolated with the surrounding pixel points of the area covered by the first text content to obtain the background image of the first text line image.
[0137] Taking the example of a generative adversarial network as an image generation model, the generative adversarial network consists of two neural networks: a generator network and a discriminator network. The generator is used to generate predicted images close to the real dataset, while the discriminator is used to distinguish whether the input image comes from the real dataset or is generated by the generator. The first text line image is input into the image generation model. When the area where the first text line content is located is determined, the generator network can generate image content that matches the background of the area outside the area based on the pixel information of the surrounding pixels of the area and the global context features, that is, the background image of the first text line image.
[0138] Here, different languages have different fonts. When the font of the first text content and the second language are determined, a target font can be determined based on the font of the text content and the second language, thereby improving the display effect of the second text line image.
[0139] In some embodiments, a first association relationship is preset that represents the correspondence between fonts in different languages. For example, the first association relationship represents: Chinese Songti corresponds to English Roman, Chinese Kaiti corresponds to English Script, Chinese Songti corresponds to Japanese Mincho, and Chinese Heiti corresponds to Korean Gothic. After the font of the first text content and the second language are determined, a target font can be determined based on the font of the first text content, the second language, and the first association relationship, thereby improving the accuracy and efficiency of determining the target font.
[0140] In other embodiments, considering that the font of the first text content is non-standard, such as handwritten or custom font, and there is no directly corresponding font in the second language for the non-standard font of the first language, the font of the first text content can be determined as the target font so that the display effect of the second text line image after translation matches the display effect of the first text line image before translation, thereby improving the image quality of the target image.
[0141] In the disclosed embodiment, the target text content can be determined based on the target font and the second text content to flexibly adapt to the user's editing operation, thereby improving the intelligence level of image text editing.
[0142] Through the technical solution disclosed in the present invention, by obtaining the background image of the first text line image, the translation traces of the second text line image can be reduced; at the same time, by determining the target font, the font of the target text content is matched with the second language, which can improve the aesthetics and cultural expression of the target text content; and based on the background image and the target text content, the second text line image is generated. Even if the background image of the first text line image is relatively complex, the background image can still be restored after the text translation, thereby significantly improving the image quality and display effect of the second text line image.
[0143] In some embodiments, translating the first text content based on the first prompt word, the second prompt word, and the third prompt word to obtain the second text content includes:
[0144] When the first text content is translated to obtain target translation content arranged according to the target style template, the target translation content is displayed in an editing box according to the reference style information; wherein the editing box is located at a preset position of the image;
[0145] In response to an editing operation on the target translation content in the editing box, obtaining second text content;
[0146] Determining target text content based on the target font and the second text content includes:
[0147] In response to a triggering operation on the edit box, target text content is determined based on the target font and the second text content.
[0148] Here, the preset position is any position in the image, such as the top position, left position, center position, right position or bottom position of the image.
[0149] It should be noted that, based on at least one of the first and second prompt words and the third prompt word, the first text content is translated to obtain target translation content arranged according to the target style template. To improve the accuracy of the second text content, once the target translation content is obtained, an edit box can be output and the target translation content can be displayed within the edit box according to the reference style information, allowing the user to verify and edit the target translation content within the edit box.
[0150] Here, the font and font size indicated by the reference style information may be the same as or different from the font and font size of the first text content, and this embodiment of the present disclosure does not limit this.
[0151] In some embodiments, considering that the first text content may have a smaller font size, if the target translation content is displayed in the edit box using the same font size as the first text content, the user may not be able to accurately identify the target translation content, thereby hindering the user's editing experience. Therefore, a font size that is convenient for human perception, such as size 4, can be pre-set, and the target translation content can be displayed using the pre-set font size.
[0152] Here, the editing operation may be a deletion operation, an insertion operation, or a modification operation, etc. For example, if the editing operation is an insertion operation, the text unit indicated by the insertion operation is inserted into the target translation content at the location triggering the insertion operation to obtain the second text content; if the editing operation is a modification operation, the text unit indicated by the modification operation in the target translation content is modified to obtain the second text content; if the editing operation is a deletion operation, the text unit indicated by the deletion operation in the target translation content is deleted to obtain the second text content.
[0153] In the embodiment of the present disclosure, the triggering operation is any touch operation on the editing box, which is used to indicate that the editing of the target translation content is completed and the second text content can be generated. Exemplarily, the triggering operation is one or more of a click operation, a long press operation, or a slide operation.
[0154] In some embodiments, a first preset model for font style transfer can be trained based on a learning framework (such as Tensor Flow, PyTorch, etc.). Based on the first preset model, the font style of the target font is applied to the second text content to obtain the target text content, so that the font of the target text content matches the second language, thereby improving the professionalism and aesthetics of the target text content.
[0155] In other embodiments, it is determined whether a second text unit identical to the first text content exists in the second text content; if the second text unit exists in the second text content, the target font is used for text units in the second text content other than the second text unit, thereby improving the efficiency of obtaining the target text content.
[0156] Through the technical solution disclosed in the present invention, when the target translation content is obtained, the target translation content is displayed in the editing box according to the reference style information, so that the user can verify and edit the target translation content, thereby improving the accuracy of the second text content; and when a trigger operation on the editing box is detected, the target text content is determined based on the target font and the second text content, which is conducive to improving the flexibility and intelligence of text translation.
[0157] In some embodiments, determining the target text content based on the target font and the second text content includes:
[0158] Determining a degree of matching between a target font and a preset font in a font library, and determining a target rendering strategy based on the degree of matching;
[0159] Rendering the second text content based on the target rendering strategy to obtain a third text line image;
[0160] The third text line image is segmented to obtain target text content.
[0161] In some embodiments, a font library is a data set containing multiple font styles and text units, allowing users to select and use different fonts in various applications. Each font in the font library is composed of a set of vector graphics or bitmap images that define the shape and appearance of the text unit.
[0162] Therefore, by comparing the degree of matching between the font of the first text content and the preset fonts in the font library, the target rendering strategy of the second text content can be determined, thereby improving the accuracy and efficiency of rendering the second text content.
[0163] Here, the number of preset fonts can be one or more.
[0164] Taking the font of a text unit as an example, the first feature information of the font of the first text content and the second feature information of a preset font in the font library are extracted, where the feature information includes outline shape, stroke thickness, or character spacing, etc.; the extracted first feature information and the second feature information are quantized, and the similarity between the quantized first feature information and the second feature information is calculated, for example, using a measurement method such as cosine similarity and Euclidean distance to obtain a comparison result; and based on the comparison result, the degree of matching between the font of the first text content and the preset font in the font library is determined. Here, in order to improve the accuracy of determining the comparison result, a third threshold can be set in advance. When the similarity is determined, the similarity is compared with the third threshold to obtain a comparison result.
[0165] Here, the target rendering strategy includes: a font library-based rendering strategy or a preset model-based processing strategy. If the style information of the first text content closely matches the preset style information of the font library, the second text content may be rendered based on the font library to obtain a third text line image. If the style information of the first text content closely matches the preset style information of the font library, the second text content may be rendered based on the preset model-based processing strategy to obtain a third text line image.
[0166] It should be noted that, considering that the style information of the background content of the third text line image does not match the background image of the first text line image, the translation traces of the first text line image will be obvious. Therefore, the third text line image can be segmented to obtain the target text content, so that the target text content and the background image can be integrated, thereby improving the image quality of the second text line image.
[0167] Here, the segmentation process may be implemented by a threshold-based segmentation method, an edge-based segmentation method, or a model-based segmentation method, which is not limited in the embodiments of the present disclosure.
[0168] For example, when the contrast between the target text content and the background content of the third text line image is high, a second threshold can be set, and pixels with pixel values above or below the second threshold are classified as foreground and background, respectively, to obtain the target text content. In another example, a classifier (e.g., a support vector machine, a neural network, etc.) is trained using the labeled foreground and background images, and the trained classifier is then used to perform pixel-level classification on the third text line image to obtain the target text content.
[0169] Through the technical solution disclosed in the present invention, the target rendering strategy for the second text content is determined based on the matching degree between the font of the first text content and the preset font of the font library, thereby improving the accuracy and efficiency of the target rendering strategy, thereby improving the accuracy of obtaining the third text line image; and the third text line image is style-processed to obtain the text foreground, so that the text foreground and the background image can be integrated, so that even if the background of the first text line image is relatively complex, the editing traces of the first text line image can be reduced, thereby improving the intelligence of text translation.
[0170] In some embodiments, determining a degree of matching between a target font and a preset font in a font library includes:
[0171] Extracting a first font feature of a target font and a second font feature of a preset font using a feature extraction network;
[0172] Determining a matching degree between the target font and the preset font based on a similarity between the first font feature and the second font feature;
[0173] Determine the target rendering strategy based on the matching degree, including:
[0174] When the matching degree between the preset font in the font library and the target font is greater than or equal to a second threshold, determining the rendering strategy of the preset font as the target rendering strategy;
[0175] When the matching degree between the preset font in the font library and the target font is less than the second threshold, the processing strategy of the first preset model is determined as the target rendering strategy.
[0176] It should be noted that font features refer to the properties of a font in terms of its form, structure, and strokes. Different fonts have different font features, such as the decorative serifs of serif fonts and the clean lines of sans serif fonts. By extracting the first font features of the target font and the second font features of the preset font through a feature extraction network, the similarity between the two fonts can be accurately identified, thereby improving the accuracy of the matching between the two fonts.
[0177] Here, the feature extraction network can be a convolutional neural network (CNN) or a transformer model, etc.; the feature extraction network includes at least a convolutional layer and a pooling layer. The number of convolutional layers and pooling layers in the feature extraction network can be arbitrarily set as needed and is not limited in this embodiment of the present disclosure.
[0178] In some embodiments, in order to facilitate the calculation of similarity, when obtaining the first font feature and the second font feature, the first font feature can be quantized to obtain a first feature vector, for example, (contour shape A, stroke thickness A, character spacing A, curve A); and the second font feature can be quantized to obtain a second feature vector, for example, (contour shape B, stroke thickness B, character spacing B, curve B); and then based on the similarity calculation formula, the similarity between the target font and the preset font is determined.
[0179] For example, the formula for cosine similarity is:
[0180]
[0181] In formula (4), F represents the first eigenvector, and G represents the second eigenvector.
[0182] Cosine similarity determines whether two vectors point in roughly the same direction by calculating the cosine of the angle between them. The closer the cosine value is to 1, the more similar the two vectors are.
[0183] As another example, the formula of Euclidean distance is:
[0184]
[0185] In formula (5), F i Represents the vector value of a certain dimension in the first eigenvector, G i Represents the vector value of a dimension in the second eigenvector. Euclidean distance is calculated by taking the square root of the sum of the squared differences between the two vectors in each dimension. The smaller the Euclidean distance, the more similar the two vectors are.
[0186] In some embodiments, after obtaining the first eigenvector and the second eigenvector, the first eigenvector and the second eigenvector may be normalized so that the normalized first eigenvector and the normalized second eigenvector are at the same scale, thereby improving the accuracy of calculating the similarity. Here, the normalization process may be Min-Max normalization or Z-score normalization, etc.
[0187] It is understandable that in order to determine the efficiency of the target rendering strategy, a second threshold can be preset, and when the matching degree between the target font and the preset font in the font library is determined, the matching degree is compared with the second threshold to obtain a comparison result to further determine the target rendering strategy.
[0188] Here, the second threshold can be set arbitrarily according to needs, such as 90%, and the embodiment of the present disclosure is not limited to this.
[0189] In an embodiment of the present disclosure, when the comparison result indicates that the matching degree between a preset font in the font library and the target font is greater than or equal to a second threshold, this preset font can be determined as a reference font; and the rendering strategy of the reference font can be determined as the target rendering strategy, thereby ensuring the rendering accuracy of the second text content.
[0190] Exemplarily, a matched reference font is read from a font library; the graphics library or rendering engine can read the meta-information of the reference font, such as the font's outline, shape, size, and tilt angle; based on the meta-information, the graphics library or rendering engine converts the second text content into pixel data, thereby obtaining a third text line image.
[0191] In some embodiments, when the degree of match between at least two preset fonts in the font library and the target font is greater than or equal to a first threshold, the preset font with the highest degree of match among the at least two preset fonts may be determined as the reference font. For example, when the degree of match between the target font and Songti in the font library and the degree of match between the target font and Xin Songti in the font library are both greater than the first threshold, the preset font with the higher degree of match between Songti and Xin Songti is determined as the reference font.
[0192] When the comparison result indicates that the matching degree between the preset fonts in the font library and the target font is less than the first threshold, it is determined that the font library-based rendering method cannot meet the requirements, and the processing strategy of the first preset model is determined as the target rendering strategy. In this way, even if the target font is a non-standard font, such as handwriting, multi-texture font or special custom font, the font style of the target font can be flexibly restored, thereby improving the generation ability of non-standard fonts, and then improving the image quality of the second text line image.
[0193] Through the technical solution disclosed in the present invention, the accuracy of determining the matching degree between the target font and the preset font is improved by calculating the similarity between the target font and the preset font; the target rendering strategy based on font library rendering and the first preset model processing is pre-set, and the target rendering strategy is flexibly determined based on the matching degree between the target font and the preset font and the second threshold, thereby improving the flexibility and intelligence of translation text generation.
[0194] In some embodiments, when the target rendering strategy is the processing strategy of the first preset model, rendering the second text content based on the target rendering strategy to obtain a third text line image includes:
[0195] The font of the first text content and the second text content are input into a first preset model, and the font of the first text content is applied to the second text content using the first preset model to obtain a third text line image.
[0196] It should be noted that a first preset model for font style transfer can be pre-trained, and the font of the first text content and the second text content can be input into the first preset model. The first preset model then processes the second text content to obtain a third text line image. In this way, even if the font of the first text content is non-standard, the font of the first text content can be used for the second text content, thereby improving the accuracy of obtaining the target text content.
[0197] Here, the first preset model includes a generator network and a discriminator network.
[0198] The training process of the first preset model is specifically as follows: obtaining training data, wherein the training data includes: a source font image and a real image of the target font, and the size of the input image of the first preset model can be unified as 256×256; a deep neural network can be determined as a generator network, which is used to transfer the font style of the target font to the source font image to obtain a generated image of the target font, wherein the structure of the generator network includes a downsampling layer or an upsampling layer, etc.; a two-classification network is determined as a discriminator network, which is used to distinguish whether the input image (i.e., the image input by the generator) is a real image of the target font or a generated image generated by the generator network; the loss function of the generator network measures the difference between the real image and the generated image, and the loss function of the discriminator network measures the accuracy of the discriminator network classification. During the actual training process, the generator network takes the source font image and the target font as conditional inputs, uses the downsampling layer to extract shallow font features, and uses the upsampling layer to restore the feature vector to an image, generating a generated image of the target font to imitate the font style of the target font to "trick" the discriminator. At the same time, the discriminator is used for adversarial training. With the help of the idea of adversarial network training, the network parameters of the generator network are optimized until the preset convergence conditions are reached. For example, the difference between the real image and the generated image is obtained based on the style loss function, and the formula of the style loss function is:
[0199] L style (G)=||S(G(x))-S(y)||2 (6);
[0200] In formula (6), G(x) is the generated image, S is the style extraction function, y is the font of the first text content, and || ||2 represents the Euclidean distance.
[0201] In some embodiments, when the font of the first text content is relatively complex, the font style of the first text content may be reversely generated, thereby reversely mapping the font style to the second text content to obtain the target text content.
[0202] In the embodiment of the present disclosure, the font of the first text content can be used in the second text content through the first preset model, so that even if the font of the first text content is a non-standard font, the font style label of the first text content can be migrated to the second text content, which is conducive to improving the accuracy of the target text content and thereby improving the flexibility and intelligence of text generation.
[0203] Figure 2 The process of an image processing method according to an exemplary embodiment is shown as follows Figure 2 ,like Figure 2 As shown, the method mainly includes the following steps:
[0204] In step 201 , in response to a translation instruction for content in an image, text recognition is performed on a first text line image in an image to obtain first text content in a first language.
[0205] In step 202, the first text content is recognized and the font of the first text content is determined.
[0206] In step 203, the first text content is eliminated to obtain a background image of the first text line image.
[0207] In step 204 , when the translation instruction indicates that the first text content is to be translated into a second language, the first text content is translated based on a translation strategy of the second language to obtain a second text content.
[0208] In some embodiments, when the translation instruction indicates that the first text content is to be translated into a second language, a first prompt word is determined from a prompt word template library based on the second language; when the first text content includes a key text unit, a second prompt word of the key text unit is determined from the prompt word template library; wherein the key text unit is used to indicate the field to which the first text content belongs; then, based on the language features indicated by the first prompt word, the first text content is translated to obtain a first translation content; then, based on the terminology information indicated by the second prompt word, the text unit associated with the key text unit in the first text content is translated to obtain a second translation content; finally, based on the first translation content and the second translation content, the second text content is obtained.
[0209] In other embodiments, a third prompt word is generated in response to an instruction to apply a target style template within a preset style template; wherein the preset style template is used to indicate the arrangement style of the translated content; and when the first text content is translated to obtain target translated content arranged according to the target style template, the target translated content is displayed within an edit box according to the reference style information; and in response to editing the target translated content within the edit box, the second text content is obtained. Here, the edit box is located at a preset position in the image.
[0210] For example, Figure 3 is a schematic diagram of a preset style template according to an exemplary embodiment. Figure 3As shown, the first line of the preset style template 30 displays: the coordinates of the first text line area are (110, 50)-(130, 110); the second to third lines display: the original text of the first text content is "The mobile phone has an extra-large display screen and fingerprint recognition function"; the fourth to sixth lines display: the translation of the first text content (i.e., the target translation content) is "The mobile phone has a large display screen and fingerprint recognition function".
[0211] As another example, Figure 4 FIG. 1 is a schematic diagram showing a method of translating the content of an image to be processed according to an exemplary embodiment. Figure 4 As shown, the first text content of the first text line image 41 in the image to be processed 40 is "The mobile phone has an extra-large display screen and a fingerprint recognition function". When the translation instruction for the first text line image 41 is detected, and Figure 3 In the case of an application instruction of the target style template in the target style template, the first text content is translated to obtain target translation content arranged according to the target style template; and the target translation content is displayed in the editing box 42 according to the reference style information.
[0212] In step 205 , a target font is determined based on the font of the first text content and the second language.
[0213] In step 206 , the degree of matching between the target font and the preset fonts in the font library is determined.
[0214] In some embodiments, a feature extraction network is used to extract a first font feature of a target font and a second font feature of a preset font; and based on the similarity between the first font feature and the second font feature, a matching degree between the target font and the preset font is determined.
[0215] In step 207 , it is determined whether there is a matching degree greater than or equal to a second threshold.
[0216] In some embodiments, if it is determined that there is a matching degree greater than or equal to the second threshold, step 208 is performed.
[0217] In some other embodiments, if it is determined that there is no matching degree greater than or equal to the second threshold, step 209 is performed.
[0218] In step 208 , the rendering strategy of the reference font is determined as the target rendering strategy.
[0219] In step 209 , the processing strategy of the first preset model is determined as the target rendering strategy.
[0220] In step 210 , the second text content is rendered based on the target rendering strategy to obtain a third text line image.
[0221] In some embodiments, when the target rendering strategy is the rendering strategy of a reference font, based on a graphics library or a rendering engine, the meta-information of the reference font can be read, such as the font's outline, shape, size, and inclination angle; based on the meta-information, the graphics library or the rendering engine converts the second text content into pixel data, thereby obtaining a third text line image.
[0222] In other embodiments, when the target rendering strategy is the processing strategy of the first preset model, the font of the first text content and the second text content are input into the first preset model, and the generator network of the first preset model is used to apply the font of the first text content to the second text content to obtain a third text line image.
[0223] In step 211 , the third text line image is segmented to obtain target text content.
[0224] In step 212 , a second text line image is generated based on the target text content and the background image.
[0225] In step 213 , a target image is obtained based on the image and the second text line image.
[0226] In some embodiments, the image to be processed can be preprocessed, for example, grayscale, edge detection, etc.; the first text line image is matched with the second text line image so as to maintain the size of the area where the first text line image is located in the preprocessed image to be processed, or the size of the area where the first text line image is located is adjusted, for example, the area where the first text line image is located is expanded; the area where the unadjusted or adjusted first text line image is located in the preprocessed image to be processed is cut to obtain the first image; the second text line image is then fused with the first image, and the fused image is optimized through the PatchGAN model to eliminate edge traces, thereby obtaining the target image.
[0227] The technical solution disclosed herein, in response to a translation instruction for content in an image, recognizes a first text line image in the image and obtains first text content in a first language and style information of the first text line image. This facilitates translation of the first text content based on the translation instruction, meets the user's translation needs, and helps accurately capture the contextual relevance of the first text content. Furthermore, based on the style information, it can lay the foundation for reducing translation traces in the first text line image. When the translation instruction instructs to translate the first text content into a second language, the first text content is translated based on a translation strategy for the second language to obtain second text content. Because the translation strategy matches the second language, the obtained second text content has natural linguistic expression, thereby improving translation accuracy. Furthermore, based on the second text content and the style information, a second text line image is generated, such that the style information of the second text line image matches the style information of the first text line image. A target image is then obtained based on the image and the second text line image. This not only ensures the image quality of the target image but is also applicable to a variety of cross-language image-to-text translation scenarios, such as advertising poster design, cultural work translation, or public sign translation, thereby significantly improving the accuracy, display effect, and intelligence level of image-to-text translation.
[0228] Figure 5 FIG. 1 is a block diagram of an image processing apparatus according to an exemplary embodiment. Figure 5 As shown, the image processing device 500 mainly includes:
[0229] A first acquisition module 501 is configured to, in response to a translation instruction for content in an image, recognize a first text line image in the image to obtain first text content in a first language and style information of the first text line image;
[0230] A translation module 502 is configured to, when the translation instruction indicates that the first text content is to be translated into a second language, translate the first text content based on a translation strategy of the second language to obtain a second text content;
[0231] A generating module 503 is configured to generate a second text line image based on the second text content and the style information of the first text line image;
[0232] The synthesis module 504 is configured to obtain a target image based on the image and the second text line image.
[0233] In some embodiments, the translation module 502 is specifically configured to:
[0234] When the translation instruction indicates that the first text content is to be translated into the second language, determining a first prompt word in the second language from a prompt word template library;
[0235] In the case where the first text content includes a key text unit, determining a second prompt word for the key text unit from the prompt word template library; wherein the key text unit is used to indicate the field to which the first text content belongs;
[0236] The first text content is translated based on the first prompt word and / or the second prompt word to obtain the second text content.
[0237] In some embodiments, the translation module 502 is further configured to:
[0238] translating the first text content based on the language feature indicated by the first prompt word to obtain first translated content;
[0239] Based on the term information indicated by the second prompt word, translating the text unit associated with the key text unit in the first text content to obtain second translated content;
[0240] The second text content is obtained based on the first translation content and / or the second translation content.
[0241] In some embodiments, the translation module 502 is further configured to:
[0242] Determining language similarity between a text unit of the first text content and a reference text unit in the second language in the first prompt word, and determining a text unit having the language similarity greater than or equal to a first threshold as a first text unit;
[0243] Based on the language feature indicated by the first prompt word, the text units in the first text content except the first text unit are translated to obtain the first translated content.
[0244] In some embodiments, the apparatus 500 further includes:
[0245] an output module configured to generate a third prompt word in response to an application instruction of a target style template in a preset style template; wherein the preset style template is used to indicate an arrangement style of the translated content;
[0246] The translation module 502 is further configured to:
[0247] The first text content is translated based on the first prompt word, the second prompt word, and the third prompt word to obtain the second text content.
[0248] In some embodiments, the first acquisition module 501 is specifically configured to:
[0249] Performing text recognition on the first text line image to obtain the first text content;
[0250] performing elimination processing on the first text content to obtain a background image of the first text line image;
[0251] The generating module 503 is specifically configured to:
[0252] determining a target font based on the second language and the font of the first text content;
[0253] Determining target text content based on the target font and the second text content;
[0254] The second text line image is generated based on the target text content and the background image.
[0255] In some embodiments, the translation module 502 is further configured to:
[0256] When the first text content is translated to obtain target translation content arranged according to the target style template, the target translation content is displayed in an editing box according to the reference style information; wherein the editing box is located at a preset position of the image;
[0257] In response to an editing operation on the target translation content in the editing box, obtaining the second text content;
[0258] The generating module 503 is further configured to:
[0259] In response to a triggering operation on the edit box, the target text content is determined based on the target font and the second text content.
[0260] In some embodiments, the generating module 503 is further configured to:
[0261] Determining a degree of matching between the target font and a preset font in a font library, and determining a target rendering strategy based on the degree of matching;
[0262] Rendering the second text content based on the target rendering strategy to obtain a third text line image;
[0263] The third text line image is segmented to obtain the target text content.
[0264] In some embodiments, the generating module 503 is further configured to:
[0265] Extracting a first font feature of the target font and a second font feature of the preset font using a feature extraction network;
[0266] Determining a matching degree between the target font and the preset font based on a similarity between the first font feature and the second font feature;
[0267] In the case of a preset font in the font library whose matching degree with the target font is greater than or equal to a second threshold, determining the rendering strategy of the preset font as the target rendering strategy;
[0268] When the matching degree between the preset font and the target font in the font library is less than the second threshold, the processing strategy of the first preset model is determined as the target rendering strategy.
[0269] In some embodiments, when the target rendering strategy is the processing strategy of the first preset model, the generating module 503 is further configured to:
[0270] The font of the first text content and the second text content are input into the first preset model, and the font of the first text content is applied to the second text content using the generator network of the first preset model to obtain the third text line image.
[0271] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0272] Based on the same inventive concept, an embodiment of the present disclosure provides an electronic device, which may be a computer or a terminal in one or more of the above embodiments. Figure 6 FIG. 1 is a schematic diagram showing the structure of an electronic device according to an exemplary embodiment. Figure 6 As shown, the electronic device 600 adopts general computer hardware and includes a processor 601 , a memory 602 , a bus 603 , an input device 604 , and an output device 605 .
[0273] In some possible implementations, the memory 602 may include computer storage media in the form of volatile and / or non-volatile memory, such as read-only memory and / or random access memory. The memory 602 may store an operating system, application programs, other program modules, executable code, program data, user data, and the like.
[0274] Input device 604 can be used to input commands and information to the electronic device, such as a keyboard or pointing device, such as a mouse, trackball, touchpad, microphone, joystick, game pad, satellite TV antenna, scanner, or similar device. Input device 604 can be connected to processor 601 via bus 603.
[0275] The output device 605 can be used to output information to the electronic device 600. In addition to the monitor, the output device 605 can also be other peripheral output devices, such as speakers and / or printing devices. The output device 605 can also be connected to the processor 601 through the bus 603.
[0276] The electronic device 600 may be connected to a network, such as a local area network (LAN), via an antenna 606. In a networked environment, executable instructions may be stored in a remote storage device, rather than being limited to local storage.
[0277] When the processor 601 in the electronic device 600 executes the executable code or application stored in the memory 602, the electronic device 600 can implement the file processing method in the above embodiment. The specific execution process can be found in the above embodiment and will not be repeated here.
[0278] The memory 602 may store information for implementing Figure 5 The executable instructions of the functions of the first acquisition module 501, the translation module 502, the generation module 503 and the synthesis module 504. Figure 5 The functions / implementation processes of the first acquisition module 501, the translation module 502, the generation module 503 and the synthesis module 504 can be realized by Figure 6 The processor 601 in the embodiment calls the executable instructions stored in the memory 602 to implement the above-mentioned embodiment. For specific implementation process and functions, please refer to the above-mentioned related embodiments.
[0279] Based on the same inventive concept, an embodiment of the present disclosure further provides a storage medium having instructions stored therein. When the instructions are executed on a computer, the instructions are used to execute the image processing method in one or more of the above embodiments.
[0280] Based on the same inventive concept, the embodiments of the present disclosure further provide a computer program or a computer program product. When the computer program product is executed on a computer, the computer implements the image processing method in one or more of the above embodiments.
[0281] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.
[0282] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An image processing method, characterized in that: The method comprises: In response to a translation instruction for content in an image, recognizing a first text line image in the image to obtain first text content in a first language and style information of the first text line image; When the translation instruction indicates that the first text content is to be translated into a second language, translating the first text content based on a translation strategy for the second language to obtain a second text content; generating a second text line image based on the second text content and the style information; A target image is obtained based on the image and the second text line image.
2. The method according to claim 1, characterized in that When the translation instruction indicates that the first text content is to be translated into a second language, translating the first text content based on the translation strategy of the second language to obtain the second text content includes: When the translation instruction indicates that the first text content is to be translated into the second language, determining a first prompt word in the second language from a prompt word template library; In the case where the first text content includes a key text unit, determining a second prompt word for the key text unit from the prompt word template library; wherein the key text unit is used to indicate the field to which the first text content belongs; The first text content is translated based on the first prompt word and / or the second prompt word to obtain the second text content.
3. The method according to claim 2, characterized in that The translating the first text content based on the first prompt word and / or the second prompt word to obtain the second text content includes: translating the first text content based on the language feature indicated by the first prompt word to obtain first translated content; Based on the term information indicated by the second prompt word, translating the text unit associated with the key text unit in the first text content to obtain second translated content; The second text content is obtained based on the first translation content and / or the second translation content.
4. The method according to claim 3, characterized in that The translating the first text content based on the language feature indicated by the first prompt word to obtain first translated content includes: Determining language similarity between a text unit of the first text content and a reference text unit in the second language in the first prompt word, and determining a text unit having the language similarity greater than or equal to a first threshold as a first text unit; Based on the language feature indicated by the first prompt word, the text units in the first text content except the first text unit are translated to obtain the first translated content.
5. The method according to claim 2, characterized in that The method further comprises: generating a third prompt word in response to an application instruction for a target style template in a preset style template; wherein the preset style template is used to indicate an arrangement style of the translated content; The first text content is translated based on the first prompt word, the second prompt word, and the third prompt word to obtain the second text content.
6. The method according to claim 5, characterized in that The step of recognizing the first text line image in the image to obtain first text content in a first language and style information of the first text line image includes: Performing text recognition on the first text line image to obtain the first text content; performing elimination processing on the first text content to obtain a background image of the first text line image; The generating a second text line image based on the second text content and the style information includes: determining a target font based on the second language and the font of the first text content; Determining target text content based on the target font and the second text content; The second text line image is generated based on the target text content and the background image.
7. The method according to claim 6, characterized in that The translating the first text content based on the first prompt word, the second prompt word, and the third prompt word to obtain the second text content includes: When the first text content is translated to obtain target translation content arranged according to the target style template, the target translation content is displayed in an editing box according to the reference style information; wherein the editing box is located at a preset position of the image; In response to an editing operation on the target translation content in the editing box, obtaining the second text content; The determining of target text content based on the target font and the second text content includes: In response to a triggering operation on the edit box, the target text content is determined based on the target font and the second text content.
8. The method according to claim 6, characterized in that The determining the target text content based on the target font and the second text content includes: Determining a degree of matching between the target font and a preset font in a font library, and determining a target rendering strategy based on the degree of matching; Rendering the second text content based on the target rendering strategy to obtain a third text line image; The third text line image is segmented to obtain the target text content.
9. The method according to claim 8, characterized in that Determining the degree of matching between the target font and a preset font in a font library includes: Extracting a first font feature of the target font and a second font feature of the preset font using a feature extraction network; Determining a matching degree between the target font and the preset font based on a similarity between the first font feature and the second font feature; Determining a target rendering strategy based on the matching degree includes: When the matching degree between the preset font and the target font in the font library is greater than or equal to a second threshold, determining the rendering strategy of the preset font as the target rendering strategy; When the matching degree between the preset font and the target font in the font library is less than the second threshold, the processing strategy of the first preset model is determined as the target rendering strategy.
10. The method according to claim 9, characterized in that In a case where the target rendering strategy is the processing strategy of the first preset model, rendering the second text content based on the target rendering strategy to obtain a third text line image includes: The font of the first text content and the second text content are input into the first preset model, and the font of the first text content is applied to the second text content using the first preset model to obtain the third text line image.
11. An image processing device, characterized in that: include: a first acquisition module configured to, in response to a translation instruction for content in an image, recognize a first text line image in the image and obtain first text content in a first language and style information of the first text line image; a translation module configured to, when the translation instruction indicates that the first text content is to be translated into a second language, translate the first text content based on a translation strategy of the second language to obtain a second text content; a generating module configured to generate a second text line image based on the second text content and style information of the first text line image; The synthesis module is configured to obtain a target image based on the image and the second text line image.
12. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
A method and device for processing image
CN111783508A
Term translation method and device after machine translation, equipment and storage medium
CN112364669A
Scene image-text generation method and system
CN114359033A
Medical machine translation method and device and electronic equipment
CN119167951A
Cited By
Model training method, translation method, electronic device and computer program product
CN121071484A