Electronic device, method, and computer-readable storage medium for processing image including characters
By employing OCR to distinguish between paragraph and line-based character blocks, the electronic device minimizes translation errors and ensures consistent rendering, improving user experience through accurate and visually coherent character display.
Patent Information
- Application Number
- PCT/KR2025/095302
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-26
- Filing Date
- 2025-05-02
- Publication Date
- 2026-01-15
AI Technical Summary
Existing character translation methods in electronic devices often result in errors due to incorrect handling of paragraph and line-based translations, leading to inconsistent character sizes and content discrepancies.
The electronic device employs optical character recognition (OCR) to identify blocks of characters, determining whether they form paragraphs or lines, and applies paragraph-based or line-based rendering accordingly to maintain context and alignment.
This approach reduces translation errors by ensuring consistent character size and content accuracy, enhancing user experience by maintaining semantic unity and visual coherence.
Smart Images

Figure KR2025095302_15012026_PF_FP_ABST
Abstract
Description
Electronic device, method, and computer-readable storage medium for processing images containing characters
[0001] The present disclosure relates to an electronic device, a method, and a computer-readable storage medium for processing an image including characters.
[0002] An electronic device can identify characters contained within an image. These characters may be written in a specific language. The electronic device can translate these characters into characters in another language. By translating these characters, the electronic device can generate characters in another language.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.
[0004] An electronic device is described. The electronic device may include at least one processor storing instructions, the electronic device including a memory, a display, and a processing circuit, the memory including one or more storage media. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input for translating characters within an image displayed through the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform optical character recognition (OCR) on the image based on the input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a block including first characters of a first language located in a plurality of lines within the image based on the OCR. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine whether an arrangement of the first characters included within the block satisfies a reference condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain second characters of a second language translated from the first characters based on the arrangement of the first characters satisfying the criterion condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the second characters within the image by performing paragraph-based rendering on the second characters.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain third characters of the second language translated from the first characters based on the arrangement of the first characters that do not satisfy the criterion condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the third characters within the image by performing line-based rendering on the third characters.
[0005] A method is described. The method can be performed in an electronic device including a display. The method can include receiving an input for translating characters within an image displayed through the display. The method can include performing optical character recognition (OCR) on the image based on the input. The method can include obtaining a block including first characters of a first language located on a plurality of lines within the image based on the OCR. The method can include identifying whether an arrangement of the first characters within the block satisfies a reference condition. The method can include obtaining second characters of a second language translated from the first characters based on an arrangement of the first characters satisfying the reference condition. The method can include displaying the second characters within the image by performing paragraph-based rendering on the second characters. The method may include an operation of obtaining third characters of the second language translated from the first characters based on the arrangement of the first characters that do not satisfy the reference condition. The method may include an operation of displaying the third characters within the image by performing line-based rendering on the third characters.
[0006] A non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device including a display, cause the electronic device to receive an input for translating characters within an image displayed through the display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform optical character recognition (OCR) on the image based on the input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a block including first characters of a first language positioned in a plurality of lines within the image based on the OCR. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify whether an arrangement of the first characters included within the block satisfies a reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain second characters of a second language translated from the first characters based on the arrangement of the first characters that satisfy the reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display the second characters within the image by performing paragraph-based rendering on the second characters.The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain third characters of the second language translated from the first characters based on the arrangement of the first characters that do not satisfy the reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display the third characters within the image by performing line-based rendering on the third characters.
[0007] An electronic device is described. The electronic device may include at least one processor storing instructions, the electronic device including a memory, a display, and a processing circuit, the memory including one or more storage media. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input for translating characters within an image displayed through the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform optical character recognition (OCR) on the image based on the input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a block including first characters of a first language located in a plurality of lines within the image based on the OCR. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine whether an arrangement of the first characters included within the block satisfies a reference condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain second characters of a second language translated from the first characters based on the arrangement of the first characters satisfying the criterion condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to change the first characters to the second characters by displaying a paragraph including the second characters positioned over the first characters in the image displayed through the display.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain third characters of the second language translated from the first characters based on the arrangement of the first characters that do not satisfy the criterion condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to change the first characters to the second characters by displaying the third characters assigned to the lines of the image displayed through the display.
[0008] A method is described. The method can be performed in an electronic device including a display. The method can include receiving an input for translating characters within an image displayed through the display. The method can include performing optical character recognition (OCR) on the image based on the input. The method can include obtaining a block including first characters of a first language positioned on a plurality of lines within the image based on the OCR. The method can include identifying whether an arrangement of the first characters included within the block satisfies a reference condition. The method can include obtaining second characters of a second language translated from the first characters based on the arrangement of the first characters satisfying the reference condition. The method can include changing the first characters to the second characters by displaying a paragraph including the second characters positioned on the first characters within the image displayed through the display. The method may include an operation of obtaining third characters of the second language translated from the first characters based on the arrangement of the first characters that do not satisfy the criterion condition. The method may include an operation of changing the first characters into the second characters by displaying the third characters assigned to the lines of the image displayed through the display.
[0009] A non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device including a display, cause the electronic device to receive an input for translating characters within an image displayed through the display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform optical character recognition (OCR) on the image based on the input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a block including first characters of a first language positioned in a plurality of lines within the image based on the OCR. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify whether an arrangement of the first characters included within the block satisfies a reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain second characters of a second language translated from the first characters based on the arrangement of the first characters that satisfy the reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to change the first characters to the second characters by displaying a paragraph including the second characters positioned on the first characters in the image displayed through the display.The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain third characters of the second language translated from the first characters based on the arrangement of the first characters that do not satisfy the reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to change the first characters to the second characters by displaying the third characters assigned to the lines of the image displayed through the display.
[0010] Figures 1a and 1b illustrate examples of characters displayed according to optical character recognition (OCR).
[0011] Figure 2 is a simplified block diagram of an electronic device.
[0012] FIG. 3 is a flowchart illustrating exemplary operations of an electronic device for performing translation according to an embodiment of the present disclosure.
[0013] Figures 4a and 4b illustrate examples of reference conditions.
[0014] Figure 5 illustrates an exemplary method for adding a newline character.
[0015] Figure 6 shows an example of an area in an image from which characters have been removed.
[0016] Figure 7 shows examples of characters displayed according to paragraph-based rendering.
[0017] Figure 8 illustrates examples of characters displayed according to line-based rendering.
[0018] Figure 9 illustrates examples of operations of at least one processor.
[0019] Figure 10a illustrates an example of the operations of a paragraph-based rendering unit.
[0020] Figure 10b illustrates an example of the operations of a line-based rendering unit.
[0021] Figure 11 illustrates examples of paragraph-based rendering and line-based rendering performed within a single block.
[0022] FIG. 12 is a block diagram of an electronic device within a network environment according to various embodiments.
[0023] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.
[0024] Figures 1a and 1b illustrate examples of characters displayed according to optical character recognition (OCR).
[0025] Referring to FIGS. 1A and 1B , an electronic device (100) may be described as a device available for performing OCR. The electronic device (100) may be one of various forms of mobile devices, such as smartphones (e.g., bar-type smartphones, foldable-type smartphones, or rollable-type smartphones), tablets, wearable devices, cellular phones, laptops, and / or other similar computing devices, having various form factors that include circuits (or circuitry) for providing an operation for performing OCR.
[0026] For convenience of explanation, embodiments of the present disclosure are described using OCR, but it will be apparent to those skilled in the art that various methods for identifying text (character or characters) from an image (e.g., scene text recognition) may be substituted.
[0027] The electronic device (100) may include a display (e.g., display (230) of FIG. 2). The display may be used to display an image.
[0028] According to one embodiment, referring to FIG. 1A, a state (101) may be described as a state in which an image (103) is displayed. Within the state (101), the electronic device (100) may display an image (103) including at least one first character (104) constituting at least one paragraph through the display.
[0029] In one embodiment, the electronic device (100) may identify whether a character (e.g., first characters (104)) is included in the image (103) based on displaying the image (103). For example, the electronic device (100) may display an executable object (105) related to a translation of the characters in the image (103) based on identifying that a character is included in the image (103).
[0030] According to one embodiment, the electronic device (100) may receive an input (106) for an executable object (105). The input (106) may be described as an input for translation of characters within an image (103). The input (106) may include a touch input having a contact point on the executable object (105).
[0031] According to one embodiment, based on the input (106), the electronic device (100) can transition from the state (101) to the state (107). The state (107) can be described as a state in which a block including first characters (104) is acquired. The electronic device (100) can perform optical character recognition (OCR) on the image (103) based on the input (106). According to the OCR, the electronic device (100) can acquire a block (108) including first characters (104) of a first language located in a plurality of lines (109) within the image (103). The block (108) can be composed of lines (109). Each of the lines (109) can be composed of at least one sentence or a part of a sentence. The sentences can be composed of first characters (104). In one embodiment, the lines (109) may be composed of a plurality of sentences that form (or can be understood as) at least one paragraph (or paragraphs). For example, a paragraph may correspond to a single semantic unit. In one embodiment, the lines (109) may be composed of at least one word, phrase, or sentence, each line forming (or can be understood as) a distinct item (e.g., a name, a title, an outline, a step, a component). For example, a single word (or phrase, or a sentence) may correspond to a single semantic unit. In one embodiment, depending on the size of the displayed area, a single item may be displayed across multiple lines. In one embodiment, except when the characters (104) are identified as paragraphs, the words, phrases, or sentences that make up the characters (104) may each be considered to correspond to a distinct item.
[0032] According to one embodiment, the electronic device (100) may perform a translation of the first characters (104) within the block (108). The first characters (104) may be recognized as constituting a plurality of sentences forming a single paragraph. In this case, for example, lines (109) within the block (e.g., block (108)) may together form a single semantic unit. The electronic device (100) may obtain second characters (112) of a second language by performing a translation of the first characters (104) constituting the paragraph according to the lines (109). Since the electronic device (100) performs a line-based translation of the first characters (104) constituting the paragraph, the context of the entire first characters (104) may be different from the context of the entire second characters (112). The first characters (104) may be mistranslated because the translation is performed by treating the lines (109) that form a single meaning unit (e.g., a paragraph) together or the meaning unit (e.g., a phrase, a sentence) that continues across at least two lines (e.g., "feet in space") as being separated into each line (109-1, 109-2, 109-3, 109-4, 109-5). The context of each of the lines (109-1, 109-2, 109-3, 109-4, 109-5) on which the first characters (104) are located may correspond to the context of each of the lines (111-1, 111-2, 111-3, 111-4, 111-5) on which the second characters (112) are located.
[0033] In one embodiment, the electronic device (100) can perform line-based rendering on the second characters (112). By performing the line-based rendering, the electronic device (100) can transition from a state (107) to a state (110) in which the second characters (112) can be displayed within the image (103).
[0034] In one embodiment, the image (103) may be described as an image including first characters (104). The electronic device (100) may display the second characters (112) within the image (103) by performing line-based rendering on the second characters (112).
[0035] Since the context of the entire first characters (104) is different from the context of the entire second characters (112), the translation of the first characters (104) may contain errors. Since the first characters (104), which should be translated as a single semantic unit (e.g., a paragraph or sentence), are translated as multiple semantic units (e.g., sentences or phrases), the second characters (112) may convey different content than the first characters (104).
[0036] To resolve such errors, the electronic device (100) may perform paragraph-based translation of the first characters (104) constituting a paragraph. According to embodiments of the present disclosure, the electronic device (100) may identify whether characters (e.g., the first characters (104)) within an image (e.g., the image (103)) constitute a paragraph, in order to perform paragraph-based translation of the first characters (104) constituting a paragraph.
[0037] Since the first characters (104) form a paragraph, the second characters (112) displayed by performing line-based rendering on the first characters (104) may not be evenly distributed within the image (103) depending on the length of the lines (111). If the first characters (104) forming a paragraph are tilted within the image (103), the first characters (104) may have different perspectives and thus different sizes. Since the first characters (104) have different sizes depending on the lines (109), the second characters (112) displayed by performing line-based rendering may have different sizes depending on the lines (111). Since the second characters (112) are displayed in different sizes depending on the lines (111), the user may feel uncomfortable that the second characters (112) appear to have different sizes depending on the lines (111). A method may be required to resolve user inconvenience caused by an electronic device (100) performing line-based rendering for second characters (112).
[0038] To resolve this inconvenience, the electronic device (100) may perform paragraph-based rendering on the second characters (112) based on the first characters (104) that constitute a paragraph. According to embodiments of the present disclosure, the electronic device (100) may identify whether characters (e.g., first characters (104)) within an image (e.g., image (103)) constitute a paragraph in order to perform paragraph-based rendering on the second characters (112).
[0039] The electronic device (100) may execute the operations exemplified in the description of FIGS. 3 to 11 to perform paragraph-based translation for the first characters (104) constituting the paragraph. The electronic device (100) may execute the operations exemplified in the description of FIGS. 3 to 11 to perform paragraph-based rendering for the second characters (112). The electronic device (100) may include components for executing the above operations. The above components may be exemplified in the description of FIG. 2.
[0040] Referring to FIG. 1B, a state (113) can be described as a state in which an image (114) is displayed. Within the state (113), the electronic device (100) can display the image (114) through a display. The image (114) can include first characters (115). The first characters (115) may not be composed of paragraphs but may be composed of individual lines. In other words, the first characters (115) can be recognized as constituting a plurality of words or phrases that form distinct items (e.g., names) in each line (or area). In this case, for example, each line (or area) within a block (e.g., block (117)) can form its own semantic unit.
[0041] In one embodiment, the electronic device (100) may identify whether a character (e.g., first characters (115)) is included in the image (114) based on displaying the image (114). For example, the electronic device (100) may display an executable object (105) through the display based on identifying that a character is included in the image (114). The executable object (105) may be associated with a translation of the characters in the image (114).
[0042] According to one embodiment, the electronic device (100) may receive an input (106) for an executable object (105). The input (106) may be described as an input for translation of characters within an image (103). The input (106) may include a touch input having a contact point on the executable object (105).
[0043] Based on the input (106), the electronic device (100) can transition from state (113) to state (116). State (113) can be described as a state in which a block (117) including first characters (115) is acquired. The electronic device (100) can perform optical character recognition (OCR) on the image (114). According to the OCR, the electronic device (100) can acquire a block (117) including first characters (115) of a first language located in a plurality of lines (118) within the image (114).
[0044] The electronic device (100) can perform a translation of the first characters (115) within the block (117). The electronic device (100) can obtain third characters (120) of the second language by performing a translation of the first characters (115) constituting lines (118) based on paragraphs. Since the electronic device (100) performs a paragraph-based translation of the first characters (115) constituting the lines, the context of the first characters (115) may be different from the context of the third characters (120). The context of the lines (118) may be different from the overall context of the first characters (115). The context of the first characters (115) located in the lines (118) may be different from the context of the third characters (120). The first characters (115) may be mistranslated because the distinct meaning units formed by each line (118-1, 118-2, 118-3, 118-4, 118-5) are processed as a single meaning unit (e.g., “history tour hamburger”) and the translation is performed.
[0045] In one embodiment, the electronic device (100) can perform paragraph-based rendering for the third characters (120). By performing paragraph-based rendering, the electronic device (100) can transition from a state (116) to a state (119) in which the third characters (120) can be displayed within the image (114).
[0046] In one embodiment, the image (114) may be described as an image including first characters (115). The electronic device (100) may display the third characters (120) within the image (114) by performing paragraph-based rendering for the third characters (120).
[0047] Since the context of the entire first characters (115) is different from the context of the entire third characters (120), the translation of the first characters (115) may have errors. Since the first characters (115), which should be translated as separate meaning units (e.g., sentences or phrases) for each line, are translated as a single meaning unit (e.g., paragraphs or sentences), the third characters (120) may convey different content from the first characters (115). To resolve such errors, the electronic device (100) may perform line-based translation of the first characters (115) constituting the lines. According to one embodiment, the electronic device (100) may identify whether characters (e.g., first characters (115)) within an image (e.g., image (114)) constitute a paragraph or whether each line constitutes distinct content (or items) to perform line-based translation of the first characters (115) constituting the lines.
[0048] The first characters (104) constituting a paragraph may be arranged unevenly within a block (117). The third characters (120) displayed by performing paragraph-based rendering may be arranged evenly within the block (117). Since the third characters (120) are arranged evenly within the block (117), the third characters (120) may be displayed at a different position from the positions of the first characters (104) that are arranged unevenly within the block (117). Since the third characters (120) are displayed at a different position from the first characters (115) within the block (117), the user may feel uncomfortable. A method for resolving the user's inconvenience caused by performing paragraph-based rendering on the third characters (120) may be required.
[0049] To resolve this inconvenience, the electronic device (100) may perform line-based rendering on the third characters (120) based on the first characters (115) that constitute the lines. To perform line-based rendering on the third characters (120), the electronic device (100) may identify whether the characters (e.g., the first characters (115)) within an image (e.g., the image (114)) constitute a paragraph or whether each line constitutes distinct content (or items).
[0050] According to one embodiment, the electronic device (100) may execute the operations exemplified in the description of FIGS. 3 to 11 to perform line-based translation for the first characters (115) constituting the lines. The electronic device (100) may execute the operations exemplified in the description of FIGS. 3 to 11 to perform line-based rendering for the third characters (120). The electronic device (100) may include components for executing the above operations. The components may be exemplified in the description of FIG. 2.
[0051] Figure 2 is a simplified block diagram of an electronic device.
[0052] Referring to FIG. 2, according to one embodiment, the electronic device (200) may be one of various forms of mobile devices, such as smartphones having various form factors (e.g., bar-type smartphones, foldable-type smartphones, or rollable-type smartphones), tablets, wearable devices, cellular phones, laptops, and / or other similar computing devices. The electronic device (200) may include the electronic device (100) of FIGS. 1A and 1B , or may correspond to the electronic device (100) of FIGS. 1A and 1B . The electronic device (200) may include at least a portion of the electronic device (1201) of FIG. 12 , or may correspond to at least a portion of the electronic device (1201) of FIG. 12 . The electronic device (100) may include at least one processor (210), a memory (220), and a display (230).
[0053] At least one processor (210) may include processing circuitry. At least one processor (210) may include a central processing unit (CPU) (e.g., including processing circuitry). At least one processor (210) may include a graphic processing unit (GPU) (e.g., including processing circuitry) and / or a neural processing unit (NPU) (e.g., including processing circuitry). At least one processor (210) may be described as an application processor. At least one processor (210) may include an artificial intelligence-based processor. At least one processor (210) may be configured to control a memory (220) and a display (230). At least one processor (210) may be configured to individually or collectively execute instructions stored in the memory (220) to cause the electronic device (200) to perform at least some of the operations illustrated in the descriptions of FIGS. 1A and 1B . At least one processor (210) may be configured to execute instructions stored in the memory (220) to cause the electronic device (200) to perform at least some of the operations illustrated in the descriptions of FIGS. 3 through 11.
[0054] The memory (220) may include one or more storage media. The memory (220) may store various data used by at least one component of the electronic device (200) (e.g., at least one processor (210) and / or display (230)). The data may include input data or output data for software and commands related thereto. The memory (220) may include volatile memory or non-volatile memory.
[0055] The display (230) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the strength of a force generated by a touch. The display (230) may further include a structure capable of detecting an input using a stylus pen, such as an electromagnetic resonance (EMR) or an active electrostatic solution (AES) type. The display (230) may be configured to display an image. The display (230) may be configured to display characters within the image. The display (230) may be configured to receive an input for translating characters within the image.
[0056] The electronic device (200) illustrated in the description of FIG. 2 can execute at least some of the operations illustrated in the description of FIGS. 3 to 11. The operations illustrated in the description of FIGS. 3 to 11 can be caused by (or within) the electronic device (200) under the control of at least one processor (210).
[0057] The electronic device (200) illustrated in the description of FIG. 2 may perform at least some of the operations illustrated in the descriptions of FIGS. 3 to 11 through an artificial intelligence model (or a deep learning model, or a machine learning model). The artificial intelligence model may be located in the electronic device (200). An artificial intelligence model located in a server may replace at least some of the functions of the artificial intelligence model stored in the electronic device (200). The artificial intelligence model of the present disclosure may be a large language model (LLM) or a large vision model (LVM). An input for requesting a translation to the artificial intelligence model may be referred to as a prompt. Text and / or images obtained as OCR results may be transmitted to the artificial intelligence model as part of the prompt.
[0058] The artificial intelligence model or machine learning model of the present disclosure may be a model trained to identify the type of language to be translated based on context. The artificial intelligence model of the present disclosure may be a model trained to merge two input contents (e.g., text and image) to create a single image. The artificial intelligence model of the present disclosure may be a model trained to identify the type of content (e.g., map, document, sign, recipe, assembly instructions, operation instructions) of an input file (e.g., image or video). The artificial intelligence model of the present disclosure may be a model trained to analyze the meaning of words, phrases, and sentences formed by input characters and determine whether it is appropriate to translate them as a single paragraph or to translate each line where the characters are located as a distinct semantic unit. For example, the machine learning model may be trained or learned through sufficient training data for contexts such as activity history, device status, and the current user interface to associate specific contexts with specific tasks and specific action sequences. The machine learning model of the present disclosure may include various transformer models. The operation of the machine learning model of the present disclosure may include a learning and inference process of finding patterns in data, storing them as models that are generalized rules, and inputting new data into the learned model to obtain results.
[0059] For convenience of explanation, the following describes the LLM used in this disclosure, but it is self-evident that the artificial intelligence neural network of this document may include not only a language model, but also various foundation models such as a code model and an image model, as well as other artificial intelligence neural network models.
[0060] The LLM mentioned in this document can refer to a language model based on an artificial neural network that has learned a large amount of text data through pre-training. An LLM can contain a relatively large number of parameters (e.g., over 10 billion) compared to conventional language models. An LLM can utilize a transformer artificial neural network structure based on the attention mechanism.
[0061] The attention mechanism is a technique that helps artificial intelligence models focus (attention) on important parts of input data. The attention mechanism can be used to predict output data by predicting the degree to which a portion of time-series input data (e.g., voice or video input data, or input data from a layer of a neural network) contributes to the intermediate or final output of the neural network. Recurrent neural networks (RNNs), which sequentially process each element of a sequence, exhibit poor prediction performance when there is information dependency between long time-series distances. However, the attention mechanism can consider information dependency between long time-series distances by controlling the degree of weight concentration (attention) within the overall (or partial) context of the input data.
[0062] A Transformer can be structured as an encoder-decoder. The encoder processes input data and outputs compressed information (e.g., a contextual representation), while the decoder processes the compressed information and outputs token-based data. Each encoder and decoder can include an independent attention network, or a cross-attention network connecting the encoder and decoder.
[0063] For example, LLM training may involve pre-training and / or fine-tuning. Pre-training involves training the LLM to acquire general linguistic knowledge using large amounts of text data. For example, this may involve self-supervised learning, where the LLM predicts the next word based on the previous word sequence in the text string. Fine-tuning involves training the LLM to be suitable for a specific domain (e.g., chatbot, translation, summarization, Q&A) or task. The LLM may undergo additional supervised learning (or adaptive learning) based on the pre-trained model using a dataset tailored to the domain's purpose. The LLM can perform tasks with text inputs containing natural language, called prompts.
[0064] For example, fine-tuning can be omitted when learning LLMs. The user can control the prompts provided to the LLM to enhance its performance on the desired task. Similar to in-context learning or zero-shot / few-shot learning, the prompts can provide additional task examples and / or guidance for performing the task. Publicly available LLMs include BERT (Bidirectional Encoder Representations from Transformer) and GPT (generative pre-trained transformer).
[0065] The term "LLM" can refer to the language neural network model itself, but it can also refer to the model of an LLM-based application (e.g., chatbot, translation, summarization, text classification, sentence generation). For example, an LLM-based chatbot like ChatGPT or an LLM-based translator can also be referred to as "LLM."
[0066] "LLM" may also include an inference engine utilizing the LLM neural network model. For example, "inputting an input prompt to the LLM" may mean "inputting the input prompt to an inference engine based on the LLM." For example, "the output of the LLM for the input prompt" may mean the output information of the last neural network layer of the LLM (or output information modified through further processing) obtained when the input prompt is input to the LLM-based inference engine.
[0067] FIG. 3 is a flowchart illustrating exemplary operations of an electronic device for performing translation according to an embodiment of the present disclosure.
[0068] Referring to FIG. 3, in operation 300, at least one processor (210) may display an image (e.g., image (103) of FIG. 1A or image (114) of FIG. 1B) via a display (230). At least one processor (210) may receive an input (e.g., input (106) of FIG. 1A or input (106) of FIG. 1B) for translation of characters (e.g., first characters (104) of FIG. 1A or first characters (115) of FIG. 1B) within the image.
[0069] In operation 310, at least one processor (210) (e.g., module OCR processor (915) of FIG. 9) may perform optical character recognition (OCR) on an image based on an input. By performing OCR, the at least one processor (210) (e.g., module OCR processor (915) of FIG. 9) may extract features of curves, straight lines, and / or intersections of characters within the image using optical techniques. The at least one processor (210) (e.g., module OCR processor (915) of FIG. 9) may identify characters within the image by comparing the features of the characters with a character database within a memory (220).
[0070] At least one processor (210) (e.g., module language analysis processor (925) of FIG. 9) may obtain a block (e.g., block (108) of FIG. 1a or block (117) of FIG. 1b) containing first characters (e.g., first characters (104) of FIG. 1a or first characters (115) of FIG. 1b) located on a plurality of lines (e.g., lines (109) of FIG. 1a) within the image, according to OCR for the image.
[0071] According to one embodiment, at least one processor (210) can display an image (e.g., image (103) of FIG. 1A) including first characters (e.g., first characters (104) of FIG. 1A) via a display (230).
[0072] At least one processor (210) (e.g., image analysis processor (905) of FIG. 9) may identify, based on displaying the image (103), whether a character (e.g., first characters (104)) is included in the image (103). At least one processor (210) (e.g., image analysis processor (905) of FIG. 9) may display, based on identifying, that a character is included in the image (103), an executable object (e.g., 105 of FIG. 1A) via the display (230). The executable object (e.g., 105 of FIG. 1A) may be related to translation of characters in the image (e.g., 103 of FIG. 1A).
[0073] At least one processor (210) can receive an input (e.g., 106 of FIG. 1A) for an executable object (e.g., 105 of FIG. 1A). The input (e.g., 106 of FIG. 1A) can be described as an input for translation of characters within an image (e.g., 103 of FIG. 1A). The input (e.g., 106 of FIG. 1A) can include a touch input for the executable object (e.g., 105 of FIG. 1A). The input (e.g., 106 of FIG. 1A) can include a touch input of tapping the executable object (e.g., 105 of FIG. 1A). The input (106) can include a touch input having a contact point on the executable object (e.g., 105 of FIG. 1A). At least one processor (210) can identify an input (e.g., 106 of FIG. 1A) via a display (230) (e.g., a touchscreen).
[0074] Based on the input (e.g., 106 of FIG. 1A), the electronic device (200) can transition from a state (e.g., 101 of FIG. 1A) to a state (e.g., 107 of FIG. 1A) by obtaining a block containing first characters (e.g., 104 of FIG. 1A). At least one processor (210) (e.g., the OCR processor (915) of FIG. 9) can perform OCR on the image (e.g., 103 of FIG. 1A). At least one processor (210) (e.g., OCR processor (915) of FIG. 9) may obtain a block (e.g., 108 of FIG. 1a) containing first characters (e.g., 104 of FIG. 1a) of a first language located in a plurality of lines (e.g., 109 of FIG. 1a) within an image (e.g., 103 of FIG. 1a), according to OCR.
[0075] According to one embodiment, at least one processor (210) may identify an area containing text information within an image based on OCR of the image. Depending on the arrangement of the text information, at least one processor (210) may acquire the entire area containing text information as a single block or as multiple blocks.
[0076] In operation 320, at least one processor (210) (e.g., the text structure analysis processor (935) of FIG. 9) can identify whether the arrangement of first characters (e.g., the first characters (104) of FIG. 1A) included in a block (e.g., the block (108) of FIG. 1A) satisfies a criteria condition. The criteria condition may relate to whether the first characters (e.g., the first characters (104) of FIG. 1A) constitute a paragraph. The criteria condition may be set to determine whether the first characters (e.g., the first characters (104) of FIG. 1A) constitute a paragraph. Examples of the criteria condition are illustrated in the descriptions of FIGS. 4A and 4B.
[0077] According to one embodiment, at least one processor (210) (e.g., the text structure analysis processor (935) of FIG. 9) can identify one or more words included in lines (e.g., lines (405) of FIG. 4A) within a block (e.g., block (400) of FIG. 4A). The at least one processor (210) (e.g., the text structure analysis processor (935) of FIG. 9) can determine whether the first characters (e.g., the first characters (104) of FIG. 1A) constitute a paragraph by analyzing the first characters (e.g., the first characters (104) of FIG. 1A) based on the number of one or more words included in the lines (e.g., the lines (405) of FIG. 4A). For example, at least one processor (210) (e.g., a text structure analysis processor (935) of FIG. 9) may determine whether first characters (e.g., first characters (104) of FIG. 1A) constitute a paragraph based on a proportion of lines (e.g., lines (405) of FIG. 4A) within a block (e.g., block (400) of FIG. 4A) having a word count greater than or equal to a reference number (e.g., three).
[0078] According to one embodiment, at least one processor (210) (e.g., the text structure analysis processor (935) of FIG. 9) can identify a ratio of a length of a side parallel to lines (e.g., lines (415) of FIG. 4B) of a block (e.g., block (410) of FIG. 4B) to a length of each of the lines. At least one processor (210) (e.g., the text structure analysis processor (935) of FIG. 9) can determine whether the first characters constitute a paragraph by analyzing the first characters based on a ratio of a length of a side parallel to lines (e.g., lines (415) of FIG. 4B) of the block (e.g., block (410) of FIG. 4B) to a length of each of the lines. For example, at least one processor (210) (e.g., text structure analysis processor (935) of FIG. 9) can determine whether the first characters constitute a paragraph based on the ratio of lines in which the ratio of the length of the area where characters are displayed in each line (e.g., lines (415-1, 415-2, 415-3, 415-4) of FIG. 4b) to the width of the block (e.g., block (410) of FIG. 4b) is greater than or equal to a reference ratio (e.g., 80%).
[0079] In one embodiment, the criteria for translation and / or rendering may be determined based on attributes associated with characters included in the image. The attributes associated with the characters may include at least one of punctuation marks, symbols, numbers, types of characters or images, relationships between words, probabilities related to the arrangement of words, or context expressed by the characters. For example, characters may be translated and / or rendered on a paragraph-by-paragraph basis (e.g., block-by-block) or line-by-line basis (e.g., line-by-line) basis, depending on the attributes associated with the characters.
[0080] According to one embodiment, the criteria for determining the method of translation and / or rendering (e.g., the criteria for determining whether to process lines as at least one paragraph unit) may be based at least in part on whether punctuation marks (e.g., periods, commas, quotation marks) included in the characters are identified. For example, if periods are identified between characters, the electronic device may determine to process (e.g., translate and / or render) the characters as paragraph units. For example, if punctuation marks are not identified in the characters, or if punctuation marks are identified below a threshold (e.g., two or fewer), the electronic device (200) may determine the method of translation and / or rendering the characters using the criteria described above (e.g., the percentage of lines with a word count greater than or equal to the threshold and / or the percentage of area occupied by characters).
[0081] According to one embodiment, the criteria for determining the method of translation and / or rendering (e.g., the criteria for determining whether to process lines as at least one paragraph unit) may be determined based at least in part on the type of character, punctuation mark, or symbol identified at the beginning of each line. For example, if each line begins with a number, if the numbers identified at the beginning of each line increase or decrease along the line, or if each line begins with a symbol, the electronic device (200) may determine to process (e.g., translate and / or render) the characters by dividing them into semantic units for each line. For example, if each line does not begin with a number or does not begin with a symbol, the electronic device (200) may determine the method of translation and / or rendering the characters using the criteria described above (e.g., the percentage of lines with a word count greater than or equal to a reference and / or the percentage of area occupied by characters).
[0082] According to one embodiment, the criteria for determining the method of translation and / or rendering (e.g., the criteria for determining whether to process lines as at least one paragraph unit) may be determined based at least in part on the type (or nature) of the text or image to be translated. For example, the electronic device (200) may transmit an image and / or text containing the text to be translated to an artificial intelligence model, thereby obtaining, through the artificial intelligence model, information as to whether the type (or nature) of the text or image to be translated falls into a specific category. The category may include a map, a document, a sign, a recipe, an assembly order, and / or an operation method. For example, if the type (or nature) of the text or image to be translated falls into a document (e.g., a thesis or article), the electronic device (200) may determine to process (e.g., translate and / or render) characters as paragraph units. For example, if the type (or nature) of the text or image to be translated falls into a map or sign, the electronic device (200) may determine to process (e.g., translate and / or render) characters by dividing them into semantic units for each line.
[0083] In one embodiment, the method of translation and / or rendering (e.g., whether to process lines as at least one paragraph unit) may be requested to an AI model for the decision, or the AI model that has been requested to perform the translation may automatically determine (or process) the decision. For example, at least one processor (210) may transmit an image and / or text containing text to be translated to the AI model, thereby obtaining information through the AI model as to whether semantic units (e.g., phrases, sentences) that extend across at least two lines have a continuous relationship with each other or have separate relationships for each line. For example, the AI model (e.g., LLM) may recognize characters within a block and, based on learned linguistic characteristics, determine whether the last character (or word) of one line and the first character (or word) of the next line have a relationship that can be arranged continuously with each other. Depending on the determination result, the lines within at least one block may be processed as a single semantic unit (e.g., translated and / or rendered), or the semantic units may be processed separately for each line (e.g., translated and / or rendered). In one embodiment, the AI model may determine whether to process lines as a single semantic unit (e.g., for translation and / or rendering) or to process each line as a separate semantic unit (e.g., for translation and / or rendering) by understanding the overall context of the content expressed by the provided characters.
[0084] In operation 330, at least one processor (210) (e.g., translation processor (955) of FIG. 9) may perform translation of the first characters based on the arrangement of the first characters (e.g., first characters (104) of FIG. 1A) that satisfy a reference condition. Since the arrangement of the first characters satisfies the reference condition, the first characters may form a paragraph. At least one processor (210) (e.g., translation processor (955) of FIG. 9) may perform paragraph-based translation of the first characters that form a paragraph. At least one processor (210) (e.g., translation processor (955) of FIG. 9) may perform translation of the entire paragraph formed by the first characters. At least one processor (210) (e.g., translation processor (955) of FIG. 9) may receive the entire block at once, thereby recognizing the entire block as a single semantic unit, and perform translation for the entire block. A single phrase or sentence within the lines of a block can be translated as a single phrase or sentence even if it spans at least two lines.
[0085] At least one processor (210) (e.g., translation processor (955) of FIG. 9) can perform paragraph-based translation of the first characters to obtain second characters (e.g., second characters (705) of FIG. 7) of a second language translated from the first characters. The second language may be different from the first language.
[0086] According to one embodiment, the language to be translated from a first language can be automatically determined based on a context identified in an electronic device (e.g., the electronic device (200) of FIG. 2). For example, a history of actions performed by a user over a certain period of time before requesting a translation function can be used as the context. For example, if an electronic device (200) whose default language is set to the first language identifies that the user has viewed a document, email, webpage, or conversation expressed in a second language, or has entered a character in the second language before requesting a translation function, the electronic device (200) can estimate the language requiring translation as the second language and either prioritize translation into the second language or designate the second language as the default language for translation and provide it to the user. The context, which includes the history of actions (e.g., the location of the electronic device, the time the electronic device was used, the type of application being used, the type of application used immediately before, and / or the type of connected network), can be transmitted as a prompt to an artificial intelligence model that performs at least a portion of the translation and rendering operations of the present disclosure, along with a target character and / or a target image for translation.
[0087] In operation 340, at least one processor (210) (e.g., paragraph rendering processor (965) of FIG. 9) may perform paragraph-based rendering on the second characters. The at least one processor (210) (e.g., paragraph rendering processor (965) of FIG. 9) may perform paragraph-based rendering on the second characters by displaying a paragraph including the second characters positioned on the first characters. The at least one processor (210) (e.g., paragraph rendering processor (965) of FIG. 9) may perform paragraph-based rendering by displaying the second characters that replace the first characters and are represented as paragraphs. The at least one processor (210) (e.g., paragraph rendering processor (965) of FIG. 9) may display the second characters within the image by performing paragraph-based rendering on the second characters. At least one processor (210) (e.g., paragraph rendering processor (965) of FIG. 9) can display second characters within the image by overlaying them on the first characters.
[0088] In one embodiment, at least one processor (210) (e.g., OCR processor (915)) can identify locations where first characters are placed within an image by performing OCR. At least one processor (210) (e.g., paragraph rendering processor (965) of FIG. 9) can perform paragraph-based rendering on second characters to delete the first characters within the image and display the second characters at locations corresponding to the locations where the first characters are placed within the image.
[0089] According to one embodiment, the second character and the image may be generated as a single image through an artificial intelligence model and stored and / or displayed through a display (230). For example, a portion of an area where the first character of the image is located may be in-painted using the second character, and another portion of the area where the first character is located may be in-painted using another area of the image. The stored image may be provided so that the user can access it through an image management application (e.g., a gallery application or a photo application) installed on the electronic device (200). The image in which the second character is in-painted through the artificial intelligence model may include, for example, text content for each of the first character and / or the second character as metadata. The electronic device (200) may obtain text content before and / or after translation from the metadata and provide it to the user (e.g., store it on a clipboard, display it as a recommended text).
[0090] According to one embodiment, an image including translated characters may be stored in the memory (220) of the electronic device or in the storage of a server connected to a user account of the electronic device (200). For example, an image including translated characters may be stored for a specified period of time (e.g., one week) and then automatically deleted. For example, the electronic device (200) may manage a translation history for an image / video for which a translation request has been made or a document or webpage including an image for which a translation request has been made, and if it determines that content with a translation history has been retranslated, it may decide to store the image including translated characters. Otherwise, if the screen changes from an image including translated characters to another screen, it may decide to remove the image including translated characters without storing it.
[0091] According to one embodiment, character information (first characters) and translated character information (second characters) extracted through object recognition or character recognition (e.g., OCR) may be stored together with or separately from the original image and / or the image provided after translation as a content type corresponding to the character. For example, the electronic device (200) manages a translation history for an image / video for which a translation request has been made or a document or webpage including an image for which a translation request has been made, and when it is determined that content with a translation history is to be translated again, the electronic device (200) may process a translation request by overlaying the stored translated character information on the original image or merging the stored translated character information with the original image as an image object without a separate translation process. For example, when a translation request is later received for another image related to an image including the first characters (e.g., an image located on the same URL or a related URL, an image including a similar image object, or an image including similar characters), the electronic device (200) may process the translation request using the stored second characters.
[0092] The first characters can comprise multiple paragraphs. If there is a line break between the multiple paragraphs, adding a newline character may be required to perform paragraph-based rendering for the second characters. Adding a newline character is exemplified in the description of Figure 5.
[0093] According to one embodiment, at least one processor (210) (e.g., line break processor (945)) can identify a line break between a first line (e.g., first line (500) of FIG. 5) and a second line based on the length of a first word of the second line (e.g., second line (520) of FIG. 5). At least one processor (210) (e.g., line break processor (945)) can add a newline character at a location corresponding to the end of the first line (e.g., location (545) of FIG. 5) based on identifying the line break between the first line and the second line.
[0094] In operation 320, at least one processor (210) (e.g., the text structure analysis processor (935) of FIG. 9) may identify whether the arrangement of first characters (e.g., the first characters (115) of FIG. 1B) included in a block (e.g., the block (117) of FIG. 1B) satisfies a criteria condition. The criteria condition may be related to whether the first characters constitute a paragraph. For examples of the criteria condition, reference may be made to the descriptions of FIGS. 4A and 4B.
[0095] At least one processor (210) (e.g., a text structure analysis processor (935) of FIG. 9) can identify whether the arrangement of the first characters (115) within the block (117) satisfies a reference condition based on the number of words located in each of the lines (118) within the block (117) of FIG. 1b.
[0096] The lines (118) may contain one or more words composed of first characters (115). At least one processor (210) may identify the number of words in each of the lines (118-1, 118-2, 118-3, 118-4, 118-5).
[0097] In one embodiment, at least one of lines (118-1, 2, 3, 5) may contain at least one word. For example, line (118-1) may contain two words, line (118-2) may contain one word, line (118-3) may contain two words, line (118-4) may contain three words, and line (118-5) may contain two words. At least one processor (210) (e.g., text structure analysis processor (935) of FIG. 9 ) may identify a number of lines among lines (118) that contain less than a reference number of words. As a non-limiting example, the reference number may be three. As a non-limiting example, five of lines (118) (e.g., line (118-1), line (118-2), line (118-3), line (118-4), and line (118-5)) may contain less than a reference number of words. As a non-limiting example, at least one processor (210) may identify a ratio (e.g., a ratio of 5 to 5) of the number of lines within a block (117) to the number of lines containing less than a reference number of words (e.g., three), but is not limited thereto.
[0098] At least one processor (210) (e.g., text structure analysis processor (935) of FIG. 9) can identify that the arrangement of the first characters (115) within the block (117) does not satisfy a reference condition based on a ratio of the number of lines containing words less than a reference number. At least one processor (210) (e.g., text structure analysis processor (935) of FIG. 9) can identify that the first characters (115) within the block (117) do not constitute a paragraph because the arrangement of the first characters (115) within the block (117) does not satisfy the reference condition.
[0099] At least one processor (210) (e.g., the text structure analysis processor (935) of FIG. 9) can identify whether the arrangement of the first characters (115) satisfies the reference condition based on the distribution of blank areas within the block (117). At least one processor (210) (e.g., the text structure analysis processor (935) of FIG. 9) can identify the distribution of the lines (118) within the block (117). Since the lines (118) are irregularly distributed within the block (117), the blank areas within the block (117) may be irregularly distributed. At least one processor (210) (e.g., the text structure analysis processor (935) of FIG. 9) can identify that the arrangement of the first characters (115) does not satisfy the reference condition based on the irregularly distributed blank areas within the block (117).
[0100] In operation 350, at least one processor (210) (e.g., translation processor (955) of FIG. 9) may perform translation of the first characters based on the arrangement of the first characters (e.g., first characters (115) of FIG. 1B) that do not satisfy a reference condition. Since the arrangement of the first characters does not satisfy the reference condition, the first characters may not form paragraphs. Since the arrangement of the first characters does not satisfy the reference condition, the first characters may form lines. At least one processor (210) (e.g., translation processor (955) of FIG. 9) may perform line-based translation of the first characters that form the lines. At least one processor (210) (e.g., translation processor (955) of FIG. 9) may perform translation of each of the lines that the first characters form. At least one processor (210) (e.g., the translation processor (955) of FIG. 9) can receive each line included in a block one at a time, thereby recognizing each line as a single semantic unit and performing translation. The last word of one line and the first word of the next line can be processed and translated as belonging to separate words, phrases, or sentences that are distinct from each other. At least one processor (210) (e.g., the translation processor (955) of FIG. 9) can also receive the lines included in a block all at once and then perform translation for each line.
[0101] At least one processor (210) (e.g., translation processor (955) of FIG. 9) can obtain third characters (e.g., third characters (810) of FIG. 8) of a second language translated from the first characters by performing line-based translation of the first characters. The second language may be different from the first language. At least one processor (210) (e.g., translation processor (955) of FIG. 9) can perform line-based translation of the first characters of a plurality of languages, thereby maintaining characters of the same language as the second language among the first characters and displaying together third characters translated from characters of a language different from the second language among the first characters. According to one embodiment, the first characters may be composed of a plurality of languages. At least one processor (210) can obtain third characters of the second language by performing translation of the first characters composed of a plurality of languages into a predetermined second language (e.g., a default language set in the system of the electronic device (200). At least one processor (210) may display a user interface (UI) for providing a function to change a second language to be used for translating the first characters together with third characters. At least one processor (210) may change the second language to be used for translating the first characters to a third language based on a user input for changing the second language to be used for translating the first characters. At least one processor (210) may obtain fourth characters in the third language by translating the first characters into the third language based on the user input. At least one processor (210) may display the fourth characters through the display (230) by obtaining the fourth characters.
[0102] At operation 360, at least one processor (210) (e.g., the line rendering processor (975) of FIG. 9) can perform line-based rendering for third characters. The at least one processor (210) (e.g., the line rendering processor (975) of FIG. 9) can perform line-based rendering for the third characters by displaying the third characters assigned to lines (e.g., the lines (805) of FIG. 8). The at least one processor (210) (e.g., the line rendering processor (975) of FIG. 9) can perform paragraph-based rendering by replacing the first characters and displaying the third characters separated by line. The at least one processor (210) (e.g., the line rendering processor (975) of FIG. 9) can display the third characters within the image by performing line-based rendering for the third characters. At least one processor (210) (e.g., line rendering processor (975) of FIG. 9) can display third characters within the image by overlaying them on the first characters.
[0103] According to one embodiment, at least one processor (210) (e.g., OCR processor (915) of FIG. 9) may identify locations where first characters are placed within an image by performing OCR. At least one processor (210) may perform line-based rendering on third characters to delete the first characters within the image and display the third characters at locations corresponding to the locations where the first characters are placed within the image. At least one processor (210) (e.g., line rendering processor (975) of FIG. 9) may remove the first characters within the image to display the third characters on top of the first characters. At least one processor (210) (e.g., line rendering processor (975) of FIG. 9) may obtain second areas corresponding to the first areas from which the first characters are removed within the image based on an artificial intelligence model. Obtaining the second areas may refer to the description of FIG. 6.
[0104] At least one processor (210) (e.g., the line rendering processor (975) of FIG. 9) may display third characters within an image by performing line-based rendering on the third characters. Third characters displayed according to line-based rendering are exemplified within the description of FIG. 8.
[0105] Figures 4a and 4b illustrate examples of reference conditions.
[0106] Referring to FIG. 4A, at least one processor (210) can identify whether the arrangement of characters within a block (400, 104) satisfies a reference condition based on the number of words located in each of the lines within the block.
[0107] In one embodiment, a block (400) may include characters within lines (405). The lines (405) may include one or more words composed of characters. At least one processor (210) may identify the number of words in each of the lines (405-1, 405-2, 405-3, 405-4).
[0108] As a non-limiting example, line (405-1) may contain two words, line (405-2) may contain three words, line (405-3) may contain two words, and line (405-4) may contain two words. However, this is not limiting.
[0109] In one embodiment, at least one processor (210) can identify a number of lines among the lines (405) that contain less than a reference number of words (e.g., three). Three lines among the lines (405) (e.g., line (405-1), line (405-3), and line (405-4)) can contain less than a reference number of words (e.g., three). At least one processor (210) can identify a ratio of the number of lines within the block (400) to the number of lines containing less than three words as a predetermined ratio (e.g., a ratio of 4 to 3), but is not limited thereto.
[0110] According to one embodiment, at least one processor (210) can identify that the arrangement of characters within a block (400) does not satisfy a reference condition based on a certain ratio. At least one processor (210) can identify that the characters within a block (400) do not form a paragraph because the arrangement of characters within the block (400) does not satisfy the reference condition.
[0111] In one embodiment, the block (108) may include characters within lines (109). The lines (109) may include one or more words composed of characters. For example, at least one processor (210) may identify the number of words in each of the lines (109-1, 109-2, 109-3, 109-4, 109-5).
[0112] As a non-limiting example, line (109-1) may contain five words, line (109-2) may contain four words, line (109-3) may contain two words, line (109-4) may contain three words, and line (109-5) may contain three words. However, this is not limited thereto.
[0113] In one embodiment, at least one processor (210) can identify a number of lines among the lines (109) that contain less than a reference number of words (e.g., three). One line among the lines (109) (e.g., line (109-3)) can contain less than the reference number of words. At least one processor (210) can identify a ratio of the number of lines within the block (108) to the number of lines containing less than the reference number of words (e.g., three) as a predetermined ratio (e.g., a ratio of 5 to 1).
[0114] According to one embodiment, at least one processor (210) can identify, based on the ratio, that the arrangement of characters within a block (108) satisfies a reference condition. At least one processor (210) can identify that the characters within a block (108) form a paragraph because the arrangement of characters within a block (108) satisfies the reference condition.
[0115] Referring to FIG. 4b, at least one processor (210) can identify whether the arrangement of characters within the block satisfies a reference condition based on a first ratio of the length of the side (420, 425) parallel to the lines of the block (410, 108) to the length of each of the lines.
[0116] In one embodiment, the block (410) may include lines (415). The length of the lines (415) may be proportional to the number of characters located in the lines (415). At least one processor (210) may identify the length of each of the lines (415-1, 415-2, 415-3, 415-4). At least one processor (210) may identify the length (420) of a side that is parallel or substantially parallel to the lines of the block (410).
[0117] As a non-limiting example, the first ratio of line (415-2) and the first ratio of line (415-4) may be relatively greater than the first ratio of line (415-1) and the first ratio of line (415-3).
[0118] In one embodiment, at least one processor (210) can identify a number of lines among the lines (415) having a first ratio above a reference ratio. As a non-limiting example, lines (415-2) and lines (415-4) can have the first ratio above the reference ratio. As a non-limiting example, at least one processor (210) can identify a second ratio of the number of lines within the block (410) to the number of lines having the first ratio above the reference ratio as a predetermined ratio (e.g., a ratio of 4 to 2), but is not limited thereto.
[0119] In one embodiment, at least one processor (210) can identify that the arrangement of characters within a block (410) does not satisfy a reference condition based on a second ratio of the block (410) that is less than another reference ratio (e.g., a ratio of 4 to 2). As a non-limiting example, the other reference ratio can be defined as a certain ratio (e.g., a ratio of 4 to 3). For example, the at least one processor (210) can identify that the characters within the block (410) do not form a paragraph because the arrangement of characters within the block (410) does not satisfy the reference condition.
[0120] In one embodiment, the block (108) may include lines (109). The length of the lines (109) may be proportional to the number of characters located in the lines (109). For example, at least one processor (210) may identify the length of each of the lines (109-1, 109-2, 109-3, 109-4, 109-5). At least one processor (210) may identify the length (425) of a side that is parallel or substantially parallel to the lines of the block (108).
[0121] As non-limiting examples, the first ratio of line (109-1), the first ratio of line (109-2), the first ratio of line (109-4), and the first ratio of line (109-5) may have ratios that are relatively larger than the first ratio of line (109-3).
[0122] In one embodiment, at least one processor (210) can identify a number of lines among the lines (109) having a first ratio above a reference ratio. As a non-limiting example, lines (109-1), (109-2), (109-4), and (109-5) can have the first ratio above the reference ratio. As a non-limiting example, at least one processor (210) can identify a second ratio of the number of lines within the block (108) to the number of lines having the first ratio below the reference ratio as a predetermined ratio (e.g., a ratio of 5 to 4), but is not limited thereto.
[0123] In one embodiment, at least one processor (210) can identify that the arrangement of characters within a block (108) satisfies a reference condition based on a second ratio of the block (108) exceeding another reference ratio (e.g., a ratio of 5 to 4). As a non-limiting example, the other reference ratio can be defined as a certain ratio (e.g., a ratio of 4 to 3). At least one processor (210) can identify that the characters within the block (108) form a paragraph because the arrangement of characters within the block (108) satisfies the reference condition.
[0124] At least one processor (210) can identify that the first characters (104) of FIG. 1A satisfy the reference condition of FIG. 4A and / or the reference condition of FIG. 4B. For example, the first characters (104) can constitute a paragraph by satisfying the reference condition of FIG. 4A and / or the reference condition of FIG. 4B.
[0125] Figure 5 illustrates an exemplary method for adding a newline character.
[0126] Referring to FIG. 5, at least one processor (210) can identify a first line (500) within a block (108). At least one processor (210) can identify a length (505) of the first line (500). At least one processor (210) can identify a length (510) of a side of the block (108) that is parallel or substantially parallel to the first line (500). At least one processor (210) can identify a difference (515) between the length (510) of the side of the block (108) and the length (505) of the first line (500).
[0127] At least one processor (210) can identify a second line (520) adjacent to the first line (500) and positioned below the first line (500). At least one processor (210) can identify a length (530) of a first word (525) of the second line (520) among words positioned in the second line (520). Even though the length (530) of the first word (525) of the second line (520) is shorter than the difference (515), when the second line (520) is positioned on a line following the first line (500), there may be a line break between the first line (500) and the second line (520). Because, although there is sufficient space within the difference (515) space of the first line (500) to place the first word (525) of the second line (520), the first word (525) of the second line (520) is not placed within the difference (515) space, at least one processor (210) can identify the second line (520) as the next line of the first line (500) by adding a newline character. At least one processor (210) can identify the length (530) of the first word (525) of the second line (520) as shorter than the difference (515) by comparing the length (530) of the first word (525) of the second line (520) with the length of the difference (515). At least one processor (210) can identify a line break between the first line (500) and the second line (520) by identifying a length (530) of a first word (525) of the second line (520) that is shorter than the difference (515).
[0128] According to one embodiment, at least one processor (210) may obtain second characters (535) by performing paragraph-based translation of first characters. The second characters (535) may be located on lines. At least one processor (210) may identify a third line (540) corresponding to the first line (500) among the lines on which the second characters (535) are located.
[0129] At least one processor (210) may add a newline character to a position (545) within the second characters (535) corresponding to the end of the first line (500) of the block (108) based on identification of a length (530) of a first word (525) of a second line (520) that is shorter than a difference (515). The position (545) within the second characters (535) corresponding to the end of the first line (500) of the block (108) may correspond to a position indicating the end of a third line (540) of the second characters (535). By adding the newline character to the position (545), the at least one processor (210) may perform paragraph-based rendering for the second characters (535).
[0130] At least one processor (210) can display the second characters (535) including a line break between a third line (540) and a fourth line (550) by performing paragraph-based rendering for the second characters (535) by adding a line break character at a location (545). The fourth line (550) can be located below and adjacent to the third line (540). At least one processor (210) can display the second characters (535) on the locations of the first characters corresponding to the second characters (535) by displaying the second characters (535) including a line break between the third line (540) and the fourth line (550).
[0131] By displaying the second characters (535) on the positions of the first characters corresponding to the second characters (535), the second characters (535) can be shown to the user on the positions of the first characters corresponding to the second characters (535). At least one processor (210) can improve the readability and visibility of the second characters by displaying the second characters (535) on the positions of the first characters corresponding to the second characters (535).
[0132] At least one processor (210) may remove the first characters from the image to display the second characters (535) on the first characters. The area in the image from which the first characters are removed may be exemplified in the description of FIG. 6.
[0133] Figure 6 shows an example of an area in an image from which characters have been removed.
[0134] Referring to FIG. 6, at least one processor (210) may remove first characters (610) from the image (600) to display second characters translated from the first characters (610) on the first characters (610). By removing the first characters (610) from the image (600), the image (600) may include areas where the first characters (610) are removed.
[0135] In order to display areas related to the background of the image (600) within areas from which the first characters (610) have been removed, a method of acquiring areas related to the background of the image (600) within the electronic device (200) may be required. The electronic device (200) may include an artificial intelligence (AI) model. The AI model may include a machine learning model, a deep learning model, and / or a generative AI model. The AI model may be trained to generate areas corresponding to areas from which characters have been removed.
[0136] At least one processor (210) can input (or provide) an image (600) and first characters (610) to an artificial intelligence model. At least one processor (210) can obtain first regions (615) corresponding to regions from which the first characters (610) have been removed from the artificial intelligence model into which the image (600) and first characters (610) have been input (or provided). The first regions (615) can be continuous with the background in the image (600) or can be inferred from the background in the image (600).
[0137] At least one processor (210) can display second characters on first characters (610) within an image (600) in which first areas (615) are displayed.
[0138] According to one embodiment, at least one processor (210) may identify locations of first areas (615) within an image where first characters are positioned by performing OCR. At least one processor (210) may delete the first characters within the image and display second characters on locations of the first areas (615) corresponding to the locations of the first characters within the image.
[0139] At least one processor (210) can display second areas that do not overlap with the second characters among the first areas (615), based on the first areas (615). By displaying the second areas together with the second characters, the at least one processor (210) can improve the readability and visibility of the second characters.
[0140] At least one processor (210) can display second characters overlapping first characters (610) within the image (600) by performing paragraph-based rendering on the second characters. At least one processor (210) can delete the first characters (610) and display the second characters at the positions where the first characters (610) were deleted by performing paragraph-based rendering on the second characters. The second characters displayed according to paragraph-based rendering are exemplified in the description of FIG. 7.
[0141] Figure 7 illustrates examples of characters displayed according to paragraph-based rendering.
[0142] Referring to FIG. 7, a state (700) can be described as a state in which second characters (705) are displayed through a display (230). By performing paragraph-based rendering on the second characters (705), the electronic device (200) can transition from the state (107) of FIG. 1A to the state (700). Within the state (700), at least one processor (210) can display the second characters (705) within an image (103) through the display (230).
[0143] At least one processor (210) can identify the color of the first characters (104). At least one processor (210) can display the second characters (705) within the image (103) according to the color of the first characters (104).
[0144] At least one processor (210) can perform paragraph-based rendering on the second characters (705) by displaying a paragraph including the second characters (705) positioned on the first characters (104). At least one processor (210) can perform paragraph-based rendering on the second characters (705) by displaying the second characters (705) that replace the first characters (104) and are represented by a paragraph. At least one processor (210) can display the second characters (705) overlapping the first characters (104) within the image (103) by performing paragraph-based rendering on the second characters (705).
[0145] Since the second characters (705) in the state (700) are obtained based on the paragraph-based translation of the first characters (104), unlike the second characters (112) in the state (110) of FIG. 1A, the context of the second characters (705) can correspond to the context of the first characters (104).
[0146] Since the second characters (705) within the state (700) are displayed within the image (103) based on paragraph-based rendering, unlike the second characters (112) within the state (110) of FIG. 1A, the second characters (705) can be displayed in the same size as each other. By having at least one processor (210) display the second characters (705) constituting the paragraph in the same size as each other, the readability and visibility of the second characters (705) can be increased.
[0147] Figure 8 shows examples of characters displayed according to line-based rendering.
[0148] Referring to FIG. 8, at least one processor (210) can perform line-based rendering on third characters (810). By performing line-based rendering by at least one processor (210), the electronic device (200) can transition from the state (116) of FIG. 1B to the state (800) in which the third characters (810) are displayed within the image (114) via the display (230).
[0149] In one embodiment, at least one processor (210) can identify the color of the first characters (115). At least one processor (210) can display the third characters (810) within the image (114) according to the color of the first characters (115).
[0150] According to one embodiment, at least one processor (210) can perform line-based rendering on the third characters (810) by displaying the third characters (810) assigned to lines (805) of the image (114). The at least one processor (210) can perform line-based rendering on the third characters (810) by replacing the first characters (115) and displaying the third characters (810) separated by line. The at least one processor (210) can display the third characters (810) superimposed on the first characters (115) within the image (114) by performing line-based rendering on the third characters (810).
[0151] According to one embodiment, the third characters (810) in the state (800) are obtained based on a line-based translation of the first characters (115), unlike the third characters (120) in the state (119) of FIG. 1B, so that the context of each of the lines (805-1, 805-2, 805-3, 805-4, 805-5) can correspond to the context of each of the lines (118-1, 118-2, 118-3, 118-4, 118-5).
[0152] According to one embodiment, the third characters (810) in the state (800) are displayed in the image (114) based on line-based rendering, unlike the third characters (120) in the state (119) of FIG. 1B, such that each of the lines (805-1, 805-2, 805-3, 805-4, 805-5) to which the third characters (810) are assigned can be displayed on the corresponding positions of each of the lines (118-1, 118-2, 118-3, 118-4, 118-5) on which the first characters (115) are positioned. Each of the lines (805-1, 805-2, 805-3, 805-4, 805-5) to which the third characters (810) are assigned is displayed on the corresponding positions of each of the lines (118-1, 118-2, 118-3, 118-4, 118-5) on which the first characters (115) are positioned, thereby increasing the readability and visibility of the third characters (810).
[0153] Figure 9 illustrates examples of operations of at least one processor.
[0154] Referring to FIG. 9, at least one processor (210) may include an image analysis processor (905), an optical character recognition (OCR) processor (915), a language analysis processor (925), a text structure processor (935), a line break processor (945), a translation processor (955), a paragraph rendering processor (965), and / or a line rendering processor (975).
[0155] According to one embodiment, the processor means a central processing unit (CPU) processor, and since it is a hardware processor, at least one processor (210) in one processor may exist as a software module, or one processor may perform the role of at least one other processor.
[0156] In one embodiment, the electronic device (200) may obtain an image (900). For example, at least one processor (210) may provide the image (900) to an image analysis processor (905). For example, the image analysis processor (905) may identify whether text is included in the image (900).
[0157] In operation 910, the image analysis processor (905) may provide the image (900) to the OCR processor (915) based on identifying the image (900) as containing text. The OCR processor (915) may perform OCR on the image (900).
[0158] In operation 920, the OCR processor (915) may provide the first characters extracted from the image (900) according to OCR to the language analysis processor (925). For example, the language analysis processor (925) may identify the language of the first characters.
[0159] In operation 930, the OCR processor (915) may provide a block including first characters positioned on a plurality of lines within an image acquired through OCR to a text structure analysis processor (935). The text structure analysis processor (935) may identify whether the first characters within the block constitute a paragraph by identifying whether the first characters within the block satisfy a reference condition. The text structure analysis processor (935) may identify whether the first characters within the block constitute a paragraph by analyzing information about the positions of the blocks including the first characters and the lines including the first characters provided from the OCR processor (915). The text structure analysis processor (935) may determine whether the first characters constitute a paragraph based on whether the arrangement of the first characters satisfies the reference condition within the description of FIGS. 4A and 4B.
[0160] In operation 940, the text structure analysis processor (935) may provide the first characters to the line break processor (945) based on the first characters constituting the paragraph. The line break processor (945) may identify a line break within the first characters constituting the paragraph. Based on the identification, the line break processor (945) may add a newline character within the first characters. The text structure analysis processor (935) may perform OCR so that the context of the second characters obtained by performing paragraph-based translation of the first characters may vary depending on whether the first characters within the block obtained by performing OCR have a line break. The text structure analysis processor (935) may process the line break information so as to maintain the context of the first characters by adding a line break character at an appropriate location within the first characters.
[0161] In operation 950, the line break processor (945) may provide first characters to the translation processor (955). The translation processor (955) may perform a paragraph-based translation of the first characters with a line break character added. For example, the translation processor (955) may obtain second characters by performing a paragraph-based translation of the first characters constituting a paragraph. The translation processor (955) may obtain third characters by performing a line-based translation of the first characters constituting lines.
[0162] At operation 960, the translation processor (955) may provide the second characters to the paragraph rendering processor (965). The paragraph rendering processor (965) may perform paragraph-based rendering on the second characters, thereby displaying the second characters overlaid on the first characters within the image.
[0163] At operation 970, the translation processor (955) may provide third characters to the line rendering processor (975). The line rendering processor (975) may perform line-based rendering on the third characters, thereby displaying the third characters overlaid on the first characters within the image.
[0164] In one embodiment, the line rendering processor (975) may perform line-based rendering on the third characters, thereby deleting the first characters within the image and displaying the third characters at positions corresponding to where the first characters are positioned within the image. The operations of the paragraph rendering processor (965) for performing paragraph-based rendering are exemplified in the description of FIG. 10A.
[0165] Figure 10a illustrates an example of the operations of a paragraph-based rendering unit.
[0166] Referring to FIG. 10A, the electronic device (200) may include a paragraph-based rendering unit (1000). The paragraph-based rendering unit (1000) may include a translation mask generation unit (1020-1), a background and text calculation unit (1025-1), a text eraser unit (1030-1), a background mask generation unit (1035-1), an off-screen rendering unit (1040-1), and a bitmap generation unit (1045-1).
[0167] The paragraph-based rendering unit (1000), the translation mask generation unit (1020-1), the background and text calculation unit (1025-1), the text erasure unit (1030-1), the background mask generation unit (1035-1), the off-screen rendering unit (1040-1), and the bitmap generation unit (1045-1) can support the function of processing images and / or characters through an algorithm stored in the memory (220). Although the paragraph-based rendering unit (1000), the translation mask generation unit (1020-1), the background and text calculation unit (1025-1), the text erasure unit (1030-1), the background mask generation unit (1035-1), the off-screen rendering unit (1040-1), and the bitmap generation unit (1045-1) are described as 'units', they can perform the above functions software-wise and / or functionally.
[0168] At least one processor (210) can input (or provide) an image (1005) and second characters (1015) to a translation mask generation unit (1020-1). The second characters (1015) can be obtained by performing a paragraph-based translation of the first characters (1010) included in the image (1005). The translation mask generation unit (1020-1) can obtain a first mask including a black background and white second characters (1015) based on the input (or provision). The translation mask generation unit (1020-1) can generate a first mask having a black background based on inputting polygon values of paragraphs of the first characters (1010) in the image (1005) and second characters (1015), which are paragraph-based translation results for the first characters (1010).
[0169] At least one processor (210) can input (or provide) an image (1005), first characters (1010), and second characters (1015) to a background and text calculation unit (1025-1). The background and text calculation unit (1025-1) can obtain a color of a background within a first mask based on the image (1005). The background and text calculation unit (1025-1) can obtain a color of the second characters (1015) within the first mask based on identifying the color of the first characters (1010). The color of the second characters (1015) can substantially correspond to the color of the first characters (1010). Since the second characters (1015) substantially correspond to the color of the first characters (1010), the second characters (1015) can be provided with relatively less incongruity to the user. The background and text calculation unit (1025-1) can obtain a second mask based on the color of the background and the color of the second characters (1015). The background and text calculation unit (1025-1) can generate the second mask by detecting the area of the first characters (1010) according to the boundary of the lines of the first characters (1010). The second mask can be used to perform paragraph-based rendering for the second characters (1015). The background and text calculation unit (1025-1) can determine the background color of the second mask by using the pixel value of the surrounding area of the block according to the edge of the block to identify the background color of the block including the first characters (1010). At least one processor (210) can provide the image (1005) and the second characters (1015) to the text eraser unit (1030-1). The text eraser unit (1030-1) can obtain a background corresponding to the area of the first characters (1010) removed from the image (1005) based on an artificial intelligence model.
[0170] At least one processor (210) can input (or provide) a second mask including a background color acquired by the background and text calculation unit (1025-1) to the background mask generation unit (1035-1). The background mask generation unit (1035-1) can superimpose the second mask on the first mask based on the background color acquired by the text calculation unit (1025-1). The background mask generation unit (1035-1) can naturally process (e.g., rounding) the edge of the second mask based on the background color acquired by the background and text calculation unit (1025-1). By the background mask generation unit (1035-1) naturally processing the edge of the second mask, the visibility and readability of the second characters (1015) can be increased.
[0171] At least one processor (210) can input (or provide) a first mask and a second mask to an off-screen rendering unit (1040-1). The off-screen rendering unit (1040-1) can perform paragraph-based rendering of the second characters (1015) by overlaying the second mask on the first mask position within the image (1005). The off-screen rendering unit (1040-1) can perform paragraph-based rendering of the second characters (1015) by adjusting the transparency of the second mask. The off-screen rendering unit (1040-1) can overlay the second characters (1015) on the second mask such that each of the second characters (1015) has a uniform size and the lines on which the second characters (1015) are positioned have uniform intervals. The off-screen rendering unit (1040-1) can perform line-based rendering on the second characters (1015) using the background color of the first mask and the color of the first characters (1010) by using the first mask obtained from the translation mask generation unit (1020-1), the color of the first characters (1010) obtained from the background and text calculation unit (1025-1), and the second mask. At least one processor (210) can input (or provide) an image (1005) on which paragraph-based rendering of the second characters (1015) is performed to the bitmap generation unit (1045-1). The bitmap generation unit (1045-1) can obtain a bitmap (1050) from the image (1005) on which paragraph-based rendering of the second characters (1015) is performed. The bitmap (1050) may be defined as the final image obtained by performing paragraph-based rendering of the second characters (1015) within the image (1005).
[0172] Figure 10b illustrates an example of the operations of a line-based rendering unit.
[0173] Referring to FIG. 10b, the electronic device (200) may include a line-based rendering unit (1055). The line-based rendering unit (1055) may include a translation mask generation unit (1020-2), a background and text calculation unit (1025-2), a text eraser unit (1030-2), a background mask generation unit (1035-2), an off-screen rendering unit (1040-2), and a bitmap generation unit (1045-2).
[0174] The line-based rendering unit (1055), the translation mask generation unit (1020-2), the background and text calculation unit (1025-2), the text erasure unit (1030-2), the background mask generation unit (1035-2), the off-screen rendering unit (1040-2), and the bitmap generation unit (1045-2) may support the function of processing images and / or characters through an algorithm stored in the memory (220). Although the line-based rendering unit (1055), the translation mask generation unit (1020-2), the background and text calculation unit (1025-2), the text erasure unit (1030-2), the background mask generation unit (1035-2), the off-screen rendering unit (1040-2), and the bitmap generation unit (1045-2) are described as 'units', they may perform the functions software-wise and / or functionally.
[0175] At least one processor (210) can input (or provide) an image (1060) and second characters (1070) to a translation mask generation unit (1020-2). The second characters (1070) can be obtained by performing a line-based translation of the first characters (1065) included in the image (1060). The translation mask generation unit (1020-2) can obtain a first mask including a black background and white second characters (1070) based on the input (or provision). The translation mask generation unit (1020-2) can generate a first mask having a black background based on inputting polygon values of paragraphs of the first characters (1065) in the image (1060) and second characters (1070), which are paragraph-based translation results for the first characters (1065).
[0176] At least one processor (210) can input (or provide) an image (1060), first characters (1065), and second characters (1070) to the background and text calculation unit (1025-2). The background and text calculation unit (1025-2) can obtain the color of the background within the first mask based on the image (1060). The background and text calculation unit (1025-2) can obtain the color of the second characters (1070) within the first mask based on identifying the color of the first characters (1065). The color of the second characters (1070) can substantially correspond to the color of the first characters (1065). Since the second characters (1070) substantially correspond to the color of the first characters (1065), the second characters (1070) can be provided with relatively less incongruity to the user. The background and text calculation unit (1025-2) can obtain a second mask based on the color of the background and the color of the second characters (1070). The background and text calculation unit (1025-2) can generate a second mask by detecting the area of the first characters (1065) according to the boundary of the lines of the first characters (1065). The second mask can be used to perform line-based rendering for the second characters (1070). The background and text calculation unit (1025-2) can determine the background color of the second mask by using the pixel values of the surrounding area of the block according to the edge of the block to identify the background color of the block including the first characters (1065).
[0177] At least one processor (210) can provide an image (1060) and second characters (1070) to a text eraser (1030-2). The text eraser (1030-2) can obtain a background corresponding to an area of the first characters (1065) removed from the image (1060) based on an artificial intelligence model.
[0178] At least one processor (210) can input (or provide) a second mask including a background color acquired by the background and text calculation unit (1025-2) to the background mask generation unit (1035-2). The background mask generation unit (1035-2) can superimpose the second mask on the first mask based on the background color acquired by the text calculation unit (1025-2). The background mask generation unit (1035-2) can naturally process (e.g., rounding) the edge of the second mask based on the background color acquired by the background and text calculation unit (1025-2). By the background mask generation unit (1035-2) naturally processing the edge of the second mask, the visibility and readability of the second characters (1070) can be increased.
[0179] At least one processor (210) can input (or provide) a first mask and a second mask to an off-screen rendering unit (1040-2). The off-screen rendering unit (1040-2) can perform line-based rendering of the second characters (1070) by overlaying the second mask on the first mask position in the image (1060). The off-screen rendering unit (1040-2) can perform line-based rendering of the second characters (1070) by adjusting the transparency of the second mask. The off-screen rendering unit (1040-2) can overlay the second characters (1070) on the second mask such that each of the second characters (1070) has a uniform size and the lines on which the second characters (1070) are positioned have uniform intervals. The off-screen rendering unit (1040-2) can perform line-based rendering on the second characters (1070) using the background color of the first mask and the color of the first characters (1065) by using the first mask obtained from the translation mask generation unit (1020-2), the color of the first characters (1065) obtained from the background and text calculation unit (1025-2), and the second mask.
[0180] At least one processor (210) can input (or provide) an image (1060) on which line-based rendering of the second characters (1070) is performed to the bitmap generation unit (1045-2). The bitmap generation unit (1045-2) can obtain a bitmap (1075) from the image (1060) on which line-based rendering of the second characters (1070) is performed. The bitmap (1075) can be defined as a final image obtained by performing line-based rendering of the second characters (1070) within the image (1060). The bitmap generation unit (1045-2) can generate the bitmap (1075) to transfer the bitmap (1075), which is the final image, to a software application for rendering.
[0181] Figure 11 illustrates examples of paragraph-based rendering and line-based rendering performed within a single block.
[0182] Referring to FIG. 11, a state (1100) may be described as a state in which an image (1105) is displayed. Within the state (1100), at least one processor (210) may display an image (1105) via a display (230). The image (1105) may include first characters (1114) and second characters (1120).
[0183] At least one processor (210) can display an executable object (105) via a display (230) based on identifying characters within an image (1105). At least one processor (210) can receive an input for the executable object (105). The input (106) can be described as an input for translation of characters within the image (1105). The input (106) can include a touch input having a contact point on the executable object (105).
[0184] At least one processor (210) can perform OCR on the image (1105) based on the input (106). The at least one processor (210) can obtain, according to the OCR, a block (1110) comprising first characters (1114) of a first language and second characters (1120) of the first language located on a plurality of lines within the image (1105).
[0185] At least one processor (210) can identify whether the arrangement of the first characters (1114) and the arrangement of the second characters (1120) included in the block (1110) satisfy a reference condition. At least one processor (210) can identify that the arrangement of the first characters (1114) does not satisfy the reference condition because the first characters (1114) form lines. At least one processor (210) can identify that the arrangement of the second characters (1120) satisfies the reference condition because the second characters (1120) form paragraphs.
[0186] At least one processor (210) can perform a line-based translation of the first characters (1114) constituting the lines based on the arrangement of the first characters (1114) that do not satisfy the reference condition. By performing the line-based translation of the first characters (1114), the at least one processor (210) can obtain third characters (1129) having a second language. At least one processor (210) performs line-based translation of the first characters (1114), so that the context of each of the lines (1115-1, 1115-2, 1115-3, 1115-4, 1115-5) on which the first characters (1114) are located can correspond to the context of each of the lines (1130-1, 1130-2, 1130-3, 1130-4, 1130-5) to which the third characters (1129) are assigned.
[0187] At least one processor (210) can perform paragraph-based translation of second characters (1120) constituting a paragraph based on the arrangement of second characters (1120) that satisfy a reference condition. At least one processor (210) can obtain fourth characters (1135) having a second language by performing paragraph-based translation of the second characters (1120). By performing paragraph-based translation of the second characters (1120) by at least one processor (210), the context of the fourth characters (1135) can correspond to the context of the second characters (1120).
[0188] At least one processor (210) can perform line-based rendering of the third characters (1129). At least one processor (210) can perform paragraph-based rendering of the fourth characters (1135). By performing line-based rendering of the third characters (1129) and paragraph-based rendering of the fourth characters (1135), the electronic device (200) can transition from state (1100) to state (1125). State (1125) can be described as a state in which the third characters (1129) and the fourth characters (1135) are displayed.
[0189] At least one processor (210) can perform line-based rendering of the first characters (1114) by displaying third characters (1129) assigned to lines (1130) of the image (1105). At least one processor (210) can perform line-based rendering of the first characters (1114) by displaying third characters (1129) that replace the first characters (1114) and are separated by line. Within the state (1125), at least one processor (210) can display third characters (1129) superimposed on the first characters (1114) within the image (1105) by performing line-based rendering of the third characters (1129).
[0190] At least one processor (210) can perform line-based rendering of the third characters (1129) so that each of the lines (1130-1, 1130-2, 1130-3, 1130-4, 1130-5) to which the third characters (1129) are assigned can be displayed on the corresponding positions of each of the lines (1115-1, 1115-2, 1115-3, 1115-4, 1115-5) on which the first characters (1114) are located. By displaying each of the lines (1130-1, 1130-2, 1130-3, 1130-4, 1130-5) to which the third characters (1129) are assigned on each of the lines (1115-1, 1115-2, 1115-3, 1115-4, 1115-5) on which the first characters (1114) are located, the readability and visibility of the third characters (1129) can be increased.
[0191] At least one processor (210) can perform paragraph-based rendering of the fourth characters (1135) by displaying a paragraph including the fourth characters (1135) positioned on the second characters (1120) within the image (1105). At least one processor (210) can perform paragraph-based rendering of the fourth characters (1135) by displaying the fourth characters (1135) that replace the second characters (1120) and are represented by a paragraph. At least one processor (210) can display the fourth characters (1135) overlapping the second characters (1120) within the image (1105) by performing paragraph-based rendering of the fourth characters (1135).
[0192] At least one processor (210) can perform paragraph-based rendering of the fourth characters (1135) to display the fourth characters (1135) in the same size as each other. By displaying the fourth characters (1135) constituting a paragraph in the same size as each other, the readability and visibility of the fourth characters (1135) can be increased.
[0193] At least one processor (210) can maintain the context of the first characters (1114) and the context of the second characters (1120) within the image (1105) by performing line-based translation and paragraph-based translation for each of the first characters (1114) and the second characters (1120) within the block (1110). By performing line-based rendering and paragraph-based rendering for each of the third characters (1129) and the fourth characters (1135), the readability and visibility of the third characters (1129) and the fourth characters (1135) can be increased.
[0194] FIG. 12 is a block diagram of an electronic device within a network environment according to various embodiments.
[0195] Referring to FIG. 12, in a network environment (1400), an electronic device (1201) may communicate with an electronic device (1202) via a first network (1298) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (1204) or a server (1208) via a second network (1299) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (1201) may communicate with the electronic device (1204) via the server (1208). According to one embodiment, the electronic device (1201) may include a processor (1220), a memory (1230), an input module (1250), an audio output module (1255), a display module (1260), an audio module (1270), a sensor module (1276), an interface (1277), a connection terminal (1278), a haptic module (1279), a camera module (1280), a power management module (1288), a battery (1289), a communication module (1290), a subscriber identification module (1296), or an antenna module (1297). In some embodiments, the electronic device (1201) may omit at least one of these components (e.g., the connection terminal (1278)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1276), camera module (1280), or antenna module (1297)) may be integrated into a single component (e.g., display module (1260)).
[0196] The processor (1220) may execute software (e.g., a program (1240)) to control at least one other component (e.g., a hardware or software component) of the electronic device (1201) connected to the processor (1220) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1220) may store commands or data received from other components (e.g., a sensor module (1276) or a communication module (1290)) in a volatile memory (1232), process the commands or data stored in the volatile memory (1232), and store result data in a non-volatile memory (1234). According to one embodiment, the processor (1220) may include a main processor (1221) (e.g., a central processing unit or an application processor) or an auxiliary processor (1223) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1221). For example, when the electronic device (1201) includes the main processor (1221) and the auxiliary processor (1223), the auxiliary processor (1223) may be configured to use less power than the main processor (1221) or to be specialized for a given function. The auxiliary processor (1223) may be implemented separately from the main processor (1221) or as a part thereof.
[0197] The auxiliary processor (1223) may control at least a portion of functions or states associated with at least one component (e.g., the display module (1260), the sensor module (1276), or the communication module (1290)) of the electronic device (1201) on behalf of the main processor (1221) while the main processor (1221) is in an inactive (e.g., sleep) state, or together with the main processor (1221) while the main processor (1221) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1223) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1280) or a communication module (1290)). In one embodiment, the auxiliary processor (1223) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1201) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1208)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0198] The memory (1230) can store various data used by at least one component (e.g., the processor (1220) or the sensor module (1276)) of the electronic device (1201). The data can include, for example, software (e.g., the program (1240)) and input data or output data for commands related thereto. The memory (1230) can include a volatile memory (1232) or a non-volatile memory (1234).
[0199] The program (1240) may be stored as software in memory (1230) and may include an operating system (1242), middleware (1244), or an application (1246).
[0200] The input module (1250) can receive commands or data to be used in a component of the electronic device (1201) (e.g., a processor (1220)) from an external source (e.g., a user) of the electronic device (1201). The input module (1250) can include a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0201] The audio output module (1255) can output audio signals to the outside of the electronic device (1201). The audio output module (1255) can include a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0202] The display module (1260) can visually provide information to an external party (e.g., a user) of the electronic device (1201). The display module (1260) may include a display, a holographic device, or a projector, and a control circuit for controlling the device. In one embodiment, the display module (1260) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0203] The audio module (1270) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (1270) can acquire sound through the input module (1250), output sound through the sound output module (1255), or an external electronic device (e.g., electronic device (1202)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1201).
[0204] The sensor module (1276) can detect the operating status (e.g., power or temperature) of the electronic device (1201) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1276) can include a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0205] The interface (1277) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1201) with an external electronic device (e.g., the electronic device (1202)). In one embodiment, the interface (1277) may include a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0206] The connection terminal (1278) may include a connector through which the electronic device (1201) may be physically connected to an external electronic device (e.g., the electronic device (1202)). According to one embodiment, the connection terminal (1278) may include an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0207] The haptic module (1279) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (1279) may include a motor, a piezoelectric element, or an electrical stimulation device.
[0208] The camera module (1280) can capture still images and videos. According to one embodiment, the camera module (1280) may include one or more lenses, image sensors, image signal processors, or flashes.
[0209] The power management module (1288) can manage the power supplied to the electronic device (1201). According to one embodiment, the power management module (1288) can be implemented as at least a part of a power management integrated circuit (PMIC).
[0210] A battery (1289) may power at least one component of the electronic device (1201). In one embodiment, the battery (1289) may include a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0211] The communication module (1290) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1201) and an external electronic device (e.g., electronic device (1202), electronic device (1204), or server (1208)), and the performance of communication through the established communication channel. The communication module (1290) may operate independently from the processor (1220) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1290) may include a wireless communication module (1292) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1294) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (1204) via a first network (1298) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1299) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1292) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1296) to verify or authenticate the electronic device (1201) within a communication network such as the first network (1298) or the second network (1299).
[0212] The wireless communication module (1292) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1292) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1292) can support various technologies for securing performance in high-frequency bands, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1292) can support various requirements specified in the electronic device (1201), an external electronic device (e.g., the electronic device (1204)), or a network system (e.g., the second network (1299)). According to one embodiment, the wireless communication module (1292) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0213] The antenna module (1297) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1297) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1297) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1298) or the second network (1299), may be selected from the plurality of antennas by the communication module (1290). A signal or power may be transmitted or received between the communication module (1290) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1297).
[0214] According to various embodiments, the antenna module (1297) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0215] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0216] According to one embodiment, commands or data may be transmitted or received between the electronic device (1201) and an external electronic device (1204) via a server (1208) connected to a second network (1299). Each of the external electronic devices (1202 or 1204) may be the same or a different type of device as the electronic device (1201). According to one embodiment, all or part of the operations executed in the electronic device (1201) may be executed in one or more of the external electronic devices (1202, 1204, or 1208). When the electronic device (1201) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1201) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1201). The electronic device (1201) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (1201) may provide an ultra-low latency service using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1204) may include an Internet of Things (IoT) device. The server (1208) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (1204) or the server (1208) may be included in the second network (1299).The electronic device (1201) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0217] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0218] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0219] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. In one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0220] Various embodiments of the present document may be implemented as software (e.g., a program (1240)) including one or more instructions stored in a storage medium (e.g., an internal memory (1236) or an external memory (1238)) readable by a machine (e.g., an electronic device (1201)). A processor (e.g., a processor (1220)) of the machine (e.g., an electronic device (1201)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0221] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0222] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0223] As described above, the electronic device (e.g., the electronic device (200) of FIG. 2) may include at least one processor (e.g., at least one processor (210) of FIG. 2) including a memory (e.g., the memory (220) of FIG. 2)) that stores instructions and includes one or more storage media, a display (e.g., the display (230) of FIG. 2), and a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input (e.g., the input (106) of FIG. 1A) for translating characters in an image (e.g., the image (103) of FIG. 1A or the image (114) of FIG. 1B) displayed via the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform optical character recognition (OCR) on the image based on the input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a block (e.g., block (108) of FIG. 1A or block (117) of FIG. 1B) comprising first characters (e.g., first characters (104) of FIG. 1A or first characters (115) of FIG. 1B) of a first language located in a plurality of lines (e.g., lines (109) of FIG. 1A or lines (118) of FIG. 1B) within the image according to the OCR. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify whether an arrangement of the first characters included within the block satisfies a reference condition.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain second characters of a second language (e.g., second characters (705) of FIG. 7) translated from the first characters based on the arrangement of the first characters that satisfy the criterion condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the second characters, within the image, by performing paragraph-based rendering on the second characters. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain third characters of the second language (e.g., third characters (810) of FIG. 8) translated from the first characters based on the arrangement of the first characters that do not satisfy the criterion condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the third characters within the image by performing line-based rendering of the third characters.
[0224] The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the second characters within the image displayed via the display by performing paragraph-based rendering by displaying a paragraph including the second characters positioned on the first characters within the image displayed via the display based on the arrangement of the first characters that satisfy the criterion condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the third characters within the image by performing line-based rendering by displaying the third characters assigned to the lines of the image displayed via the display based on the arrangement of the first characters that do not satisfy the criterion condition.
[0225] The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, within the image, the second characters by replacing the first characters and displaying the second characters represented by paragraphs based on the arrangement of the first characters that satisfy the reference condition, thereby performing paragraph-based rendering. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, within the image, the third characters by replacing the first characters and displaying the third characters separated by lines based on the arrangement of the first characters that do not satisfy the reference condition, thereby performing line-based rendering.
[0226] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to in-paint a portion of an area of the image where the first character is located using the second character. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to in-paint another portion of the area where the first character is located using another area of the image.
[0227] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to store text content for each of the first characters and the second characters as metadata in the memory. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to provide the text content obtained from the metadata to a user.
[0228] The above criteria may relate to whether the first characters constitute a paragraph.
[0229] The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify whether the arrangement of the first characters satisfies the criteria condition based on the number of words located in each of the lines.
[0230] The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify whether the arrangement of the first characters satisfies the reference condition based on a ratio of the length of a side parallel to the lines of the block to the length of each of the lines.
[0231] The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, based on a distribution of blank areas within the block, whether the arrangement of the first characters satisfies the reference condition.
[0232] The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the second characters by performing paragraph-based translation of the first characters based on the arrangement of the first characters that satisfies the reference condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the third characters by performing line-based translation of the first characters based on the arrangement of the first characters that do not satisfy the reference condition.
[0233] The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, based on the arrangement of the first characters satisfying the reference condition, a length of a first word of a second line located below and adjacent to the first line among the lines, the second line having a length less than a difference between a length of a side parallel to the lines of the block and a length of the first line among the lines. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform paragraph-based rendering on the second characters by adding a newline character at a position within the second characters corresponding to an end of the first line of the block, based on the identification of the length of the first word of the second line.
[0234] The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain first areas from the image and the artificial intelligence model into which the first characters are input, by removing first characters within the block based on the arrangement of the first characters that satisfy the reference condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, within the image, second areas that do not overlap with the second characters among the first areas. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the first areas from the image and the artificial intelligence model into which the first characters are input, by removing first characters within the block based on the arrangement of the first characters that do not satisfy the reference condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, within the image, third areas of the first areas that do not overlap with the third characters.
[0235] The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify colors of the first characters based on the arrangement of the first characters that satisfy the criteria. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the second characters, within the image, according to the colors. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the colors of the first characters based on the arrangement of the first characters that do not satisfy the criteria. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the third characters, within the image, according to the colors.
[0236] As described above, the method may be performed in an electronic device including a display. The method may include receiving an input for translating characters within an image displayed through the display. The method may include performing optical character recognition (OCR) on the image based on the input. The method may include obtaining a block including first characters of a first language positioned on a plurality of lines within the image based on the OCR. The method may include identifying whether an arrangement of the first characters within the block satisfies a reference condition. The method may include obtaining second characters of a second language translated from the first characters based on the arrangement of the first characters satisfying the reference condition. The method may include displaying the second characters within the image by performing paragraph-based rendering on the second characters. The method may include an operation of obtaining third characters of the second language translated from the first characters based on the arrangement of the first characters that do not satisfy the reference condition. The method may include an operation of displaying the third characters within the image by performing line-based rendering on the third characters.
[0237] The method may include an action of displaying the second characters within the image by performing paragraph-based rendering by displaying a paragraph including the second characters positioned on the first characters within the image displayed through the display based on the arrangement of the first characters that satisfy the reference condition. The method may include an action of displaying the third characters within the image by performing line-based rendering by displaying the third characters assigned to the lines of the image displayed through the display based on the arrangement of the first characters that do not satisfy the reference condition.
[0238] The method may include an action of displaying the second characters, within the image, by performing paragraph-based rendering based on the arrangement of the first characters that satisfy the reference condition, by replacing the first characters and displaying the second characters represented by paragraphs. The method may include an action of displaying the third characters, within the image, by performing line-based rendering based on the arrangement of the first characters that do not satisfy the reference condition, by replacing the first characters and displaying the third characters separated by lines.
[0239] For example, the method may include an operation of in-painting a portion of an area of the image where the first character is located using the second character. The method may include an operation of in-painting another portion of the area where the first character is located using another area of the image.
[0240] For example, the method may include an operation of storing text content for each of the first characters and the second characters as metadata in the memory. The method may include an operation of providing the text content obtained from the metadata to a user. The reference condition may be related to whether the first characters constitute a paragraph.
[0241] The method may include an operation of identifying whether the arrangement of the first characters satisfies the reference condition based on the number of words located in each of the lines.
[0242] The method may include an operation of identifying whether the arrangement of the first characters satisfies the reference condition based on a ratio of the length of the side parallel to the lines of the block to the length of each of the lines.
[0243] The method may include an operation of identifying whether the arrangement of the first characters satisfies the reference condition based on the distribution of blank areas within the block.
[0244] The method may include an operation of obtaining the second characters by performing a paragraph-based translation of the first characters based on the arrangement of the first characters that satisfy the reference condition. The method may include an operation of obtaining the third characters by performing a line-based translation of the first characters based on the arrangement of the first characters that do not satisfy the reference condition.
[0245] The method may include, based on the arrangement of the first characters satisfying the reference condition, an operation of identifying a length of a first word of a second line located below the first line and adjacent to the first line among the lines having a length shorter than a difference between a length of a side parallel to the lines of the block and a length of the first line among the lines. The method may include, based on the identification of the length of the first word of the second line, an operation of performing paragraph-based rendering for the second characters by adding a newline character to a position within the second characters corresponding to an end of the first line of the block.
[0246] The method may include an operation of obtaining first areas from the image and the artificial intelligence model into which the first characters are input by removing the first characters within the block based on the arrangement of the first characters that satisfy the reference condition. The method may include an operation of displaying, within the image, second areas that do not overlap with the second characters among the first areas. The method may include an operation of obtaining the first areas from the image and the artificial intelligence model into which the first characters are input by removing the first characters within the block based on the arrangement of the first characters that do not satisfy the reference condition. The method may include an operation of displaying, within the image, third areas that do not overlap with the third characters among the first areas.
[0247] The method may include an operation of identifying the colors of the first characters based on the arrangement of the first characters that satisfy the reference condition. The method may include an operation of displaying the second characters within the image according to the colors. The method may include an operation of identifying the colors of the first characters based on the arrangement of the first characters that do not satisfy the reference condition. The method may include an operation of displaying the third characters within the image according to the colors.
[0248] As described above, the non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device including a display, cause the electronic device to receive an input for translating characters within an image displayed through the display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform optical character recognition (OCR) on the image based on the input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a block including first characters of a first language positioned on a plurality of lines within the image based on the OCR. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify whether an arrangement of the first characters included within the block satisfies a reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain second characters of a second language translated from the first characters based on the arrangement of the first characters that satisfy the reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display the second characters within the image by performing paragraph-based rendering on the second characters.The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain third characters of the second language translated from the first characters based on the arrangement of the first characters that do not satisfy the reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display the third characters within the image by performing line-based rendering on the third characters.
[0249] The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display, within the image, the second characters by performing paragraph-based rendering by displaying a paragraph including the second characters positioned on the first characters within the image displayed through the display based on the arrangement of the first characters that satisfy the reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display, within the image, the third characters by performing line-based rendering by displaying, within the image, the third characters by displaying, based on the arrangement of the first characters that do not satisfy the reference condition.
[0250] The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display, within the image, the second characters by replacing the first characters and displaying the second characters represented by paragraphs based on the arrangement of the first characters that satisfy the reference condition, thereby performing paragraph-based rendering. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display, within the image, the third characters by replacing the first characters and displaying the third characters separated by lines based on the arrangement of the first characters that do not satisfy the reference condition, thereby performing line-based rendering.
[0251] For example, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to in-paint a portion of the area where the first character is located in the image using the second character. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to in-paint another portion of the area where the first character is located using another area of the image.
[0252] For example, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to store text content for each of the first characters and the second characters as metadata in the memory. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to provide the text content obtained from the metadata to a user.
[0253] The above criteria may relate to whether the first characters constitute a paragraph. The one or more programs, when executed by the electronic device, may include instructions that cause the electronic device to determine whether the arrangement of the first characters satisfies the criteria based on the number of words located in each of the lines.
[0254] The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify whether the arrangement of the first characters satisfies the reference condition based on a ratio of the length of the side parallel to the lines of the block to the length of each of the lines.
[0255] The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify whether the arrangement of the first characters satisfies the reference condition based on a distribution of blank areas within the block.
[0256] The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain the second characters by performing paragraph-based translation of the first characters based on the arrangement of the first characters that satisfies the reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain the third characters by performing line-based translation of the first characters based on the arrangement of the first characters that do not satisfy the reference condition.
[0257] The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, based on the arrangement of the first characters that satisfy the reference condition, a length of a first word of a second line located below the first line and adjacent to the first line among the lines, the length of the first word being shorter than a difference between a length of a side parallel to the lines of the block and a length of the first line among the lines. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform paragraph-based rendering on the second characters by adding a newline character to a position within the second characters corresponding to an end of the first line of the block, based on the identification of the length of the first word of the second line.
[0258] The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain first areas from the image and the artificial intelligence model into which the first characters are input, by removing first characters within the block based on the arrangement of the first characters that satisfy the reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display, within the image, second areas that do not overlap with the second characters among the first areas. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain the first areas from the image and the artificial intelligence model into which the first characters are input, by removing first characters within the block based on the arrangement of the first characters that do not satisfy the reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display, within the image, third areas of the first areas that do not overlap with the third characters.
[0259] The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify colors of the first characters based on the arrangement of the first characters that satisfy the reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display the second characters within the image according to the colors. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify the colors of the first characters based on the arrangement of the first characters that do not satisfy the reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display the third characters within the image according to the colors.
[0260] The electronic device, as described above, may include at least one processor that stores instructions and includes a memory, a display, and a processing circuit, including one or more storage media. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input for translating characters within an image displayed through the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform optical character recognition (OCR) on the image based on the input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a block comprising first characters of a first language positioned in a plurality of lines within the image based on the OCR. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify whether an arrangement of the first characters included within the block satisfies a reference condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain second characters of a second language translated from the first characters based on the arrangement of the first characters satisfying the criterion condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to change the first characters to the second characters by displaying a paragraph including the second characters positioned over the first characters in the image displayed through the display.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain third characters of the second language translated from the first characters based on the arrangement of the first characters that do not satisfy the criterion condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to change the first characters to the second characters by displaying the third characters assigned to the lines of the image displayed through the display.
[0261] As described above, the method can be performed in an electronic device including a display. The method can include an operation of receiving an input for translating characters within an image displayed through the display. The method can include an operation of performing optical character recognition (OCR) on the image based on the input. The method can include an operation of obtaining a block including first characters of a first language positioned on a plurality of lines within the image based on the OCR. The method can include an operation of identifying whether an arrangement of the first characters included within the block satisfies a reference condition. The method can include an operation of obtaining second characters of a second language translated from the first characters based on the arrangement of the first characters satisfying the reference condition. The method can include an operation of changing the first characters to the second characters by displaying a paragraph including the second characters positioned on the first characters within the image displayed through the display. The method may include an operation of obtaining third characters of the second language translated from the first characters based on the arrangement of the first characters that do not satisfy the criterion condition. The method may include an operation of changing the first characters into the second characters by displaying the third characters assigned to the lines of the image displayed through the display.
[0262] As described above, the non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device including a display, cause the electronic device to receive an input for translating characters within an image displayed through the display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to perform optical character recognition (OCR) on the image based on the input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a block including first characters of a first language positioned on a plurality of lines within the image based on the OCR. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify whether an arrangement of the first characters included within the block satisfies a reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain second characters of a second language translated from the first characters based on the arrangement of the first characters that satisfy the reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to change the first characters to the second characters by displaying a paragraph including the second characters positioned on the first characters in the image displayed through the display.The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain third characters of the second language translated from the first characters based on the arrangement of the first characters that do not satisfy the reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to change the first characters to the second characters by displaying the third characters assigned to the lines of the image displayed through the display.
[0263] The electronic device, as described above, may include at least one processor, which stores instructions and includes a memory, a display, and a processing circuit, including one or more storage media. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input for translating characters within an image displayed through the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain first characters of a first language located in a plurality of lines within the image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify whether an attribute associated with the first characters satisfies a specified condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to translate the first characters into a second language so as to correspond to at least one semantic unit displayed across the plurality of lines, based on the attributes of the first characters that satisfy the designated condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to translate the first characters into the second language so as to correspond to a plurality of semantic units, each of which is separately displayed across each line of the plurality of lines, based on the attributes of the first characters that do not satisfy the designated condition.
[0264] For example, the attribute may be associated with the number of words composed of the first characters. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the attribute as satisfying the specified condition based on a ratio of lines among the plurality of lines in which the number of words contained in each line is greater than or equal to a reference number being greater than or equal to a reference ratio.
[0265] For example, the attribute may be associated with the length of the area in which the first characters are displayed. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the attribute as satisfying the specified condition based on a ratio of lines among the plurality of lines in which the length of the area in which characters are displayed is greater than or equal to a reference length being greater than or equal to a reference length.
[0266] For example, the attribute may be associated with a symbol or number included in the first characters. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the attribute as not satisfying the specified condition based on the placement of the symbol or number at the beginning of each line among the plurality of lines.
[0267] For example, the attribute may be associated with a type of content identified from at least one of the first characters or the image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the attribute as not satisfying the specified condition based on the type of the content being the specified type. For example, the attribute may be associated with a placement relationship of words composed of the first characters. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the attribute as satisfying the specified condition based on a probability that the last word of a first line among the plurality of lines and the first word of a second line, which is a line following the first line, are arranged consecutively is greater than or equal to a reference value, using an artificial intelligence model that has learned language.
[0268] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to store second characters of the second language translated from the first characters as character information within the memory of the electronic device.
[0269] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to provide the stored second characters as at least part of a response to another input for translation of the characters of the image or characters of another image related to the image.
[0270] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate another image based on the image, wherein the first characters are removed and the second characters are included. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to store the another image so that it is accessible via an image management application installed on the electronic device.
[0271] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine the second language based on contextual information related to operations performed on the electronic device prior to receiving the input for translation.
[0272] As described above, the method may be performed in an electronic device including a display. The method may include an operation of receiving an input for translating characters within an image displayed through the display. The method may include an operation of obtaining first characters of a first language located on a plurality of lines within the image. The method may include an operation of identifying whether an attribute associated with the first characters satisfies a specified condition. The method may include an operation of translating the first characters into a second language so as to correspond to at least one semantic unit displayed across the plurality of lines, based on the attribute of the first characters that satisfies the specified condition. The method may include an operation of translating the first characters into the second language so as to correspond to a plurality of semantic units, each of which is separately displayed on each line of the plurality of lines, based on the attribute of the first characters that do not satisfy the specified condition.
[0273] For example, the property may be associated with the number of words composed of the first characters. The method may include an operation of identifying the property as satisfying the specified condition based on the fact that, among the plurality of lines, the number of lines in which the number of words included in each line is greater than or equal to a reference number is greater than or equal to a reference ratio.
[0274] For example, the property may be associated with the length of the area where the first characters are displayed. The method may include an operation of identifying the property as satisfying the specified condition based on a ratio of lines among the plurality of lines in which the length of the area where the characters are displayed is greater than or equal to a reference length being greater than or equal to a reference ratio.
[0275] For example, the attribute may be associated with a symbol or number included in the first characters. The method may include an operation of identifying the attribute as not satisfying the specified condition based on the symbol or number being placed at the beginning of each line among the plurality of lines.
[0276] For example, the attribute may be associated with a type of content identified from at least one of the first characters or the image. The method may include an operation of identifying the attribute as not satisfying the specified condition based on the type of the content being a specified type. For example, the attribute may be associated with a placement relationship of words composed of the first characters. The method may include an operation of identifying the attribute as satisfying the specified condition based on a probability that the last word of the first line among the plurality of lines and the first word of the second line, which is the line following the first line, are arranged consecutively is greater than or equal to a reference value, using an artificial intelligence model that has learned language.
[0277] For example, the method may include an operation of storing second characters of the second language translated from the first characters as character information in the memory of the electronic device.
[0278] For example, the method may include providing the stored second characters as at least part of a response to another input based on receiving another input for translation of the characters of the image or characters of another image related to the image.
[0279] For example, the method may include an operation of generating another image based on the image, wherein the first characters are removed and the second characters are included. The method may include an operation of storing the other image so that it can be accessed through an image management application installed on the electronic device.
[0280] For example, the method may include determining the second language based on contextual information related to operations performed on the electronic device prior to receiving the input for the translation.
[0281] As described above, the non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device including a display, cause the electronic device to receive an input for translating characters within an image displayed through the display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain first characters of a first language located on a plurality of lines within the image. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify whether an attribute associated with the first characters satisfies a specified condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to translate the first characters into a second language corresponding to at least one semantic unit displayed across the plurality of lines, based on the attribute of the first character satisfying the specified condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to translate the first characters into the second language so that each of the first characters corresponds to a plurality of semantic units, each of which is displayed separately on each line of the plurality of lines, based on the attributes of the first characters that do not satisfy the specified condition.
[0282] For example, the property may be associated with the number of words composed of the first characters. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify the property as satisfying the specified condition based on a ratio of lines among the plurality of lines in which the number of words included in each line is greater than or equal to a reference number being greater than or equal to a reference ratio.
[0283] For example, the property may be associated with the length of the area where the first characters are displayed. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify the property as satisfying the specified condition based on a ratio of lines among the plurality of lines whose lengths of the area where characters are displayed are greater than or equal to a reference length being greater than or equal to a reference ratio.
[0284] For example, the attribute may be associated with a symbol or number included in the first characters. The one or more programs, when executed by the electronic device, may include instructions that cause the electronic device to identify the attribute as not satisfying the specified condition based on the placement of the symbol or number at the beginning of each line among the plurality of lines.
[0285] For example, the attribute may be associated with a type of content identified from at least one of the first characters or the image. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify the attribute as not satisfying the specified condition based on the type of the content being the specified type. For example, the attribute may be associated with a placement relationship of words composed of the first characters. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify the attribute as satisfying the specified condition based on a probability that the last word of a first line among the plurality of lines and the first word of a second line, which is a line following the first line, are arranged consecutively is greater than or equal to a reference value using an artificial intelligence model that has learned language.
[0286] For example, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to store second characters of the second language translated from the first characters as character information within the memory of the electronic device.
[0287] For example, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to provide the stored second characters as at least part of a response to the other input based on receiving another input for translation of the characters of the image or characters of another image related to the image.
[0288] For example, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate another image based on the image, wherein the first characters are removed and the second characters are included. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to store the other image so that it is accessible via an image management application installed on the electronic device.
[0289] For example, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to determine the second language based on contextual information related to operations performed by the electronic device prior to receiving the input for translation.
[0290] The electronic device, as described above, may include at least one processor, which stores instructions and includes a memory, a display, and a processing circuit, which includes one or more storage media. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input for translating characters within an image displayed through the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the characters from the image based on the input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a block comprising first characters of a first language positioned on a plurality of lines within the image based on the identification. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify whether an arrangement of the first characters included within the block satisfies a reference condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain second characters of a second language translated from the first characters based on the arrangement of the first characters satisfying the criterion condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the second characters within the image by performing paragraph-based rendering on the second characters.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain third characters of the second language translated from the first characters based on the arrangement of the first characters that do not satisfy the criterion condition. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the third characters within the image by performing line-based rendering on the third characters.
[0291] As described above, the method may be performed in an electronic device including a display. The method may include receiving an input for translating characters within an image displayed through the display. The method may include identifying the characters from the image based on the input. The method may include obtaining a block including first characters of a first language positioned on a plurality of lines within the image based on the identification. The method may include identifying whether an arrangement of the first characters included within the block satisfies a reference condition. The method may include obtaining second characters of a second language translated from the first characters based on an arrangement of the first characters satisfying the reference condition. The method may include displaying the second characters within the image by performing paragraph-based rendering on the second characters. The method may include an operation of obtaining third characters of the second language translated from the first characters based on the arrangement of the first characters that do not satisfy the reference condition. The method may include an operation of displaying the third characters within the image by performing line-based rendering on the third characters.
[0292] The one or more programs as described above may include instructions that, when executed by the electronic device including the display, cause the electronic device to receive an input for translating characters within an image displayed through the display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify the characters from the image based on the input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a block comprising first characters of a first language positioned on a plurality of lines within the image based on the identification. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify whether an arrangement of the first characters included within the block satisfies a reference condition. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain second characters of a second language translated from the first characters based on the arrangement of the first characters that satisfy the criteria. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display the second characters within the image by performing paragraph-based rendering on the second characters. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain third characters of the second language translated from the first characters based on the arrangement of the first characters that do not satisfy the criteria.The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display the third characters within the image by performing line-based rendering of the third characters.
Claims
1. In an electronic device (200), A memory (220) storing instructions and including one or more storage media; display (230); and At least one processor (210) comprising a processing circuit, The above instructions, when executed individually or collectively by the at least one processor (210), Receive input for translation of characters in the image (103, 114) displayed through the above display (230), Based on the above input, perform OCR (optical character recognition) on the image (103, 114), According to the above OCR, a block (108, 117) containing first characters (104, 115) of a first language located in a plurality of lines (109, 118) within the image (103, 114) is obtained, Identifying whether the arrangement of the first characters (104, 115) included in the above block (108, 117) satisfies the reference condition, and Based on the arrangement of the first characters (104) that satisfy the above criteria, second characters (705) of a second language translated from the first characters (104) are obtained, and paragraph-based rendering is performed on the second characters (705), thereby displaying the second characters (705) within the image (103), and Based on the arrangement of the first characters (115) that do not satisfy the above criteria, third characters (810) of the second language translated from the first characters (115) are obtained, and line-based rendering is performed on the third characters (810) to display the third characters (810) within the image (114). causing the above electronic device (200), Electronic device (200).
2. In claim 1, The above instructions, when executed individually or collectively by the at least one processor (210), By displaying a paragraph including the second characters (705) positioned on the first characters (104) in the image (103) displayed through the display (230) based on the arrangement of the first characters (104) that satisfy the above criteria, the second characters (705) are displayed in the image (103) by performing the paragraph-based rendering, and By displaying the third characters (810) assigned to the lines (805) of the image (114) displayed through the display (230) based on the arrangement of the first characters (115) that do not satisfy the above criteria, the line-based rendering is performed, so that the third characters (810) are displayed within the image (114). causing the above electronic device (200), Electronic device (200).
3. In claim 1, The above criteria are: Regarding whether the above first characters (104, 115) constitute a paragraph, Electronic device (200).
4. In claim 1, The above instructions, when executed individually or collectively by the at least one processor (210), Based on the number of words located in each of the above lines (109, 118), to identify whether the arrangement of the first characters (104, 115) satisfies the above criteria, causing the above electronic device (200), Electronic device (200).
5. In claim 1, The above instructions, when executed individually or collectively by the at least one processor (210), To identify whether the arrangement of the first characters (104, 115) satisfies the reference condition based on the ratio of the length of the side parallel to the lines (109, 118) of the blocks (108, 117) to the length of each of the lines (109, 118). causing the above electronic device (200), Electronic device (200).
6. In claim 1, The above instructions, when executed individually or collectively by the at least one processor (210), Based on the distribution of blank areas within the above blocks (108, 117), to identify whether the arrangement of the first characters (104, 115) satisfies the above criteria, causing the above electronic device (200), Electronic device (200).
7. In claim 1, The above instructions, when executed individually or collectively by the at least one processor (210), Based on the arrangement of the first characters (104) that satisfy the above criteria, the second characters (705) are obtained by performing a paragraph-based translation of the first characters (104), and Based on the arrangement of the first characters (115) that do not satisfy the above criteria, the third characters (810) are obtained by performing a line-based translation of the first characters (115). causing the above electronic device (200), Electronic device (200).
8. In claim 1, The above instructions, when executed individually or collectively by the at least one processor (210), Based on the arrangement of the first characters (104) that satisfy the above criteria, identifying the length of the first word of a second line located below the first line and adjacent to the first line among the lines (109) having a length shorter than the difference between the length of the side parallel to the lines (109) of the block (108) and the length of the first line among the lines (109), and Based on the identification of the length of the first word of the second line, paragraph-based rendering for the second characters (705) is performed by adding a newline character at a position within the second characters (705) corresponding to the end of the first line of the block (108). causing the above electronic device (200), Electronic device (200).
9. In claim 1, The above instructions, when executed individually or collectively by the at least one processor (210), Based on the arrangement of the first characters (104) that satisfy the above criteria, by removing the first characters (104) within the block (108), first areas from which the first characters (104) within the block (108) are removed are obtained from the image (103) and the artificial intelligence model into which the first characters (104) are input, and second areas that do not overlap with the second characters (705) among the first areas are displayed within the image (103), and Based on the arrangement of the first characters (115) that do not satisfy the above criteria, the first areas are obtained from the artificial intelligence model into which the image (114) and the first characters (115) are input by removing the first characters (115) within the block (117), and third areas that do not overlap with the third characters (810) among the first areas are displayed within the image (114). causing the above electronic device (200), Electronic device (200).
10. A method for executing within an electronic device (200) including a display (230), the method comprising: An operation of receiving input for translation of characters in an image (103, 114) displayed through the above display (230), An operation of performing OCR (optical character recognition) on the image (103, 114) based on the above input; According to the above OCR, an operation of obtaining a block (108, 117) including first characters (104, 115) of a first language located in a plurality of lines (109, 118) within the image (103, 114), An operation for identifying whether the arrangement of the first characters (104, 115) included in the above block (108, 117) satisfies a reference condition; An operation of obtaining second characters (705) of a second language translated from the first characters (104) based on the arrangement of the first characters (104) that satisfy the above criteria, and an operation of changing the first characters (104) to the second characters (705) by displaying a paragraph including the second characters (705) positioned on the first characters (104) in the image (103) displayed through the display (230), and An operation of obtaining third characters (810) of the second language translated from the first characters (115) based on the arrangement of the first characters (115) that do not satisfy the above criteria, and an operation of changing the first characters (115) to the third characters (810) by displaying the third characters (810) assigned to the lines (118) of the image (114) displayed through the display (230), method.
11. In the electronic device (200), A memory (220) storing instructions and including one or more storage media; display (230); and At least one processor (210) comprising a processing circuit, The above instructions, when executed individually or collectively by the at least one processor (210), Receive input for translation of characters in the image (103, 114) displayed through the above display (230), Obtaining first characters (104, 115) of a first language located in multiple lines (109, 118) within the above image (103, 114), Identify whether the attributes related to the above first characters (104, 115) satisfy the specified conditions, Based on the properties of the first characters (104) that satisfy the above-mentioned conditions, the first characters (104) are translated into a second language to correspond to at least one semantic unit displayed across the plurality of lines, and Based on the above properties of the first characters (115) that do not satisfy the above specified conditions, the first characters (115) are translated into the second language so that each of the first characters (115) corresponds to a plurality of semantic units that are displayed separately on each line of the plurality of lines. causing the above electronic device (200), Electronic device (200).
12. In claim 11, The above properties are, Associated with the number of words composed of the above first characters, The above instructions, when executed individually or collectively by the at least one processor (210), Among the above multiple lines, the ratio of lines in which the number of words included in each line is greater than or equal to a standard number is greater than or equal to a standard ratio, so that the property is identified as satisfying the specified condition. causing the above electronic device, Electronic devices.
13. In claim 11, The above properties are, The above first characters are related to the length of the displayed area, The above instructions, when executed individually or collectively by the at least one processor (210), Among the above multiple lines, the ratio of lines in which characters are displayed with a length greater than or equal to a standard length is greater than or equal to a standard ratio, so that the property is identified as satisfying the specified condition. causing the above electronic device, Electronic devices.
14. In claim 11, The above properties are, Associated with the symbols or numbers included in the above first characters, The above instructions, when executed individually or collectively by the at least one processor (210), Among the above multiple lines, based on the placement of the symbol or the number at the beginning of each line, to identify the property as not satisfying the specified condition, causing the above electronic device, Electronic devices.
15. In claim 11, The above properties are, Associated with the type of content identified from at least one of the first characters or the images, The above instructions, when executed individually or collectively by the at least one processor (210), Based on the above content having the specified type, identify the property as not satisfying the above specified condition; causing the above electronic device, Electronic devices.
Citation Information
Patent Citations
Electronic document creation system, and program
JP2015204075A
Translation display device
JP2918114B2
Rendering a text image following a line
KR1020140073480A
Device for protecting over-load in yaw drive for wind power generation
KR1020240065589A
A method and an apparatus for creating translated images while maintaining the style of text
KR102586180B1